ren worker folder adding orig, mv new one into root
This commit is contained in:
@@ -0,0 +1,46 @@
|
||||
---
|
||||
name: detector-run-behaviors
|
||||
description: |
|
||||
Self-check the diversity of your task's reference runs by pulling out a
|
||||
small set of discriminating behavior axes — named behaviors that
|
||||
distinguish runs from each other (framing choices, hallucinations,
|
||||
citation style, etc.) and emitting a structured behaviors × runs
|
||||
matrix. Useful as a sanity check before submission: if your reference
|
||||
runs all behave identically along every dimension you can name, the
|
||||
task probably isn't discriminating enough.
|
||||
allowed-tools: Bash, Read, Write
|
||||
---
|
||||
|
||||
# Run-behaviors extractor
|
||||
|
||||
This skill helps you see how your reference runs differ from each other.
|
||||
It pulls out 5–10 behavior axes — named behaviors that distinguish
|
||||
runs from each other (framing, investigation depth, hallucinations,
|
||||
citation style, hedging) — and writes a structured matrix you can use
|
||||
to confirm your task is producing genuinely diverse failure modes.
|
||||
|
||||
**This skill needs at least 2 reference runs.** Run your task with
|
||||
`scripts/harbor-run harbor-tasks/<slug> -k 4` (or similar) first so there
|
||||
are multiple `grade.md` and `answer.md` files to compare; with fewer
|
||||
than 2 runs there's nothing to discriminate against and the detector
|
||||
returns `not-applicable`.
|
||||
|
||||
Read these before starting:
|
||||
|
||||
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs. Detectors with structured payloads (this one's `runBehaviors` matrix) embed them in the same frontmatter block as `detector`/`verdict`/`confidence`.
|
||||
2. `.claude/skills/detector-run-behaviors/core.md` — what makes a good behavior axis, the structured `runBehaviors` payload shape, verdict enums, body sections.
|
||||
|
||||
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
|
||||
|
||||
## Acting on the verdict
|
||||
|
||||
- **`summary`** with `HIGH` confidence — your runs differ along clear,
|
||||
named axes. Good: that's the signal that says your task is
|
||||
discriminating enough to produce a useful score distribution.
|
||||
- **`summary`** with `MEDIUM` or `LOW` confidence — your runs look
|
||||
similar to each other and the axes you pulled out feel forced. The
|
||||
task may not be producing enough diversity to be a meaningful
|
||||
benchmark. Consider whether the prompt is too prescriptive, or
|
||||
whether more reference runs would surface real variation.
|
||||
- **`not-applicable`** — fewer than 2 reference runs. Run more trials
|
||||
first.
|
||||
@@ -0,0 +1,276 @@
|
||||
# Run-behaviors extractor — core
|
||||
|
||||
This file is the canonical, context-neutral content for the detector-run-behaviors
|
||||
detector. It defines what makes a good behavior axis, the structured
|
||||
`runBehaviors` payload schema, the verdict enums, and the body shape. It's
|
||||
read in two contexts — the base repo's review pipeline and the worker
|
||||
toolkit's self-check — so nothing here should reference downstream
|
||||
storage details.
|
||||
|
||||
## What this detector is for
|
||||
|
||||
Reviewers want to know quickly **how much diversity** a slug's reference runs
|
||||
have. Did every run miss the same point? Did one run hallucinate something
|
||||
none of the others did? Did the worst-scoring run fail in a fundamentally
|
||||
different way than the best? The existing per-run "notes" column captures
|
||||
some of this, but it's freeform — you have to read four cells of prose to
|
||||
notice that exactly one run hallucinated a UI and exactly two ran into the
|
||||
auth-edge case.
|
||||
|
||||
This detector pulls out a small set of **discriminating axes** — named
|
||||
behaviors that distinguish runs from each other — and emits a matrix that
|
||||
downstream tooling renders as a grid. Each row is a run; each column is a
|
||||
behavior; a filled cell means the run exhibits it. Outlier behaviors (only
|
||||
one run has them, or all-but-one do) get visual emphasis.
|
||||
|
||||
The point isn't to grade runs; it's to make run diversity (and the shape
|
||||
of that diversity) glanceable.
|
||||
|
||||
## Inputs
|
||||
|
||||
Read whatever you need from `harbor-tasks/<slug>/`. The load-bearing
|
||||
artifacts:
|
||||
|
||||
- `reference-runs/<run-id>/grade.md` — the grader's per-run reasoning. The
|
||||
primary input. Read every run. Look for what each run *did differently*
|
||||
from the others — which rubric items it hit, which it missed, what
|
||||
framing it brought to the prompt that the others didn't.
|
||||
- `reference-runs/<run-id>/agent-output/answer.md` — what the agent
|
||||
actually wrote. Useful when `grade.md` reasoning is terse and you need
|
||||
to confirm what the agent did, or to find behaviors the grader didn't
|
||||
call out (e.g., "this run cites file paths; the others don't").
|
||||
- The grader guidance — the rubric. Resolve which guidance file the grader
|
||||
actually reads (`bash scripts/guidance-target.sh <slug>` — the worker
|
||||
shell's guidance-target resolution) and read that file, never its sibling.
|
||||
Use this to *avoid* including
|
||||
behaviors that just restate "did the agent satisfy rubric item N." The
|
||||
rubric items are already a column-set; we want axes the rubric *doesn't*
|
||||
capture — tactical choices, framing, hallucinations, anything that
|
||||
distinguishes one run from another along a dimension the rubric doesn't
|
||||
score directly.
|
||||
|
||||
## What makes a good behavior axis
|
||||
|
||||
The single most useful behavior to surface is one that **one run exhibits
|
||||
and the others don't** (or one run lacks while the others share). That's
|
||||
the outlier signal the reviewer is hunting. Aim for:
|
||||
|
||||
- **5 to 10 behaviors total.** The grid is rows-by-runs, so the row
|
||||
count is where the visual scales. Fewer than 5 and the matrix has no
|
||||
shape; more than ~10 and the legend below gets cluttered.
|
||||
- **Behaviors that genuinely discriminate.** A row where every run is
|
||||
filled (all runs missed the same rubric item) or every run is empty
|
||||
*usually* carries no diversity signal. **Exception: when the always-
|
||||
exhibited (or always-missed) behavior is a core thing the
|
||||
grader guidance is looking for**, include it anyway. A small set of
|
||||
"every run failed here" rows can be load-bearing context — they show
|
||||
the reader at a glance that the task's primary failure mode is
|
||||
reliably reproducing, not just a one-off. Cap these at 1-3 core
|
||||
rows; the rest should be true outliers. Universal claims are also
|
||||
the ones most likely to be wrong: before shipping an "every run"
|
||||
(or "no run") row, re-check the claim against each run's `grade.md`
|
||||
— the top-scoring run is where it most often breaks.
|
||||
- **Short, scannable labels.** ≤ 6 words. Render-time the row label
|
||||
has more room than a column header, but legend cards repeat them, so
|
||||
brevity still pays.
|
||||
- **Behavior-shaped, not score-shaped.** Prefer "Hallucinated withdraw
|
||||
UI" over "Failed Issue 3." The rubric-issue × run grid already exists
|
||||
(see `RubricHeatmap` / `rubricVerdicts`); this matrix is *complementary*
|
||||
— it captures things the rubric doesn't grade, plus the small set of
|
||||
core rubric concerns where the diversity signal is "all runs missed
|
||||
this" (and that fact is itself the headline).
|
||||
- **Defined precisely enough to apply consistently.** The `description`
|
||||
field is your operational definition. A reader should be able to read
|
||||
it and re-apply the same label to a new run without ambiguity.
|
||||
|
||||
Bad behavior axes:
|
||||
|
||||
- "Wrote a thorough answer" — vague, not falsifiable, every run is
|
||||
somewhere on the spectrum.
|
||||
- "Missed Issue 4" — already captured by the rubric × runs grid; you're
|
||||
not adding signal.
|
||||
- "Got the right answer" — score-shaped, not behavior-shaped, and already
|
||||
captured by `reward`.
|
||||
- "Used the word 'security'" — too fine-grained to be a useful axis.
|
||||
|
||||
## Patterns to look for in `grade.md`
|
||||
|
||||
Behaviors that show up across many slugs and tend to be discriminating:
|
||||
|
||||
- **Framing choice** — did the agent treat this as a security audit, a
|
||||
refactor proposal, a compliance review, an incident postmortem? Runs
|
||||
that frame the same prompt differently will produce structurally
|
||||
different answers.
|
||||
- **Investigation depth** — did the agent read 2 files, 12 files, 50
|
||||
files? Does the grader specifically note which files were/weren't
|
||||
opened?
|
||||
- **Hallucinations** — did the agent describe a function/file/UI that
|
||||
doesn't exist? This is almost always a useful axis when at least one
|
||||
run does it.
|
||||
- **Citation style** — did the agent cite file:line, just file paths, or
|
||||
no paths at all? Often correlates with reward.
|
||||
- **Hedging vs. confident assertion** — same answer can be marked up or
|
||||
down depending on whether the agent hedged appropriately.
|
||||
- **Self-correction within the run** — did the agent backtrack mid-answer
|
||||
("actually, looking more carefully…") or commit to the first read?
|
||||
- **Topic-area coverage** — for multi-issue rubrics, did the agent split
|
||||
attention evenly or skip a whole topic area?
|
||||
|
||||
Behaviors to *avoid* listing (already captured elsewhere):
|
||||
|
||||
- "Got rubric item N right/wrong" — see `rubricVerdicts`.
|
||||
- "Scored above 0.5" — see `reward`.
|
||||
- "Took a long time" — not stable / not behavior-shaped.
|
||||
|
||||
## Verdict and confidence
|
||||
|
||||
- `verdict`: `summary` when you produced a matrix (this detector is
|
||||
descriptive, not pass/fail; `summary` signals "no judgment, just an
|
||||
extraction"). Use `not-applicable` instead when the matrix can't be
|
||||
built — see "What about `not-applicable`?" at the bottom. Those are
|
||||
the only two values.
|
||||
- `confidence`: `HIGH` | `MEDIUM` | `LOW` — how confident you are that
|
||||
these axes are the *most* discriminating ones (vs. better axes you
|
||||
might have missed). `HIGH` for runs whose differences are stark and
|
||||
easy to articulate; `MEDIUM` when runs are similar enough that the
|
||||
axes you chose feel forced; `LOW` when you only had partial data
|
||||
(e.g., missing `answer.md` files).
|
||||
|
||||
## Frontmatter and body schema
|
||||
|
||||
The detector report is YAML frontmatter (with the structured
|
||||
`runBehaviors` matrix inline) followed by a markdown body. Both contexts
|
||||
produce the same shape; only the *sink* differs.
|
||||
|
||||
**Frontmatter** — exactly these top-level keys:
|
||||
|
||||
```yaml
|
||||
---
|
||||
detector: detector-run-behaviors
|
||||
verdict: summary | not-applicable
|
||||
confidence: HIGH | MEDIUM | LOW
|
||||
runBehaviors:
|
||||
behaviors:
|
||||
- id: b01
|
||||
kind: failure
|
||||
label: "Tunnel-vision on legal framing"
|
||||
description: "Frames the whole answer as a compliance/legal question and never opens any client-side code."
|
||||
- id: b02
|
||||
kind: failure
|
||||
label: "Hallucinates withdraw UI"
|
||||
description: "Describes a withdraw-flow UI component (e.g. demos a button or modal) that does not exist anywhere in the codebase."
|
||||
- id: b03
|
||||
kind: target
|
||||
label: "Cites file paths"
|
||||
description: "Cites paths with file:line precision when making load-bearing claims about the codebase."
|
||||
perRun:
|
||||
"reward-0.44-LQrU9Cg": [b01]
|
||||
"reward-0.47-JWSWFw3": [b02]
|
||||
"reward-0.49-A7Mte9P": [b03]
|
||||
"reward-0.56-SCZ7wSC": [b03]
|
||||
---
|
||||
```
|
||||
|
||||
Field rules:
|
||||
|
||||
- `behaviors[].id`: stable string like `"b01"`. Just an identifier — must
|
||||
be unique within the matrix and must match the ids you reference in
|
||||
`perRun`. Validation rejects unknown ids.
|
||||
- `behaviors[].kind`: `"target"` or `"failure"`. **Polarity matters** —
|
||||
the grid renders green for a `target` cell that the run hit, red for
|
||||
a `failure` cell that the run exhibited. Pick the framing that makes
|
||||
the axis sharpest: "Cites file paths" (target, green when present) vs.
|
||||
"Doesn't cite file paths" (failure, red when present) — generally the
|
||||
rarer half should be the named axis so cells fill more sparsely. Use
|
||||
`failure` for things the agent shouldn't do, `target` for things the
|
||||
agent should do. A filled cell asserts the polarity *for that run*,
|
||||
not just factual presence: a behavior can be true of a run and still
|
||||
not be a fault for it — a run that avoided the underlying issue by
|
||||
construction had nothing to surface, and a `failure` cell there paints
|
||||
the strongest run red for doing the right thing. Likewise don't fill a
|
||||
`target` cell for work that's actually off-target scope (edits to a
|
||||
lookalike flow the prompt never asked about). If the polarity doesn't
|
||||
hold for every run you'd mark, reframe the axis or leave that run's
|
||||
cell empty.
|
||||
- `behaviors[].label`: ≤ 6 words, render-time column header. Sentence
|
||||
case ("Hallucinates withdraw UI"), not Title Case.
|
||||
- `behaviors[].description`: 1-2 sentences. The operational definition
|
||||
the reader can re-apply. Render-time tooltip.
|
||||
- `perRun`: keyed on the **run directory name** (e.g.
|
||||
`"reward-0.44-LQrU9Cg"`), value is an array of behavior ids. Empty
|
||||
array is fine — it means "this run exhibits none of the listed
|
||||
behaviors," which is itself a signal.
|
||||
|
||||
Before you build the matrix, list the actual `reference-runs/<run-id>/`
|
||||
directories and take your `perRun` keys from that listing verbatim. Every
|
||||
run in `reference-runs/` should appear in `perRun`, and every `perRun`
|
||||
key must match one of those directories exactly. A report whose keys
|
||||
cite run ids that don't exist on disk is describing an earlier
|
||||
generation of runs — it's invalid no matter how good the axes look, so
|
||||
re-derive the matrix from the current runs rather than ship it. Runs you
|
||||
don't list will render as empty rows.
|
||||
|
||||
Cell values need the same discipline as the keys. A filled cell is a
|
||||
claim about a specific run: before you emit it, ground it in a specific
|
||||
quote or line from *that run's* `grade.md` or `answer.md` that you
|
||||
actually read. Check the run's own framing — agents often explicitly
|
||||
disclaim a behavior (a "Not covered" section, "static linting is not a
|
||||
full audit") that a skim of the diff would credit them with, and a run
|
||||
that hedges its scope is different from one that declares the work
|
||||
"complete and verified." Check how the run ended, too: a run cut off
|
||||
mid-work (crash, API error partway through implementing) never got to
|
||||
decide what to omit, so don't read its omissions as final behavioral
|
||||
choices. A cell you can't ground in the run's own text stays **empty**
|
||||
— an unmarked cell is neutral; note the ambiguity in the per-behavior
|
||||
notes as unclear rather than guessing, because a guessed cell is a
|
||||
false claim about a run the reader can check.
|
||||
|
||||
**Body sections**, in this order:
|
||||
|
||||
```markdown
|
||||
# Run-behaviors extraction: <slug>
|
||||
|
||||
## How the runs differ
|
||||
|
||||
2-4 paragraphs. Articulate the *shape* of the diversity — "two runs
|
||||
attack the prompt from a compliance angle, one writes a workspace audit,
|
||||
one hallucinates a UI" — before showing the matrix. The body is what a
|
||||
reader gets if they want the qualitative narrative; the matrix is what
|
||||
they glance at. When the shape is convergence — no run demonstrates the
|
||||
strong path, or every run lands on the same failure — say so plainly as
|
||||
an observation. Convergence is often the intended shape of the task, so
|
||||
describe it; don't label it a defect.
|
||||
|
||||
## Per-behavior notes
|
||||
|
||||
For each behavior you pulled out, give 1-2 sentences explaining what
|
||||
counts as exhibiting it and which run is the canonical example. Quote
|
||||
from `grade.md` or `answer.md` when the line between "exhibits" and
|
||||
"doesn't" is subtle. If you left a run's cell empty because you couldn't
|
||||
ground it either way, say so here ("unclear for reward-0.53-…: neither
|
||||
the grade nor the answer addresses it") instead of silently omitting —
|
||||
the empty cell and the note together are the honest representation.
|
||||
|
||||
> "We're going to defer demoing the withdraw flow to a follow-up turn"
|
||||
> — reward-0.47-JWSWFw3, answer.md ¶3 (the prior phrase being the
|
||||
> load-bearing tell)
|
||||
```
|
||||
|
||||
Don't restate the rubric. If a behavior column lines up with a rubric
|
||||
issue, the reader will see that from the rubric-issue grid — your column
|
||||
is adding new signal, not redundant signal.
|
||||
|
||||
## What about `not-applicable`?
|
||||
|
||||
If `reference-runs/` is empty or has only one run, there's nothing to
|
||||
build a discrimination matrix from. Emit:
|
||||
|
||||
```yaml
|
||||
verdict: not-applicable
|
||||
confidence: HIGH
|
||||
```
|
||||
|
||||
… with a body that explains which trigger fired ("only one reference
|
||||
run") and stop. Don't try to find behaviors a single run "exhibits" — a
|
||||
1-row matrix is noise, and the outlier highlights need ≥ 2 rows to
|
||||
compute against.
|
||||
Reference in New Issue
Block a user