remove folder - untrustworthy

This commit is contained in:
2026-08-11 14:12:23 -04:00
parent f18bb0a146
commit 0012380fd3
140 changed files with 0 additions and 33136 deletions

View File

@@ -1,46 +0,0 @@
---
name: detector-run-behaviors
description: |
Self-check the diversity of your task's reference runs by pulling out a
small set of discriminating behavior axes — named behaviors that
distinguish runs from each other (framing choices, hallucinations,
citation style, etc.) and emitting a structured behaviors × runs
matrix. Useful as a sanity check before submission: if your reference
runs all behave identically along every dimension you can name, the
task probably isn't discriminating enough.
allowed-tools: Bash, Read, Write
---
# Run-behaviors extractor
This skill helps you see how your reference runs differ from each other.
It pulls out 5–10 behavior axes — named behaviors that distinguish
runs from each other (framing, investigation depth, hallucinations,
citation style, hedging) — and writes a structured matrix you can use
to confirm your task is producing genuinely diverse failure modes.
**This skill needs at least 2 reference runs.** Run your task with
`scripts/harbor-run harbor-tasks/<slug> -k 4` (or similar) first so there
are multiple `grade.md` and `answer.md` files to compare; with fewer
than 2 runs there's nothing to discriminate against and the detector
returns `not-applicable`.
Read these before starting:
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs. Detectors with structured payloads (this one's `runBehaviors` matrix) embed them in the same frontmatter block as `detector`/`verdict`/`confidence`.
2. `.claude/skills/detector-run-behaviors/core.md` — what makes a good behavior axis, the structured `runBehaviors` payload shape, verdict enums, body sections.
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
## Acting on the verdict
- **`summary`** with `HIGH` confidence — your runs differ along clear,
named axes. Good: that's the signal that says your task is
discriminating enough to produce a useful score distribution.
- **`summary`** with `MEDIUM` or `LOW` confidence — your runs look
similar to each other and the axes you pulled out feel forced. The
task may not be producing enough diversity to be a meaningful
benchmark. Consider whether the prompt is too prescriptive, or
whether more reference runs would surface real variation.
- **`not-applicable`** — fewer than 2 reference runs. Run more trials
first.

View File

@@ -1,276 +0,0 @@
# Run-behaviors extractor — core
This file is the canonical, context-neutral content for the detector-run-behaviors
detector. It defines what makes a good behavior axis, the structured
`runBehaviors` payload schema, the verdict enums, and the body shape. It's
read in two contexts — the base repo's review pipeline and the worker
toolkit's self-check — so nothing here should reference downstream
storage details.
## What this detector is for
Reviewers want to know quickly **how much diversity** a slug's reference runs
have. Did every run miss the same point? Did one run hallucinate something
none of the others did? Did the worst-scoring run fail in a fundamentally
different way than the best? The existing per-run "notes" column captures
some of this, but it's freeform — you have to read four cells of prose to
notice that exactly one run hallucinated a UI and exactly two ran into the
auth-edge case.
This detector pulls out a small set of **discriminating axes** — named
behaviors that distinguish runs from each other — and emits a matrix that
downstream tooling renders as a grid. Each row is a run; each column is a
behavior; a filled cell means the run exhibits it. Outlier behaviors (only
one run has them, or all-but-one do) get visual emphasis.
The point isn't to grade runs; it's to make run diversity (and the shape
of that diversity) glanceable.
## Inputs
Read whatever you need from `harbor-tasks/<slug>/`. The load-bearing
artifacts:
- `reference-runs/<run-id>/grade.md` — the grader's per-run reasoning. The
primary input. Read every run. Look for what each run *did differently*
from the others — which rubric items it hit, which it missed, what
framing it brought to the prompt that the others didn't.
- `reference-runs/<run-id>/agent-output/answer.md` — what the agent
actually wrote. Useful when `grade.md` reasoning is terse and you need
to confirm what the agent did, or to find behaviors the grader didn't
call out (e.g., "this run cites file paths; the others don't").
- The grader guidance — the rubric. Resolve which guidance file the grader
actually reads (`bash scripts/guidance-target.sh <slug>` — the worker
shell's guidance-target resolution) and read that file, never its sibling.
Use this to *avoid* including
behaviors that just restate "did the agent satisfy rubric item N." The
rubric items are already a column-set; we want axes the rubric *doesn't*
capture — tactical choices, framing, hallucinations, anything that
distinguishes one run from another along a dimension the rubric doesn't
score directly.
## What makes a good behavior axis
The single most useful behavior to surface is one that **one run exhibits
and the others don't** (or one run lacks while the others share). That's
the outlier signal the reviewer is hunting. Aim for:
- **5 to 10 behaviors total.** The grid is rows-by-runs, so the row
count is where the visual scales. Fewer than 5 and the matrix has no
shape; more than ~10 and the legend below gets cluttered.
- **Behaviors that genuinely discriminate.** A row where every run is
filled (all runs missed the same rubric item) or every run is empty
*usually* carries no diversity signal. **Exception: when the always-
exhibited (or always-missed) behavior is a core thing the
grader guidance is looking for**, include it anyway. A small set of
"every run failed here" rows can be load-bearing context — they show
the reader at a glance that the task's primary failure mode is
reliably reproducing, not just a one-off. Cap these at 1-3 core
rows; the rest should be true outliers. Universal claims are also
the ones most likely to be wrong: before shipping an "every run"
(or "no run") row, re-check the claim against each run's `grade.md`
— the top-scoring run is where it most often breaks.
- **Short, scannable labels.** ≤ 6 words. Render-time the row label
has more room than a column header, but legend cards repeat them, so
brevity still pays.
- **Behavior-shaped, not score-shaped.** Prefer "Hallucinated withdraw
UI" over "Failed Issue 3." The rubric-issue × run grid already exists
(see `RubricHeatmap` / `rubricVerdicts`); this matrix is *complementary*
— it captures things the rubric doesn't grade, plus the small set of
core rubric concerns where the diversity signal is "all runs missed
this" (and that fact is itself the headline).
- **Defined precisely enough to apply consistently.** The `description`
field is your operational definition. A reader should be able to read
it and re-apply the same label to a new run without ambiguity.
Bad behavior axes:
- "Wrote a thorough answer" — vague, not falsifiable, every run is
somewhere on the spectrum.
- "Missed Issue 4" — already captured by the rubric × runs grid; you're
not adding signal.
- "Got the right answer" — score-shaped, not behavior-shaped, and already
captured by `reward`.
- "Used the word 'security'" — too fine-grained to be a useful axis.
## Patterns to look for in `grade.md`
Behaviors that show up across many slugs and tend to be discriminating:
- **Framing choice** — did the agent treat this as a security audit, a
refactor proposal, a compliance review, an incident postmortem? Runs
that frame the same prompt differently will produce structurally
different answers.
- **Investigation depth** — did the agent read 2 files, 12 files, 50
files? Does the grader specifically note which files were/weren't
opened?
- **Hallucinations** — did the agent describe a function/file/UI that
doesn't exist? This is almost always a useful axis when at least one
run does it.
- **Citation style** — did the agent cite file:line, just file paths, or
no paths at all? Often correlates with reward.
- **Hedging vs. confident assertion** — same answer can be marked up or
down depending on whether the agent hedged appropriately.
- **Self-correction within the run** — did the agent backtrack mid-answer
("actually, looking more carefully…") or commit to the first read?
- **Topic-area coverage** — for multi-issue rubrics, did the agent split
attention evenly or skip a whole topic area?
Behaviors to *avoid* listing (already captured elsewhere):
- "Got rubric item N right/wrong" — see `rubricVerdicts`.
- "Scored above 0.5" — see `reward`.
- "Took a long time" — not stable / not behavior-shaped.
## Verdict and confidence
- `verdict`: `summary` when you produced a matrix (this detector is
descriptive, not pass/fail; `summary` signals "no judgment, just an
extraction"). Use `not-applicable` instead when the matrix can't be
built — see "What about `not-applicable`?" at the bottom. Those are
the only two values.
- `confidence`: `HIGH` | `MEDIUM` | `LOW` — how confident you are that
these axes are the *most* discriminating ones (vs. better axes you
might have missed). `HIGH` for runs whose differences are stark and
easy to articulate; `MEDIUM` when runs are similar enough that the
axes you chose feel forced; `LOW` when you only had partial data
(e.g., missing `answer.md` files).
## Frontmatter and body schema
The detector report is YAML frontmatter (with the structured
`runBehaviors` matrix inline) followed by a markdown body. Both contexts
produce the same shape; only the *sink* differs.
**Frontmatter** — exactly these top-level keys:
```yaml
---
detector: detector-run-behaviors
verdict: summary | not-applicable
confidence: HIGH | MEDIUM | LOW
runBehaviors:
behaviors:
- id: b01
kind: failure
label: "Tunnel-vision on legal framing"
description: "Frames the whole answer as a compliance/legal question and never opens any client-side code."
- id: b02
kind: failure
label: "Hallucinates withdraw UI"
description: "Describes a withdraw-flow UI component (e.g. demos a button or modal) that does not exist anywhere in the codebase."
- id: b03
kind: target
label: "Cites file paths"
description: "Cites paths with file:line precision when making load-bearing claims about the codebase."
perRun:
"reward-0.44-LQrU9Cg": [b01]
"reward-0.47-JWSWFw3": [b02]
"reward-0.49-A7Mte9P": [b03]
"reward-0.56-SCZ7wSC": [b03]
---
```
Field rules:
- `behaviors[].id`: stable string like `"b01"`. Just an identifier — must
be unique within the matrix and must match the ids you reference in
`perRun`. Validation rejects unknown ids.
- `behaviors[].kind`: `"target"` or `"failure"`. **Polarity matters** —
the grid renders green for a `target` cell that the run hit, red for
a `failure` cell that the run exhibited. Pick the framing that makes
the axis sharpest: "Cites file paths" (target, green when present) vs.
"Doesn't cite file paths" (failure, red when present) — generally the
rarer half should be the named axis so cells fill more sparsely. Use
`failure` for things the agent shouldn't do, `target` for things the
agent should do. A filled cell asserts the polarity *for that run*,
not just factual presence: a behavior can be true of a run and still
not be a fault for it — a run that avoided the underlying issue by
construction had nothing to surface, and a `failure` cell there paints
the strongest run red for doing the right thing. Likewise don't fill a
`target` cell for work that's actually off-target scope (edits to a
lookalike flow the prompt never asked about). If the polarity doesn't
hold for every run you'd mark, reframe the axis or leave that run's
cell empty.
- `behaviors[].label`: ≤ 6 words, render-time column header. Sentence
case ("Hallucinates withdraw UI"), not Title Case.
- `behaviors[].description`: 1-2 sentences. The operational definition
the reader can re-apply. Render-time tooltip.
- `perRun`: keyed on the **run directory name** (e.g.
`"reward-0.44-LQrU9Cg"`), value is an array of behavior ids. Empty
array is fine — it means "this run exhibits none of the listed
behaviors," which is itself a signal.
Before you build the matrix, list the actual `reference-runs/<run-id>/`
directories and take your `perRun` keys from that listing verbatim. Every
run in `reference-runs/` should appear in `perRun`, and every `perRun`
key must match one of those directories exactly. A report whose keys
cite run ids that don't exist on disk is describing an earlier
generation of runs — it's invalid no matter how good the axes look, so
re-derive the matrix from the current runs rather than ship it. Runs you
don't list will render as empty rows.
Cell values need the same discipline as the keys. A filled cell is a
claim about a specific run: before you emit it, ground it in a specific
quote or line from *that run's* `grade.md` or `answer.md` that you
actually read. Check the run's own framing — agents often explicitly
disclaim a behavior (a "Not covered" section, "static linting is not a
full audit") that a skim of the diff would credit them with, and a run
that hedges its scope is different from one that declares the work
"complete and verified." Check how the run ended, too: a run cut off
mid-work (crash, API error partway through implementing) never got to
decide what to omit, so don't read its omissions as final behavioral
choices. A cell you can't ground in the run's own text stays **empty**
— an unmarked cell is neutral; note the ambiguity in the per-behavior
notes as unclear rather than guessing, because a guessed cell is a
false claim about a run the reader can check.
**Body sections**, in this order:
```markdown
# Run-behaviors extraction: <slug>
## How the runs differ
2-4 paragraphs. Articulate the *shape* of the diversity — "two runs
attack the prompt from a compliance angle, one writes a workspace audit,
one hallucinates a UI" — before showing the matrix. The body is what a
reader gets if they want the qualitative narrative; the matrix is what
they glance at. When the shape is convergence — no run demonstrates the
strong path, or every run lands on the same failure — say so plainly as
an observation. Convergence is often the intended shape of the task, so
describe it; don't label it a defect.
## Per-behavior notes
For each behavior you pulled out, give 1-2 sentences explaining what
counts as exhibiting it and which run is the canonical example. Quote
from `grade.md` or `answer.md` when the line between "exhibits" and
"doesn't" is subtle. If you left a run's cell empty because you couldn't
ground it either way, say so here ("unclear for reward-0.53-…: neither
the grade nor the answer addresses it") instead of silently omitting —
the empty cell and the note together are the honest representation.
> "We're going to defer demoing the withdraw flow to a follow-up turn"
> — reward-0.47-JWSWFw3, answer.md ¶3 (the prior phrase being the
> load-bearing tell)
```
Don't restate the rubric. If a behavior column lines up with a rubric
issue, the reader will see that from the rubric-issue grid — your column
is adding new signal, not redundant signal.
## What about `not-applicable`?
If `reference-runs/` is empty or has only one run, there's nothing to
build a discrimination matrix from. Emit:
```yaml
verdict: not-applicable
confidence: HIGH
```
… with a body that explains which trigger fired ("only one reference
run") and stop. Don't try to find behaviors a single run "exhibits" — a
1-row matrix is noise, and the outlier highlights need ≥ 2 rows to
compute against.