--- name: detector-run-behaviors description: | Self-check the diversity of your task's reference runs by pulling out a small set of discriminating behavior axes — named behaviors that distinguish runs from each other (framing choices, hallucinations, citation style, etc.) and emitting a structured behaviors × runs matrix. Useful as a sanity check before submission: if your reference runs all behave identically along every dimension you can name, the task probably isn't discriminating enough. allowed-tools: Bash, Read, Write --- # Run-behaviors extractor This skill helps you see how your reference runs differ from each other. It pulls out 5–10 behavior axes — named behaviors that distinguish runs from each other (framing, investigation depth, hallucinations, citation style, hedging) — and writes a structured matrix you can use to confirm your task is producing genuinely diverse failure modes. **This skill needs at least 2 reference runs.** Run your task with `scripts/harbor-run harbor-tasks/ -k 4` (or similar) first so there are multiple `grade.md` and `answer.md` files to compare; with fewer than 2 runs there's nothing to discriminate against and the detector returns `not-applicable`. Read these before starting: 1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs. Detectors with structured payloads (this one's `runBehaviors` matrix) embed them in the same frontmatter block as `detector`/`verdict`/`confidence`. 2. `.claude/skills/detector-run-behaviors/core.md` — what makes a good behavior axis, the structured `runBehaviors` payload shape, verdict enums, body sections. Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`. ## Acting on the verdict - **`summary`** with `HIGH` confidence — your runs differ along clear, named axes. Good: that's the signal that says your task is discriminating enough to produce a useful score distribution. - **`summary`** with `MEDIUM` or `LOW` confidence — your runs look similar to each other and the axes you pulled out feel forced. The task may not be producing enough diversity to be a meaningful benchmark. Consider whether the prompt is too prescriptive, or whether more reference runs would surface real variation. - **`not-applicable`** — fewer than 2 reference runs. Run more trials first.