2.4 KiB
name, description, allowed-tools
| name | description | allowed-tools |
|---|---|---|
| detector-run-behaviors | Self-check the diversity of your task's reference runs by pulling out a small set of discriminating behavior axes — named behaviors that distinguish runs from each other (framing choices, hallucinations, citation style, etc.) and emitting a structured behaviors × runs matrix. Useful as a sanity check before submission: if your reference runs all behave identically along every dimension you can name, the task probably isn't discriminating enough. | Bash, Read, Write |
Run-behaviors extractor
This skill helps you see how your reference runs differ from each other. It pulls out 5–10 behavior axes — named behaviors that distinguish runs from each other (framing, investigation depth, hallucinations, citation style, hedging) — and writes a structured matrix you can use to confirm your task is producing genuinely diverse failure modes.
This skill needs at least 2 reference runs. Run your task with
scripts/harbor-run harbor-tasks/<slug> -k 4 (or similar) first so there
are multiple grade.md and answer.md files to compare; with fewer
than 2 runs there's nothing to discriminate against and the detector
returns not-applicable.
Read these before starting:
.claude/skills/_detector-worker-shell.md— where to write the report and how to handle re-runs. Detectors with structured payloads (this one'srunBehaviorsmatrix) embed them in the same frontmatter block asdetector/verdict/confidence..claude/skills/detector-run-behaviors/core.md— what makes a good behavior axis, the structuredrunBehaviorspayload shape, verdict enums, body sections.
Compose the report per the schema in core.md and write it per _detector-worker-shell.md.
Acting on the verdict
summarywithHIGHconfidence — your runs differ along clear, named axes. Good: that's the signal that says your task is discriminating enough to produce a useful score distribution.summarywithMEDIUMorLOWconfidence — your runs look similar to each other and the axes you pulled out feel forced. The task may not be producing enough diversity to be a meaningful benchmark. Consider whether the prompt is too prescriptive, or whether more reference runs would surface real variation.not-applicable— fewer than 2 reference runs. Run more trials first.