Files
Eric Bell 392781f7aa chore: init commit
in worker.../repo/GITFOLDER.zip is the .git folder.
2026-08-11 14:44:09 -04:00

2.4 KiB
Raw Permalink Blame History

name, description, allowed-tools
name description allowed-tools
detector-run-behaviors Self-check the diversity of your task's reference runs by pulling out a small set of discriminating behavior axes — named behaviors that distinguish runs from each other (framing choices, hallucinations, citation style, etc.) and emitting a structured behaviors × runs matrix. Useful as a sanity check before submission: if your reference runs all behave identically along every dimension you can name, the task probably isn't discriminating enough. Bash, Read, Write

Run-behaviors extractor

This skill helps you see how your reference runs differ from each other. It pulls out 5–10 behavior axes — named behaviors that distinguish runs from each other (framing, investigation depth, hallucinations, citation style, hedging) — and writes a structured matrix you can use to confirm your task is producing genuinely diverse failure modes.

This skill needs at least 2 reference runs. Run your task with scripts/harbor-run harbor-tasks/<slug> -k 4 (or similar) first so there are multiple grade.md and answer.md files to compare; with fewer than 2 runs there's nothing to discriminate against and the detector returns not-applicable.

Read these before starting:

  1. .claude/skills/_detector-worker-shell.md — where to write the report and how to handle re-runs. Detectors with structured payloads (this one's runBehaviors matrix) embed them in the same frontmatter block as detector/verdict/confidence.
  2. .claude/skills/detector-run-behaviors/core.md — what makes a good behavior axis, the structured runBehaviors payload shape, verdict enums, body sections.

Compose the report per the schema in core.md and write it per _detector-worker-shell.md.

Acting on the verdict

  • summary with HIGH confidence — your runs differ along clear, named axes. Good: that's the signal that says your task is discriminating enough to produce a useful score distribution.
  • summary with MEDIUM or LOW confidence — your runs look similar to each other and the axes you pulled out feel forced. The task may not be producing enough diversity to be a meaningful benchmark. Consider whether the prompt is too prescriptive, or whether more reference runs would surface real variation.
  • not-applicable — fewer than 2 reference runs. Run more trials first.