added potion-polyglot worker folder w/o repos

This commit is contained in:
2026-09-08 21:58:19 -04:00
parent f10303b8c2
commit a16a457669
211 changed files with 41614 additions and 0 deletions

View File

@@ -0,0 +1,46 @@
---
name: detector-run-behaviors
description: |
Self-check the diversity of your task's reference runs by pulling out a
small set of discriminating behavior axes — named behaviors that
distinguish runs from each other (framing choices, hallucinations,
citation style, etc.) and emitting a structured behaviors × runs
matrix. Useful as a sanity check before submission: if your reference
runs all behave identically along every dimension you can name, the
task probably isn't discriminating enough.
allowed-tools: Bash, Read, Write
---
# Run-behaviors extractor
This skill helps you see how your reference runs differ from each other.
It pulls out 5–10 behavior axes — named behaviors that distinguish
runs from each other (framing, investigation depth, hallucinations,
citation style, hedging) — and writes a structured matrix you can use
to confirm your task is producing genuinely diverse failure modes.
**This skill needs at least 2 reference runs.** Run your task with
`scripts/harbor-run harbor-tasks/<slug> -k 4` (or similar) first so there
are multiple `grade.md` and `answer.md` files to compare; with fewer
than 2 runs there's nothing to discriminate against and the detector
returns `not-applicable`.
Read these before starting:
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs. Detectors with structured payloads (this one's `runBehaviors` matrix) embed them in the same frontmatter block as `detector`/`verdict`/`confidence`.
2. `.claude/skills/detector-run-behaviors/core.md` — what makes a good behavior axis, the structured `runBehaviors` payload shape, verdict enums, body sections.
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
## Acting on the verdict
- **`summary`** with `HIGH` confidence — your runs differ along clear,
named axes. Good: that's the signal that says your task is
discriminating enough to produce a useful score distribution.
- **`summary`** with `MEDIUM` or `LOW` confidence — your runs look
similar to each other and the axes you pulled out feel forced. The
task may not be producing enough diversity to be a meaningful
benchmark. Consider whether the prompt is too prescriptive, or
whether more reference runs would surface real variation.
- **`not-applicable`** — fewer than 2 reference runs. Run more trials
first.