added stocks app codebase and md
This commit is contained in:
@@ -0,0 +1,46 @@
|
||||
---
|
||||
name: detector-run-behaviors
|
||||
description: |
|
||||
Self-check the diversity of your task's reference runs by pulling out a
|
||||
small set of discriminating behavior axes — named behaviors that
|
||||
distinguish runs from each other (framing choices, hallucinations,
|
||||
citation style, etc.) and emitting a structured behaviors × runs
|
||||
matrix. Useful as a sanity check before submission: if your reference
|
||||
runs all behave identically along every dimension you can name, the
|
||||
task probably isn't discriminating enough.
|
||||
allowed-tools: Bash, Read, Write
|
||||
---
|
||||
|
||||
# Run-behaviors extractor
|
||||
|
||||
This skill helps you see how your reference runs differ from each other.
|
||||
It pulls out 5–10 behavior axes — named behaviors that distinguish
|
||||
runs from each other (framing, investigation depth, hallucinations,
|
||||
citation style, hedging) — and writes a structured matrix you can use
|
||||
to confirm your task is producing genuinely diverse failure modes.
|
||||
|
||||
**This skill needs at least 2 reference runs.** Run your task with
|
||||
`scripts/harbor-run harbor-tasks/<slug> -k 4` (or similar) first so there
|
||||
are multiple `grade.md` and `answer.md` files to compare; with fewer
|
||||
than 2 runs there's nothing to discriminate against and the detector
|
||||
returns `not-applicable`.
|
||||
|
||||
Read these before starting:
|
||||
|
||||
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs. Detectors with structured payloads (this one's `runBehaviors` matrix) embed them in the same frontmatter block as `detector`/`verdict`/`confidence`.
|
||||
2. `.claude/skills/detector-run-behaviors/core.md` — what makes a good behavior axis, the structured `runBehaviors` payload shape, verdict enums, body sections.
|
||||
|
||||
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
|
||||
|
||||
## Acting on the verdict
|
||||
|
||||
- **`summary`** with `HIGH` confidence — your runs differ along clear,
|
||||
named axes. Good: that's the signal that says your task is
|
||||
discriminating enough to produce a useful score distribution.
|
||||
- **`summary`** with `MEDIUM` or `LOW` confidence — your runs look
|
||||
similar to each other and the axes you pulled out feel forced. The
|
||||
task may not be producing enough diversity to be a meaningful
|
||||
benchmark. Consider whether the prompt is too prescriptive, or
|
||||
whether more reference runs would surface real variation.
|
||||
- **`not-applicable`** — fewer than 2 reference runs. Run more trials
|
||||
first.
|
||||
Reference in New Issue
Block a user