added potion-polyglot worker folder w/o repos
This commit is contained in:
@@ -0,0 +1,87 @@
|
||||
# Worker detector shell
|
||||
|
||||
This file is the shared boilerplate for all detector self-check skills
|
||||
under `.claude/skills/` in the worker toolkit. Each detector's `SKILL.md`
|
||||
points at this file plus its own `core.md` so the output path convention,
|
||||
the directory-creation step, and the overwrite semantics don't duplicate
|
||||
across detectors.
|
||||
|
||||
## Resolve the rubric target first
|
||||
|
||||
Wherever your detector's `core.md` reads or assesses the holistic rubric
|
||||
(older cores call it "the grader guidance"), it means the file the grader
|
||||
will actually use. Resolve it before reading anything:
|
||||
|
||||
```
|
||||
bash scripts/guidance-target.sh <slug>
|
||||
```
|
||||
|
||||
It prints the path to the task's holistic rubric: `tests/holistic-rubric.md`
|
||||
on a task created with this toolkit. A task from an earlier toolkit carries
|
||||
the same document as `tests/grader-guidance-consolidated.md`, or as
|
||||
`tests/grader-guidance.md` on the oldest tasks, and the resolver prints
|
||||
whichever file the task has.
|
||||
Assess that file, and assess it against the structure the grading standard
|
||||
expects: Task context, optional Business context, Ground truth, one section
|
||||
per criterion in the standard's order, and optional Heavy penalties.
|
||||
|
||||
Never assess a file the resolver did not name.
|
||||
|
||||
Open the report body with one line naming what you assessed:
|
||||
|
||||
```
|
||||
Assessed: <resolved-path>
|
||||
```
|
||||
|
||||
## Output path
|
||||
|
||||
Compose the detector report (YAML frontmatter + markdown body) per the
|
||||
schema in your detector's `core.md`. The frontmatter must include
|
||||
`detector`, `verdict`, and `confidence`; detectors with structured
|
||||
payloads (`detector-fact-check-rubric-claims` → `claims`, `detector-run-behaviors` →
|
||||
`runBehaviors`) embed them in the same frontmatter block as the other
|
||||
fields.
|
||||
|
||||
Write the report to:
|
||||
|
||||
```
|
||||
harbor-tasks/<slug>/detectors/<detector-name>.md
|
||||
```
|
||||
|
||||
`<detector-name>` is the value of the `detector:` frontmatter field —
|
||||
e.g. `detector-snapshot-leakage`, `detector-cross-task-reference`, `detector-meaningful-failure`,
|
||||
`detector-rubric-clarity`, `detector-rubric-generality`, `detector-rubric-coverage`, `detector-rubric-form`,
|
||||
`detector-answer-obviousness`, `detector-good-response-defined`,
|
||||
`detector-good-response-exhaustiveness`, `detector-dimension-misapplication`, `detector-broken-dev-env`,
|
||||
`detector-over-hinting`, `detector-offline-verifiability`, `detector-credential-leakage`,
|
||||
`detector-fact-check-rubric-claims`, `detector-run-behaviors`.
|
||||
|
||||
## Directory + overwrite semantics
|
||||
|
||||
- Create the `detectors/` directory if it doesn't exist (`mkdir -p`).
|
||||
- Overwrite the file if it already exists from a prior run. Detector
|
||||
skills are meant to be re-runnable — every time you tweak the rubric,
|
||||
the snapshot, or anything else this detector reads, re-run the skill
|
||||
and read the fresh report.
|
||||
|
||||
## Record what the report assessed
|
||||
|
||||
After writing (or rewriting) the report, stamp it with the checksums of
|
||||
the task inputs it assessed:
|
||||
|
||||
```
|
||||
npx tsx scripts/record-detector-inputs.ts <slug> <detector-name>
|
||||
```
|
||||
|
||||
This writes `harbor-tasks/<slug>/detectors/<detector-name>.inputs.json`.
|
||||
`submit-task.ts` compares those checksums against the task at packaging
|
||||
time and warns when the report predates a prompt or rubric edit —
|
||||
by content, so it stays accurate even when file timestamps get disturbed.
|
||||
An unstamped report falls back to the less reliable timestamp comparison.
|
||||
Re-run the command after every re-run of the detector.
|
||||
|
||||
## Verdict and body schema live in core.md
|
||||
|
||||
Each detector's `core.md` is the source of truth for its verdict enum and
|
||||
the body section structure. Don't restate them in `SKILL.md` — read
|
||||
`core.md` and use the values it specifies.
|
||||
Reference in New Issue
Block a user