3.5 KiB
Worker detector shell
This file is the shared boilerplate for all detector self-check skills
under .claude/skills/ in the worker toolkit. Each detector's SKILL.md
points at this file plus its own core.md so the output path convention,
the directory-creation step, and the overwrite semantics don't duplicate
across detectors.
Resolve the rubric target first
Wherever your detector's core.md reads or assesses the holistic rubric
(older cores call it "the grader guidance"), it means the file the grader
will actually use. Resolve it before reading anything:
bash scripts/guidance-target.sh <slug>
It prints the path to the task's holistic rubric: tests/holistic-rubric.md
on a task created with this toolkit. A task from an earlier toolkit carries
the same document as tests/grader-guidance-consolidated.md, or as
tests/grader-guidance.md on the oldest tasks, and the resolver prints
whichever file the task has.
Assess that file, and assess it against the structure the grading standard
expects: Task context, optional Business context, Ground truth, one section
per criterion in the standard's order, and optional Heavy penalties.
Never assess a file the resolver did not name.
Open the report body with one line naming what you assessed:
Assessed: <resolved-path>
Output path
Compose the detector report (YAML frontmatter + markdown body) per the
schema in your detector's core.md. The frontmatter must include
detector, verdict, and confidence; detectors with structured
payloads (detector-fact-check-rubric-claims → claims, detector-run-behaviors →
runBehaviors) embed them in the same frontmatter block as the other
fields.
Write the report to:
harbor-tasks/<slug>/detectors/<detector-name>.md
<detector-name> is the value of the detector: frontmatter field —
e.g. detector-snapshot-leakage, detector-cross-task-reference, detector-meaningful-failure,
detector-rubric-clarity, detector-rubric-generality, detector-rubric-coverage, detector-rubric-form,
detector-answer-obviousness, detector-good-response-defined,
detector-good-response-exhaustiveness, detector-dimension-misapplication, detector-broken-dev-env,
detector-over-hinting, detector-offline-verifiability, detector-credential-leakage,
detector-fact-check-rubric-claims, detector-run-behaviors.
Directory + overwrite semantics
- Create the
detectors/directory if it doesn't exist (mkdir -p). - Overwrite the file if it already exists from a prior run. Detector skills are meant to be re-runnable — every time you tweak the rubric, the snapshot, or anything else this detector reads, re-run the skill and read the fresh report.
Record what the report assessed
After writing (or rewriting) the report, stamp it with the checksums of the task inputs it assessed:
npx tsx scripts/record-detector-inputs.ts <slug> <detector-name>
This writes harbor-tasks/<slug>/detectors/<detector-name>.inputs.json.
submit-task.ts compares those checksums against the task at packaging
time and warns when the report predates a prompt or rubric edit —
by content, so it stays accurate even when file timestamps get disturbed.
An unstamped report falls back to the less reliable timestamp comparison.
Re-run the command after every re-run of the detector.
Verdict and body schema live in core.md
Each detector's core.md is the source of truth for its verdict enum and
the body section structure. Don't restate them in SKILL.md — read
core.md and use the values it specifies.