# Worker detector shell This file is the shared boilerplate for all detector self-check skills under `.claude/skills/` in the worker toolkit. Each detector's `SKILL.md` points at this file plus its own `core.md` so the output path convention, the directory-creation step, and the overwrite semantics don't duplicate across detectors. ## Resolve the guidance target first A task directory can carry two guidance files: `tests/grader-guidance-consolidated.md` (the Consolidated Grading Standard) and `tests/grader-guidance.md` (the legacy standard). Wherever your detector's `core.md` reads or assesses "the grader guidance", it means the file the grader will actually use. Resolve it before reading anything: ``` bash scripts/guidance-target.sh ``` It prints the guidance file's path and the standard it grades under (`consolidated` or `legacy`), using the same rule as the grader itself. Assess that file, and assess it against its own standard's structure: - **Consolidated**: Task context, optional Business context, Ground truth, one section per criterion in the standard's order, optional Heavy penalties. - **Legacy**: Task context, Business context, what strong and weak responses look like, Ground truth, optional Supporting evidence, optional Correctness, optional Heavy penalties. Never flag a document for not following the other standard's structure, and never assess the file the resolver did not name. To assess the legacy file deliberately (for a task being graded with `GRADING_STANDARD=legacy`), set that variable when running the resolver. Open the report body with one line naming what you assessed: ``` Assessed: ( standard) ``` ## Output path Compose the detector report (YAML frontmatter + markdown body) per the schema in your detector's `core.md`. The frontmatter must include `detector`, `verdict`, and `confidence`; detectors with structured payloads (`detector-fact-check-rubric-claims` → `claims`, `detector-run-behaviors` → `runBehaviors`) embed them in the same frontmatter block as the other fields. Write the report to: ``` harbor-tasks//detectors/.md ``` `` is the value of the `detector:` frontmatter field — e.g. `detector-snapshot-leakage`, `detector-cross-task-reference`, `detector-meaningful-failure`, `detector-rubric-clarity`, `detector-rubric-generality`, `detector-answer-obviousness`, `detector-good-response-defined`, `detector-good-response-exhaustiveness`, `detector-dimension-misapplication`, `detector-broken-dev-env`, `detector-over-hinting`, `detector-offline-verifiability`, `detector-credential-leakage`, `detector-fact-check-rubric-claims`, `detector-run-behaviors`. ## Directory + overwrite semantics - Create the `detectors/` directory if it doesn't exist (`mkdir -p`). - Overwrite the file if it already exists from a prior run. Detector skills are meant to be re-runnable — every time you tweak the rubric, the snapshot, or anything else this detector reads, re-run the skill and read the fresh report. ## Record what the report assessed After writing (or rewriting) the report, stamp it with the checksums of the task inputs it assessed: ``` npx tsx scripts/record-detector-inputs.ts ``` This writes `harbor-tasks//detectors/.inputs.json`. `submit-task.ts` compares those checksums against the task at packaging time and warns when the report predates a prompt or grader-guidance edit — by content, so it stays accurate even when file timestamps get disturbed. An unstamped report falls back to the less reliable timestamp comparison. Re-run the command after every re-run of the detector. ## Verdict and body schema live in core.md Each detector's `core.md` is the source of truth for its verdict enum and the body section structure. Don't restate them in `SKILL.md` — read `core.md` and use the values it specifies.