Files
project-work/worker-toolkit-stocks-in-the-future/.claude/skills/_detector-worker-shell.md

3.9 KiB

Worker detector shell

This file is the shared boilerplate for all detector self-check skills under .claude/skills/ in the worker toolkit. Each detector's SKILL.md points at this file plus its own core.md so the output path convention, the directory-creation step, and the overwrite semantics don't duplicate across detectors.

Resolve the guidance target first

A task directory can carry two guidance files: tests/grader-guidance-consolidated.md (the Consolidated Grading Standard) and tests/grader-guidance.md (the legacy standard). Wherever your detector's core.md reads or assesses "the grader guidance", it means the file the grader will actually use. Resolve it before reading anything:

bash scripts/guidance-target.sh <slug>

It prints the guidance file's path and the standard it grades under (consolidated or legacy), using the same rule as the grader itself. Assess that file, and assess it against its own standard's structure:

  • Consolidated: Task context, optional Business context, Ground truth, one section per criterion in the standard's order, optional Heavy penalties.
  • Legacy: Task context, Business context, what strong and weak responses look like, Ground truth, optional Supporting evidence, optional Correctness, optional Heavy penalties.

Never flag a document for not following the other standard's structure, and never assess the file the resolver did not name. To assess the legacy file deliberately (for a task being graded with GRADING_STANDARD=legacy), set that variable when running the resolver.

Open the report body with one line naming what you assessed:

Assessed: <resolved-path> (<consolidated|legacy> standard)

Output path

Compose the detector report (YAML frontmatter + markdown body) per the schema in your detector's core.md. The frontmatter must include detector, verdict, and confidence; detectors with structured payloads (detector-fact-check-rubric-claims → claims, detector-run-behaviors → runBehaviors) embed them in the same frontmatter block as the other fields.

Write the report to:

harbor-tasks/<slug>/detectors/<detector-name>.md

<detector-name> is the value of the detector: frontmatter field — e.g. detector-snapshot-leakage, detector-cross-task-reference, detector-meaningful-failure, detector-rubric-clarity, detector-rubric-generality, detector-answer-obviousness, detector-good-response-defined, detector-good-response-exhaustiveness, detector-dimension-misapplication, detector-broken-dev-env, detector-over-hinting, detector-offline-verifiability, detector-credential-leakage, detector-fact-check-rubric-claims, detector-run-behaviors.

Directory + overwrite semantics

  • Create the detectors/ directory if it doesn't exist (mkdir -p).
  • Overwrite the file if it already exists from a prior run. Detector skills are meant to be re-runnable — every time you tweak the rubric, the snapshot, or anything else this detector reads, re-run the skill and read the fresh report.

Record what the report assessed

After writing (or rewriting) the report, stamp it with the checksums of the task inputs it assessed:

npx tsx scripts/record-detector-inputs.ts <slug> <detector-name>

This writes harbor-tasks/<slug>/detectors/<detector-name>.inputs.json. submit-task.ts compares those checksums against the task at packaging time and warns when the report predates a prompt or grader-guidance edit — by content, so it stays accurate even when file timestamps get disturbed. An unstamped report falls back to the less reliable timestamp comparison. Re-run the command after every re-run of the detector.

Verdict and body schema live in core.md

Each detector's core.md is the source of truth for its verdict enum and the body section structure. Don't restate them in SKILL.md — read core.md and use the values it specifies.