added 260907 version of worker toolkit

This commit is contained in:
2026-09-08 20:30:51 -04:00
parent cbd3f0f8ca
commit 97aca37663
1283 changed files with 142951 additions and 0 deletions

View File

@@ -0,0 +1,87 @@
# Worker detector shell
This file is the shared boilerplate for all detector self-check skills
under `.claude/skills/` in the worker toolkit. Each detector's `SKILL.md`
points at this file plus its own `core.md` so the output path convention,
the directory-creation step, and the overwrite semantics don't duplicate
across detectors.
## Resolve the rubric target first
Wherever your detector's `core.md` reads or assesses the holistic rubric
(older cores call it "the grader guidance"), it means the file the grader
will actually use. Resolve it before reading anything:
```
bash scripts/guidance-target.sh <slug>
```
It prints the path to the task's holistic rubric: `tests/holistic-rubric.md`
on a task created with this toolkit. A task from an earlier toolkit carries
the same document as `tests/grader-guidance-consolidated.md`, or as
`tests/grader-guidance.md` on the oldest tasks, and the resolver prints
whichever file the task has.
Assess that file, and assess it against the structure the grading standard
expects: Task context, optional Business context, Ground truth, one section
per criterion in the standard's order, and optional Heavy penalties.
Never assess a file the resolver did not name.
Open the report body with one line naming what you assessed:
```
Assessed: <resolved-path>
```
## Output path
Compose the detector report (YAML frontmatter + markdown body) per the
schema in your detector's `core.md`. The frontmatter must include
`detector`, `verdict`, and `confidence`; detectors with structured
payloads (`detector-fact-check-rubric-claims` → `claims`, `detector-run-behaviors` →
`runBehaviors`) embed them in the same frontmatter block as the other
fields.
Write the report to:
```
harbor-tasks/<slug>/detectors/<detector-name>.md
```
`<detector-name>` is the value of the `detector:` frontmatter field —
e.g. `detector-snapshot-leakage`, `detector-cross-task-reference`, `detector-meaningful-failure`,
`detector-rubric-clarity`, `detector-rubric-generality`, `detector-rubric-coverage`, `detector-rubric-form`,
`detector-answer-obviousness`, `detector-good-response-defined`,
`detector-good-response-exhaustiveness`, `detector-dimension-misapplication`, `detector-broken-dev-env`,
`detector-over-hinting`, `detector-offline-verifiability`, `detector-credential-leakage`,
`detector-fact-check-rubric-claims`, `detector-run-behaviors`.
## Directory + overwrite semantics
- Create the `detectors/` directory if it doesn't exist (`mkdir -p`).
- Overwrite the file if it already exists from a prior run. Detector
skills are meant to be re-runnable — every time you tweak the rubric,
the snapshot, or anything else this detector reads, re-run the skill
and read the fresh report.
## Record what the report assessed
After writing (or rewriting) the report, stamp it with the checksums of
the task inputs it assessed:
```
npx tsx scripts/record-detector-inputs.ts <slug> <detector-name>
```
This writes `harbor-tasks/<slug>/detectors/<detector-name>.inputs.json`.
`submit-task.ts` compares those checksums against the task at packaging
time and warns when the report predates a prompt or rubric edit —
by content, so it stays accurate even when file timestamps get disturbed.
An unstamped report falls back to the less reliable timestamp comparison.
Re-run the command after every re-run of the detector.
## Verdict and body schema live in core.md
Each detector's `core.md` is the source of truth for its verdict enum and
the body section structure. Don't restate them in `SKILL.md` — read
`core.md` and use the values it specifies.