2.9 KiB
2.9 KiB
name, description, allowed-tools
| name | description | allowed-tools |
|---|---|---|
| detector-rubric-generality | Self-check your holistic rubric for whether it describes, in general, what makes a response strong or weak — so a grader can apply it to any agent — or whether it speaks too much in terms of your reference runs ("clarity is reliably high on this task", "agents will fail here", "all four trials hit 85+"). Identifying failure modes as general response properties is good; leaning on what the observed runs did as the scoring basis is what this catches. Doesn't flag illustrative pointers to runs or describing failure modes — only run-anchoring that gates scoring. Also flags a rubric that names the framework your task runs on (Harbor, Pier, the sandbox) instead of describing the task in its own terms. | Bash, Read, Write |
Rubric-generality detector
This skill checks whether your holistic rubric (the file
bash scripts/guidance-target.sh <slug> resolves) describes response quality in
general terms — so the task works for any agent, not just the ones whose
reference runs you have today — or whether it leans too much on what the
observed runs happened to do ("reliably high on this task," "agents will," "all
N trials," tiers keyed to a specific run). It also flags a rubric that names the
framework your task runs on (Harbor, Pier, the sandbox) instead of the task's
own terms — "the final Harbor instruction" should just read "the final
instruction."
Read these before deciding:
.claude/skills/_detector-worker-shell.md— where to write the report and how to handle re-runs..claude/skills/detector-rubric-generality/core.md— what counts as run-anchored scoring vs. general response-quality description, verdict definitions, frontmatter/body schema.
Compose the report per the schema in core.md and write it per _detector-worker-shell.md.
Acting on the verdict
generalizes— the main thrust describes what makes a response strong or weak in general terms; any run-references are illustrative. Good.minor-issues— the core scoring is general, but some phrasings lean on observed-run statistics or "agents tend to" framing, or name the framework your task runs on. Look at the "run-anchored phrasings" and "infra-framework references" lists in the report and reframe each as a general property of a response (or, for an infra name, reword to the task's own terms). No need to rebuild the rubric.material-issues— the load-bearing scoring criteria are defined by what the reference runs did, so a grader couldn't score a new agent that fails differently. Look at the "load-bearing run-dependence" section — rewrite those criteria to describe what a strong/weak response looks like in general, then re-run this skill.not-applicable— the resolved rubric file is missing, empty, or template-only. Write the rubric first, then come back to this skill.