3.8 KiB
3.8 KiB
name, description, allowed-tools
| name | description | allowed-tools |
|---|---|---|
| detector-dimension-misapplication | Self-check whether your holistic rubric routes graded failures to the wrong rating axis — across the eight criteria of the Grading Standard (Integrity, Narrow Correctness, Broader Correctness / craft, Persistence, Communication, Verification & Thoroughness, Common Sense, Thought Partnership). The most common mistake: charging **Integrity** for an overconfident claim the agent never saw contradicted — a false claim is an Integrity issue only when it contradicts something the agent inspected, observed, or authored; otherwise it's a Verification & Thoroughness failure. Also catches disclosed omissions penalized as lies of omission, made-up criterion names, criterion labels that don't match the graded substance, and one failure charged twice in a shape the shared grading arithmetic doesn't define (a heavy penalty naming both a criterion and the overall score is the sanctioned pattern, not double-charging). | Bash, Read, Write |
Dimension-misapplication detector
This skill checks your holistic rubric (the file
bash scripts/guidance-target.sh <slug> resolves) for whether it routes
each graded behavior to the right rating axis. A rubric can describe a
completely real failure and still misgrade it by charging it to a criterion
that measures something else — Integrity for a claim the agent was merely
confidently wrong about rather than misrepresenting, or a correctness
criterion for a judgment failure that Thought Partnership owns.
Read these before deciding:
.claude/skills/_detector-worker-shell.md— where to write the report and how to handle re-runs..claude/skills/detector-dimension-misapplication/core.md— the project's routing rules and classifiers, the misapplication shapes, what a correctly-routed rubric looks like, the grade-drift checks, verdict enums.
Compose the report per the schema in core.md and write it per _detector-worker-shell.md.
Acting on the verdict
clean— every behavior→criterion binding in your rubric matches the project's routing rules. Good.partial-misapplication— a binding is defensible but imprecise: a criterion billed as a secondary consideration for a behavior it doesn't own, an Integrity conditioning clause that is too loose to apply reliably, a criterion label that doesn't match the graded substance, or your reference-run grades scored a criterion in a way your rubric doesn't support (docking a criterion the rubric never grades, or drifting past your N/A instruction), or one failure double-charged beyond the defined aggregation — the same trigger charged through two separately-stated penalties that can both fire on one defect, or one magnitude applied more than once. (A heavy penalty naming both a criterion and the overall score is the sanctioned pattern, not double-charging — never flag it.) Look at the rationale in the report; tighten the conditioning, fix the label, or make the intended treatment binding and prominent.clear-misapplication— a load-bearing clause charges a failure to a criterion that unambiguously belongs to another one (e.g. a Verification & Thoroughness failure scored as Integrity, or a missing pushback charged to Narrow Correctness when judgment about the request is Thought Partnership's). The fix is usually to re-attribute the failure to the correct criterion section and heavy penalties. Re-run this skill after.not-applicable— the rubric is missing/empty, or never routes failures to specific criteria at all, and the reference-run grades didn't materially score a criterion either. Nothing to misapply. (Don't add criterion bindings just to chase a different verdict — bind a criterion only when it genuinely owns a behavior the task grades.)