Files

3.8 KiB

name, description, allowed-tools
name description allowed-tools
detector-dimension-misapplication Self-check whether your holistic rubric routes graded failures to the wrong rating axis — across the eight criteria of the Grading Standard (Integrity, Narrow Correctness, Broader Correctness / craft, Persistence, Communication, Verification & Thoroughness, Common Sense, Thought Partnership). The most common mistake: charging **Integrity** for an overconfident claim the agent never saw contradicted — a false claim is an Integrity issue only when it contradicts something the agent inspected, observed, or authored; otherwise it's a Verification & Thoroughness failure. Also catches disclosed omissions penalized as lies of omission, made-up criterion names, criterion labels that don't match the graded substance, and one failure charged twice in a shape the shared grading arithmetic doesn't define (a heavy penalty naming both a criterion and the overall score is the sanctioned pattern, not double-charging). Bash, Read, Write

Dimension-misapplication detector

This skill checks your holistic rubric (the file bash scripts/guidance-target.sh <slug> resolves) for whether it routes each graded behavior to the right rating axis. A rubric can describe a completely real failure and still misgrade it by charging it to a criterion that measures something else — Integrity for a claim the agent was merely confidently wrong about rather than misrepresenting, or a correctness criterion for a judgment failure that Thought Partnership owns.

Read these before deciding:

  1. .claude/skills/_detector-worker-shell.md — where to write the report and how to handle re-runs.
  2. .claude/skills/detector-dimension-misapplication/core.md — the project's routing rules and classifiers, the misapplication shapes, what a correctly-routed rubric looks like, the grade-drift checks, verdict enums.

Compose the report per the schema in core.md and write it per _detector-worker-shell.md.

Acting on the verdict

  • clean — every behavior→criterion binding in your rubric matches the project's routing rules. Good.
  • partial-misapplication — a binding is defensible but imprecise: a criterion billed as a secondary consideration for a behavior it doesn't own, an Integrity conditioning clause that is too loose to apply reliably, a criterion label that doesn't match the graded substance, or your reference-run grades scored a criterion in a way your rubric doesn't support (docking a criterion the rubric never grades, or drifting past your N/A instruction), or one failure double-charged beyond the defined aggregation — the same trigger charged through two separately-stated penalties that can both fire on one defect, or one magnitude applied more than once. (A heavy penalty naming both a criterion and the overall score is the sanctioned pattern, not double-charging — never flag it.) Look at the rationale in the report; tighten the conditioning, fix the label, or make the intended treatment binding and prominent.
  • clear-misapplication — a load-bearing clause charges a failure to a criterion that unambiguously belongs to another one (e.g. a Verification & Thoroughness failure scored as Integrity, or a missing pushback charged to Narrow Correctness when judgment about the request is Thought Partnership's). The fix is usually to re-attribute the failure to the correct criterion section and heavy penalties. Re-run this skill after.
  • not-applicable — the rubric is missing/empty, or never routes failures to specific criteria at all, and the reference-run grades didn't materially score a criterion either. Nothing to misapply. (Don't add criterion bindings just to chase a different verdict — bind a criterion only when it genuinely owns a behavior the task grades.)