Files
project-work/worker-toolkit-potion-polyglot/.claude/skills/detector-rubric-form/SKILL.md

3.9 KiB

name, description, allowed-tools
name description allowed-tools
detector-rubric-form Self-check that your atomic rubric is well-formed. A deterministic contract checks the artifact: the file parses against the criterion schema, criteria number 2 to 24, ids are kebab-case and unique, category and severity use the defined vocabularies, extra_credit criteria carry no severity, at most 2 criteria are crux, `dimensions` names grading-standard criteria, and no text states a numeric penalty amount. A judgment layer checks the writing: each guideline is one positively phrased, independently judgeable requirement, criteria stand alone, factual criteria carry their answer key inline in bold, and elaborations clarify the guideline instead of adding requirements. Reads `tests/atomic-rubric.yaml` (or `tests/rubrics.yaml`) and `tests/grader-context.md`. Emits `not-applicable` when the task has no atomic rubric yet. Bash, Read, Write

Rubric-form detector

This skill checks your atomic rubric as an artifact. Each criterion is scored on its own, and the aggregate score is computed from category and severity. That only works when the file obeys the schema and each criterion states one requirement a grader can judge independently.

The failure shapes to catch:

  • Schema violations. The file fails to parse, ids repeat or are not kebab-case, a category or severity value is outside the vocabulary, an extra_credit criterion carries a severity, more than 2 criteria are crux, or dimensions is empty.
  • Numeric penalty language. A guideline, elaboration, or tests/grader-context.md sentence states a penalty amount, such as "subtract roughly 0.35". Penalty weight is expressed through category and severity. Sizing the subtraction is the grading machinery's job.
  • Negation-phrased guidelines. A guideline says "should not" or "must not" instead of stating the requirement positively. Use "The response should avoid X" for prohibitions.
  • Bundled or fragmentary criteria. One criterion packs several independent requirements, so a grader must improvise a partial verdict. Or a criterion cannot be judged without reading a sibling criterion. Parallel facts from one derivation may share a criterion.
  • Missing answer keys. A criterion grades the response for surfacing a specific fact, and the fact is not stated inline in bold in the guideline.
  • Requirements hidden in elaborations. An elaboration adds a requirement the guideline never states.
  • Unfair grading shapes. Criteria spent on trivially-satisfied properties, two criteria that both fire on one defect with no note saying which one charges, phrasing that forecloses an approach the rubric's own text treats as acceptable, or a requirement the task's environment cannot satisfy.

Read these before deciding:

  1. .claude/skills/_detector-worker-shell.md — where to write the report and how to handle re-runs.
  2. .claude/skills/detector-rubric-form/core.md — the deterministic contract with its pattern sweeps, the judgment checks, what is deliberately not a finding, verdict definitions, and the body schema.

Compose the report per the schema in core.md and write it per _detector-worker-shell.md.

Acting on the verdict

  • clear — the file passes the deterministic contract and the criteria read as a working rubric. Good.
  • minor-issues — the contract passes, and the findings are polish-level. Read the findings list and tighten the criteria. There is no need to rebuild the rubric.
  • material-issues — the file breaks the deterministic contract, or at least one criterion cannot be graded as written. Fix every finding in the deterministic-contract section first, then the judgment findings. Re-run this skill after editing.
  • not-applicable — the task has no atomic rubric yet. Write the atomic rubric first, then come back to this skill.