Files
project-work/worker-toolkit-potion-polyglot/.claude/skills/write-atomic-rubric/SKILL.md

11 KiB

name, description
name description
write-atomic-rubric Convert a task's finished holistic rubric into the atomic rubric package — tests/atomic-rubric.yaml (criteria with id, category, severity, dimensions, guideline, elaboration) plus tests/grader-context.md (task context, business context, and ground truth, extracted verbatim). Covers Maximum Viable Atomicity, positive guideline phrasing with bold inline answer keys, conditional criteria, dodged-bullet escalation pairs, Crux designation from the holistic rubric's heavy penalties (at most two per task), the schema rules (2-24 criteria; kebab-case ids; no numeric penalty language; no severity on extra_credit), and staging and validation. Use after the holistic rubric is final.

Writing the Atomic Rubric

What this is

The atomic rubric restates a task's holistic rubric as a list of small, independently judgeable criteria. A rubric grader reads each criterion, investigates the run, and emits one verdict per criterion; the per-criterion verdicts combine into the task score. The conversion produces two files in the task's tests/ directory:

  • tests/atomic-rubric.yaml — every task-specific requirement as an atomic criterion.
  • tests/grader-context.md — the generalized sections the grader reads once: task context, business context, and ground truth.

The source is the task's holistic rubric: tests/holistic-rubric.md, or on older tasks tests/grader-guidance-consolidated.md or tests/grader-guidance.md. Older tasks also carry the atomic file under its earlier name, tests/rubrics.yaml; tools read both names, and a task keeps the file name it already has. Never rename a committed file, and never edit the source document during conversion; the conversion is a restatement, not a revision. If you find a defect in the source, fix the source first under the write-holistic-rubric skill, then convert.

grader-context.md

Extract the source's Task context, Business context, and Ground truth sections verbatim. Title the file # Grader Context — <task-slug>. The one sanctioned rewording is an internal cross-reference: where the source text points at a section that no longer exists as a section ("see Heavy penalties"), point it at the criterion that now owns the rule. If the source has no Business context section, extract what exists. Never invent content, and never summarize: a grader calibrated by a paraphrase is calibrated wrong.

atomic-rubric.yaml

Top-level keys:

task: <task-slug>
source: harbor-tasks/<task-slug>/tests/holistic-rubric.md
context: grader-context.md
criteria:
  - ...

task is the slug exactly. source is the repo-relative path of the document you converted from, under whichever name the task carries. Write guideline and elaboration as YAML literal block scalars (|) so markdown survives intact.

Each criterion carries:

  • id — a kebab-case slug, unique within the file, stable once written, and descriptive enough to be quoted on its own ("names-the-injected-config-key").
  • category — one of three values. primary_intent marks a requirement at the heart of what the task asks for. extra_credit marks a valuable behavior beyond the task's requirements; it can only raise the score, and a response that does not earn it loses nothing. dodged_bullet marks a specific failure the response must avoid; a response that avoids it passes the criterion.
  • severity — how heavily a failed criterion weighs in the score: crux, certain_dealbreaker, possible_dealbreaker, or unlikely_dealbreaker (displayed as Crux, Critical, Major, Minor). Required on every criterion except extra_credit, which never carries one. The grader never sees severity; it judges each criterion on its own terms, and severity applies afterward.
  • dimensions — the criterion or criteria of the Grading Standard this item targets, at least one, named exactly as the standard names them: Integrity, Narrow Correctness, Broader Correctness / the craft of software engineering, Persistence, Communication, Verification & Thoroughness, Common Sense, Thought Partnership.
  • guideline — one positively phrased statement of the requirement.
  • elaboration — optional judgment guidance for the grader.

Writing criteria

  • One criterion per smallest meaningful unit. Convert at Maximum Viable Atomicity: each criterion covers one requirement that can be judged on its own. Do not chop a requirement into fragments that cannot be judged alone, and do not bundle requirements that can pass or fail independently. Parallel facts derived the same way, such as the values of one calculated column, may share a criterion. Never group facts in a way designed to over-penalize a response.
  • Phrase requirements positively. Write "The response should ..." or "The response should avoid ..."; never write "should not". Factual criteria carry their answer key inline, in bold, so the criterion is judgeable without opening another document.
  • Keep each criterion self-contained. Never reference one criterion from another. A criterion may briefly restate a fact that also lives in grader-context.md so that it stands alone; that duplication is intended, and it is the one exception to the source's say-each-thing-once rule.
  • Write conditionals as conditionals. "If the response includes a migration, it should ...". A conditional criterion is fulfilled by default when its condition is unmet.
  • Describe only the response. Every criterion states a property of the response. Notes on how to verify a claim, which evidence to trust, or how to calibrate judgment fold into the elaboration of the criterion they support; they are never criteria of their own.
  • Put judgment guidance in the elaboration. State what fulfills the criterion and what fails it, with concrete examples from the source. Where several kinds of response are acceptable, list them. Where the source names behavior that must not trip the rule (the honest or flagged variant), carry that non-trigger into the elaboration.
  • Give a strictly worse failure its own criterion. Where the source ranks one failure clearly worse than a related one, encode the worse variant as a separate dodged_bullet that fails in addition to the base criterion, so a response committing the worse failure fails both and the score reflects the difference.
  • Write criteria for likely failures. A criterion earns its place by catching behavior responses actually get wrong. Skip trivial properties every response satisfies, and never penalize behavior outside the agent's control, such as a tooling failure.
  • No numeric penalty language. Severity and category carry the weight; the text never does. No "subtract 0.35", no points, no "out of 100", in guidelines or elaborations. Validation rejects numeric penalty phrasing.
  • No generic scoring mechanics. Flooring, how verdicts aggregate, and how penalties combine live in the shared grader prompt, never in a criterion.
  • Preserve the source's facts exactly. Keep every load-bearing fact, path and line citation, and code quotation, with markdown formatting (backticks, bold, fences) intact. Never invent facts, paths, or requirements the source does not carry.

The file carries between 2 and 24 criteria; most tasks land in the teens. Every scoring-relevant rule of the source lands in exactly one criterion's guideline or elaboration. Content that is context rather than a requirement belongs in grader-context.md, not in a criterion.

Crux designation

crux is the top severity tier, reserved for the task's defining cliff. Derive it from the source's Heavy penalties section, and only from there.

  • Write one Crux criterion per heavy penalty that targets the overall score, carrying that penalty's fire conditions and its stated non-triggers.
  • A heavy penalty that targets only a criterion of the standard, not the overall score, converts at certain_dealbreaker, not Crux.
  • When one penalty fires only on a conjunction (the response did A and also claimed B), write a single criterion covering the whole conjunction, phrased so it passes or fails outright; splitting it, or leaving room for partial fulfillment, lets partial credit dilute a dealbreaker.
  • When the source spells one dealbreaker out as several facets of the same failure, merge them into one Crux criterion; never write one Crux per facet.
  • A task carries at most two Crux criteria. Where the source has more overall-score penalties than that, keep Crux on the two that define the task's failure mode and convert the rest at certain_dealbreaker.
  • Designate Crux only from the source document. Never promote a criterion to Crux because runs that failed it happened to score low.

Alignment with the holistic rubric

The two rubrics grade the same task, and their scores should agree. A run graded under the atomic rubric should land near the score the holistic rubric gives it, and runs should keep their relative order: a run the holistic rubric places far below another belongs far below it under the atomic rubric too. When atomic scores compress a gap the source creates, the missing lever is almost always Crux designation on the dealbreaker involved, not more criteria.

Validate, stage, self-check

Run the two rubric detectors after generating the package, and again after any edit:

  • /detector-rubric-coverage checks that every scoring-relevant rule of the source document lands in a criterion.
  • /detector-rubric-form checks that every criterion follows the form rules in this skill.

Fix what they flag before packaging the task; the package ships tests/atomic-rubric.yaml and tests/grader-context.md alongside the task's other files.

To grade under the atomic rubric inside the worker toolkit, stage the grading copies with npx tsx scripts/stage-atomic-rubric.ts <task-slug>. Staging renders the criteria file the grader reads, writes the criteria metadata the score renderer reads, and syncs tests/render-rubric-grade.py from task-shared/. Re-run it after every rubric edit. Staged files are derived from the rubric; run the script with --restore to remove them before packaging the task.

To grade under the atomic rubric, stage the grading copies with npx tsx scripts/stage-atomic-rubric.ts <task-slug> inside the devcontainer: staging checks the package's structure (a task key, a criteria list, a unique id plus a guideline and a category on every criterion, at most two Crux criteria), renders the criteria file the grader reads, and installs the rubric-aware harness. Staged files are working-tree only; never commit them. The /detector-rubric-form and /detector-rubric-coverage skills check the content rules (severity vocabulary, the numeric-penalty ban, coverage of the holistic rubric).

Reviewers working in a repo checkout also run npx tsx scripts/validate-rubrics-cli.ts --slug <task-slug>, which enforces the same schema, the criteria count, the Crux cap, and the numeric-penalty ban. That script is part of the review pipeline and does not ship in the toolkit.

  • .claude/skills/write-holistic-rubric/SKILL.md — the source document this skill converts; its prose ground rules and penalty phrasing apply to the source, and its attribution rules decide which dimension a criterion targets.