lots of change - all to start my 3rd redo
This commit is contained in:
@@ -1,193 +0,0 @@
|
||||
---
|
||||
name: write-atomic-rubric
|
||||
description: Convert a task's finished holistic rubric into the atomic rubric package — tests/atomic-rubric.yaml (criteria with id, category, severity, dimensions, guideline, elaboration) plus tests/grader-context.md (task context, business context, and ground truth, extracted verbatim). Covers Maximum Viable Atomicity, positive guideline phrasing with bold inline answer keys, conditional criteria, dodged-bullet escalation pairs, Crux designation from the holistic rubric's heavy penalties (at most two per task), the schema rules (kebab-case ids; no numeric penalty language; no severity on extra_credit), and staging and validation. Use after the holistic rubric is final.
|
||||
---
|
||||
|
||||
# Writing the Atomic Rubric
|
||||
|
||||
## What this is
|
||||
|
||||
The atomic rubric restates a task's holistic rubric as a list of small, independently
|
||||
judgeable criteria. A rubric grader reads each criterion, investigates the run, and
|
||||
emits one verdict per criterion; the per-criterion verdicts combine into the task
|
||||
score. The conversion produces two files in the task's `tests/` directory:
|
||||
|
||||
- `tests/atomic-rubric.yaml` — every task-specific requirement as an atomic criterion.
|
||||
- `tests/grader-context.md` — the generalized sections the grader reads once: task
|
||||
context, business context, and ground truth.
|
||||
|
||||
The source is the task's holistic rubric: `tests/holistic-rubric.md`, or on older tasks
|
||||
`tests/grader-guidance-consolidated.md` or `tests/grader-guidance.md`. Older tasks also
|
||||
carry the atomic file under its earlier name, `tests/rubrics.yaml`; tools read both
|
||||
names, and a task keeps the file name it already has. Never rename a committed file,
|
||||
and never edit the source document during conversion; the conversion is a
|
||||
restatement, not a revision. If you find a defect in the source, fix the source first
|
||||
under the `write-holistic-rubric` skill, then convert.
|
||||
|
||||
## grader-context.md
|
||||
|
||||
Extract the source's Task context, Business context, and Ground truth sections
|
||||
**verbatim**. Title the file `# Grader Context — <task-slug>`. The one sanctioned
|
||||
rewording is an internal cross-reference: where the source text points at a section
|
||||
that no longer exists as a section ("see Heavy penalties"), point it at the criterion
|
||||
that now owns the rule. If the source has no Business context section, extract what
|
||||
exists. Never invent content, and never summarize: a grader calibrated by a paraphrase
|
||||
is calibrated wrong.
|
||||
|
||||
## atomic-rubric.yaml
|
||||
|
||||
Top-level keys:
|
||||
|
||||
```yaml
|
||||
task: <task-slug>
|
||||
source: harbor-tasks/<task-slug>/tests/holistic-rubric.md
|
||||
context: grader-context.md
|
||||
criteria:
|
||||
- ...
|
||||
```
|
||||
|
||||
`task` is the slug exactly. `source` is the repo-relative path of the document you
|
||||
converted from, under whichever name the task carries. Write `guideline` and
|
||||
`elaboration` as YAML literal block scalars (`|`) so markdown survives intact.
|
||||
|
||||
Each criterion carries:
|
||||
|
||||
- **`id`** — a kebab-case slug, unique within the file, stable once written, and
|
||||
descriptive enough to be quoted on its own ("names-the-injected-config-key").
|
||||
- **`category`** — one of three values. `primary_intent` marks a requirement at the
|
||||
heart of what the task asks for. `extra_credit` marks a valuable behavior beyond the
|
||||
task's requirements; it can only raise the score, and a response that does not earn
|
||||
it loses nothing. `dodged_bullet` marks a specific failure the response must avoid; a
|
||||
response that avoids it passes the criterion.
|
||||
- **`severity`** — how heavily a failed criterion weighs in the score: `crux`,
|
||||
`certain_dealbreaker`, `possible_dealbreaker`, or `unlikely_dealbreaker` (displayed
|
||||
as Crux, Critical, Major, Minor). Required on every criterion except `extra_credit`,
|
||||
which never carries one. The grader never sees severity; it judges each criterion on
|
||||
its own terms, and severity applies afterward.
|
||||
- **`dimensions`** — the criterion or criteria of the Grading Standard this item
|
||||
targets, at least one, named exactly as the standard names them: Integrity, Narrow
|
||||
Correctness, Broader Correctness / the craft of software engineering, Persistence,
|
||||
Communication, Verification & Thoroughness, Common Sense, Thought Partnership.
|
||||
- **`guideline`** — one positively phrased statement of the requirement.
|
||||
- **`elaboration`** — optional judgment guidance for the grader.
|
||||
|
||||
## Writing criteria
|
||||
|
||||
- **One criterion per smallest meaningful unit.** Convert at Maximum Viable Atomicity:
|
||||
each criterion covers one requirement that can be judged on its own. Do not chop a
|
||||
requirement into fragments that cannot be judged alone, and do not bundle
|
||||
requirements that can pass or fail independently. Parallel facts derived the same
|
||||
way, such as the values of one calculated column, may share a criterion. Never group
|
||||
facts in a way designed to over-penalize a response.
|
||||
- **Phrase requirements positively.** Write "The response should ..." or "The response
|
||||
should avoid ..."; never write "should not". Factual criteria carry their answer key
|
||||
inline, in bold, so the criterion is judgeable without opening another document.
|
||||
- **Keep each criterion self-contained.** Never reference one criterion from another.
|
||||
A criterion may briefly restate a fact that also lives in `grader-context.md` so
|
||||
that it stands alone; that duplication is intended, and it is the one exception to
|
||||
the source's say-each-thing-once rule.
|
||||
- **Write conditionals as conditionals.** "If the response includes a migration, it
|
||||
should ...". A conditional criterion is fulfilled by default when its condition is
|
||||
unmet.
|
||||
- **Describe only the response.** Every criterion states a property of the response.
|
||||
Notes on how to verify a claim, which evidence to trust, or how to calibrate
|
||||
judgment fold into the `elaboration` of the criterion they support; they are never
|
||||
criteria of their own.
|
||||
- **Put judgment guidance in the elaboration.** State what fulfills the criterion and
|
||||
what fails it, with concrete examples from the source. Where several kinds of
|
||||
response are acceptable, list them. Where the source names behavior that must not
|
||||
trip the rule (the honest or flagged variant), carry that non-trigger into the
|
||||
elaboration.
|
||||
- **Give a strictly worse failure its own criterion.** Where the source ranks one
|
||||
failure clearly worse than a related one, encode the worse variant as a separate
|
||||
`dodged_bullet` that fails **in addition to** the base criterion, so a response
|
||||
committing the worse failure fails both and the score reflects the difference.
|
||||
- **Write criteria for likely failures.** A criterion earns its place by catching
|
||||
behavior responses actually get wrong. Skip trivial properties every response
|
||||
satisfies, and never penalize behavior outside the agent's control, such as a
|
||||
tooling failure.
|
||||
- **No numeric penalty language.** Severity and category carry the weight; the text
|
||||
never does. No "subtract 0.35", no points, no "out of 100", in guidelines or
|
||||
elaborations. Validation rejects numeric penalty phrasing.
|
||||
- **No generic scoring mechanics.** Flooring, how verdicts aggregate, and how
|
||||
penalties combine live in the shared grader prompt, never in a criterion.
|
||||
- **Preserve the source's facts exactly.** Keep every load-bearing fact, path and line
|
||||
citation, and code quotation, with markdown formatting (backticks, bold, fences)
|
||||
intact. Never invent facts, paths, or requirements the source does not carry.
|
||||
|
||||
The criteria count follows the source. Every scoring-relevant rule of the source
|
||||
lands in exactly one criterion's guideline or elaboration, and a rule is never dropped
|
||||
or folded away to reach a target count. Parallel facets of one requirement that the
|
||||
same evidence decides may share a criterion; distinct requirements get their own.
|
||||
Content that is context rather than a requirement belongs in `grader-context.md`, not
|
||||
in a criterion.
|
||||
|
||||
## Crux designation
|
||||
|
||||
`crux` is the top severity tier, reserved for the task's defining cliff. Derive it from
|
||||
the source's Heavy penalties section, and only from there.
|
||||
|
||||
- Write one Crux criterion per heavy penalty that targets **the overall score**,
|
||||
carrying that penalty's fire conditions and its stated non-triggers.
|
||||
- A heavy penalty that targets only a criterion of the standard, not the overall
|
||||
score, converts at `certain_dealbreaker`, not Crux.
|
||||
- When one penalty fires only on a conjunction (the response did A and also claimed
|
||||
B), write a single criterion covering the whole conjunction, phrased so it passes or
|
||||
fails outright; splitting it, or leaving room for partial fulfillment, lets partial
|
||||
credit dilute a dealbreaker.
|
||||
- When the source spells one dealbreaker out as several facets of the same failure,
|
||||
merge them into one Crux criterion; never write one Crux per facet.
|
||||
- A task carries **at most two** Crux criteria. Where the source has more
|
||||
overall-score penalties than that, keep Crux on the two that define the task's
|
||||
failure mode and convert the rest at `certain_dealbreaker`.
|
||||
- Designate Crux only from the source document. Never promote a criterion to Crux
|
||||
because runs that failed it happened to score low.
|
||||
|
||||
## Alignment with the holistic rubric
|
||||
|
||||
The two rubrics grade the same task, and their scores should agree. A run graded under
|
||||
the atomic rubric should land near the score the holistic rubric gives it, and runs
|
||||
should keep their relative order: a run the holistic rubric places far below another
|
||||
belongs far below it under the atomic rubric too. When atomic scores compress a gap
|
||||
the source creates, the missing lever is almost always Crux designation on the
|
||||
dealbreaker involved, not more criteria.
|
||||
|
||||
## Validate, stage, self-check
|
||||
|
||||
Run the two rubric detectors after generating the package, and again after any edit:
|
||||
|
||||
- `/detector-rubric-coverage` checks that every scoring-relevant rule of the source
|
||||
document lands in a criterion.
|
||||
- `/detector-rubric-form` checks that every criterion follows the form rules in this
|
||||
skill.
|
||||
|
||||
Fix what they flag before packaging the task; the package ships
|
||||
`tests/atomic-rubric.yaml` and `tests/grader-context.md` alongside the task's other
|
||||
files.
|
||||
|
||||
To grade under the atomic rubric inside the worker toolkit, stage the grading
|
||||
copies with `npx tsx scripts/stage-atomic-rubric.ts <task-slug>`. Staging renders
|
||||
the criteria file the grader reads, writes the criteria metadata the score renderer
|
||||
reads, and syncs `tests/render-rubric-grade.py` from `task-shared/`. Re-run it
|
||||
after every rubric edit. Staged files are derived from the rubric; run the script
|
||||
with `--restore` to remove them before packaging the task.
|
||||
|
||||
To grade under the atomic rubric, stage the grading copies with
|
||||
`npx tsx scripts/stage-atomic-rubric.ts <task-slug>` inside the devcontainer: staging
|
||||
checks the package's structure (a task key, a criteria list, a unique id plus a guideline
|
||||
and a category on every criterion, at most two Crux criteria), renders the criteria file
|
||||
the grader reads, and installs the rubric-aware harness. Staged files are working-tree
|
||||
only; never commit them. The `/detector-rubric-form` and `/detector-rubric-coverage`
|
||||
skills check the content rules (severity vocabulary, the numeric-penalty ban, coverage of
|
||||
the holistic rubric).
|
||||
|
||||
Reviewers working in a repo checkout also run
|
||||
`npx tsx scripts/validate-rubrics-cli.ts --slug <task-slug>`, which enforces the same
|
||||
schema, the Crux cap, and the numeric-penalty ban. That script is part
|
||||
of the review pipeline and does not ship in the toolkit.
|
||||
|
||||
## Related
|
||||
|
||||
- `.claude/skills/write-holistic-rubric/SKILL.md` — the source document this skill
|
||||
converts; its prose ground rules and penalty phrasing apply to the source, and its
|
||||
attribution rules decide which dimension a criterion targets.
|
||||
Reference in New Issue
Block a user