chore: init commit

in worker.../repo/GITFOLDER.zip is the .git folder.
This commit is contained in:
2026-08-11 14:44:09 -04:00
parent 0012380fd3
commit 2854619bc9
782 changed files with 65944 additions and 0 deletions

View File

@@ -0,0 +1,42 @@
---
name: detector-rubric-clarity
description: |
Self-check your grader guidance for whether it's well-written
enough for a grader to apply consistently. Catches material ambiguity in
scoring tiers, heavy penalties, and pass/fail criteria, plus typos, grammar
errors, and disfluent prose that interrupt the reader. Doesn't flag
every microscopic ambiguity or awkward sentence — only what would
actually affect grading or block use.
allowed-tools: Bash, Read, Write
---
# Rubric-clarity detector
This skill checks the prose of your grader guidance (the file
`bash scripts/guidance-target.sh <slug>` resolves) for two failure shapes:
material ambiguity in load-bearing wording (scoring tiers, heavy penalties,
pass/fail criteria that two reasonable graders could apply differently)
and copy-edit issues (typos, grammar errors, disfluent sentences) that
make the doc fail to read professionally.
Read these before deciding:
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
2. `.claude/skills/detector-rubric-clarity/core.md` — what counts as material ambiguity vs. copy-edit issues, verdict definitions, frontmatter/body schema.
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
## Acting on the verdict
- **`clear`** — the rubric reads professionally and the scoring criteria
pin down what a grader should look for. Good.
- **`minor-issues`** — small copy-edit nits worth polishing but no
load-bearing ambiguity. Look at the issues list in the report; tighten
them up. No need to rebuild the rubric.
- **`material-issues`** — at least one load-bearing scoring criterion is
ambiguous, OR the prose has enough errors that the doc doesn't read
professionally. Look at the "material ambiguities" section in the
report — those are the wordings to rewrite. Re-run this skill after
rewriting.
- **`not-applicable`** — the rubric is missing, empty, or template-only.
Write the rubric first, then come back to this skill.

View File

@@ -0,0 +1,208 @@
# Rubric-clarity detector — core
This file is the canonical, context-neutral content for the detector-rubric-clarity
detector. It defines what counts as material ambiguity vs. copy-edit
issues, the verdict enums, the patterns to recognize, and the output
schema. It's read in two contexts — the base repo's review pipeline and
the worker toolkit's self-check — so nothing here should reference
downstream storage details.
## What this detector is for
The rubric — the grader-guidance file the grader actually reads (see Inputs for how to resolve it) — is hand-authored prose that a grader reads at scoring time. Two things can go wrong with the prose itself, independent of whether the rubric's *substance* is right (other detectors cover that):
1. **Material ambiguity in load-bearing wording.** A scoring tier says "traces the flow accurately" — but "accurately" isn't defined. A heavy penalty says "if the agent dismisses the concern" — but what counts as "dismissing"? Two graders looking at the same answer can land in different tiers because the rubric's wording doesn't pin the criterion down.
2. **Copy-edit issues that interrupt the reader.** Typos, broken grammar, sentences that don't parse on first read, prose that's so disfluent the grader stalls trying to figure out what's meant. The rubric is a working document a grader has to use under time pressure; a doc that doesn't read professionally throws sand in the gears.
This detector flags both. It does **not** flag:
- Every minor wording quirk. Natural language is inherently ambiguous; pedantic interpretations of fully-readable sentences are noise.
- Stylistic preferences (passive voice, semicolon usage, oxford commas).
- Awkward but understandable phrasing where the meaning lands cleanly on the first read.
The operational test for ambiguity: **would two reasonable graders apply this scoring criterion differently because of the wording?** If yes, flag it. If no, leave it.
The operational test for copy-edit issues: **does the doc read professionally, or does the prose interrupt the reader?** A typo or two in body sentences with otherwise solid prose: tolerable. Multiple typos, broken sentences, or disfluent phrasing throughout: not tolerable.
## Inputs
Read whatever you need from `harbor-tasks/<slug>/`. The load-bearing artifacts are:
- The grader guidance — the primary input. Read every line. A task directory can carry two guidance files (`tests/grader-guidance-consolidated.md`, graded under the Consolidated Grading Standard, and the legacy `tests/grader-guidance.md`); resolve which one the grader actually reads (`bash scripts/guidance-target.sh <slug>` — the worker shell's guidance-target resolution) and assess that file, never its sibling. Judge the document against its own standard's structure (consolidated: task context, ground truth, one section per criterion; legacy: task and business context, strong/weak response descriptions, ground truth) — never flag it for not following the other standard's structure.
- `instruction.md` — secondary. Use to confirm that an ambiguity in the rubric matters because the prompt depends on the rubric's interpretation. (An ambiguity buried in background context that no scoring criterion touches isn't material.)
- `reference-runs/*/grade.md` — when present, read them. The operational test for ambiguity is "would two reasonable graders apply this differently?" — and the grade files are a record of graders actually applying this rubric. For each heavy deduction, tier boundary, and pass/fail rule, check whether the grades applied it the same way: did one grade apply a deduction that another skipped on similar behavior; did one N/A an axis that another scored; did grades read the same clause in incompatible ways? When they diverged, trace the divergence back to the specific sentence that permits both readings — that sentence is a material ambiguity, and the divergent grades are your evidence. Divergence alone isn't sufficient proof (graders are somewhat stochastic even on unambiguous rubrics), so always pair it with a concrete competing-readings analysis of the wording; but a sentence you'd have shrugged at in isolation becomes a confirmed problem when the grades demonstrably split on it.
You do not need to read the workspace or source repo — this detector judges the prose, not the substance. But never *assert* anything about how the runs were graded ("all five grades applied the deduction consistently") unless you actually read the grade files. If no reference runs exist, judge the prose on its own and say so in the body.
## Verdict definitions
- **`not-applicable`** — the resolved guidance file is missing, empty, or contains only the unmodified template content (the header scaffolding without scored issues, all-TODO stubs, the default that ships with the task harness). There's no prose to evaluate; emit this and stop. A doc that was clearly *authored* but arrives incomplete is not `not-applicable` — that's `material-issues` (see below).
- **`clear`** — the rubric reads professionally throughout. No material ambiguities in scoring criteria, heavy penalties, or pass/fail rules. A typo or two in body prose with otherwise solid sentences is fine — the bar is "reads professionally," not "is perfect." Two reasonable graders working from this rubric would apply each tier the same way.
- **`minor-issues`** — small copy-edit issues exist (a few typos scattered through, one or two awkward-but-understandable sentences) but the doc still reads professionally and there are no material ambiguities in load-bearing wording. The grader can apply each scoring criterion consistently; the reviewer should still get a heads-up so the worker can polish the prose.
- **`material-issues`** — at least one of:
- **Load-bearing ambiguity present.** A scoring-tier definition, heavy penalty, or pass/fail criterion uses wording that two reasonable graders would apply differently. The rubric needs the wording pinned down before it can grade consistently.
- **Doc no longer reads professionally.** Enough typos, grammar errors, or disfluent sentences that the prose interrupts the reader. The volume and shape of the issues add up to a doc that needs a copy-editing pass before it can be shipped.
- **Doc is structurally incomplete.** The file is truncated (ends mid-sentence or mid-code-block), or its own structure promises content that isn't there — a heading with nothing under it, a "see the heavy deductions below" pointing at a section that doesn't exist. A grader cannot apply a rubric that isn't all there. This is strictly about the doc's *own* promises going unfulfilled: a deliberately lean rubric that never promised more is fine, and whether a rubric defines "good" richly enough is a different detector's concern.
## Confidence
- **HIGH** — the call is unambiguous. Either the rubric is clearly clean, or the load-bearing ambiguity / copy-edit volume is plain to see.
- **MEDIUM** — at least one finding is genuinely a judgment call. A different reviewer might read the same sentence as clear enough.
- **LOW** — limited information (the rubric is very short, the prompt context is thin, or the criterion is in an unfamiliar domain). Verdict is best-guess.
## What counts as "material ambiguity"
The test isn't whether a word *looks* fuzzy — it's whether the rubric supplies enough privileged information (an answer key, a list of expected facts, the specific behaviors that count) for the grader to apply that word consistently. "Accurately traces" backed by an enumerated answer key is fine: the grader compares the answer to the key. "Accurately traces" with no key is not fine: the grader has nothing to check against. Trust the grader's judgment when ground truth is in the rubric; flag when it isn't.
Concretely, the patterns that gate scoring without supporting ground truth:
- **Judgment terms in scoring criteria, unsupported by ground truth.** "Thoroughly analyzes," "accurately traces," "appropriately balances," "substantially addresses." These are fine when the rubric has spelled out what the analysis must cover, what the trace looks like, or what a substantial answer includes (an answer key, a list of expected citations, the specific facts that mark each tier). They're material ambiguity when the rubric leans on the term to do the work and never spells out the standard — the grader has no way to apply it consistently.
- **Behavioral verbs in heavy penalties, unsupported by ground truth.** "If the agent dismisses the concern," "if the agent fails to acknowledge X," "if the agent overstates Y." These work when the rubric has named the specific shape of dismissing/acknowledging/overstating (example phrasings, the load-bearing concern by name with citations, what an overstatement of this risk would sound like). They're material when the rubric uses the verb to gate scoring without supplying examples or specifics — the grader has to guess where the line is.
- **All-or-nothing penalty triggers over in-between behavior.** A heavy penalty whose trigger reads as binary ("if the agent does not surface this gap") when real responses can land partway — the agent mentions the gap but mischaracterizes its consequence, or surfaces it wrapped in reassurance the user could miss. If the rubric doesn't say how the middle case scores, each grader improvises a partial penalty of their own size. Never describe a trigger as "mechanical" or "not a subjective call" without checking the runs (when present) for behavior that partially satisfies it — the unhandled middle case is usually sitting in the grades.
- **"Must mention X" where X is itself undefined.** A criterion like "must mention the race condition" works when "the race condition" has been concretely identified earlier (specific file/line, the mechanism). It doesn't work when the rubric introduces "the race condition" without first defining which race, which line, which mechanism — the grader can't tell whether a tangential mention satisfies the criterion.
- **Unclear pronoun referents in scoring-determining sentences.** "If the agent says this is fine, that's a B-tier response" — what is "this"? In a sentence that gates scoring, pronouns with multiple plausible antecedents make the call non-mechanical.
- **Tier descriptions that overlap.** A-tier and B-tier descriptions that share most of their language without naming the specific difference that distinguishes them. The grader can't tell which tier a borderline answer belongs in.
- **Conditional scope ambiguity.** "If A, then B unless C" sentences where the scope of "unless C" is unclear (does it modify B or the whole if-then?). Common in dense rubric prose.
- **Deduction arithmetic the grading model can't apply.** Each scored axis — a behavioral dimension under the legacy standard, a criterion under the consolidated standard — is scored 0.0–1.0, and the overall score is the mean of the non-N/A axes minus any heavy penalties the guidance directs at "the overall score" (each applied after the mean is computed, floored at 0.0 — the arithmetic the resolved standard's grader system prompt defines) — there is no 0–100 scale anywhere; magnitudes are fractions. Check every numeric score value against that model. A cap, deduction, or tier boundary written outside [0, 1] when used as a score value ("Confidence ≤ 20", "subtract roughly 45 points") is inert or ambiguous as written: one grader rescales by ÷100, another ignores the clause, a third guesses. An instruction to subtract from "the overall score" IS applyable — the grader subtracts it from the computed mean, and a penalty naming both a dimension and the overall applies in both places by design — so never flag overall-directed penalties as such; flag their *magnitudes* when they're off-scale, and check the stacking rules below. Mixed scales in one document (some clauses on 0–1, others on 0–100) force the grader to guess clause-by-clause.
- **Penalty machinery the grading model can't apply: hard gates, caps, and pins.** Dealbreakers belong in a rubric as heavy point deductions, not as hard gates, score caps, or pinned values ("hard gate: overall ≤ 0.3", "pin Confidence at 0.1"). A rubric built on gate/cap/pin machinery uses a shape the grading model does not support, leaving each grader to improvise a translation — flag it and suggest re-expressing each gate as a heavy deduction on the axes it concerns.
- **Deduction stacking ambiguity.** When a rubric attaches two effects to one defect (a deduction plus a floor, or two separately-stated deductions), it must say whether they're one penalty or two. Wording that can be read either way splits graders: some apply both halves, some drop one. (A single penalty naming both an axis and the overall score is not this — the grader system prompt defines that pairing: the axis subtraction attributes the failure, the overall subtraction applies after the mean.)
- **Overlapping deductions without a count-once rule.** Two separately-stated deductions that can both fire on the same single defect. Unless the rubric says which one applies — or that the second fires only when it represents a genuinely distinct miss — graders double-count inconsistently.
- **Asymmetric anchoring across tiers.** Failure outcomes carry concrete numbers while the strong outcome says only "very high" (or vice versa). Two graders can land 0.75 vs 0.95 on the same strong response because the strong end is unanchored.
- **Penalty magnitude that empirically destroys ordering.** A heavy deduction so large that, in the reference-run grades, post-penalty scores no longer order responses by quality — a run that was stronger before the penalty finishes below a weaker one, or every response floored to a single value with no pre-penalty spread carrying the discrimination. Flag this only with run evidence: penalty size itself is the author's design prerogative, and a big number is never a finding on taste alone. An all-runs-penalized band whose pre-penalty axis scores still discriminate is a task working as designed, not a finding.
- **Internal contradiction between sections.** One section permits or credits what another section deducts for or forbids — e.g., the calibration notes say an alternative load-bearing finding can clear the bar, while a later rule says holistic findings are not substitutes for the expected analysis. Two graders anchor on different halves of the contradiction and score the same response differently. Also confirm a deduction's stated value agrees everywhere it appears (including wherever `instruction.md` or the reference-run grades quote it).
The kinds of ambiguity that are **not** material:
- Judgment terms backed by ground truth. "Accurately," "thoroughly," "appropriately," and similar words are fine when the rubric has supplied the answer key, expected facts, or specific behaviors that let the grader recognize when the term applies. The grader is the wise human in the loop; we trust them to apply backed-up terms.
- Mild verbal hedging in background prose that doesn't gate scoring ("the codebase generally uses…", "this pattern is usually…").
- Genre-standard verbal shortcuts where the meaning is fixed by context (everyone knows what "a senior engineer would flag this" means in a rubric, even though "senior" isn't defined).
- Ambiguities in rubric prose that the scoring tiers don't depend on.
- Numbers greater than 1 that are not score values: counts ("misses 3 of the 4 call sites"), behavior thresholds ("if fewer than 80% of the tests pass"), line numbers, run counts, dollar amounts in the scenario. Only numbers that set or adjust an axis score (or the overall) get checked against the 0.0–1.0 scale.
## What counts as "copy-edit issues"
Things that interrupt the reader and make the doc fail to read professionally:
- **Typos.** Misspellings, transpositions, missing/extra letters. ("addtional", "transfter", "stipulates" when "stipulate" was meant.)
- **Grammar errors.** Subject-verb disagreement, wrong tense, mismatched plurality, broken constructions ("the agent provide" / "agents was").
- **Disfluent sentences.** Sentences that don't parse on first read, or read like they were transcribed mid-thought. Run-ons that combine three ideas without punctuation. Sentence fragments masquerading as full sentences.
- **Excessive verbatim repetition that adds noise.** A rubric is a formal pedantic working document, and parallel phrasing is often intentional — re-using the same construction across scoring tiers makes them easier to compare, and stable terminology helps the grader. Only flag repetition when the same sentence appears so often that the reader skims past it and the document would clearly read better with the boilerplate cut.
- **Sentence-level awkwardness that interrupts the reader.** Clunky constructions where the reader has to back up and re-read to figure out what's meant. (Mild awkwardness is fine — the bar is whether the reader stalls.)
- **Leftover toolkit-template content.** The grader-guidance file ships with a scaffold containing instructions like `<!-- REPLACE everything below this line with actual grader guidance. -->` and similar HTML-comment blocks. A finalized submission with that scaffold still present is shipping a doc that explicitly tells the grader the worker didn't finish — the scaffold itself says so. Treat trailing TODOs the same way: a "TODO: add scoring tier definitions" at the bottom of a finalized rubric is the worker telegraphing incompleteness.
Things that are **not** copy-edit issues worth flagging:
- Stylistic preferences (oxford commas, em-dash vs en-dash, passive voice, sentence length).
- Fully-readable sentences with mild clunkiness.
- Code-block formatting choices (backticks vs HTML `<code>`; bold via `**` vs `<b>`).
- The rubric's overall structure (sections, headings, length) — that's a different kind of issue.
## Verdict reduction in practice
Apply the operational tests:
1. **Did you find any material ambiguity in scoring-determining wording?** If yes → `material-issues`. Stop.
2. **Is the doc structurally incomplete — truncated, or missing sections its own structure promises?** If yes → `material-issues`. Stop.
3. **Did you find enough copy-edit issues that the doc no longer reads professionally?** If yes → `material-issues`. Stop.
4. **Did you find a few minor copy-edit nits but the doc still reads professionally?** → `minor-issues`.
5. **Is the rubric template/empty?** → `not-applicable`.
6. **Otherwise** → `clear`.
The threshold between `minor-issues` and `material-issues` on the copy-edit axis is a judgment call. Anchor on: a single typo in a long doc with otherwise tight prose is `minor`; a paragraph where every other sentence has a typo or grammar error is `material`. When in doubt, lean `minor-issues` for copy-edit-only findings — material-issues is for issues that actually block use, and we'd rather not cry wolf.
The presence of load-bearing ambiguity escalates straight to `material-issues` regardless of copy-edit state. A pristinely-typed rubric whose A+/A tiers can't be distinguished is still not usable.
Deduction-arithmetic findings follow the same logic, with one calibrated exception. A mismatch on a load-bearing deduction — mixed scales in one document, gate/cap/pin machinery, or grades showing the clause applied inconsistently — is `material-issues`. A single out-of-range number whose conversion is unambiguous in context (one "25" in a doc where every other value is correctly 0.xx) can go under `minor-issues` with MEDIUM confidence: graders sometimes rescale such a doc consistently, but the clause is still unapplyable as written and the worker should fix it.
Magnitude is never the materiality test for arithmetic divergence. When the grades reconcile a clause several different ways, do not talk yourself out of the finding because "the wobble is bounded to a few hundredths" or "no reading crosses a scoring boundary" — check the *orderings* instead. If competing readings can reorder responses (a run that was stronger before the penalty finishing below a weaker one), the ambiguity is material no matter how small each individual reading's effect looks: a ranking inversion is the most damaging grading failure a rubric can produce.
## Patterns to look for
When reading the resolved guidance file, walk it in this order:
1. **Scoring structure first.** A legacy doc defines tiers (A+ through D, or pass/fail); a consolidated doc defines a section per criterion, each with its own scoring guidance. Read the scoring bands back-to-back and ask: can I tell, from these descriptions alone, where a borderline answer would land? If two adjacent bands share most of their language without naming a specific distinguishing fact, that's material ambiguity. Apply the test to whatever scoring structure the resolved standard uses — a consolidated doc without a tier ladder, or a legacy doc without per-criterion sections, is following its own standard, not exhibiting an issue.
2. **Heavy deductions next.** "If the agent does X, subtract roughly N." Is "X" defined with enough specificity that a grader can mechanically check whether the agent did it? If "X" is "dismisses the concern" or "overstates the risk" without examples of what dismissing/overstating look like, that's material ambiguity. Then ask whether the trigger handles the middle case: responses that partially satisfy it (mention-but-mischaracterize, hedge-but-surface) — an all-or-nothing trigger over gradable behavior leaves the partial case to grader improvisation. Then check the *number*: is it on the 0.0–1.0 scale, and expressed as something the scoring model supports — a heavy deduction on named axis scores (dimensions or criteria, per the resolved standard) and/or the overall score (the grader system prompt defines overall-directed subtractions: applied after the axis mean, floored at 0.0), not legacy gate/cap/pin machinery? Finally check the deduction set as a whole for stacking and overlap: can two deductions be read as both firing on one defect?
3. **"What a good response says" / "What a bad response says" pairs.** Are the criteria in these sentences load-bearing for tier placement? If yes, apply the same ambiguity test. Vague criteria here propagate into the tier definitions.
4. **The document against itself.** With the tiers and deductions fresh, sweep for cross-section contradictions: does a section's closing rule match its lead sentence; does any tier bullet endorse behavior another section deducts for; do two sections give incompatible answers on whether one finding suffices; does every stated deduction value agree everywhere it's quoted? Internal contradiction is material ambiguity by definition — two graders anchor on different halves.
5. **The grades, when present.** Read `reference-runs/*/grade.md` and check each heavy deduction and tier boundary for consistent application across runs (see Inputs). Divergence that traces to a specific sentence upgrades that sentence from "arguably fine" to confirmed material ambiguity.
6. **Body prose throughout.** Skim for typos, broken grammar, and disfluent sentences. Group similar issues. Confirm the doc is structurally complete — it doesn't end mid-sentence or mid-code-block, and every section its own structure promises is present.
## Anti-patterns: do not do these
- **Don't flag every instance of natural-language ambiguity.** A rubric that says "the codebase generally uses Pulumi" doesn't need to define "generally." Body context isn't load-bearing; only scoring-determining wording is.
- **Don't list every typo individually.** Group by paragraph or section. Five typos in one paragraph is one finding, not five.
- **Don't flag stylistic preferences.** Passive voice, semicolons, em-dashes: none of these are copy-edit issues.
- **Don't critique the rubric's substance.** "This criterion is too lenient" or "this heavy deduction is calibrated wrong" are meaningfulness or fact-check concerns — they belong in those detectors, not here. The arithmetic check asks only "can the grader apply this number in the scoring model as written?", never "is this number well chosen?"
- **Don't convert grader stochasticity into findings.** Grades that differ in *score* while applying every criterion the same way are noise, not ambiguity. Only cite grade divergence when you can name the specific sentence whose competing readings produced it.
- **Don't paraphrase away a clause's conditions.** When the body characterizes a penalty or criterion — conditional vs. unconditional, scoped vs. blanket, one-shot vs. per-instance — quote the clause verbatim and keep its qualifiers. Describing a conditionally-applied penalty as unconditional is a factual error in the report, and reviewers check.
- **Don't propose major restructuring as a finding.** "The whole rubric should be reorganized" isn't a copy-edit issue — that's a separate concern. Stay scoped to wording-level issues.
- **Don't flag the document for not following the other standard's structure.** A consolidated doc has no tier ladder or strong/weak-response sections and a legacy doc has no per-criterion sections; each shape is that standard working as designed. Judge the resolved file against its own standard only.
- **Don't escalate `minor-issues` to `material-issues` for cosmetic reasons.** The verdict gates whether the worker should rewrite vs polish; `material-issues` should mean "rewrite needed," not "could be tightened."
## Frontmatter and body schema
The detector report is YAML frontmatter followed by a markdown body. Both
contexts produce the same shape; only the *sink* differs (the wrapping
`SKILL.md` tells you where to send the report).
**Frontmatter** — exactly these keys, exactly these enum values:
```yaml
---
detector: detector-rubric-clarity
verdict: clear | minor-issues | material-issues | not-applicable
confidence: HIGH | MEDIUM | LOW
---
```
**Body sections**, in this order:
```markdown
# Rubric-clarity check: <slug>
## Material ambiguities
For each load-bearing ambiguity (wording in a scoring tier, heavy penalty, or
pass/fail criterion that two reasonable graders could apply differently),
write a short block:
### <short label>
- **Where:** quote the verbatim sentence from the resolved guidance file and
name its location (scoring tier, heavy penalty, "what a good response says",
etc.).
- **Why it's ambiguous:** 1–2 sentences naming the specific competing
readings a grader could land on, and why those readings would produce
different scores.
- **Grade evidence (when reference runs exist):** if the reference-run
grades applied this criterion divergently, say how (which runs, which
readings). If they applied it consistently, you may say so — but only
after actually reading the grade files.
- **Suggested rewrite (optional):** one concrete phrasing that pins the
criterion down. Skip if the right rewrite depends on the rubric author's
intent and you can't infer it from context.
If there are no material ambiguities, write "None found." and move on.
## Copy-edit issues
A bulleted list of typos, grammar errors, and disfluent sentences. For each:
quote the verbatim phrase and (if not obvious) one-line correction. Group
similar issues — don't list ten typos one per line if they're scattered
through a single paragraph; cite the paragraph once.
If the doc reads professionally throughout, write "None found." Don't list
stylistic preferences (passive voice, semicolon usage) — only things that
are clearly errors or that interrupt the reader.
## Overall verdict
1–2 paragraphs synthesizing the above into the chosen verdict. Be explicit
about which of (material ambiguity / copy-edit volume) drove the call.
```
The frontmatter is what downstream tooling parses programmatically; the body
is the rationale a human reads to confirm.