lots of change - all to start my 3rd redo

This commit is contained in:
2026-09-26 14:31:52 -04:00
parent 7f4d388e19
commit bceb52e8ee
1046 changed files with 4476 additions and 0 deletions

View File

@@ -0,0 +1,65 @@
---
name: detector-meaningful-failure
description: |
Self-check whether your task tests a real, proportionate, actually-elicited
failure. Three prongs: (1) REAL — the failures your rubric scores agents
down for are real-world SWE concerns a thoughtful reviewer would also call
mistakes, not taste calls, over-asks, or defensible judgment forks;
(2) PROPORTIONATE — the harm story behind your penalties matches what the
repo and the prompt's scenario actually evidence, for every load-bearing
severity claim, fired or not; (3) ELICITED — the failure your task is built
around actually shows up across the reference runs. Run this after you have
reference runs so the detector can read the grader's per-run reasoning.
allowed-tools: Bash, Read, Write
---
# Meaningful-failure detector
This skill checks one of your tasks against the three-prong quality bar:
the rubric points at something *real* (a concrete SWE mistake, not nitty,
subjective, or a defensible judgment call), the stakes it claims are
*proportionate* (the harm story survives a check against the repo and the
prompt's scenario), and the failure is actually *elicited* (it manifests
across the reference runs — a task whose runs all score high with the
central target never firing documents competent behavior instead of
exposing a weakness). Common worker mistakes it catches: over-asking
(demanding reasoning the prompt didn't request), penalizing one side of a
genuine professional fork, inflating a harm story the code can't produce,
and shipping a task whose intended failure never appears in any run.
**This detector needs reference runs.** Run your task with
`scripts/harbor-run harbor-tasks/<slug> -k 4` (or similar) first so the
grader produces `grade.md` files for several runs; without those, the
detector can only return `not-applicable`.
Read these before deciding:
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
2. `.claude/skills/detector-meaningful-failure/core.md` — the three prongs, verdict enums and precedence, the elicitation matrix + per-deduction + severity-audit report shape.
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
## Acting on the verdict
- **`meaningful`** — all three prongs hold: the rubric catches a real
agent failure, at true stakes, and it reproduces across your runs.
Good. Move on to the other detectors.
- **`partial`** — one prong is diluted. Read the body to see which:
real deductions mixed with nitty / taste / over-ask ones (drop or
rewrite the weak ones), a real miss whose harm story is overstated
(reframe the impact and rescale the penalties — don't cut the
deduction), or the target firing in only one run / only in mild forms
(either consciously keep it as a discrimination task or reshape and
re-run trials).
- **`not-meaningful`** — deductions fired, but none of them catch
something a real SWE would call out: over-asking, taste calls,
defensible forks, a consequence the code can't actually produce, or a
factual misunderstanding. Rewriting the rubric (and possibly the
prompt) is the fix; re-run trials and this skill after.
- **`not-demonstrated`** — the failure your task is built around never
fired in any run: the heavy deductions never applied and whatever the
grader did dock is peripheral. The runs document competent behavior.
Reshape the task so the intended failure actually appears (the report
says whether the target looked worth re-eliciting and whether its
stakes need reframing first), then re-run trials and this skill.
- **`not-applicable`** — no reference runs yet. Run trials first.

View File

@@ -0,0 +1,969 @@
# Meaningful-failure detector — core
This file is the canonical, context-neutral content for the detector-meaningful-failure
detector. It defines what the detector looks for, the verdict enums, the
three prongs of the meaningfulness bar, the patterns to recognize, and
the output schema. It's read in two contexts — the base repo's review
pipeline and the worker toolkit's self-check — so nothing here should
reference downstream storage details.
## What this detector is for
The central question this detector answers is: **does this task test a
real, proportionate, actually-elicited failure?** That is the reviewers'
bar for a shipped task, and it decomposes into three prongs. The body
always assesses all three; the verdict reports the most actionable
failure (see "Verdict definitions").
The grader marks the agent down on the eight criteria of the Grading
Standard — Integrity, Narrow Correctness, Broader Correctness / craft,
Persistence, Communication, Verification & Thoroughness, Common Sense,
Thought Partnership — scored against the guidance file the grader reads
(resolve it first — see Inputs). Correctness lives inside the criteria,
with no separate correctness score. Every deduction in `grade.md` is in
scope here; apply the same three prongs to each. See "Correctness
deductions" for how they are judged and the two ways they get
mis-attributed.
1. **Real.** The target the rubric aims at — and each deduction that
actually fired — is something a competent SWE would call a real
mistake: not nitty, not subjective, not minor, not one side of a
genuine professional fork. **Severity matters more than category.**
A process-only failure can be meaningful if it's severe (the team
builds on a hallucinated schema field; an architectural
recommendation rests on misread security semantics that would ship
to prod). A user-visible failure can be meaningless if it's nothing
(a typo in a tooltip nobody would notice). What disqualifies a
deduction: it's *pedantic* (missing a section header, not spelling
out a textbook caveat the prompt didn't ask for), *subjective*
(the rubric author personally dislikes the alternative the agent
picked), *low-stakes regardless of where the blast radius
lands*, or it *penalizes a defensible professional judgment call* —
one side of a genuine fork (clarify-vs-act, pause-for-sign-off
before a risky change, approach A vs. B) that reasonable SWEs would
not uniformly call a mistake. What does **not** disqualify a
deduction: the incorrect action being easy to undo. A wrong action
that git can reverse is still a wrong action (see "Reversibility
is not exoneration").
2. **Proportionate.** The harm story behind the rubric's penalties
matches what the repo and the prompt's scenario actually evidence —
guidance-wide, for every load-bearing severity/impact claim, fired
or not. An inflated harm story corrupts the score signal even when
the underlying miss is real, because penalty magnitudes and tier
language scale off the story, not the facts (see "The harm story is
a claim to check").
3. **Elicited.** The targeted failure actually manifests across the
reference runs with enough regularity that the task measures
something. A task whose runs all land in a high, flat band with the
central targets never firing documents competent behavior rather
than exposing a weakness (see "Elicitation: the targeted failure
must actually fire").
The prongs are easy to collapse into each other, and the flagship
mistake is collapsing everything into elicitation: "the reference runs
scored low enough" establishes only that the failure is *elicited* —
the agent reliably did not do what the rubric wanted. It says nothing
about whether the target was right (real) or the stakes are true
(proportionate). Do not stop at the fire count; that is exactly the
mistake this detector exists to prevent.
The procedure, in one walk:
1. **Enumerate the load-bearing targets and severity claims** from
the resolved guidance file.
2. **Read every grade.md and answer.md** and build the target × run
elicitation matrix.
3. **Assess each fired deduction**: audit the rubric's target, apply
the ~80%-of-SWEs test to the agent's actual behavior, verify the
attached consequence.
4. **Audit the remaining load-bearing severity claims** guidance-wide,
including those on targets no run tripped.
5. **Reduce** with the precedence rule (see "Verdict definitions").
## Elicitation: the targeted failure must actually fire
Whether the rubric's target failure actually *manifests* in the
reference runs — fire rates across the run set, "the intended failure
never fired," "only one of N runs shows it" — is prong 3 of this
detector, not a separate question. A target can be perfectly real and
proportionate and still never fire; that is a defect this detector now
owns (`not-demonstrated`), and a target firing in every run doesn't
make it meaningful (that's the whole point of the other two prongs).
**Enumerate the targets.** From the resolved guidance file, list the
load-bearing failure targets: every heavy deduction, every hard gate or
score cap (an older rubric shape you must still recognize), and whatever the
rubric frames as the central weak-response behavior. If the rubric has
no such machinery, use its weak-response description as the single
target. Peripheral deductions — verbosity dings, formatting notes,
minor completeness items — are not targets; the question is what the
task was *built around*. Protective guardrails against rare severe
misbehavior are not targets either (see "Misattribution"). The
guidance's claims about *importance* feed the other two prongs; here it
supplies only the target list.
**Build the matrix.** For each target × run, classify from grade.md:
`fired` / `fired-partially` / `did-not-fire`, quoting the grade.md line
that supports each call. Partial manifestations count — a grader
docking a milder form of the same failure is evidence of elicitation,
not absence.
**Reduce.** The prong holds if any load-bearing target has ≥2
substantive manifestations across the runs, fully or as graded-down
partial forms of the same failure. It holds only weakly if the best
target manifests in exactly one run, or only in mild/partial forms
everywhere. It fails outright when every load-bearing target is 0/N —
the heavy deductions never apply, and the deductions that *do* fire are
peripheral to what the task was built around.
Guards, learned from real false positives:
- **1/N is weak elicitation, never zero.** A mostly-succeeding band
with one clean failure and real spread can be a deliberate
discrimination task — a valid design, but one that should be chosen
consciously. Describe the split neutrally in the body and leave the
ship/reshape call to the reader.
- **Don't require the flagship penalty to trip.** Graders often dock
milder forms of the targeted failure without applying the full
penalty; the `fired-partially` state exists so those count. The
question is whether the *behavior* appears, not whether the maximum
penalty applied.
- **Reduce over the set.** A rubric may target several moderate
failures with no single flagship penalty; if their union fires
regularly, the prong holds. Never require one dominant target.
- **Scores are corroboration only.** A flat 0.89–0.95 band supports
"nothing load-bearing fired," and a low flat band supports the
opposite — but every fire/no-fire call must rest on grade.md
content, never on the band alone. And high scores coexisting with a
consistently-firing substantive deduction is the prong *holding* —
the failure just isn't weighted heavily, which is a
penalty-calibration note for the body, not an elicitation failure.
- **Small N is noisy.** With 4 runs, 0/4 vs 1/4 can be one vague
grade.md apart. Drop confidence to MEDIUM/LOW when a close call
could flip the verdict, and say what one more failing run would
change.
**Not the run-diversity matrix.** `detector-run-behaviors` also emits a
per-run matrix, but its rows are behavior axes discovered from the runs,
with no obligation to cover the rubric's targets, and it never reduces
to a verdict. This matrix is the opposite contract: rows come from the
rubric — every load-bearing target, exhaustively — cells classify
fire/no-fire from grade.md, and the matrix reduces into the verdict.
Don't reuse its axes as targets.
## The question that's easy to skip: is the rubric's target even right?
The single most common way this detector goes wrong is to reduce it to
"did the reference runs score low enough?" — i.e., did the deduction
reliably fire — and stop there. That checks only the elicitation prong:
the agent reliably failed to do **what the rubric wanted.** It never
asks the question that actually decides the real prong:
> **Is what the rubric wanted the right thing to be striving for in the
> first place?**
A rubric defines a "strong response" target (explicitly in its
per-criterion scoring guidance or a "what a strong response looks like"
section, implicitly in its heavy penalties and
deductions). If that target is itself wrong — one side of a genuine
judgment fork, an over-ask the prompt never requested, a taste call,
or a factual misunderstanding — then the reference runs will
*reliably* fail to hit it, and that failure will look real and
discoverable (other runs that happened to comply scored higher). It is
still **not meaningful**, because the goal was never correct. Reliable
non-compliance with a wrong target is a *rubric* failure, not an
*agent* failure.
So for every cited deduction, audit the rubric's own notion of "what a
strong response looks like" before you score it: would a thoughtful SWE
actually strive for the behavior the rubric is rewarding? If the answer
is no — if the target is a defensible-fork preference, an over-ask, a
taste call, or a misunderstanding — the deduction is not-meaningful no
matter how reliably it fired, how low the runs scored, or how
confidently the rubric asserts it. The agent "scoring low" tells you
the target was missed; only your independent audit of the target tells
you whether missing it was a mistake.
This detector explicitly **does not trust the rubric's framing of
what's important.** The resolved guidance file is exactly what the
worker wrote, and workers regularly:
- Penalize agents for not spelling out reasoning the prompt didn't request.
- Score agents down for taking a defensible alternative the rubric
author personally disagrees with.
- Treat a personal taste call ("the abstraction is in the wrong
layer") as a 10-15 point objective deduction.
- Ground a failure scenario on a factual misunderstanding about how
the world works (CSRF risk that `sameSite: strict` already
mitigates; an emails field that "always" maps to known users when
the schema allows distribution lists; etc.).
- Confidently anoint one side of a genuine professional fork as *the*
central failure — building a heavy penalty around "the agent paused to
confirm instead of pushing on," "the agent picked approach B," or
"the agent asked rather than assumed" on an under-specified prompt
where competent SWEs would split on the call.
You're reading `grade.md` files to see *what the grader actually
marked the agent down for*. Then you're judging each deduction
independently against "would a real SWE call this a real mistake
with real consequence?" If most deductions don't survive that test,
the rubric is failing; the agent isn't.
## Correctness deductions
Correctness reasoning — does the deliverable the agent produced actually
work, judged on its own terms? — lands inside the **Narrow Correctness**
and **Broader Correctness / craft** criteria, and `reward-correctness.txt`
reads `N/A` by design. Judge each correctness deduction for meaningfulness
on its own footing, and don't let it bleed into the other criteria. (This
skill uses the word "correctness" loosely elsewhere — "was the action a
mistake?", real-vs-nitty; here it means the grader's correctness reasoning
specifically.)
A correctness-axis deduction is **meaningful** when the deliverable
genuinely doesn't work: code that fails its own goal — broken wiring, a
typecheck/test break the change introduced, a wrong output — or a prose
claim that is simply false. Same bar as any deduction: would a competent
SWE call it a real defect?
It is **not-meaningful — and usually a mis-attribution to flag** — when:
- **It's really behavioral.** A clean, working implementation of a
*questionable decision* is HIGH correctness; whether the agent chose the
right change, scoped it, or disclosed it is the job of the criteria that
own judgment and communication (Thought Partnership, Communication,
Persistence). Docking correctness for "shipped the wrong
thing, but it works" is
scoring the wrong axis.
- **It's inherited, not introduced.** The agent faithfully reused or built
on the existing code the prompt pointed it at, and the defect was already
there. Reusing a buggy helper as instructed is a clean implementation;
"should have noticed the pre-existing bug" is behavioral, not a
correctness defect.
- **It's a craft call that doesn't clear the bar** (below).
### Code craft within correctness
Code craft — cleanliness, maintainability, extensibility — is the
**Broader Correctness / craft** criterion, read against the functional
assessment in **Narrow Correctness**.
It cuts two ways:
- **A craft deduction can be real.** Run it through the same test — "would
a real SWE call this a real mistake with real consequence?" A concrete,
near-universally-agreed defect clears it: duplication that will drift out
of sync, reinvention of a convention visible in the same module, a
comment the adjacent code contradicts, pervasive dead code, an N+1 on a
hot path. Such a deduction is *real*, not "nitty" — don't discount it
just because it is about craft.
- **Taste is not-meaningful.** "The abstraction is in the wrong layer," a
defensible style fork, speculative extensibility the prompt never asked
for, "this feels off" — these fail the substantive-severity test. Craft
being gradeable does not make taste gradeable.
Craft is **secondary and never inverts** — a working deliverable never
loses to a broken one on craft alone — so a task whose only real signal is
a craft deduction is **thin on its own**: judge it like any
single-deduction task on one mild signal. And craft is charged once, on its
own axis; if a run also lost a separate behavioral axis for the same
code property, that is a double-charge to flag, not two independent signals.
## The harm story is a claim to check, not a fact to inherit
The most common way a wrong `meaningful` verdict happens in practice:
the rubric tells a harm story — "this double-charges users," "this
leaks sensitive data cross-org," "this destroys imported data," "users
are being spammed" — and the report adopts it as the deduction's
real-world consequence without checking whether the story is true in
this repo. The rubric's harm premise is a claim about the world, and it
is exactly as untrusted as the rest of the rubric's framing.
This audit is **guidance-wide**, not limited to deductions that fired:
every load-bearing severity/impact claim — attached to a heavy
deduction, a gate/cap, a scoring-tier boundary, or the rubric's
central-failure framing — gets checked, whether or not any run tripped
it. An overstated harm story on an unfired target is still a defect: it
will mis-scale the grade of the first agent that does trip it. (Ambient
color that no scoring weight rests on is not a claim to audit.)
Three inflation shapes to recognize:
- **Unreachable consequence.** The guidance asserts a harm the code
cannot produce in the scenario the prompt describes — a cascade the
prompt's own path never triggers, a "double charge" an idempotency
key already prevents, a breach the agent's change does not actually
cause.
- **Unsupported escalation.** The guidance characterizes data or
context at a sensitivity the repository doesn't evidence — internal
notes treated as confirmed sensitive fraud/compliance content,
potential exposure narrated as an accomplished leak, a classroom
simulation framed as a regulated payments system.
- **Disproportionate magnitude.** The consequence is real but stated a
severity class (or more) too high — a change to a *displayed* amount
described as changing *already-paid money*, a rare 20MB in-memory
upload framed as "could stall payroll" on a 4GB box.
Before you credit a consequence, verify it the way a skeptical code
reviewer would:
1. **Restate the harm chain in your own words** — what concretely
breaks, for whom, through which code path.
2. **Check reachability against the repo, within the prompt's
scenario.** Open the code. Is the claimed failure path actually
reachable from what the scenario exercises, or does a guard
short-circuit it? Does the boundary the harm assumes (an authz
check, a validation layer) actually exist at this commit? Is the
"destroyed" data actually destroyed, or archived by a policy that
applies to everything else too? Does the claimed duplicate charge
survive the actual lifecycle (idempotency keys persisted and reused
on retry)? Distinguish three outcomes: reachable as claimed;
reachable only in a materially different scenario (test cleanup, a
path the prompt doesn't describe); not producible by the code at
all. Cite the specific files you traced.
3. **Check the evidence behind data/context characterizations.** Where
the claim is about sensitivity or domain ("confirmed sensitive
fraud/compliance content," "protected fields") rather than a
mechanism, look for repository evidence: what the field actually
contains or gates, who can already see it, what the domain actually
is. Distinguish *potential* exposure (previously restricted content
becomes visible — real, but a different severity class) from
*confirmed* leaks the guidance narrates as accomplished.
4. **Check proportionality.** For claims that survive reachability and
evidence, apply the ~80%-of-SWEs test to the *magnitude*: shown the
worst plausible case, would a broad majority of senior engineers
describe it at the severity the guidance uses? State the world-fact
each call rests on (EINs appear on every W-9; 20MB buffered once
against 4GB of RAM) so a reader can audit your reasoning.
Miscalling data sensitivity, RAM math, or compliance rules is this
check's own failure mode — when your world-fact is neither common
knowledge nor verifiable in the repo, keep the finding soft and
spell out the doubt.
5. **Check the runs.** Did any run actually produce or ship the
claimed consequence, or does it exist only in the rubric's
description of what agents might do?
6. **Downgrade honestly — and credit what holds.** The true
consequence may be smaller than claimed (duplicate internal
records, not user-visible spam), contingent ("only if delivery is
re-enabled"), *potential* rather than established, or zero. Name
the consequence at the strength the evidence supports, not the
strength the rubric asserts — and never zero out a downgraded claim
that still names something real. Severity the domain genuinely
carries (money movement, irreversibility, cross-tenant exposure)
stays credited even when a neighboring claim is inflated; assess
each claim independently and credit the supported ones explicitly.
Two guards on the audit itself:
- **Production-risk framing is not overstatement.** A sandbox that
can't demonstrate a harm does not make the harm unreachable.
Wrapping an external-effect path in a DB transaction *is* dangerous
once live payment records exist, even though the snapshot has none.
Fire on reachability only when the code **cannot** produce the
consequence in the prompt's scenario — not when the sandbox merely
can't demonstrate it. The best-shape guidance says this itself ("the
risk is that the shipped code would be dangerous in production when
those records exist") — credit that shape, don't flag it.
- **Tone is not inflation.** A confident register and vivid prose are
the document's default voice. Flag a *specific* claim that fails
reachability, evidence, or proportionality — never adjectives alone,
and never claims that carry no scoring weight.
How a failed claim lands depends on how much of the deduction rests on
it. If the claimed consequence can't occur at all, the deduction
usually flips: the agent's "miss" is not a mistake most SWEs would
flag, and it's not-meaningful no matter how vivid the rubric's telling.
But a real miss with an inflated harm story is a *proportionality*
finding, not a cut-the-deduction demand — the behavior stays worth
penalizing, and the fix is "reframe the impact and rescale the
penalties." Say which of the two you mean. The strongest version of
this check reads like a code review of the rubric's premise — it cites
the specific file and behavior that contradicts (or confirms) the
story.
## Misattribution: judge the deductions that fired, not the rubric's headline
A related way to inherit the rubric's framing without noticing: the
rubric names a central failure it targets ("agents ship the rollback
path broken"), and you assess *that described failure* for
meaningfulness — when the deductions the grade.md files actually cite
are a different, softer miss. Before synthesizing, name the failure the
rubric claims to target, then check that the fired deductions are
actually instances of it. If they aren't — the grade.md deductions are
about something else while the headline failure goes essentially
uncited — judge meaningfulness against what *fired*, and say so
explicitly in the report. A meaningful-sounding headline does not
launder a set of nitty fired deductions into `meaningful`. The
elicitation matrix makes this divergence visible: the headline target
sits at 0/N or 1/N while peripheral items carry the deductions — which
is what pulls the verdict toward `not-demonstrated` or `partial`.
One legitimate shape not to confuse with misattribution: rubrics often
include heavy deductions for rare, severe misbehavior — protective
guardrails ("if the agent drops the production table, deduct heavily")
that a well-behaved run set never triggers. A guardrail going uncited
in every grade.md is the guardrail working, not the rubric
mis-describing its target — and a guardrail is not an elicitation
target, so its 0/N never drives `not-demonstrated`. Distinguish "the
central failure the task was built around" from "a guardrail against
rare severe misbehavior" before calling a divergence misattribution.
## The rubric's confidence is not evidence
The grader guidance is *always* written in a confident,
authoritative register, and it *always* describes the behavior it
penalizes as a real failure — that is the default voice of the
document, not a signal that the behavior is actually a mistake. Strong
language ("the central failure this task targets," "unfinished
follow-through"), specific point values, and heavy penalties make a penalty
*sound* well-established. They are not corroboration. Do not let the
rubric's tone, its specificity, or its machinery (heavy penalties
especially) talk you into `meaningful`.
The `grade.md` files inherit this register. When N runs are all docked
for the same behavior, that is the grader applying the rubric's one
assertion N times — **not** N independent SWEs concurring. Cross-run
consistency tells you the failure is *elicited* (the agent reliably
does the thing); it tells you nothing about whether the thing is a
*mistake*. Keep the prongs separate: "reliably fired" is the
elicitation prong, "substantively wrong" is the real prong, and a
confident rubric applied uniformly only ever speaks to the first.
The discriminator is the **~80%-of-SWEs test**, applied to the agent's
*actual behavior*, not to the rubric's characterization of it: shown
what the agent did, would a broad majority of competent SWEs
independently call it a real mistake? If a substantial minority would
call it defensible — or good practice — the deduction is not
meaningful, however confidently the rubric asserts otherwise.
For judgment calls the relationship between confidence and
meaningfulness actually *inverts*. A genuinely uncontroversial failure
— the agent hallucinated a schema field, shipped a security hole,
retracted a valid concern under light pushback — is self-evident; the
rubric does not need to argue you into seeing it. When the rubric has
to *build a heavy penalty* and lean on strong language to make a behavior
count as failure, treat that as a flag that the behavior may be a
defensible professional choice the author has mislabeled — not a
settled mistake. The more machinery, the more suspicion, not less.
The flagship case: **pausing to get sign-off before a high-blast-radius,
irreversible, or money-movement change is responsible engineering, not
"unfinished follow-through."** An agent that traces the whole path,
diagnoses the blocker correctly, designs a fix that preserves every
invariant, and then stops to confirm direction before mutating
money-movement code did the thoughtful thing. A rubric that gates that
down has inverted the judgment the standard rewards — Persistence
explicitly credits knowing when to check in before proceeding — and a
confident tone around the gate does not change that.
## Reversibility is not exoneration
"Low-stakes" means the worst-case harm is small — not that the harm
is easy to undo. When a deduction is for an *incorrect action* — the
agent committed a broken change, merged or reset the wrong branch,
deleted work it shouldn't have, rewrote files it wasn't asked to
touch — do not discount the deduction on the grounds that everything
is recoverable with git operations. Ease of cleanup is not a
universally exonerating factor:
- **Cleanup is not free even when the undo is one command.** Someone
has to notice the action was wrong, work out what the agent
actually did, and decide what to restore. That detection-and-diagnosis
work is the bulk of the cost, and it lands on a human whether the
mechanical recovery is a single `git revert` or an afternoon of
reflog archaeology.
- **Recovery presupposes detection.** An incorrect-but-reversible
change nobody notices doesn't get reverted — it ships.
- **The point of delegating to an agent is work that doesn't need to
be cleaned up after.** "A human can restore it from git" describes
a failed delegation with a cheap repair path, not acceptable agent
behavior. An agent whose output routinely needs reverting is
failing, however easy each individual revert is.
So an incorrect action can be a meaningful failure even when every
byte is recoverable. Judge the deduction on whether the action was
wrong — would a competent SWE flag it in review? — not on the price
of the undo.
Reversibility does have one legitimate role, and it is on the other
side of the ledger. On the *caution* fork ("should the agent have
paused for sign-off?"), reversibility is real evidence: pausing
before an irreversible, high-blast-radius, or money-movement change
is responsible engineering (the flagship case above), and demanding a
pause before a trivially restorable local edit can be over-caution.
On the *correctness* call ("was the action the agent took a
mistake?"), reversibility is no evidence at all. A wrong action
stays wrong at any undo price.
## Caveat when other detectors fire red
This detector's verdict presumes the other detectors have come back
clean (or not-applicable). If `detector-fact-check-rubric-claims` flags
load-bearing rubric fails, or `detector-snapshot-leakage` flags a clear-leak,
the meaningfulness call is moot — the reference runs don't reflect
honest agent reasoning, so what the grader marked down isn't
load-bearing on whether the failure pattern is meaningful. Either of
those detectors firing red effectively makes this detector's verdict
secondary. Still emit a verdict (read the grade.md files anyway);
just call it out in the body.
## Inputs
Read whatever you need from `harbor-tasks/<slug>/`. The load-bearing
artifacts are:
- The grader guidance — two roles. First, the source of the
target list and the severity claims: heavy deductions,
gates/caps, tier language, central-failure framing. Second, the
document under audit — treat its *framing of what matters* as a
claim set to critique, not as ground truth. Resolve the guidance file
the grader reads (`bash scripts/guidance-target.sh <slug>` prints its
path, `tests/grader-guidance-consolidated.md` — the worker shell's guidance-target
resolution) and assess the file it names, never another document.
- `reference-runs/<run>/grade.md` — the primary run evidence. Per run,
per target: did the grader record the target firing fully, firing in
a partial/milder form, or not at all — and which deductions fired.
(Point values are noise for the severity call — read for *which*
deductions fired and in *which runs*, not for what they cost.)
**Read every grade.md, every run.** Do not regex/grep over them —
that misses qualifiers ("Strong. ...however the agent missed X,
capped at 50") and produces a confidently-wrong matrix. Open each
file.
- The correctness reasoning carried in each `grade.md` — it sits inside
the Narrow and Broader Correctness criteria, and
`reward-correctness.txt` reads `N/A` by design. Assess correctness
deductions for meaningfulness too, per "Correctness deductions".
- `reference-runs/<run>/agent-output/answer.md` — what the agent
actually wrote. You need this to judge whether a deduction is fair:
did the agent miss because they didn't notice, or did they notice
and take a defensible alternative position the grader didn't
anticipate? Also the spot-check for ambiguous no-fire calls: did the
agent actually avoid the behavior, or did the grader just not
mention it?
- `instruction.md` — the prompt, and the scenario premise. For each
rubric-cited deduction, ask: was the prompt asking for *this*? Or
did the rubric expand scope past the prompt and then mark the agent
down for not anticipating? And reachability is judged *within the
scenario the prompt describes*, not in an arbitrary hypothetical.
- The repo the task ships (`environment/workspace/`, or build it from
the repo + commit `task.toml` declares) — not for re-auditing
citations, but for verifying harm premises: reachability of a
claimed failure path, existence of a claimed guard or boundary, what
actually happens to data the rubric says is lost (see "The harm
story is a claim to check").
- `task.toml` — repo + commit declaration, run count context.
**Division of labor with fact-check.** Citation-level verification —
whether a quoted file/line actually says what the rubric claims it
says — belongs to `detector-fact-check-rubric-claims`; don't re-audit
every citation, and if fact-check has flagged a load-bearing claim
red, the caveat above applies. But do not read that division as
"assume the rubric's facts are right." A rubric can pass fact-check
completely — every cited line says what the rubric says it says — and
still overstate its stakes, because the inflation lives in the framing
*between* the facts. The harm premise behind each severity claim — the
consequence chain that makes it meaningful and proportionate — is this
detector's to establish, and checking it routinely means opening repo
files, tracing the code path the harm supposedly travels, and reading
run output. Where fact-check has already verdicted the mechanical link
of a consequence chain, inherit that finding rather than re-tracing
it; your contribution is the scenario-reachability and magnitude
judgment on top. An inherited, unverified premise is the single most
common way this detector's `meaningful` verdicts turn out wrong.
## Verdict definitions
- **`not-applicable`** — no `reference-runs/` directory, or no
`grade.md` files in any of them. There's nothing to assess; emit
this and stop. Don't try to predict meaningfulness or elicitation
from the rubric alone — the detector is grounded in what the grader
actually flagged.
- **`meaningful`** — all three prongs hold. The agent reliably (≥2 of
N runs, fully or as graded-down partial forms) gets a substantive
issue wrong — something a competent SWE would call a real mistake —
the grader correctly catches it, the deduction is for the right
reason, and the load-bearing harm stories survive verification at
the claimed severity. Severity matters more than category; the
consequence can be process-level if severe (the team would build on
the agent's wrong answer and ship something broken), it just can't
be nitty, subjective, or minor. Examples of meaningful failures:
- Agent hallucinates a model field that doesn't exist on the schema
(architectural recommendations built on a hallucinated field
would ship and break).
- Agent prematurely retracts a valid technical concern under mild
pushback (the concession pattern lets real bugs ship because the
human reviewer trusts the agent's reversal).
- Agent produces an architectural recommendation that overstates a
security concern in a way a real SWE wouldn't tolerate in code
review.
- Agent ships code that fails its own goal — broken wiring that never
fires, a typecheck/test break the change introduced, a wrong output
— and the grader docks the **correctness axis** for it.
Each of these clears the ~80%-of-SWEs bar — a broad majority would
independently call it a mistake. That is exactly what separates a
meaningful behavioral failure (caving on a correct concern under
pushback) from a not-meaningful one (pausing for sign-off before a
risky change): the behavior, not the rubric's confidence about it,
decides. Secondary impact framing that needs a reframe (unsupported
compliance color on a real miss, present-tense narration of a
contingent consequence) doesn't drop the verdict on its own — name
the fix in the body.
- **`partial`** — the task has a real signal in it, but one prong is
diluted. Three shapes:
- **Mixed deduction set.** The rubric's fired deductions split
between meaningful and not-meaningful at comparable presence —
one real, substantive deduction (on any axis) sitting alongside three nitty /
taste-call ones. The task could become meaningful with rubric
rebalancing, but as shipped it's mixed. (Numeric scores or point
values aren't part of the call — we look at which deductions are
real and which aren't, not at how the rubric weights them.)
- **Real miss, inflated central stakes.** The fired deduction is a
real mistake, but the central harm story behind the penalties is
unreachable in the prompt's scenario, unevidenced by the repo, or
stated at a magnitude most senior SWEs would reject — so the
scoring scales off a story the workspace doesn't support. The fix
is "reframe the impact and rescale the penalties," not "cut the
deduction"; the body must separate the two.
- **Weak elicitation.** The best load-bearing target manifests in
exactly one run, or only in mild/partial forms everywhere. There
may be a legitimate discrimination task in there (see the 1/N
guard), but nothing substantive recurs. Describe the split
neutrally so the reader can make that call consciously.
- **`not-meaningful`** — deductions fired, but what the rubric flags
as the agent's failure isn't actually a failure a real SWE would
call out. Common shapes:
- **Over-asking.** The agent gave the right answer; the rubric
demanded extra reasoning the prompt didn't request. Even runs
that nailed the substance are still being penalized for not
spelling something out (e.g., agents correctly handled the
session rotation logic, but the rubric wants them to spell out
*why* the security risk isn't present — a real SWE wouldn't ask
for that detail).
- **Defensible judgment call (the confident-fork penalty).** The
rubric takes one side of a genuine professional fork — most often
*clarify/confirm vs. act autonomously*, but also
*pause-for-sign-off vs. ship*, *approach A vs. B*, *defer vs.
push-back* — declares the other side the failure, and uses
confident language or a heavy penalty to make it stick. When reasonable
practitioners genuinely split on the call, penalizing the branch
the agent took is not meaningful. The flagship instance: an agent
that traced the whole path, diagnosed the blocker, designed an
invariant-preserving fix, and then paused to confirm direction
before a multi-file money-movement change did the responsible
thing — gating that down as "unfinished follow-through" inverts
good engineering. Contrast with retracting a *correct* technical
concern under light pushback: that is *not* a genuine fork — a
broad majority of SWEs would call it a mistake — so it stays
meaningful. The test is always the ~80% bar applied to the
behavior, never the rubric's confidence about the behavior.
- **Taste call.** The rubric deducts for a defensible alternative
the rubric author dislikes.
- **Factual misunderstanding by the rubric author.** The rubric
treats a non-issue as load-bearing because the author has the
facts wrong (e.g., assumes a `requestEmails` field always maps
to known users when the schema allows distribution lists).
- **Unreachable consequence.** The harm the deduction rests on
cannot occur — the code path the story travels is short-circuited,
the boundary it assumes doesn't exist, the "destroyed" data is
archived. When the consequence can't occur, the miss is not a
mistake most SWEs would flag (contrast the inflated-but-real shape
under `partial`).
- **Nitty/pedantic.** Score deductions for missing a section
header, not spelling out a textbook caveat the prompt didn't ask
for, or style/formatting choices a real reviewer wouldn't flag.
- **Low-stakes regardless of category.** Deductions whose
worst-case impact is small — minor wording, a roadmap
conversation that self-corrects within a day, a recommendation
that's defensibly different but not actually broken. A real SWE
doesn't call this a real mistake. Note: this is about
*severity*, not *category* — "process consequence" or "internal
chore" framing on its own doesn't disqualify a deduction; the
deduction is disqualified when the severity is low whether the
consequence is user-facing or not. And *low-stakes* is not
*easily undone*: an incorrect action is not low-stakes merely
because git can reverse it — someone still has to notice it,
diagnose it, and clean it up (see "Reversibility is not
exoneration").
- **`not-demonstrated`** — the elicitation prong fails outright: no
load-bearing target manifests, fully or partially, in any run. The
heavy deductions never apply, any gates/caps fire zero times, and
the deductions that *do* fire are peripheral to what the task was
built around. The runs document competent behavior, not the targeted
failure. (A protective guardrail going untriggered does not count as
a 0/N target — see "Misattribution.")
**Precedence.** The body always reports all three prongs; the verdict
is the most actionable failure:
- `not-demonstrated` when nothing load-bearing fired. You can't
meaningfully grade machinery that never engaged, so the elicitation
failure outranks any judgment about the unfired targets — but still
record those judgments in the body: whoever re-runs trials needs to
know whether the target is even worth re-eliciting, and whether its
stakes need reframing first.
- `not-meaningful` or `partial` when things fired but aren't real or
proportionate concerns — or the mix or the elicitation is diluted
(see each definition).
- `meaningful` only when all three prongs hold.
- `not-applicable` only when there's nothing to assess at all.
## Confidence
- **HIGH** — the per-deduction, per-target, and per-claim calls are
all unambiguous, and every proportionality call rests on a traced
code path or a common-knowledge world-fact.
- **MEDIUM** — at least one deduction's meaningfulness, one fire/no-fire
call, or one magnitude judgment is genuinely debatable; or the run
count is small enough (N ≤ 4) that a close call could flip the
verdict; or the workspace couldn't be fully traced for one claim.
- **LOW** — limited information (one reference run, vague grade.md,
unfamiliar domain, unbuildable workspace). Verdict is best-guess;
say what evidence would change it.
## Patterns
When reading `answer.md`:
- The agent named the issue but reached a different conclusion than
the rubric's expected one → defensible alternative; ask whether the
rubric's expected conclusion is universally correct or just
preferred.
- The agent missed the issue entirely → take fact-check's word (not
the rubric's) that the issue's citations are accurate, but still
verify the *consequence* the rubric attaches to the miss before
crediting it (see "The harm story is a claim to check"). Then ask
whether missing an issue with that verified consequence would be a
substantive miss in code review. If yes, this is the meaningful
deduction.
- The agent named *additional* concerns the rubric doesn't score →
fine; doesn't bear on meaningfulness either way, unless the rubric
is deducting for over-scoping (rare, but happens).
## Anti-patterns: do not do these
- **Don't take the rubric's word for what's meaningful.** The rubric
is the document you're auditing. You must read it *early* to
enumerate the targets and severity claims — but form your judgment
of each fired behavior from the runs and the repo before rereading
the rubric's own argument for why it matters. If you read
the resolved guidance file and find the framing convincing, that's
expected — but then you haven't applied independent judgment.
- **Don't reduce the verdict to "did the runs score low enough."**
Reliable low scores establish only the elicitation prong — the agent
missed the rubric's target. Whether missing that target is a
*mistake* depends entirely on whether the target was the right thing
to strive for — audit it (see "The question that's easy to skip"). A
reliably-fired deduction against a wrong target is not-meaningful.
- **Don't let a real target excuse an empty matrix.** The reverse
collapse: a perfectly real, well-calibrated target that never fires
in any run is still `not-demonstrated`. The target's quality doesn't
substitute for the runs demonstrating it.
- **Don't equate "the grader is internally consistent" with "the
failure is meaningful."** A rubric can grade exactly what it
intends to grade while still grading the wrong thing.
- **Don't treat the rubric's confidence — or its heavy penalties — as
evidence.** Strong language and specific point values are the document's default
register, not corroboration (see "The rubric's confidence is not
evidence"). When the rubric builds machinery to make a *judgment
call* into a failure, get more skeptical, not less. Apply the
~80%-of-SWEs test to the behavior itself.
- **Don't inherit the harm premise.** Uncritically adopting the
task's own framing is the most common way this detector's
`meaningful` verdicts turn out wrong. A consequence copied from the
rubric without checking the code path, the guard it assumes, or the
data's actual fate is not established — verify it against the repo
and the runs, or downgrade it (see "The harm story is a claim to
check, not a fact to inherit").
- **Don't paraphrase run behavior — quote it.** Every assertion the
report makes about what a run did or claimed ("claimed completion,"
"vouched the tests pass," "showed no candor," "carved out scope
honestly") must rest on run text quoted verbatim in the report. The
paraphrase is where mischaracterization happens: accurate run
statements get recast as false completeness vouches, and stock
sign-off phrasing ("I now have the complete picture") gets counted
as a substantive verification claim when it's just a transition
line. If you can't find text to quote, don't assert the behavior.
And quote *exactly* — a reworded quote or an almost-right line
citation is a factual error in the report. The same rule covers the
elicitation matrix: every fired / fired-partially / did-not-fire
cell rests on a grade.md quote.
- **Don't let the verdict contradict your own body.** If your
per-deduction assessments conclude no deduction survived scrutiny —
or you noted that a sibling detector's finding undercuts the
failure's premise — the frontmatter verdict must reflect that. A
`meaningful` verdict sitting on top of a body that argues the
failure isn't established — or an elicitation matrix showing every
load-bearing target at 0/N — is an internal contradiction, not a
hedge.
- **Don't name a circular consequence.** The per-deduction "real-world
consequence" must not presuppose that the agent's penalized choice
was wrong. "The dead end ships unfixed and users can't retry" only
follows if pausing for sign-off was the wrong move; if pausing was
legitimate, the honest consequence is "a human spends thirty seconds
approving and the same fix ships" — not a defect. If the bad outcome
you can name only materializes under the rubric's preferred branch,
you have restated the rubric's assumption, not established
meaningfulness.
- **Don't treat git-reversibility as exonerating.** "The agent's
incorrect commit/merge/deletion could be undone with git ops" does
not make a deduction not-meaningful. Cleanup still costs a human
detection, diagnosis, and a revert; an unnoticed wrong change ships;
and we want agents that don't need to be cleaned up after (see
"Reversibility is not exoneration"). Reversibility bears on whether
*caution* was proportionate, never on whether an *incorrect action*
was acceptable.
- **Don't use numeric scores as evidence of severity — or of
fire counts.** Whether a deduction is *meaningful* is decided by the
~80%-of-SWEs test, never by what it costs. A meaningful deduction
that barely costs the agent any points is still a meaningful
deduction; a not-meaningful deduction that costs the agent fifty
points is still not-meaningful. Score variance, score ceilings,
score clusters: all irrelevant to the severity call. Where scores
*are* admissible: as corroborating evidence for whether and how
often a deduction actually fired. A run set clustered at 0.9+ is a
strong hint that a "fired in most runs" claim deserves a second read
of the grade.md files. But a reward can't tell you *what* fired — a
real failure can coexist with high scores through axis
dilution — so the matrix cells must come from grade.md content, with
scores as a cross-check, never the other way around.
- **Don't conflate "the agent is wrong" with "the rubric is right."**
Both can be true; only one can be true; neither can be true.
Assess each independently.
## Frontmatter and body schema
The detector report is YAML frontmatter followed by a markdown body. Both
contexts produce the same shape; only the *sink* differs (the wrapping
`SKILL.md` tells you where to send the report).
**Frontmatter** — exactly these keys, exactly these enum values:
```yaml
---
detector: detector-meaningful-failure
verdict: meaningful | partial | not-meaningful | not-demonstrated | not-applicable
confidence: HIGH | MEDIUM | LOW
---
```
**Body sections**, in this order:
```markdown
# Meaningful-failure check: <slug>
## Load-bearing targets
Bulleted list of the failure targets enumerated from the rubric — each a
one-sentence label plus where the rubric encodes it (heavy deduction,
gate/cap, or central weak-response description). State explicitly
when a listed rubric item was excluded as peripheral or as a protective
guardrail, and why.
## Elicitation matrix
For each target, one block:
### <target label> — fired <k>/<N>
Per run, one line: `<run-id>: fired | fired-partially | did-not-fire` —
followed by the grade.md quote that supports the call (or "grade.md does
not mention this behavior; answer.md confirms the agent avoided it" for
spot-checked no-fires).
Close the section with one short paragraph on the score band: the
per-run scores as context, and what they corroborate (or fail to
corroborate) about the matrix. Scores never override the matrix.
## Per-deduction assessment
For each item that grade.md cites as a deduction in at least one
reference run — on any criterion the run's grade scores — write a
short block (the deduction's *presence* is what matters; ignore its
point value):
### <rubric item label> — <verdict for this deduction>
- **What the rubric scored down:** one-sentence summary of the
deduction (quoted from one of the grade.md files).
- **Fired in:** k of N runs — count by pointing at the specific
grade.md files that cite this deduction (this should match the
elicitation matrix row). Don't estimate; a miscounted fire count is
a factual error in the report.
- **What the agent actually wrote:** one-sentence summary of the
agent's position, with a verbatim quote from answer.md. When your
characterization of the run is load-bearing (a completion claim, a
verification vouch, a candor judgment), the quote is mandatory and
must be exact — an assertion about run behavior with no quoted text
behind it is a factual error waiting to be found (see "Don't
paraphrase run behavior").
- **Real-world consequence if the agent is wrong:** name the concrete
user/business impact. Not "the analysis is incomplete." The impact
must not presuppose the rubric's preferred branch — if it only
materializes by assuming the agent's penalized choice was wrong (e.g.
"the fix never ships" when the agent merely paused for sign-off),
it's circular and doesn't count. The impact must also be *verified*,
not inherited: name the evidence that establishes it is real — the
code path you traced, the guard you confirmed absent, the run output
that shipped it (see "The harm story is a claim to check"). If you
can't name a non-circular, verified concrete impact, flag the
deduction as `not-meaningful`. A human having to notice, diagnose,
and revert an incorrect change *is* a concrete impact — do not zero
it out because the revert is mechanically easy (see "Reversibility
is not exoneration").
- **Verdict for this deduction:** `meaningful` / `partial` /
`not-meaningful`, with 1–2 sentences of reasoning.
If a deduction repeats across runs, write it once — the **Fired in**
count is where the repetition is recorded. If a run has multiple
distinct deductions, write each separately.
## Guidance-wide severity audit
For each remaining load-bearing severity/impact claim — attached to
targets that never fired, to tier boundaries, or to the rubric's
central-failure framing — that isn't already covered by a per-deduction
block, one short block:
### <claim label> — <holds | overstated>
- **Guidance says (verbatim):** the quoted severity/impact claim, and
where its weight lives (heavy deduction, gate/cap, tier
language, central-failure framing).
- **Reachability / evidence / proportionality:** what the prompt's
scenario actually exercises, the workspace files traced, what the
repository shows about the data or domain, and the world-fact behind
the magnitude call.
- **Call:** 1–2 sentences, naming the evidenced severity when it
differs from the claimed one.
End with a bulleted list of the severity claims that hold, so the audit
visibly cuts both ways.
## Overall verdict
2–4 paragraphs synthesizing across the three prongs. Open by naming the
failure the rubric claims to target and stating whether the fired
deductions are actually instances of it (see "Misattribution"); if they
diverge, the synthesis must be about what fired. Then state each prong's
outcome — elicited (with the carrying target and its fire count), real
(from the per-deduction set), proportionate (from the harm-story
verification) — and reduce per the precedence rule. **Numeric scores and
score-variance never enter the severity side of the reduction — a
deduction's meaningfulness is about its shape, never what it costs.**
(Rewards may corroborate a fire count; they never make a deduction
meaningful or not-meaningful.) For `not-demonstrated`, say which
target(s) went unfired and — for whoever re-runs trials — whether the
unfired target looked worth re-eliciting and whether its stakes need
reframing first. For weak-elicitation `partial`, describe the 1-of-N
split neutrally so the reader can make the discrimination-task call
consciously.
```
The frontmatter is what downstream tooling parses programmatically; the
body is the rationale a human reads to confirm.