Loaded up for the 3rd redo

Still on potion-voice
This commit is contained in:
2026-09-26 14:57:10 -04:00
parent bceb52e8ee
commit e55ccea018
216 changed files with 43127 additions and 41 deletions

View File

@@ -0,0 +1,95 @@
---
name: detector-broken-dev-env
description: |
Self-check whether your task presents an *incidentally* broken local dev
environment that the test agent has to awkwardly work around. The workspace
doesn't build/install/run, a dependency or service is missing, or the test
suite has pre-existing failures or flakes unrelated to your task — and the
agent burns effort coping with that instead of doing what your prompt asks.
These tasks are weak: the existing test suite is the main verifier, so env
noise lands straight in the grade. The one allowed shape is intentional
breakage — a task whose subject IS the broken env ("my dev env is broken, fix
it"). A pre-existing app bug that your prompt asks the agent to find or fix is
the task working, not breakage. Also checks that the workspace is actually in
the state your prompt (or snapshot) says it's in — a promised uncommitted
change that's already committed, a "build X" ask where X already ships, or
referenced data that isn't there is a premise mismatch even when everything
builds green. And checks that everything you package reflects the same
revision of your task — runs graded under an earlier prompt or rubric, a
reward.txt that no longer matches its grade.md, or a re-uploaded older
archive is package drift even when every artifact is individually healthy.
allowed-tools: Bash, Read, Write
---
# Broken-dev-env detector
This skill checks whether your task hands the agent a dev environment that is
broken for reasons unrelated to what you're asking it to do — a build that won't
run, a missing dependency, or a test suite with pre-existing failures the prompt
never mentions. If the agent has to fight that breakage to make progress, the
task is testing "can the agent cope with a broken env" instead of the behavior
you meant to grade, and the verifier signal gets noisy. It also checks that
your scored reference runs are valid samples of agent behavior — a run killed
mid-work by an API error, truncated, or missing its output snapshot reflects
infrastructure, not the agent, and shouldn't ship as evidence. And it checks
that the workspace matches your task's stated premise: if your prompt or
snapshot asserts something about the workspace ("review my uncommitted
change", "there's already data in the repo", a previous turn's fix) that the
shipped state contradicts, the agent responds to the workspace as it actually
is and your rubric grades a task that can't happen. Finally, it checks that
your package is one coherent revision of the task: if you polish the prompt or
rubric after generating runs, the shipped runs and grades must be regenerated
or regraded to match — a package whose parts describe different versions of
the task can't evidence it.
Read these before deciding:
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
2. `.claude/skills/detector-broken-dev-env/core.md` — intentional-vs-incidental, the three shapes breakage takes plus the runs-corrupted, premise-mismatch, and package-drift shapes, what counts as legitimate task difficulty, verdict enums.
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
## Acting on the verdict
- **`clean`** — your environment runs fine (or the only red tests are the bug
your prompt is about). Good.
- **`partial`** — there's some env friction, but it's minor or borderline. Read
the rationale; either smooth the setup so the agent never hits it, or confirm
it's cosmetic enough not to distort the run.
- **`incidental-breakage`** — the agent has to work around a broken setup your
prompt didn't ask it to fix. Fix the environment (repair the Dockerfile,
pin deps, remove the unrelated failing/flaky tests) so the agent starts from a
working baseline, then re-run this skill. Don't try to rescue it by reframing
the breakage as the task — see `intentional`.
- **`runs-corrupted`** — one or more of your scored reference runs was ended or
distorted by infrastructure rather than by the agent (an API/model error
mid-run, a truncated trajectory, a missing output snapshot, a verifier
timeout), so it isn't a valid sample of agent behavior. Your workspace may be
perfectly healthy. Re-run the affected trials, replace the corrupted runs,
repackage, then re-run this skill.
- **`premise-mismatch`** — the shipped workspace contradicts what your prompt
or snapshot asserts (the promised uncommitted change is already committed,
the feature you ask the agent to build already exists, referenced data is
absent, a prior turn's state was reset away), and your holistic rubric
assumes the premise holds. Either fix the workspace so the premise is true
(workspace.patch, seeds, snapshot end-state), or — if the false premise is
deliberate — make the rubric grade the agent on surfacing it, then re-run
this skill and re-collect reference runs.
- **`package-drift`** — your packaged artifacts don't all reflect the same
revision of the task: runs were graded under an earlier prompt or rubric, a
reward.txt no longer matches its grade.md, runs record conflicting task
versions, or you re-uploaded an older archive after making fixes. Nothing
may be broken — but the runs no longer demonstrate the shipped task. Apply
the smallest coherent fix: regenerate runs against the current prompt,
regrade against the current rubric (see `/regrade-reference-run`), re-copy
the runs so each reward.txt matches its grade.md, or rebuild and re-upload
the archive — then re-run this skill.
- **`intentional`** — your task is explicitly about fixing the environment. That
is a valid task; nothing to change. (Only legitimate if your *prompt* asks for
the repair — not if the agent merely ended up coping with a broken env.)
- **`not-applicable`** — the task has no runnable environment (pure analysis /
writing) AND the prompt/snapshot make no workspace-checkable assertions, or
there's no evidence yet (no reference runs and the rubric says
nothing about the env). Re-run once you have reference runs. (Runs that exist
but died on infrastructure are `runs-corrupted`, not this; a prose prompt
that asserts workspace state can still earn `premise-mismatch`.)

View File

@@ -0,0 +1,683 @@
# Broken-dev-env detector — core
This file is the canonical, context-neutral content for the detector-broken-dev-env
detector. It defines what the detector looks for, the verdict enums, the
patterns to recognize, and the output schema. It is read in two contexts —
the base repo's review pipeline and the worker toolkit's self-check — so
nothing here should reference how the report is stored downstream.
## What this detector is for
A task hands the agent a repo at a chosen commit plus a prompt, and the agent
works in a local dev environment (the task workspace). Sometimes that
environment is **broken in a way that has nothing to do with the prompt's
ask**: the workspace doesn't install or build out of the box, a binary or
dependency is missing, env vars aren't set, a service won't start, a migration
is wedged, or the test suite has pre-existing failures or flakes unrelated to
the task. The agent then burns effort diagnosing and working around that
breakage instead of (or on top of) doing the work the prompt actually asked
for.
We do not want tasks where the broken environment is **incidental** — an
accidental artifact of how the task was extracted, left in the workspace, and
silently presented to the agent as just one more hazard to fight through. It
makes the task noisy: the agent's score then partly reflects whether it could
push through a broken setup, not whether it did the intended work. The existing
test suite is the primary verifier for these tasks, so environment noise in the
build or the tests directly muddies the signal the task is supposed to produce.
There is one legitimate shape: **intentional** breakage. If the task is
*explicitly about* the broken environment — "my local dev env is broken, can
you fix it", "the test suite won't run, figure out why", "the build is red, get
it green" — then a broken environment is the deliberate subject of the task, not
a hazard. That is fine and must not be flagged.
This detector also owns an adjacent defect in the same "the grade reflects
infrastructure, not the work" family: **reference runs corrupted by the
execution infrastructure** rather than by anything the agent did. An API or
model error kills a run mid-implementation, a trajectory is truncated so the
grader scores a transcript the agent never produced, an output snapshot is
missing, a verifier times out — and the run is scored and shipped as if it
showed real agent behavior. The workspace can be perfectly healthy in every one
of these cases; see the dedicated shape section below.
And it owns a third defect in the same family where nothing is broken at all:
the shipped workspace **contradicts the premise the task states**. The prompt
(or a snapshot's prior turns) asserts something concrete about the workspace —
"review my uncommitted change," "take a first pass at building X," "there's
already data seeded in the repo," "in the previous turn you fixed Y" — and the
workspace the agent actually receives doesn't honor it: the promised diff is
already committed, the feature to build already ships complete, the referenced
data is absent, the prior turn's state was reset away. Everything installs and
tests green, yet the environment is wrong *for the task*: agents reasonably
respond to the workspace as it actually is, and the rubric — written as if the
premise held — either can't be applied or is applied unfairly. See the
dedicated shape section below.
The last defect in the family is **package drift**: the shipped artifacts
don't all reflect the same revision of the task. Workers iterate — polish the
prompt after generating runs, rewrite the rubric after grading, re-upload an
archive after feedback without rebuilding it — and each of those steps can
ship a package whose parts disagree about which version of the task they
belong to: runs graded under an earlier prompt (sometimes still scaffold
placeholder text), grades produced against an earlier rubric, an old archive
resubmitted wholesale. Every artifact can be individually healthy and the
package still fails to evidence its own task. See the dedicated shape section
below.
Taken together, this is the "is the submission package itself sound?" check —
the environment, the runs, the workspace-vs-premise fit, and the version
coherence of the shipped artifacts. This detector decides: does *this*
submission present an incidentally broken dev environment that the agent has
to awkwardly work around — a reference-run set corrupted by infrastructure
rather than agent behavior — a workspace that contradicts the premise the
prompt or snapshot asserts — or a package whose artifacts ship from different
revisions of the task?
## Inputs
Read whatever you need from `harbor-tasks/<slug>/`. The load-bearing artifacts:
- `instruction.md` — the prompt the agent received. **This is what decides
intentional vs. incidental.** Does the prompt ask the agent to diagnose, fix,
or repair the environment / build / dependencies / tooling / failing setup? If
yes, breakage is the subject (intentional). If the prompt asks for something
else entirely (implement feature X, audit module Y, write design doc Z) and
the environment is nonetheless broken, the breakage is incidental.
- `reference-runs/<run>/agent-output/answer.md` and the run's trajectory — the
test agent's actual behavior. Sample 2–3 runs. Look for the agent spending
turns getting to a runnable baseline: install/build failures, missing
binaries, "the tests won't run so I…", patching config unrelated to the task,
re-running with workarounds, or prose in the answer noting that the
environment was broken. This is the strongest evidence that the breakage
actually distorted the run. Also check each run's **terminal state** (the
last few events of the trajectory, whether `agent-output/` exists and is
non-empty, and whether `grade.md` itself notices an abrupt ending) — that is
what decides `runs-corrupted`, per the shape section below.
- Per run, the version-coherence artifacts: `grade.md` (the scoring structure
the grader actually applied), `reward.txt` and `reward-correctness.txt`
(the recorded scores; under the Grading Standard
`reward-correctness.txt` legitimately reads `N/A`),
`result.json` / `config.json` when present (recorded task name/checksum),
and any transcript/session artifact that records the prompt the agent
actually received. These are what the `package-drift` shape joins against
the shipped `instruction.md` and the resolved guidance file — see the shape
section below.
- `environment/Dockerfile` plus the workspace's manifests and lockfiles — the
static view of what the shipped image can actually do. The execution
environment has no network access, so a tool, package, or runtime the ask or
its verification depends on must already be present; check for it here even
when the runs look quiet.
- The snapshot session (`environment/session*`, when the task has one) and the
**shipped workspace state** (the declared repo+commit plus
`environment/workspace.patch`) — the two halves of the premise check.
The prompt and the snapshot's prior turns are the source of
workspace-checkable assertions; the built workspace (git status/diff, file
and branch existence, seed contents, whether a named feature or fix is
present) is the ground truth they are checked against. See the
`premise-mismatch` shape section below.
- The grader guidance — the rubric. Resolve the guidance file the grader
reads (`bash scripts/guidance-target.sh <slug>` prints its path,
`tests/grader-guidance-consolidated.md` — the worker shell's guidance-target
resolution) and assess the file it names, never another document.
Sometimes the rubric itself reveals
the environment is broken: "note that the suite has a pre-existing failure in
X, ignore it", "the dev server doesn't start; a strong agent works around
it", "don't penalize the agent for the broken migration." A rubric that treats
env breakage as an obstacle course the agent must navigate (rather than the
thing to fix) is a strong incidental-breakage signal.
- `task.toml` — for the source repo and commit, when you need to confirm whether
a failure the agent hit is pre-existing in the workspace vs. introduced by the
agent.
## Incidental vs. legitimate task difficulty
The hard part of this detector is not mistaking the task working as intended for
incidental breakage. Keep these straight:
- **The task's own bug or failing test is not breakage.** If the prompt is "fix
the failing `X` test" or "the agent's change should make the suite pass", then
a red suite at the start is the *subject* of the task. Breakage only counts as
incidental when it is **unrelated to the prompt's ask**.
- **TDD is not breakage.** An agent writing code and watching tests go red→green
as it works is the loop functioning, not a broken environment.
- **A pre-existing failing test unrelated to the task is incidental.** The agent
can't trust the suite's signal and has to reason about which failures are
"expected" — noise the prompt never asked it to deal with.
- **Setup the agent must repair just to reach a working baseline is incidental**
when the prompt didn't ask for it. Pinning a dependency version, recreating a
missing file, or hand-fixing a config to get install/build/run to succeed —
all unrelated to the actual deliverable — is the classic shape.
## Three shapes incidental breakage takes
Any one of them establishes that incidental breakage *exists* — but presence
alone earns at most `partial`. Escalate to `incidental-breakage` only when the
breakage **materially distorted the task's signal**: it blocked or aborted a
run, consumed a substantial share of the agent's effort (well beyond confirming
a known-unrelated failure and moving on — agents spending a modest slice of a
run establishing the baseline and then proceeding unimpeded is `partial`
territory), or plausibly changed what the grader saw or the score. Genuine
breakage that the agents note, route around, and that leaves no trace in the
grade stays `partial` — worth fixing, not task-disqualifying.
**Shape 1 — setup/build breakage before the work can start.** The workspace
doesn't install, build, or start out of the box for reasons unrelated to the
task. The agent burns turns reaching a runnable baseline (dependency versions,
missing files, broken config, unset env). The prompt never asked for any of it.
Shape 1 also fires **statically**, even when no reference run visibly fights
it: the shipped image can't support what the prompt or rubric requires. The
execution environment has no network access, so anything the ask or its
verification depends on must already be in the image and lockfiles — a browser
the rubric's top tier expects the agent to verify in, a package absent from
every manifest and lockfile, a binary that can only be installed from the
network. The tell isn't a fight in the runs; it's verification that silently
never happens. Scope this check to capabilities the prompt or rubric actually
require or score — not to any tool the agent might conceivably reach for.
**Shape 2 — pre-existing failing or flaky tests the agent must navigate.** The
test suite has failures or flakes unrelated to the task. The agent can't trust
green/red, has to retry or guess which failures are "expected," and the
verifier's own signal is polluted. This is the most corrosive shape because the
existing suite is the task's primary verifier — noise here lands straight in the
grade.
**Shape 3 — broken tooling framed as a hazard in the rubric.** The
grader-guidance explicitly tells the grader the environment is broken and that
the agent should (or shouldn't) be distracted by it — "ignore the failing lint
step", "the server won't start; a strong agent works around it." The rubric
treats env breakage as an obstacle course rather than the deliverable.
The shapes can co-occur; cite each one you see.
## A fourth shape — reference runs corrupted by infrastructure (`runs-corrupted`)
A submission's reference runs can be invalidated by the machinery *around* the
agent even when the workspace itself is perfectly healthy: an API or model
error kills a run mid-implementation, a headless plan-mode ending strands the
agent waiting for an approval that never comes, a trajectory is truncated
mid-tool-call so the grader scores a transcript the agent never produced, an
agent-output snapshot is missing or corrupt, or a verifier timeout is counted
as a scored run. The damage takes two forms: the run set no longer evidences
the task (a low score reflects the infrastructure, not the agent), and the
grader can actively mis-grade — e.g. a completion-honesty penalty fired on a
final message the truncation ate.
This is not env breakage, and the intentional-vs-incidental test doesn't apply
(no prompt makes an API error the subject). It earns its own verdict,
`runs-corrupted` — never `incidental-breakage`, which would misdescribe a
healthy workspace.
Per scored run, check the terminal state:
- Does the trajectory end with a complete final assistant message — or
mid-tool-call, on an unparseable tail, or with an infrastructure error string
(an API 4xx, an invalid-model error, a plan-mode exit that errored with no
subsequent agent turn) as the last event?
- Is `agent-output/` present and non-empty?
- Does `grade.md` itself notice the incompleteness ("the run ends abruptly",
"no final summary") — or, worse, score the truncated state as if it were the
agent's behavior?
The load-bearing boundary is whether the run reached a **gradable state before
the infrastructure event**. A run killed mid-investigation with a clean tree
and no answer never became a valid sample of agent behavior — that fires. An
error that only ate the closing summary *after* the fix, tests, and substance
had all landed leaves the run usable — note it in the body as `partial`-grade
noise, not corruption.
Guards against overfiring:
- **Match infrastructure signatures only in a run's terminal events**, never by
searching the whole transcript — agents quote error text while debugging, and
repos discuss API errors in prose.
- **A deliberate stop is legitimate behavior, not corruption.** An agent that
presents a plan or asks a question as its chosen ending — a shape the rubric
credits — ended naturally. The corruption case is the run trying to continue
and being unable to: the plan-mode exit returns an error, no agent turn
follows, and nothing ships.
- **Brevity is not truncation.** Truncation needs structural evidence — a last
event that is a tool call, an unparseable tail, or a missing final message
the grade itself trips over — not a stylistic judgment about a terse ending.
## A fifth shape — the workspace contradicts the task's premise (`premise-mismatch`)
Sometimes the environment installs, builds, and tests green — nothing is
"broken" in the workaround sense — yet the workspace is not in the state the
task *says* it is in. The prompt, and any prior snapshot turns, make concrete
assertions about the workspace, and the graded workspace either honors them or
it doesn't. When it doesn't, and the rubric was written assuming it does, the
task exercises something other than what it describes: runs sail past the
intended difficulty, improvise a different task than the one described, or get
penalized for reasonably responding to the environment as it actually is.
The recurring premise types, each checkable against the shipped workspace:
- **Pending-change** — the prompt promises uncommitted edits ("review my
uncommitted change", "the diff on my branch"), but the tree is clean and the
change is folded into an existing commit, so `git diff HEAD` is empty.
- **Absence** — the prompt asks the agent to "add" / "build" / "take a first
pass at" a capability that the workspace (including `workspace.patch`)
already ships substantially complete, so most of the prompt isn't actionable
as written.
- **Continuity** — the snapshot's prior turns leave the tree in a state (a fix
landed, a breakage present) that the graded workspace does not carry: the
checkout was reset or repaired between turns, so the agent replays history
that no longer matches the tree it is acting on.
- **Presence** — the prompt references load-bearing data or files ("there's
already history data in the repo", a named branch or config file) that the
shipped state doesn't have: zero seeded rows, no such file.
- **Reproducibility** — the incident the prompt reports cannot occur in the
shipped configuration: the symptom only manifests in a test double, or the
code path the described failure depends on isn't wired.
The decision procedure: **extract** every workspace-checkable assertion from
`instruction.md` and the snapshot session; **verify** each against the built
workspace (git status/diff for pending-change, code search and reading for
absence/presence, the snapshot's implied end-state vs. the shipped tree for
continuity, the configuration and code path for reproducibility); then
**classify** each failed premise against the resolved guidance file — does the
rubric assume the premise holds (grades content only reachable if it holds,
describes the task in the premise's terms), or does it know the true state and
credit the agent for surfacing the discrepancy?
That last question is the shape's carve-out, the analog of the
intentional/incidental test (which itself doesn't apply here — no workaround is
involved): **a deliberately false premise is a core, legitimate task design.**
Many good tasks hand the agent a wrong user belief on purpose and grade whether
the agent surfaces it. Never fire on "the premise is false" alone — fire only
when the rubric itself assumes the premise holds, or nowhere credits
discovering that it doesn't.
Guards against overfiring:
- **"Already exists" is a judgment call on partial implementations.** An ask to
add a capability when a half-wired helper exists may legitimately mean
"finish it." Treat an absence premise as violated only when the existing code
*substantially fulfills the ask* — feature-complete, tested, or explicitly
documented as done. Partial overlap is `partial`, not `premise-mismatch`.
- **Snapshot-vs-workspace drift can be benign.** Timestamps, lockfiles, and the
prior agent's exploratory scratch are not continuity violations. Only
load-bearing state counts — an edit the snapshot's turns present as done and
that the prompt or rubric relies on. Corroborate with the runs (agents
confused by the reset) before HIGH confidence.
- **Data-presence claims can be satisfied at runtime.** Seeds may be empty
while a setup script or fixture factory creates the data on boot. Check the
full bring-up path (Dockerfile, setup scripts, test fixtures), not just seed
files, before declaring data absent.
- **Reproducibility tracing is the deepest and most error-prone check.** Cap it
at what reading the configuration and the relevant code path can establish,
with citations; when the trace is inconclusive, report `partial` at
LOW/MEDIUM confidence rather than asserting the incident cannot occur.
The reference runs corroborate but are not required — the workspace check
stands alone. Where runs exist, look for agents reporting an empty diff, "this
already exists," missing data, or phantom workarounds for state that isn't
there, and for grades improvising anchors the rubric never defined.
## A sixth shape — artifacts from mixed revisions (`package-drift`)
A submission ships as one package: prompt, rubric, reference runs (each with
its grade and recorded score), workspace definition, snapshot. Nothing in it
needs to be broken for the package to be unsound: if the artifacts don't all
reflect the same revision of the task, the runs don't demonstrate the shipped
prompt and the shipped rubric would not produce the shipped scores — the
package cannot evidence its own task, and reviewers burn whole feedback
rounds on "you uploaded the old version." This earns its own verdict,
`package-drift`; the environment may build and test perfectly, and the
intentional-vs-incidental test doesn't apply (no prompt makes staleness the
subject).
Three sub-shapes, each a mostly mechanical join over artifacts already in the
package — the judgment call is confined to "is this divergence load-bearing
or cosmetic":
- **Stale re-upload.** The whole archive is an older revision than the
current round: prior-version artifacts throughout, the last round's
feedback visibly unaddressed even though the resubmission claims otherwise,
every pairwise comparison drifting in the same direction (all artifacts
current-minus-one). The fix is "rebuild and re-upload," not five separate
regenerations — say so.
- **Half-updated revision.** One artifact was refreshed and its counterpart
wasn't. The recurring joins:
- *Prompt ↔ runs.* Each run records the prompt the agent actually received
(a transcript/session artifact, or the prompt as quoted in `grade.md`).
Normalize away harness preamble and formatting, then compare the
task-content core against the shipped `instruction.md`. Fires on
substantive divergence — a different ask, missing or extra requirements,
or scaffold placeholder text ("# Replace this with your refined task
instruction") in the run-time prompt. Runs that record no prompt are
not-checkable, not evidence.
- *Rubric ↔ grades.* Extract the scoring structure each `grade.md`
applies — scored axes, heavy deductions and their magnitudes, any hard
gate/cap invoked (an older rubric shape: current guidance expresses
dealbreakers as heavy penalties, but you must still recognize cap
language in grades), tier names, quoted rubric phrases — and check each
load-bearing element exists in the shipped rubric (the resolved guidance
file).
The operative question: **would the shipped rubric, applied to this run,
plausibly produce this grade?** Fires on a clear no — e.g. every grade
"caps the overall score at 0.25" while the shipped rubric subtracts a
penalty instead.
- *Reward ↔ grade.* Each `reward.txt` should match the overall score its
`grade.md` arrives at. Under the Grading Standard there is no separate
correctness score — `reward-correctness.txt` legitimately reads `N/A`
and the grade has no `## Correctness` heading, which is the standard
working as designed, not drift. When a run's `grade.md` does carry a
`## Correctness` heading (runs graded under earlier toolkit releases),
its `reward-correctness.txt` should match the score (or `N/A`) under
that heading. Check the join against the shape the grade actually has. A
package-wide mismatch usually means the grades were revised after the runs
were scored and never re-copied — the half-updated signature in miniature.
One axis updated and the other left behind is the same shape: where a
grade writes both scores, they are written together, so they should never
disagree about which `grade.md` they came from.
- **Internal version drift.** The prompt, workspace, and snapshot record
states that cannot all be the same revision of the task: runs carrying
conflicting recorded task checksums (`result.json`) were generated against
different versions and cannot jointly evidence the shipped one; a snapshot
recorded against a workspace revision the shipped `workspace.patch` no
longer produces.
Lane lines, so this shape stays mechanical:
- **A stale run is not a corrupted run.** `runs-corrupted` owns runs killed
by the machinery around the agent; `package-drift` owns healthy runs that
evidence a different revision.
- **Premise-mismatch owns workspace-vs-prompt-assertion; package-drift owns
artifact-vs-artifact revision disagreement.** "The prompt promises an
uncommitted diff that isn't there" is premise; "the runs were generated
before the prompt said that" is drift.
- **Never audit the guidance's run citations here.** Guidance that describes
the observed runs at all — their count, scores, or behaviors — is
`detector-rubric-generality`'s flag, whether the citations are stale or
current. This shape joins the runs against the prompt and rubric, not
against the guidance's prose about runs.
- **Sibling reports under `detectors/` are out of scope** — they are
regenerated downstream, so staleness there is self-healing. Note it in one
sentence if you see it; don't fire on it.
Guards against overfiring:
- **Regrading is the fix, not the bug.** A run regraded against the final
rubric legitimately pairs an older transcript with a current `grade.md` —
that is exactly the remediation this shape's findings prescribe. Never fire
merely because a transcript predates the rubric; fire only when the grade's
*mechanism* isn't in the shipped rubric, or the transcript's recorded
prompt itself diverges from the shipped one.
- **Post-run copy edits are normal.** Workers are encouraged to polish rubric
wording after grading, and graders paraphrase rather than quote. Anchor on
named mechanisms and numbers (gate conditions, penalty sizes, tier
boundaries), which survive paraphrase — never require verbatim matches,
and fire only on structural divergence.
- **Harness framing isn't drift.** A run-recorded prompt may wrap a verbatim
`instruction.md` in preamble or formatting; require substantive content
divergence before firing.
- **A single cosmetic lag is `partial`.** One reward off by a rounding step,
wording lag with no scoring consequence — real, absorbable, worth a
sentence, not the verdict.
When firing, name the smallest coherent fix aimed at the join that failed:
regenerate runs against the shipped prompt, regrade against the shipped
rubric, re-copy the runs so each `reward.txt` matches its `grade.md`, or
rebuild and re-upload the archive.
If more than one shape is present (env breakage, corrupted runs, premise
mismatch, package drift), verdict whichever defect most invalidates the
submission's evidence and name the others in the Rationale.
## Verdict definitions
- **`not-applicable`** — there's no way to decide from this submission. Two
triggers:
- **No runnable environment in play**: the task is pure static analysis,
code review, or technical writing — the agent is never expected to build,
run, or test anything, so there is no dev environment that could be broken.
`instruction.md` asks only for prose/analysis and the reference runs show no
build/test/run attempts. **The premise check still applies here**: a
review/audit prompt can assert workspace state ("review my uncommitted
change") that the shipped tree contradicts. Only conclude `not-applicable`
when the prompt and snapshot also make no workspace-checkable assertions.
- **No evidence available**: there are no reference runs (or empty ones) AND
the resolved guidance file gives no signal about the environment, so there's
nothing to ground a breakage call on. Re-run once reference runs land.
Runs that **exist but are infrastructure-broken are not an evidence gap** —
that is `runs-corrupted`, a defect, not `not-applicable`.
- **`incidental-breakage`** — clear evidence (Shape 1, 2, or 3) that the local
dev environment is broken in a way **unrelated to the prompt's ask**, AND the
breakage materially distorted the task's signal: a run was blocked or
aborted, a substantial share of agent effort went to the breakage, or what
the grader saw (or the score) plausibly changed. The prompt does not ask the
agent to fix the environment. This is the verdict we do not want a task to
earn.
- **`runs-corrupted`** — at least one *scored, packaged* reference run never
reached a gradable state because of infrastructure: killed mid-work by an
API/model error, stranded in an unapprovable plan-mode ending with no shipped
work, truncated so the grader scored a transcript the agent didn't produce,
missing its output snapshot, or a verifier timeout counted as a run. The
workspace may be perfectly healthy — this verdict is about the run set, not
the env. List every affected run id and the signature found.
- **`premise-mismatch`** — at least one load-bearing premise the prompt or
snapshot asserts about the workspace does not hold in the shipped state, AND
the resolved guidance file assumes the premise holds (or nowhere credits
surfacing the discrepancy). The task as graded cannot exercise what it
describes. The environment may build and test perfectly — this verdict is
about the workspace being *wrong for the task*, not broken. Quote the
premise and the contradicting workspace evidence.
- **`package-drift`** — at least one load-bearing revision disagreement
between shipped artifacts: the archive is a pre-feedback revision
re-uploaded wholesale, a run's recorded prompt substantively diverges from
the shipped `instruction.md` (scaffold placeholder text included), grades
apply a scoring mechanism the shipped rubric does not contain, `reward.txt`
systematically disagrees with `grade.md`, or runs carry conflicting
recorded task checksums. Each artifact may be individually healthy — this
verdict is about the package's parts describing different revisions of the
task. Quote the divergent strings from both sides of the join and name the
smallest coherent fix.
- **`partial`** — breakage or friction is present but did not materially
distort the task's signal: a pre-existing unrelated failure the agents
confirm and route around, a single flaky retry, a one-line config nudge, or
infrastructure noise that only arrived after a run's substance had landed.
A genuine defect the task would be better without — worth naming so the
author can smooth it — but no run was blocked and the grade was unaffected.
If the friction plausibly changed how the agent spent its effort or what the
grader saw, escalate to `incidental-breakage`; if it's cosmetic, lean
`clean`. Also the verdict for a **weak or peripheral premise contradiction**:
the referenced data exists but is thinner than implied, the "new" feature
exists in a clearly-incomplete form the prompt could plausibly mean to
extend, or the mismatch is real but peripheral to what the rubric grades.
And for **cosmetic revision lag**: grades paraphrasing rubric wording that
was later lightly copy-edited, a single reward off by a rounding step —
divergence that would not change a score or mislead a reviewer.
- **`intentional`** — the environment breakage IS the subject of the task. The
prompt explicitly asks the agent to diagnose, fix, or repair the environment,
build, dependencies, tooling, or failing setup. The breakage is the point, so
it is not a hazard and not flagged. (The premise-mismatch analog — a
deliberately false premise the guidance grades surfacing — maps to `clean`,
not `intentional`; say so in the Rationale.)
- **`clean`** — no evidence the dev environment is incidentally broken,
every workspace-checkable premise in the prompt and snapshot holds in the
shipped state (or is deliberately false with the rubric grading its
discovery), and the checkable artifacts agree on one revision of the task.
The agent operated against a working baseline (or the task
depends on one and nothing in the runs or rubric shows unrelated env
friction). Tests failing because of the agent's own in-progress work, or
because the prompt's bug is the subject, are `clean`, not breakage.
## Confidence
- **HIGH** — grounding is unambiguous. The reference runs (or the rubric) show
the agent fighting a broken setup that the prompt plainly didn't ask about;
for `runs-corrupted`, a run's terminal events (or its grade) show the
infrastructure failure verbatim; for `premise-mismatch`, the check is
mechanical (an empty `git diff HEAD` against a promised uncommitted change,
a fully-shipped implementation against a "build X" ask) and the rubric
plainly assumes the premise; for `package-drift`, the divergence is
quotable from both sides of the join (the scaffold text in the run's
recorded prompt, cap language in every grade while the shipped rubric has
none); or, for `intentional`, the prompt explicitly
asks to fix the environment.
- **MEDIUM** — the pattern is present but interpretation is debatable. A
reasonable reviewer might read the friction as ordinary task difficulty.
- **LOW** — limited information; verdict is a best guess (often because the
reference runs are thin or the rubric is silent on the environment).
## Patterns to look for
In the reference runs:
- **Install / build / start failures early in the run**, followed by the agent
patching things the prompt never mentioned, just to get going.
- **The agent retrying the test suite**, or reasoning aloud about which
pre-existing failures are "expected" vs. caused by its change.
- **Answer prose that complains about or footnotes the environment** — "note the
suite had unrelated failures", "I couldn't run X so I worked around it."
- **Time/turns spent on tooling unrelated to the deliverable** — a large share
of the run going to environment repair rather than the actual ask.
- **Terminal events that are infrastructure, not behavior** — the last event is
an API/model error, an errored plan-mode exit with nothing after it, a tool
call with no result, or the trajectory just stops; the output snapshot is
missing → `runs-corrupted` territory.
In the shipped environment (statically — even when the runs look quiet):
- **A capability the prompt or rubric requires that the image can't provide** —
a browser the rubric expects verification in that was never installed, a
package the deliverable imports that is absent from every manifest and
lockfile, a tool that can only be installed from the network. Check the
Dockerfile and lockfiles against what the ask and its verification assume.
In the workspace, checked against the prompt and snapshot (the premise check):
- **A promised pending change that isn't pending** — the prompt says "review my
uncommitted change" and `git status` / `git diff HEAD` come back clean.
- **The ask already delivered** — the prompt asks to build/add/first-pass a
capability and the workspace (including `workspace.patch`) ships it
substantially complete, with tests or docs presenting it as done.
- **Prior-turn state that didn't survive** — the snapshot's turns fixed (or
broke) something the shipped tree doesn't reflect.
- **Referenced data or files absent** — seeds create zero rows of the data the
prompt says is "already there"; a named branch/file doesn't exist, and no
bring-up step creates it.
- **Run corroboration** — agents reporting an empty diff or "this already
exists," burning turns on workarounds for state that isn't there, grades
improvising anchors.
Across the shipped artifacts (the version-coherence check):
- **Grades invoking a mechanism the shipped rubric lacks** — cap/gate
language, tier names, or penalty magnitudes absent from
the resolved guidance file.
- **A run-recorded prompt that isn't the shipped prompt** — scaffold
placeholder text, or a substantively different ask.
- **`reward.txt` disagreeing with `grade.md` across the run set** — the
grades were revised and the scores never re-copied.
- **Conflicting recorded task checksums across runs**, or the prior round's
feedback still visibly unaddressed in a resubmitted archive.
In `instruction.md` (to separate intentional from incidental):
- Asks to **fix / repair / debug the env, build, deps, or failing setup** →
lean `intentional`.
- Asks for a **feature, audit, trace, design, or fix to specific app behavior**,
with breakage showing up anyway → lean `incidental-breakage`.
In the resolved guidance file:
- Instructions to the grader to **discount, ignore, or expect** environment
failures the agent shouldn't be blamed for → the env is broken and the rubric
is papering over it (incidental).
## What you are NOT doing
- **Not flagging a task whose subject is the broken environment** — that's
`intentional`. Read `instruction.md` before deciding.
- **Not flagging legitimate red tests** — the agent's own in-progress work, or a
failing test the prompt asks the agent to fix, is the task working.
- **Not flagging every imperfect run as corrupted** — `runs-corrupted` requires
an infrastructure event at the run's terminal state, not a low score, a terse
ending, or a deliberate stop-and-ask the rubric credits.
- **Not flagging a deliberately false premise the rubric grades.** A task built
around a wrong user belief, where the guidance knows the true workspace state
and credits the agent for surfacing it, is a valid design — the
premise-mismatch shape fires only when the rubric assumes the premise holds.
- **Not flagging wording drift between rubric and grades.** Paraphrase is
normal; structure and numbers are the signal. And an older-but-regraded run
is the prescribed remediation, not drift — check the transcript's recorded
prompt, not its age.
- **Not auditing the guidance's descriptions of runs.** Run-anchored guidance
— stale or current — is `detector-rubric-generality`'s lane; the
package-drift shape joins the runs against the prompt and rubric only.
- **Not grading the agent's submission** or re-deriving any other detector's
call. This detector is solely about whether the environment is incidentally
broken and worked-around, whether the scored runs are valid samples of
agent behavior, whether the workspace matches the task's stated premise,
and whether the shipped artifacts agree on one revision of the task.
## Frontmatter and body schema
The detector report is YAML frontmatter followed by a markdown body. Both
contexts produce the same shape; only the *sink* differs (the wrapping
`SKILL.md` tells you where to send the report).
**Frontmatter** — exactly these keys, exactly these enum values:
```yaml
---
detector: detector-broken-dev-env
verdict: incidental-breakage | runs-corrupted | premise-mismatch | package-drift | partial | intentional | clean | not-applicable
confidence: HIGH | MEDIUM | LOW
---
```
**Body sections**, in this order:
```markdown
# Broken-dev-env check: <slug>
## Verbatim grounding
Pull the load-bearing quotes that justify the verdict. Quote them inline as
blockquotes — don't paraphrase. For `incidental-breakage` / `partial`: quote the
reference-run text (or rubric line) that shows the agent hitting / working around
the broken environment, AND quote the part of `instruction.md` that shows the
prompt did NOT ask for it. For `runs-corrupted`: quote the terminal trajectory
events (or the `grade.md` text) that show the infrastructure failure — the
error string, the truncation point, the grader tripping over the missing
ending — and name each affected run id. For `premise-mismatch`: quote the
premise verbatim from `instruction.md` or the snapshot AND the workspace
evidence contradicting it (the `git diff HEAD` output, the file/commit that
already ships the ask, the empty seed, the missing prior-turn state), plus the
guidance line showing the rubric assumes the premise holds. For
`package-drift`: quote the exact divergent strings from **both sides** of the
join — the run-recorded prompt line next to the shipped `instruction.md`
line, the grade's cap/penalty language next to the shipped rubric's
mechanism, the `reward.txt` value next to the `grade.md` score line, the
conflicting checksums — and name each affected run id. A drift call asserted
without paired quotes is unreviewable. For `intentional`: quote the part of
`instruction.md` that asks the agent to fix the environment. For `clean`: quote
what the runs / rubric DO show (a working baseline, or task-intrinsic red
tests) so the reader can confirm. For `not-applicable`: quote the artifact
showing the trigger (the prompt asking only for prose, or the missing
reference runs).
## Rationale
2–4 paragraphs explaining what is broken (or why nothing is), tied to the
grounding above. Be specific: which shape (1/2/3, the fourth runs-corrupted,
the fifth premise-mismatch, or the sixth package-drift)? Which run shows the
workaround — or, for
`runs-corrupted`, which runs are invalid and whether each reached a gradable
state before the infrastructure event — or, for `premise-mismatch`, which
premise type failed and whether the rubric assumes it holds or credits its
discovery — or, for `package-drift`, which join failed, whether the
divergence is load-bearing or cosmetic, and the smallest coherent fix
(regenerate runs, regrade, re-copy rewards, or rebuild and re-upload)? Why is
the breakage unrelated to the prompt's ask (or,
for `intentional`, why it IS the ask)? For `not-applicable`, explain which
trigger fired and what would make the detector runnable.
```
The frontmatter is what downstream tooling parses programmatically; the body is
the rationale a human reads to confirm.