Loaded up for the 3rd redo
Still on potion-voice
This commit is contained in:
@@ -0,0 +1,95 @@
|
||||
---
|
||||
name: detector-broken-dev-env
|
||||
description: |
|
||||
Self-check whether your task presents an *incidentally* broken local dev
|
||||
environment that the test agent has to awkwardly work around. The workspace
|
||||
doesn't build/install/run, a dependency or service is missing, or the test
|
||||
suite has pre-existing failures or flakes unrelated to your task — and the
|
||||
agent burns effort coping with that instead of doing what your prompt asks.
|
||||
These tasks are weak: the existing test suite is the main verifier, so env
|
||||
noise lands straight in the grade. The one allowed shape is intentional
|
||||
breakage — a task whose subject IS the broken env ("my dev env is broken, fix
|
||||
it"). A pre-existing app bug that your prompt asks the agent to find or fix is
|
||||
the task working, not breakage. Also checks that the workspace is actually in
|
||||
the state your prompt (or snapshot) says it's in — a promised uncommitted
|
||||
change that's already committed, a "build X" ask where X already ships, or
|
||||
referenced data that isn't there is a premise mismatch even when everything
|
||||
builds green. And checks that everything you package reflects the same
|
||||
revision of your task — runs graded under an earlier prompt or rubric, a
|
||||
reward.txt that no longer matches its grade.md, or a re-uploaded older
|
||||
archive is package drift even when every artifact is individually healthy.
|
||||
allowed-tools: Bash, Read, Write
|
||||
---
|
||||
|
||||
# Broken-dev-env detector
|
||||
|
||||
This skill checks whether your task hands the agent a dev environment that is
|
||||
broken for reasons unrelated to what you're asking it to do — a build that won't
|
||||
run, a missing dependency, or a test suite with pre-existing failures the prompt
|
||||
never mentions. If the agent has to fight that breakage to make progress, the
|
||||
task is testing "can the agent cope with a broken env" instead of the behavior
|
||||
you meant to grade, and the verifier signal gets noisy. It also checks that
|
||||
your scored reference runs are valid samples of agent behavior — a run killed
|
||||
mid-work by an API error, truncated, or missing its output snapshot reflects
|
||||
infrastructure, not the agent, and shouldn't ship as evidence. And it checks
|
||||
that the workspace matches your task's stated premise: if your prompt or
|
||||
snapshot asserts something about the workspace ("review my uncommitted
|
||||
change", "there's already data in the repo", a previous turn's fix) that the
|
||||
shipped state contradicts, the agent responds to the workspace as it actually
|
||||
is and your rubric grades a task that can't happen. Finally, it checks that
|
||||
your package is one coherent revision of the task: if you polish the prompt or
|
||||
rubric after generating runs, the shipped runs and grades must be regenerated
|
||||
or regraded to match — a package whose parts describe different versions of
|
||||
the task can't evidence it.
|
||||
|
||||
Read these before deciding:
|
||||
|
||||
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
|
||||
2. `.claude/skills/detector-broken-dev-env/core.md` — intentional-vs-incidental, the three shapes breakage takes plus the runs-corrupted, premise-mismatch, and package-drift shapes, what counts as legitimate task difficulty, verdict enums.
|
||||
|
||||
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
|
||||
|
||||
## Acting on the verdict
|
||||
|
||||
- **`clean`** — your environment runs fine (or the only red tests are the bug
|
||||
your prompt is about). Good.
|
||||
- **`partial`** — there's some env friction, but it's minor or borderline. Read
|
||||
the rationale; either smooth the setup so the agent never hits it, or confirm
|
||||
it's cosmetic enough not to distort the run.
|
||||
- **`incidental-breakage`** — the agent has to work around a broken setup your
|
||||
prompt didn't ask it to fix. Fix the environment (repair the Dockerfile,
|
||||
pin deps, remove the unrelated failing/flaky tests) so the agent starts from a
|
||||
working baseline, then re-run this skill. Don't try to rescue it by reframing
|
||||
the breakage as the task — see `intentional`.
|
||||
- **`runs-corrupted`** — one or more of your scored reference runs was ended or
|
||||
distorted by infrastructure rather than by the agent (an API/model error
|
||||
mid-run, a truncated trajectory, a missing output snapshot, a verifier
|
||||
timeout), so it isn't a valid sample of agent behavior. Your workspace may be
|
||||
perfectly healthy. Re-run the affected trials, replace the corrupted runs,
|
||||
repackage, then re-run this skill.
|
||||
- **`premise-mismatch`** — the shipped workspace contradicts what your prompt
|
||||
or snapshot asserts (the promised uncommitted change is already committed,
|
||||
the feature you ask the agent to build already exists, referenced data is
|
||||
absent, a prior turn's state was reset away), and your holistic rubric
|
||||
assumes the premise holds. Either fix the workspace so the premise is true
|
||||
(workspace.patch, seeds, snapshot end-state), or — if the false premise is
|
||||
deliberate — make the rubric grade the agent on surfacing it, then re-run
|
||||
this skill and re-collect reference runs.
|
||||
- **`package-drift`** — your packaged artifacts don't all reflect the same
|
||||
revision of the task: runs were graded under an earlier prompt or rubric, a
|
||||
reward.txt no longer matches its grade.md, runs record conflicting task
|
||||
versions, or you re-uploaded an older archive after making fixes. Nothing
|
||||
may be broken — but the runs no longer demonstrate the shipped task. Apply
|
||||
the smallest coherent fix: regenerate runs against the current prompt,
|
||||
regrade against the current rubric (see `/regrade-reference-run`), re-copy
|
||||
the runs so each reward.txt matches its grade.md, or rebuild and re-upload
|
||||
the archive — then re-run this skill.
|
||||
- **`intentional`** — your task is explicitly about fixing the environment. That
|
||||
is a valid task; nothing to change. (Only legitimate if your *prompt* asks for
|
||||
the repair — not if the agent merely ended up coping with a broken env.)
|
||||
- **`not-applicable`** — the task has no runnable environment (pure analysis /
|
||||
writing) AND the prompt/snapshot make no workspace-checkable assertions, or
|
||||
there's no evidence yet (no reference runs and the rubric says
|
||||
nothing about the env). Re-run once you have reference runs. (Runs that exist
|
||||
but died on infrastructure are `runs-corrupted`, not this; a prose prompt
|
||||
that asserts workspace state can still earn `premise-mismatch`.)
|
||||
@@ -0,0 +1,683 @@
|
||||
# Broken-dev-env detector — core
|
||||
|
||||
This file is the canonical, context-neutral content for the detector-broken-dev-env
|
||||
detector. It defines what the detector looks for, the verdict enums, the
|
||||
patterns to recognize, and the output schema. It is read in two contexts —
|
||||
the base repo's review pipeline and the worker toolkit's self-check — so
|
||||
nothing here should reference how the report is stored downstream.
|
||||
|
||||
## What this detector is for
|
||||
|
||||
A task hands the agent a repo at a chosen commit plus a prompt, and the agent
|
||||
works in a local dev environment (the task workspace). Sometimes that
|
||||
environment is **broken in a way that has nothing to do with the prompt's
|
||||
ask**: the workspace doesn't install or build out of the box, a binary or
|
||||
dependency is missing, env vars aren't set, a service won't start, a migration
|
||||
is wedged, or the test suite has pre-existing failures or flakes unrelated to
|
||||
the task. The agent then burns effort diagnosing and working around that
|
||||
breakage instead of (or on top of) doing the work the prompt actually asked
|
||||
for.
|
||||
|
||||
We do not want tasks where the broken environment is **incidental** — an
|
||||
accidental artifact of how the task was extracted, left in the workspace, and
|
||||
silently presented to the agent as just one more hazard to fight through. It
|
||||
makes the task noisy: the agent's score then partly reflects whether it could
|
||||
push through a broken setup, not whether it did the intended work. The existing
|
||||
test suite is the primary verifier for these tasks, so environment noise in the
|
||||
build or the tests directly muddies the signal the task is supposed to produce.
|
||||
|
||||
There is one legitimate shape: **intentional** breakage. If the task is
|
||||
*explicitly about* the broken environment — "my local dev env is broken, can
|
||||
you fix it", "the test suite won't run, figure out why", "the build is red, get
|
||||
it green" — then a broken environment is the deliberate subject of the task, not
|
||||
a hazard. That is fine and must not be flagged.
|
||||
|
||||
This detector also owns an adjacent defect in the same "the grade reflects
|
||||
infrastructure, not the work" family: **reference runs corrupted by the
|
||||
execution infrastructure** rather than by anything the agent did. An API or
|
||||
model error kills a run mid-implementation, a trajectory is truncated so the
|
||||
grader scores a transcript the agent never produced, an output snapshot is
|
||||
missing, a verifier times out — and the run is scored and shipped as if it
|
||||
showed real agent behavior. The workspace can be perfectly healthy in every one
|
||||
of these cases; see the dedicated shape section below.
|
||||
|
||||
And it owns a third defect in the same family where nothing is broken at all:
|
||||
the shipped workspace **contradicts the premise the task states**. The prompt
|
||||
(or a snapshot's prior turns) asserts something concrete about the workspace —
|
||||
"review my uncommitted change," "take a first pass at building X," "there's
|
||||
already data seeded in the repo," "in the previous turn you fixed Y" — and the
|
||||
workspace the agent actually receives doesn't honor it: the promised diff is
|
||||
already committed, the feature to build already ships complete, the referenced
|
||||
data is absent, the prior turn's state was reset away. Everything installs and
|
||||
tests green, yet the environment is wrong *for the task*: agents reasonably
|
||||
respond to the workspace as it actually is, and the rubric — written as if the
|
||||
premise held — either can't be applied or is applied unfairly. See the
|
||||
dedicated shape section below.
|
||||
|
||||
The last defect in the family is **package drift**: the shipped artifacts
|
||||
don't all reflect the same revision of the task. Workers iterate — polish the
|
||||
prompt after generating runs, rewrite the rubric after grading, re-upload an
|
||||
archive after feedback without rebuilding it — and each of those steps can
|
||||
ship a package whose parts disagree about which version of the task they
|
||||
belong to: runs graded under an earlier prompt (sometimes still scaffold
|
||||
placeholder text), grades produced against an earlier rubric, an old archive
|
||||
resubmitted wholesale. Every artifact can be individually healthy and the
|
||||
package still fails to evidence its own task. See the dedicated shape section
|
||||
below.
|
||||
|
||||
Taken together, this is the "is the submission package itself sound?" check —
|
||||
the environment, the runs, the workspace-vs-premise fit, and the version
|
||||
coherence of the shipped artifacts. This detector decides: does *this*
|
||||
submission present an incidentally broken dev environment that the agent has
|
||||
to awkwardly work around — a reference-run set corrupted by infrastructure
|
||||
rather than agent behavior — a workspace that contradicts the premise the
|
||||
prompt or snapshot asserts — or a package whose artifacts ship from different
|
||||
revisions of the task?
|
||||
|
||||
## Inputs
|
||||
|
||||
Read whatever you need from `harbor-tasks/<slug>/`. The load-bearing artifacts:
|
||||
|
||||
- `instruction.md` — the prompt the agent received. **This is what decides
|
||||
intentional vs. incidental.** Does the prompt ask the agent to diagnose, fix,
|
||||
or repair the environment / build / dependencies / tooling / failing setup? If
|
||||
yes, breakage is the subject (intentional). If the prompt asks for something
|
||||
else entirely (implement feature X, audit module Y, write design doc Z) and
|
||||
the environment is nonetheless broken, the breakage is incidental.
|
||||
- `reference-runs/<run>/agent-output/answer.md` and the run's trajectory — the
|
||||
test agent's actual behavior. Sample 2–3 runs. Look for the agent spending
|
||||
turns getting to a runnable baseline: install/build failures, missing
|
||||
binaries, "the tests won't run so I…", patching config unrelated to the task,
|
||||
re-running with workarounds, or prose in the answer noting that the
|
||||
environment was broken. This is the strongest evidence that the breakage
|
||||
actually distorted the run. Also check each run's **terminal state** (the
|
||||
last few events of the trajectory, whether `agent-output/` exists and is
|
||||
non-empty, and whether `grade.md` itself notices an abrupt ending) — that is
|
||||
what decides `runs-corrupted`, per the shape section below.
|
||||
- Per run, the version-coherence artifacts: `grade.md` (the scoring structure
|
||||
the grader actually applied), `reward.txt` and `reward-correctness.txt`
|
||||
(the recorded scores; under the Grading Standard
|
||||
`reward-correctness.txt` legitimately reads `N/A`),
|
||||
`result.json` / `config.json` when present (recorded task name/checksum),
|
||||
and any transcript/session artifact that records the prompt the agent
|
||||
actually received. These are what the `package-drift` shape joins against
|
||||
the shipped `instruction.md` and the resolved guidance file — see the shape
|
||||
section below.
|
||||
- `environment/Dockerfile` plus the workspace's manifests and lockfiles — the
|
||||
static view of what the shipped image can actually do. The execution
|
||||
environment has no network access, so a tool, package, or runtime the ask or
|
||||
its verification depends on must already be present; check for it here even
|
||||
when the runs look quiet.
|
||||
- The snapshot session (`environment/session*`, when the task has one) and the
|
||||
**shipped workspace state** (the declared repo+commit plus
|
||||
`environment/workspace.patch`) — the two halves of the premise check.
|
||||
The prompt and the snapshot's prior turns are the source of
|
||||
workspace-checkable assertions; the built workspace (git status/diff, file
|
||||
and branch existence, seed contents, whether a named feature or fix is
|
||||
present) is the ground truth they are checked against. See the
|
||||
`premise-mismatch` shape section below.
|
||||
- The grader guidance — the rubric. Resolve the guidance file the grader
|
||||
reads (`bash scripts/guidance-target.sh <slug>` prints its path,
|
||||
`tests/grader-guidance-consolidated.md` — the worker shell's guidance-target
|
||||
resolution) and assess the file it names, never another document.
|
||||
Sometimes the rubric itself reveals
|
||||
the environment is broken: "note that the suite has a pre-existing failure in
|
||||
X, ignore it", "the dev server doesn't start; a strong agent works around
|
||||
it", "don't penalize the agent for the broken migration." A rubric that treats
|
||||
env breakage as an obstacle course the agent must navigate (rather than the
|
||||
thing to fix) is a strong incidental-breakage signal.
|
||||
- `task.toml` — for the source repo and commit, when you need to confirm whether
|
||||
a failure the agent hit is pre-existing in the workspace vs. introduced by the
|
||||
agent.
|
||||
|
||||
## Incidental vs. legitimate task difficulty
|
||||
|
||||
The hard part of this detector is not mistaking the task working as intended for
|
||||
incidental breakage. Keep these straight:
|
||||
|
||||
- **The task's own bug or failing test is not breakage.** If the prompt is "fix
|
||||
the failing `X` test" or "the agent's change should make the suite pass", then
|
||||
a red suite at the start is the *subject* of the task. Breakage only counts as
|
||||
incidental when it is **unrelated to the prompt's ask**.
|
||||
- **TDD is not breakage.** An agent writing code and watching tests go red→green
|
||||
as it works is the loop functioning, not a broken environment.
|
||||
- **A pre-existing failing test unrelated to the task is incidental.** The agent
|
||||
can't trust the suite's signal and has to reason about which failures are
|
||||
"expected" — noise the prompt never asked it to deal with.
|
||||
- **Setup the agent must repair just to reach a working baseline is incidental**
|
||||
when the prompt didn't ask for it. Pinning a dependency version, recreating a
|
||||
missing file, or hand-fixing a config to get install/build/run to succeed —
|
||||
all unrelated to the actual deliverable — is the classic shape.
|
||||
|
||||
## Three shapes incidental breakage takes
|
||||
|
||||
Any one of them establishes that incidental breakage *exists* — but presence
|
||||
alone earns at most `partial`. Escalate to `incidental-breakage` only when the
|
||||
breakage **materially distorted the task's signal**: it blocked or aborted a
|
||||
run, consumed a substantial share of the agent's effort (well beyond confirming
|
||||
a known-unrelated failure and moving on — agents spending a modest slice of a
|
||||
run establishing the baseline and then proceeding unimpeded is `partial`
|
||||
territory), or plausibly changed what the grader saw or the score. Genuine
|
||||
breakage that the agents note, route around, and that leaves no trace in the
|
||||
grade stays `partial` — worth fixing, not task-disqualifying.
|
||||
|
||||
**Shape 1 — setup/build breakage before the work can start.** The workspace
|
||||
doesn't install, build, or start out of the box for reasons unrelated to the
|
||||
task. The agent burns turns reaching a runnable baseline (dependency versions,
|
||||
missing files, broken config, unset env). The prompt never asked for any of it.
|
||||
|
||||
Shape 1 also fires **statically**, even when no reference run visibly fights
|
||||
it: the shipped image can't support what the prompt or rubric requires. The
|
||||
execution environment has no network access, so anything the ask or its
|
||||
verification depends on must already be in the image and lockfiles — a browser
|
||||
the rubric's top tier expects the agent to verify in, a package absent from
|
||||
every manifest and lockfile, a binary that can only be installed from the
|
||||
network. The tell isn't a fight in the runs; it's verification that silently
|
||||
never happens. Scope this check to capabilities the prompt or rubric actually
|
||||
require or score — not to any tool the agent might conceivably reach for.
|
||||
|
||||
**Shape 2 — pre-existing failing or flaky tests the agent must navigate.** The
|
||||
test suite has failures or flakes unrelated to the task. The agent can't trust
|
||||
green/red, has to retry or guess which failures are "expected," and the
|
||||
verifier's own signal is polluted. This is the most corrosive shape because the
|
||||
existing suite is the task's primary verifier — noise here lands straight in the
|
||||
grade.
|
||||
|
||||
**Shape 3 — broken tooling framed as a hazard in the rubric.** The
|
||||
grader-guidance explicitly tells the grader the environment is broken and that
|
||||
the agent should (or shouldn't) be distracted by it — "ignore the failing lint
|
||||
step", "the server won't start; a strong agent works around it." The rubric
|
||||
treats env breakage as an obstacle course rather than the deliverable.
|
||||
|
||||
The shapes can co-occur; cite each one you see.
|
||||
|
||||
## A fourth shape — reference runs corrupted by infrastructure (`runs-corrupted`)
|
||||
|
||||
A submission's reference runs can be invalidated by the machinery *around* the
|
||||
agent even when the workspace itself is perfectly healthy: an API or model
|
||||
error kills a run mid-implementation, a headless plan-mode ending strands the
|
||||
agent waiting for an approval that never comes, a trajectory is truncated
|
||||
mid-tool-call so the grader scores a transcript the agent never produced, an
|
||||
agent-output snapshot is missing or corrupt, or a verifier timeout is counted
|
||||
as a scored run. The damage takes two forms: the run set no longer evidences
|
||||
the task (a low score reflects the infrastructure, not the agent), and the
|
||||
grader can actively mis-grade — e.g. a completion-honesty penalty fired on a
|
||||
final message the truncation ate.
|
||||
|
||||
This is not env breakage, and the intentional-vs-incidental test doesn't apply
|
||||
(no prompt makes an API error the subject). It earns its own verdict,
|
||||
`runs-corrupted` — never `incidental-breakage`, which would misdescribe a
|
||||
healthy workspace.
|
||||
|
||||
Per scored run, check the terminal state:
|
||||
|
||||
- Does the trajectory end with a complete final assistant message — or
|
||||
mid-tool-call, on an unparseable tail, or with an infrastructure error string
|
||||
(an API 4xx, an invalid-model error, a plan-mode exit that errored with no
|
||||
subsequent agent turn) as the last event?
|
||||
- Is `agent-output/` present and non-empty?
|
||||
- Does `grade.md` itself notice the incompleteness ("the run ends abruptly",
|
||||
"no final summary") — or, worse, score the truncated state as if it were the
|
||||
agent's behavior?
|
||||
|
||||
The load-bearing boundary is whether the run reached a **gradable state before
|
||||
the infrastructure event**. A run killed mid-investigation with a clean tree
|
||||
and no answer never became a valid sample of agent behavior — that fires. An
|
||||
error that only ate the closing summary *after* the fix, tests, and substance
|
||||
had all landed leaves the run usable — note it in the body as `partial`-grade
|
||||
noise, not corruption.
|
||||
|
||||
Guards against overfiring:
|
||||
|
||||
- **Match infrastructure signatures only in a run's terminal events**, never by
|
||||
searching the whole transcript — agents quote error text while debugging, and
|
||||
repos discuss API errors in prose.
|
||||
- **A deliberate stop is legitimate behavior, not corruption.** An agent that
|
||||
presents a plan or asks a question as its chosen ending — a shape the rubric
|
||||
credits — ended naturally. The corruption case is the run trying to continue
|
||||
and being unable to: the plan-mode exit returns an error, no agent turn
|
||||
follows, and nothing ships.
|
||||
- **Brevity is not truncation.** Truncation needs structural evidence — a last
|
||||
event that is a tool call, an unparseable tail, or a missing final message
|
||||
the grade itself trips over — not a stylistic judgment about a terse ending.
|
||||
|
||||
## A fifth shape — the workspace contradicts the task's premise (`premise-mismatch`)
|
||||
|
||||
Sometimes the environment installs, builds, and tests green — nothing is
|
||||
"broken" in the workaround sense — yet the workspace is not in the state the
|
||||
task *says* it is in. The prompt, and any prior snapshot turns, make concrete
|
||||
assertions about the workspace, and the graded workspace either honors them or
|
||||
it doesn't. When it doesn't, and the rubric was written assuming it does, the
|
||||
task exercises something other than what it describes: runs sail past the
|
||||
intended difficulty, improvise a different task than the one described, or get
|
||||
penalized for reasonably responding to the environment as it actually is.
|
||||
|
||||
The recurring premise types, each checkable against the shipped workspace:
|
||||
|
||||
- **Pending-change** — the prompt promises uncommitted edits ("review my
|
||||
uncommitted change", "the diff on my branch"), but the tree is clean and the
|
||||
change is folded into an existing commit, so `git diff HEAD` is empty.
|
||||
- **Absence** — the prompt asks the agent to "add" / "build" / "take a first
|
||||
pass at" a capability that the workspace (including `workspace.patch`)
|
||||
already ships substantially complete, so most of the prompt isn't actionable
|
||||
as written.
|
||||
- **Continuity** — the snapshot's prior turns leave the tree in a state (a fix
|
||||
landed, a breakage present) that the graded workspace does not carry: the
|
||||
checkout was reset or repaired between turns, so the agent replays history
|
||||
that no longer matches the tree it is acting on.
|
||||
- **Presence** — the prompt references load-bearing data or files ("there's
|
||||
already history data in the repo", a named branch or config file) that the
|
||||
shipped state doesn't have: zero seeded rows, no such file.
|
||||
- **Reproducibility** — the incident the prompt reports cannot occur in the
|
||||
shipped configuration: the symptom only manifests in a test double, or the
|
||||
code path the described failure depends on isn't wired.
|
||||
|
||||
The decision procedure: **extract** every workspace-checkable assertion from
|
||||
`instruction.md` and the snapshot session; **verify** each against the built
|
||||
workspace (git status/diff for pending-change, code search and reading for
|
||||
absence/presence, the snapshot's implied end-state vs. the shipped tree for
|
||||
continuity, the configuration and code path for reproducibility); then
|
||||
**classify** each failed premise against the resolved guidance file — does the
|
||||
rubric assume the premise holds (grades content only reachable if it holds,
|
||||
describes the task in the premise's terms), or does it know the true state and
|
||||
credit the agent for surfacing the discrepancy?
|
||||
|
||||
That last question is the shape's carve-out, the analog of the
|
||||
intentional/incidental test (which itself doesn't apply here — no workaround is
|
||||
involved): **a deliberately false premise is a core, legitimate task design.**
|
||||
Many good tasks hand the agent a wrong user belief on purpose and grade whether
|
||||
the agent surfaces it. Never fire on "the premise is false" alone — fire only
|
||||
when the rubric itself assumes the premise holds, or nowhere credits
|
||||
discovering that it doesn't.
|
||||
|
||||
Guards against overfiring:
|
||||
|
||||
- **"Already exists" is a judgment call on partial implementations.** An ask to
|
||||
add a capability when a half-wired helper exists may legitimately mean
|
||||
"finish it." Treat an absence premise as violated only when the existing code
|
||||
*substantially fulfills the ask* — feature-complete, tested, or explicitly
|
||||
documented as done. Partial overlap is `partial`, not `premise-mismatch`.
|
||||
- **Snapshot-vs-workspace drift can be benign.** Timestamps, lockfiles, and the
|
||||
prior agent's exploratory scratch are not continuity violations. Only
|
||||
load-bearing state counts — an edit the snapshot's turns present as done and
|
||||
that the prompt or rubric relies on. Corroborate with the runs (agents
|
||||
confused by the reset) before HIGH confidence.
|
||||
- **Data-presence claims can be satisfied at runtime.** Seeds may be empty
|
||||
while a setup script or fixture factory creates the data on boot. Check the
|
||||
full bring-up path (Dockerfile, setup scripts, test fixtures), not just seed
|
||||
files, before declaring data absent.
|
||||
- **Reproducibility tracing is the deepest and most error-prone check.** Cap it
|
||||
at what reading the configuration and the relevant code path can establish,
|
||||
with citations; when the trace is inconclusive, report `partial` at
|
||||
LOW/MEDIUM confidence rather than asserting the incident cannot occur.
|
||||
|
||||
The reference runs corroborate but are not required — the workspace check
|
||||
stands alone. Where runs exist, look for agents reporting an empty diff, "this
|
||||
already exists," missing data, or phantom workarounds for state that isn't
|
||||
there, and for grades improvising anchors the rubric never defined.
|
||||
|
||||
## A sixth shape — artifacts from mixed revisions (`package-drift`)
|
||||
|
||||
A submission ships as one package: prompt, rubric, reference runs (each with
|
||||
its grade and recorded score), workspace definition, snapshot. Nothing in it
|
||||
needs to be broken for the package to be unsound: if the artifacts don't all
|
||||
reflect the same revision of the task, the runs don't demonstrate the shipped
|
||||
prompt and the shipped rubric would not produce the shipped scores — the
|
||||
package cannot evidence its own task, and reviewers burn whole feedback
|
||||
rounds on "you uploaded the old version." This earns its own verdict,
|
||||
`package-drift`; the environment may build and test perfectly, and the
|
||||
intentional-vs-incidental test doesn't apply (no prompt makes staleness the
|
||||
subject).
|
||||
|
||||
Three sub-shapes, each a mostly mechanical join over artifacts already in the
|
||||
package — the judgment call is confined to "is this divergence load-bearing
|
||||
or cosmetic":
|
||||
|
||||
- **Stale re-upload.** The whole archive is an older revision than the
|
||||
current round: prior-version artifacts throughout, the last round's
|
||||
feedback visibly unaddressed even though the resubmission claims otherwise,
|
||||
every pairwise comparison drifting in the same direction (all artifacts
|
||||
current-minus-one). The fix is "rebuild and re-upload," not five separate
|
||||
regenerations — say so.
|
||||
- **Half-updated revision.** One artifact was refreshed and its counterpart
|
||||
wasn't. The recurring joins:
|
||||
- *Prompt ↔ runs.* Each run records the prompt the agent actually received
|
||||
(a transcript/session artifact, or the prompt as quoted in `grade.md`).
|
||||
Normalize away harness preamble and formatting, then compare the
|
||||
task-content core against the shipped `instruction.md`. Fires on
|
||||
substantive divergence — a different ask, missing or extra requirements,
|
||||
or scaffold placeholder text ("# Replace this with your refined task
|
||||
instruction") in the run-time prompt. Runs that record no prompt are
|
||||
not-checkable, not evidence.
|
||||
- *Rubric ↔ grades.* Extract the scoring structure each `grade.md`
|
||||
applies — scored axes, heavy deductions and their magnitudes, any hard
|
||||
gate/cap invoked (an older rubric shape: current guidance expresses
|
||||
dealbreakers as heavy penalties, but you must still recognize cap
|
||||
language in grades), tier names, quoted rubric phrases — and check each
|
||||
load-bearing element exists in the shipped rubric (the resolved guidance
|
||||
file).
|
||||
The operative question: **would the shipped rubric, applied to this run,
|
||||
plausibly produce this grade?** Fires on a clear no — e.g. every grade
|
||||
"caps the overall score at 0.25" while the shipped rubric subtracts a
|
||||
penalty instead.
|
||||
- *Reward ↔ grade.* Each `reward.txt` should match the overall score its
|
||||
`grade.md` arrives at. Under the Grading Standard there is no separate
|
||||
correctness score — `reward-correctness.txt` legitimately reads `N/A`
|
||||
and the grade has no `## Correctness` heading, which is the standard
|
||||
working as designed, not drift. When a run's `grade.md` does carry a
|
||||
`## Correctness` heading (runs graded under earlier toolkit releases),
|
||||
its `reward-correctness.txt` should match the score (or `N/A`) under
|
||||
that heading. Check the join against the shape the grade actually has. A
|
||||
package-wide mismatch usually means the grades were revised after the runs
|
||||
were scored and never re-copied — the half-updated signature in miniature.
|
||||
One axis updated and the other left behind is the same shape: where a
|
||||
grade writes both scores, they are written together, so they should never
|
||||
disagree about which `grade.md` they came from.
|
||||
- **Internal version drift.** The prompt, workspace, and snapshot record
|
||||
states that cannot all be the same revision of the task: runs carrying
|
||||
conflicting recorded task checksums (`result.json`) were generated against
|
||||
different versions and cannot jointly evidence the shipped one; a snapshot
|
||||
recorded against a workspace revision the shipped `workspace.patch` no
|
||||
longer produces.
|
||||
|
||||
Lane lines, so this shape stays mechanical:
|
||||
|
||||
- **A stale run is not a corrupted run.** `runs-corrupted` owns runs killed
|
||||
by the machinery around the agent; `package-drift` owns healthy runs that
|
||||
evidence a different revision.
|
||||
- **Premise-mismatch owns workspace-vs-prompt-assertion; package-drift owns
|
||||
artifact-vs-artifact revision disagreement.** "The prompt promises an
|
||||
uncommitted diff that isn't there" is premise; "the runs were generated
|
||||
before the prompt said that" is drift.
|
||||
- **Never audit the guidance's run citations here.** Guidance that describes
|
||||
the observed runs at all — their count, scores, or behaviors — is
|
||||
`detector-rubric-generality`'s flag, whether the citations are stale or
|
||||
current. This shape joins the runs against the prompt and rubric, not
|
||||
against the guidance's prose about runs.
|
||||
- **Sibling reports under `detectors/` are out of scope** — they are
|
||||
regenerated downstream, so staleness there is self-healing. Note it in one
|
||||
sentence if you see it; don't fire on it.
|
||||
|
||||
Guards against overfiring:
|
||||
|
||||
- **Regrading is the fix, not the bug.** A run regraded against the final
|
||||
rubric legitimately pairs an older transcript with a current `grade.md` —
|
||||
that is exactly the remediation this shape's findings prescribe. Never fire
|
||||
merely because a transcript predates the rubric; fire only when the grade's
|
||||
*mechanism* isn't in the shipped rubric, or the transcript's recorded
|
||||
prompt itself diverges from the shipped one.
|
||||
- **Post-run copy edits are normal.** Workers are encouraged to polish rubric
|
||||
wording after grading, and graders paraphrase rather than quote. Anchor on
|
||||
named mechanisms and numbers (gate conditions, penalty sizes, tier
|
||||
boundaries), which survive paraphrase — never require verbatim matches,
|
||||
and fire only on structural divergence.
|
||||
- **Harness framing isn't drift.** A run-recorded prompt may wrap a verbatim
|
||||
`instruction.md` in preamble or formatting; require substantive content
|
||||
divergence before firing.
|
||||
- **A single cosmetic lag is `partial`.** One reward off by a rounding step,
|
||||
wording lag with no scoring consequence — real, absorbable, worth a
|
||||
sentence, not the verdict.
|
||||
|
||||
When firing, name the smallest coherent fix aimed at the join that failed:
|
||||
regenerate runs against the shipped prompt, regrade against the shipped
|
||||
rubric, re-copy the runs so each `reward.txt` matches its `grade.md`, or
|
||||
rebuild and re-upload the archive.
|
||||
|
||||
If more than one shape is present (env breakage, corrupted runs, premise
|
||||
mismatch, package drift), verdict whichever defect most invalidates the
|
||||
submission's evidence and name the others in the Rationale.
|
||||
|
||||
## Verdict definitions
|
||||
|
||||
- **`not-applicable`** — there's no way to decide from this submission. Two
|
||||
triggers:
|
||||
- **No runnable environment in play**: the task is pure static analysis,
|
||||
code review, or technical writing — the agent is never expected to build,
|
||||
run, or test anything, so there is no dev environment that could be broken.
|
||||
`instruction.md` asks only for prose/analysis and the reference runs show no
|
||||
build/test/run attempts. **The premise check still applies here**: a
|
||||
review/audit prompt can assert workspace state ("review my uncommitted
|
||||
change") that the shipped tree contradicts. Only conclude `not-applicable`
|
||||
when the prompt and snapshot also make no workspace-checkable assertions.
|
||||
- **No evidence available**: there are no reference runs (or empty ones) AND
|
||||
the resolved guidance file gives no signal about the environment, so there's
|
||||
nothing to ground a breakage call on. Re-run once reference runs land.
|
||||
Runs that **exist but are infrastructure-broken are not an evidence gap** —
|
||||
that is `runs-corrupted`, a defect, not `not-applicable`.
|
||||
- **`incidental-breakage`** — clear evidence (Shape 1, 2, or 3) that the local
|
||||
dev environment is broken in a way **unrelated to the prompt's ask**, AND the
|
||||
breakage materially distorted the task's signal: a run was blocked or
|
||||
aborted, a substantial share of agent effort went to the breakage, or what
|
||||
the grader saw (or the score) plausibly changed. The prompt does not ask the
|
||||
agent to fix the environment. This is the verdict we do not want a task to
|
||||
earn.
|
||||
- **`runs-corrupted`** — at least one *scored, packaged* reference run never
|
||||
reached a gradable state because of infrastructure: killed mid-work by an
|
||||
API/model error, stranded in an unapprovable plan-mode ending with no shipped
|
||||
work, truncated so the grader scored a transcript the agent didn't produce,
|
||||
missing its output snapshot, or a verifier timeout counted as a run. The
|
||||
workspace may be perfectly healthy — this verdict is about the run set, not
|
||||
the env. List every affected run id and the signature found.
|
||||
- **`premise-mismatch`** — at least one load-bearing premise the prompt or
|
||||
snapshot asserts about the workspace does not hold in the shipped state, AND
|
||||
the resolved guidance file assumes the premise holds (or nowhere credits
|
||||
surfacing the discrepancy). The task as graded cannot exercise what it
|
||||
describes. The environment may build and test perfectly — this verdict is
|
||||
about the workspace being *wrong for the task*, not broken. Quote the
|
||||
premise and the contradicting workspace evidence.
|
||||
- **`package-drift`** — at least one load-bearing revision disagreement
|
||||
between shipped artifacts: the archive is a pre-feedback revision
|
||||
re-uploaded wholesale, a run's recorded prompt substantively diverges from
|
||||
the shipped `instruction.md` (scaffold placeholder text included), grades
|
||||
apply a scoring mechanism the shipped rubric does not contain, `reward.txt`
|
||||
systematically disagrees with `grade.md`, or runs carry conflicting
|
||||
recorded task checksums. Each artifact may be individually healthy — this
|
||||
verdict is about the package's parts describing different revisions of the
|
||||
task. Quote the divergent strings from both sides of the join and name the
|
||||
smallest coherent fix.
|
||||
- **`partial`** — breakage or friction is present but did not materially
|
||||
distort the task's signal: a pre-existing unrelated failure the agents
|
||||
confirm and route around, a single flaky retry, a one-line config nudge, or
|
||||
infrastructure noise that only arrived after a run's substance had landed.
|
||||
A genuine defect the task would be better without — worth naming so the
|
||||
author can smooth it — but no run was blocked and the grade was unaffected.
|
||||
If the friction plausibly changed how the agent spent its effort or what the
|
||||
grader saw, escalate to `incidental-breakage`; if it's cosmetic, lean
|
||||
`clean`. Also the verdict for a **weak or peripheral premise contradiction**:
|
||||
the referenced data exists but is thinner than implied, the "new" feature
|
||||
exists in a clearly-incomplete form the prompt could plausibly mean to
|
||||
extend, or the mismatch is real but peripheral to what the rubric grades.
|
||||
And for **cosmetic revision lag**: grades paraphrasing rubric wording that
|
||||
was later lightly copy-edited, a single reward off by a rounding step —
|
||||
divergence that would not change a score or mislead a reviewer.
|
||||
- **`intentional`** — the environment breakage IS the subject of the task. The
|
||||
prompt explicitly asks the agent to diagnose, fix, or repair the environment,
|
||||
build, dependencies, tooling, or failing setup. The breakage is the point, so
|
||||
it is not a hazard and not flagged. (The premise-mismatch analog — a
|
||||
deliberately false premise the guidance grades surfacing — maps to `clean`,
|
||||
not `intentional`; say so in the Rationale.)
|
||||
- **`clean`** — no evidence the dev environment is incidentally broken,
|
||||
every workspace-checkable premise in the prompt and snapshot holds in the
|
||||
shipped state (or is deliberately false with the rubric grading its
|
||||
discovery), and the checkable artifacts agree on one revision of the task.
|
||||
The agent operated against a working baseline (or the task
|
||||
depends on one and nothing in the runs or rubric shows unrelated env
|
||||
friction). Tests failing because of the agent's own in-progress work, or
|
||||
because the prompt's bug is the subject, are `clean`, not breakage.
|
||||
|
||||
## Confidence
|
||||
|
||||
- **HIGH** — grounding is unambiguous. The reference runs (or the rubric) show
|
||||
the agent fighting a broken setup that the prompt plainly didn't ask about;
|
||||
for `runs-corrupted`, a run's terminal events (or its grade) show the
|
||||
infrastructure failure verbatim; for `premise-mismatch`, the check is
|
||||
mechanical (an empty `git diff HEAD` against a promised uncommitted change,
|
||||
a fully-shipped implementation against a "build X" ask) and the rubric
|
||||
plainly assumes the premise; for `package-drift`, the divergence is
|
||||
quotable from both sides of the join (the scaffold text in the run's
|
||||
recorded prompt, cap language in every grade while the shipped rubric has
|
||||
none); or, for `intentional`, the prompt explicitly
|
||||
asks to fix the environment.
|
||||
- **MEDIUM** — the pattern is present but interpretation is debatable. A
|
||||
reasonable reviewer might read the friction as ordinary task difficulty.
|
||||
- **LOW** — limited information; verdict is a best guess (often because the
|
||||
reference runs are thin or the rubric is silent on the environment).
|
||||
|
||||
## Patterns to look for
|
||||
|
||||
In the reference runs:
|
||||
|
||||
- **Install / build / start failures early in the run**, followed by the agent
|
||||
patching things the prompt never mentioned, just to get going.
|
||||
- **The agent retrying the test suite**, or reasoning aloud about which
|
||||
pre-existing failures are "expected" vs. caused by its change.
|
||||
- **Answer prose that complains about or footnotes the environment** — "note the
|
||||
suite had unrelated failures", "I couldn't run X so I worked around it."
|
||||
- **Time/turns spent on tooling unrelated to the deliverable** — a large share
|
||||
of the run going to environment repair rather than the actual ask.
|
||||
- **Terminal events that are infrastructure, not behavior** — the last event is
|
||||
an API/model error, an errored plan-mode exit with nothing after it, a tool
|
||||
call with no result, or the trajectory just stops; the output snapshot is
|
||||
missing → `runs-corrupted` territory.
|
||||
|
||||
In the shipped environment (statically — even when the runs look quiet):
|
||||
|
||||
- **A capability the prompt or rubric requires that the image can't provide** —
|
||||
a browser the rubric expects verification in that was never installed, a
|
||||
package the deliverable imports that is absent from every manifest and
|
||||
lockfile, a tool that can only be installed from the network. Check the
|
||||
Dockerfile and lockfiles against what the ask and its verification assume.
|
||||
|
||||
In the workspace, checked against the prompt and snapshot (the premise check):
|
||||
|
||||
- **A promised pending change that isn't pending** — the prompt says "review my
|
||||
uncommitted change" and `git status` / `git diff HEAD` come back clean.
|
||||
- **The ask already delivered** — the prompt asks to build/add/first-pass a
|
||||
capability and the workspace (including `workspace.patch`) ships it
|
||||
substantially complete, with tests or docs presenting it as done.
|
||||
- **Prior-turn state that didn't survive** — the snapshot's turns fixed (or
|
||||
broke) something the shipped tree doesn't reflect.
|
||||
- **Referenced data or files absent** — seeds create zero rows of the data the
|
||||
prompt says is "already there"; a named branch/file doesn't exist, and no
|
||||
bring-up step creates it.
|
||||
- **Run corroboration** — agents reporting an empty diff or "this already
|
||||
exists," burning turns on workarounds for state that isn't there, grades
|
||||
improvising anchors.
|
||||
|
||||
Across the shipped artifacts (the version-coherence check):
|
||||
|
||||
- **Grades invoking a mechanism the shipped rubric lacks** — cap/gate
|
||||
language, tier names, or penalty magnitudes absent from
|
||||
the resolved guidance file.
|
||||
- **A run-recorded prompt that isn't the shipped prompt** — scaffold
|
||||
placeholder text, or a substantively different ask.
|
||||
- **`reward.txt` disagreeing with `grade.md` across the run set** — the
|
||||
grades were revised and the scores never re-copied.
|
||||
- **Conflicting recorded task checksums across runs**, or the prior round's
|
||||
feedback still visibly unaddressed in a resubmitted archive.
|
||||
|
||||
In `instruction.md` (to separate intentional from incidental):
|
||||
|
||||
- Asks to **fix / repair / debug the env, build, deps, or failing setup** →
|
||||
lean `intentional`.
|
||||
- Asks for a **feature, audit, trace, design, or fix to specific app behavior**,
|
||||
with breakage showing up anyway → lean `incidental-breakage`.
|
||||
|
||||
In the resolved guidance file:
|
||||
|
||||
- Instructions to the grader to **discount, ignore, or expect** environment
|
||||
failures the agent shouldn't be blamed for → the env is broken and the rubric
|
||||
is papering over it (incidental).
|
||||
|
||||
## What you are NOT doing
|
||||
|
||||
- **Not flagging a task whose subject is the broken environment** — that's
|
||||
`intentional`. Read `instruction.md` before deciding.
|
||||
- **Not flagging legitimate red tests** — the agent's own in-progress work, or a
|
||||
failing test the prompt asks the agent to fix, is the task working.
|
||||
- **Not flagging every imperfect run as corrupted** — `runs-corrupted` requires
|
||||
an infrastructure event at the run's terminal state, not a low score, a terse
|
||||
ending, or a deliberate stop-and-ask the rubric credits.
|
||||
- **Not flagging a deliberately false premise the rubric grades.** A task built
|
||||
around a wrong user belief, where the guidance knows the true workspace state
|
||||
and credits the agent for surfacing it, is a valid design — the
|
||||
premise-mismatch shape fires only when the rubric assumes the premise holds.
|
||||
- **Not flagging wording drift between rubric and grades.** Paraphrase is
|
||||
normal; structure and numbers are the signal. And an older-but-regraded run
|
||||
is the prescribed remediation, not drift — check the transcript's recorded
|
||||
prompt, not its age.
|
||||
- **Not auditing the guidance's descriptions of runs.** Run-anchored guidance
|
||||
— stale or current — is `detector-rubric-generality`'s lane; the
|
||||
package-drift shape joins the runs against the prompt and rubric only.
|
||||
- **Not grading the agent's submission** or re-deriving any other detector's
|
||||
call. This detector is solely about whether the environment is incidentally
|
||||
broken and worked-around, whether the scored runs are valid samples of
|
||||
agent behavior, whether the workspace matches the task's stated premise,
|
||||
and whether the shipped artifacts agree on one revision of the task.
|
||||
|
||||
## Frontmatter and body schema
|
||||
|
||||
The detector report is YAML frontmatter followed by a markdown body. Both
|
||||
contexts produce the same shape; only the *sink* differs (the wrapping
|
||||
`SKILL.md` tells you where to send the report).
|
||||
|
||||
**Frontmatter** — exactly these keys, exactly these enum values:
|
||||
|
||||
```yaml
|
||||
---
|
||||
detector: detector-broken-dev-env
|
||||
verdict: incidental-breakage | runs-corrupted | premise-mismatch | package-drift | partial | intentional | clean | not-applicable
|
||||
confidence: HIGH | MEDIUM | LOW
|
||||
---
|
||||
```
|
||||
|
||||
**Body sections**, in this order:
|
||||
|
||||
```markdown
|
||||
# Broken-dev-env check: <slug>
|
||||
|
||||
## Verbatim grounding
|
||||
|
||||
Pull the load-bearing quotes that justify the verdict. Quote them inline as
|
||||
blockquotes — don't paraphrase. For `incidental-breakage` / `partial`: quote the
|
||||
reference-run text (or rubric line) that shows the agent hitting / working around
|
||||
the broken environment, AND quote the part of `instruction.md` that shows the
|
||||
prompt did NOT ask for it. For `runs-corrupted`: quote the terminal trajectory
|
||||
events (or the `grade.md` text) that show the infrastructure failure — the
|
||||
error string, the truncation point, the grader tripping over the missing
|
||||
ending — and name each affected run id. For `premise-mismatch`: quote the
|
||||
premise verbatim from `instruction.md` or the snapshot AND the workspace
|
||||
evidence contradicting it (the `git diff HEAD` output, the file/commit that
|
||||
already ships the ask, the empty seed, the missing prior-turn state), plus the
|
||||
guidance line showing the rubric assumes the premise holds. For
|
||||
`package-drift`: quote the exact divergent strings from **both sides** of the
|
||||
join — the run-recorded prompt line next to the shipped `instruction.md`
|
||||
line, the grade's cap/penalty language next to the shipped rubric's
|
||||
mechanism, the `reward.txt` value next to the `grade.md` score line, the
|
||||
conflicting checksums — and name each affected run id. A drift call asserted
|
||||
without paired quotes is unreviewable. For `intentional`: quote the part of
|
||||
`instruction.md` that asks the agent to fix the environment. For `clean`: quote
|
||||
what the runs / rubric DO show (a working baseline, or task-intrinsic red
|
||||
tests) so the reader can confirm. For `not-applicable`: quote the artifact
|
||||
showing the trigger (the prompt asking only for prose, or the missing
|
||||
reference runs).
|
||||
|
||||
## Rationale
|
||||
|
||||
2–4 paragraphs explaining what is broken (or why nothing is), tied to the
|
||||
grounding above. Be specific: which shape (1/2/3, the fourth runs-corrupted,
|
||||
the fifth premise-mismatch, or the sixth package-drift)? Which run shows the
|
||||
workaround — or, for
|
||||
`runs-corrupted`, which runs are invalid and whether each reached a gradable
|
||||
state before the infrastructure event — or, for `premise-mismatch`, which
|
||||
premise type failed and whether the rubric assumes it holds or credits its
|
||||
discovery — or, for `package-drift`, which join failed, whether the
|
||||
divergence is load-bearing or cosmetic, and the smallest coherent fix
|
||||
(regenerate runs, regrade, re-copy rewards, or rebuild and re-upload)? Why is
|
||||
the breakage unrelated to the prompt's ask (or,
|
||||
for `intentional`, why it IS the ask)? For `not-applicable`, explain which
|
||||
trigger fired and what would make the detector runnable.
|
||||
```
|
||||
|
||||
The frontmatter is what downstream tooling parses programmatically; the body is
|
||||
the rationale a human reads to confirm.
|
||||
Reference in New Issue
Block a user