lots of change - all to start my 3rd redo
This commit is contained in:
@@ -1,95 +0,0 @@
|
||||
---
|
||||
name: detector-broken-dev-env
|
||||
description: |
|
||||
Self-check whether your task presents an *incidentally* broken local dev
|
||||
environment that the test agent has to awkwardly work around. The workspace
|
||||
doesn't build/install/run, a dependency or service is missing, or the test
|
||||
suite has pre-existing failures or flakes unrelated to your task — and the
|
||||
agent burns effort coping with that instead of doing what your prompt asks.
|
||||
These tasks are weak: the existing test suite is the main verifier, so env
|
||||
noise lands straight in the grade. The one allowed shape is intentional
|
||||
breakage — a task whose subject IS the broken env ("my dev env is broken, fix
|
||||
it"). A pre-existing app bug that your prompt asks the agent to find or fix is
|
||||
the task working, not breakage. Also checks that the workspace is actually in
|
||||
the state your prompt (or snapshot) says it's in — a promised uncommitted
|
||||
change that's already committed, a "build X" ask where X already ships, or
|
||||
referenced data that isn't there is a premise mismatch even when everything
|
||||
builds green. And checks that everything you package reflects the same
|
||||
revision of your task — runs graded under an earlier prompt or rubric, a
|
||||
reward.txt that no longer matches its grade.md, or a re-uploaded older
|
||||
archive is package drift even when every artifact is individually healthy.
|
||||
allowed-tools: Bash, Read, Write
|
||||
---
|
||||
|
||||
# Broken-dev-env detector
|
||||
|
||||
This skill checks whether your task hands the agent a dev environment that is
|
||||
broken for reasons unrelated to what you're asking it to do — a build that won't
|
||||
run, a missing dependency, or a test suite with pre-existing failures the prompt
|
||||
never mentions. If the agent has to fight that breakage to make progress, the
|
||||
task is testing "can the agent cope with a broken env" instead of the behavior
|
||||
you meant to grade, and the verifier signal gets noisy. It also checks that
|
||||
your scored reference runs are valid samples of agent behavior — a run killed
|
||||
mid-work by an API error, truncated, or missing its output snapshot reflects
|
||||
infrastructure, not the agent, and shouldn't ship as evidence. And it checks
|
||||
that the workspace matches your task's stated premise: if your prompt or
|
||||
snapshot asserts something about the workspace ("review my uncommitted
|
||||
change", "there's already data in the repo", a previous turn's fix) that the
|
||||
shipped state contradicts, the agent responds to the workspace as it actually
|
||||
is and your rubric grades a task that can't happen. Finally, it checks that
|
||||
your package is one coherent revision of the task: if you polish the prompt or
|
||||
rubric after generating runs, the shipped runs and grades must be regenerated
|
||||
or regraded to match — a package whose parts describe different versions of
|
||||
the task can't evidence it.
|
||||
|
||||
Read these before deciding:
|
||||
|
||||
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
|
||||
2. `.claude/skills/detector-broken-dev-env/core.md` — intentional-vs-incidental, the three shapes breakage takes plus the runs-corrupted, premise-mismatch, and package-drift shapes, what counts as legitimate task difficulty, verdict enums.
|
||||
|
||||
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
|
||||
|
||||
## Acting on the verdict
|
||||
|
||||
- **`clean`** — your environment runs fine (or the only red tests are the bug
|
||||
your prompt is about). Good.
|
||||
- **`partial`** — there's some env friction, but it's minor or borderline. Read
|
||||
the rationale; either smooth the setup so the agent never hits it, or confirm
|
||||
it's cosmetic enough not to distort the run.
|
||||
- **`incidental-breakage`** — the agent has to work around a broken setup your
|
||||
prompt didn't ask it to fix. Fix the environment (repair the Dockerfile,
|
||||
pin deps, remove the unrelated failing/flaky tests) so the agent starts from a
|
||||
working baseline, then re-run this skill. Don't try to rescue it by reframing
|
||||
the breakage as the task — see `intentional`.
|
||||
- **`runs-corrupted`** — one or more of your scored reference runs was ended or
|
||||
distorted by infrastructure rather than by the agent (an API/model error
|
||||
mid-run, a truncated trajectory, a missing output snapshot, a verifier
|
||||
timeout), so it isn't a valid sample of agent behavior. Your workspace may be
|
||||
perfectly healthy. Re-run the affected trials, replace the corrupted runs,
|
||||
repackage, then re-run this skill.
|
||||
- **`premise-mismatch`** — the shipped workspace contradicts what your prompt
|
||||
or snapshot asserts (the promised uncommitted change is already committed,
|
||||
the feature you ask the agent to build already exists, referenced data is
|
||||
absent, a prior turn's state was reset away), and your holistic rubric
|
||||
assumes the premise holds. Either fix the workspace so the premise is true
|
||||
(workspace.patch, seeds, snapshot end-state), or — if the false premise is
|
||||
deliberate — make the rubric grade the agent on surfacing it, then re-run
|
||||
this skill and re-collect reference runs.
|
||||
- **`package-drift`** — your packaged artifacts don't all reflect the same
|
||||
revision of the task: runs were graded under an earlier prompt or rubric, a
|
||||
reward.txt no longer matches its grade.md, runs record conflicting task
|
||||
versions, or you re-uploaded an older archive after making fixes. Nothing
|
||||
may be broken — but the runs no longer demonstrate the shipped task. Apply
|
||||
the smallest coherent fix: regenerate runs against the current prompt,
|
||||
regrade against the current rubric (see `/regrade-reference-run`), re-copy
|
||||
the runs so each reward.txt matches its grade.md, or rebuild and re-upload
|
||||
the archive — then re-run this skill.
|
||||
- **`intentional`** — your task is explicitly about fixing the environment. That
|
||||
is a valid task; nothing to change. (Only legitimate if your *prompt* asks for
|
||||
the repair — not if the agent merely ended up coping with a broken env.)
|
||||
- **`not-applicable`** — the task has no runnable environment (pure analysis /
|
||||
writing) AND the prompt/snapshot make no workspace-checkable assertions, or
|
||||
there's no evidence yet (no reference runs and the rubric says
|
||||
nothing about the env). Re-run once you have reference runs. (Runs that exist
|
||||
but died on infrastructure are `runs-corrupted`, not this; a prose prompt
|
||||
that asserts workspace state can still earn `premise-mismatch`.)
|
||||
Reference in New Issue
Block a user