added 260907 version of worker toolkit

This commit is contained in:
2026-09-08 20:30:51 -04:00
parent cbd3f0f8ca
commit 97aca37663
1283 changed files with 142951 additions and 0 deletions

View File

@@ -0,0 +1,95 @@
---
name: detector-broken-dev-env
description: |
Self-check whether your task presents an *incidentally* broken local dev
environment that the test agent has to awkwardly work around. The workspace
doesn't build/install/run, a dependency or service is missing, or the test
suite has pre-existing failures or flakes unrelated to your task — and the
agent burns effort coping with that instead of doing what your prompt asks.
These tasks are weak: the existing test suite is the main verifier, so env
noise lands straight in the grade. The one allowed shape is intentional
breakage — a task whose subject IS the broken env ("my dev env is broken, fix
it"). A pre-existing app bug that your prompt asks the agent to find or fix is
the task working, not breakage. Also checks that the workspace is actually in
the state your prompt (or snapshot) says it's in — a promised uncommitted
change that's already committed, a "build X" ask where X already ships, or
referenced data that isn't there is a premise mismatch even when everything
builds green. And checks that everything you package reflects the same
revision of your task — runs graded under an earlier prompt or rubric, a
reward.txt that no longer matches its grade.md, or a re-uploaded older
archive is package drift even when every artifact is individually healthy.
allowed-tools: Bash, Read, Write
---
# Broken-dev-env detector
This skill checks whether your task hands the agent a dev environment that is
broken for reasons unrelated to what you're asking it to do — a build that won't
run, a missing dependency, or a test suite with pre-existing failures the prompt
never mentions. If the agent has to fight that breakage to make progress, the
task is testing "can the agent cope with a broken env" instead of the behavior
you meant to grade, and the verifier signal gets noisy. It also checks that
your scored reference runs are valid samples of agent behavior — a run killed
mid-work by an API error, truncated, or missing its output snapshot reflects
infrastructure, not the agent, and shouldn't ship as evidence. And it checks
that the workspace matches your task's stated premise: if your prompt or
snapshot asserts something about the workspace ("review my uncommitted
change", "there's already data in the repo", a previous turn's fix) that the
shipped state contradicts, the agent responds to the workspace as it actually
is and your rubric grades a task that can't happen. Finally, it checks that
your package is one coherent revision of the task: if you polish the prompt or
rubric after generating runs, the shipped runs and grades must be regenerated
or regraded to match — a package whose parts describe different versions of
the task can't evidence it.
Read these before deciding:
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
2. `.claude/skills/detector-broken-dev-env/core.md` — intentional-vs-incidental, the three shapes breakage takes plus the runs-corrupted, premise-mismatch, and package-drift shapes, what counts as legitimate task difficulty, verdict enums.
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
## Acting on the verdict
- **`clean`** — your environment runs fine (or the only red tests are the bug
your prompt is about). Good.
- **`partial`** — there's some env friction, but it's minor or borderline. Read
the rationale; either smooth the setup so the agent never hits it, or confirm
it's cosmetic enough not to distort the run.
- **`incidental-breakage`** — the agent has to work around a broken setup your
prompt didn't ask it to fix. Fix the environment (repair the Dockerfile,
pin deps, remove the unrelated failing/flaky tests) so the agent starts from a
working baseline, then re-run this skill. Don't try to rescue it by reframing
the breakage as the task — see `intentional`.
- **`runs-corrupted`** — one or more of your scored reference runs was ended or
distorted by infrastructure rather than by the agent (an API/model error
mid-run, a truncated trajectory, a missing output snapshot, a verifier
timeout), so it isn't a valid sample of agent behavior. Your workspace may be
perfectly healthy. Re-run the affected trials, replace the corrupted runs,
repackage, then re-run this skill.
- **`premise-mismatch`** — the shipped workspace contradicts what your prompt
or snapshot asserts (the promised uncommitted change is already committed,
the feature you ask the agent to build already exists, referenced data is
absent, a prior turn's state was reset away), and your holistic rubric
assumes the premise holds. Either fix the workspace so the premise is true
(workspace.patch, seeds, snapshot end-state), or — if the false premise is
deliberate — make the rubric grade the agent on surfacing it, then re-run
this skill and re-collect reference runs.
- **`package-drift`** — your packaged artifacts don't all reflect the same
revision of the task: runs were graded under an earlier prompt or rubric, a
reward.txt no longer matches its grade.md, runs record conflicting task
versions, or you re-uploaded an older archive after making fixes. Nothing
may be broken — but the runs no longer demonstrate the shipped task. Apply
the smallest coherent fix: regenerate runs against the current prompt,
regrade against the current rubric (see `/regrade-reference-run`), re-copy
the runs so each reward.txt matches its grade.md, or rebuild and re-upload
the archive — then re-run this skill.
- **`intentional`** — your task is explicitly about fixing the environment. That
is a valid task; nothing to change. (Only legitimate if your *prompt* asks for
the repair — not if the agent merely ended up coping with a broken env.)
- **`not-applicable`** — the task has no runnable environment (pure analysis /
writing) AND the prompt/snapshot make no workspace-checkable assertions, or
there's no evidence yet (no reference runs and the rubric says
nothing about the env). Re-run once you have reference runs. (Runs that exist
but died on infrastructure are `runs-corrupted`, not this; a prose prompt
that asserts workspace state can still earn `premise-mismatch`.)