Files
project-work/worker-toolkit-stocks-in-the-future/.claude/skills/detector-broken-dev-env/SKILL.md

6.2 KiB

name, description, allowed-tools
name description allowed-tools
detector-broken-dev-env Self-check whether your task presents an *incidentally* broken local dev environment that the test agent has to awkwardly work around. The workspace doesn't build/install/run, a dependency or service is missing, or the test suite has pre-existing failures or flakes unrelated to your task — and the agent burns effort coping with that instead of doing what your prompt asks. These tasks are weak: the existing test suite is the main verifier, so env noise lands straight in the grade. The one allowed shape is intentional breakage — a task whose subject IS the broken env ("my dev env is broken, fix it"). A pre-existing app bug that your prompt asks the agent to find or fix is the task working, not breakage. Also checks that the workspace is actually in the state your prompt (or snapshot) says it's in — a promised uncommitted change that's already committed, a "build X" ask where X already ships, or referenced data that isn't there is a premise mismatch even when everything builds green. And checks that everything you package reflects the same revision of your task — runs graded under an earlier prompt or rubric, a reward.txt that no longer matches its grade.md, or a re-uploaded older archive is package drift even when every artifact is individually healthy. Bash, Read, Write

Broken-dev-env detector

This skill checks whether your task hands the agent a dev environment that is broken for reasons unrelated to what you're asking it to do — a build that won't run, a missing dependency, or a test suite with pre-existing failures the prompt never mentions. If the agent has to fight that breakage to make progress, the task is testing "can the agent cope with a broken env" instead of the behavior you meant to grade, and the verifier signal gets noisy. It also checks that your scored reference runs are valid samples of agent behavior — a run killed mid-work by an API error, truncated, or missing its output snapshot reflects infrastructure, not the agent, and shouldn't ship as evidence. And it checks that the workspace matches your task's stated premise: if your prompt or snapshot asserts something about the workspace ("review my uncommitted change", "there's already data in the repo", a previous turn's fix) that the shipped state contradicts, the agent responds to the workspace as it actually is and your rubric grades a task that can't happen. Finally, it checks that your package is one coherent revision of the task: if you polish the prompt or rubric after generating runs, the shipped runs and grades must be regenerated or regraded to match — a package whose parts describe different versions of the task can't evidence it.

Read these before deciding:

  1. .claude/skills/_detector-worker-shell.md — where to write the report and how to handle re-runs.
  2. .claude/skills/detector-broken-dev-env/core.md — intentional-vs-incidental, the three shapes breakage takes plus the runs-corrupted, premise-mismatch, and package-drift shapes, what counts as legitimate task difficulty, verdict enums.

Compose the report per the schema in core.md and write it per _detector-worker-shell.md.

Acting on the verdict

  • clean — your environment runs fine (or the only red tests are the bug your prompt is about). Good.
  • partial — there's some env friction, but it's minor or borderline. Read the rationale; either smooth the setup so the agent never hits it, or confirm it's cosmetic enough not to distort the run.
  • incidental-breakage — the agent has to work around a broken setup your prompt didn't ask it to fix. Fix the environment (repair the Dockerfile, pin deps, remove the unrelated failing/flaky tests) so the agent starts from a working baseline, then re-run this skill. Don't try to rescue it by reframing the breakage as the task — see intentional.
  • runs-corrupted — one or more of your scored reference runs was ended or distorted by infrastructure rather than by the agent (an API/model error mid-run, a truncated trajectory, a missing output snapshot, a verifier timeout), so it isn't a valid sample of agent behavior. Your workspace may be perfectly healthy. Re-run the affected trials, replace the corrupted runs, repackage, then re-run this skill.
  • premise-mismatch — the shipped workspace contradicts what your prompt or snapshot asserts (the promised uncommitted change is already committed, the feature you ask the agent to build already exists, referenced data is absent, a prior turn's state was reset away), and your grader guidance assumes the premise holds. Either fix the workspace so the premise is true (workspace.patch, seeds, snapshot end-state), or — if the false premise is deliberate — make the guidance grade the agent on surfacing it, then re-run this skill and re-collect reference runs.
  • package-drift — your packaged artifacts don't all reflect the same revision of the task: runs were graded under an earlier prompt or rubric, a reward.txt no longer matches its grade.md, runs record conflicting task versions, or you re-uploaded an older archive after making fixes. Nothing may be broken — but the runs no longer demonstrate the shipped task. Apply the smallest coherent fix: regenerate runs against the current prompt, regrade against the current rubric (see /regrade-reference-run), re-copy the runs so each reward.txt matches its grade.md, or rebuild and re-upload the archive — then re-run this skill.
  • intentional — your task is explicitly about fixing the environment. That is a valid task; nothing to change. (Only legitimate if your prompt asks for the repair — not if the agent merely ended up coping with a broken env.)
  • not-applicable — the task has no runnable environment (pure analysis / writing) AND the prompt/snapshot make no workspace-checkable assertions, or there's no evidence yet (no reference runs and the rubric says nothing about the env). Re-run once you have reference runs. (Runs that exist but died on infrastructure are runs-corrupted, not this; a prose prompt that asserts workspace state can still earn premise-mismatch.)