Files
Eric Bell 2854619bc9 chore: init commit
in worker.../repo/GITFOLDER.zip is the .git folder.
2026-08-11 14:44:09 -04:00

3.7 KiB

name, description, allowed-tools
name description allowed-tools
detector-good-response-exhaustiveness Self-check whether your grader guidance covers all the *plausible* types of strong response — the big-picture approaches a broad majority (~80%) of SWEs would consider reasonable for your prompt — or whether it only credits a subset, so an agent taking a reasonable-but-uncredited approach gets unfairly marked down — or sweeps a legitimate shape into a penalty aimed at something else (honest disclosure of incomplete work; an approach a reference run actually took), run-evidenced only. Flagship case: the clarify-vs-act fork — when a prompt has a real ambiguity, both "flag the issue and ask" and "flag the issue, state an assumption, act, and report" are usually legitimate, and the rubric should credit both. Not about crazy exhaustiveness — just the major forks (clarify-vs-act, build-vs-buy, assess-vs-fix, defer-vs-push-back). Reads instruction.md + the grader guidance file that `bash scripts/guidance-target.sh <slug>` resolves (reference runs optional, except penalty-side findings which require them). Bash, Read, Write

Good-response-exhaustiveness detector

This skill checks whether your grader guidance credits all the plausible ways a competent SWE could respond well to your prompt — not just your preferred path.

For many prompts there's more than one legitimate strong answer. The flagship case is the clarify-vs-act fork: when your prompt has a real ambiguity or decision point, both

  1. "flag the issue and ask for clarification," and
  2. "flag the issue, make a reasonable assumption (state it), act, and report what you did,"

are usually legitimate. If ~80% of SWEs would accept both, your rubric should credit both — otherwise an agent that takes the uncredited path gets marked down for picking a reasonable approach you happened not to list.

Other big forks to check: build-vs-buy / extend-vs-replace (design prompts), assess vs. answer-plus-fix (question prompts), and defer vs. push back (when the prompt frames a decision as already made).

The bar is not crazy exhaustiveness — just the few big-picture approaches a broad majority of SWEs would agree are reasonable. A favorite/A+ approach plus acceptable alternatives is great; the problem is excluding a reasonable one (often via a one-sided heavy penalty — "heavily penalize unless the agent asks," which dings a reasonable act-on-assumption answer, or vice versa).

Read these before deciding:

  1. .claude/skills/_detector-worker-shell.md — where to write the report and how to handle re-runs.
  2. .claude/skills/detector-good-response-exhaustiveness/core.md — the big forks, the ~80% bar, what is NOT a gap, the boundaries against detector-good-response-defined and detector-answer-obviousness, verdict enums.

Compose the report per the schema in core.md and write it per _detector-worker-shell.md.

Acting on the verdict

  • exhaustive — your rubric credits the big-picture plausible approaches (or the prompt has one reasonable shape and you cover it). Good. Move on.
  • partial — you cover the main approach but miss a secondary plausible one. Read the per-approach assessment; add a tier/criterion that credits it.
  • has-gaps — you're missing a major reasonable approach (often one side of the clarify-vs-act fork, or a build-vs-buy alternative). Add explicit credit for it — e.g. "a strong response either asks for clarification on ABC, or states the assumption that XYZ and proceeds, then reports it" — and relax any heavy penalty that forces one side of a legitimate fork. Re-run after.
  • not-applicable — no grader guidance to assess yet. Draft it first.