lots of change - all to start my 3rd redo

This commit is contained in:
2026-09-26 14:31:52 -04:00
parent 7f4d388e19
commit bceb52e8ee
1046 changed files with 4476 additions and 0 deletions

View File

@@ -1,66 +0,0 @@
---
name: detector-good-response-exhaustiveness
description: |
Self-check whether your holistic rubric covers all the *plausible* types of
strong response — the big-picture approaches a broad majority (~80%) of SWEs would
consider reasonable for your prompt — or whether it only credits a subset, so an agent
taking a reasonable-but-uncredited approach gets unfairly marked down — or sweeps a
legitimate shape into a penalty aimed at something else (honest disclosure of incomplete
work; an approach a reference run actually took), run-evidenced only. Flagship case:
the clarify-vs-act fork — when a prompt has a real ambiguity, both "flag the issue and
ask" and "flag the issue, state an assumption, act, and report" are usually legitimate,
and the rubric should credit both. Not about crazy exhaustiveness — just the major forks
(clarify-vs-act, build-vs-buy, assess-vs-fix, defer-vs-push-back). Reads instruction.md
+ the holistic rubric file that `bash scripts/guidance-target.sh <slug>` resolves
(reference runs optional, except penalty-side findings which
require them).
allowed-tools: Bash, Read, Write
---
# Good-response-exhaustiveness detector
This skill checks whether your holistic rubric credits **all the plausible
ways a competent SWE could respond well** to your prompt — not just your
preferred path.
For many prompts there's more than one legitimate strong answer. The flagship
case is the **clarify-vs-act fork**: when your prompt has a real ambiguity or
decision point, both
1. "flag the issue and **ask for clarification**," and
2. "flag the issue, **make a reasonable assumption (state it), act, and report
what you did**,"
are usually legitimate. If ~80% of SWEs would accept both, your rubric should
credit both — otherwise an agent that takes the uncredited path gets marked
down for picking a reasonable approach you happened not to list.
Other big forks to check: **build-vs-buy / extend-vs-replace** (design
prompts), **assess vs. answer-plus-fix** (question prompts), and **defer vs.
push back** (when the prompt frames a decision as already made).
The bar is **not** crazy exhaustiveness — just the few big-picture approaches a
broad majority of SWEs would agree are reasonable. A favorite/A+ approach plus
acceptable alternatives is great; the problem is *excluding* a reasonable one
(often via a one-sided heavy penalty — "heavily penalize unless the agent asks,"
which dings a reasonable act-on-assumption answer, or vice versa).
Read these before deciding:
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
2. `.claude/skills/detector-good-response-exhaustiveness/core.md` — the big forks, the ~80% bar, what is NOT a gap, the boundaries against detector-good-response-defined and detector-answer-obviousness, verdict enums.
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
## Acting on the verdict
- **`exhaustive`** — your rubric credits the big-picture plausible approaches
(or the prompt has one reasonable shape and you cover it). Good. Move on.
- **`partial`** — you cover the main approach but miss a secondary plausible
one. Read the per-approach assessment; add a tier/criterion that credits it.
- **`has-gaps`** — you're missing a major reasonable approach (often one side
of the clarify-vs-act fork, or a build-vs-buy alternative). Add explicit
credit for it — e.g. "a strong response either asks for clarification on ABC,
or states the assumption that XYZ and proceeds, then reports it" — and relax
any heavy penalty that forces one side of a legitimate fork. Re-run after.
- **`not-applicable`** — no holistic rubric to assess yet. Draft it first.

View File

@@ -1,361 +0,0 @@
# Good-response-exhaustiveness detector — core
This file is the canonical, context-neutral content for the
detector-good-response-exhaustiveness detector. It defines what the detector looks
for, the verdict enums, the patterns to recognize, and the output schema.
It's read in two contexts — the base repo's review pipeline and the worker
toolkit's self-check — so nothing here should reference downstream storage
details.
## What this detector is for
The grader scores an agent's answer against the rubric's notion of a strong
response. For many prompts there is **more than one legitimate way a
competent SWE would respond**, and the rubric should credit each *plausible*
approach — not just the author's preferred one. This detector asks:
**Does the rubric's set of accepted strong responses cover the big-picture
approaches that a broad majority (~80%) of SWEs would consider reasonable?**
When it doesn't, an agent that takes a perfectly reasonable but uncredited
approach gets unfairly marked down — the task ends up grading "did you pick
the author's path" instead of "did you respond well."
The fairness frame cuts both ways. A legitimate response shape must be
**credited** — the coverage question above — and it must not be **swept into
a penalty** aimed at a different behavior. When the reference runs show a
penalty for overclaiming landing on a run that honestly scoped its claim, or
a penalty written for one implementation catching a run that reasonably took
another, that is the same unfairness from the penalty side: a legitimate
shape scored as a failure. These penalty-side shapes are **run-evidenced
only** — see patterns 8–9.
The flagship case is the **clarify-vs-act fork.** When a prompt has a real
ambiguity or judgment call, two responses are usually both legitimate:
1. **Flag the issue and ask for clarification** before acting, or
2. **Flag the issue, make a reasonable assumption (state it), act on it, and
report what was done.**
If ~80% of SWEs would accept *both*, the rubric should credit both. A rubric
that credits only one — or heavily penalizes the other — has a coverage gap.
The bar is deliberately not "crazy exhaustive": you are looking for the few
**big-picture** approaches most SWEs would agree are reasonable, not every
micro-variation. Cover the major forks, not the long tail.
## What "covering the plausible approaches" means
- The rubric **credits, or at least leaves room for, each major reasonable
approach** — not just the author's pick.
- It's fine to have a *preferred* / A+ approach **plus** acceptable
alternatives at the same or a slightly lower tier. What matters is that a
reasonable approach isn't left **uncredited or penalized**.
- "Credit" can be explicit ("a strong response either asks for clarification
**or** states an assumption and proceeds") or structural (tiers/criteria
that a reasonable alternative could satisfy). What you're checking is
whether a grader, holding a reasonable-but-different answer, would find a
basis to score it well.
## The big forks to check
Walk the prompt and ask which broadly-accepted approaches exist. The common
ones:
1. **Clarify vs. act-on-assumption** — the flagship. Real ambiguity / a
decision the prompt leaves open: both "ask first" and "state an assumption
and proceed, then report" are usually legitimate. Does the rubric credit
both, or does it reward only asking (and ding acting) or only acting (and
ding asking)?
2. **Assess/answer vs. answer-plus-fix** — for a question or assessment
prompt, both "answer what was asked" and "answer + propose a remedy" can
be reasonable. (Note the boundary with detector-answer-obviousness: *requiring* a
fix the prompt didn't ask for is detector-answer-obviousness's unrequested-scope;
here the concern is the mirror image — failing to credit a reasonable
answer-only response, or a reasonable answer-plus-fix response.)
3. **Build vs. buy / extend vs. replace** — for design/architecture prompts,
several approaches are often defensible (keep the in-house system and
extend it; or step back and recommend an off-the-shelf platform). Does the
rubric credit the reasonable alternatives or canonize one?
4. **Defer vs. push back** — when the prompt frames a decision as already
made by the team, both "accept the stated decision and proceed" and
"register a concern" can be reasonable. Does the rubric credit the one it
doesn't prefer?
Coverage gaps are not always fork-shaped. A rubric can credit both sides of
every fork and still leave a plausible response with nowhere to land — see
patterns 4–7 below for the rubric-visible shapes: a strong-response set
limited to a single credited path, a plausible middle/hybrid response that
falls between the credited tier and a penalty, and the two recurring named
shapes (**comply-and-flag** and **do-what-was-asked-without-extras**) that
rubrics most often leave uncredited.
Not every prompt has multiple plausible approaches — a factual question or a
prompt with one obviously-correct design has a single strong shape, and then
exhaustiveness is trivially met. Only flag a gap when there is a **real,
broadly-agreed alternative the rubric omits.**
## What is NOT a gap
- **Niche approaches.** Something only a minority of SWEs would do is not a
required coverage item. The bar is ~80% agreement.
- **A ranked-but-inclusive rubric.** Crediting several shapes and ranking
them (preferred A+ + acceptable alternatives) is *good coverage*, not a
gap. Don't flag a rubric for having a favorite — flag it for excluding a
reasonable approach.
- **Genuinely single-approach prompts.** If there's one reasonable strong
shape, `exhaustive` is the right call.
- **Don't re-litigate adjacent detectors.** Whether the rubric defines good
*at all* is detector-good-response-defined; whether it canonizes a *non-obvious*
answer or demands *unrequested scope* is detector-answer-obviousness. This detector
assumes a positive target exists and asks whether the accepted **set** is
complete.
- **Your own taste.** The test is "would ~80% of SWEs accept this approach,"
not "would I have done it this way." Don't invent alternatives a broad
majority wouldn't actually endorse.
- **Strict-by-design rubrics.** A rubric that explicitly and deliberately
penalizes honest-incomplete work — because completeness itself is the
deliverable the prompt asked for, and it says so — made a design choice,
not a coverage error. Pattern 8 applies only when the rubric frames the
penalty as targeting overclaiming / confidence / honesty.
- **Penalty magnitude.** How big a deduction is, is the author's design call.
The penalty-side shapes are about *which responses* a penalty catches, as
shown by the runs — never "this deduction feels too heavy."
## Inputs
Read whatever you need from the task directory. The load-bearing artifacts:
- `instruction.md` — **read it first.** Establish the plausible strong-response
approaches a competent SWE could reasonably take to *this* prompt
(especially: does it contain a real ambiguity / decision point that opens
the clarify-vs-act fork or a build-vs-buy choice?). This is the baseline the
rubric's coverage is measured against.
- The grader guidance — the rubric. Which approaches does it credit?
Do any tiers / heavy penalties penalize a reasonable approach? Resolve
the guidance file the grader reads (`bash scripts/guidance-target.sh
<slug>` prints its path, `tests/grader-guidance-consolidated.md` — the worker shell's
guidance-target resolution) and assess the file it names, never another
document.
- `reference-runs/<run>/agent-output/answer.md` + `grade.md` — *optional,
supporting evidence.* If a run took a reasonable-but-uncredited approach and
the grader dinged it, that confirms a real gap. Not required. When you do
cite a run, characterize it accurately: a run that lost points to a
one-sided criterion is evidence the gap *exists*, not evidence it doesn't;
and check the run actually did what you say it did before crediting it as
an honest instance of an approach. The one exception to "optional": the
penalty-side shapes (patterns 8–9) exist only as run evidence — without a
`grade.md` showing the penalty landing, they are not findings.
## Verdict definitions
- **`not-applicable`** — the resolved guidance file is missing, empty, or only
the unmodified template scaffold. Emit this and stop.
- **`exhaustive`** — the rubric credits all the big-picture plausible
strong-response approaches for this prompt (or the prompt has a single
reasonable strong shape and the rubric covers it). A grader holding any
~80%-reasonable answer would find a basis to score it on its merits.
- **`partial`** — the rubric covers the primary approach(es) but misses a
*secondary* plausible one that a meaningful minority of SWEs would take.
There's a real coverage gap, but it's not the central fork the task hinges
on.
- **`has-gaps`** — the rubric misses a **major** plausible approach (one
~80% of SWEs would consider reasonable for this prompt) — most often one
side of the clarify-vs-act fork, or a build-vs-buy alternative the prompt
invites. A reasonable response taking the uncredited approach would be
unfairly marked down, so the task is partly grading path-selection rather
than response quality.
## Confidence
- **HIGH** — the missing (or present) approach is clearly something a broad
majority of SWEs would accept; the call isn't a close one.
- **MEDIUM** — whether the omitted approach clears the ~80% bar is genuinely
debatable; a reasonable reviewer might call it niche.
- **LOW** — limited info (terse prompt, unfamiliar domain) makes the set of
plausible approaches hard to enumerate confidently. Best-guess.
## Patterns to look for
1. **Enumerate the prompt's plausible approaches first**, before reading the
rubric — so the rubric doesn't anchor you to only the approaches it
happened to consider. Name the big forks (clarify-vs-act, build-vs-buy,
assess-vs-fix, defer-vs-push-back) that genuinely apply, and the named
shapes from patterns 6–7 where the prompt invites them. Anchoring cuts
both ways: if the rubric *frames* an approach as categorically weak (e.g.
treats clarification-first as a scoping failure), check independently
whether that approach was reasonable for this prompt rather than adopting
the rubric's framing.
2. **Map each onto the rubric.** For each plausible approach, is there a tier
/ criterion / "good response" statement it would satisfy? Or is it
unmentioned / implicitly excluded? When you conclude an approach *is*
credited, verify the text you're pointing at actually credits that
approach — a line crediting an adjacent behavior is not credit for a
different response type.
3. **Check heavy penalties for one-sidedness.** "Heavily penalize unless the agent
asks for clarification" dings a reasonable act-on-assumption answer;
"heavily penalize unless the agent ships a fix" dings a reasonable
clarify-first or assess-only answer. A heavy deduction that forces one side
of a legitimate fork is the classic gap.
4. **Check the A+/"strong response" enumeration for single-track framing** —
one prescribed path when the prompt clearly admits several. This includes
crediting only one failure path or finding when the prompt and rubric
themselves put several on the table. Judge coverage from what the prompt
and rubric surface — whether *additional* uncredited paths exist in the
source repo is detector-fact-check-rubric-claims's verification, not yours.
5. **Construct the plausible middle/hybrid response and find its tier.** When
the rubric describes a credited behavior and a penalized behavior, a real
agent often lands between them — keeps the questioned choice but surfaces
the alternative and asks for confirmation, or ships the workaround while
flagging the breakage. Build that blend and check it maps onto some tier or
criterion. A plausible response that falls between two penalties — or
between the strong tier and a heavy deduction — with no scored home is a
coverage gap.
6. **Check the comply-and-flag response for credit.** When the user
explicitly insists on a risky or questionable action, "do what was asked
and flag the concern in a line" is usually as legitimate as "push back /
guard / decline" — often more so, since the user made the call knowingly.
Rubrics recurringly score this fork asymmetrically: the cautious side
(block, guard, refuse) earns the credit while complying with the explicit
instruction draws a heavy deduction even when the concern was surfaced.
Check that a comply-and-flag response has a scored home. Credit can be
structural — a tier such an answer would satisfy counts; don't demand an
approach-specific sentence.
7. **Check the do-what-was-asked-without-extras response for credit.** When
the literal request is coherent and the workspace supports it (the change
is small and safe, or the asked-for feature already exists and passes
tests), "do exactly what was requested, competently, and report" is often
the plausible majority answer. A rubric that names plain compliance as the
core failure — crediting only responses that first discover and surface a
deeper concern the prompt never raised — leaves that majority shape
uncredited. The deeper-discovery path can still be the A+; the question is
whether competent literal compliance has a tier to land on. (Boundary:
*requiring* the extras is detector-answer-obviousness's unrequested-scope;
the coverage question here is whether plain compliance is credited at
all.)
8. **Check penalties aimed at overclaiming against the runs that disclosed
honestly (run-evidenced only).** The mirror image of the comply-and-flag
credit check: a penalty targeting overclaiming or unearned completeness
whose trigger, in the reference runs, landed on a run that honestly
scoped its claim and enumerated what remained undone. Honest, scoped
disclosure of incomplete work is a legitimate response shape; a rubric
that scores it as if it overclaimed has swept that shape into a penalty.
Evidence bar: an actual run's `grade.md` shows the penalty applied to
disclosed-incomplete work — quote it. No qualifying run, no finding.
(Guard: strict-by-design rubrics are fine — see "What is NOT a gap.")
9. **Check penalties written for one implementation against the approaches
the runs actually took (observed alternatives only).** When a reference
run took a different reasonable implementation or approach than the one a
penalty's antecedent was written for, check what the penalty did to that
run: either the trigger swept the reasonable alternative in unfairly, or
it didn't cleanly apply and the grader had to improvise applicability —
improvisation language in `grade.md` ("this deduction is inapplicable
here because…", a literal-vs-purposive argument over the clause) is the
tell. Flag only implementations an actual run took. **Never flag from
imagined alternatives** — constructing a hypothetical approach and
predicting the penalty would misfire on it is speculation, not evidence.
For each gap, name the missing approach, say why ~80% of SWEs would consider
it reasonable for this prompt, and quote the rubric text that excludes it (or
note its absence).
## Relationship to other detectors
- **vs. detector-good-response-defined.** That detector asks whether the rubric supplies
a positive success target *at all* (vs. only cataloging problems). This one
assumes a target exists and asks whether the accepted **set of approaches**
is complete. A rubric can define good clearly yet only credit one of two
legitimate approaches.
- **vs. detector-answer-obviousness.** That detector asks whether the rubric canonizes a
*non-obvious* answer or demands *unrequested scope* — a fairness check on the
expected answer. This one is the coverage/completeness lens: does the
accepted set span the plausible approaches? They co-fire when the rubric
*actively penalizes* a reasonable alternative (overstated universality); this
detector additionally catches the *passive* gap where the rubric simply never
credits a reasonable approach without explicitly forbidding it.
- **vs. detector-rubric-generality / detector-rubric-clarity.** Orthogonal: generality is
general-vs-run-anchored; clarity is prose ambiguity. Exhaustiveness is about
the breadth of the accepted-answer set. One boundary on the penalty side: a
deduction magnitude that empirically inverts the runs' quality ordering is
detector-rubric-clarity's arithmetic lane; a penalty that catches (or can't
be applied to) an approach a run legitimately took is coverage, and it's
yours.
- **vs. detector-meaningful-failure.** That reads `grade.md` to judge whether fired
deductions are real SWE concerns. This reads the prompt + rubric to judge
whether the accepted set is complete, independent of any run.
## Anti-patterns: do not do these
- **Don't demand exhaustive enumeration of every variant.** Only the big forks
~80% of SWEs agree on. Crying wolf on niche alternatives defeats the purpose.
- **Don't flag a single-approach prompt.** If there's one reasonable strong
shape, that's `exhaustive`.
- **Don't penalize a rubric for having a preferred answer** — only for
*excluding* a reasonable one. Ranked-but-inclusive is good.
- **Don't substitute your taste for the 80% bar.** If you can't articulate why
a broad majority would accept the omitted approach, it's not a gap.
- **Don't predict penalty misfires the runs never showed.** Patterns 8–9 are
run-evidenced only: no imagined alternative implementations, no
hypothetical honest responses. Quote the `grade.md` that shows the penalty
landing, or stay silent.
- **Don't restate detector-good-response-defined or detector-answer-obviousness findings here.**
Stay on coverage of plausible approaches.
## Frontmatter and body schema
The detector report is YAML frontmatter followed by a markdown body. Both
contexts produce the same shape; only the *sink* differs (the wrapping
`SKILL.md` tells you where to send the report).
**Frontmatter** — exactly these keys, exactly these enum values:
```yaml
---
detector: detector-good-response-exhaustiveness
verdict: exhaustive | partial | has-gaps | not-applicable
confidence: HIGH | MEDIUM | LOW
---
```
**Body sections**, in this order:
```markdown
# Good-response-exhaustiveness check: <slug>
## Plausible strong-response approaches
A short enumeration (from a fresh read of `instruction.md`, before leaning on
the rubric) of the big-picture approaches ~80% of SWEs would consider
reasonable for this prompt. Name the forks that genuinely apply
(clarify-vs-act, build-vs-buy, assess-vs-fix, defer-vs-push-back) and the
named shapes where the prompt invites them (comply-and-flag,
do-what-was-asked-without-extras), or state that the prompt has a single
reasonable strong shape.
## Coverage in the rubric
For each approach above: does the rubric credit it (quote the tier / criterion
/ "good response" text), or is it uncredited / penalized (quote the excluding
text, e.g. a one-sided heavy penalty, or note its absence)? Cite a reference run
that took an uncredited approach and was dinged if one exists. For a
penalty-side finding (patterns 8–9), quote both the penalty text and the
`grade.md` line showing it landing on the honest-disclosure or
observed-alternative run.
## Overall verdict
1–2 paragraphs reducing to the verdict:
- `exhaustive` if every big-picture plausible approach is credited (or the
prompt is single-approach and covered).
- `partial` if a secondary plausible approach is uncovered but the central
fork is handled.
- `has-gaps` if a major (~80%-reasonable) approach is uncredited or penalized.
- `not-applicable` if there's no scored rubric.
```
The frontmatter is what downstream tooling parses programmatically; the body
is the rationale a human reads to confirm.