Loaded up for the 3rd redo
Still on potion-voice
This commit is contained in:
@@ -0,0 +1,60 @@
|
||||
---
|
||||
name: detector-good-response-defined
|
||||
description: |
|
||||
Self-check whether your holistic rubric makes it easy for the grader to tell
|
||||
what a strong response looks like — a positive success target ("what a good response
|
||||
says," an answer key of the findings a top answer surfaces, a worked example, or tiers
|
||||
that enumerate concrete positive content) — or whether it only catalogs problems
|
||||
(failure scenarios, "what a bad response says," deductions, heavy penalties), leaving the
|
||||
grader to infer "good" from the absence of listed problems. Multiple acceptable "good"
|
||||
shapes are fine and are never penalized. Reads the holistic rubric file that
|
||||
`bash scripts/guidance-target.sh <slug>` resolves (instruction.md for
|
||||
context).
|
||||
allowed-tools: Bash, Read, Write
|
||||
---
|
||||
|
||||
# Good-response-defined detector
|
||||
|
||||
This skill checks whether your holistic rubric gives the grader a **positive
|
||||
picture of success** — what a strong response actually says, contains, or
|
||||
does — or whether it only lists the ways a response can go wrong.
|
||||
|
||||
A grader needs to recognize a strong answer on its own terms. If your rubric
|
||||
is all "concrete failure scenario," "what a bad response says," deductions,
|
||||
and heavy penalties, the grader can only score by *absence of listed problems* —
|
||||
which over-credits a hollow answer that happens to dodge every trap, and
|
||||
under-serves a genuinely strong answer that does something you didn't
|
||||
anticipate. The fix is to state, affirmatively, what a good answer
|
||||
establishes — per issue or overall.
|
||||
|
||||
**Multiple "good" options are fine — encouraged.** "A strong response either
|
||||
defends the current design with sound reasoning, or proposes a migration
|
||||
with explicit tradeoffs — both acceptable" *defines good* perfectly well.
|
||||
The skill never penalizes you for allowing several strong shapes; it only
|
||||
flags never describing any.
|
||||
|
||||
The canonical format already asks for this. The rubric structure
|
||||
(`/write-holistic-rubric`) carries the positive target in its Ground truth
|
||||
and per-criterion sections. Keeping the failure half but dropping the good
|
||||
half is the `problems-only` shape this catches.
|
||||
|
||||
Read these before deciding:
|
||||
|
||||
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
|
||||
2. `.claude/skills/detector-good-response-defined/core.md` — what counts as defining good vs. problems-only, why multiple "good" options are fine, the boundary against detector-rubric-clarity, verdict enums.
|
||||
|
||||
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
|
||||
|
||||
## Acting on the verdict
|
||||
|
||||
- **`defines-good`** — your rubric gives the grader a clear positive target
|
||||
(one shape or several). Good. Move on.
|
||||
- **`partial`** — you've defined good for part of the task but the central
|
||||
thing it tests is left as failure scenarios. Add a "what a good response
|
||||
says" / answer-key treatment for the load-bearing issue, then re-run.
|
||||
- **`problems-only`** — your rubric is a catalog of problems with no
|
||||
affirmative success target. For each issue, add what a strong response
|
||||
establishes (it's fine to list more than one acceptable shape), or add an
|
||||
answer key of the findings a top answer surfaces, so the grader can
|
||||
recognize "good" directly. Re-run after.
|
||||
- **`not-applicable`** — no holistic rubric to assess yet. Draft it first.
|
||||
@@ -0,0 +1,262 @@
|
||||
# Good-response-defined detector — core
|
||||
|
||||
This file is the canonical, context-neutral content for the
|
||||
detector-good-response-defined detector. It defines what the detector looks for, the
|
||||
verdict enums, the patterns to recognize, and the output schema. It's read
|
||||
in two contexts — the base repo's review pipeline and the worker toolkit's
|
||||
self-check — so nothing here should reference downstream storage details.
|
||||
|
||||
## What this detector is for
|
||||
|
||||
The central question this detector answers is: **reading the grader
|
||||
guidance, can the grader easily tell what a strong
|
||||
response looks like — or does the rubric only catalog the ways a response
|
||||
can go wrong?**
|
||||
|
||||
A grader scores an agent's answer against the rubric. To do that well, the
|
||||
grader needs a positive picture of success: what a strong response
|
||||
actually says, contains, or does. When the rubric supplies that — "a good
|
||||
response establishes X, cites Y, and recommends Z" / an answer key of the
|
||||
facts an A+ surfaces / a worked example of the target answer — the grader
|
||||
can recognize a strong answer directly, including a strong answer that
|
||||
takes a route the rubric author didn't personally anticipate.
|
||||
|
||||
When the rubric instead reads as a pile of problems — failure scenarios,
|
||||
"what a bad response says," deductions, heavy penalties that subtract points —
|
||||
with no affirmative statement of what good looks like, the grader is left
|
||||
to infer success from the *absence* of listed problems. That's a weak
|
||||
basis for grading: a response can dodge every enumerated failure and still
|
||||
be hollow, and a genuinely strong response that does something the rubric
|
||||
never imagined has nothing positive to be matched against. The grader ends
|
||||
up reverse-engineering the target from the list of traps, which is exactly
|
||||
the inconsistency this detector exists to surface.
|
||||
|
||||
**Multiple "good" options are fine — encouraged, even.** The bar is not "a
|
||||
single canonical answer." A rubric that says "a strong response either
|
||||
defends the current design with sound reasoning, or proposes a migration
|
||||
with explicit tradeoffs — both are acceptable" has *defined good* perfectly
|
||||
well. Do not penalize a rubric for admitting several strong shapes; only
|
||||
penalize it for never affirmatively describing any of them.
|
||||
|
||||
## What counts as "defining good"
|
||||
|
||||
Any affirmative specification of the success target the grader can match an
|
||||
answer against:
|
||||
|
||||
- **"What a good/strong response says/contains/does"** sections, per issue
|
||||
or overall.
|
||||
- **An answer key / ground-truth findings list** — the specific facts,
|
||||
citations, or conclusions a top-tier answer surfaces, so tiers map onto
|
||||
presence/absence of those facts.
|
||||
- **A worked exemplar** of the target answer (or a clear sketch of one).
|
||||
- **Tier descriptions that enumerate concrete positive content** — e.g.
|
||||
"A+ identifies the 100x precision risk in `delete(',.').to_i`, the race
|
||||
condition from the missing lock, and the nil-user audit gap" names what a
|
||||
strong answer contains, not just what a weak one misses.
|
||||
- **Multiple acceptable shapes**, each described — a menu of strong answers.
|
||||
|
||||
The test is functional: **if a grader read only this rubric, would it have
|
||||
a concrete positive target to compare the answer against?** If yes →
|
||||
defined. The positive target can be terse; it just has to exist and be
|
||||
specific enough to recognize.
|
||||
|
||||
## What does NOT count
|
||||
|
||||
- **Only failure scenarios / "what a bad response says" / deductions /
|
||||
heavy penalties.** A rubric written entirely as a catalog of mistakes defines
|
||||
*bad*, not *good*. "Deduct 20 if the agent misses the race condition"
|
||||
tells the grader what to subtract for; it never states what a strong
|
||||
answer affirmatively establishes.
|
||||
- **A "ground truth" section that states facts but never says the response
|
||||
should surface them.** Listing the repo's actual behavior is necessary
|
||||
context, but on its own it leaves "so what should a good answer *do* with
|
||||
this?" to the grader's imagination. (It edges toward `defines-good` when
|
||||
paired with tiers or a findings list that say the answer must surface
|
||||
those facts.)
|
||||
- **A positive sentence so vacuous it conveys no checkable target** — "a
|
||||
good response is thorough and senior-level" with nothing concrete behind
|
||||
it. Note: if the positive target *exists* but its *wording* is ambiguous
|
||||
("traces the flow accurately" with no key), that vagueness is
|
||||
detector-rubric-clarity's call; this detector cares whether a positive target is
|
||||
*present at all*. The two co-fire when the only attempt at a target is an
|
||||
empty phrase.
|
||||
|
||||
## Inputs
|
||||
|
||||
Read whatever you need from the task directory. The load-bearing artifact:
|
||||
|
||||
- The grader guidance — **the primary input; read every line.** Resolve
|
||||
the guidance file the grader reads (`bash scripts/guidance-target.sh
|
||||
<slug>` prints its path, `tests/grader-guidance-consolidated.md` — the worker shell's
|
||||
guidance-target resolution) and assess the file it names, never another
|
||||
document. You are judging whether it gives the grader a positive model
|
||||
of success.
|
||||
- `instruction.md` — secondary, for context on what the task asks (so you
|
||||
can tell whether the rubric's positive target, if any, actually addresses
|
||||
the request). You are not judging the prompt here.
|
||||
|
||||
You do not need the workspace, source repo, or reference runs. This
|
||||
detector judges what the rubric supplies the grader, not whether its claims
|
||||
are true (fact-check), nor whether its expectations are fair
|
||||
(detector-answer-obviousness), nor whether observed runs hit them (detector-meaningful-failure).
|
||||
|
||||
## Verdict definitions
|
||||
|
||||
- **`not-applicable`** — the resolved guidance file is missing, empty, or
|
||||
only the unmodified template scaffold (no authored content to assess).
|
||||
Emit this and stop.
|
||||
|
||||
- **`defines-good`** — the rubric gives the grader a clear positive model
|
||||
of what a strong response looks like: a "what a good response says"
|
||||
treatment, an answer key of expected findings, a worked exemplar, or
|
||||
tiers that enumerate concrete positive content. One acceptable shape or
|
||||
several — either way, a grader reading only this rubric could recognize a
|
||||
strong answer on its own terms, not just by absence of problems.
|
||||
|
||||
- **`partial`** — there's some positive signal, but it's thin or
|
||||
incomplete: a positive target for secondary issues but not the
|
||||
load-bearing one; a ground-truth section that gestures at the facts
|
||||
without saying a good answer must surface them; an A+ tier that names a
|
||||
couple of positive elements while the rest of the rubric is failure
|
||||
scenarios. The grader can tell what good looks like for part of the task
|
||||
and has to infer it for the rest.
|
||||
|
||||
- **`problems-only`** — the rubric is essentially a catalog of problems:
|
||||
failure scenarios, "what a bad response says," deductions, and hard
|
||||
gates, with no affirmative statement of what a strong response contains
|
||||
or does. The grader can only recognize "good" as "didn't trip the listed
|
||||
problems," which leaves strong-but-unanticipated answers unmatched and
|
||||
hollow-but-clean answers over-credited.
|
||||
|
||||
## Confidence
|
||||
|
||||
- **HIGH** — the call is unambiguous. The rubric plainly has (or plainly
|
||||
lacks) a positive success target.
|
||||
- **MEDIUM** — genuinely borderline: e.g., a ground-truth section that a
|
||||
reasonable reviewer might read as an implicit positive target or might
|
||||
not. Typical of `partial` calls.
|
||||
- **LOW** — limited or confusing material (very short rubric, unusual
|
||||
structure). Verdict is best-guess.
|
||||
|
||||
## Patterns to look for
|
||||
|
||||
Walk the rubric and tally positive vs. negative content:
|
||||
|
||||
1. **Scan for affirmative target language** — "what a good/strong response
|
||||
says," "a sound answer establishes," "the response should surface," an
|
||||
"answer key" / "expected findings" / "ground truth the answer must
|
||||
identify," a worked example. Presence of any concrete one points to
|
||||
`defines-good`.
|
||||
2. **Scan the tiers.** Does the top tier *enumerate what a strong answer
|
||||
contains* (positive), or only *what lower tiers miss* (negative framed
|
||||
relative to failures)? An A+ that lists concrete findings is a positive
|
||||
target even without a separate "good response" section.
|
||||
3. **Tally the negative-only structures** — "concrete failure scenario,"
|
||||
"what a bad response says," "deduct/cap if the agent fails to…," hard
|
||||
gates. A rubric that is *only* these, with nothing from step 1 or a
|
||||
positive step-2 tier, is `problems-only`.
|
||||
4. **Check coverage, not just presence.** If the positive target exists for
|
||||
minor points but the central thing the task tests has only failure
|
||||
framing, that's `partial`.
|
||||
5. **Honor multiple-good.** If the rubric describes more than one acceptable
|
||||
strong shape, that is *defining good*, not ambiguity — score it
|
||||
`defines-good`, never penalize the plurality.
|
||||
|
||||
## Relationship to other detectors
|
||||
|
||||
- **vs. detector-rubric-clarity.** detector-rubric-clarity asks, *for each criterion the
|
||||
rubric states, can a grader apply it consistently?* (wording, ambiguity,
|
||||
prose quality). This detector asks, *does the rubric state a positive
|
||||
success target at all, or only failures?* (orientation/coverage). The
|
||||
clean separating cases: a rubric with crisp, unambiguous failure
|
||||
scenarios and no "what good looks like" → detector-rubric-clarity `clear`,
|
||||
detector-good-response-defined `problems-only`; a rubric with a clear positive
|
||||
target whose tier wording is fuzzy → detector-good-response-defined `defines-good`,
|
||||
detector-rubric-clarity `material-issues`. They co-fire when the only attempt at a
|
||||
positive target is an empty phrase.
|
||||
- **vs. detector-answer-obviousness.** detector-answer-obviousness asks whether the rubric's
|
||||
expected answer is the obviously-right thing to do *given the prompt*
|
||||
(fairness — does it canonize a defensible alternative or demand
|
||||
unrequested scope). This detector doesn't judge whether the target is
|
||||
*right* or *fair*; only whether a positive target is *present* for the
|
||||
grader to use. A rubric can define good clearly (this detector passes) yet
|
||||
canonize a non-obvious answer (detector-answer-obviousness fires), and vice versa.
|
||||
- **vs. detector-rubric-generality.** detector-rubric-generality asks whether the rubric
|
||||
describes strong/weak *in general* vs. anchoring on the observed reference
|
||||
runs. A rubric can define good in run-anchored terms (generality fires,
|
||||
this passes) or fail to define good at all (this fires, generality may be
|
||||
moot). Related lenses, different defects.
|
||||
- **vs. detector-meaningful-failure.** detector-meaningful-failure reads `grade.md` and asks
|
||||
whether the deductions that fired are real SWE concerns. This detector
|
||||
doesn't look at runs; it asks whether the rubric supplies a positive
|
||||
target irrespective of what any run did.
|
||||
|
||||
## Anti-patterns: do not do these
|
||||
|
||||
- **Don't require a single canonical answer.** Multiple described strong
|
||||
shapes is `defines-good`. Penalizing plurality is the exact mistake to
|
||||
avoid — the user explicitly wants room for more than one "good."
|
||||
- **Don't double-count detector-rubric-clarity.** If a positive target is present but
|
||||
its wording is ambiguous, that's clarity's finding; here it still counts
|
||||
as *defined* (unless the wording is so empty it specifies nothing).
|
||||
- **Don't reward a wall of failure scenarios because it's thorough.** A long,
|
||||
detailed catalog of everything that can go wrong is still `problems-only`
|
||||
if it never says what a strong answer affirmatively does.
|
||||
- **Don't demand an exemplar.** A concrete answer key or positive tier
|
||||
content is enough; a fully worked sample answer is nice but not required.
|
||||
- **Don't judge whether the target is correct or fair** — that's
|
||||
fact-check / detector-answer-obviousness / detector-meaningful-failure. Only whether it's
|
||||
present and usable.
|
||||
|
||||
## Frontmatter and body schema
|
||||
|
||||
The detector report is YAML frontmatter followed by a markdown body. Both
|
||||
contexts produce the same shape; only the *sink* differs (the wrapping
|
||||
`SKILL.md` tells you where to send the report).
|
||||
|
||||
**Frontmatter** — exactly these keys, exactly these enum values:
|
||||
|
||||
```yaml
|
||||
---
|
||||
detector: detector-good-response-defined
|
||||
verdict: defines-good | partial | problems-only | not-applicable
|
||||
confidence: HIGH | MEDIUM | LOW
|
||||
---
|
||||
```
|
||||
|
||||
**Body sections**, in this order:
|
||||
|
||||
```markdown
|
||||
# Good-response-defined check: <slug>
|
||||
|
||||
## Positive target present?
|
||||
|
||||
2–4 sentences: does the rubric affirmatively describe what a strong
|
||||
response looks like, and where? Quote the positive-target language verbatim
|
||||
if present ("What a good response says: …", an answer-key bullet, a positive
|
||||
A+ enumeration). If the rubric describes more than one acceptable strong
|
||||
shape, note that — it counts in favor of `defines-good`.
|
||||
|
||||
## What the grader has to infer
|
||||
|
||||
2–4 sentences: name the negative-only structures (failure scenarios, "what
|
||||
a bad response says," deductions, heavy penalties) and, for `partial` /
|
||||
`problems-only`, state exactly what part of "good" the grader is left to
|
||||
reverse-engineer from the absence of problems — especially for the central
|
||||
thing the task tests.
|
||||
|
||||
## Overall verdict
|
||||
|
||||
1–2 paragraphs reducing to the verdict:
|
||||
|
||||
- `defines-good` if a grader reading only this rubric has a concrete
|
||||
positive target (one shape or several) for the load-bearing parts.
|
||||
- `partial` if the positive target covers some of the task but the grader
|
||||
must infer "good" for the central part.
|
||||
- `problems-only` if the rubric is essentially a catalog of problems with no
|
||||
affirmative success target.
|
||||
- `not-applicable` if there's no authored rubric.
|
||||
```
|
||||
|
||||
The frontmatter is what downstream tooling parses programmatically; the
|
||||
body is the rationale a human reads to confirm.
|
||||
Reference in New Issue
Block a user