ren worker folder adding orig, mv new one into root

This commit is contained in:
2026-09-25 10:34:29 -04:00
parent 10f0668e32
commit 5b010039d7
1308 changed files with 44597 additions and 1511 deletions

View File

@@ -0,0 +1,60 @@
---
name: detector-good-response-defined
description: |
Self-check whether your holistic rubric makes it easy for the grader to tell
what a strong response looks like — a positive success target ("what a good response
says," an answer key of the findings a top answer surfaces, a worked example, or tiers
that enumerate concrete positive content) — or whether it only catalogs problems
(failure scenarios, "what a bad response says," deductions, heavy penalties), leaving the
grader to infer "good" from the absence of listed problems. Multiple acceptable "good"
shapes are fine and are never penalized. Reads the holistic rubric file that
`bash scripts/guidance-target.sh <slug>` resolves (instruction.md for
context).
allowed-tools: Bash, Read, Write
---
# Good-response-defined detector
This skill checks whether your holistic rubric gives the grader a **positive
picture of success** — what a strong response actually says, contains, or
does — or whether it only lists the ways a response can go wrong.
A grader needs to recognize a strong answer on its own terms. If your rubric
is all "concrete failure scenario," "what a bad response says," deductions,
and heavy penalties, the grader can only score by *absence of listed problems* —
which over-credits a hollow answer that happens to dodge every trap, and
under-serves a genuinely strong answer that does something you didn't
anticipate. The fix is to state, affirmatively, what a good answer
establishes — per issue or overall.
**Multiple "good" options are fine — encouraged.** "A strong response either
defends the current design with sound reasoning, or proposes a migration
with explicit tradeoffs — both acceptable" *defines good* perfectly well.
The skill never penalizes you for allowing several strong shapes; it only
flags never describing any.
The canonical format already asks for this. The rubric structure
(`/write-holistic-rubric`) carries the positive target in its Ground truth
and per-criterion sections. Keeping the failure half but dropping the good
half is the `problems-only` shape this catches.
Read these before deciding:
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
2. `.claude/skills/detector-good-response-defined/core.md` — what counts as defining good vs. problems-only, why multiple "good" options are fine, the boundary against detector-rubric-clarity, verdict enums.
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
## Acting on the verdict
- **`defines-good`** — your rubric gives the grader a clear positive target
(one shape or several). Good. Move on.
- **`partial`** — you've defined good for part of the task but the central
thing it tests is left as failure scenarios. Add a "what a good response
says" / answer-key treatment for the load-bearing issue, then re-run.
- **`problems-only`** — your rubric is a catalog of problems with no
affirmative success target. For each issue, add what a strong response
establishes (it's fine to list more than one acceptable shape), or add an
answer key of the findings a top answer surfaces, so the grader can
recognize "good" directly. Re-run after.
- **`not-applicable`** — no holistic rubric to assess yet. Draft it first.

View File

@@ -0,0 +1,262 @@
# Good-response-defined detector — core
This file is the canonical, context-neutral content for the
detector-good-response-defined detector. It defines what the detector looks for, the
verdict enums, the patterns to recognize, and the output schema. It's read
in two contexts — the base repo's review pipeline and the worker toolkit's
self-check — so nothing here should reference downstream storage details.
## What this detector is for
The central question this detector answers is: **reading the grader
guidance, can the grader easily tell what a strong
response looks like — or does the rubric only catalog the ways a response
can go wrong?**
A grader scores an agent's answer against the rubric. To do that well, the
grader needs a positive picture of success: what a strong response
actually says, contains, or does. When the rubric supplies that — "a good
response establishes X, cites Y, and recommends Z" / an answer key of the
facts an A+ surfaces / a worked example of the target answer — the grader
can recognize a strong answer directly, including a strong answer that
takes a route the rubric author didn't personally anticipate.
When the rubric instead reads as a pile of problems — failure scenarios,
"what a bad response says," deductions, heavy penalties that subtract points —
with no affirmative statement of what good looks like, the grader is left
to infer success from the *absence* of listed problems. That's a weak
basis for grading: a response can dodge every enumerated failure and still
be hollow, and a genuinely strong response that does something the rubric
never imagined has nothing positive to be matched against. The grader ends
up reverse-engineering the target from the list of traps, which is exactly
the inconsistency this detector exists to surface.
**Multiple "good" options are fine — encouraged, even.** The bar is not "a
single canonical answer." A rubric that says "a strong response either
defends the current design with sound reasoning, or proposes a migration
with explicit tradeoffs — both are acceptable" has *defined good* perfectly
well. Do not penalize a rubric for admitting several strong shapes; only
penalize it for never affirmatively describing any of them.
## What counts as "defining good"
Any affirmative specification of the success target the grader can match an
answer against:
- **"What a good/strong response says/contains/does"** sections, per issue
or overall.
- **An answer key / ground-truth findings list** — the specific facts,
citations, or conclusions a top-tier answer surfaces, so tiers map onto
presence/absence of those facts.
- **A worked exemplar** of the target answer (or a clear sketch of one).
- **Tier descriptions that enumerate concrete positive content** — e.g.
"A+ identifies the 100x precision risk in `delete(',.').to_i`, the race
condition from the missing lock, and the nil-user audit gap" names what a
strong answer contains, not just what a weak one misses.
- **Multiple acceptable shapes**, each described — a menu of strong answers.
The test is functional: **if a grader read only this rubric, would it have
a concrete positive target to compare the answer against?** If yes →
defined. The positive target can be terse; it just has to exist and be
specific enough to recognize.
## What does NOT count
- **Only failure scenarios / "what a bad response says" / deductions /
heavy penalties.** A rubric written entirely as a catalog of mistakes defines
*bad*, not *good*. "Deduct 20 if the agent misses the race condition"
tells the grader what to subtract for; it never states what a strong
answer affirmatively establishes.
- **A "ground truth" section that states facts but never says the response
should surface them.** Listing the repo's actual behavior is necessary
context, but on its own it leaves "so what should a good answer *do* with
this?" to the grader's imagination. (It edges toward `defines-good` when
paired with tiers or a findings list that say the answer must surface
those facts.)
- **A positive sentence so vacuous it conveys no checkable target** — "a
good response is thorough and senior-level" with nothing concrete behind
it. Note: if the positive target *exists* but its *wording* is ambiguous
("traces the flow accurately" with no key), that vagueness is
detector-rubric-clarity's call; this detector cares whether a positive target is
*present at all*. The two co-fire when the only attempt at a target is an
empty phrase.
## Inputs
Read whatever you need from the task directory. The load-bearing artifact:
- The grader guidance — **the primary input; read every line.** Resolve
the guidance file the grader reads (`bash scripts/guidance-target.sh
<slug>` prints its path, `tests/grader-guidance-consolidated.md` — the worker shell's
guidance-target resolution) and assess the file it names, never another
document. You are judging whether it gives the grader a positive model
of success.
- `instruction.md` — secondary, for context on what the task asks (so you
can tell whether the rubric's positive target, if any, actually addresses
the request). You are not judging the prompt here.
You do not need the workspace, source repo, or reference runs. This
detector judges what the rubric supplies the grader, not whether its claims
are true (fact-check), nor whether its expectations are fair
(detector-answer-obviousness), nor whether observed runs hit them (detector-meaningful-failure).
## Verdict definitions
- **`not-applicable`** — the resolved guidance file is missing, empty, or
only the unmodified template scaffold (no authored content to assess).
Emit this and stop.
- **`defines-good`** — the rubric gives the grader a clear positive model
of what a strong response looks like: a "what a good response says"
treatment, an answer key of expected findings, a worked exemplar, or
tiers that enumerate concrete positive content. One acceptable shape or
several — either way, a grader reading only this rubric could recognize a
strong answer on its own terms, not just by absence of problems.
- **`partial`** — there's some positive signal, but it's thin or
incomplete: a positive target for secondary issues but not the
load-bearing one; a ground-truth section that gestures at the facts
without saying a good answer must surface them; an A+ tier that names a
couple of positive elements while the rest of the rubric is failure
scenarios. The grader can tell what good looks like for part of the task
and has to infer it for the rest.
- **`problems-only`** — the rubric is essentially a catalog of problems:
failure scenarios, "what a bad response says," deductions, and hard
gates, with no affirmative statement of what a strong response contains
or does. The grader can only recognize "good" as "didn't trip the listed
problems," which leaves strong-but-unanticipated answers unmatched and
hollow-but-clean answers over-credited.
## Confidence
- **HIGH** — the call is unambiguous. The rubric plainly has (or plainly
lacks) a positive success target.
- **MEDIUM** — genuinely borderline: e.g., a ground-truth section that a
reasonable reviewer might read as an implicit positive target or might
not. Typical of `partial` calls.
- **LOW** — limited or confusing material (very short rubric, unusual
structure). Verdict is best-guess.
## Patterns to look for
Walk the rubric and tally positive vs. negative content:
1. **Scan for affirmative target language** — "what a good/strong response
says," "a sound answer establishes," "the response should surface," an
"answer key" / "expected findings" / "ground truth the answer must
identify," a worked example. Presence of any concrete one points to
`defines-good`.
2. **Scan the tiers.** Does the top tier *enumerate what a strong answer
contains* (positive), or only *what lower tiers miss* (negative framed
relative to failures)? An A+ that lists concrete findings is a positive
target even without a separate "good response" section.
3. **Tally the negative-only structures** — "concrete failure scenario,"
"what a bad response says," "deduct/cap if the agent fails to…," hard
gates. A rubric that is *only* these, with nothing from step 1 or a
positive step-2 tier, is `problems-only`.
4. **Check coverage, not just presence.** If the positive target exists for
minor points but the central thing the task tests has only failure
framing, that's `partial`.
5. **Honor multiple-good.** If the rubric describes more than one acceptable
strong shape, that is *defining good*, not ambiguity — score it
`defines-good`, never penalize the plurality.
## Relationship to other detectors
- **vs. detector-rubric-clarity.** detector-rubric-clarity asks, *for each criterion the
rubric states, can a grader apply it consistently?* (wording, ambiguity,
prose quality). This detector asks, *does the rubric state a positive
success target at all, or only failures?* (orientation/coverage). The
clean separating cases: a rubric with crisp, unambiguous failure
scenarios and no "what good looks like" → detector-rubric-clarity `clear`,
detector-good-response-defined `problems-only`; a rubric with a clear positive
target whose tier wording is fuzzy → detector-good-response-defined `defines-good`,
detector-rubric-clarity `material-issues`. They co-fire when the only attempt at a
positive target is an empty phrase.
- **vs. detector-answer-obviousness.** detector-answer-obviousness asks whether the rubric's
expected answer is the obviously-right thing to do *given the prompt*
(fairness — does it canonize a defensible alternative or demand
unrequested scope). This detector doesn't judge whether the target is
*right* or *fair*; only whether a positive target is *present* for the
grader to use. A rubric can define good clearly (this detector passes) yet
canonize a non-obvious answer (detector-answer-obviousness fires), and vice versa.
- **vs. detector-rubric-generality.** detector-rubric-generality asks whether the rubric
describes strong/weak *in general* vs. anchoring on the observed reference
runs. A rubric can define good in run-anchored terms (generality fires,
this passes) or fail to define good at all (this fires, generality may be
moot). Related lenses, different defects.
- **vs. detector-meaningful-failure.** detector-meaningful-failure reads `grade.md` and asks
whether the deductions that fired are real SWE concerns. This detector
doesn't look at runs; it asks whether the rubric supplies a positive
target irrespective of what any run did.
## Anti-patterns: do not do these
- **Don't require a single canonical answer.** Multiple described strong
shapes is `defines-good`. Penalizing plurality is the exact mistake to
avoid — the user explicitly wants room for more than one "good."
- **Don't double-count detector-rubric-clarity.** If a positive target is present but
its wording is ambiguous, that's clarity's finding; here it still counts
as *defined* (unless the wording is so empty it specifies nothing).
- **Don't reward a wall of failure scenarios because it's thorough.** A long,
detailed catalog of everything that can go wrong is still `problems-only`
if it never says what a strong answer affirmatively does.
- **Don't demand an exemplar.** A concrete answer key or positive tier
content is enough; a fully worked sample answer is nice but not required.
- **Don't judge whether the target is correct or fair** — that's
fact-check / detector-answer-obviousness / detector-meaningful-failure. Only whether it's
present and usable.
## Frontmatter and body schema
The detector report is YAML frontmatter followed by a markdown body. Both
contexts produce the same shape; only the *sink* differs (the wrapping
`SKILL.md` tells you where to send the report).
**Frontmatter** — exactly these keys, exactly these enum values:
```yaml
---
detector: detector-good-response-defined
verdict: defines-good | partial | problems-only | not-applicable
confidence: HIGH | MEDIUM | LOW
---
```
**Body sections**, in this order:
```markdown
# Good-response-defined check: <slug>
## Positive target present?
2–4 sentences: does the rubric affirmatively describe what a strong
response looks like, and where? Quote the positive-target language verbatim
if present ("What a good response says: …", an answer-key bullet, a positive
A+ enumeration). If the rubric describes more than one acceptable strong
shape, note that — it counts in favor of `defines-good`.
## What the grader has to infer
2–4 sentences: name the negative-only structures (failure scenarios, "what
a bad response says," deductions, heavy penalties) and, for `partial` /
`problems-only`, state exactly what part of "good" the grader is left to
reverse-engineer from the absence of problems — especially for the central
thing the task tests.
## Overall verdict
1–2 paragraphs reducing to the verdict:
- `defines-good` if a grader reading only this rubric has a concrete
positive target (one shape or several) for the load-bearing parts.
- `partial` if the positive target covers some of the task but the grader
must infer "good" for the central part.
- `problems-only` if the rubric is essentially a catalog of problems with no
affirmative success target.
- `not-applicable` if there's no authored rubric.
```
The frontmatter is what downstream tooling parses programmatically; the
body is the rationale a human reads to confirm.