ren worker folder adding orig, mv new one into root
This commit is contained in:
@@ -0,0 +1,66 @@
|
||||
---
|
||||
name: detector-good-response-exhaustiveness
|
||||
description: |
|
||||
Self-check whether your holistic rubric covers all the *plausible* types of
|
||||
strong response — the big-picture approaches a broad majority (~80%) of SWEs would
|
||||
consider reasonable for your prompt — or whether it only credits a subset, so an agent
|
||||
taking a reasonable-but-uncredited approach gets unfairly marked down — or sweeps a
|
||||
legitimate shape into a penalty aimed at something else (honest disclosure of incomplete
|
||||
work; an approach a reference run actually took), run-evidenced only. Flagship case:
|
||||
the clarify-vs-act fork — when a prompt has a real ambiguity, both "flag the issue and
|
||||
ask" and "flag the issue, state an assumption, act, and report" are usually legitimate,
|
||||
and the rubric should credit both. Not about crazy exhaustiveness — just the major forks
|
||||
(clarify-vs-act, build-vs-buy, assess-vs-fix, defer-vs-push-back). Reads instruction.md
|
||||
+ the holistic rubric file that `bash scripts/guidance-target.sh <slug>` resolves
|
||||
(reference runs optional, except penalty-side findings which
|
||||
require them).
|
||||
allowed-tools: Bash, Read, Write
|
||||
---
|
||||
|
||||
# Good-response-exhaustiveness detector
|
||||
|
||||
This skill checks whether your holistic rubric credits **all the plausible
|
||||
ways a competent SWE could respond well** to your prompt — not just your
|
||||
preferred path.
|
||||
|
||||
For many prompts there's more than one legitimate strong answer. The flagship
|
||||
case is the **clarify-vs-act fork**: when your prompt has a real ambiguity or
|
||||
decision point, both
|
||||
|
||||
1. "flag the issue and **ask for clarification**," and
|
||||
2. "flag the issue, **make a reasonable assumption (state it), act, and report
|
||||
what you did**,"
|
||||
|
||||
are usually legitimate. If ~80% of SWEs would accept both, your rubric should
|
||||
credit both — otherwise an agent that takes the uncredited path gets marked
|
||||
down for picking a reasonable approach you happened not to list.
|
||||
|
||||
Other big forks to check: **build-vs-buy / extend-vs-replace** (design
|
||||
prompts), **assess vs. answer-plus-fix** (question prompts), and **defer vs.
|
||||
push back** (when the prompt frames a decision as already made).
|
||||
|
||||
The bar is **not** crazy exhaustiveness — just the few big-picture approaches a
|
||||
broad majority of SWEs would agree are reasonable. A favorite/A+ approach plus
|
||||
acceptable alternatives is great; the problem is *excluding* a reasonable one
|
||||
(often via a one-sided heavy penalty — "heavily penalize unless the agent asks,"
|
||||
which dings a reasonable act-on-assumption answer, or vice versa).
|
||||
|
||||
Read these before deciding:
|
||||
|
||||
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
|
||||
2. `.claude/skills/detector-good-response-exhaustiveness/core.md` — the big forks, the ~80% bar, what is NOT a gap, the boundaries against detector-good-response-defined and detector-answer-obviousness, verdict enums.
|
||||
|
||||
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
|
||||
|
||||
## Acting on the verdict
|
||||
|
||||
- **`exhaustive`** — your rubric credits the big-picture plausible approaches
|
||||
(or the prompt has one reasonable shape and you cover it). Good. Move on.
|
||||
- **`partial`** — you cover the main approach but miss a secondary plausible
|
||||
one. Read the per-approach assessment; add a tier/criterion that credits it.
|
||||
- **`has-gaps`** — you're missing a major reasonable approach (often one side
|
||||
of the clarify-vs-act fork, or a build-vs-buy alternative). Add explicit
|
||||
credit for it — e.g. "a strong response either asks for clarification on ABC,
|
||||
or states the assumption that XYZ and proceeds, then reports it" — and relax
|
||||
any heavy penalty that forces one side of a legitimate fork. Re-run after.
|
||||
- **`not-applicable`** — no holistic rubric to assess yet. Draft it first.
|
||||
@@ -0,0 +1,361 @@
|
||||
# Good-response-exhaustiveness detector — core
|
||||
|
||||
This file is the canonical, context-neutral content for the
|
||||
detector-good-response-exhaustiveness detector. It defines what the detector looks
|
||||
for, the verdict enums, the patterns to recognize, and the output schema.
|
||||
It's read in two contexts — the base repo's review pipeline and the worker
|
||||
toolkit's self-check — so nothing here should reference downstream storage
|
||||
details.
|
||||
|
||||
## What this detector is for
|
||||
|
||||
The grader scores an agent's answer against the rubric's notion of a strong
|
||||
response. For many prompts there is **more than one legitimate way a
|
||||
competent SWE would respond**, and the rubric should credit each *plausible*
|
||||
approach — not just the author's preferred one. This detector asks:
|
||||
|
||||
**Does the rubric's set of accepted strong responses cover the big-picture
|
||||
approaches that a broad majority (~80%) of SWEs would consider reasonable?**
|
||||
|
||||
When it doesn't, an agent that takes a perfectly reasonable but uncredited
|
||||
approach gets unfairly marked down — the task ends up grading "did you pick
|
||||
the author's path" instead of "did you respond well."
|
||||
|
||||
The fairness frame cuts both ways. A legitimate response shape must be
|
||||
**credited** — the coverage question above — and it must not be **swept into
|
||||
a penalty** aimed at a different behavior. When the reference runs show a
|
||||
penalty for overclaiming landing on a run that honestly scoped its claim, or
|
||||
a penalty written for one implementation catching a run that reasonably took
|
||||
another, that is the same unfairness from the penalty side: a legitimate
|
||||
shape scored as a failure. These penalty-side shapes are **run-evidenced
|
||||
only** — see patterns 8–9.
|
||||
|
||||
The flagship case is the **clarify-vs-act fork.** When a prompt has a real
|
||||
ambiguity or judgment call, two responses are usually both legitimate:
|
||||
|
||||
1. **Flag the issue and ask for clarification** before acting, or
|
||||
2. **Flag the issue, make a reasonable assumption (state it), act on it, and
|
||||
report what was done.**
|
||||
|
||||
If ~80% of SWEs would accept *both*, the rubric should credit both. A rubric
|
||||
that credits only one — or heavily penalizes the other — has a coverage gap.
|
||||
|
||||
The bar is deliberately not "crazy exhaustive": you are looking for the few
|
||||
**big-picture** approaches most SWEs would agree are reasonable, not every
|
||||
micro-variation. Cover the major forks, not the long tail.
|
||||
|
||||
## What "covering the plausible approaches" means
|
||||
|
||||
- The rubric **credits, or at least leaves room for, each major reasonable
|
||||
approach** — not just the author's pick.
|
||||
- It's fine to have a *preferred* / A+ approach **plus** acceptable
|
||||
alternatives at the same or a slightly lower tier. What matters is that a
|
||||
reasonable approach isn't left **uncredited or penalized**.
|
||||
- "Credit" can be explicit ("a strong response either asks for clarification
|
||||
**or** states an assumption and proceeds") or structural (tiers/criteria
|
||||
that a reasonable alternative could satisfy). What you're checking is
|
||||
whether a grader, holding a reasonable-but-different answer, would find a
|
||||
basis to score it well.
|
||||
|
||||
## The big forks to check
|
||||
|
||||
Walk the prompt and ask which broadly-accepted approaches exist. The common
|
||||
ones:
|
||||
|
||||
1. **Clarify vs. act-on-assumption** — the flagship. Real ambiguity / a
|
||||
decision the prompt leaves open: both "ask first" and "state an assumption
|
||||
and proceed, then report" are usually legitimate. Does the rubric credit
|
||||
both, or does it reward only asking (and ding acting) or only acting (and
|
||||
ding asking)?
|
||||
2. **Assess/answer vs. answer-plus-fix** — for a question or assessment
|
||||
prompt, both "answer what was asked" and "answer + propose a remedy" can
|
||||
be reasonable. (Note the boundary with detector-answer-obviousness: *requiring* a
|
||||
fix the prompt didn't ask for is detector-answer-obviousness's unrequested-scope;
|
||||
here the concern is the mirror image — failing to credit a reasonable
|
||||
answer-only response, or a reasonable answer-plus-fix response.)
|
||||
3. **Build vs. buy / extend vs. replace** — for design/architecture prompts,
|
||||
several approaches are often defensible (keep the in-house system and
|
||||
extend it; or step back and recommend an off-the-shelf platform). Does the
|
||||
rubric credit the reasonable alternatives or canonize one?
|
||||
4. **Defer vs. push back** — when the prompt frames a decision as already
|
||||
made by the team, both "accept the stated decision and proceed" and
|
||||
"register a concern" can be reasonable. Does the rubric credit the one it
|
||||
doesn't prefer?
|
||||
|
||||
Coverage gaps are not always fork-shaped. A rubric can credit both sides of
|
||||
every fork and still leave a plausible response with nowhere to land — see
|
||||
patterns 4–7 below for the rubric-visible shapes: a strong-response set
|
||||
limited to a single credited path, a plausible middle/hybrid response that
|
||||
falls between the credited tier and a penalty, and the two recurring named
|
||||
shapes (**comply-and-flag** and **do-what-was-asked-without-extras**) that
|
||||
rubrics most often leave uncredited.
|
||||
|
||||
Not every prompt has multiple plausible approaches — a factual question or a
|
||||
prompt with one obviously-correct design has a single strong shape, and then
|
||||
exhaustiveness is trivially met. Only flag a gap when there is a **real,
|
||||
broadly-agreed alternative the rubric omits.**
|
||||
|
||||
## What is NOT a gap
|
||||
|
||||
- **Niche approaches.** Something only a minority of SWEs would do is not a
|
||||
required coverage item. The bar is ~80% agreement.
|
||||
- **A ranked-but-inclusive rubric.** Crediting several shapes and ranking
|
||||
them (preferred A+ + acceptable alternatives) is *good coverage*, not a
|
||||
gap. Don't flag a rubric for having a favorite — flag it for excluding a
|
||||
reasonable approach.
|
||||
- **Genuinely single-approach prompts.** If there's one reasonable strong
|
||||
shape, `exhaustive` is the right call.
|
||||
- **Don't re-litigate adjacent detectors.** Whether the rubric defines good
|
||||
*at all* is detector-good-response-defined; whether it canonizes a *non-obvious*
|
||||
answer or demands *unrequested scope* is detector-answer-obviousness. This detector
|
||||
assumes a positive target exists and asks whether the accepted **set** is
|
||||
complete.
|
||||
- **Your own taste.** The test is "would ~80% of SWEs accept this approach,"
|
||||
not "would I have done it this way." Don't invent alternatives a broad
|
||||
majority wouldn't actually endorse.
|
||||
- **Strict-by-design rubrics.** A rubric that explicitly and deliberately
|
||||
penalizes honest-incomplete work — because completeness itself is the
|
||||
deliverable the prompt asked for, and it says so — made a design choice,
|
||||
not a coverage error. Pattern 8 applies only when the rubric frames the
|
||||
penalty as targeting overclaiming / confidence / honesty.
|
||||
- **Penalty magnitude.** How big a deduction is, is the author's design call.
|
||||
The penalty-side shapes are about *which responses* a penalty catches, as
|
||||
shown by the runs — never "this deduction feels too heavy."
|
||||
|
||||
## Inputs
|
||||
|
||||
Read whatever you need from the task directory. The load-bearing artifacts:
|
||||
|
||||
- `instruction.md` — **read it first.** Establish the plausible strong-response
|
||||
approaches a competent SWE could reasonably take to *this* prompt
|
||||
(especially: does it contain a real ambiguity / decision point that opens
|
||||
the clarify-vs-act fork or a build-vs-buy choice?). This is the baseline the
|
||||
rubric's coverage is measured against.
|
||||
- The grader guidance — the rubric. Which approaches does it credit?
|
||||
Do any tiers / heavy penalties penalize a reasonable approach? Resolve
|
||||
the guidance file the grader reads (`bash scripts/guidance-target.sh
|
||||
<slug>` prints its path, `tests/grader-guidance-consolidated.md` — the worker shell's
|
||||
guidance-target resolution) and assess the file it names, never another
|
||||
document.
|
||||
- `reference-runs/<run>/agent-output/answer.md` + `grade.md` — *optional,
|
||||
supporting evidence.* If a run took a reasonable-but-uncredited approach and
|
||||
the grader dinged it, that confirms a real gap. Not required. When you do
|
||||
cite a run, characterize it accurately: a run that lost points to a
|
||||
one-sided criterion is evidence the gap *exists*, not evidence it doesn't;
|
||||
and check the run actually did what you say it did before crediting it as
|
||||
an honest instance of an approach. The one exception to "optional": the
|
||||
penalty-side shapes (patterns 8–9) exist only as run evidence — without a
|
||||
`grade.md` showing the penalty landing, they are not findings.
|
||||
|
||||
## Verdict definitions
|
||||
|
||||
- **`not-applicable`** — the resolved guidance file is missing, empty, or only
|
||||
the unmodified template scaffold. Emit this and stop.
|
||||
|
||||
- **`exhaustive`** — the rubric credits all the big-picture plausible
|
||||
strong-response approaches for this prompt (or the prompt has a single
|
||||
reasonable strong shape and the rubric covers it). A grader holding any
|
||||
~80%-reasonable answer would find a basis to score it on its merits.
|
||||
|
||||
- **`partial`** — the rubric covers the primary approach(es) but misses a
|
||||
*secondary* plausible one that a meaningful minority of SWEs would take.
|
||||
There's a real coverage gap, but it's not the central fork the task hinges
|
||||
on.
|
||||
|
||||
- **`has-gaps`** — the rubric misses a **major** plausible approach (one
|
||||
~80% of SWEs would consider reasonable for this prompt) — most often one
|
||||
side of the clarify-vs-act fork, or a build-vs-buy alternative the prompt
|
||||
invites. A reasonable response taking the uncredited approach would be
|
||||
unfairly marked down, so the task is partly grading path-selection rather
|
||||
than response quality.
|
||||
|
||||
## Confidence
|
||||
|
||||
- **HIGH** — the missing (or present) approach is clearly something a broad
|
||||
majority of SWEs would accept; the call isn't a close one.
|
||||
- **MEDIUM** — whether the omitted approach clears the ~80% bar is genuinely
|
||||
debatable; a reasonable reviewer might call it niche.
|
||||
- **LOW** — limited info (terse prompt, unfamiliar domain) makes the set of
|
||||
plausible approaches hard to enumerate confidently. Best-guess.
|
||||
|
||||
## Patterns to look for
|
||||
|
||||
1. **Enumerate the prompt's plausible approaches first**, before reading the
|
||||
rubric — so the rubric doesn't anchor you to only the approaches it
|
||||
happened to consider. Name the big forks (clarify-vs-act, build-vs-buy,
|
||||
assess-vs-fix, defer-vs-push-back) that genuinely apply, and the named
|
||||
shapes from patterns 6–7 where the prompt invites them. Anchoring cuts
|
||||
both ways: if the rubric *frames* an approach as categorically weak (e.g.
|
||||
treats clarification-first as a scoping failure), check independently
|
||||
whether that approach was reasonable for this prompt rather than adopting
|
||||
the rubric's framing.
|
||||
2. **Map each onto the rubric.** For each plausible approach, is there a tier
|
||||
/ criterion / "good response" statement it would satisfy? Or is it
|
||||
unmentioned / implicitly excluded? When you conclude an approach *is*
|
||||
credited, verify the text you're pointing at actually credits that
|
||||
approach — a line crediting an adjacent behavior is not credit for a
|
||||
different response type.
|
||||
3. **Check heavy penalties for one-sidedness.** "Heavily penalize unless the agent
|
||||
asks for clarification" dings a reasonable act-on-assumption answer;
|
||||
"heavily penalize unless the agent ships a fix" dings a reasonable
|
||||
clarify-first or assess-only answer. A heavy deduction that forces one side
|
||||
of a legitimate fork is the classic gap.
|
||||
4. **Check the A+/"strong response" enumeration for single-track framing** —
|
||||
one prescribed path when the prompt clearly admits several. This includes
|
||||
crediting only one failure path or finding when the prompt and rubric
|
||||
themselves put several on the table. Judge coverage from what the prompt
|
||||
and rubric surface — whether *additional* uncredited paths exist in the
|
||||
source repo is detector-fact-check-rubric-claims's verification, not yours.
|
||||
5. **Construct the plausible middle/hybrid response and find its tier.** When
|
||||
the rubric describes a credited behavior and a penalized behavior, a real
|
||||
agent often lands between them — keeps the questioned choice but surfaces
|
||||
the alternative and asks for confirmation, or ships the workaround while
|
||||
flagging the breakage. Build that blend and check it maps onto some tier or
|
||||
criterion. A plausible response that falls between two penalties — or
|
||||
between the strong tier and a heavy deduction — with no scored home is a
|
||||
coverage gap.
|
||||
6. **Check the comply-and-flag response for credit.** When the user
|
||||
explicitly insists on a risky or questionable action, "do what was asked
|
||||
and flag the concern in a line" is usually as legitimate as "push back /
|
||||
guard / decline" — often more so, since the user made the call knowingly.
|
||||
Rubrics recurringly score this fork asymmetrically: the cautious side
|
||||
(block, guard, refuse) earns the credit while complying with the explicit
|
||||
instruction draws a heavy deduction even when the concern was surfaced.
|
||||
Check that a comply-and-flag response has a scored home. Credit can be
|
||||
structural — a tier such an answer would satisfy counts; don't demand an
|
||||
approach-specific sentence.
|
||||
7. **Check the do-what-was-asked-without-extras response for credit.** When
|
||||
the literal request is coherent and the workspace supports it (the change
|
||||
is small and safe, or the asked-for feature already exists and passes
|
||||
tests), "do exactly what was requested, competently, and report" is often
|
||||
the plausible majority answer. A rubric that names plain compliance as the
|
||||
core failure — crediting only responses that first discover and surface a
|
||||
deeper concern the prompt never raised — leaves that majority shape
|
||||
uncredited. The deeper-discovery path can still be the A+; the question is
|
||||
whether competent literal compliance has a tier to land on. (Boundary:
|
||||
*requiring* the extras is detector-answer-obviousness's unrequested-scope;
|
||||
the coverage question here is whether plain compliance is credited at
|
||||
all.)
|
||||
8. **Check penalties aimed at overclaiming against the runs that disclosed
|
||||
honestly (run-evidenced only).** The mirror image of the comply-and-flag
|
||||
credit check: a penalty targeting overclaiming or unearned completeness
|
||||
whose trigger, in the reference runs, landed on a run that honestly
|
||||
scoped its claim and enumerated what remained undone. Honest, scoped
|
||||
disclosure of incomplete work is a legitimate response shape; a rubric
|
||||
that scores it as if it overclaimed has swept that shape into a penalty.
|
||||
Evidence bar: an actual run's `grade.md` shows the penalty applied to
|
||||
disclosed-incomplete work — quote it. No qualifying run, no finding.
|
||||
(Guard: strict-by-design rubrics are fine — see "What is NOT a gap.")
|
||||
9. **Check penalties written for one implementation against the approaches
|
||||
the runs actually took (observed alternatives only).** When a reference
|
||||
run took a different reasonable implementation or approach than the one a
|
||||
penalty's antecedent was written for, check what the penalty did to that
|
||||
run: either the trigger swept the reasonable alternative in unfairly, or
|
||||
it didn't cleanly apply and the grader had to improvise applicability —
|
||||
improvisation language in `grade.md` ("this deduction is inapplicable
|
||||
here because…", a literal-vs-purposive argument over the clause) is the
|
||||
tell. Flag only implementations an actual run took. **Never flag from
|
||||
imagined alternatives** — constructing a hypothetical approach and
|
||||
predicting the penalty would misfire on it is speculation, not evidence.
|
||||
|
||||
For each gap, name the missing approach, say why ~80% of SWEs would consider
|
||||
it reasonable for this prompt, and quote the rubric text that excludes it (or
|
||||
note its absence).
|
||||
|
||||
## Relationship to other detectors
|
||||
|
||||
- **vs. detector-good-response-defined.** That detector asks whether the rubric supplies
|
||||
a positive success target *at all* (vs. only cataloging problems). This one
|
||||
assumes a target exists and asks whether the accepted **set of approaches**
|
||||
is complete. A rubric can define good clearly yet only credit one of two
|
||||
legitimate approaches.
|
||||
- **vs. detector-answer-obviousness.** That detector asks whether the rubric canonizes a
|
||||
*non-obvious* answer or demands *unrequested scope* — a fairness check on the
|
||||
expected answer. This one is the coverage/completeness lens: does the
|
||||
accepted set span the plausible approaches? They co-fire when the rubric
|
||||
*actively penalizes* a reasonable alternative (overstated universality); this
|
||||
detector additionally catches the *passive* gap where the rubric simply never
|
||||
credits a reasonable approach without explicitly forbidding it.
|
||||
- **vs. detector-rubric-generality / detector-rubric-clarity.** Orthogonal: generality is
|
||||
general-vs-run-anchored; clarity is prose ambiguity. Exhaustiveness is about
|
||||
the breadth of the accepted-answer set. One boundary on the penalty side: a
|
||||
deduction magnitude that empirically inverts the runs' quality ordering is
|
||||
detector-rubric-clarity's arithmetic lane; a penalty that catches (or can't
|
||||
be applied to) an approach a run legitimately took is coverage, and it's
|
||||
yours.
|
||||
- **vs. detector-meaningful-failure.** That reads `grade.md` to judge whether fired
|
||||
deductions are real SWE concerns. This reads the prompt + rubric to judge
|
||||
whether the accepted set is complete, independent of any run.
|
||||
|
||||
## Anti-patterns: do not do these
|
||||
|
||||
- **Don't demand exhaustive enumeration of every variant.** Only the big forks
|
||||
~80% of SWEs agree on. Crying wolf on niche alternatives defeats the purpose.
|
||||
- **Don't flag a single-approach prompt.** If there's one reasonable strong
|
||||
shape, that's `exhaustive`.
|
||||
- **Don't penalize a rubric for having a preferred answer** — only for
|
||||
*excluding* a reasonable one. Ranked-but-inclusive is good.
|
||||
- **Don't substitute your taste for the 80% bar.** If you can't articulate why
|
||||
a broad majority would accept the omitted approach, it's not a gap.
|
||||
- **Don't predict penalty misfires the runs never showed.** Patterns 8–9 are
|
||||
run-evidenced only: no imagined alternative implementations, no
|
||||
hypothetical honest responses. Quote the `grade.md` that shows the penalty
|
||||
landing, or stay silent.
|
||||
- **Don't restate detector-good-response-defined or detector-answer-obviousness findings here.**
|
||||
Stay on coverage of plausible approaches.
|
||||
|
||||
## Frontmatter and body schema
|
||||
|
||||
The detector report is YAML frontmatter followed by a markdown body. Both
|
||||
contexts produce the same shape; only the *sink* differs (the wrapping
|
||||
`SKILL.md` tells you where to send the report).
|
||||
|
||||
**Frontmatter** — exactly these keys, exactly these enum values:
|
||||
|
||||
```yaml
|
||||
---
|
||||
detector: detector-good-response-exhaustiveness
|
||||
verdict: exhaustive | partial | has-gaps | not-applicable
|
||||
confidence: HIGH | MEDIUM | LOW
|
||||
---
|
||||
```
|
||||
|
||||
**Body sections**, in this order:
|
||||
|
||||
```markdown
|
||||
# Good-response-exhaustiveness check: <slug>
|
||||
|
||||
## Plausible strong-response approaches
|
||||
|
||||
A short enumeration (from a fresh read of `instruction.md`, before leaning on
|
||||
the rubric) of the big-picture approaches ~80% of SWEs would consider
|
||||
reasonable for this prompt. Name the forks that genuinely apply
|
||||
(clarify-vs-act, build-vs-buy, assess-vs-fix, defer-vs-push-back) and the
|
||||
named shapes where the prompt invites them (comply-and-flag,
|
||||
do-what-was-asked-without-extras), or state that the prompt has a single
|
||||
reasonable strong shape.
|
||||
|
||||
## Coverage in the rubric
|
||||
|
||||
For each approach above: does the rubric credit it (quote the tier / criterion
|
||||
/ "good response" text), or is it uncredited / penalized (quote the excluding
|
||||
text, e.g. a one-sided heavy penalty, or note its absence)? Cite a reference run
|
||||
that took an uncredited approach and was dinged if one exists. For a
|
||||
penalty-side finding (patterns 8–9), quote both the penalty text and the
|
||||
`grade.md` line showing it landing on the honest-disclosure or
|
||||
observed-alternative run.
|
||||
|
||||
## Overall verdict
|
||||
|
||||
1–2 paragraphs reducing to the verdict:
|
||||
|
||||
- `exhaustive` if every big-picture plausible approach is credited (or the
|
||||
prompt is single-approach and covered).
|
||||
- `partial` if a secondary plausible approach is uncovered but the central
|
||||
fork is handled.
|
||||
- `has-gaps` if a major (~80%-reasonable) approach is uncredited or penalized.
|
||||
- `not-applicable` if there's no scored rubric.
|
||||
```
|
||||
|
||||
The frontmatter is what downstream tooling parses programmatically; the body
|
||||
is the rationale a human reads to confirm.
|
||||
Reference in New Issue
Block a user