lots of change - all to start my 3rd redo
This commit is contained in:
@@ -1,60 +0,0 @@
|
||||
---
|
||||
name: detector-good-response-defined
|
||||
description: |
|
||||
Self-check whether your holistic rubric makes it easy for the grader to tell
|
||||
what a strong response looks like — a positive success target ("what a good response
|
||||
says," an answer key of the findings a top answer surfaces, a worked example, or tiers
|
||||
that enumerate concrete positive content) — or whether it only catalogs problems
|
||||
(failure scenarios, "what a bad response says," deductions, heavy penalties), leaving the
|
||||
grader to infer "good" from the absence of listed problems. Multiple acceptable "good"
|
||||
shapes are fine and are never penalized. Reads the holistic rubric file that
|
||||
`bash scripts/guidance-target.sh <slug>` resolves (instruction.md for
|
||||
context).
|
||||
allowed-tools: Bash, Read, Write
|
||||
---
|
||||
|
||||
# Good-response-defined detector
|
||||
|
||||
This skill checks whether your holistic rubric gives the grader a **positive
|
||||
picture of success** — what a strong response actually says, contains, or
|
||||
does — or whether it only lists the ways a response can go wrong.
|
||||
|
||||
A grader needs to recognize a strong answer on its own terms. If your rubric
|
||||
is all "concrete failure scenario," "what a bad response says," deductions,
|
||||
and heavy penalties, the grader can only score by *absence of listed problems* —
|
||||
which over-credits a hollow answer that happens to dodge every trap, and
|
||||
under-serves a genuinely strong answer that does something you didn't
|
||||
anticipate. The fix is to state, affirmatively, what a good answer
|
||||
establishes — per issue or overall.
|
||||
|
||||
**Multiple "good" options are fine — encouraged.** "A strong response either
|
||||
defends the current design with sound reasoning, or proposes a migration
|
||||
with explicit tradeoffs — both acceptable" *defines good* perfectly well.
|
||||
The skill never penalizes you for allowing several strong shapes; it only
|
||||
flags never describing any.
|
||||
|
||||
The canonical format already asks for this. The rubric structure
|
||||
(`/write-holistic-rubric`) carries the positive target in its Ground truth
|
||||
and per-criterion sections. Keeping the failure half but dropping the good
|
||||
half is the `problems-only` shape this catches.
|
||||
|
||||
Read these before deciding:
|
||||
|
||||
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
|
||||
2. `.claude/skills/detector-good-response-defined/core.md` — what counts as defining good vs. problems-only, why multiple "good" options are fine, the boundary against detector-rubric-clarity, verdict enums.
|
||||
|
||||
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
|
||||
|
||||
## Acting on the verdict
|
||||
|
||||
- **`defines-good`** — your rubric gives the grader a clear positive target
|
||||
(one shape or several). Good. Move on.
|
||||
- **`partial`** — you've defined good for part of the task but the central
|
||||
thing it tests is left as failure scenarios. Add a "what a good response
|
||||
says" / answer-key treatment for the load-bearing issue, then re-run.
|
||||
- **`problems-only`** — your rubric is a catalog of problems with no
|
||||
affirmative success target. For each issue, add what a strong response
|
||||
establishes (it's fine to list more than one acceptable shape), or add an
|
||||
answer key of the findings a top answer surfaces, so the grader can
|
||||
recognize "good" directly. Re-run after.
|
||||
- **`not-applicable`** — no holistic rubric to assess yet. Draft it first.
|
||||
Reference in New Issue
Block a user