Loaded up for the 3rd redo
Still on potion-voice
This commit is contained in:
@@ -0,0 +1,98 @@
|
||||
---
|
||||
name: detector-answer-obviousness
|
||||
description: |
|
||||
Self-check whether the answer your rubric expects is *fairly* obvious given your
|
||||
prompt — neither so non-obvious that your holistic rubric penalizes the agent for
|
||||
mind-reading, nor so cued that your prompt hands the answer over. Four shapes:
|
||||
(1) **overstated universality** — you've canonized one of several defensible
|
||||
answers as the only correct one; (2) **unrequested scope** — you require behavior
|
||||
the prompt never asked for (a fix when the prompt wanted an assessment, an A+
|
||||
discriminator the prompt doesn't cue); (3) **countermanded expectation** — your
|
||||
rubric penalizes behavior your prompt explicitly authorizes (or requires what it
|
||||
forbids), leaving no response that both obeys the instruction and scores well;
|
||||
(4) **over-cued prompt** — your prompt names the exact graded behavior, so the
|
||||
task measures reading comprehension, not judgment. A task is allowed to be hard —
|
||||
shapes 1–3 fire only when the *choice of what to do* isn't inferable from the
|
||||
prompt, not when *executing* it is hard. Reads instruction.md + the holistic
|
||||
rubric file that `bash scripts/guidance-target.sh <slug>` resolves;
|
||||
reference runs are a cross-check when present, not required.
|
||||
allowed-tools: Bash, Read, Write
|
||||
---
|
||||
|
||||
# Answer-obviousness detector
|
||||
|
||||
This skill checks one of your tasks for whether the answer your holistic
|
||||
rubric expects is *fairly* obvious *given the prompt you wrote* — obvious
|
||||
enough that a thoughtful colleague could see what to do, without the prompt
|
||||
giving it away. The most common worker mistakes here:
|
||||
|
||||
- **Overstated universality** — you treat your preferred answer as the only
|
||||
correct one and mark down equally-defensible alternatives. ("The correct
|
||||
fix is X" when X is *a* fix, not *the* fix.) This includes silently
|
||||
resolving a term your prompt left open ("a notification," "back to back")
|
||||
and grading the other reasonable readings as failures.
|
||||
- **Unrequested scope** — you require something the prompt doesn't ask for.
|
||||
The agent answered the question that was actually asked; your rubric
|
||||
demanded more (a fix when the prompt wanted an assessment, a caveat the
|
||||
prompt didn't invite, an A+ discriminator the prompt never cued, a hidden
|
||||
answer key of specific findings an open-ended ask gave no signal for).
|
||||
- **Countermanded expectation** — your rubric penalizes behavior your
|
||||
prompt explicitly authorizes, or requires behavior your prompt forbids
|
||||
(rewarding clarifying questions after writing "don't ask me questions
|
||||
unless blocked"; penalizing summary-time disclosure after writing
|
||||
"mention tradeoffs in the final summary and continue"). There's no
|
||||
response that both follows your instruction and scores well.
|
||||
- **Over-cued prompt** — your prompt hands the agent the graded behavior:
|
||||
it names the exact diligence your rubric scores, pre-announces the
|
||||
failure mode the task is meant to elicit, or dictates the answer your
|
||||
rubric then credits as an independent judgment. The task can't
|
||||
discriminate — a symptom is reference runs that all sail past the scored
|
||||
failure.
|
||||
|
||||
Crucially, **a hard task is fine.** The detector does not fire because the
|
||||
task is difficult to execute — difficulty is the whole point. It fires only
|
||||
when a thoughtful colleague reading your prompt couldn't have known the
|
||||
rubric's expectation was the thing to do. Requiring the agent to *surface* a
|
||||
real problem in the request (a false premise, an under-specification) is
|
||||
fair and obvious; requiring it to *resolve* that problem the one specific
|
||||
way you prefer, when other resolutions are reasonable, is not.
|
||||
|
||||
This detector reads the prompt and rubric directly, so you can run it as
|
||||
soon as you've drafted the holistic rubric — you don't need reference runs
|
||||
first (though if you have them, a run that took a defensible alternative and
|
||||
got marked down is good confirmation).
|
||||
|
||||
Read these before deciding:
|
||||
|
||||
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
|
||||
2. `.claude/skills/detector-answer-obviousness/core.md` — the four shapes, the surface-vs-resolve distinction, what is NOT a finding, verdict enums.
|
||||
|
||||
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
|
||||
|
||||
## Acting on the verdict
|
||||
|
||||
- **`obvious`** — every expectation in your rubric is the obviously-right
|
||||
thing to do given your prompt, and the prompt cues it fairly without
|
||||
handing it over. Good. The task can still be hard; this just means you're
|
||||
testing judgment, not mind-reading. Move on.
|
||||
- **`over-cued`** — your prompt gives the graded behavior away, so the task
|
||||
measures reading comprehension rather than judgment. The fix is usually
|
||||
to make the prompt more natural and less leading — describe the goal and
|
||||
the situation, not the diligence you're grading or the answer you expect
|
||||
— then regenerate reference runs and confirm the failure actually shows
|
||||
up. Re-run this skill after.
|
||||
- **`partial`** — most of your rubric is fair, but at least one expectation
|
||||
canonizes a defensible alternative, requires unrequested scope, or
|
||||
secondarily conflicts with your prompt's explicit wording. Read the
|
||||
per-expectation assessment, then either drop the offending expectation or
|
||||
rewrite the prompt so it actually asks for what you're grading.
|
||||
- **`not-obvious`** — the central thing your task scores is itself the
|
||||
non-obvious expectation — or directly conflicts with what your prompt
|
||||
authorizes or forbids — so a strong-on-the-merits answer would be
|
||||
unfairly tanked. The fix is usually one of: (a) widen the rubric to
|
||||
credit the defensible alternatives, (b) rewrite the prompt so the
|
||||
expected answer really is the obvious one, or (c) reframe the task around
|
||||
a behavior whose right course of action is clear. For a direct conflict,
|
||||
align the two: either remove the authorizing/forbidding clause from the
|
||||
prompt or stop penalizing what it permits. Re-run this skill after.
|
||||
- **`not-applicable`** — no holistic rubric to assess yet. Draft it first.
|
||||
Reference in New Issue
Block a user