--- name: detector-rubric-generality description: | Self-check your grader guidance for whether it describes, in general, what makes a response strong or weak — so a grader can apply it to any agent — or whether it speaks too much in terms of your reference runs ("clarity is reliably high on this task", "agents will fail here", "all four trials hit 85+"). Identifying failure modes as general response properties is good; leaning on what the observed runs did as the scoring basis is what this catches. Doesn't flag illustrative pointers to runs or describing failure modes — only run-anchoring that gates scoring. Also flags guidance that names the framework your task runs on (Harbor, Pier, the sandbox) instead of describing the task in its own terms. allowed-tools: Bash, Read, Write --- # Rubric-generality detector This skill checks whether your grader guidance (the file `bash scripts/guidance-target.sh ` resolves) describes response quality in general terms — so the task works for any agent, not just the ones whose reference runs you have today — or whether it leans too much on what the observed runs happened to do ("reliably high on this task," "agents will," "all N trials," tiers keyed to a specific run). It also flags guidance that names the framework your task runs on (Harbor, Pier, the sandbox) instead of the task's own terms — "the final Harbor instruction" should just read "the final instruction." Read these before deciding: 1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs. 2. `.claude/skills/detector-rubric-generality/core.md` — what counts as run-anchored scoring vs. general response-quality description, verdict definitions, frontmatter/body schema. Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`. ## Acting on the verdict - **`generalizes`** — the main thrust describes what makes a response strong or weak in general terms; any run-references are illustrative. Good. - **`minor-issues`** — the core scoring is general, but some phrasings lean on observed-run statistics or "agents tend to" framing, or name the framework your task runs on. Look at the "run-anchored phrasings" and "infra-framework references" lists in the report and reframe each as a general property of a response (or, for an infra name, reword to the task's own terms). No need to rebuild the rubric. - **`material-issues`** — the load-bearing scoring criteria are defined by what the reference runs did, so a grader couldn't score a new agent that fails differently. Look at the "load-bearing run-dependence" section — rewrite those criteria to describe what a strong/weak response looks like in general, then re-run this skill. - **`not-applicable`** — the resolved guidance file is missing, empty, or template-only. Write the guidance first, then come back to this skill.