Files
project-work/worker-toolkit-potion-polyglot/.claude/skills/detector-rubric-generality/SKILL.md

2.9 KiB

name, description, allowed-tools
name description allowed-tools
detector-rubric-generality Self-check your holistic rubric for whether it describes, in general, what makes a response strong or weak — so a grader can apply it to any agent — or whether it speaks too much in terms of your reference runs ("clarity is reliably high on this task", "agents will fail here", "all four trials hit 85+"). Identifying failure modes as general response properties is good; leaning on what the observed runs did as the scoring basis is what this catches. Doesn't flag illustrative pointers to runs or describing failure modes — only run-anchoring that gates scoring. Also flags a rubric that names the framework your task runs on (Harbor, Pier, the sandbox) instead of describing the task in its own terms. Bash, Read, Write

Rubric-generality detector

This skill checks whether your holistic rubric (the file bash scripts/guidance-target.sh <slug> resolves) describes response quality in general terms — so the task works for any agent, not just the ones whose reference runs you have today — or whether it leans too much on what the observed runs happened to do ("reliably high on this task," "agents will," "all N trials," tiers keyed to a specific run). It also flags a rubric that names the framework your task runs on (Harbor, Pier, the sandbox) instead of the task's own terms — "the final Harbor instruction" should just read "the final instruction."

Read these before deciding:

  1. .claude/skills/_detector-worker-shell.md — where to write the report and how to handle re-runs.
  2. .claude/skills/detector-rubric-generality/core.md — what counts as run-anchored scoring vs. general response-quality description, verdict definitions, frontmatter/body schema.

Compose the report per the schema in core.md and write it per _detector-worker-shell.md.

Acting on the verdict

  • generalizes — the main thrust describes what makes a response strong or weak in general terms; any run-references are illustrative. Good.
  • minor-issues — the core scoring is general, but some phrasings lean on observed-run statistics or "agents tend to" framing, or name the framework your task runs on. Look at the "run-anchored phrasings" and "infra-framework references" lists in the report and reframe each as a general property of a response (or, for an infra name, reword to the task's own terms). No need to rebuild the rubric.
  • material-issues — the load-bearing scoring criteria are defined by what the reference runs did, so a grader couldn't score a new agent that fails differently. Look at the "load-bearing run-dependence" section — rewrite those criteria to describe what a strong/weak response looks like in general, then re-run this skill.
  • not-applicable — the resolved rubric file is missing, empty, or template-only. Write the rubric first, then come back to this skill.