Files
Eric Bell 2854619bc9 chore: init commit
in worker.../repo/GITFOLDER.zip is the .git folder.
2026-08-11 14:44:09 -04:00

2.9 KiB

name, description, allowed-tools
name description allowed-tools
detector-rubric-generality Self-check your grader guidance for whether it describes, in general, what makes a response strong or weak — so a grader can apply it to any agent — or whether it speaks too much in terms of your reference runs ("clarity is reliably high on this task", "agents will fail here", "all four trials hit 85+"). Identifying failure modes as general response properties is good; leaning on what the observed runs did as the scoring basis is what this catches. Doesn't flag illustrative pointers to runs or describing failure modes — only run-anchoring that gates scoring. Also flags guidance that names the framework your task runs on (Harbor, Pier, the sandbox) instead of describing the task in its own terms. Bash, Read, Write

Rubric-generality detector

This skill checks whether your grader guidance (the file bash scripts/guidance-target.sh <slug> resolves) describes response quality in general terms — so the task works for any agent, not just the ones whose reference runs you have today — or whether it leans too much on what the observed runs happened to do ("reliably high on this task," "agents will," "all N trials," tiers keyed to a specific run). It also flags guidance that names the framework your task runs on (Harbor, Pier, the sandbox) instead of the task's own terms — "the final Harbor instruction" should just read "the final instruction."

Read these before deciding:

  1. .claude/skills/_detector-worker-shell.md — where to write the report and how to handle re-runs.
  2. .claude/skills/detector-rubric-generality/core.md — what counts as run-anchored scoring vs. general response-quality description, verdict definitions, frontmatter/body schema.

Compose the report per the schema in core.md and write it per _detector-worker-shell.md.

Acting on the verdict

  • generalizes — the main thrust describes what makes a response strong or weak in general terms; any run-references are illustrative. Good.
  • minor-issues — the core scoring is general, but some phrasings lean on observed-run statistics or "agents tend to" framing, or name the framework your task runs on. Look at the "run-anchored phrasings" and "infra-framework references" lists in the report and reframe each as a general property of a response (or, for an infra name, reword to the task's own terms). No need to rebuild the rubric.
  • material-issues — the load-bearing scoring criteria are defined by what the reference runs did, so a grader couldn't score a new agent that fails differently. Look at the "load-bearing run-dependence" section — rewrite those criteria to describe what a strong/weak response looks like in general, then re-run this skill.
  • not-applicable — the resolved guidance file is missing, empty, or template-only. Write the guidance first, then come back to this skill.