restored entire zip and config'd

This commit is contained in:
2026-10-07 15:37:45 -04:00
parent 4bd9264a29
commit 5ca605dcce
219 changed files with 43370 additions and 0 deletions

View File

@@ -0,0 +1,52 @@
---
name: detector-rubric-generality
description: |
Self-check your holistic rubric for whether it describes, in
general, what makes a response strong or weak — so a grader can apply it to
any agent — or whether it speaks too much in terms of your reference runs
("clarity is reliably high on this task", "agents will fail here", "all four
trials hit 85+"). Identifying failure modes as general response properties is
good; leaning on what the observed runs did as the scoring basis is what this
catches. Doesn't flag illustrative pointers to runs or describing failure
modes — only run-anchoring that gates scoring. Also flags a rubric that names
the framework your task runs on (Harbor, Pier, the sandbox) instead of
describing the task in its own terms.
allowed-tools: Bash, Read, Write
---
# Rubric-generality detector
This skill checks whether your holistic rubric (the file
`bash scripts/guidance-target.sh <slug>` resolves) describes response quality in
general terms — so the task works for any agent, not just the ones whose
reference runs you have today — or whether it leans too much on what the
observed runs happened to do ("reliably high on this task," "agents will," "all
N trials," tiers keyed to a specific run). It also flags a rubric that names the
framework your task runs on (Harbor, Pier, the sandbox) instead of the task's
own terms — "the final Harbor instruction" should just read "the final
instruction."
Read these before deciding:
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
2. `.claude/skills/detector-rubric-generality/core.md` — what counts as run-anchored scoring vs. general response-quality description, verdict definitions, frontmatter/body schema.
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
## Acting on the verdict
- **`generalizes`** — the main thrust describes what makes a response strong or
weak in general terms; any run-references are illustrative. Good.
- **`minor-issues`** — the core scoring is general, but some phrasings lean on
observed-run statistics or "agents tend to" framing, or name the framework
your task runs on. Look at the "run-anchored phrasings" and "infra-framework
references" lists in the report and reframe each as a general property of a
response (or, for an infra name, reword to the task's own terms). No need to
rebuild the rubric.
- **`material-issues`** — the load-bearing scoring criteria are defined by what
the reference runs did, so a grader couldn't score a new agent that fails
differently. Look at the "load-bearing run-dependence" section — rewrite those
criteria to describe what a strong/weak response looks like in general, then
re-run this skill.
- **`not-applicable`** — the resolved rubric file is missing, empty, or template-only.
Write the rubric first, then come back to this skill.