ren worker folder adding orig, mv new one into root
This commit is contained in:
@@ -0,0 +1,52 @@
|
||||
---
|
||||
name: detector-rubric-generality
|
||||
description: |
|
||||
Self-check your holistic rubric for whether it describes, in
|
||||
general, what makes a response strong or weak — so a grader can apply it to
|
||||
any agent — or whether it speaks too much in terms of your reference runs
|
||||
("clarity is reliably high on this task", "agents will fail here", "all four
|
||||
trials hit 85+"). Identifying failure modes as general response properties is
|
||||
good; leaning on what the observed runs did as the scoring basis is what this
|
||||
catches. Doesn't flag illustrative pointers to runs or describing failure
|
||||
modes — only run-anchoring that gates scoring. Also flags a rubric that names
|
||||
the framework your task runs on (Harbor, Pier, the sandbox) instead of
|
||||
describing the task in its own terms.
|
||||
allowed-tools: Bash, Read, Write
|
||||
---
|
||||
|
||||
# Rubric-generality detector
|
||||
|
||||
This skill checks whether your holistic rubric (the file
|
||||
`bash scripts/guidance-target.sh <slug>` resolves) describes response quality in
|
||||
general terms — so the task works for any agent, not just the ones whose
|
||||
reference runs you have today — or whether it leans too much on what the
|
||||
observed runs happened to do ("reliably high on this task," "agents will," "all
|
||||
N trials," tiers keyed to a specific run). It also flags a rubric that names the
|
||||
framework your task runs on (Harbor, Pier, the sandbox) instead of the task's
|
||||
own terms — "the final Harbor instruction" should just read "the final
|
||||
instruction."
|
||||
|
||||
Read these before deciding:
|
||||
|
||||
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
|
||||
2. `.claude/skills/detector-rubric-generality/core.md` — what counts as run-anchored scoring vs. general response-quality description, verdict definitions, frontmatter/body schema.
|
||||
|
||||
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
|
||||
|
||||
## Acting on the verdict
|
||||
|
||||
- **`generalizes`** — the main thrust describes what makes a response strong or
|
||||
weak in general terms; any run-references are illustrative. Good.
|
||||
- **`minor-issues`** — the core scoring is general, but some phrasings lean on
|
||||
observed-run statistics or "agents tend to" framing, or name the framework
|
||||
your task runs on. Look at the "run-anchored phrasings" and "infra-framework
|
||||
references" lists in the report and reframe each as a general property of a
|
||||
response (or, for an infra name, reword to the task's own terms). No need to
|
||||
rebuild the rubric.
|
||||
- **`material-issues`** — the load-bearing scoring criteria are defined by what
|
||||
the reference runs did, so a grader couldn't score a new agent that fails
|
||||
differently. Look at the "load-bearing run-dependence" section — rewrite those
|
||||
criteria to describe what a strong/weak response looks like in general, then
|
||||
re-run this skill.
|
||||
- **`not-applicable`** — the resolved rubric file is missing, empty, or template-only.
|
||||
Write the rubric first, then come back to this skill.
|
||||
Reference in New Issue
Block a user