remove folder - untrustworthy

This commit is contained in:
2026-08-11 14:12:23 -04:00
parent f18bb0a146
commit 0012380fd3
140 changed files with 0 additions and 33136 deletions

View File

@@ -1,52 +0,0 @@
---
name: detector-rubric-generality
description: |
Self-check your grader guidance for whether it describes, in
general, what makes a response strong or weak — so a grader can apply it to
any agent — or whether it speaks too much in terms of your reference runs
("clarity is reliably high on this task", "agents will fail here", "all four
trials hit 85+"). Identifying failure modes as general response properties is
good; leaning on what the observed runs did as the scoring basis is what this
catches. Doesn't flag illustrative pointers to runs or describing failure
modes — only run-anchoring that gates scoring. Also flags guidance that names
the framework your task runs on (Harbor, Pier, the sandbox) instead of
describing the task in its own terms.
allowed-tools: Bash, Read, Write
---
# Rubric-generality detector
This skill checks whether your grader guidance (the file
`bash scripts/guidance-target.sh <slug>` resolves) describes response quality in
general terms — so the task works for any agent, not just the ones whose
reference runs you have today — or whether it leans too much on what the
observed runs happened to do ("reliably high on this task," "agents will," "all
N trials," tiers keyed to a specific run). It also flags guidance that names the
framework your task runs on (Harbor, Pier, the sandbox) instead of the task's
own terms — "the final Harbor instruction" should just read "the final
instruction."
Read these before deciding:
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
2. `.claude/skills/detector-rubric-generality/core.md` — what counts as run-anchored scoring vs. general response-quality description, verdict definitions, frontmatter/body schema.
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
## Acting on the verdict
- **`generalizes`** — the main thrust describes what makes a response strong or
weak in general terms; any run-references are illustrative. Good.
- **`minor-issues`** — the core scoring is general, but some phrasings lean on
observed-run statistics or "agents tend to" framing, or name the framework
your task runs on. Look at the "run-anchored phrasings" and "infra-framework
references" lists in the report and reframe each as a general property of a
response (or, for an infra name, reword to the task's own terms). No need to
rebuild the rubric.
- **`material-issues`** — the load-bearing scoring criteria are defined by what
the reference runs did, so a grader couldn't score a new agent that fails
differently. Look at the "load-bearing run-dependence" section — rewrite those
criteria to describe what a strong/weak response looks like in general, then
re-run this skill.
- **`not-applicable`** — the resolved guidance file is missing, empty, or template-only.
Write the guidance first, then come back to this skill.