lots of change - all to start my 3rd redo

This commit is contained in:
2026-09-26 14:31:52 -04:00
parent 7f4d388e19
commit bceb52e8ee
1046 changed files with 4476 additions and 0 deletions

View File

@@ -0,0 +1,52 @@
---
name: detector-rubric-generality
description: |
Self-check your holistic rubric for whether it describes, in
general, what makes a response strong or weak — so a grader can apply it to
any agent — or whether it speaks too much in terms of your reference runs
("clarity is reliably high on this task", "agents will fail here", "all four
trials hit 85+"). Identifying failure modes as general response properties is
good; leaning on what the observed runs did as the scoring basis is what this
catches. Doesn't flag illustrative pointers to runs or describing failure
modes — only run-anchoring that gates scoring. Also flags a rubric that names
the framework your task runs on (Harbor, Pier, the sandbox) instead of
describing the task in its own terms.
allowed-tools: Bash, Read, Write
---
# Rubric-generality detector
This skill checks whether your holistic rubric (the file
`bash scripts/guidance-target.sh <slug>` resolves) describes response quality in
general terms — so the task works for any agent, not just the ones whose
reference runs you have today — or whether it leans too much on what the
observed runs happened to do ("reliably high on this task," "agents will," "all
N trials," tiers keyed to a specific run). It also flags a rubric that names the
framework your task runs on (Harbor, Pier, the sandbox) instead of the task's
own terms — "the final Harbor instruction" should just read "the final
instruction."
Read these before deciding:
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
2. `.claude/skills/detector-rubric-generality/core.md` — what counts as run-anchored scoring vs. general response-quality description, verdict definitions, frontmatter/body schema.
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
## Acting on the verdict
- **`generalizes`** — the main thrust describes what makes a response strong or
weak in general terms; any run-references are illustrative. Good.
- **`minor-issues`** — the core scoring is general, but some phrasings lean on
observed-run statistics or "agents tend to" framing, or name the framework
your task runs on. Look at the "run-anchored phrasings" and "infra-framework
references" lists in the report and reframe each as a general property of a
response (or, for an infra name, reword to the task's own terms). No need to
rebuild the rubric.
- **`material-issues`** — the load-bearing scoring criteria are defined by what
the reference runs did, so a grader couldn't score a new agent that fails
differently. Look at the "load-bearing run-dependence" section — rewrite those
criteria to describe what a strong/weak response looks like in general, then
re-run this skill.
- **`not-applicable`** — the resolved rubric file is missing, empty, or template-only.
Write the rubric first, then come back to this skill.

View File

@@ -0,0 +1,414 @@
# Rubric-generality detector — core
This file is the canonical, context-neutral content for the detector-rubric-generality
detector. It defines what counts as run-anchored scoring — and naming the
infrastructure the task runs on — vs. general response-quality description, the
verdict enums, the patterns to recognize,
and the output schema. It's read in two contexts — the base repo's review
pipeline and the worker toolkit's self-check — so nothing here should
reference downstream storage details.
## What this detector is for
We are building a benchmark that should work for **any** agent, not just the
handful of agents whose reference runs we happen to have on hand today. The
grader reads the task's grader guidance to score a response. For the benchmark
to generalize, the guidance's main thrust has to describe — in general terms —
what makes a response **strong or weak**, grounded in the task and the code, so
a grader can apply it to a response no reference run produced.
The failure this detector catches is grader guidance that instead **speaks too
much in terms of the observed agent runs.** Phrasings like "clarity is reliably
high on this task," "agents will fail here," "all four reference trials hit
85+," or "the runs that missed the population gap" describe *what the agents we
already watched happened to do*. The more the guidance leans on those
observations to do the scoring work, the more it targets the specific set of
failures we see today — and the less it tells a grader how to score a new agent
that fails (or succeeds) in a way none of the reference runs did.
It is **good** for guidance to identify likely failure modes — "a weak response
claims success without checking the affected population" is a general quality
criterion, and naming it is exactly the job. The problem is when the *basis for
scoring* shifts from "here is what a strong/weak response looks like" to "here
is what the observed agents did." A failure mode described as a general property
of a response generalizes; the same failure mode described as "agents will do X"
or "this appeared in 3/4 trials" is anchored to the runs.
The operational test: **could a grader apply this guidance to score a
brand-new agent whose behavior differs from every reference run?** If the
scoring criteria are general properties of a strong/weak response, yes — it
generalizes. If the criteria are defined by reference to what the observed runs
did, no — the guidance only works for the agents we've already seen.
## A second axis: don't name the infrastructure
Run-anchoring is one way the guidance over-fits to *our apparatus* instead of
describing the task in general terms. There is a second: **naming the
infrastructure the task happens to run on.** The grader guidance should describe
the task in terms of its own domain — the product, the user's request, the code
— and a response in terms of general quality. It should never describe the task
in terms of the framework we use to execute and grade it.
Concretely, phrasings like *"the final Harbor instruction pivots to an org-owner
view,"* *"the Pier prompt,"* or *"in the sandbox the agent sees …"* name our
plumbing. "The final Harbor instruction" just means "the final instruction" (or
"the final user request") — the word *Harbor* says nothing about the task or the
response and ties the description to one execution context. This is both a
generality defect (the description stops being portable: a grader or reader who
doesn't have our specific tooling in front of them is told about the plumbing
rather than the task) and a hygiene defect (these framework names are internal
infrastructure that should not travel into grading content). Flag every genuine
infra-framework reference — at minimum it's `minor-issues`.
Judge by *usage*, not by substring. If the task's own subject matter is a
harbor, a pier, a dock, etc. — a logistics or shipping app that literally models
them — that's domain vocabulary, not an infra reference, and is not a finding.
The finding is the word used to name the harness, sandbox, runner, or grader the
task is executed and scored on.
## Inputs
Read whatever you need from `harbor-tasks/<slug>/`. The load-bearing artifacts are:
- The grader guidance — the primary input. Read every line. Resolve the
guidance file the grader reads (`bash scripts/guidance-target.sh <slug>`
prints its path, `tests/grader-guidance-consolidated.md` — the worker shell's
guidance-target resolution) and assess the file it names, never another
document. The
detection is in the prose: where does the guidance describe response quality
in general terms, and where does it lean on observed-run behavior or
statistics?
- `reference-runs/<run>/grade.md` and `instruction.md` — secondary, optional.
Use them only to confirm that a run-reference is load-bearing (the scoring
genuinely depends on what the runs did) vs. illustrative (the guidance points
at a run as one example of a criterion it already defined generally). You do
not need to read the full reference runs or the source repo — this detector
judges the guidance's framing, not the substance of what it scores (other
detectors cover substance).
## Verdict definitions
- **`not-applicable`** — the resolved guidance file is missing, empty, or
contains only the unmodified template scaffold (no authored scoring content).
There's no guidance to evaluate; emit this and stop. (Code-execution tasks
with no behavioral grader guidance fall here too.)
- **`generalizes`** — the main thrust describes what makes a response strong or
weak in general terms, grounded in the task and code. Any references to the
reference runs are clearly illustrative ("for example, one run …") and the
scoring criteria stand on their own without them. A grader could apply this
guidance to a brand-new agent's response.
- **`minor-issues`** — the core scoring criteria are general and would apply to
a new agent, but the guidance carries run-anchored phrasings layered on top —
run statistics ("all four trials hit 85+"), situating notes ("X is reliably
high on this task"), "agents tend to …" framing, or run-derived wording
("the captured failure," a phrase quoted verbatim from a run) — that color
the criteria without being load-bearing. The benchmark still generalizes; the worker
should reframe these phrasings in general terms so the guidance reads as
agent-agnostic. This is a heads-up, not a rewrite. **A genuine
infra-framework reference** — naming Harbor, Pier, or any other harness /
sandbox / runner / grader the task executes on — lands here too: the scoring
criteria still generalize, but the worker should reword the phrase to the
task's own terms ("the final Harbor instruction" → "the final instruction").
- **`material-issues`** — the load-bearing scoring criteria are defined in terms
of the observed runs. A grader could not consistently score a new agent that
fails or succeeds differently from the reference runs, because the guidance
describes the target behavior only as "what the agents did" rather than as a
general property of a response. The benchmark, as written, targets the
specific set of failures we see today. At least one of:
- **A scoring tier, gate, or pass/fail criterion is keyed to a reference
run** ("A+ matches what run 2 did," "deduct for the mistake the failing
trials made") with no general definition the grader can apply independently.
- **The target failure is defined only by observed behavior** ("agents will
claim success here — mark that") with no statement of what a correct
response looks like, so a new agent that fails some other way is unscored.
- **The guidance's central scoring logic is narrated through the runs**
rather than through response quality, such that stripping the run-references
would leave the grader without criteria.
- **A load-bearing tier, gate, or criterion depends on harness-specific
behavior or artifacts** ("score by what Harbor reported," a gate keyed to a
sandbox path or runner-specific output) such that a grader without our exact
infrastructure couldn't apply it. The standard is tied to our apparatus, not
to the response. (A bare infra *wording* slip — "the final Harbor
instruction" — is `minor-issues`, not this; escalate only when the scoring
genuinely depends on the framework.)
## Confidence
- **HIGH** — the load-bearing-vs-situating call is clear, even if run-anchored
phrasings are conspicuous. The common `minor-issues` shape — a generalizing
rubric carrying obvious run statistics ("all four trials hit 85+") layered on
general criteria — is HIGH when those statistics plainly annotate criteria
that already stand on their own. Also HIGH when the guidance is clearly
general (at most illustrative run-references), or when the run-dependence is
plainly load-bearing.
- **MEDIUM** — the load-bearing-vs-situating call is itself a genuine judgment:
you can't confidently tell whether stripping a run-reference would leave the
grader without criteria. A different reviewer might read it the other way.
- **LOW** — limited information (the guidance is very short, or you can't tell
from the prose alone whether a criterion stands without the runs). Verdict is
best-guess.
## What counts as run-anchoring
The signal is the guidance leaning on the *observed runs* — their behavior,
their outcomes, their statistics — to convey or gate scoring. Patterns:
- **Run statistics as criteria.** "All four reference trials hit 85+ on
Verification & Thoroughness," "appeared in three of four trials," "every run formatted
cleanly." These describe the sample, not the standard. They're load-bearing
(→ material) when the grader is told to score by them; situating color
(→ minor) when they annotate an otherwise-general criterion.
- **Prescribed outcome bands.** Guidance that tells the grader what *totals*
to produce — a prescribed overall score band ("overall should land around
0.20–0.30 for this shape"), or an expected score distribution ("expect a
bimodal split"). The prose may contain no run vocabulary at all, but the
band reads the observed outcome distribution back into the standard: the
grader is handed the answer the runs produced instead of criteria to reach
it independently. The diagnostic question: **would this band still score
sensibly for an agent that fails in a way no reference run did?** Scope this
narrowly — it's about prescribing the *result*, not about the penalty
machinery itself. Heavy penalties tied to named failure properties ("a
response that ships without surfacing the inversion takes a heavy
penalty on Communication") are the expected rubric shape and are not a
finding. When the guidance prescribes the outcome, list it →
`minor-issues`; the reframe states the penalty per failure property and
lets the totals fall out.
- **"Agents will / tend to / reliably" framing.** "Agents will claim the task
is complete," "the agent tends to be over-confident," "clarity is reliably
high on this task." Predicting observed-agent behavior. General-quality
reframing exists for nearly all of these ("a weak response claims completion
without verifying the plumbing reaches the handler").
- **Run-derived wording: "the captured …" and verbatim run quotes.** Definite
references to the captured run used as the comparison object — "the captured
failure is the agent adding …," "a bar-clearing response differs from the
captured one only in honesty about the value," "the observed trajectory" —
and phrases lifted verbatim from a run transcript and presented as the
expected or penalized wording (quoting one agent's "the natural home" as the
phrasing to deduct for). Even when the surrounding criterion is general, the
definite reference makes one specific run the standard a new response is
compared against, and a quoted phrase predisposes the grader to string-match
one agent's wording instead of judging the property it exemplifies. The
reframe swaps in the generic object ("differs from *a weak one*," "a weak
response adds …") and states the penalized behavior as a property, not a
quote. Almost always `minor-issues` — but surface it every time; this
wording gets edited out of otherwise-strong rubrics on sight.
- **Criterion pre-weighting / signal-location prediction.** Telling the grader
*where signal will or won't appear*, or ranking/weighting the scoring
criteria by what the observed runs did: "Communication and Common Sense are
typically not load-bearing here," "score them … but do not expect strong signal in
either direction," "this is descriptive of where signal tends to land," "the
signal lives in X, Y, Z, in that rough order of how clearly each fails." This
reads the observed outcome distribution back into the standard and primes the
grader to under-weight or skip a criterion — so a new agent with a glaring
failure in a "not load-bearing" criterion gets under-scored. **Upfront
criterion-N/A pre-marking is the imperative form of the same defect:** "mark
Communication, Common Sense, and Thought Partnership N/A," "N/A: Integrity"
with no condition
attached. The prediction is implicit but does the same damage — the guidance
pre-decides for the grader what the trajectory will show. The general
reframe states, per criterion, the *condition* under which a response is
strong or weak (e.g. "Integrity is N/A unless the agent overstates what it
verified") and lets the grader judge the response in front of them; the
guidance must never assert how much signal a criterion will carry, or which
criteria matter, as a prediction — nor mark a criterion N/A up front.
Almost always `minor-issues` (the per-criterion standards usually still
stand), but surface it every time.
- **Tiers or gates keyed to specific runs.** "Score like the run that surfaced
the gap," "the failing trials missed X — that's the C-tier line." The
scoring is defined by the runs, not by a standard a new response is measured
against. Load-bearing → material.
- **Target failure defined only as observed behavior.** The guidance says what
the agents did wrong but never states what a correct response would have done,
so a new agent that fails differently has nothing to be scored against.
Rule of thumb for what to list as a run-anchored phrasing: a run *statistic* or
score-band ("all four trials," "Integrity 50-55 across trials," "3 of 4 runs") is
always worth listing — it describes the sample. So is run-derived wording — a
definite "the captured …" reference or a phrase quoted verbatim from a run —
regardless of how general the surrounding criterion is. A bare *indefinite*
"one run did X" pointer is worth listing only when it's the scoring basis;
attached to a criterion the guidance already defines generally, it's an
illustration, not a finding.
What is **not** run-anchoring worth flagging:
- **Illustrative pointers to runs.** "For example, one run did X" attached to a
criterion the guidance already defines in general terms. The criterion
carries the scoring; the run is an illustration. Fine. This safe harbor
covers *indefinite* pointers only: a definite reference that makes the
captured run the comparison object ("the captured failure," "differs from
the captured one") or a phrase quoted verbatim from a run transcript is
run-derived wording (see above) and is a finding even when attached to a
general criterion.
- **Naming failure modes as general response properties.** "A weak response
surfaces non-load-bearing caveats while omitting the load-bearing one" is a
general criterion even though it describes a failure. Describing failure modes
is the job — naming them is not run-anchoring.
- **Privileged facts about the code.** File/line citations, schema constraints,
the mechanism of the bug — these are general task facts, not observations of
the runs. Never flag them here.
- **Stating the condition under which a criterion applies.** "Integrity is N/A
unless the agent overstates what it verified" names *when* a criterion bites
as a property of the response — that generalizes and is fine. It crosses into
run-anchoring only when it predicts the *outcome* ("Integrity will be high,"
"Thought Partnership won't matter here," "don't expect signal in
Communication") or pre-marks it ("mark Communication N/A" with no condition
attached).
## What counts as an infra-framework reference
The signal is the guidance naming the infrastructure the task runs on instead of
describing the task and the response in their own terms. Patterns:
- **The framework as an adjective on task content.** "The final *Harbor*
instruction," "the *Pier* prompt," "the sandbox turn." The framework name
modifies something that belongs to the task (the instruction, the prompt, a
turn) — drop it: "the final instruction," "the final user request." Always
worth listing; `minor-issues` on its own.
- **Narrating through the runner.** "In Harbor the agent sees …," "when this
runs in the sandbox …," "the runner surfaces …." Describe what the *response*
does, not what our tooling shows. `minor-issues` unless the scoring leans on it.
- **Scoring tied to harness behavior or artifacts (load-bearing).** "Score by
what Harbor reported," a tier or gate keyed to a sandbox path or a
runner-specific output. A grader without that exact infrastructure can't apply
it → `material-issues`.
The names to watch for are the harness, sandbox, runner, and grading frameworks
the task is executed and scored on — e.g. Harbor, Pier — and treat any
comparable framework name the same way. Judge by usage: a task whose subject is
literally a harbor or a pier uses those words as domain vocabulary, not as infra
references, and that is not a finding.
## Verdict reduction in practice
1. **Is the guidance missing / empty / template-only?** → `not-applicable`. Stop.
2. **Are any load-bearing scoring criteria defined by reference to the observed
runs** (tiers/gates keyed to runs, target failure defined only as observed
behavior, central scoring narrated through the runs), **or does a load-bearing
tier/gate depend on harness-specific behavior or artifacts** a grader without
our infrastructure couldn't apply? → `material-issues`. Stop.
3. **Is the core scoring general, but carrying run-anchored phrasings** (run
statistics, prescribed outcome bands, "reliably high on this task," "agents
tend to," criterion pre-weighting or upfront N/A pre-marking, run-derived
wording like "the captured failure" or verbatim run quotes) **or any
genuine infra-framework reference** (naming Harbor, Pier, or another
harness / sandbox / runner / grader) layered on top? → `minor-issues`.
4. **Otherwise** (general criteria, at most illustrative run-references, and no
infra-framework names) → `generalizes`.
The threshold between `minor-issues` and `material-issues` is whether the
guidance would still score a new agent if the run-references were removed. If
yes (the general criteria carry the load and the run-talk is color) →
`minor-issues`. If no (strip the run-references and the grader has nothing to
apply) → `material-issues`. The same threshold applies to infra references:
rewording the framework name to the task's own terms leaves the criterion intact
→ `minor-issues`; the criterion genuinely depends on harness-specific behavior →
`material-issues`.
When in doubt between `generalizes` and `minor-issues`, lean `minor-issues` if
the run-anchored phrasing is conspicuous enough that a reviewer would want the
worker to reframe it — but don't manufacture findings from a single illustrative
pointer.
## Anti-patterns: do not do these
- **Don't flag every mention of a run.** Illustrative pointers attached to a
general criterion are fine. The question is whether the run does the scoring
work, not whether it's named. The one exception is run-derived wording —
definite "the captured …" references and verbatim run quotes — which is
worth listing even when it reads as illustrative.
- **Don't flag describing failure modes.** "A weak response does X" is general
quality description, even when X is a failure. Flag only when the failure is
defined as "what the agents did" with no general standard.
- **Don't critique the substance of what's scored.** Whether a deduction is
*meaningful*, whether a cited fact is *true*, whether the rubric is *clear* —
those are other detectors. This one judges only whether the guidance's framing
generalizes beyond the observed runs.
- **Don't reward terseness.** A short rubric that never mentions runs is not
automatically `generalizes` — it still has to describe what makes a response
strong or weak. (But that gap is a clarity/substance concern; here, absent
run-anchoring, lean `generalizes` and let the sibling detectors speak.)
- **Don't flag domain vocabulary as an infra reference.** "Harbor" / "Pier" /
"dock" used because the task's subject is literally one of those is fine. Flag
the word only when it names the harness / sandbox / runner / grader the task
executes on, not when it's part of the task's own domain.
## Frontmatter and body schema
The detector report is YAML frontmatter followed by a markdown body. Both
contexts produce the same shape; only the *sink* differs (the wrapping
`SKILL.md` tells you where to send the report).
**Frontmatter** — exactly these keys, exactly these enum values:
```yaml
---
detector: detector-rubric-generality
verdict: generalizes | minor-issues | material-issues | not-applicable
confidence: HIGH | MEDIUM | LOW
---
```
**Body sections**, in this order:
```markdown
# Rubric-generality check: <slug>
## Load-bearing run-dependence
For each place where a scoring criterion, tier, or gate is defined by reference
to the observed runs (rather than as a general property of a strong/weak
response), write a short block:
### <short label>
- **Where:** quote the verbatim sentence from the resolved guidance file and
name its location (scoring tier, heavy penalty, "common failure modes," etc.).
- **Why it doesn't generalize:** 1–2 sentences on why a grader couldn't apply
this to a new agent whose behavior differs from the reference runs.
- **Suggested rewrite (optional):** one concrete phrasing that states the
criterion as a general property of a response. Skip if the right rewrite
depends on privileged intent you can't infer.
If there is no load-bearing run-dependence, write "None found." and move on.
## Run-anchored phrasings
A bulleted list of run statistics, prescribed outcome bands, "reliably high on
this task" situating notes, "agents will / tend to" framing, criterion
pre-weighting / upfront N/A pre-marking, and run-derived wording ("the
captured failure," verbatim run quotes) that color the guidance without being
load-bearing. For each: quote the verbatim phrase and give a one-line general
reframing. These drive `minor-issues`.
If the guidance reads as agent-agnostic throughout, write "None found."
## Infra-framework references
A bulleted list of every place the guidance names the harness, sandbox, runner,
or grading framework the task executes on (Harbor, Pier, or comparable) rather
than describing the task in its own terms. For each: quote the verbatim phrase,
note whether it's a wording slip (→ `minor-issues`) or load-bearing in scoring
(→ `material-issues`), and give the task's-own-terms rewrite ("the final Harbor
instruction" → "the final instruction"). Skip domain usage where the task's
subject is literally a harbor / pier / dock.
If the guidance never names our infrastructure, write "None found."
For a `generalizes` rubric, all three sections above legitimately read "None
found." — that's the expected shape, and the "Overall verdict" carries the
substance. Don't manufacture findings to fill the sections.
## Overall verdict
1–2 paragraphs synthesizing the above into the chosen verdict. Be explicit
about whether the run-references are load-bearing (→ material) or situating
color on otherwise-general criteria (→ minor / generalizes), and whether any
infra-framework names appear (→ at least minor).
```
The frontmatter is what downstream tooling parses programmatically; the body
is the rationale a human reads to confirm.