restored entire zip and config'd

This commit is contained in:
2026-10-07 15:37:45 -04:00
parent 4bd9264a29
commit 5ca605dcce
219 changed files with 43370 additions and 0 deletions

View File

@@ -0,0 +1,67 @@
---
name: detector-rubric-coverage
description: |
Self-check that your atomic rubric fully captures your holistic rubric.
Verifies four things. Every load-bearing requirement, penalty, and
"do not penalize" rule in the holistic rubric maps to a criterion. No
criterion invents a requirement or an answer-key fact the holistic rubric
does not support. The holistic rubric's context sections survive in
`tests/grader-context.md`. Every heavy penalty that targets the overall
score is encoded as a crux criterion, or at `certain_dealbreaker` once two
criteria already carry crux. Restructuring is never flagged; only
content differences that change scoring are. Reads the holistic rubric,
`tests/atomic-rubric.yaml` (or `tests/rubrics.yaml`), and
`tests/grader-context.md`. Emits `not-applicable` when the task has no
atomic rubric yet.
allowed-tools: Bash, Read, Write
---
# Rubric-coverage detector
This skill checks that your atomic rubric and your holistic rubric express the
same task. The atomic rubric restructures the holistic rubric into criteria.
It must not lose scoring content, and it must not add scoring content.
The failure shapes to catch:
- **A lost requirement or penalty.** The holistic rubric requires something,
or penalizes something, and no criterion captures it. A response the
holistic rubric would mark down now scores clean.
- **A lost "do not penalize" rule.** The holistic rubric protects a behavior,
and the criteria drop the protection. The atomic rubric now penalizes what
the holistic rubric permits.
- **Invented content.** A criterion requires something the holistic rubric
never asks for, or states an answer-key fact with no source in the holistic
rubric or the context document.
- **Lost context.** A ground-truth fact that criteria rely on is missing from
both `tests/grader-context.md` and the criteria themselves.
- **A crux mismatch.** The holistic rubric applies a heavy penalty against
the overall score, and no criterion carries `severity: crux` to encode it.
A task carries at most two crux criteria; once two are designated, a
further overall-score penalty is correctly encoded at `certain_dealbreaker`.
Read these before deciding:
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
2. `.claude/skills/detector-rubric-coverage/core.md` — what counts as a coverage gap versus invented content, the crux-alignment rule, what is deliberately not a finding, verdict definitions, and the body schema.
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
## Acting on the verdict
- **`clear`** — the atomic rubric fully captures the holistic rubric. A
grader scoring from either form would land in the same place.
- **`minor-issues`** — the load-bearing mapping is sound, but some
non-load-bearing content drifted. Read the findings and tighten the
conversion. There is no need to rebuild the rubric.
- **`material-issues`** — a load-bearing requirement, penalty, or protection
is missing, a criterion invents content, needed context is gone, or a
heavy penalty against the overall score has no criterion encoding it at
`crux` (or at `certain_dealbreaker` once two crux criteria exist). Fix the
named findings in the atomic rubric. If a finding reveals that the
holistic rubric itself needs the change, edit the holistic rubric first
and then re-convert, so the two forms stay in agreement. Re-run this
skill after editing either file.
- **`not-applicable`** — the task has no atomic rubric yet, or no holistic
rubric to compare it against. Write the missing rubric first, then come
back to this skill.

View File

@@ -0,0 +1,318 @@
# Rubric-coverage detector — core
This file is the canonical, context-neutral content for the detector-rubric-coverage
detector. It defines what counts as a coverage gap between a task's holistic
rubric and its atomic rubric, what counts as invented content, the verdict
enum, and the output schema. It is read in two contexts — the base repo's
review pipeline and the worker toolkit's self-check — so nothing here should
reference downstream storage details.
## What this detector is for
A task carries its grading requirements in two forms. The **holistic rubric** is
the prose document the grader reads. The **atomic rubric** is the same
requirements expressed as a list of criteria in `tests/atomic-rubric.yaml`,
each one independently judgeable, with the generalized context sections
preserved in the companion document `tests/grader-context.md`. The two forms
must express the same task. The atomic rubric restructures the holistic
rubric; it does not extend it, and it does not shrink it.
This detector verifies that equivalence in both directions:
1. **Nothing load-bearing is lost.** Every requirement, penalty, and
non-trigger in the holistic rubric that affects scoring maps to a criterion,
or to a criterion's elaboration.
2. **Nothing is invented.** No criterion introduces a requirement, an
answer-key fact, or a severity that the holistic rubric does not support.
3. **Context survives.** The holistic rubric's context sections (task context,
business context, ground truth) are preserved in `tests/grader-context.md`,
so criteria that lean on those facts still have them available.
4. **Crux designations match.** The `crux` severity tier is reserved for a
criterion that encodes a heavy penalty of the holistic rubric targeting the
overall score, and a task carries at most two crux criteria. A heavy
penalty against the overall score with no criterion encoding it is a
material gap. When the holistic rubric carries more overall-score heavy
penalties than the cap allows, the two that define the task's failure mode
carry `crux` and the rest carry `certain_dealbreaker`; a surplus penalty
encoded that way is covered, not mismatched.
This detector does **not** judge:
- Whether the criteria are well-formed as artifacts. Schema validity,
atomicity, and phrasing belong to the detector-rubric-form detector.
- Whether the holistic rubric's substance is right. Meaningfulness, factual
accuracy, prose clarity, and generality belong to their own detectors.
- Style differences between the two forms. Restructuring is the point of the
conversion. A coverage finding requires a scoring-relevant difference in
content, never a difference in shape.
## Inputs
Read from `harbor-tasks/<slug>/`:
- The holistic rubric — primary. Resolve it with
`bash scripts/guidance-target.sh <slug>`, which prints the path to the file
the grader reads (`tests/holistic-rubric.md`; a task packaged under an
earlier release carries it as `tests/grader-guidance-consolidated.md` or
`tests/grader-guidance.md`). Read every line of the file the resolver names,
and never assess a different document.
- `tests/atomic-rubric.yaml` — primary. A task packaged under an earlier
release carries the same artifact as `tests/rubrics.yaml`; when
`tests/atomic-rubric.yaml` is absent, assess `tests/rubrics.yaml`.
- `tests/grader-context.md` — the atomic rubric's companion context document.
Read it in full; it is where dropped holistic context is supposed to have
landed.
- `instruction.md` — secondary. Use it to confirm that a holistic requirement
is load-bearing for scoring before flagging its absence as material.
You do not need the workspace, the reference runs, or the source repo. This
detector compares two documents; it does not verify their claims against code.
## Verdict definitions
- **`not-applicable`** — there is no atomic rubric to assess (neither
`tests/atomic-rubric.yaml` nor `tests/rubrics.yaml` exists), or there is no
holistic rubric to compare it against. Name the missing side in the body,
emit this verdict, and stop.
- **`clear`** — the atomic rubric fully captures the holistic rubric. Every
load-bearing requirement, penalty, and non-trigger maps to a criterion; no
criterion invents content; the context sections survive in
`tests/grader-context.md`; crux designations line up with the holistic
rubric's overall-score heavy penalties within the two-crux cap.
- **`minor-issues`** — the mapping is sound where it matters, but
non-load-bearing content drifted: background nuance was condensed away, a
fulfillment shape from the holistic prose did not make it into an
elaboration, or a criterion carries harmless connective prose with no
holistic source. A grader scoring from either form would land in the same
place; the worker should still tighten the conversion.
- **`material-issues`** — at least one of:
- **A load-bearing gap.** A requirement, penalty, or non-trigger that
affects scoring in the holistic rubric has no criterion that captures it.
- **Invented content.** A criterion requires something the holistic rubric
never requires, or states an answer-key fact with no basis in the holistic
rubric or the context document.
- **Context loss criteria depend on.** A ground-truth or context fact that
criteria lean on is present in the holistic rubric but absent from both
`tests/grader-context.md` and the criteria themselves.
- **A crux mismatch.** A heavy penalty in the holistic rubric that targets
the overall score has no crux criterion encoding it, unless two criteria
already carry `crux` and the penalty is encoded at `certain_dealbreaker`.
## Confidence
- **HIGH** — the mapping is unambiguous in both directions, or a gap is plain
to see (a whole heavy penalty with no criterion anywhere near it).
- **MEDIUM** — at least one call rests on judging whether a clause is
load-bearing or whether an elaboration's coverage of it is close enough.
- **LOW** — limited information (a very short holistic rubric, an unfamiliar
domain, or heavy restructuring that makes the mapping genuinely hard to
trace).
## What counts as a coverage gap (holistic → atomic)
Walk the holistic rubric clause by clause and locate each of these in the
atomic rubric:
- **Requirements.** Everything the holistic rubric says a response should do,
surface, state, or include. Tier prose counts: the content of a strong-tier
description is a set of requirements, and each load-bearing one needs a
criterion. The tier scaffolding itself does not need to survive; its content
does.
- **Penalties.** Every deduction the holistic rubric directs at a criterion or
at the overall score. The penalty's *trigger* must be captured by a
criterion whose failure corresponds to it. The penalty's *magnitude* does
not survive, by design — the atomic rubric expresses weight through
`category` and `severity`, so check that the assigned severity is
proportionate to the holistic penalty's weight. A penalty that names both a
criterion and the overall score is one dealbreaker, not two; one criterion
captures it.
- **Non-triggers.** Statements that protect behavior from penalties: "do not
penalize X", "X is acceptable", "either A or B clears the bar", "when the
condition is unmet, this does not apply". These prevent over-penalizing.
When a non-trigger is dropped, the atomic rubric penalizes what the holistic
rubric permits — a criterion phrased without the exception, or missing the
either/or fork, is a gap even though every requirement is present. Look for
the protection in the criterion's guideline (conditional or either/or
phrasing) or its elaboration (fulfillment shapes, does-not-fire notes).
- **Answer-key facts.** The specific facts, citations, and mechanisms the
holistic rubric supplies as ground truth. Each must survive either inline in
the criterion that grades it or in `tests/grader-context.md`. A criterion
that says "the response should identify the defect" whose defect is defined
nowhere in the atomic package has lost its key.
- **Conditions and qualifiers.** A penalty the holistic rubric applies
conditionally must not become an unconditional criterion, and a scoped
requirement must not become a blanket one. Compare qualifiers clause by
clause.
## What counts as invented content (atomic → holistic)
Walk the criteria and check each against the holistic rubric and the context
document:
- **New requirements.** A guideline requiring something the holistic rubric
never asks for. The conversion is not the place to add scope; a genuinely
missing requirement belongs in the holistic rubric first, so both forms stay
in agreement.
- **New answer-key facts.** A bolded key, citation, or mechanism stated in a
criterion with no support in the holistic rubric or the context document.
Whether such a fact is *true* is a different detector's job; here the
finding is that the two forms no longer say the same thing. Tightening an
existing fact (adding a file and line to a mechanism the holistic rubric
already names) is not invention.
- **Severity without basis.** A `crux` criterion with no heavy penalty against
the overall score behind it in the holistic rubric. Crux weighting dominates
the aggregate score, so an unsupported crux re-weights the whole rubric;
treat it as material when it dominates scoring and as minor when the backing
penalty is arguable (for example, a moderate overall-score penalty, which
belongs at a normal severity tier rather than crux).
- **New requirements smuggled into elaboration.** An elaboration is for
fulfillment shapes and clarification. When it adds a requirement, check the
holistic rubric for it; content with no holistic basis is a coverage finding
here, and the guideline-vs-elaboration placement is the
detector-rubric-form detector's lane.
## What is NOT a finding
- **Restructuring.** Tiers dissolving into criteria, strong/weak prose
becoming fulfillment shapes in elaborations, one holistic paragraph
collapsing into one criterion, or one holistic penalty becoming a base
criterion plus a worse-variant criterion that fails in addition to it
(paired escalation is a sanctioned encoding of "this variant is strictly
worse").
- **Dropped penalty magnitudes.** The atomic rubric carries no numeric
penalty amounts by design. A "subtract roughly 0.35" that survives only as
a severity tier is the conversion working.
- **Dropped generic scoring mechanics.** Floor-at-zero notes, "penalties are
never ceilings", and similar task-independent mechanics belong to the shared
grading machinery, not to per-task criteria.
- **Condensed context.** `tests/grader-context.md` may compress the holistic
rubric's context prose. The finding is a lost *fact* that criteria rely on,
never lost word count.
- **Wording differences with the same scoring effect.** Judge what a grader
would do, not whether the sentences match.
- **A duplicated file set.** Both rubric forms sitting side by side in
`tests/` is the intended package shape, not redundancy.
## How to work
1. Read the holistic rubric end to end and list its load-bearing clauses:
requirements, penalties (with their targets and conditions), non-triggers,
and answer-key facts.
2. Read `tests/atomic-rubric.yaml` (or `tests/rubrics.yaml`) end to end,
guideline and elaboration both, and `tests/grader-context.md` in full.
3. Map each holistic clause to the criterion or context section that captures
it. Record the criterion `id`. A clause may map to several criteria and
several clauses may map to one criterion; what matters is that the scoring
content lands somewhere.
4. Sweep the reverse direction: for each criterion, find its holistic source.
5. Check the crux designations against the holistic rubric's heavy penalties
that target the overall score, in both directions, allowing for the
two-crux cap: once two criteria carry `crux`, a further overall-score
penalty is correctly encoded at `certain_dealbreaker`.
6. Reduce to a verdict per the definitions above.
Never assert a mapping you have not traced. If you claim a clause is covered,
name the criterion id that covers it.
## Anti-patterns: do not do these
- **Don't flag the restructuring itself.** The two forms are supposed to look
different. Only content differences with scoring effect are findings.
- **Don't demand one criterion per holistic sentence.** Several parallel facts
from one derivation may live in one criterion, and one dense holistic
paragraph may fan out into several criteria.
- **Don't paraphrase away qualifiers.** Quote the holistic clause verbatim,
conditions included, and quote the criterion text verbatim next to it.
Describing a conditionally-applied penalty as unconditional is a factual
error in the report.
- **Don't re-litigate substance.** "This requirement is an over-ask" is the
meaningfulness detector's lane. Here the holistic rubric is the reference,
right or wrong.
- **Don't treat sharpened citations as invention.** A criterion may pin an
existing holistic fact to a file and line. Invention means a *new* fact or
requirement, not a more precise statement of an existing one.
- **Don't count a both-targets penalty twice.** A holistic dealbreaker may
direct its penalty at a criterion and at the overall score together; that is
one dealbreaker, encoded once.
## Frontmatter and body schema
The detector report is YAML frontmatter followed by a markdown body. Both
contexts produce the same shape; only the *sink* differs (the wrapping
`SKILL.md` tells you where to send the report).
**Frontmatter** — exactly these keys, exactly these enum values:
```yaml
---
detector: detector-rubric-coverage
verdict: clear | minor-issues | material-issues | not-applicable
confidence: HIGH | MEDIUM | LOW
---
```
**Body sections**, in this order:
```markdown
# Rubric-coverage check: <slug>
Assessed: <resolved holistic rubric path> against <atomic rubric path> and tests/grader-context.md
## Coverage map
One table row per load-bearing holistic clause (requirement, penalty, or
non-trigger):
| Holistic clause (short, verbatim key phrase) | Criterion id(s) | Status |
| --- | --- | --- |
| "…" | criterion-id | covered / partial / missing |
## Coverage gaps
One block per `partial` or `missing` row:
### <short label>
- **Holistic clause:** the verbatim sentence(s) and their location (section
or heading in the holistic rubric).
- **Closest criterion:** the criterion id that comes nearest, quoted, or a
statement that none exists.
- **What is lost:** 1-2 sentences on the scoring effect of the gap — which
responses now score differently under the atomic rubric.
- **Suggested criterion (optional):** a concrete guideline that would close
the gap.
If there are no gaps, write "None found." and move on.
## Invented content
One block per criterion (or elaboration) with content the holistic rubric
does not support: quote the criterion text verbatim, state what was searched
for in the holistic rubric and the context document, and name the scoring
effect. If there is none, write "None found."
## Context integrity
Whether the holistic rubric's context sections survive in
tests/grader-context.md. Name any fact that criteria rely on that is missing
from both the context document and the criteria. If everything survives,
say so.
## Crux alignment
List every heavy penalty in the holistic rubric that targets the overall
score and the criterion encoding it (`crux`, or `certain_dealbreaker` once
two crux criteria are designated), and every crux criterion and the penalty
backing it. Flag mismatches in either direction.
## Overall verdict
1-2 paragraphs reducing the findings to the chosen verdict. Be explicit about
which direction (gap, invention, context loss, crux mismatch) drove the call.
```
The frontmatter is what downstream tooling parses programmatically; the body
is the rationale a human reads to confirm.