added 260907 version of worker toolkit
This commit is contained in:
@@ -0,0 +1,67 @@
|
||||
---
|
||||
name: detector-rubric-coverage
|
||||
description: |
|
||||
Self-check that your atomic rubric fully captures your holistic rubric.
|
||||
Verifies four things. Every load-bearing requirement, penalty, and
|
||||
"do not penalize" rule in the holistic rubric maps to a criterion. No
|
||||
criterion invents a requirement or an answer-key fact the holistic rubric
|
||||
does not support. The holistic rubric's context sections survive in
|
||||
`tests/grader-context.md`. Every heavy penalty that targets the overall
|
||||
score is encoded as a crux criterion, or at `certain_dealbreaker` once two
|
||||
criteria already carry crux. Restructuring is never flagged; only
|
||||
content differences that change scoring are. Reads the holistic rubric,
|
||||
`tests/atomic-rubric.yaml` (or `tests/rubrics.yaml`), and
|
||||
`tests/grader-context.md`. Emits `not-applicable` when the task has no
|
||||
atomic rubric yet.
|
||||
allowed-tools: Bash, Read, Write
|
||||
---
|
||||
|
||||
# Rubric-coverage detector
|
||||
|
||||
This skill checks that your atomic rubric and your holistic rubric express the
|
||||
same task. The atomic rubric restructures the holistic rubric into criteria.
|
||||
It must not lose scoring content, and it must not add scoring content.
|
||||
|
||||
The failure shapes to catch:
|
||||
|
||||
- **A lost requirement or penalty.** The holistic rubric requires something,
|
||||
or penalizes something, and no criterion captures it. A response the
|
||||
holistic rubric would mark down now scores clean.
|
||||
- **A lost "do not penalize" rule.** The holistic rubric protects a behavior,
|
||||
and the criteria drop the protection. The atomic rubric now penalizes what
|
||||
the holistic rubric permits.
|
||||
- **Invented content.** A criterion requires something the holistic rubric
|
||||
never asks for, or states an answer-key fact with no source in the holistic
|
||||
rubric or the context document.
|
||||
- **Lost context.** A ground-truth fact that criteria rely on is missing from
|
||||
both `tests/grader-context.md` and the criteria themselves.
|
||||
- **A crux mismatch.** The holistic rubric applies a heavy penalty against
|
||||
the overall score, and no criterion carries `severity: crux` to encode it.
|
||||
A task carries at most two crux criteria; once two are designated, a
|
||||
further overall-score penalty is correctly encoded at `certain_dealbreaker`.
|
||||
|
||||
Read these before deciding:
|
||||
|
||||
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
|
||||
2. `.claude/skills/detector-rubric-coverage/core.md` — what counts as a coverage gap versus invented content, the crux-alignment rule, what is deliberately not a finding, verdict definitions, and the body schema.
|
||||
|
||||
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
|
||||
|
||||
## Acting on the verdict
|
||||
|
||||
- **`clear`** — the atomic rubric fully captures the holistic rubric. A
|
||||
grader scoring from either form would land in the same place.
|
||||
- **`minor-issues`** — the load-bearing mapping is sound, but some
|
||||
non-load-bearing content drifted. Read the findings and tighten the
|
||||
conversion. There is no need to rebuild the rubric.
|
||||
- **`material-issues`** — a load-bearing requirement, penalty, or protection
|
||||
is missing, a criterion invents content, needed context is gone, or a
|
||||
heavy penalty against the overall score has no criterion encoding it at
|
||||
`crux` (or at `certain_dealbreaker` once two crux criteria exist). Fix the
|
||||
named findings in the atomic rubric. If a finding reveals that the
|
||||
holistic rubric itself needs the change, edit the holistic rubric first
|
||||
and then re-convert, so the two forms stay in agreement. Re-run this
|
||||
skill after editing either file.
|
||||
- **`not-applicable`** — the task has no atomic rubric yet, or no holistic
|
||||
rubric to compare it against. Write the missing rubric first, then come
|
||||
back to this skill.
|
||||
@@ -0,0 +1,318 @@
|
||||
# Rubric-coverage detector — core
|
||||
|
||||
This file is the canonical, context-neutral content for the detector-rubric-coverage
|
||||
detector. It defines what counts as a coverage gap between a task's holistic
|
||||
rubric and its atomic rubric, what counts as invented content, the verdict
|
||||
enum, and the output schema. It is read in two contexts — the base repo's
|
||||
review pipeline and the worker toolkit's self-check — so nothing here should
|
||||
reference downstream storage details.
|
||||
|
||||
## What this detector is for
|
||||
|
||||
A task carries its grading requirements in two forms. The **holistic rubric** is
|
||||
the prose document the grader reads. The **atomic rubric** is the same
|
||||
requirements expressed as a list of criteria in `tests/atomic-rubric.yaml`,
|
||||
each one independently judgeable, with the generalized context sections
|
||||
preserved in the companion document `tests/grader-context.md`. The two forms
|
||||
must express the same task. The atomic rubric restructures the holistic
|
||||
rubric; it does not extend it, and it does not shrink it.
|
||||
|
||||
This detector verifies that equivalence in both directions:
|
||||
|
||||
1. **Nothing load-bearing is lost.** Every requirement, penalty, and
|
||||
non-trigger in the holistic rubric that affects scoring maps to a criterion,
|
||||
or to a criterion's elaboration.
|
||||
2. **Nothing is invented.** No criterion introduces a requirement, an
|
||||
answer-key fact, or a severity that the holistic rubric does not support.
|
||||
3. **Context survives.** The holistic rubric's context sections (task context,
|
||||
business context, ground truth) are preserved in `tests/grader-context.md`,
|
||||
so criteria that lean on those facts still have them available.
|
||||
4. **Crux designations match.** The `crux` severity tier is reserved for a
|
||||
criterion that encodes a heavy penalty of the holistic rubric targeting the
|
||||
overall score, and a task carries at most two crux criteria. A heavy
|
||||
penalty against the overall score with no criterion encoding it is a
|
||||
material gap. When the holistic rubric carries more overall-score heavy
|
||||
penalties than the cap allows, the two that define the task's failure mode
|
||||
carry `crux` and the rest carry `certain_dealbreaker`; a surplus penalty
|
||||
encoded that way is covered, not mismatched.
|
||||
|
||||
This detector does **not** judge:
|
||||
|
||||
- Whether the criteria are well-formed as artifacts. Schema validity,
|
||||
atomicity, and phrasing belong to the detector-rubric-form detector.
|
||||
- Whether the holistic rubric's substance is right. Meaningfulness, factual
|
||||
accuracy, prose clarity, and generality belong to their own detectors.
|
||||
- Style differences between the two forms. Restructuring is the point of the
|
||||
conversion. A coverage finding requires a scoring-relevant difference in
|
||||
content, never a difference in shape.
|
||||
|
||||
## Inputs
|
||||
|
||||
Read from `harbor-tasks/<slug>/`:
|
||||
|
||||
- The holistic rubric — primary. Resolve it with
|
||||
`bash scripts/guidance-target.sh <slug>`, which prints the path to the file
|
||||
the grader reads (`tests/holistic-rubric.md`; a task packaged under an
|
||||
earlier release carries it as `tests/grader-guidance-consolidated.md` or
|
||||
`tests/grader-guidance.md`). Read every line of the file the resolver names,
|
||||
and never assess a different document.
|
||||
- `tests/atomic-rubric.yaml` — primary. A task packaged under an earlier
|
||||
release carries the same artifact as `tests/rubrics.yaml`; when
|
||||
`tests/atomic-rubric.yaml` is absent, assess `tests/rubrics.yaml`.
|
||||
- `tests/grader-context.md` — the atomic rubric's companion context document.
|
||||
Read it in full; it is where dropped holistic context is supposed to have
|
||||
landed.
|
||||
- `instruction.md` — secondary. Use it to confirm that a holistic requirement
|
||||
is load-bearing for scoring before flagging its absence as material.
|
||||
|
||||
You do not need the workspace, the reference runs, or the source repo. This
|
||||
detector compares two documents; it does not verify their claims against code.
|
||||
|
||||
## Verdict definitions
|
||||
|
||||
- **`not-applicable`** — there is no atomic rubric to assess (neither
|
||||
`tests/atomic-rubric.yaml` nor `tests/rubrics.yaml` exists), or there is no
|
||||
holistic rubric to compare it against. Name the missing side in the body,
|
||||
emit this verdict, and stop.
|
||||
|
||||
- **`clear`** — the atomic rubric fully captures the holistic rubric. Every
|
||||
load-bearing requirement, penalty, and non-trigger maps to a criterion; no
|
||||
criterion invents content; the context sections survive in
|
||||
`tests/grader-context.md`; crux designations line up with the holistic
|
||||
rubric's overall-score heavy penalties within the two-crux cap.
|
||||
|
||||
- **`minor-issues`** — the mapping is sound where it matters, but
|
||||
non-load-bearing content drifted: background nuance was condensed away, a
|
||||
fulfillment shape from the holistic prose did not make it into an
|
||||
elaboration, or a criterion carries harmless connective prose with no
|
||||
holistic source. A grader scoring from either form would land in the same
|
||||
place; the worker should still tighten the conversion.
|
||||
|
||||
- **`material-issues`** — at least one of:
|
||||
- **A load-bearing gap.** A requirement, penalty, or non-trigger that
|
||||
affects scoring in the holistic rubric has no criterion that captures it.
|
||||
- **Invented content.** A criterion requires something the holistic rubric
|
||||
never requires, or states an answer-key fact with no basis in the holistic
|
||||
rubric or the context document.
|
||||
- **Context loss criteria depend on.** A ground-truth or context fact that
|
||||
criteria lean on is present in the holistic rubric but absent from both
|
||||
`tests/grader-context.md` and the criteria themselves.
|
||||
- **A crux mismatch.** A heavy penalty in the holistic rubric that targets
|
||||
the overall score has no crux criterion encoding it, unless two criteria
|
||||
already carry `crux` and the penalty is encoded at `certain_dealbreaker`.
|
||||
|
||||
## Confidence
|
||||
|
||||
- **HIGH** — the mapping is unambiguous in both directions, or a gap is plain
|
||||
to see (a whole heavy penalty with no criterion anywhere near it).
|
||||
- **MEDIUM** — at least one call rests on judging whether a clause is
|
||||
load-bearing or whether an elaboration's coverage of it is close enough.
|
||||
- **LOW** — limited information (a very short holistic rubric, an unfamiliar
|
||||
domain, or heavy restructuring that makes the mapping genuinely hard to
|
||||
trace).
|
||||
|
||||
## What counts as a coverage gap (holistic → atomic)
|
||||
|
||||
Walk the holistic rubric clause by clause and locate each of these in the
|
||||
atomic rubric:
|
||||
|
||||
- **Requirements.** Everything the holistic rubric says a response should do,
|
||||
surface, state, or include. Tier prose counts: the content of a strong-tier
|
||||
description is a set of requirements, and each load-bearing one needs a
|
||||
criterion. The tier scaffolding itself does not need to survive; its content
|
||||
does.
|
||||
- **Penalties.** Every deduction the holistic rubric directs at a criterion or
|
||||
at the overall score. The penalty's *trigger* must be captured by a
|
||||
criterion whose failure corresponds to it. The penalty's *magnitude* does
|
||||
not survive, by design — the atomic rubric expresses weight through
|
||||
`category` and `severity`, so check that the assigned severity is
|
||||
proportionate to the holistic penalty's weight. A penalty that names both a
|
||||
criterion and the overall score is one dealbreaker, not two; one criterion
|
||||
captures it.
|
||||
- **Non-triggers.** Statements that protect behavior from penalties: "do not
|
||||
penalize X", "X is acceptable", "either A or B clears the bar", "when the
|
||||
condition is unmet, this does not apply". These prevent over-penalizing.
|
||||
When a non-trigger is dropped, the atomic rubric penalizes what the holistic
|
||||
rubric permits — a criterion phrased without the exception, or missing the
|
||||
either/or fork, is a gap even though every requirement is present. Look for
|
||||
the protection in the criterion's guideline (conditional or either/or
|
||||
phrasing) or its elaboration (fulfillment shapes, does-not-fire notes).
|
||||
- **Answer-key facts.** The specific facts, citations, and mechanisms the
|
||||
holistic rubric supplies as ground truth. Each must survive either inline in
|
||||
the criterion that grades it or in `tests/grader-context.md`. A criterion
|
||||
that says "the response should identify the defect" whose defect is defined
|
||||
nowhere in the atomic package has lost its key.
|
||||
- **Conditions and qualifiers.** A penalty the holistic rubric applies
|
||||
conditionally must not become an unconditional criterion, and a scoped
|
||||
requirement must not become a blanket one. Compare qualifiers clause by
|
||||
clause.
|
||||
|
||||
## What counts as invented content (atomic → holistic)
|
||||
|
||||
Walk the criteria and check each against the holistic rubric and the context
|
||||
document:
|
||||
|
||||
- **New requirements.** A guideline requiring something the holistic rubric
|
||||
never asks for. The conversion is not the place to add scope; a genuinely
|
||||
missing requirement belongs in the holistic rubric first, so both forms stay
|
||||
in agreement.
|
||||
- **New answer-key facts.** A bolded key, citation, or mechanism stated in a
|
||||
criterion with no support in the holistic rubric or the context document.
|
||||
Whether such a fact is *true* is a different detector's job; here the
|
||||
finding is that the two forms no longer say the same thing. Tightening an
|
||||
existing fact (adding a file and line to a mechanism the holistic rubric
|
||||
already names) is not invention.
|
||||
- **Severity without basis.** A `crux` criterion with no heavy penalty against
|
||||
the overall score behind it in the holistic rubric. Crux weighting dominates
|
||||
the aggregate score, so an unsupported crux re-weights the whole rubric;
|
||||
treat it as material when it dominates scoring and as minor when the backing
|
||||
penalty is arguable (for example, a moderate overall-score penalty, which
|
||||
belongs at a normal severity tier rather than crux).
|
||||
- **New requirements smuggled into elaboration.** An elaboration is for
|
||||
fulfillment shapes and clarification. When it adds a requirement, check the
|
||||
holistic rubric for it; content with no holistic basis is a coverage finding
|
||||
here, and the guideline-vs-elaboration placement is the
|
||||
detector-rubric-form detector's lane.
|
||||
|
||||
## What is NOT a finding
|
||||
|
||||
- **Restructuring.** Tiers dissolving into criteria, strong/weak prose
|
||||
becoming fulfillment shapes in elaborations, one holistic paragraph
|
||||
collapsing into one criterion, or one holistic penalty becoming a base
|
||||
criterion plus a worse-variant criterion that fails in addition to it
|
||||
(paired escalation is a sanctioned encoding of "this variant is strictly
|
||||
worse").
|
||||
- **Dropped penalty magnitudes.** The atomic rubric carries no numeric
|
||||
penalty amounts by design. A "subtract roughly 0.35" that survives only as
|
||||
a severity tier is the conversion working.
|
||||
- **Dropped generic scoring mechanics.** Floor-at-zero notes, "penalties are
|
||||
never ceilings", and similar task-independent mechanics belong to the shared
|
||||
grading machinery, not to per-task criteria.
|
||||
- **Condensed context.** `tests/grader-context.md` may compress the holistic
|
||||
rubric's context prose. The finding is a lost *fact* that criteria rely on,
|
||||
never lost word count.
|
||||
- **Wording differences with the same scoring effect.** Judge what a grader
|
||||
would do, not whether the sentences match.
|
||||
- **A duplicated file set.** Both rubric forms sitting side by side in
|
||||
`tests/` is the intended package shape, not redundancy.
|
||||
|
||||
## How to work
|
||||
|
||||
1. Read the holistic rubric end to end and list its load-bearing clauses:
|
||||
requirements, penalties (with their targets and conditions), non-triggers,
|
||||
and answer-key facts.
|
||||
2. Read `tests/atomic-rubric.yaml` (or `tests/rubrics.yaml`) end to end,
|
||||
guideline and elaboration both, and `tests/grader-context.md` in full.
|
||||
3. Map each holistic clause to the criterion or context section that captures
|
||||
it. Record the criterion `id`. A clause may map to several criteria and
|
||||
several clauses may map to one criterion; what matters is that the scoring
|
||||
content lands somewhere.
|
||||
4. Sweep the reverse direction: for each criterion, find its holistic source.
|
||||
5. Check the crux designations against the holistic rubric's heavy penalties
|
||||
that target the overall score, in both directions, allowing for the
|
||||
two-crux cap: once two criteria carry `crux`, a further overall-score
|
||||
penalty is correctly encoded at `certain_dealbreaker`.
|
||||
6. Reduce to a verdict per the definitions above.
|
||||
|
||||
Never assert a mapping you have not traced. If you claim a clause is covered,
|
||||
name the criterion id that covers it.
|
||||
|
||||
## Anti-patterns: do not do these
|
||||
|
||||
- **Don't flag the restructuring itself.** The two forms are supposed to look
|
||||
different. Only content differences with scoring effect are findings.
|
||||
- **Don't demand one criterion per holistic sentence.** Several parallel facts
|
||||
from one derivation may live in one criterion, and one dense holistic
|
||||
paragraph may fan out into several criteria.
|
||||
- **Don't paraphrase away qualifiers.** Quote the holistic clause verbatim,
|
||||
conditions included, and quote the criterion text verbatim next to it.
|
||||
Describing a conditionally-applied penalty as unconditional is a factual
|
||||
error in the report.
|
||||
- **Don't re-litigate substance.** "This requirement is an over-ask" is the
|
||||
meaningfulness detector's lane. Here the holistic rubric is the reference,
|
||||
right or wrong.
|
||||
- **Don't treat sharpened citations as invention.** A criterion may pin an
|
||||
existing holistic fact to a file and line. Invention means a *new* fact or
|
||||
requirement, not a more precise statement of an existing one.
|
||||
- **Don't count a both-targets penalty twice.** A holistic dealbreaker may
|
||||
direct its penalty at a criterion and at the overall score together; that is
|
||||
one dealbreaker, encoded once.
|
||||
|
||||
## Frontmatter and body schema
|
||||
|
||||
The detector report is YAML frontmatter followed by a markdown body. Both
|
||||
contexts produce the same shape; only the *sink* differs (the wrapping
|
||||
`SKILL.md` tells you where to send the report).
|
||||
|
||||
**Frontmatter** — exactly these keys, exactly these enum values:
|
||||
|
||||
```yaml
|
||||
---
|
||||
detector: detector-rubric-coverage
|
||||
verdict: clear | minor-issues | material-issues | not-applicable
|
||||
confidence: HIGH | MEDIUM | LOW
|
||||
---
|
||||
```
|
||||
|
||||
**Body sections**, in this order:
|
||||
|
||||
```markdown
|
||||
# Rubric-coverage check: <slug>
|
||||
|
||||
Assessed: <resolved holistic rubric path> against <atomic rubric path> and tests/grader-context.md
|
||||
|
||||
## Coverage map
|
||||
|
||||
One table row per load-bearing holistic clause (requirement, penalty, or
|
||||
non-trigger):
|
||||
|
||||
| Holistic clause (short, verbatim key phrase) | Criterion id(s) | Status |
|
||||
| --- | --- | --- |
|
||||
| "…" | criterion-id | covered / partial / missing |
|
||||
|
||||
## Coverage gaps
|
||||
|
||||
One block per `partial` or `missing` row:
|
||||
|
||||
### <short label>
|
||||
|
||||
- **Holistic clause:** the verbatim sentence(s) and their location (section
|
||||
or heading in the holistic rubric).
|
||||
- **Closest criterion:** the criterion id that comes nearest, quoted, or a
|
||||
statement that none exists.
|
||||
- **What is lost:** 1-2 sentences on the scoring effect of the gap — which
|
||||
responses now score differently under the atomic rubric.
|
||||
- **Suggested criterion (optional):** a concrete guideline that would close
|
||||
the gap.
|
||||
|
||||
If there are no gaps, write "None found." and move on.
|
||||
|
||||
## Invented content
|
||||
|
||||
One block per criterion (or elaboration) with content the holistic rubric
|
||||
does not support: quote the criterion text verbatim, state what was searched
|
||||
for in the holistic rubric and the context document, and name the scoring
|
||||
effect. If there is none, write "None found."
|
||||
|
||||
## Context integrity
|
||||
|
||||
Whether the holistic rubric's context sections survive in
|
||||
tests/grader-context.md. Name any fact that criteria rely on that is missing
|
||||
from both the context document and the criteria. If everything survives,
|
||||
say so.
|
||||
|
||||
## Crux alignment
|
||||
|
||||
List every heavy penalty in the holistic rubric that targets the overall
|
||||
score and the criterion encoding it (`crux`, or `certain_dealbreaker` once
|
||||
two crux criteria are designated), and every crux criterion and the penalty
|
||||
backing it. Flag mismatches in either direction.
|
||||
|
||||
## Overall verdict
|
||||
|
||||
1-2 paragraphs reducing the findings to the chosen verdict. Be explicit about
|
||||
which direction (gap, invention, context loss, crux mismatch) drove the call.
|
||||
```
|
||||
|
||||
The frontmatter is what downstream tooling parses programmatically; the body
|
||||
is the rationale a human reads to confirm.
|
||||
Reference in New Issue
Block a user