lots of change - all to start my 3rd redo
This commit is contained in:
@@ -1,73 +0,0 @@
|
||||
---
|
||||
name: detector-rubric-form
|
||||
description: |
|
||||
Self-check that your atomic rubric is well-formed. A deterministic contract
|
||||
checks the artifact: the file parses against the criterion schema,
|
||||
ids are kebab-case and unique, category and
|
||||
severity use the defined vocabularies, extra_credit criteria carry no
|
||||
severity, at most 2 criteria are crux, `dimensions` names grading-standard
|
||||
criteria, and no text states a numeric penalty amount. A judgment layer
|
||||
checks the writing: each guideline is one positively phrased,
|
||||
independently judgeable requirement, criteria stand alone, factual
|
||||
criteria carry their answer key inline in bold, and elaborations clarify
|
||||
the guideline instead of adding requirements. Reads
|
||||
`tests/atomic-rubric.yaml` (or `tests/rubrics.yaml`) and
|
||||
`tests/grader-context.md`. Emits `not-applicable` when the task has no
|
||||
atomic rubric yet.
|
||||
allowed-tools: Bash, Read, Write
|
||||
---
|
||||
|
||||
# Rubric-form detector
|
||||
|
||||
This skill checks your atomic rubric as an artifact. Each criterion is scored
|
||||
on its own, and the aggregate score is computed from `category` and
|
||||
`severity`. That only works when the file obeys the schema and each criterion
|
||||
states one requirement a grader can judge independently.
|
||||
|
||||
The failure shapes to catch:
|
||||
|
||||
- **Schema violations.** The file fails to parse, ids repeat or are not
|
||||
kebab-case, a category or severity value is outside the vocabulary, an
|
||||
extra_credit criterion carries a severity, more than 2 criteria are crux,
|
||||
or `dimensions` is empty.
|
||||
- **Numeric penalty language.** A guideline, elaboration, or
|
||||
`tests/grader-context.md` sentence states a penalty amount, such as
|
||||
"subtract roughly 0.35". Penalty weight is expressed through category and
|
||||
severity. Sizing the subtraction is the grading machinery's job.
|
||||
- **Negation-phrased guidelines.** A guideline says "should not" or "must
|
||||
not" instead of stating the requirement positively. Use "The response
|
||||
should avoid X" for prohibitions.
|
||||
- **Bundled or fragmentary criteria.** One criterion packs several
|
||||
independent requirements, so a grader must improvise a partial verdict.
|
||||
Or a criterion cannot be judged without reading a sibling criterion.
|
||||
Parallel facts from one derivation may share a criterion.
|
||||
- **Missing answer keys.** A criterion grades the response for surfacing a
|
||||
specific fact, and the fact is not stated inline in bold in the guideline.
|
||||
- **Requirements hidden in elaborations.** An elaboration adds a requirement
|
||||
the guideline never states.
|
||||
- **Unfair grading shapes.** Criteria spent on trivially-satisfied
|
||||
properties, two criteria that both fire on one defect with no note saying
|
||||
which one charges, phrasing that forecloses an approach the rubric's own
|
||||
text treats as acceptable, or a requirement the task's environment cannot
|
||||
satisfy.
|
||||
|
||||
Read these before deciding:
|
||||
|
||||
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
|
||||
2. `.claude/skills/detector-rubric-form/core.md` — the deterministic contract with its pattern sweeps, the judgment checks, what is deliberately not a finding, verdict definitions, and the body schema.
|
||||
|
||||
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
|
||||
|
||||
## Acting on the verdict
|
||||
|
||||
- **`clear`** — the file passes the deterministic contract and the criteria
|
||||
read as a working rubric. Good.
|
||||
- **`minor-issues`** — the contract passes, and the findings are
|
||||
polish-level. Read the findings list and tighten the criteria. There is no
|
||||
need to rebuild the rubric.
|
||||
- **`material-issues`** — the file breaks the deterministic contract, or at
|
||||
least one criterion cannot be graded as written. Fix every finding in the
|
||||
deterministic-contract section first, then the judgment findings. Re-run
|
||||
this skill after editing.
|
||||
- **`not-applicable`** — the task has no atomic rubric yet. Write the atomic
|
||||
rubric first, then come back to this skill.
|
||||
@@ -1,301 +0,0 @@
|
||||
# Rubric-form detector — core
|
||||
|
||||
This file is the canonical, context-neutral content for the detector-rubric-form
|
||||
detector. It defines the deterministic contract an atomic rubric must satisfy,
|
||||
the judgment checks on top of it, the verdict enum, and the output schema. It
|
||||
is read in two contexts — the base repo's review pipeline and the worker
|
||||
toolkit's self-check — so nothing here should reference downstream storage
|
||||
details.
|
||||
|
||||
## What this detector is for
|
||||
|
||||
The **atomic rubric** (`tests/atomic-rubric.yaml`) expresses a task's grading
|
||||
requirements as a list of criteria. Each criterion is scored on its own, and
|
||||
the aggregate score is computed from the per-criterion verdicts using the
|
||||
criterion's `category` and `severity`. That machinery only works when the
|
||||
artifact is well-formed: the file must obey the criterion schema, and each
|
||||
criterion must state one requirement a grader can judge independently.
|
||||
|
||||
This detector checks the artifact itself, in two layers:
|
||||
|
||||
1. **A deterministic contract.** Schema and vocabulary rules that either hold
|
||||
or do not. Spelled out below; the list is the contract.
|
||||
2. **Judgment checks.** Atomicity, self-containment, phrasing, answer-key
|
||||
placement, elaboration discipline, and fair-grading properties that need a
|
||||
reader, not a validator.
|
||||
|
||||
It does **not** judge whether the criteria match the task's holistic rubric —
|
||||
the detector-rubric-coverage detector owns content equivalence — and it does
|
||||
not verify factual claims against the source repo, route failures to grading
|
||||
criteria, or weigh whether the tested failure matters. Those belong to their
|
||||
own detectors.
|
||||
|
||||
## Inputs
|
||||
|
||||
Read from `harbor-tasks/<slug>/`:
|
||||
|
||||
- `tests/atomic-rubric.yaml` — the primary input. A task packaged under an
|
||||
earlier release carries the same artifact as `tests/rubrics.yaml`; when
|
||||
`tests/atomic-rubric.yaml` is absent, assess `tests/rubrics.yaml`. Read
|
||||
every criterion, guideline and elaboration both.
|
||||
- `tests/grader-context.md` — the companion context document. The
|
||||
numeric-penalty rule below applies to it too, and the self-containment
|
||||
check needs to know what context the criteria can legitimately lean on.
|
||||
- `instruction.md` — secondary. Use it to judge whether a criterion's
|
||||
requirement is within reach of a response produced in this task's
|
||||
environment, and whether an either/or fork is warranted.
|
||||
|
||||
You do not need the workspace, the reference runs, or the holistic rubric.
|
||||
|
||||
## The deterministic contract
|
||||
|
||||
Every check in this list either passes or fails on the file as written.
|
||||
Report each failure with the offending text quoted verbatim.
|
||||
|
||||
1. **Parses as YAML.** The file loads as a YAML document with a top-level
|
||||
`task` string and a `criteria` list. A file that does not parse is a
|
||||
broken artifact; report the parse error and verdict `material-issues`.
|
||||
2. **`task` names this task.** The `task` field equals the task's slug.
|
||||
3. **Ids are kebab-case and unique.** Each `id` matches
|
||||
`^[a-z0-9]+(-[a-z0-9]+)*$` and appears once.
|
||||
4. **`category` vocabulary.** One of `primary_intent`, `extra_credit`,
|
||||
`dodged_bullet`.
|
||||
5. **`severity` vocabulary and placement.** One of `crux`,
|
||||
`certain_dealbreaker`, `possible_dealbreaker`, `unlikely_dealbreaker`.
|
||||
Required on `primary_intent` and `dodged_bullet` criteria. Forbidden on
|
||||
`extra_credit` criteria.
|
||||
6. **Crux cap.** At most 2 criteria carry `severity: crux`.
|
||||
7. **`dimensions` names at least one grading-standard criterion.** Each entry
|
||||
is one of the eight, exactly as the grading standard names them:
|
||||
`Integrity`, `Narrow Correctness`,
|
||||
`Broader Correctness / the craft of software engineering`, `Persistence`,
|
||||
`Communication`, `Verification & Thoroughness`, `Common Sense`,
|
||||
`Thought Partnership`.
|
||||
8. **`guideline` is non-empty** on every criterion.
|
||||
9. **Zero numeric penalty language.** Penalty weight is expressed through
|
||||
`category` and `severity`; sizing the subtraction is the grading
|
||||
machinery's job. No guideline, elaboration, or context-document sentence
|
||||
may state a numeric penalty amount. Run these over the atomic rubric AND
|
||||
`tests/grader-context.md`; the pattern list is the contract:
|
||||
|
||||
```bash
|
||||
TESTS=harbor-tasks/<slug>/tests
|
||||
RUBRIC="$TESTS/atomic-rubric.yaml"; [ -f "$RUBRIC" ] || RUBRIC="$TESTS/rubrics.yaml"
|
||||
|
||||
# Subtraction verbs with an amount: "subtract roughly 0.35", "deduct 5", "dock 40-45"
|
||||
grep -inE '(subtract|deduct|dock)[a-z]*[[:space:]]+((roughly|about|around|approximately|up[[:space:]]+to|at[[:space:]]+least)[[:space:]]+)?[0-9]' "$RUBRIC" "$TESTS/grader-context.md"
|
||||
|
||||
# An amount attached to a penalty noun: "a 0.35 penalty", "a 20% penalty", "0.1-0.4 deduction"
|
||||
grep -inE '[0-9]+(\.[0-9]+)?([[:space:]]*(-|to|–|—)[[:space:]]*[0-9]+(\.[0-9]+)?)?[[:space:]]*(%|percent)?[[:space:]]*(point[[:space:]]+)?(penalt|deduction)' "$RUBRIC" "$TESTS/grader-context.md"
|
||||
|
||||
# A penalty noun with an amount: "penalty of 0.35", "penalize by 20%", "deduction of 0.1"
|
||||
grep -inE '(penalt[a-z]*|penali[sz][a-z]*|deduction)[[:space:]]+(of|by)[[:space:]]+((roughly|about|around|approximately|up[[:space:]]+to|at[[:space:]]+least)[[:space:]]+)?[0-9]' "$RUBRIC" "$TESTS/grader-context.md"
|
||||
|
||||
# Score adjustments by amount: "lower the score by 0.2"
|
||||
grep -inE 'score[[:space:]]+by[[:space:]]+((roughly|about|around|approximately)[[:space:]]+)?[0-9]' "$RUBRIC" "$TESTS/grader-context.md"
|
||||
|
||||
# Point values and out-of-100 scales: "5 points", "1 pt", "out of 100"
|
||||
grep -inE '[0-9]+(\.[0-9]+)?[[:space:]]+(points?|pts)([^a-z]|$)|out[[:space:]]+of[[:space:]]+100' "$RUBRIC" "$TESTS/grader-context.md"
|
||||
```
|
||||
|
||||
Every hit is a candidate, not automatically a finding: confirm the number
|
||||
sizes a penalty or a score before reporting. Counts ("misses 3 of the 4
|
||||
call sites"), behavior thresholds ("fewer than 80% of the tests pass"),
|
||||
line numbers, dollar amounts, and version numbers never count.
|
||||
Qualitative penalty phrasing ("this is a certain dealbreaker") never
|
||||
matches and is the sanctioned form.
|
||||
11. **Positively phrased guidelines.** A guideline is one positively-phrased
|
||||
statement of the requirement: "The response should …", the conditional
|
||||
form "If the response includes X, it should …", or "The response should
|
||||
avoid …" for prohibitions. Negation words in the requirement itself —
|
||||
"should not", "must not", "may not", "does not", "never" — are the
|
||||
non-sanctioned form; "avoid" replaces them. Candidates:
|
||||
|
||||
```bash
|
||||
grep -inE '(should|must|may|shall)[[:space:]]+not[[:space:]]|do(es)?[[:space:]]+not[[:space:]]|never[[:space:]]' "$RUBRIC"
|
||||
```
|
||||
|
||||
Confirm each hit phrases the *requirement* before reporting. Negation
|
||||
inside an answer key describing the state of the code ("a constant that
|
||||
does not exist"), or inside an elaboration describing what a failing
|
||||
response looks like, is not a finding.
|
||||
|
||||
## Judgment checks
|
||||
|
||||
- **Atomicity.** Each criterion states one requirement that can be judged
|
||||
independently. Flag two shapes:
|
||||
- **Bundles of independent requirements.** A guideline a grader could
|
||||
reasonably half-pass — the response did A but not B, and A and B stand or
|
||||
fall separately — forces an improvised partial verdict. Split it.
|
||||
- **Fragments that cannot be judged alone.** A criterion whose pass/fail
|
||||
condition only makes sense while reading a sibling criterion or a
|
||||
document the grader does not have.
|
||||
Parallel facts from the same derivation MAY bundle: when several claims
|
||||
stand or fall together because they come from one piece of evidence or one
|
||||
mechanism, one criterion carrying all of them is sanctioned, and so is an
|
||||
enumerated answer key inside one criterion when the facts form one finding.
|
||||
- **Self-containment.** Each criterion is judgeable from its own text plus
|
||||
`tests/grader-context.md`. Flag a criterion whose requirement depends on
|
||||
another criterion's content ("the same standard as the criterion above",
|
||||
"see `other-criterion-id` for the definition"). A routing note in an
|
||||
elaboration that names a sibling criterion id to prevent double-charging is
|
||||
acceptable; the requirement itself must still stand alone.
|
||||
- **Answer keys inline and bold.** A factual criterion — one that grades the
|
||||
response for surfacing or stating a specific fact — carries its answer key
|
||||
inside the guideline, in bold, with citations where they exist. A key that
|
||||
lives only in `tests/grader-context.md` makes the grader hunt; a key that
|
||||
exists nowhere makes the criterion ungradeable.
|
||||
- **Elaboration discipline.** An elaboration clarifies its guideline: what
|
||||
fulfills it, what fails it, tricky-concept clarification, charge-once
|
||||
routing. Flag an elaboration that adds a requirement the guideline does not
|
||||
state — a grader reading guidelines alone would miss it, and requirements
|
||||
belong in guidelines.
|
||||
- **Weight on behavior that can meaningfully fail.** Criteria should target
|
||||
behavior a real response can get wrong in a way that matters. A rubric
|
||||
padded with trivially-satisfied properties (the response is in English, the
|
||||
response mentions the file it edited) dilutes the weight of the criteria
|
||||
that matter, because every criterion carries weight in the aggregate.
|
||||
- **No over-penalizing bundles.** One defect should not fail several criteria
|
||||
at once unless each represents a genuinely distinct miss. A base criterion
|
||||
plus a strictly-worse-variant criterion that fails in addition to it is a
|
||||
sanctioned escalation pair; two near-duplicate criteria that both fire on
|
||||
the same single defect, with no routing note saying which one charges, is
|
||||
double-counting built into the artifact.
|
||||
- **Room for defensible judgment calls.** Where the task admits more than one
|
||||
defensible approach, the criterion should accommodate it with either/or
|
||||
phrasing ("The response should either flag the discrepancy and ask, or
|
||||
proceed under a stated assumption") or a conditional. Flag a criterion
|
||||
phrased as the one true path when the rubric's own elaborations or the
|
||||
context document acknowledge an alternative as acceptable. Whether an
|
||||
uncredited alternative *is* defensible against the prompt is the
|
||||
answer-obviousness detector's lane; here the flag is phrasing that
|
||||
forecloses what the atomic package itself treats as acceptable.
|
||||
- **Within the response's reach.** Criteria must be satisfiable by a response
|
||||
produced in the task's environment. Flag a criterion that requires actions
|
||||
the environment does not support (reaching the network, running a service
|
||||
the sandbox does not have) or that grades infrastructure failures — a tool
|
||||
crash, a harness timeout — as if they were response behavior.
|
||||
|
||||
## Verdict definitions
|
||||
|
||||
- **`not-applicable`** — there is no atomic rubric to assess: neither
|
||||
`tests/atomic-rubric.yaml` nor `tests/rubrics.yaml` exists. Emit this and
|
||||
stop. A file that exists but does not parse is NOT `not-applicable` — that
|
||||
is a broken authored artifact, and it is `material-issues`.
|
||||
|
||||
- **`clear`** — the deterministic contract passes in full, and the criteria
|
||||
read as a working rubric: atomic, self-contained, positively phrased,
|
||||
factual keys inline and bold, elaborations clarifying rather than adding.
|
||||
|
||||
- **`minor-issues`** — the deterministic contract passes, and the judgment
|
||||
findings are polish-level: an awkward-but-judgeable bundle, an answer key
|
||||
parked in the context document instead of inline, mild padding, a single
|
||||
negation-phrased guideline whose pass/fail direction is still plain.
|
||||
|
||||
- **`material-issues`** — at least one of:
|
||||
- **A deterministic-contract violation.** The file fails schema,
|
||||
vocabulary, cap, or numeric-penalty rules as written. Validation gates on
|
||||
these, so the artifact is broken until fixed.
|
||||
- **A load-bearing judgment failure.** A bundle a grader must half-pass on
|
||||
realistic responses; a criterion that cannot be judged alone; a factual
|
||||
criterion with no answer key anywhere; a requirement that exists only in
|
||||
an elaboration; a criterion outside the response's reach; double-counting
|
||||
built into near-duplicate criteria; negation phrasing that leaves the
|
||||
pass/fail direction genuinely unclear.
|
||||
|
||||
## Confidence
|
||||
|
||||
- **HIGH** — the deterministic results are unambiguous and the judgment calls
|
||||
are plain (most runs of this detector, by construction).
|
||||
- **MEDIUM** — at least one finding is genuinely a judgment call: a bundle
|
||||
that could be read as one derivation, a key whose inline-ness is arguable.
|
||||
- **LOW** — limited information (an unfamiliar domain where "can this be
|
||||
judged alone" is hard to tell, or a very large rubric only sampled).
|
||||
|
||||
## Anti-patterns: do not do these
|
||||
|
||||
- **Don't report raw grep hits as findings.** The patterns generate
|
||||
candidates; the confirmed penalty-sizing or requirement-negation reading is
|
||||
the finding. Quote the confirmed text verbatim, with the criterion id.
|
||||
- **Don't flag sanctioned bundles.** Parallel same-derivation facts in one
|
||||
criterion, enumerated keys forming one finding, and base + worse-variant
|
||||
escalation pairs are the format working.
|
||||
- **Don't flag charge-once routing notes as cross-references.** Naming a
|
||||
sibling criterion id to prevent double-charging is discipline, not
|
||||
dependence.
|
||||
- **Don't re-litigate content.** Whether a requirement matches the holistic
|
||||
rubric is coverage's lane; whether a stated fact is true is fact-check's;
|
||||
whether the targeted failure matters is meaningfulness's. Judge the
|
||||
artifact, not the task.
|
||||
- **Don't demand splitting past judgeability.** Maximum viable atomicity
|
||||
means the smallest *meaningful* unit. A criterion is small enough when a
|
||||
grader can pass or fail it in one decision; pushing further fragments it.
|
||||
- **Don't treat `dimensions` routing as this detector's call.** The
|
||||
deterministic check is vocabulary only. Whether a failure is routed to the
|
||||
right grading criterion belongs to the dimension-misapplication detector.
|
||||
|
||||
## Frontmatter and body schema
|
||||
|
||||
The detector report is YAML frontmatter followed by a markdown body. Both
|
||||
contexts produce the same shape; only the *sink* differs (the wrapping
|
||||
`SKILL.md` tells you where to send the report).
|
||||
|
||||
**Frontmatter** — exactly these keys, exactly these enum values:
|
||||
|
||||
```yaml
|
||||
---
|
||||
detector: detector-rubric-form
|
||||
verdict: clear | minor-issues | material-issues | not-applicable
|
||||
confidence: HIGH | MEDIUM | LOW
|
||||
---
|
||||
```
|
||||
|
||||
**Body sections**, in this order:
|
||||
|
||||
```markdown
|
||||
# Rubric-form check: <slug>
|
||||
|
||||
Assessed: <atomic rubric path>
|
||||
|
||||
## Deterministic contract
|
||||
|
||||
One line per check (1-11), pass or FAIL. For each FAIL: the offending text
|
||||
quoted verbatim, the criterion id (or file location), and the rule it
|
||||
breaks. For the pattern checks, state that the sweeps ran and what they
|
||||
matched; a candidate hit cleared as a non-finding gets one line saying why.
|
||||
|
||||
## Atomicity and self-containment
|
||||
|
||||
One block per finding:
|
||||
|
||||
### <short label>
|
||||
|
||||
- **Criterion:** the criterion id.
|
||||
- **Where:** the guideline or elaboration text, quoted verbatim.
|
||||
- **Why:** 1-2 sentences — which independent requirements are bundled, or
|
||||
what the criterion depends on that it does not contain.
|
||||
- **Suggested split or rewrite:** concrete replacement criteria or phrasing.
|
||||
|
||||
If there are none, write "None found."
|
||||
|
||||
## Phrasing and answer keys
|
||||
|
||||
Findings on positive phrasing, inline/bold answer keys, and elaboration
|
||||
discipline, same block shape as above. If there are none, write
|
||||
"None found."
|
||||
|
||||
## Fair-grading findings
|
||||
|
||||
Findings on trivially-satisfied criteria, over-penalizing bundles, missing
|
||||
either/or accommodation, and requirements outside the response's reach,
|
||||
same block shape. If there are none, write "None found."
|
||||
|
||||
## Overall verdict
|
||||
|
||||
1-2 paragraphs reducing the findings to the chosen verdict. Be explicit
|
||||
about whether the deterministic contract or the judgment layer drove the
|
||||
call.
|
||||
```
|
||||
|
||||
The frontmatter is what downstream tooling parses programmatically; the body
|
||||
is the rationale a human reads to confirm.
|
||||
Reference in New Issue
Block a user