6.5 KiB
name, description, allowed-tools
| name | description | allowed-tools |
|---|---|---|
| detector-fact-check-rubric-claims | Self-check every load-bearing factual claim in your grader guidance against the source repo at the commit declared in `task.toml`. Catches stale citations, dead-code-as-load-bearing assertions, schema-constraint claims that don't hold, behaviors mis-attributed to a file or line, and rubric self-contradictions. Each claim is checked on two axes: is it TRUE against the workspace, and — for facts your rubric grades the response for knowing or finding — is it REACHABLE from what the test agent is given (the prompt, the snapshot session, and the workspace)? A true fact the agent has no way to learn is a fairness defect, not a knowledge test. The most common worker mistakes: citing files or lines from memory and letting the rubric drift from what the code actually shows, and gating the score on privileged context (provider behavior, policy thresholds) that lives only in your head. | Bash, Read, Write |
Fact-check rubric claims
This skill checks every factual claim in your grader guidance (the file
bash scripts/guidance-target.sh <slug> resolves)
that the rubric's score depends on, against the actual source repo at the
commit your task.toml declares — and, for facts the rubric grades the
response for knowing or finding, whether the test agent could actually
reach them from the package it is given.
Read these before starting:
.claude/skills/_detector-worker-shell.md— where to write the report and how to handle re-runs. Detectors with structured payloads (this one'sclaimsarray) embed them in the same frontmatter block asdetector/verdict/confidence..claude/skills/detector-fact-check-rubric-claims/core.md— what counts as a load-bearing factual claim, the per-claim verdict enums, the top-level reduction, the frontmatter/body schema.
How to do it
Work through the rubric one claim at a time:
- Read the rubric. Open the guidance file
bash scripts/guidance-target.sh <slug>resolves. Identify every load-bearing factual claim (file paths, line ranges, function names, schema constraints, runtime behaviors that the rubric says are load-bearing for some scoring criterion). Assign each claim an id (c01,c02, …). Also mark which claims gate scoring on knowledge — facts the response is graded for knowing or finding, as opposed to background that only justifies the rubric to the grader (seecore.md, "The second axis"). - Make sure the patched workspace exists. Your rubric describes what the test agent sees, and the test agent sees
git archive <commit>plusenvironment/workspace.patchapplied — the workspace atharbor-tasks/<slug>/environment/workspace/. The directory is gitignored; if it's missing, runscripts/build-workspace.sh <slug>to rebuild it. Reading the baregit show <commit>:<path>instead would miss any files your snapshot session added/modified/deleted, producing false fails on every patched file. - For each claim, verify against the patched workspace. Open
harbor-tasks/<slug>/environment/workspace/<path>and compare what the rubric claims against what's actually there. For symbol-existence / call-site / dead-code checks, grep the workspace tree (rg '<symbol>' harbor-tasks/<slug>/environment/workspace/). Markpass/unclear/partial/failpercore.md's verdict definitions. - For each scoring-gate claim, run the reachability check. Where your rubric grades the response for knowing or finding a fact (external-provider behavior, business context, a policy threshold, a canonical root cause), trace where in the package the test agent could learn it — the prompt, the snapshot session, or the patched workspace — per
core.md's "The second axis" ladder. Hard-to-find is reachable; nowhere-in-the-package means the claim'snoteleads with anunreachable:marker plus the searches you ran. A claim can be true and still unreachable — that's exactly the defect this step catches. - Compose the report. Embed all per-claim records in the YAML frontmatter alongside
detector/verdict/confidence. The body is a short provenance summary (workspace path, commit, count of claims checked); the substance is in the inline claims array. If any claim is unreachable, name those claims in the body. - Write per
_detector-worker-shell.mdtoharbor-tasks/<slug>/detectors/detector-fact-check-rubric-claims.md.
The top-level verdict reduces from the per-claim verdicts using the rules in core.md — fail if any load-bearing claim is fail; not-applicable if there are no claims or every claim is unclear; partial if any load-bearing claim is partial / unclear, or any claim at all is fail, or any claim's note leads with unreachable:; else pass (non-load-bearing partial/unclear drift doesn't change the color — the per-claim list still shows it).
Acting on the verdict
pass— every load-bearing claim survived verification and every scoring-gate fact is reachable (any remaining drift is non-load-bearing and listed per-claim). Good. Move on.partial— a load-bearing claim has real drift or couldn't be verified, OR a non-load-bearing claim is outright false, OR a fact your rubric grades the response for knowing isn't reachable from the package. Look at the per-claim list in the report: fix any incorrect citations, even non-load-bearing ones, since they make the rubric harder to trust. For anunreachable:claim the fix is one of two moves: put the fact in the materials (state it in the prompt, plant a reachable signal in the repo), or stop gating the score on it (grade the overclaim — the agent asserting what the evidence can't support — instead of the hidden answer).fail— at least one load-bearing claim doesn't survive verification at the declared commit. The fix is to either (a) rewrite the rubric so the load-bearing claim matches the source, or (b) changetask.toml'scommitto one where the claim holds. Re-run this skill after.not-applicable— the rubric is empty / template, ortask.tomlis missing repo/commit. Write the rubric first (and confirm the commit), then come back.
Single-session approach
This worker version does the whole thing in one Claude session (read rubric → extract claims → verify each → write the markdown). Take time on each claim — read the source, quote the relevant lines, write a specific note. Skimming claims wholesale is the failure mode this skill exists to prevent in your own rubric.