Files
project-work/worker-toolkit-potion-polyglot/.claude/skills/regrade-reference-run/SKILL.md
Eric Bell e55ccea018 Loaded up for the 3rd redo
Still on potion-voice
2026-09-26 14:57:10 -04:00

7.1 KiB
Raw Blame History

name, description, allowed-tools
name description allowed-tools
regrade-reference-run Re-run a task's verifier (the grader) against a reference run you already captured, skipping the agent. Use when iterating on tests/holistic-rubric.md, measuring grader variance, or sanity-checking a verifier change — anywhere you'd otherwise re-spend minutes of agent runtime just to get a fresh grade against the same agent behavior. Bash, Read, Glob, Grep

Re-grade a reference run without re-running the agent

When to use this

You have a harbor-tasks/<slug>/reference-runs/<run-id>/ directory captured by an earlier real trial — its agent-output/, agent/trajectory.json, grade.md, reward.txt, and reward-correctness.txt are all on disk. You want to grade that captured run again. Most common reason: you edited tests/holistic-rubric.md and want to see how the new wording changes the scores against the same agent behavior, without paying for a fresh agent run.

Regrading re-derives the full grade — every criterion's reasoning in grade.md and the reward — so it is the right tool for iterating on any part of your rubric.

Other good fits:

  • Grader variance. Run the same reference 10× in parallel, look at the spread in reward.txt. Useful when you suspect the grader is non-deterministic on a borderline call.
  • Sanity-check a verifier change. If you patched tests/test.sh itself, regrade an existing reference run to confirm the patch produces the same grade against the same agent behavior.

How it works

scripts/harbor-regrade invokes the standard harbor run plumbing but plugs in a replay adapter (scripts/replay_agent.py) instead of an agent. The adapter:

  1. Uploads your captured agent-output/ into the trial container's /workspace — overlays the agent's surviving file edits on top of the base workspace built by the task's Dockerfile.
  2. If agent-output/_HARBOR_DELETIONS.txt exists (records of any files the agent deleted), rms each listed path so the workspace state ends up identical to where the original agent left it.
  3. Uploads the captured agent/trajectory.json so the grader reads the same transcript it would have on the original run.

Then the verifier (tests/test.sh) runs exactly as it does for any other trial. Same git ls-files/git diff workspace capture, same deterministic checks, same grader, same reward.txt/reward-correctness.txt/grade.md output. The only difference is that the agent phase is now seconds of file overlay instead of minutes of agent work.

How to invoke

scripts/harbor-regrade <task-dir> <reference-run-dir>... [-k N] [extra harbor args]
scripts/harbor-regrade <task-dir> --all [--jobs N] [extra harbor args]
  • <task-dir>: harbor-tasks/<slug> — same dir you'd pass to scripts/harbor-run.
  • <reference-run-dir>: harbor-tasks/<slug>/reference-runs/<run-id>. Name several to re-grade them all.
  • --all: re-grade every run under harbor-tasks/<slug>/reference-runs/. This is what you usually want after editing your rubric.
  • --jobs N: how many to re-grade at once, default 2. Each one is a container, so raise it only as far as your machine comfortably allows.
  • -k N: N independent regrades of a single run, for variance measurement.

-k and --all are easy to mix up. -k 4 grades one captured run four times; --all grades each of your captured runs once. If you want fresh grades for all four of your reference runs, that's --all, not -k 4.

What grades the run

The grader scores the eight criteria of the Grading Standard against the task's holistic rubric (tests/holistic-rubric.md; a task started on an earlier toolkit carries the same document as tests/grader-guidance-consolidated.md). The reward is the mean of the non-N/A criteria minus any overall penalties, floored at 0. reward-correctness.txt always reads N/A by design — correctness lives inside the criteria rather than as a separate score — so only the reward and the criterion reasoning move when you edit the rubric.

The regrade uses the grader assets already in the task's tests/ directory, so a run regrades under the same standard it was originally graded with.

Output lands in harbor-jobs/<job>/<trial-id>/ like any other harbor trial — <job> is a timestamp for a single regrade, and regrade-<n>-<run-id> for each run under --all, so you can tell at a glance which reference run a result came from — verifier/reward.txt, verifier/reward-correctness.txt, verifier/reward.json, verifier/grade.md, verifier/test-stdout.txt, trial.log. To see how the new grade diverges from the original:

diff harbor-tasks/<slug>/reference-runs/<run-id>/grade.md \
     harbor-jobs/<job>/<trial-id>/verifier/grade.md

For the number alone, the tail of verifier/test-stdout.txt prints it, or compare directly:

echo "before: $(cat harbor-tasks/<slug>/reference-runs/<run-id>/reward.txt)"
echo "after:  $(cat harbor-jobs/<job>/<trial-id>/verifier/reward.txt)"

Typical iteration loop

  1. Run a few real trials to capture reference runs: scripts/harbor-run harbor-tasks/<slug> -k 4, then npx tsx scripts/copy-reference-run.ts harbor-jobs/<job>/<trial> for each one you want to keep.
  2. Read the captured grade.md files — every criterion section, not just the headline score — and find places where the grader's judgment doesn't match what you'd say as the task author.
  3. Edit tests/holistic-rubric.md to clarify the points the grader got wrong.
  4. scripts/harbor-regrade harbor-tasks/<slug> --all to re-grade every captured run in one go.
  5. Diff the new grade.md files vs the originals. Repeat until the grader is reasoning correctly on each captured behavior.

Leave the regrade output where it lands. It is there for you to read and compare, not to copy back over your reference runs — a regrade is a new trial with a new id, so copy-reference-run.ts would add a run rather than update one, leaving you with twice as many and no way to tell which grades came from which version of your rubric. Your reference runs should stay as the real trials you captured.

This is much faster (and cheaper) than re-running scripts/harbor-run after every grader edit, because each agent run takes minutes and produces a different trajectory anyway — so re-running confounds "is the grader better?" with "is the agent behavior different?".

Caveat: old reference runs

If your reference run was captured before this toolkit version, its agent-output/ won't include _HARBOR_DELETIONS.txt. The replay still works, but any file deletions the agent made in that run can't be reproduced (the original capture only preserved files the agent created or modified, not the ones it removed). For tasks where the agent doesn't delete anything (most behavioral-rating tasks where the agent just writes answer.md), this doesn't matter at all. For tasks where the agent edits code and may have deleted files, you may want to re-capture a fresh reference run after the next time you run scripts/harbor-run. The same applies to a run captured before this version in which the agent renamed a file with git mv: the rename's source path is missing from _HARBOR_DELETIONS.txt, so the replay keeps both copies.