Files
project-work/worker-toolkit-stocks-in-the-future/.claude/skills/regrade-reference-run/SKILL.md
Eric Bell 2854619bc9 chore: init commit
in worker.../repo/GITFOLDER.zip is the .git folder.
2026-08-11 14:44:09 -04:00

7.1 KiB
Raw Blame History

name, description, allowed-tools
name description allowed-tools
regrade-reference-run Re-run a task's verifier (the grader) against a reference run you already captured, skipping the agent. Use when iterating on tests/grader-guidance.md, measuring grader variance, or sanity-checking a verifier change — anywhere you'd otherwise re-spend minutes of agent runtime just to get a fresh grade against the same agent behavior. Bash, Read, Glob, Grep

Re-grade a reference run without re-running the agent

When to use this

You have a harbor-tasks/<slug>/reference-runs/<run-id>/ directory captured by an earlier real trial — its agent-output/, agent/trajectory.json, grade.md, reward.txt, and reward-correctness.txt are all on disk. You want to grade that captured run again. Most common reason: you edited tests/grader-guidance.md and want to see how the new wording changes the scores against the same agent behavior, without paying for a fresh agent run.

Regrading re-derives both scores — behavioral (reward.txt) and correctness (reward-correctness.txt) — so it's the right tool for iterating on your correctness guidance too, not just the behavioral half.

Other good fits:

  • Grader variance. Run the same reference 10× in parallel, look at the spread in reward.txt (and in reward-correctness.txt — the two axes don't necessarily have the same variance). Useful when you suspect the grader is non-deterministic on a borderline call.
  • Sanity-check a verifier change. If you patched tests/test.sh itself, regrade an existing reference run to confirm the patch produces the same grade against the same agent behavior.

How it works

scripts/harbor-regrade invokes the standard harbor run plumbing but plugs in a replay adapter (scripts/replay_agent.py) instead of an agent. The adapter:

  1. Uploads your captured agent-output/ into the trial container's /workspace — overlays the agent's surviving file edits on top of the base workspace built by the task's Dockerfile.
  2. If agent-output/_HARBOR_DELETIONS.txt exists (records of any files the agent deleted), rms each listed path so the workspace state ends up identical to where the original agent left it.
  3. Uploads the captured agent/trajectory.json so the grader reads the same transcript it would have on the original run.

Then the verifier (tests/test.sh) runs exactly as it does for any other trial. Same git ls-files/git diff workspace capture, same deterministic checks, same grader, same reward.txt/reward-correctness.txt/grade.md output. The only difference is that the agent phase is now seconds of file overlay instead of minutes of agent work.

How to invoke

scripts/harbor-regrade <task-dir> <reference-run-dir> [-k N] [extra harbor args]
  • <task-dir>: harbor-tasks/<slug> — same dir you'd pass to scripts/harbor-run.
  • <reference-run-dir>: harbor-tasks/<slug>/reference-runs/<run-id> — must contain agent-output/.
  • -k N: N independent regrades against the same captured state. Use for variance measurement.

Which standard grades the run

Regrades default to the Consolidated Grading Standard: the grader scores the eight criteria against tests/grader-guidance-consolidated.md, and the reward is the mean of the non-N/A criteria minus any overall penalties, floored at 0. Under this standard reward-correctness.txt always reads N/A — correctness lives inside the criteria rather than as a separate score.

To grade the way earlier releases did — seven behavioral dimensions plus a separate correctness score, against tests/grader-guidance.md — set HARBOR_GRADING_STANDARD=legacy:

HARBOR_GRADING_STANDARD=legacy scripts/harbor-regrade harbor-tasks/<slug> harbor-tasks/<slug>/reference-runs/<run-id>

If your task's tests/ directory predates the consolidated assets, the regrade says so in its output and grades under the legacy standard. To grade consolidated, copy the current shared assets in first:

cp task-shared/test.sh task-shared/grader-system-prompt*.md task-shared/render-grade*.py harbor-tasks/<slug>/tests/

Output lands in harbor-jobs/<timestamp>/<trial-id>/ like any other harbor trial — verifier/reward.txt, verifier/reward-correctness.txt, verifier/reward.json, verifier/grade.md, verifier/test-stdout.txt, trial.log. To see how the new grade diverges from the original:

diff harbor-tasks/<slug>/reference-runs/<run-id>/grade.md \
     harbor-jobs/<timestamp>/<trial-id>/verifier/grade.md

For the numbers alone, compare both axes side by side — the tail of verifier/test-stdout.txt prints them as behavioral reward: … correctness: …:

echo "before: $(cat harbor-tasks/<slug>/reference-runs/<run-id>/reward.txt) / $(cat harbor-tasks/<slug>/reference-runs/<run-id>/reward-correctness.txt)"
echo "after:  $(cat harbor-jobs/<timestamp>/<trial-id>/verifier/reward.txt) / $(cat harbor-jobs/<timestamp>/<trial-id>/verifier/reward-correctness.txt)"

A correctness score that moves when you only edited behavioral guidance (or vice versa) is worth a look — under the legacy standard the two axes are meant to be independent, and a rubric edit that drags both usually means the guidance is charging one fault to both. (Under the consolidated standard the correctness slot reads N/A by design, so only the reward moves.)

Typical iteration loop

  1. Run a few real trials to capture reference runs: scripts/harbor-run harbor-tasks/<slug> -k 4, then npx tsx scripts/copy-reference-run.ts harbor-jobs/<job>/<trial> for each one you want to keep.
  2. Read the captured grade.md files — both the behavioral paragraphs and the ## Correctness section — and find places where the grader's judgment doesn't match what you'd say as the task author.
  3. Edit the guidance file the run grades against — tests/grader-guidance-consolidated.md by default, tests/grader-guidance.md under HARBOR_GRADING_STANDARD=legacy — to clarify the points the grader got wrong.
  4. scripts/harbor-regrade harbor-tasks/<slug> harbor-tasks/<slug>/reference-runs/<run-id> for each captured run you care about.
  5. Diff the new grade.md files vs the originals. Repeat until the grader is reasoning correctly on each captured behavior.

This is much faster (and cheaper) than re-running scripts/harbor-run after every grader edit, because each agent run takes minutes and produces a different trajectory anyway — so re-running confounds "is the grader better?" with "is the agent behavior different?".

Caveat: old reference runs

If your reference run was captured before this toolkit version, its agent-output/ won't include _HARBOR_DELETIONS.txt. The replay still works, but any file deletions the agent made in that run can't be reproduced (the original capture only preserved files the agent created or modified, not the ones it removed). For tasks where the agent doesn't delete anything (most behavioral-rating tasks where the agent just writes answer.md), this doesn't matter at all. For tasks where the agent edits code and may have deleted files, you may want to re-capture a fresh reference run after the next time you run scripts/harbor-run.