6.1 KiB
name, description, allowed-tools
| name | description | allowed-tools |
|---|---|---|
| regrade-reference-run | Re-run a task's verifier (the grader) against a reference run you already captured, skipping the agent. Use when iterating on tests/holistic-rubric.md, measuring grader variance, or sanity-checking a verifier change — anywhere you'd otherwise re-spend minutes of agent runtime just to get a fresh grade against the same agent behavior. | Bash, Read, Glob, Grep |
Re-grade a reference run without re-running the agent
When to use this
You have a harbor-tasks/<slug>/reference-runs/<run-id>/ directory captured by an earlier real trial — its agent-output/, agent/trajectory.json, grade.md, reward.txt, and reward-correctness.txt are all on disk. You want to grade that captured run again. Most common reason: you edited tests/holistic-rubric.md and want to see how the new wording changes the scores against the same agent behavior, without paying for a fresh agent run.
Regrading re-derives the full grade — every criterion's reasoning in grade.md and the reward — so it is the right tool for iterating on any part of your rubric.
Other good fits:
- Grader variance. Run the same reference 10× in parallel, look at the spread in
reward.txt. Useful when you suspect the grader is non-deterministic on a borderline call. - Sanity-check a verifier change. If you patched
tests/test.shitself, regrade an existing reference run to confirm the patch produces the same grade against the same agent behavior.
How it works
scripts/harbor-regrade invokes the standard harbor run plumbing but plugs in a replay adapter (scripts/replay_agent.py) instead of an agent. The adapter:
- Uploads your captured
agent-output/into the trial container's/workspace— overlays the agent's surviving file edits on top of the base workspace built by the task'sDockerfile. - If
agent-output/_HARBOR_DELETIONS.txtexists (records of any files the agent deleted),rms each listed path so the workspace state ends up identical to where the original agent left it. - Uploads the captured
agent/trajectory.jsonso the grader reads the same transcript it would have on the original run.
Then the verifier (tests/test.sh) runs exactly as it does for any other trial. Same git ls-files/git diff workspace capture, same deterministic checks, same grader, same reward.txt/reward-correctness.txt/grade.md output. The only difference is that the agent phase is now seconds of file overlay instead of minutes of agent work.
How to invoke
scripts/harbor-regrade <task-dir> <reference-run-dir> [-k N] [extra harbor args]
<task-dir>:harbor-tasks/<slug>— same dir you'd pass toscripts/harbor-run.<reference-run-dir>:harbor-tasks/<slug>/reference-runs/<run-id>— must containagent-output/.-k N: N independent regrades against the same captured state. Use for variance measurement.
What grades the run
The grader scores the eight criteria of the Grading Standard against the task's holistic rubric (tests/holistic-rubric.md; a task started on an earlier toolkit carries the same document as tests/grader-guidance-consolidated.md). The reward is the mean of the non-N/A criteria minus any overall penalties, floored at 0. reward-correctness.txt always reads N/A by design — correctness lives inside the criteria rather than as a separate score — so only the reward and the criterion reasoning move when you edit the rubric.
The regrade uses the grader assets already in the task's tests/ directory, so a run regrades under the same standard it was originally graded with.
Output lands in harbor-jobs/<timestamp>/<trial-id>/ like any other harbor trial — verifier/reward.txt, verifier/reward-correctness.txt, verifier/reward.json, verifier/grade.md, verifier/test-stdout.txt, trial.log. To see how the new grade diverges from the original:
diff harbor-tasks/<slug>/reference-runs/<run-id>/grade.md \
harbor-jobs/<timestamp>/<trial-id>/verifier/grade.md
For the number alone, the tail of verifier/test-stdout.txt prints it, or compare directly:
echo "before: $(cat harbor-tasks/<slug>/reference-runs/<run-id>/reward.txt)"
echo "after: $(cat harbor-jobs/<timestamp>/<trial-id>/verifier/reward.txt)"
Typical iteration loop
- Run a few real trials to capture reference runs:
scripts/harbor-run harbor-tasks/<slug> -k 4, thennpx tsx scripts/copy-reference-run.ts harbor-jobs/<job>/<trial>for each one you want to keep. - Read the captured
grade.mdfiles — every criterion section, not just the headline score — and find places where the grader's judgment doesn't match what you'd say as the task author. - Edit
tests/holistic-rubric.mdto clarify the points the grader got wrong. scripts/harbor-regrade harbor-tasks/<slug> harbor-tasks/<slug>/reference-runs/<run-id>for each captured run you care about.- Diff the new
grade.mdfiles vs the originals. Repeat until the grader is reasoning correctly on each captured behavior.
This is much faster (and cheaper) than re-running scripts/harbor-run after every grader edit, because each agent run takes minutes and produces a different trajectory anyway — so re-running confounds "is the grader better?" with "is the agent behavior different?".
Caveat: old reference runs
If your reference run was captured before this toolkit version, its agent-output/ won't include _HARBOR_DELETIONS.txt. The replay still works, but any file deletions the agent made in that run can't be reproduced (the original capture only preserved files the agent created or modified, not the ones it removed). For tasks where the agent doesn't delete anything (most behavioral-rating tasks where the agent just writes answer.md), this doesn't matter at all. For tasks where the agent edits code and may have deleted files, you may want to re-capture a fresh reference run after the next time you run scripts/harbor-run. The same applies to a run captured before this version in which the agent renamed a file with git mv: the rename's source path is missing from _HARBOR_DELETIONS.txt, so the replay keeps both copies.