Files
project-work/worker-toolkit-stocks-in-the-future/.claude/skills/regrade-reference-run/SKILL.md
Eric Bell 2854619bc9 chore: init commit
in worker.../repo/GITFOLDER.zip is the .git folder.
2026-08-11 14:44:09 -04:00

84 lines
7.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
name: regrade-reference-run
description: Re-run a task's verifier (the grader) against a reference run you already captured, skipping the agent. Use when iterating on tests/grader-guidance.md, measuring grader variance, or sanity-checking a verifier change — anywhere you'd otherwise re-spend minutes of agent runtime just to get a fresh grade against the same agent behavior.
allowed-tools: Bash, Read, Glob, Grep
---
# Re-grade a reference run without re-running the agent
## When to use this
You have a `harbor-tasks/<slug>/reference-runs/<run-id>/` directory captured by an earlier real trial — its `agent-output/`, `agent/trajectory.json`, `grade.md`, `reward.txt`, and `reward-correctness.txt` are all on disk. You want to grade that captured run again. Most common reason: you edited `tests/grader-guidance.md` and want to see how the new wording changes the scores against the same agent behavior, without paying for a fresh agent run.
Regrading re-derives **both** scores — behavioral (`reward.txt`) and correctness (`reward-correctness.txt`) — so it's the right tool for iterating on your correctness guidance too, not just the behavioral half.
Other good fits:
- **Grader variance.** Run the same reference 10× in parallel, look at the spread in `reward.txt` (and in `reward-correctness.txt` — the two axes don't necessarily have the same variance). Useful when you suspect the grader is non-deterministic on a borderline call.
- **Sanity-check a verifier change.** If you patched `tests/test.sh` itself, regrade an existing reference run to confirm the patch produces the same grade against the same agent behavior.
## How it works
`scripts/harbor-regrade` invokes the standard `harbor run` plumbing but plugs in a replay adapter (`scripts/replay_agent.py`) instead of an agent. The adapter:
1. Uploads your captured `agent-output/` into the trial container's `/workspace` — overlays the agent's surviving file edits on top of the base workspace built by the task's `Dockerfile`.
2. If `agent-output/_HARBOR_DELETIONS.txt` exists (records of any files the agent deleted), `rm`s each listed path so the workspace state ends up identical to where the original agent left it.
3. Uploads the captured `agent/trajectory.json` so the grader reads the same transcript it would have on the original run.
Then the verifier (`tests/test.sh`) runs exactly as it does for any other trial. Same `git ls-files`/`git diff` workspace capture, same deterministic checks, same grader, same `reward.txt`/`reward-correctness.txt`/`grade.md` output. The only difference is that the agent phase is now seconds of file overlay instead of minutes of agent work.
## How to invoke
```sh
scripts/harbor-regrade <task-dir> <reference-run-dir> [-k N] [extra harbor args]
```
- `<task-dir>`: `harbor-tasks/<slug>` — same dir you'd pass to `scripts/harbor-run`.
- `<reference-run-dir>`: `harbor-tasks/<slug>/reference-runs/<run-id>` — must contain `agent-output/`.
- `-k N`: N independent regrades against the same captured state. Use for variance measurement.
## Which standard grades the run
Regrades default to the **Consolidated Grading Standard**: the grader scores the eight criteria against `tests/grader-guidance-consolidated.md`, and the reward is the mean of the non-N/A criteria minus any overall penalties, floored at 0. Under this standard `reward-correctness.txt` always reads `N/A` — correctness lives inside the criteria rather than as a separate score.
To grade the way earlier releases did — seven behavioral dimensions plus a separate correctness score, against `tests/grader-guidance.md` — set `HARBOR_GRADING_STANDARD=legacy`:
```sh
HARBOR_GRADING_STANDARD=legacy scripts/harbor-regrade harbor-tasks/<slug> harbor-tasks/<slug>/reference-runs/<run-id>
```
If your task's `tests/` directory predates the consolidated assets, the regrade says so in its output and grades under the legacy standard. To grade consolidated, copy the current shared assets in first:
```sh
cp task-shared/test.sh task-shared/grader-system-prompt*.md task-shared/render-grade*.py harbor-tasks/<slug>/tests/
```
Output lands in `harbor-jobs/<timestamp>/<trial-id>/` like any other harbor trial — `verifier/reward.txt`, `verifier/reward-correctness.txt`, `verifier/reward.json`, `verifier/grade.md`, `verifier/test-stdout.txt`, `trial.log`. To see how the new grade diverges from the original:
```sh
diff harbor-tasks/<slug>/reference-runs/<run-id>/grade.md \
harbor-jobs/<timestamp>/<trial-id>/verifier/grade.md
```
For the numbers alone, compare both axes side by side — the tail of `verifier/test-stdout.txt` prints them as `behavioral reward: … correctness: …`:
```sh
echo "before: $(cat harbor-tasks/<slug>/reference-runs/<run-id>/reward.txt) / $(cat harbor-tasks/<slug>/reference-runs/<run-id>/reward-correctness.txt)"
echo "after: $(cat harbor-jobs/<timestamp>/<trial-id>/verifier/reward.txt) / $(cat harbor-jobs/<timestamp>/<trial-id>/verifier/reward-correctness.txt)"
```
A correctness score that moves when you only edited behavioral guidance (or vice versa) is worth a look — under the legacy standard the two axes are meant to be independent, and a rubric edit that drags both usually means the guidance is charging one fault to both. (Under the consolidated standard the correctness slot reads `N/A` by design, so only the reward moves.)
## Typical iteration loop
1. Run a few real trials to capture reference runs: `scripts/harbor-run harbor-tasks/<slug> -k 4`, then `npx tsx scripts/copy-reference-run.ts harbor-jobs/<job>/<trial>` for each one you want to keep.
2. Read the captured `grade.md` files — both the behavioral paragraphs and the `## Correctness` section — and find places where the grader's judgment doesn't match what you'd say as the task author.
3. Edit the guidance file the run grades against — `tests/grader-guidance-consolidated.md` by default, `tests/grader-guidance.md` under `HARBOR_GRADING_STANDARD=legacy` — to clarify the points the grader got wrong.
4. **`scripts/harbor-regrade harbor-tasks/<slug> harbor-tasks/<slug>/reference-runs/<run-id>`** for each captured run you care about.
5. Diff the new `grade.md` files vs the originals. Repeat until the grader is reasoning correctly on each captured behavior.
This is much faster (and cheaper) than re-running `scripts/harbor-run` after every grader edit, because each agent run takes minutes and produces a *different* trajectory anyway — so re-running confounds "is the grader better?" with "is the agent behavior different?".
## Caveat: old reference runs
If your reference run was captured before this toolkit version, its `agent-output/` won't include `_HARBOR_DELETIONS.txt`. The replay still works, but any file *deletions* the agent made in that run can't be reproduced (the original capture only preserved files the agent created or modified, not the ones it removed). For tasks where the agent doesn't delete anything (most behavioral-rating tasks where the agent just writes `answer.md`), this doesn't matter at all. For tasks where the agent edits code and may have deleted files, you may want to re-capture a fresh reference run after the next time you run `scripts/harbor-run`.