restored entire zip and config'd

This commit is contained in:
2026-10-07 15:37:45 -04:00
parent 4bd9264a29
commit 5ca605dcce
219 changed files with 43370 additions and 0 deletions

View File

@@ -0,0 +1,78 @@
---
name: regrade-reference-run
description: Re-run a task's verifier (the grader) against a reference run you already captured, skipping the agent. Use when iterating on tests/holistic-rubric.md, measuring grader variance, or sanity-checking a verifier change — anywhere you'd otherwise re-spend minutes of agent runtime just to get a fresh grade against the same agent behavior.
allowed-tools: Bash, Read, Glob, Grep
---
# Re-grade a reference run without re-running the agent
## When to use this
You have a `harbor-tasks/<slug>/reference-runs/<run-id>/` directory captured by an earlier real trial — its `agent-output/`, `agent/trajectory.json`, `grade.md`, `reward.txt`, and `reward-correctness.txt` are all on disk. You want to grade that captured run again. Most common reason: you edited `tests/holistic-rubric.md` and want to see how the new wording changes the scores against the same agent behavior, without paying for a fresh agent run.
Regrading re-derives the full grade — every criterion's reasoning in `grade.md` and the reward — so it is the right tool for iterating on any part of your rubric.
Other good fits:
- **Grader variance.** Run the same reference 10× in parallel, look at the spread in `reward.txt`. Useful when you suspect the grader is non-deterministic on a borderline call.
- **Sanity-check a verifier change.** If you patched `tests/test.sh` itself, regrade an existing reference run to confirm the patch produces the same grade against the same agent behavior.
## How it works
`scripts/harbor-regrade` invokes the standard `harbor run` plumbing but plugs in a replay adapter (`scripts/replay_agent.py`) instead of an agent. The adapter:
1. Uploads your captured `agent-output/` into the trial container's `/workspace` — overlays the agent's surviving file edits on top of the base workspace built by the task's `Dockerfile`.
2. If `agent-output/_HARBOR_DELETIONS.txt` exists (records of any files the agent deleted), `rm`s each listed path so the workspace state ends up identical to where the original agent left it.
3. Uploads the captured `agent/trajectory.json` so the grader reads the same transcript it would have on the original run.
Then the verifier (`tests/test.sh`) runs exactly as it does for any other trial. Same `git ls-files`/`git diff` workspace capture, same deterministic checks, same grader, same `reward.txt`/`reward-correctness.txt`/`grade.md` output. The only difference is that the agent phase is now seconds of file overlay instead of minutes of agent work.
## How to invoke
```sh
scripts/harbor-regrade <task-dir> <reference-run-dir>... [-k N] [extra harbor args]
scripts/harbor-regrade <task-dir> --all [--jobs N] [extra harbor args]
```
- `<task-dir>`: `harbor-tasks/<slug>` — same dir you'd pass to `scripts/harbor-run`.
- `<reference-run-dir>`: `harbor-tasks/<slug>/reference-runs/<run-id>`. Name several to re-grade them all.
- `--all`: re-grade every run under `harbor-tasks/<slug>/reference-runs/`. This is what you usually want after editing your rubric.
- `--jobs N`: how many to re-grade at once, default 2. Each one is a container, so raise it only as far as your machine comfortably allows.
- `-k N`: N independent regrades **of a single run**, for variance measurement.
`-k` and `--all` are easy to mix up. `-k 4` grades **one** captured run four times; `--all` grades **each** of your captured runs once. If you want fresh grades for all four of your reference runs, that's `--all`, not `-k 4`.
## What grades the run
The grader scores the eight criteria of the Grading Standard against the task's holistic rubric (`tests/holistic-rubric.md`; a task started on an earlier toolkit carries the same document as `tests/grader-guidance-consolidated.md`). The reward is the mean of the non-N/A criteria minus any overall penalties, floored at 0. `reward-correctness.txt` always reads `N/A` by design — correctness lives inside the criteria rather than as a separate score — so only the reward and the criterion reasoning move when you edit the rubric.
The regrade uses the grader assets already in the task's `tests/` directory, so a run regrades under the same standard it was originally graded with.
Output lands in `harbor-jobs/<job>/<trial-id>/` like any other harbor trial — `<job>` is a timestamp for a single regrade, and `regrade-<n>-<run-id>` for each run under `--all`, so you can tell at a glance which reference run a result came from — `verifier/reward.txt`, `verifier/reward-correctness.txt`, `verifier/reward.json`, `verifier/grade.md`, `verifier/test-stdout.txt`, `trial.log`. To see how the new grade diverges from the original:
```sh
diff harbor-tasks/<slug>/reference-runs/<run-id>/grade.md \
harbor-jobs/<job>/<trial-id>/verifier/grade.md
```
For the number alone, the tail of `verifier/test-stdout.txt` prints it, or compare directly:
```sh
echo "before: $(cat harbor-tasks/<slug>/reference-runs/<run-id>/reward.txt)"
echo "after: $(cat harbor-jobs/<job>/<trial-id>/verifier/reward.txt)"
```
## Typical iteration loop
1. Run a few real trials to capture reference runs: `scripts/harbor-run harbor-tasks/<slug> -k 4`, then `npx tsx scripts/copy-reference-run.ts harbor-jobs/<job>/<trial>` for each one you want to keep.
2. Read the captured `grade.md` files — every criterion section, not just the headline score — and find places where the grader's judgment doesn't match what you'd say as the task author.
3. Edit `tests/holistic-rubric.md` to clarify the points the grader got wrong.
4. **`scripts/harbor-regrade harbor-tasks/<slug> --all`** to re-grade every captured run in one go.
5. Diff the new `grade.md` files vs the originals. Repeat until the grader is reasoning correctly on each captured behavior.
Leave the regrade output where it lands. It is there for you to read and compare, not to copy back over your reference runs — a regrade is a new trial with a new id, so `copy-reference-run.ts` would *add* a run rather than update one, leaving you with twice as many and no way to tell which grades came from which version of your rubric. Your reference runs should stay as the real trials you captured.
This is much faster (and cheaper) than re-running `scripts/harbor-run` after every grader edit, because each agent run takes minutes and produces a *different* trajectory anyway — so re-running confounds "is the grader better?" with "is the agent behavior different?".
## Caveat: old reference runs
If your reference run was captured before this toolkit version, its `agent-output/` won't include `_HARBOR_DELETIONS.txt`. The replay still works, but any file *deletions* the agent made in that run can't be reproduced (the original capture only preserved files the agent created or modified, not the ones it removed). For tasks where the agent doesn't delete anything (most behavioral-rating tasks where the agent just writes `answer.md`), this doesn't matter at all. For tasks where the agent edits code and may have deleted files, you may want to re-capture a fresh reference run after the next time you run `scripts/harbor-run`. The same applies to a run captured before this version in which the agent renamed a file with `git mv`: the rename's source path is missing from `_HARBOR_DELETIONS.txt`, so the replay keeps both copies.