ren worker folder adding orig, mv new one into root
This commit is contained in:
@@ -29,12 +29,17 @@ Then the verifier (`tests/test.sh`) runs exactly as it does for any other trial.
|
||||
## How to invoke
|
||||
|
||||
```sh
|
||||
scripts/harbor-regrade <task-dir> <reference-run-dir> [-k N] [extra harbor args]
|
||||
scripts/harbor-regrade <task-dir> <reference-run-dir>... [-k N] [extra harbor args]
|
||||
scripts/harbor-regrade <task-dir> --all [--jobs N] [extra harbor args]
|
||||
```
|
||||
|
||||
- `<task-dir>`: `harbor-tasks/<slug>` — same dir you'd pass to `scripts/harbor-run`.
|
||||
- `<reference-run-dir>`: `harbor-tasks/<slug>/reference-runs/<run-id>` — must contain `agent-output/`.
|
||||
- `-k N`: N independent regrades against the same captured state. Use for variance measurement.
|
||||
- `<reference-run-dir>`: `harbor-tasks/<slug>/reference-runs/<run-id>`. Name several to re-grade them all.
|
||||
- `--all`: re-grade every run under `harbor-tasks/<slug>/reference-runs/`. This is what you usually want after editing your rubric.
|
||||
- `--jobs N`: how many to re-grade at once, default 2. Each one is a container, so raise it only as far as your machine comfortably allows.
|
||||
- `-k N`: N independent regrades **of a single run**, for variance measurement.
|
||||
|
||||
`-k` and `--all` are easy to mix up. `-k 4` grades **one** captured run four times; `--all` grades **each** of your captured runs once. If you want fresh grades for all four of your reference runs, that's `--all`, not `-k 4`.
|
||||
|
||||
## What grades the run
|
||||
|
||||
@@ -42,18 +47,18 @@ The grader scores the eight criteria of the Grading Standard against the task's
|
||||
|
||||
The regrade uses the grader assets already in the task's `tests/` directory, so a run regrades under the same standard it was originally graded with.
|
||||
|
||||
Output lands in `harbor-jobs/<timestamp>/<trial-id>/` like any other harbor trial — `verifier/reward.txt`, `verifier/reward-correctness.txt`, `verifier/reward.json`, `verifier/grade.md`, `verifier/test-stdout.txt`, `trial.log`. To see how the new grade diverges from the original:
|
||||
Output lands in `harbor-jobs/<job>/<trial-id>/` like any other harbor trial — `<job>` is a timestamp for a single regrade, and `regrade-<n>-<run-id>` for each run under `--all`, so you can tell at a glance which reference run a result came from — `verifier/reward.txt`, `verifier/reward-correctness.txt`, `verifier/reward.json`, `verifier/grade.md`, `verifier/test-stdout.txt`, `trial.log`. To see how the new grade diverges from the original:
|
||||
|
||||
```sh
|
||||
diff harbor-tasks/<slug>/reference-runs/<run-id>/grade.md \
|
||||
harbor-jobs/<timestamp>/<trial-id>/verifier/grade.md
|
||||
harbor-jobs/<job>/<trial-id>/verifier/grade.md
|
||||
```
|
||||
|
||||
For the number alone, the tail of `verifier/test-stdout.txt` prints it, or compare directly:
|
||||
|
||||
```sh
|
||||
echo "before: $(cat harbor-tasks/<slug>/reference-runs/<run-id>/reward.txt)"
|
||||
echo "after: $(cat harbor-jobs/<timestamp>/<trial-id>/verifier/reward.txt)"
|
||||
echo "after: $(cat harbor-jobs/<job>/<trial-id>/verifier/reward.txt)"
|
||||
```
|
||||
|
||||
## Typical iteration loop
|
||||
@@ -61,9 +66,11 @@ echo "after: $(cat harbor-jobs/<timestamp>/<trial-id>/verifier/reward.txt)"
|
||||
1. Run a few real trials to capture reference runs: `scripts/harbor-run harbor-tasks/<slug> -k 4`, then `npx tsx scripts/copy-reference-run.ts harbor-jobs/<job>/<trial>` for each one you want to keep.
|
||||
2. Read the captured `grade.md` files — every criterion section, not just the headline score — and find places where the grader's judgment doesn't match what you'd say as the task author.
|
||||
3. Edit `tests/holistic-rubric.md` to clarify the points the grader got wrong.
|
||||
4. **`scripts/harbor-regrade harbor-tasks/<slug> harbor-tasks/<slug>/reference-runs/<run-id>`** for each captured run you care about.
|
||||
4. **`scripts/harbor-regrade harbor-tasks/<slug> --all`** to re-grade every captured run in one go.
|
||||
5. Diff the new `grade.md` files vs the originals. Repeat until the grader is reasoning correctly on each captured behavior.
|
||||
|
||||
Leave the regrade output where it lands. It is there for you to read and compare, not to copy back over your reference runs — a regrade is a new trial with a new id, so `copy-reference-run.ts` would *add* a run rather than update one, leaving you with twice as many and no way to tell which grades came from which version of your rubric. Your reference runs should stay as the real trials you captured.
|
||||
|
||||
This is much faster (and cheaper) than re-running `scripts/harbor-run` after every grader edit, because each agent run takes minutes and produces a *different* trajectory anyway — so re-running confounds "is the grader better?" with "is the agent behavior different?".
|
||||
|
||||
## Caveat: old reference runs
|
||||
|
||||
Reference in New Issue
Block a user