ren worker folder adding orig, mv new one into root

This commit is contained in:
2026-09-25 10:34:29 -04:00
parent 10f0668e32
commit 5b010039d7
1308 changed files with 44597 additions and 1511 deletions

View File

@@ -17,10 +17,15 @@ the workspace, and nothing outside it. A task fits that world when everything
is totally verifiable from within the repo: offline-completable and
offline-verifiable, because the setup happened before the network went away.
Setup installs what the repo's own manifests and lockfiles declare at the
pinned commit — nothing more. A library the ask requires the agent to *add*
was never installed, so acquiring it means `bundle add`, `npm install <pkg>`,
`pip install` — a registry fetch, mid-task.
Setup installs what the repo's own manifests and lockfiles declare **after
`environment/workspace.patch` has been applied** — the image copies the patched
workspace in and only *then* runs the dependency install. So a package the
author added, upgraded, downgraded or re-pinned in the patch is present in the
sandbox, and is never a completability finding; judge the manifests as the
patch leaves them, not as the pinned commit left them. What is never installed
is a library the ask requires the *agent* to add: acquiring that means `bundle
add`, `npm install <pkg>`, `pip install` — a registry fetch, mid-task, after
the network is gone.
**Do not consider the task's network policy. At all.** `task.toml`'s
`allow_internet` / `network_mode` / `allowed_hosts` fields are not about the
@@ -96,8 +101,8 @@ Break that into the two halves:
**This half has a mechanical check, and it is not optional.** List every
library, framework, runner, or binary the ask or the rubric's criteria
name, then check each against every manifest and lockfile in the repo
(`Gemfile`/`Gemfile.lock`, `package.json` + its lockfile,
name, then check each against every manifest and lockfile in the repo **as
the workspace patch leaves it** (`Gemfile`/`Gemfile.lock`, `package.json` + its lockfile,
`pyproject.toml`/`requirements*.txt`/`uv.lock`, `go.mod`, the Dockerfile).
Read the files — never settle this from knowledge of what the framework
supports. When a name is absent from all of them, the deciding question is

View File

@@ -3,7 +3,7 @@ name: detector-rubric-form
description: |
Self-check that your atomic rubric is well-formed. A deterministic contract
checks the artifact: the file parses against the criterion schema,
criteria number 2 to 24, ids are kebab-case and unique, category and
ids are kebab-case and unique, category and
severity use the defined vocabularies, extra_credit criteria carry no
severity, at most 2 criteria are crux, `dimensions` names grading-standard
criteria, and no text states a numeric penalty amount. A judgment layer

View File

@@ -56,24 +56,23 @@ Report each failure with the offending text quoted verbatim.
`task` string and a `criteria` list. A file that does not parse is a
broken artifact; report the parse error and verdict `material-issues`.
2. **`task` names this task.** The `task` field equals the task's slug.
3. **Criteria count is 2 to 24.**
4. **Ids are kebab-case and unique.** Each `id` matches
3. **Ids are kebab-case and unique.** Each `id` matches
`^[a-z0-9]+(-[a-z0-9]+)*$` and appears once.
5. **`category` vocabulary.** One of `primary_intent`, `extra_credit`,
4. **`category` vocabulary.** One of `primary_intent`, `extra_credit`,
`dodged_bullet`.
6. **`severity` vocabulary and placement.** One of `crux`,
5. **`severity` vocabulary and placement.** One of `crux`,
`certain_dealbreaker`, `possible_dealbreaker`, `unlikely_dealbreaker`.
Required on `primary_intent` and `dodged_bullet` criteria. Forbidden on
`extra_credit` criteria.
7. **Crux cap.** At most 2 criteria carry `severity: crux`.
8. **`dimensions` names at least one grading-standard criterion.** Each entry
6. **Crux cap.** At most 2 criteria carry `severity: crux`.
7. **`dimensions` names at least one grading-standard criterion.** Each entry
is one of the eight, exactly as the grading standard names them:
`Integrity`, `Narrow Correctness`,
`Broader Correctness / the craft of software engineering`, `Persistence`,
`Communication`, `Verification & Thoroughness`, `Common Sense`,
`Thought Partnership`.
9. **`guideline` is non-empty** on every criterion.
10. **Zero numeric penalty language.** Penalty weight is expressed through
8. **`guideline` is non-empty** on every criterion.
9. **Zero numeric penalty language.** Penalty weight is expressed through
`category` and `severity`; sizing the subtraction is the grading
machinery's job. No guideline, elaboration, or context-document sentence
may state a numeric penalty amount. Run these over the atomic rubric AND

View File

@@ -29,12 +29,17 @@ Then the verifier (`tests/test.sh`) runs exactly as it does for any other trial.
## How to invoke
```sh
scripts/harbor-regrade <task-dir> <reference-run-dir> [-k N] [extra harbor args]
scripts/harbor-regrade <task-dir> <reference-run-dir>... [-k N] [extra harbor args]
scripts/harbor-regrade <task-dir> --all [--jobs N] [extra harbor args]
```
- `<task-dir>`: `harbor-tasks/<slug>` — same dir you'd pass to `scripts/harbor-run`.
- `<reference-run-dir>`: `harbor-tasks/<slug>/reference-runs/<run-id>` — must contain `agent-output/`.
- `-k N`: N independent regrades against the same captured state. Use for variance measurement.
- `<reference-run-dir>`: `harbor-tasks/<slug>/reference-runs/<run-id>`. Name several to re-grade them all.
- `--all`: re-grade every run under `harbor-tasks/<slug>/reference-runs/`. This is what you usually want after editing your rubric.
- `--jobs N`: how many to re-grade at once, default 2. Each one is a container, so raise it only as far as your machine comfortably allows.
- `-k N`: N independent regrades **of a single run**, for variance measurement.
`-k` and `--all` are easy to mix up. `-k 4` grades **one** captured run four times; `--all` grades **each** of your captured runs once. If you want fresh grades for all four of your reference runs, that's `--all`, not `-k 4`.
## What grades the run
@@ -42,18 +47,18 @@ The grader scores the eight criteria of the Grading Standard against the task's
The regrade uses the grader assets already in the task's `tests/` directory, so a run regrades under the same standard it was originally graded with.
Output lands in `harbor-jobs/<timestamp>/<trial-id>/` like any other harbor trial — `verifier/reward.txt`, `verifier/reward-correctness.txt`, `verifier/reward.json`, `verifier/grade.md`, `verifier/test-stdout.txt`, `trial.log`. To see how the new grade diverges from the original:
Output lands in `harbor-jobs/<job>/<trial-id>/` like any other harbor trial — `<job>` is a timestamp for a single regrade, and `regrade-<n>-<run-id>` for each run under `--all`, so you can tell at a glance which reference run a result came from — `verifier/reward.txt`, `verifier/reward-correctness.txt`, `verifier/reward.json`, `verifier/grade.md`, `verifier/test-stdout.txt`, `trial.log`. To see how the new grade diverges from the original:
```sh
diff harbor-tasks/<slug>/reference-runs/<run-id>/grade.md \
harbor-jobs/<timestamp>/<trial-id>/verifier/grade.md
harbor-jobs/<job>/<trial-id>/verifier/grade.md
```
For the number alone, the tail of `verifier/test-stdout.txt` prints it, or compare directly:
```sh
echo "before: $(cat harbor-tasks/<slug>/reference-runs/<run-id>/reward.txt)"
echo "after: $(cat harbor-jobs/<timestamp>/<trial-id>/verifier/reward.txt)"
echo "after: $(cat harbor-jobs/<job>/<trial-id>/verifier/reward.txt)"
```
## Typical iteration loop
@@ -61,9 +66,11 @@ echo "after: $(cat harbor-jobs/<timestamp>/<trial-id>/verifier/reward.txt)"
1. Run a few real trials to capture reference runs: `scripts/harbor-run harbor-tasks/<slug> -k 4`, then `npx tsx scripts/copy-reference-run.ts harbor-jobs/<job>/<trial>` for each one you want to keep.
2. Read the captured `grade.md` files — every criterion section, not just the headline score — and find places where the grader's judgment doesn't match what you'd say as the task author.
3. Edit `tests/holistic-rubric.md` to clarify the points the grader got wrong.
4. **`scripts/harbor-regrade harbor-tasks/<slug> harbor-tasks/<slug>/reference-runs/<run-id>`** for each captured run you care about.
4. **`scripts/harbor-regrade harbor-tasks/<slug> --all`** to re-grade every captured run in one go.
5. Diff the new `grade.md` files vs the originals. Repeat until the grader is reasoning correctly on each captured behavior.
Leave the regrade output where it lands. It is there for you to read and compare, not to copy back over your reference runs — a regrade is a new trial with a new id, so `copy-reference-run.ts` would *add* a run rather than update one, leaving you with twice as many and no way to tell which grades came from which version of your rubric. Your reference runs should stay as the real trials you captured.
This is much faster (and cheaper) than re-running `scripts/harbor-run` after every grader edit, because each agent run takes minutes and produces a *different* trajectory anyway — so re-running confounds "is the grader better?" with "is the agent behavior different?".
## Caveat: old reference runs

View File

@@ -1,6 +1,6 @@
---
name: write-atomic-rubric
description: Convert a task's finished holistic rubric into the atomic rubric package — tests/atomic-rubric.yaml (criteria with id, category, severity, dimensions, guideline, elaboration) plus tests/grader-context.md (task context, business context, and ground truth, extracted verbatim). Covers Maximum Viable Atomicity, positive guideline phrasing with bold inline answer keys, conditional criteria, dodged-bullet escalation pairs, Crux designation from the holistic rubric's heavy penalties (at most two per task), the schema rules (2-24 criteria; kebab-case ids; no numeric penalty language; no severity on extra_credit), and staging and validation. Use after the holistic rubric is final.
description: Convert a task's finished holistic rubric into the atomic rubric package — tests/atomic-rubric.yaml (criteria with id, category, severity, dimensions, guideline, elaboration) plus tests/grader-context.md (task context, business context, and ground truth, extracted verbatim). Covers Maximum Viable Atomicity, positive guideline phrasing with bold inline answer keys, conditional criteria, dodged-bullet escalation pairs, Crux designation from the holistic rubric's heavy penalties (at most two per task), the schema rules (kebab-case ids; no numeric penalty language; no severity on extra_credit), and staging and validation. Use after the holistic rubric is final.
---
# Writing the Atomic Rubric
@@ -115,10 +115,12 @@ Each criterion carries:
citation, and code quotation, with markdown formatting (backticks, bold, fences)
intact. Never invent facts, paths, or requirements the source does not carry.
The file carries between 2 and 24 criteria; most tasks land in the teens. Every
scoring-relevant rule of the source lands in exactly one criterion's guideline or
elaboration. Content that is context rather than a requirement belongs in
`grader-context.md`, not in a criterion.
The criteria count follows the source. Every scoring-relevant rule of the source
lands in exactly one criterion's guideline or elaboration, and a rule is never dropped
or folded away to reach a target count. Parallel facets of one requirement that the
same evidence decides may share a criterion; distinct requirements get their own.
Content that is context rather than a requirement belongs in `grader-context.md`, not
in a criterion.
## Crux designation
@@ -181,7 +183,7 @@ the holistic rubric).
Reviewers working in a repo checkout also run
`npx tsx scripts/validate-rubrics-cli.ts --slug <task-slug>`, which enforces the same
schema, the criteria count, the Crux cap, and the numeric-penalty ban. That script is part
schema, the Crux cap, and the numeric-penalty ban. That script is part
of the review pipeline and does not ship in the toolkit.
## Related