ren worker folder adding orig, mv new one into root

This commit is contained in:
2026-09-25 10:34:29 -04:00
parent 10f0668e32
commit 5b010039d7
1308 changed files with 44597 additions and 1511 deletions

View File

@@ -18,10 +18,19 @@ The criteria are defined in the grader system prompt (`harbor-tasks/<slug>/tests
## Context — two paths to a task
**Snapshot path:** The worker explored the codebase in the Explore container, found a behavior worth grading, and captured it with `$snapshot`. The snapshot (in `explore/snapshots/`) contains the full conversation transcript (`session-full.jsonl`) and worker annotations describing what behavior they observed and why it matters. If the worker asks you to help with the holistic rubric, start by reading these and invoking the `$write-holistic-rubric` skill.
**Snapshot path:** The worker explored the codebase in the Explore container, found a behavior worth grading, and captured it with `$snapshot`. The snapshot (in `explore/snapshots/`) contains the conversation transcript and worker annotations describing what behavior they observed and why it matters. Building the task copies the whole transcript to `harbor-tasks/<slug>/session-full.jsonl`. If the worker asks you to help with the holistic rubric, start by reading that and `annotation.json`, and invoke the `$write-holistic-rubric` skill — but read the next section before you write a criterion against anything in it.
**Manual path:** The worker is building a task from scratch — they will have explored on their own and have a specific behavior in mind. Follow their lead.
### What the test agent inherits, and what it doesn't
A snapshot task carries two copies of the conversation. `session-full.jsonl` at the task root is the whole capture — it exists for you and the worker to read. `environment/session.jsonl` is the one the test agent resumes from, and it stops at the last clean assistant turn before the worker's final message: that message becomes `instruction.md`, and the reply to it is dropped so the test agent has to produce its own.
`environment/session.jsonl` is therefore the authority on what the test agent knows. Two things to check against it before the rubric is written:
- **Does the prompt still make sense on its own?** When `instruction.md` answers something ("yes, do 1, 2 and 4", "go with that approach"), confirm the thing it answers survived the cut. If it didn't, the test agent is replying to a plan it can't see, and the task needs a rewritten prompt rather than a rubric.
- **Can the test agent reach every fact you grade?** A criterion drafted from `session-full.jsonl` can quietly require knowledge that only exists past the cut. Anything you expect the response to know has to be in the injected session, the prompt, or the workspace.
## Architecture
1. **Explore container** (`explore/`) — Where codebase exploration happened. Snapshots saved to `explore/snapshots/`.
@@ -66,6 +75,7 @@ Keep this in mind when writing tasks and rubrics: judge the agent on what it doe
- `npx tsx scripts/copy-reference-run.ts harbor-jobs/<job>/<trial>` — Copy a single reference run
- `npx tsx scripts/copy-reference-run.ts harbor-jobs/<job>/<slug>__*` — Copy all trials from a `-k 4` run (recommended; `submit-task.ts` expects ≥4 reference runs)
- `scripts/harbor-regrade harbor-tasks/<slug> harbor-tasks/<slug>/reference-runs/<run-id>` — Re-grade a captured reference run without re-running the agent. Use after editing `tests/holistic-rubric.md`. See the `$regrade-reference-run` skill.
- `scripts/harbor-regrade harbor-tasks/<slug> --all` — Re-grade every captured reference run, 2 at a time (`--jobs N` to change). Note `-k` re-grades ONE run N times rather than N different runs.
- `npx tsx scripts/submit-task.ts <slug>` — Validate and package for submission
## Toolkit-managed files — never edit these
@@ -186,7 +196,7 @@ When the worker asks "is my task ready to submit?" or hits a specific concern (r
**For snapshot-based tasks:**
- **Read the snapshot context** — start with `session-full.jsonl` and `annotation.json` in the snapshot directory to understand what behavior the worker thought was worth grading
- **Read the snapshot context** — start with `harbor-tasks/<slug>/session-full.jsonl` and the snapshot's `annotation.json` to understand what behavior the worker thought was worth grading, then read `environment/session.jsonl` for what the test agent actually inherits
- **Verify factual claims** — the worker knows what they observed. Read the specific files they point to and confirm their claims about the code are accurate
- **Draft the holistic rubric** — use the `$write-holistic-rubric` skill, which will guide the conversation toward eliciting the worker's privileged information