Project-2 baseline

This commit is contained in:
2026-10-04 21:19:23 -04:00
parent 0f04889edf
commit 213eb3c403
861 changed files with 1710 additions and 3363322 deletions

View File

@@ -52,7 +52,7 @@ The five shapes can co-occur, and any one of them gets verdicted as a leak. Shap
- **`not-applicable`** — There is no way to decide leakage from this submission. Three triggers:
- **No snapshot**: `harbor-tasks/<slug>/environment/session.jsonl` does not exist. The task isn't a snapshot task; there's nothing for the snapshot to leak. Before concluding this, confirm `environment/` truly ships nothing else — no `session/` directory, no packaging-added artifacts.
- **Empty snapshot**: `environment/session.jsonl` exists but is empty (zero bytes or whitespace-only). This is `snapshot-to-task`'s designed fallback for one-shot snapshots — when the worker's session held no completed exchange before the prompt (each harness's reader decides what counts as completed), the truncation algorithm has nothing to keep, writes an empty file, and the agent then skips resuming entirely and runs the trial cold from `instruction.md`. An empty `session.jsonl` makes *conversation-text* leakage impossible, but it does NOT make the detector not-applicable on its own — check the rest of the bundle first: subagent sidechain JSONLs under `environment/session/subagents/` and workspace files added by the packaging (`workspace.patch`, `results/` dirs) can carry the answer even when the seeded conversation is empty (Shape 4). Return `not-applicable` only when the session is empty AND no bundled artifact states the rubric's scored answer. (See `plugins/create-snapshot/snapshot-to-task.ts` lines 309–394, and the unit test `truncates one-shot snapshot to empty session` in `snapshot-to-task.test.ts`.) A non-empty `session-full.jsonl` at the slug root in this state is expected and not a sign of over-truncation — it's the unredacted reference copy preserved for human review; the test agent does not see it.
- **Empty snapshot**: `environment/session.jsonl` exists but is empty (zero bytes or whitespace-only). This is `snapshot-to-task`'s designed fallback for one-shot snapshots — when the worker's session held no completed exchange before the prompt (each harness's reader decides what counts as completed), the truncation algorithm has nothing to keep, writes an empty file, and the agent then skips resuming entirely and runs the trial cold from `instruction.md`. An empty `session.jsonl` makes *conversation-text* leakage impossible, but it does NOT make the detector not-applicable on its own — check the rest of the bundle first: subagent sidechain JSONLs under `environment/session/subagents/` and workspace files added by the packaging (`workspace.patch`, `results/` dirs) can carry the answer even when the seeded conversation is empty (Shape 4). Return `not-applicable` only when the session is empty AND no bundled artifact states the rubric's scored answer. (See `toolkit/plugins/create-snapshot/snapshot-to-task.ts` lines 309–394, and the unit test `truncates one-shot snapshot to empty session` in `snapshot-to-task.test.ts`.) A non-empty `session-full.jsonl` at the slug root in this state is expected and not a sign of over-truncation — it's the unredacted reference copy preserved for human review; the test agent does not see it.
- **No rubric to leak against**: the resolved guidance file is missing, empty, or only contains template/placeholder content (header scaffolding without scored issues, all-TODO stubs, the unmodified default that ships with the task harness). Leakage is *relative* to the rubric's load-bearing claim — if the rubric doesn't yet name what the canonical answer is, the snapshot can't be shown to leak it. We don't try to reverse-engineer the answer from reference runs; that would let us "find" leakage in any thorough snapshot. Wait for the rubric to land, then re-run.
- **`clear-leak`** — Shape 1, strong Shape 2, Shape 3, strong Shape 4, or strong Shape 5. Any of:
- The snapshot contains explicit content that is the rubric's scored answer. Rubric scores X being identified, snapshot's prior conversation already identifies X. Rubric scores calibrated hedging (the agent should state its uncertainty plainly), snapshot ends with the calibrated hedge. Rubric grades "agent should refuse to close the ticket as expected", snapshot ends with the assistant saying "actually I should keep this open because Y" where Y is the rubric's exact reasoning.