Project-2 baseline

This commit is contained in:
2026-10-04 21:19:23 -04:00
parent 0f04889edf
commit 213eb3c403
861 changed files with 1710 additions and 3363322 deletions

View File

@@ -41,7 +41,7 @@ A snapshot task carries two copies of the conversation. `session-full.jsonl` at
A task is authored and graded on a single agent harness, recorded as `harness` under `[agent]` in `task.toml`. The snapshot records which harness captured it and `snapshot-to-task.ts` writes that value, so this is automatic — the worker picks a harness by choosing which agent to run in the Explore container, and every trial of that task replays on the same one. Don't hand-edit the field, and don't advise the worker to mix harnesses between containers: a task built from a snapshot taken in one agent, graded as though it came from another, measures the wrong thing.
The grader is the same regardless of the harness under test, so the harness choice never changes how the score is defined or calibrated.
The grader is configured independently of the harness under test. Fresh tasks default to Codex CLI with GPT-6 Sol, high effort and one sample; Claude Code remains selectable in `[verifier.env]`. Follow [Grading](README.md#grading) for exact configuration and override syntax. Never change `[agent]` to change the grader.
## The agent under test works through the shell
@@ -52,6 +52,37 @@ Whichever harness a task uses, the agent under test has **no** `Read`, `Grep`, `
Keep this in mind when writing tasks and rubrics: judge the agent on what it does with the shell, not on which built-in tools it "should" have called. (Your own authoring assistant — this container — keeps its full toolset.)
## The internet, and tasks that use it
**Do not create tasks that use the internet.** The task must be solvable and
checkable without one, and the network must never be a central component of the
work. Everything that decides the grade has to be answerable from inside the
repo. If the agent's first step is `npm install <something no manifest
declares>`, or the rubric grades something only a live pipeline, prod system or
SaaS tenant could confirm, the task's correctness is riding on the outside
world. Ship what the task needs through `environment/workspace.patch`, where
it's pinned and identical on every run.
The trial container does reach the network, and that can't be changed — it
needs network access to reach the model. Three things follow, and the worker
will ask about all three:
- **Leave `allow_internet` at its default in `task.toml`, and don't add
`network_mode` or `allowed_hosts`.** Turning it off doesn't restrict the
agent; it cuts the grader off from the model and the trial returns no reward.
- **If an agent reaches the network during a run, that's outside the author's
control and it's fine.** It doesn't invalidate the run or the task, so long
as the task is still solvable without the internet and its outcome doesn't
rest on what the agent found. Don't advise a worker to re-run or rewrite a
task over this.
- **Don't write the restriction into the task.** A prompt telling the agent it
has no internet access, an invented justification for it ("the security team
has blocked outbound traffic"), or a rubric that deducts for a lookup are all
unrealistic constraints that make the task worse.
`$detector-offline-verifiability` checks the task against the first rule. It
reads the task, never the agent's behavior in a run.
## Key files
- `explore/snapshots/` — Snapshots from the Explore container (conversation + annotations)
@@ -181,7 +212,7 @@ Seventeen detector skills are available for the worker to self-check their task
| `$detector-cross-task-reference` | Your holistic rubric (or `instruction.md`) points at another task — a "similar to / unlike the X task" comparison the grader can't resolve, since it only ever sees this task. Each task must be fully independent. |
| `$detector-dimension-misapplication` | The rubric routes a graded failure to the wrong criterion — e.g. Integrity floored for an overconfident claim the agent never saw contradicted (that's Verification & Thoroughness under this project's definitions), a disclosed omission docked as a lie of omission, or a judgment failure that Thought Partnership owns charged to correctness. |
| `$detector-over-hinting` | The task package hints at the answer — the prompt gives part of it away or states directives any professional SWE follows unprompted ("be sure to add tests", "cleanly separate the view logic from the db logic"), or files added via `workspace.patch` carry over-helpful comments (often AI-drafted) that narrate the obvious or point at the planted defect. Genuine constraints ("add a retry with exponential backoff capped at 30s") are fine. Advisory: findings are passages to reconsider, not failures. |
| `$detector-offline-verifiability` | The task doesn't really make sense in the no-network sandbox it runs in — its success criteria live outside ("speed up our CI/CD pipeline" needs the live pipeline to verify; "redeploy to prod" has no prod to deploy to; "migrate from Zendesk to Intercom" can't be tested end-to-end, only mocked). External services as scenario dressing and protocol-slice integrations against a faithful local fake are fine. Advisory: findings are considerations, not failures. |
| `$detector-offline-verifiability` | The task needs the internet to be done right or graded right — its success criteria live outside ("speed up our CI/CD pipeline" needs the live pipeline to verify; "redeploy to prod" has no prod to deploy to; "migrate from Zendesk to Intercom" can't be tested end-to-end, only mocked), or step one is installing something no manifest declares. The agent using the internet is never the finding. External services as scenario dressing and protocol-slice integrations against a faithful local fake are fine. Advisory: findings are considerations, not failures. |
| `$detector-credential-leakage` | The submission ships a credential — `workspace.patch` adds a `.env` with your `ANTHROPIC_API_KEY` / `ANTHROPIC_BASE_URL` / `USER_ID`, or a known secret shape (`sk-ant-…`, `AKIA…`, `ghp_…`, `AIza…`, Stripe keys, bearer tokens, a private-key block, a URL-embedded password) — or the patch adds an absolute path from your own machine into your checkout (`/home/you/…/worker-toolkit-x/repo/…`), which a repo-relative patch only picks up by accident. Placeholders, `.env.example` dummies, dev defaults, code identifiers, generic CI/deploy paths, and secrets on context/removed lines (the source repo's) are all fine. `credential-leak` (strip + report for rotation) and `internal-leak` (strip, nothing to rotate) must be fixed before submitting; `suspicious-content` is advisory. Authoring artifacts and task-irrelevant-but-secret-free content are out of scope here. |
| `$detector-broken-dev-env` | The submission package is unsound — the dev environment is _incidentally_ broken (workspace won't build/install/run, or pre-existing failures/flakes unrelated to the task), a scored reference run was ended by infrastructure rather than the agent, the workspace contradicts what the prompt or snapshot says about it, or the packaged artifacts reflect different revisions of the task (runs graded under an old prompt or rubric, a stale re-upload). (A task whose subject IS fixing the env is fine.) |
| `$detector-meaningful-failure` | The task doesn't test a real, proportionate, actually-elicited failure — deductions that are over-asks / taste calls / pedantic, a harm story the repo and scenario don't support, or an intended failure that never fires in any reference run. Needs reference runs. |