20 KiB
Changelog
7b6b67ea3d
- Fixed: the breezy-complete and zeta toolkits build their containers again. The Debian release they are built on left long-term support and its package mirror is being retired, so building an Explore container or a task image failed part-way with a "404 Not Found" on a system package; those packages now come from Debian's archive instead.
- Fixed: on the breezy-complete toolkit, the Explore container now prepares its database reliably. A boot-time cache could corrupt itself while loading one of the app's larger dependencies, which left the database setup failing and the app with nothing to run against; that cache is now off in Explore, as it already was for task images.
- Fixed:
codexno longer fails to authenticate when your.envwas saved on Windows. Windows (CRLF) line endings left a stray character on the end of your key and codex was rejected with an API-key error; the key is now cleaned wherever it is read, so your.envneeds no change.
fa77be2885
- Grading no longer fails silently when your task image carries an older Claude Code. The grader model needs Claude Code 2.1.251 or newer. A task image installs Claude Code when it is first built and keeps that copy on later rebuilds, so an image built before that version failed every grade with "does not support this model" and the trial ended with no reward file.
harbor-runnow checks your task images before a local trial and rebuilds any that are too old, task images verify the version when they build, and the grader stops with a clear message if an old copy still reaches it. - Toolkit documents no longer point at files that ship only in our review pipeline. The atomic-rubric skill describes the validation the staging script performs in the toolkit, the fact-check detector names
scripts/build-workspace.sh, and the corpus-viewer notes say they apply to zeta toolkits only.
d7edb3d5c1
-
The toolkit's grading documents are now named the holistic rubric and the atomic rubric. The holistic rubric is the per-task grading document the grader reads alongside the shared Grading Standard; earlier releases called it the grader guidance. The atomic rubric is a YAML companion that restates the same requirements as separately judgeable criteria. The content rules for both are unchanged. This release adopts the names, renames the files that new tasks create, and ships rubric grading in the toolkit.
- New tasks write
tests/holistic-rubric.mdandtests/atomic-rubric.yaml. A task created on this toolkit scaffoldstests/holistic-rubric.mdas its holistic rubric. The atomic rubric package istests/atomic-rubric.yamlplustests/grader-context.md, authored after the holistic rubric is final. - A task created on an earlier toolkit version keeps its existing filenames and stays fully supported. The filename-stability promise carries forward for every existing task: grading, the detector skills,
scripts/harbor-regrade, andsubmit-taskreadtests/grader-guidance-consolidated.md, legacytests/grader-guidance.md, andtests/rubrics.yamlwherever a task carries them, indefinitely, so moving an existing task between toolkit versions still never means renaming files. Never rename a committed task file. Only new tasks use the new names. /write-holistic-rubricreplaces/write-grader-guidance-consolidated($write-holistic-rubricin codex). It is the same authoring skill under the current name, and it now also teaches length discipline: a finished holistic rubric lands near 1,500 words; a 4,000-to-5,000-word draft is repetition, not thoroughness; an edit never grows the document.- New:
/write-atomic-rubric($write-atomic-rubricin codex) converts a finished holistic rubric intotests/atomic-rubric.yamlplustests/grader-context.md. Every task-specific requirement becomes one separately judgeable criterion, and the context and ground truth those criteria rely on are extracted alongside. - Rubric grading ships in the toolkit. The rubric renderer (
render-rubric-grade.py) is included undertask-shared/and scaffolded into new tasks. Once a task's atomic rubric is written, stage its grading copies withnpx tsx scripts/stage-atomic-rubric.ts <task-slug>;scripts/harbor-regradethen re-grades a captured run in rubric mode with no patch. Run the staging script with--restoreto remove the staged copies before packaging. - Grading runs on
claude-fable-5-1. New tasks and freshly staged rubric assets grade withclaude-fable-5-1by default. A task that shipped with an earlier grader keeps that grader unless you override it, so existing scores stay comparable. Override either way withGRADER_MODEL=.... - Two new detector self-checks:
/detector-rubric-coverageand/detector-rubric-form. Coverage checks that your atomic rubric tracks your holistic rubric, so no load-bearing requirement, penalty, or "do not penalize" rule is missing from the criteria and no criterion invents one. Form checks the atomic rubric as an artifact: the criterion schema, atomicity, positive phrasing, and inline answer keys.
- New tasks write
-
Fixed: on the stocks-in-the-future toolkit, a re-graded run's minitest check now actually runs the suite. The container used to build its databases at start-up, so a check running soon after could hit a missing
stocks_in_the_future_test; both databases now ship inside the image. -
Fixed: on the zeta toolkits,
run-appno longer leaves a.venvbehind for the Python members. Dependencies now install into the container's Python, matching the graded image — so if you switch between Python members, re-runrun-appfor the one you're working on. -
The note at the top of
tests/test-commands.shno longer tells you not to edit it. Task-specific checks there are expected and kept. -
Fixed:
run-app potion-multi-dsr-watchernow boots. It had no database URL and started a cron job that never opened a port, sorun-apptimed out waiting for one; it now serves its HTTP entrypoint on port 3000. -
Codex (gpt-5.6-sol) is now the default agent. A manual task now scaffolds with
harness = "codex", and the docs start you incodex; Claude Code remains fully supported, and a task keeps whichever agent authored it. -
Fixed: re-grading a run where your agent renamed a file with
git mvno longer brings the old file back. The verifier recorded the rename as a new file only, so the re-graded workspace held both copies and the stale one broke the type-check or test suite — failures no agent caused. -
Fixed: a file your agent wrote at a path it had just removed or renamed away no longer disappears when the run is re-graded. The verifier listed that path as deleted even though the new file was sitting there, so the re-graded workspace lost it.
-
Fixed:
codexnow picks up a rotatedANTHROPIC_API_KEYwithout a container rebuild. It read its key from a file written when the container was created, so a key changed in.envafterwards left it failing to authenticate; each launch now re-reads.envfirst (in Explore, from the container's next start).claudewas never affected. -
Fixed: an Explore container that came up with an empty
/workspace/repos(or/workspace/repo) now repairs itself on the nextup. Unzipping a new toolkit over an old install could leave the container pointed at nothing, sorun-app <repo>failed withcheckout <sha> failedand rebuilding the container did not help. Reported by a worker. -
Containers now come up with their database already loaded. On the human-essentials and awbw toolkits the image used to build the database when the container started, so a trial could reach the test database before it was ready. The schema now ships inside the image, which also cuts container start-up time noticeably on awbw.
-
Fixed: on the human-essentials, zeta-platform and flaredown toolkits, a re-graded run's rspec check now actually runs the suite. The check could start before the container had finished loading the test database, in which case rspec aborted at load time and reported zero examples — which read as ordinary test failures. The verifier now waits for the schema before running any check.
-
Fixed: the same on the breezy-complete toolkit, where the container builds its databases for longer. The rspec check could report zero examples, or a missing
socratic_systems_test, on a run graded soon after the container started; the databases now ship inside the image. -
Fixed: the breezy-complete Explore container no longer seeds its database twice.
db:preparealready seeds the database it creates, so the second pass aborted partway on a duplicate record; seeding now runs only when the database has none. -
Fixed: on the awbw toolkit, restarting a container no longer leaves the test database half-loaded. Reloading the schema over an existing one failed on a foreign-key ordering in
db/schema.rb(MySQL error 3730), and the container hid the error, so a laterrspechit a broken test database instead. Reported by a worker. -
Fixed: on the Palolo toolkit, the eslint check no longer runs out of memory on the largest packages. The check now runs with a larger Node heap, and two server specs that fail intermittently on an unmodified tree are listed as known baseline failures, so the grader does not hold them against your agent.
-
Fixed: on macOS,
snapshot-to-taskno longer fails withEACCESwhile copying the snapshot's session folder. It used to die before writingtask.tomlandinstruction.mdwhen the toolkit folder was bind-mounted into the Authoring container. -
Task images now fail to build when a dependency install fails. A failed
pnpm installoryarn installused to print a warning and leave the image with missingnode_modules, so every trial ran against a broken workspace. The build now stops so you see the problem when the image is built. -
Fixed: the message printed when rubric-mode grading runs without staged files now names the kit's staging script,
npx tsx scripts/stage-atomic-rubric.ts <task-slug>. -
submit-tasknow counts only reference runs that finished cleanly toward the four it asks for. A run cut short by an API error, a non-zero agent exit or the agent timeout never finished its turn, so it doesn't show what the agent would have done: if you ship four or more runs and fewer than four of them are clean, packaging stops and asks you to re-run the failed trials. Fewer than four runs in total is still just a warning, and a verifier-side timeout still counts as clean. -
harbor-runnow names the missing file when your task directory is incomplete. A task withouttests/test.sh,instruction.mdor a parseabletask.tomlused to fail with Harbor'sEither datasets or tasks must be provided., which named neither the path nor the file; the run now stops up front and tells you which one to restore fromharbor-tasks/_task-scaffold/. -
Fixed:
run-app potion-webnow comes up with a rendered page. The app reads four environment variables at boot that it has no committed env file to supply, so the client bundle threw on the first undefined one and the page stayed blank; the container now supplies dummy values for them.
1be774e26e
-
The toolkit ships one grading standard. Every trial grades under the Grading Standard: eight criteria (Integrity, Narrow Correctness, Broader Correctness / craft, Persistence, Communication, Verification & Thoroughness, Common Sense, Thought Partnership) that produce one score. The reward is the mean of the non-N/A criteria, minus any heavy penalties your guidance directs at the overall score, floored at 0.0. The full standard ships at
task-shared/grading-standard.mdand is embedded in the grader system prompt.- Grader assets keep their
-consolidatedfilenames. A new task scaffoldstests/grader-system-prompt-consolidated.md,tests/render-grade-consolidated.py, and one guidance file,tests/grader-guidance-consolidated.md— the same filenames on every toolkit version, so moving between toolkits never means renaming files. Author the guidance with the grader-guidance skill (/write-grader-guidance-consolidatedin claude,$write-grader-guidance-consolidatedin codex), and phrase any heavy penalty qualitatively ("apply a heavy penalty to<criterion>"). The detector self-check skills assess the same file. verifier/reward-correctness.txtreadsN/Aon every trial. Correctness is scored inside the criteria (Narrow Correctness, Broader Correctness), not as a separate score.submit-taskreads theN/Aas expected and prints its reward summary underScore distribution.- A submission started on an earlier toolkit version is completed on that version. A task keeps the grader assets it was created with, and you finish and submit it on the toolkit you started it with. Start every new task on this toolkit.
- Grader assets keep their
-
/detector-credential-leakagenow reports credentials, not authoring cruft. It used to also flag things like.raccoon-setup-doneor patch content it judged unrelated to the task, so a 0-byte marker file could come back as a blocking leak; those are out of scope now. It still flags an absolute path from your own machine into your checkout (/home/you/…/worker-toolkit-x/repo/…) if your patch adds one. -
Fixed: the session a snapshot task resumes no longer carries your own machine's paths.
snapshot-to-tasknow rewrites your checkout path to the trial's/workspace, so the agent under test reads a working directory that matches where it is actually running instead of a directory from your laptop that does not exist in the trial. -
Fixed: files under a directory whose name contains an emoji or other non-ASCII character now reach the grader. On zeta-dbt (
models/🥇/,🥈,🥉) the verifier silently dropped every such file when collecting your agent's changes, so work in those directories could be graded as if it had never happened;check-workspace-syncnow prints those paths readably too. -
zeta-platform and zeta-wasabi-platform now open at an earlier commit where the app is fully wired up. Several integrations used to be disabled in the code, so a task touching one of them couldn't be exercised at all. On zeta-platform this also revives 41 specs the old skip-list had to skip; the remaining skips moved to
spec/support/known_failing_specs.rb. -
Fixed: creating a task from a snapshot no longer fails with "No user text turn found in session" / "Could not extract instruction" when your explore session has compacted (the "This session is being continued from a previous conversation…" turn). Re-running
snapshot-to-taskon an affected snapshot now fills ininstruction.mdand the seeded session normally. -
New:
scripts/harbor-run <task> --fastruns the trial agent with Claude's fast mode — same model, toolset, and grading, just faster output, so trial turnaround drops. Claude-only: other harnesses refuse the flag. -
Fixed:
/fastin the Explore and Authoring containers' interactiveclaudeno longer reports "unavailable due to network connectivity issues" — it now toggles normally. Fast mode stays off until you turn it on, per container. -
Fixed:
submit-taskno longer warns that a reference run "ran an unregistered agent". It fired once per run — most often after you re-graded a run more than once — for something only we can fix, and it counted toward the warning total without being printed, so the total didn't match what was on screen. -
submit-tasknow lists every warning it counts in its packaging summary, so the total always matches what you can read. -
harbor-runandsubmit-tasknow tell you when a toolkit script underscripts/has been edited, the way they already do for a task'senvironment/Dockerfileandtests/test.sh. Nothing blocks; scripts you add yourself are never reported. -
Fixed: potion-app now builds on a case-sensitive filesystem.
plugins/clientTheme.jsimportedcomponents/PotionBottle.jswhile the file on disk waspotionBottle.js, so webpack failed and no page mounted at all — on Linux, where a case-only difference is a different file. The same mismatch is fixed inpotion-custom-domain-appand the two dynamic-screen-recording members. -
potion-polyglot: the estate's own deployed hostnames now dead-end at localhost in the Explore container. Booting
potion-appby hand with a non-localPOTION_APP_ENVaimed the browser — login form included — at a live host, so anything typed into the app left the container; now nothing does. -
potion-polyglot caveat: several members' Dockerfiles fetch ffmpeg binaries and an ML model from the source company's S3 buckets. Nothing in the toolkit runs those fetches — read them as deployment history rather than steps to reproduce.
-
Fixed: five swingbell-polyglot members no longer serve unstyled. An anonymization pass in the source had replaced the CSS keyword
sansthroughout, including atailwind.config.jskey — so loading the config failed, Tailwind never compiled, and the app came up with no styling and nothing on the page to say why.patient-care,on-boarding-ui,on-boarding-ui-ssr,book-my-minutes-app-expertappointmentandbook-my-minutes-onboardingare all fixed.
136d19f82
- Fixed:
repo/no longer opens with changes you didn't make. Symlinks in the source repo were being unpacked as ordinary files, sogit statusshowed them as modified or deleted from the moment you downloaded the toolkit — and a snapshot taken afterwards carried them into its patch. - Heavy penalties in
tests/grader-guidance-consolidated.mdare now phrased qualitatively — write "apply a heavy penalty to<criterion>" instead of a numeric subtraction like "subtract roughly 0.40"; the grader sizes the deduction itself. The/write-grader-guidance-consolidatedskill, the task scaffold, and the grader prompt are updated to match; existing docs with numeric magnitudes still grade as written. - New: a task can give the agent under test a real browser — set
browser = trueunder[metadata]intask.tomland its trial gets Playwright with Chromium, driven bypw <script.js>. On claude it also enables theReadtool, so the agent can view a screenshot it takes; codex needs nothing extra, since it already views images with its own tool. - Leave
browseroff (the default) and the trial has no browser at all, which is what you want when the point of the task is that something can't be verified. Every new task starts withbrowser = false, whether you build it from a snapshot or by hand. - The Explore container always has the browser, whether or not your task opts in. Start your session with
RACCOON_BROWSER_TASK=1 claudeto explore under the same toolset abrowser = truetask runs. On codex the toolset is the same either way, so the flag is only for claude. - Fixed: on a multi-repo toolkit,
run-app <member>no longer ends in "didn't come up in time" after you rebuild the Explore container or start a second one against the same toolkit folder. A member's dependencies are now tracked per container, so a new container reinstalls what it is missing instead of assuming an earlier one's setup carried over. - Fixed: on the palolo-031 toolkit, creating the Explore container no longer prints a
PrismaClientKnownRequestError/P2028("Unable to start a transaction in the given time") partway through seeding the dev database. The seed now builds a smaller set of members — every organization it created before is still there, the largest capped at 10 members per status instead of 200 — so it stays inside the database connection pool on a machine with few cores, finishes the perk activation it used to die before reaching, and completes noticeably faster. Log in exactly as before (zaniyah@exhalefi.com/test). - Fixed: on the stocks-in-the-future, endsideout, and community-foundation toolkits,
run-appno longer serves the app with its styling missing — oversized images, no page layout. These apps compile their CSS with Tailwind, which the Explore container now builds when it is created. - Fixed: write-only files (
--w-------) a trial leaves behind no longer need a manualchmod.copy-reference-runnow repairs the trial directory before reading it, so the copy no longer dies withEACCESand such a file can no longer reach your task directory, where it made every later run abort at startup with aPermissionError. Packaging repairs the task directory up front too, so the tarball has nothing unreadable in it.RACCOON_SKIP_PERMISSION_REPAIR=1turns all of this off.
Earlier releases predate the Grading Standard.