Files
2026-10-04 21:19:23 -04:00

38 KiB

Changelog

1.0.0

  • New tasks are graded with GPT-6 Sol / high / one sample.
  • Local trials automatically rebuild cached task images whose Codex installation is too old for grading.
  • Fixed: a partly earned extra credit no longer lowers an atomic score. A partial verdict on an extra-credit criterion counts as a pass at half weight, so extra credit only ever raises the score, as /write-atomic-rubric describes. Before, a partial counted as half a point at full weight, which pulled down any run already scoring above 0.5. The effect on stored scores is small, so there is no need to regrade for this.
  • Packaging now warns when a detector report names a reference run that is no longer in your task. This happens after a holistic regrade adopted with --replace, which renames the run folder; submit-task.ts names each report to re-run.

157fc82e20

  • Fixed: on the casa toolkit, a re-grade no longer intermittently reports PG::UndefinedTable or NoEnvironmentInSchemaError failures. The database schema is now loaded when the task image is built rather than at container start, so the grader's setup no longer races it; a task you created on an earlier release keeps its own environment/Dockerfile. Reported by a worker.
  • Fixed: on the zeta toolkits, the grader's rspec check no longer intermittently fails with PG::UndefinedTable on zeta-platform, zeta-heimdall, zeta-hook, zeta-mule-deprecated and zeta-px-api. The grader now waits for the container's startup schema load to finish before preparing the test database; a task you created on an earlier release keeps its own tests/test-commands.sh. Reported by a worker.
  • Fixed: on the potion-polyglot toolkit, the browser-extensions task image no longer rewrites chrome/recorder/yarn.lock into a file yarn can't read, and webpack 4 now builds on its Node 18 base. A task you created on an earlier release keeps its own environment/Dockerfile. Reported by a worker.
  • On the potion-polyglot toolkit, ffmpeg is now installed in the Explore container and in the potion-app, potion-api and potion-video-processing task images. Recording, trimming, GIF and thumbnail steps in those apps now produce real output instead of silently returning nothing; AI voice and face training still cannot run offline. A task you created on an earlier release keeps its own environment/Dockerfile.
  • write-atomic-rubric now allows a scope note naming a sibling criterion, matching the docs and detector-rubric-form; a criterion's verdict still must not depend on another's.
  • submit-task.ts now refuses a workspace.patch of 49 MB or more, and warns from 10 MB and for any binary file over 1 MB in the patch; use a tiny stand-in fixture instead of real model weights or datasets.
  • Fixed: harbor-regrade now reports a run as FAILED when its trial produced no grade, such as when the image build fails, rather than ok.
  • /create-snapshot now warns when snapshot.patch includes the agent's edits from the captured turn. This happens when no turn checkpoint was recorded; check the patch and strip those edits before converting.
  • Fixed: on the freeitsm toolkit, the app can now write ticket attachments, asset imports and document uploads. Those four folders came up owned by root while Apache runs as www-data, so storing an attachment failed with a permission error rather than saving the file. Reported by a worker.
  • New toolkit: frepple, open-source supply chain planning — demand forecasting, production planning, material and capacity constraints, distribution and inventory. Three languages in one codebase: a C++ planning engine, a Django web app and a Vue frontend. run-app starts it at the port the welcome banner prints; sign in as admin / frepple, and a planned demo dataset is already loaded, so the forecasting and planning screens have data in them.
  • Fixed: codex no longer fails to authenticate when your .env has Windows (CRLF) line endings. codex now reads its key from your environment rather than a file, and the stray carriage return was reaching it again; every launcher cleans the key before starting an agent.
  • On the speedwell-polyglot toolkit, run-app strongsuit-app now also starts the Phoenix backend the app calls on port 4201, and seeds the recommendations its personalization pages read. Without the backend the member and MSS home pages and both important-date-recommendation pages threw instead of rendering, and the URL run-app prints is now the dev-login route that actually signs you in rather than one that redirects to a login you cannot reach offline.

0e2d5cce66

  • New toolkit: freeitsm, an IT service desk (tickets, CMDB, change management, assets, contracts, self-service) in plain PHP with no framework. run-app boots it at the port the welcome banner prints; sign in as admin / freeitsm123, and demo data is already seeded for every module except the LMS.
  • On the freeitsm toolkit, git status is now clean when the container comes up. Setup used to overwrite a tracked config.php, so every worker started with a modified file they had not touched.
  • Fixed: codex now works in the Authoring container. It authenticated but every shell command it tried failed with a sandbox error, because the alias was missing the two bypass flags the Explore launcher passes.
  • Fixed: harbor-regrade --all no longer re-uses the job directories from your last round. A second --all on the same task produced the same directory names, so harbor refused to write into them and the previous round's grades stayed where they were, reading as fresh ones; each round now gets its own directories. Reported by a worker.
  • The holistic rubric's title is the task's slug without your worker-id prefix. The /write-holistic-rubric template says so, and a task built from a snapshot now fills the title in for you; a submitted title that still carries the prefix is not a defect.
  • Fixed: on the potion-polyglot toolkit, run-app potion-api now starts the app. It exited on a missing Segment write key before binding a port, so run-app reported that the app did not come up in time; setting the member up now writes its own .env.local with local-only placeholder values. Sign-up and the other unauthenticated routes work; logging in does not, because the app will not issue a token until an account is verified by email and that mail cannot be delivered offline. Reported by a worker.

4ead3c97eb

  • An atomic rubric carries as many criteria as its holistic rubric needs. The /write-atomic-rubric skill, the /detector-rubric-form self-check and the review validator no longer bound the criteria count; write one criterion per scoring-relevant rule and let the count follow the source.
  • A task built from a snapshot now scaffolds the full holistic-rubric template. It used to hand you an empty file with none of the section headings and a stale instruction to write penalties as numeric subtractions; you now get the same template the manual scaffold uses, with the current qualitative penalty phrasing.
  • Fixed: writing your atomic rubric no longer marks every earlier detector report stale. Only /detector-rubric-coverage and /detector-rubric-form read that file, so the other fifteen reports stay fresh, and the staleness warning now names each report against the input it actually reads.
  • New tasks now grade once rather than averaging three samples, so a run finishes sooner and its reward moves around more between runs. Set GRADER_SAMPLES=3 — a prefix on harbor-run/harbor-regrade, or a line in .env — to average again; tasks built on an earlier release keep the sample count they were created with.
  • harbor-regrade can now re-grade all your reference runs in one go. Pass --all instead of a single run directory and it works through every run you have captured, 2 at a time (--jobs N to change), which is the usual thing to want after editing your rubric.
  • A regrade now files itself, and --replace adopts one that would overwrite an existing grade. An atomic regrade lands in your task's rubric-regrades/<run>/ with no copying or naming on your part. When a grade is already filed for that run — always the case for a holistic regrade, since the run has one — the result is left in its job directory so you can compare it first, and --replace writes it back.
  • Fixed: your task images no longer report their own setup as changes your agent made. The image finished preparing dependencies after taking its baseline snapshot, so files like Gemfile.lock and poetry.lock reached the grader as part of your agent's work in every run — including runs where the agent changed nothing at all.
  • The grader now sees when a task's setup step failed before its checks ran. Some toolkits run a preparation command, such as a dependency install or a code-generation step, before the test, lint and type-check commands in tests/test-commands.sh. When that command failed, the failure reached only the verifier log, so the grader could not tell whether a red check followed from it or from your agent's work, and a holistic rubric had to hedge with an instruction like "if code generation failed, treat the type-check failures as environmental". The failed command, its exit code and the end of its output now appear with the check results the grader reads, together with a note that failures which follow from it are not the response's doing. Nothing is added when setup succeeds. A task you created on an earlier release keeps its own tests/test.sh, so this reaches the tasks you create on this toolkit.
  • A new task's task.toml now says to leave allow_internet = true as it is. Setting it false, or adding network_mode or allowed_hosts, cuts the grader off from the model API and the trial comes back with no reward.
  • Fixed: re-running run-app after a failed setup now works on the Python members. A failed first attempt left a half-built environment that every later run refused to replace, so the member stayed broken until you rebuilt the container. Setup also installs pkg_resources now, which some older Python libraries import at start-up.
  • A Python member whose dependencies cannot install no longer fails the whole run-app. You get a line saying that member is explore-only in this container, and how to retry it, instead of a wall of installer errors and a bare "setup failed".
  • run-app with no repo now says which member it picked. On a polyglot toolkit it quietly started the default member; it now names it and lists the others. The README also uses the run-app <repo> form and no longer describes a single-repo layout.
  • run-app now points you at the error on screen when setup fails. It used to send you to run-app --logs, which is empty at that point because the app has not started yet.
  • On the Palolo toolkit, the test suite runs against its own database again, and a bare pnpm test needs --run. The image set CI=true so plain vitest would not sit in watch mode forever, but the app reads that variable in its own code: it was pointing the suite at the main development database and switching off one of the app's environment guards.
  • Fixed: on the ZenBill toolkit, the jest check no longer runs out of memory partway through the suite. Node sizes its heap from the machine's memory rather than from the container's limit, so the run kept growing until the container killed it, and roughly one graded run in ten ended with no frontend test evidence at all. The check now runs with an explicit heap cap, which makes Node collect garbage in time to finish all 134 suites inside the limit. A task you created on an earlier release keeps its own tests/test-commands.sh; the new jest line is in task-shared/test-commands.sh.
  • On the human-essentials toolkit, one flaky spec no longer counts towards your task's test signal. A request spec asserts two names in an order the app does not guarantee, so the suite went red on roughly one run in ten with nothing actually wrong; it is now skipped.
  • On the human-essentials toolkit, a task container's development database now comes up with the app's demo data loaded. It was created empty, so the app had no account you could sign in with and anything reading it saw no organizations, users or requests.
  • On the endsideout toolkit, the minitest check now installs any gems your agent added before it runs. A response that changed the Gemfile left the lockfile needing an install, and the check errored outright instead of running, so a task where adding a gem is the right answer had no test signal at all, and one graded run was marked down for a lockfile the verifier's own setup could not satisfy. The check now runs bundle check first and installs only when the lockfile calls for it, so nothing changes when the lockfile is untouched. A task you created on an earlier release keeps its own tests/test-commands.sh; the new setup line is in task-shared/test-commands.sh.
  • On the stocks-in-the-future toolkit, the minitest check now runs with a single test worker. Parallel workers each restarted the same FactoryBot level sequence, which collided with the grade levels several tests hard-code, so tests could fail with "Level has already been taken" on nothing your agent did. The baseline entry that excused this failure now matches on that error message rather than on a whole test class, so a real failure in ClassroomsControllerTest counts again. A task you created on an earlier release keeps its own tests/test-commands.sh; the new command is in task-shared/test-commands.sh.
  • On the flaredown toolkit, the rspec check now prepares its own test database and skips two specs that fail on their own. The container's start-up script creates the test database after its own readiness wait, so a check that started first found no database and could not create one; the check now runs db:test:prepare, which creates the database when it is missing. Two specs are excluded because they fail intermittently on an unmodified checkout: a food search spec whose two sample records can produce identical search vectors with nothing to break the tie, and a collection retriever spec whose random sample can draw counts that violate its strict inequality. A red rspec result is now the response's doing. A task you created on an earlier release keeps its own tests/test-commands.sh; the new lines are in task-shared/test-commands.sh.
  • On the zeta-polyglot toolkit, a task image for the zeta-wasabi-platform member no longer switches on two integrations the app's own test setup leaves off. The image builds the app's .env from the repository's .env.example, which sets HT_ENABLED to true, so any spec that created an internal account sent a request to a service that was not running and raised an unhandled request error. The image now sets HT_ENABLED=false and SAM_ONLINE=false, which is how the app's own test suite expects to run. A task you created on an earlier release keeps its own environment/Dockerfile; the updated image is task-shared/Dockerfile.zeta-wasabi-platform.
  • Fixed: on the zeta-polyglot toolkit, setting up a second Python member no longer breaks Poetry. Installing one member's dependencies downgraded a package Poetry itself needs, so every later run-app on a Python member failed with No module named 'packaging.licenses' and only a container rebuild cleared it.
  • Fixed: on the potion-polyglot toolkit, yarn test and npm test on the potion-app member now run the same suites as the graded jest check. The member's package.json test script passed a path-ignore flag on the command line, and jest treats that flag as a replacement for the ignore list in jest.config.js rather than an addition to it, so yarn test ran thirty-one suites the config skips, went red, and took far longer than npx jest, which stayed green. Agents under test hit this, spent turns untangling it, and some abandoned the command. The script no longer passes the flag, and the member is pinned to that commit. A task you created on an earlier release keeps the potion-app commit it was pinned to. Reported by a worker.
  • Fixed: on the potion-polyglot toolkit, seven members' task images now build and come with their dependencies installed. The potion-video-processing-devops image did not build at all, because a Terraform module it uses was not pinned to a version and the module's current release needs a newer provider than the one the repository locks. Six more members built but were allowed to skip a failed dependency install with only a warning, so a task on any of them met a missing-module error at the first import: browser-extensions, lambda-datadog-forwarder, potion-voice-utils, sentence-split-service, microservice-dynamic-screen-recording and potion-voice-dataset. Each now installs what it needs, and a failed install on those members fails the image build so the problem is visible when it happens.
  • Fixed: on the potion-polyglot toolkit, unzipping the toolkit on macOS or Windows no longer stops to ask about duplicate files or quietly leaves eight of them out. The potion-website member tracked four poster and video pairs under two spellings of the same filename that differ only in letter case. A filesystem that ignores case cannot hold both, so unzip either paused mid-extraction to ask which copy to keep, which reads as a hang on an archive this large, or, when told to overwrite, dropped the second copy without saying so. The duplicate entries are removed and the files themselves are unchanged.

7b6b67ea3d

  • Fixed: the breezy-complete and zeta toolkits build their containers again. The Debian release they are built on left long-term support and its package mirror is being retired, so building an Explore container or a task image failed part-way with a "404 Not Found" on a system package; those packages now come from Debian's archive instead.
  • Fixed: on the breezy-complete toolkit, the Explore container now prepares its database reliably. A boot-time cache could corrupt itself while loading one of the app's larger dependencies, which left the database setup failing and the app with nothing to run against; that cache is now off in Explore, as it already was for task images.
  • Fixed: codex no longer fails to authenticate when your .env was saved on Windows. Windows (CRLF) line endings left a stray character on the end of your key and codex was rejected with an API-key error; the key is now cleaned wherever it is read, so your .env needs no change.

fa77be2885

  • Grading no longer fails silently when your task image carries an older Claude Code. The grader model needs Claude Code 2.1.251 or newer. A task image installs Claude Code when it is first built and keeps that copy on later rebuilds, so an image built before that version failed every grade with "does not support this model" and the trial ended with no reward file. harbor-run now checks your task images before a local trial and rebuilds any that are too old, task images verify the version when they build, and the grader stops with a clear message if an old copy still reaches it.
  • Toolkit documents no longer point at files that ship only in our review pipeline. The atomic-rubric skill describes the validation the staging script performs in the toolkit, the fact-check detector names scripts/build-workspace.sh, and the corpus-viewer notes say they apply to zeta toolkits only.

d7edb3d5c1

  • The toolkit's grading documents are now named the holistic rubric and the atomic rubric. The holistic rubric is the per-task grading document the grader reads alongside the shared Grading Standard; earlier releases called it the grader guidance. The atomic rubric is a YAML companion that restates the same requirements as separately judgeable criteria. The content rules for both are unchanged. This release adopts the names, renames the files that new tasks create, and ships rubric grading in the toolkit.

    • New tasks write tests/holistic-rubric.md and tests/atomic-rubric.yaml. A task created on this toolkit scaffolds tests/holistic-rubric.md as its holistic rubric. The atomic rubric package is tests/atomic-rubric.yaml plus tests/grader-context.md, authored after the holistic rubric is final.
    • A task created on an earlier toolkit version keeps its existing filenames and stays fully supported. The filename-stability promise carries forward for every existing task: grading, the detector skills, scripts/harbor-regrade, and submit-task read tests/grader-guidance-consolidated.md, legacy tests/grader-guidance.md, and tests/rubrics.yaml wherever a task carries them, indefinitely, so moving an existing task between toolkit versions still never means renaming files. Never rename a committed task file. Only new tasks use the new names.
    • /write-holistic-rubric replaces /write-grader-guidance-consolidated ($write-holistic-rubric in codex). It is the same authoring skill under the current name, and it now also teaches length discipline: a finished holistic rubric lands near 1,500 words; a 4,000-to-5,000-word draft is repetition, not thoroughness; an edit never grows the document.
    • New: /write-atomic-rubric ($write-atomic-rubric in codex) converts a finished holistic rubric into tests/atomic-rubric.yaml plus tests/grader-context.md. Every task-specific requirement becomes one separately judgeable criterion, and the context and ground truth those criteria rely on are extracted alongside.
    • Rubric grading ships in the toolkit. The rubric renderer (render-rubric-grade.py) is included under task-shared/ and scaffolded into new tasks. Once a task's atomic rubric is written, stage its grading copies with npx tsx scripts/stage-atomic-rubric.ts <task-slug>; scripts/harbor-regrade then re-grades a captured run in rubric mode with no patch. Run the staging script with --restore to remove the staged copies before packaging.
    • Grading runs on claude-fable-5-1. New tasks and freshly staged rubric assets grade with claude-fable-5-1 by default. A task that shipped with an earlier grader keeps that grader unless you override it, so existing scores stay comparable. Override either way with GRADER_MODEL=....
    • Two new detector self-checks: /detector-rubric-coverage and /detector-rubric-form. Coverage checks that your atomic rubric tracks your holistic rubric, so no load-bearing requirement, penalty, or "do not penalize" rule is missing from the criteria and no criterion invents one. Form checks the atomic rubric as an artifact: the criterion schema, atomicity, positive phrasing, and inline answer keys.
  • Fixed: on the stocks-in-the-future toolkit, a re-graded run's minitest check now actually runs the suite. The container used to build its databases at start-up, so a check running soon after could hit a missing stocks_in_the_future_test; both databases now ship inside the image.

  • Fixed: on the zeta toolkits, run-app no longer leaves a .venv behind for the Python members. Dependencies now install into the container's Python, matching the graded image — so if you switch between Python members, re-run run-app for the one you're working on.

  • The note at the top of tests/test-commands.sh no longer tells you not to edit it. Task-specific checks there are expected and kept.

  • Fixed: run-app potion-multi-dsr-watcher now boots. It had no database URL and started a cron job that never opened a port, so run-app timed out waiting for one; it now serves its HTTP entrypoint on port 3000.

  • Codex (gpt-5.6-sol) is now the default agent. A manual task now scaffolds with harness = "codex", and the docs start you in codex; Claude Code remains fully supported, and a task keeps whichever agent authored it.

  • Fixed: re-grading a run where your agent renamed a file with git mv no longer brings the old file back. The verifier recorded the rename as a new file only, so the re-graded workspace held both copies and the stale one broke the type-check or test suite — failures no agent caused.

  • Fixed: a file your agent wrote at a path it had just removed or renamed away no longer disappears when the run is re-graded. The verifier listed that path as deleted even though the new file was sitting there, so the re-graded workspace lost it.

  • Fixed: codex now picks up a rotated ANTHROPIC_API_KEY without a container rebuild. It read its key from a file written when the container was created, so a key changed in .env afterwards left it failing to authenticate; each launch now re-reads .env first (in Explore, from the container's next start). claude was never affected.

  • Fixed: an Explore container that came up with an empty /workspace/repos (or /workspace/repo) now repairs itself on the next up. Unzipping a new toolkit over an old install could leave the container pointed at nothing, so run-app <repo> failed with checkout <sha> failed and rebuilding the container did not help. Reported by a worker.

  • Containers now come up with their database already loaded. On the human-essentials and awbw toolkits the image used to build the database when the container started, so a trial could reach the test database before it was ready. The schema now ships inside the image, which also cuts container start-up time noticeably on awbw.

  • Fixed: on the human-essentials, zeta-platform and flaredown toolkits, a re-graded run's rspec check now actually runs the suite. The check could start before the container had finished loading the test database, in which case rspec aborted at load time and reported zero examples — which read as ordinary test failures. The verifier now waits for the schema before running any check.

  • Fixed: the same on the breezy-complete toolkit, where the container builds its databases for longer. The rspec check could report zero examples, or a missing socratic_systems_test, on a run graded soon after the container started; the databases now ship inside the image.

  • Fixed: the breezy-complete Explore container no longer seeds its database twice. db:prepare already seeds the database it creates, so the second pass aborted partway on a duplicate record; seeding now runs only when the database has none.

  • Fixed: on the awbw toolkit, restarting a container no longer leaves the test database half-loaded. Reloading the schema over an existing one failed on a foreign-key ordering in db/schema.rb (MySQL error 3730), and the container hid the error, so a later rspec hit a broken test database instead. Reported by a worker.

  • Fixed: on the Palolo toolkit, the eslint check no longer runs out of memory on the largest packages. The check now runs with a larger Node heap, and two server specs that fail intermittently on an unmodified tree are listed as known baseline failures, so the grader does not hold them against your agent.

  • Fixed: on macOS, snapshot-to-task no longer fails with EACCES while copying the snapshot's session folder. It used to die before writing task.toml and instruction.md when the toolkit folder was bind-mounted into the Authoring container.

  • Task images now fail to build when a dependency install fails. A failed pnpm install or yarn install used to print a warning and leave the image with missing node_modules, so every trial ran against a broken workspace. The build now stops so you see the problem when the image is built.

  • Fixed: the message printed when rubric-mode grading runs without staged files now names the kit's staging script, npx tsx scripts/stage-atomic-rubric.ts <task-slug>.

  • submit-task now counts only reference runs that finished cleanly toward the four it asks for. A run cut short by an API error, a non-zero agent exit or the agent timeout never finished its turn, so it doesn't show what the agent would have done: if you ship four or more runs and fewer than four of them are clean, packaging stops and asks you to re-run the failed trials. Fewer than four runs in total is still just a warning, and a verifier-side timeout still counts as clean.

  • harbor-run now names the missing file when your task directory is incomplete. A task without tests/test.sh, instruction.md or a parseable task.toml used to fail with Harbor's Either datasets or tasks must be provided., which named neither the path nor the file; the run now stops up front and tells you which one to restore from harbor-tasks/_task-scaffold/.

  • Fixed: run-app potion-web now comes up with a rendered page. The app reads four environment variables at boot that it has no committed env file to supply, so the client bundle threw on the first undefined one and the page stayed blank; the container now supplies dummy values for them.

1be774e26e

  • The toolkit ships one grading standard. Every trial grades under the Grading Standard: eight criteria (Integrity, Narrow Correctness, Broader Correctness / craft, Persistence, Communication, Verification & Thoroughness, Common Sense, Thought Partnership) that produce one score. The reward is the mean of the non-N/A criteria, minus any heavy penalties your guidance directs at the overall score, floored at 0.0. The full standard ships at task-shared/grading-standard.md and is embedded in the grader system prompt.

    • Grader assets keep their -consolidated filenames. A new task scaffolds tests/grader-system-prompt-consolidated.md, tests/render-grade-consolidated.py, and one guidance file, tests/grader-guidance-consolidated.md — the same filenames on every toolkit version, so moving between toolkits never means renaming files. Author the guidance with the grader-guidance skill (/write-grader-guidance-consolidated in claude, $write-grader-guidance-consolidated in codex), and phrase any heavy penalty qualitatively ("apply a heavy penalty to <criterion>"). The detector self-check skills assess the same file.
    • verifier/reward-correctness.txt reads N/A on every trial. Correctness is scored inside the criteria (Narrow Correctness, Broader Correctness), not as a separate score. submit-task reads the N/A as expected and prints its reward summary under Score distribution.
    • A submission started on an earlier toolkit version is completed on that version. A task keeps the grader assets it was created with, and you finish and submit it on the toolkit you started it with. Start every new task on this toolkit.
  • /detector-credential-leakage now reports credentials, not authoring cruft. It used to also flag things like .raccoon-setup-done or patch content it judged unrelated to the task, so a 0-byte marker file could come back as a blocking leak; those are out of scope now. It still flags an absolute path from your own machine into your checkout (/home/you/…/worker-toolkit-x/repo/…) if your patch adds one.

  • Fixed: the session a snapshot task resumes no longer carries your own machine's paths. snapshot-to-task now rewrites your checkout path to the trial's /workspace, so the agent under test reads a working directory that matches where it is actually running instead of a directory from your laptop that does not exist in the trial.

  • Fixed: files under a directory whose name contains an emoji or other non-ASCII character now reach the grader. On zeta-dbt (models/🥇/, 🥈, 🥉) the verifier silently dropped every such file when collecting your agent's changes, so work in those directories could be graded as if it had never happened; check-workspace-sync now prints those paths readably too.

  • zeta-platform and zeta-wasabi-platform now open at an earlier commit where the app is fully wired up. Several integrations used to be disabled in the code, so a task touching one of them couldn't be exercised at all. On zeta-platform this also revives 41 specs the old skip-list had to skip; the remaining skips moved to spec/support/known_failing_specs.rb.

  • Fixed: creating a task from a snapshot no longer fails with "No user text turn found in session" / "Could not extract instruction" when your explore session has compacted (the "This session is being continued from a previous conversation…" turn). Re-running snapshot-to-task on an affected snapshot now fills in instruction.md and the seeded session normally.

  • New: scripts/harbor-run <task> --fast runs the trial agent with Claude's fast mode — same model, toolset, and grading, just faster output, so trial turnaround drops. Claude-only: other harnesses refuse the flag.

  • Fixed: /fast in the Explore and Authoring containers' interactive claude no longer reports "unavailable due to network connectivity issues" — it now toggles normally. Fast mode stays off until you turn it on, per container.

  • Fixed: submit-task no longer warns that a reference run "ran an unregistered agent". It fired once per run — most often after you re-graded a run more than once — for something only we can fix, and it counted toward the warning total without being printed, so the total didn't match what was on screen.

  • submit-task now lists every warning it counts in its packaging summary, so the total always matches what you can read.

  • harbor-run and submit-task now tell you when a toolkit script under scripts/ has been edited, the way they already do for a task's environment/Dockerfile and tests/test.sh. Nothing blocks; scripts you add yourself are never reported.

  • Fixed: potion-app now builds on a case-sensitive filesystem. plugins/clientTheme.js imported components/PotionBottle.js while the file on disk was potionBottle.js, so webpack failed and no page mounted at all — on Linux, where a case-only difference is a different file. The same mismatch is fixed in potion-custom-domain-app and the two dynamic-screen-recording members.

  • potion-polyglot: the estate's own deployed hostnames now dead-end at localhost in the Explore container. Booting potion-app by hand with a non-local POTION_APP_ENV aimed the browser — login form included — at a live host, so anything typed into the app left the container; now nothing does.

  • potion-polyglot caveat: several members' Dockerfiles fetch ffmpeg binaries and an ML model from the source company's S3 buckets. Nothing in the toolkit runs those fetches — read them as deployment history rather than steps to reproduce.

  • Fixed: five swingbell-polyglot members no longer serve unstyled. An anonymization pass in the source had replaced the CSS keyword sans throughout, including a tailwind.config.js key — so loading the config failed, Tailwind never compiled, and the app came up with no styling and nothing on the page to say why. patient-care, on-boarding-ui, on-boarding-ui-ssr, book-my-minutes-app-expertappointment and book-my-minutes-onboarding are all fixed.

136d19f82

  • Fixed: repo/ no longer opens with changes you didn't make. Symlinks in the source repo were being unpacked as ordinary files, so git status showed them as modified or deleted from the moment you downloaded the toolkit — and a snapshot taken afterwards carried them into its patch.
  • Heavy penalties in tests/grader-guidance-consolidated.md are now phrased qualitatively — write "apply a heavy penalty to <criterion>" instead of a numeric subtraction like "subtract roughly 0.40"; the grader sizes the deduction itself. The /write-grader-guidance-consolidated skill, the task scaffold, and the grader prompt are updated to match; existing docs with numeric magnitudes still grade as written.
  • New: a task can give the agent under test a real browser — set browser = true under [metadata] in task.toml and its trial gets Playwright with Chromium, driven by pw <script.js>. On claude it also enables the Read tool, so the agent can view a screenshot it takes; codex needs nothing extra, since it already views images with its own tool.
  • Leave browser off (the default) and the trial has no browser at all, which is what you want when the point of the task is that something can't be verified. Every new task starts with browser = false, whether you build it from a snapshot or by hand.
  • The Explore container always has the browser, whether or not your task opts in. Start your session with RACCOON_BROWSER_TASK=1 claude to explore under the same toolset a browser = true task runs. On codex the toolset is the same either way, so the flag is only for claude.
  • Fixed: on a multi-repo toolkit, run-app <member> no longer ends in "didn't come up in time" after you rebuild the Explore container or start a second one against the same toolkit folder. A member's dependencies are now tracked per container, so a new container reinstalls what it is missing instead of assuming an earlier one's setup carried over.
  • Fixed: on the palolo-031 toolkit, creating the Explore container no longer prints a PrismaClientKnownRequestError / P2028 ("Unable to start a transaction in the given time") partway through seeding the dev database. The seed now builds a smaller set of members — every organization it created before is still there, the largest capped at 10 members per status instead of 200 — so it stays inside the database connection pool on a machine with few cores, finishes the perk activation it used to die before reaching, and completes noticeably faster. Log in exactly as before (zaniyah@exhalefi.com / test).
  • Fixed: on the stocks-in-the-future, endsideout, and community-foundation toolkits, run-app no longer serves the app with its styling missing — oversized images, no page layout. These apps compile their CSS with Tailwind, which the Explore container now builds when it is created.
  • Fixed: write-only files (--w-------) a trial leaves behind no longer need a manual chmod. copy-reference-run now repairs the trial directory before reading it, so the copy no longer dies with EACCES and such a file can no longer reach your task directory, where it made every later run abort at startup with a PermissionError. Packaging repairs the task directory up front too, so the tarball has nothing unreadable in it. RACCOON_SKIP_PERMISSION_REPAIR=1 turns all of this off.

Earlier releases predate the Grading Standard.