Files
project-work/worker-toolkit-stocks-in-the-future/CHANGELOG.md

42 KiB
Raw Blame History

Changelog

1f3264fa7

  • The detector self-check skills now assess the guidance file the grader actually uses, on any task shape. A task can carry both tests/grader-guidance-consolidated.md and the legacy tests/grader-guidance.md; bash scripts/guidance-target.sh <slug> prints the file the grader reads and the standard it grades under, and every detector reads that file, judges it against its own standard's structure, and opens its report by naming what it assessed.
  • After your snapshot is converted into a task, the next-steps message names tests/grader-guidance-consolidated.md as the rubric file to write.

f9a5e5c51

  • You can now author with Codex instead of Claude Code. Both CLIs are installed and pre-configured in the Explore and Authoring containers — start either one with no arguments (claude or codex) and the model, reasoning effort and tool set are already set for you.

    • Pick one and use it for the whole task, in both containers. The task records which agent authored it and every trial replays on that same agent, so exploring in one and snapshotting in the other measures the wrong thing.
    • The snapshot command differs: /create-snapshot:snapshot in claude, $snapshot in codex. Same for the authoring skills — /write-grader-guidance in claude, $write-grader-guidance in codex.
    • Codex has no conversation rewind. In claude you can undo back a few turns and then snapshot; in codex, snapshot as soon as you notice the behavior you want to capture.
    • The toolkit's project instructions now ship twice: CLAUDE.md (which claude reads) and AGENTS.md (which codex reads). AGENTS.md is generated from CLAUDE.md — read either, edit neither.
  • A new repository estate is available for task authoring: clockwise-polyglot (Clockwise — AI calendar / smart-scheduling). It's a polyglot toolkit — one download hosts the whole estate and you switch between members with run-app <repo>. You can author tasks against every repo in the kit. Seven members ship real, validated test suites: rust-defrag (Cargo workspace, 232+ tests via cargo test --workspace — Rust is newly supported by run-app first-use setup), cql-v2 (custom scheduling-DSL interpreter, 59 pytest tests), preference-schedule-demo (FastAPI scheduler, 27 pytest tests under server/), defrag-powers-of-ten (Vite calendar-defrag visualization, 13 vitest tests — also the default run-app member, so it boots as a dev server out of the box), dbt (17 pytest tests for its event-classification scripts; the dbt models themselves target an external warehouse), python-analysis (794 pytest tests under services_client/ — run them from that directory), and interpreter_interview (build first: npm run build && npm test). The rest of the estate — an MJML email template set, an MCP server, a Slack deploy bot, a Next.js AI chat app, prototypes, notebooks and infra — you read and edit rather than run test suites against; tasks there are rubric-scored like any other member. Three members have suites or builds blocked by private npm packages that no longer exist in any registry (webserver, calendar-classic, and mcp-server's test run) — their dependencies cannot install, so treat them as read-and-edit.

  • A new repository estate is available for task authoring: swingbell-polyglot (SwingBell / Prescrable — a digital-health EHR and telemedicine stack: Java/Spring services, a parallel .NET migration, and React/Next frontends). It's a polyglot toolkit — one download hosts the whole estate and you switch between members with run-app <repo>. You can author tasks against every repo in the kit. A heads-up on verification style: jwt-encryption-decryption is the one member with a real green inherited suite (mvn test); everywhere else the generated context-loads skeletons are deliberately @Disabled (their Spring contexts need live Postgres/Kafka that the offline environment doesn't run — the annotation on each says so), so mvn test passes cleanly estate-wide and tasks are rubric-scored or verified with build checks / suites you author. The React frontends boot via run-app (ABDM-FE is the default member); the Java services build with their Maven wrappers — they depend on sibling shared libraries — build common-repository, common-aws-service, and reports first (mvn -DskipTests install in each) before compiling a service like hip; graded-task images do this for you automatically. The .NET twins build with dotnet build.

  • Grading now runs under the Consolidated Grading Standard by default. The grader scores eight criteria — Integrity, Narrow Correctness, Broader Correctness / craft, Persistence, Communication, Verification & Thoroughness, Common Sense, Thought Partnership — against tests/grader-guidance-consolidated.md, and the reward is the mean of the non-N/A criteria minus any heavy penalties your guidance defines, floored at 0.0. The full standard ships at task-shared/grading-standard.md and is embedded in the new tests/grader-system-prompt-consolidated.md. Under this standard verifier/reward-correctness.txt reads N/A — correctness lives inside the criteria, not as a separate score. To grade the way earlier releases did (seven behavioral dimensions plus the separate correctness score, against tests/grader-guidance.md), add --verifier-env GRADING_STANDARD=legacy to your harbor-run or harbor-regrade command (harbor-regrade also accepts HARBOR_GRADING_STANDARD=legacy as an environment variable).

    • A task created from an earlier toolkit doesn't have the consolidated files in its tests/; grading says so and continues under the legacy standard. To grade it consolidated, copy the current shared assets in: cp task-shared/test.sh task-shared/grader-system-prompt*.md task-shared/render-grade*.py harbor-tasks/<slug>/tests/.
    • New tasks scaffold tests/grader-guidance-consolidated.md as the graded guidance file — author it with /write-grader-guidance-consolidated, with penalty magnitudes as 0.0–1.0 fractions. The legacy tests/grader-guidance.md is still scaffolded, and the detector self-check skills still assess it.
    • submit-task.ts accepts a task carrying either guidance file (or both), and the toolkit-managed-files check now also covers tests/grader-system-prompt-consolidated.md.
  • A new repository estate is available for task authoring: potion-polyglot (Potion — AI personalised-video / avatar-generation SaaS). One download hosts 52 repos and you switch between them with run-app <repo>; you can author tasks against every one of them. The default, potion-app (Nuxt 2 + Express), boots with run-app potion-app — it builds, seeds a verified local user and starts the server, so you can log in at /auth/login with dev@example.com / devpassword123 (real sign-in goes through Google/LinkedIn, which no offline container can reach). potion-web (Nuxt 3) boots with no setup. MongoDB and Postgres both run in the container.

    • potion-app's jest suite is green at the pin: 13 suites / 27 tests; suites that never passed on a clean checkout are skipped in jest.config.js with their reasons, so red there means something regressed. Most other members have no inherited tests — normal here: you author the verifier with the task, and every member ships its own harbor Dockerfile. potion-qa (Selenium) and potion-snapshot-testing (Playwright) target a production app that no longer exists, so read them rather than run them.
    • potion-wp-site: composer lint:php is a real check and green at the pin. Its committed lint:wpcs ruleset is not wired — the theme has ~114 pre-existing violations, so red there says nothing about your change.
    • potion-custom-domain-app: dependencies do not install (tree from ~2013, engines pinned to Node 0.8). Read-and-edit substrate; relax them in the task if you need them.
  • flaredown: the source repo now tracks a newer upstream master. Upstream's own Docker setup was broken at the previous pin (frontend/Dockerfile forced npm 7 while the client's package.json requires npm 6 with engine-strict, so the image couldn't build) and is fixed at the new one; upstream also swapped the Ember test browser from PhantomJS to headless Chrome and added a root CLAUDE.md. That CLAUDE.md describes running the app with make + docker compose — that's the upstream workflow, for your own machine. There's no Docker daemon inside the Explore container, so in there keep using run-app and cd backend && bundle exec rspec; the dependencies are already installed for you.

  • Fixed: first-use setup for a Node repo could fail on a dependency's postinstall (Cannot find module '/opt/raccoon-node-modules/<repo>/package.json') — re-run run-app <repo> on the new toolkit.

  • The reference-data corpus is now searchable: corpus-shipping toolkits add a local viewer that starts with the Explore container (the banner prints the URL; manage it with view-corpus) giving full-text search, per-person activity, ticket ↔ chat cross-references and timeline views across every source; the index behind it is also queryable with sqlite3 (see explore/corpus-viewer/README.md).

  • Graded trials are unchanged — only data/zeta-corpus/ is staged into trial images, so the agent under test still explores it with grep/find — but a container created from an older toolkit needs recreating (npx @devcontainers/cli up --remove-existing-container) before the viewer is reachable.

  • The LLM grader's guidance was updated (internal calibration); nothing changes about how you run tasks or read grades.

  • Runs are graded by new internal infrastructure that adds a machine-readable verifier/grade.json, plus a tests/render-grade.py shipped in the scaffold alongside test.sh.

  • If a run finishes with no reward and an error about only some samples being usable, that isn't a problem with your task — just re-run (if it repeats, verifier/render-stderr-*.log and verifier/grader-stderr-*.log show what happened).

  • Fixed: /create-snapshot could write a workspace.patch that wouldn't apply (cannot apply binary patch to '<file>' without full index line) whenever a binary file differed from the repo's committed state, often a stray committed .DS_Store; for a task that's already stuck, drop that file's section from the patch and re-run bash scripts/build-workspace.sh <slug>.

  • New estate available for task authoring: speedwell-polyglot (StrongSuit / Speedwell — member-coaching / wellness apps), a polyglot toolkit where one download hosts every member, you switch with run-app <repo>, and you can author tasks against all of them.

  • Four members boot a dev server — strongsuit-app (Remix + Prisma, vitest), strongsuit-client (razzle, jest), strongsuit-server (Express, jest), speedwell-react (CRA, whose one smoke test is skipped for a pre-existing Jest/ESM issue, so npm run build is the check there) — and the rest you read and edit in place, including strongsuit_phx (mix test), jobs (go test ./...) and terraform (validate + fmt -check).

  • The estate's remaining ten members (two Chrome extensions and an assistant plugin, Zingle/Typeform→Airtable scripts, a Firebase agenda builder, a Node sync service, a Botkit bot, the shared HTML templates, App Engine Python scripts) ship no test suite, so tasks against them are rubric-scored and packaged like any other member.

  • Fixed: run-app zeta-plastic reported the app was up with every database empty; setup now creates and loads both databases for development and test (delete repos/zeta-plastic/.raccoon-setup-done and re-run to pick this up).

  • run-app no longer leaves a stray .raccoon-setup-done showing in git status after a polyglot member's first-use setup, and snapshot patches no longer pick it up.

  • Polyglot run-app strongsuit-app now generates the app's Tailwind CSS before starting the dev server, which previously stopped the Remix build resolving ./styles/tailwind.css.

  • Fixed: submit-task.ts no longer dies with Cannot stat: Permission denied while building the tarball, since copy-reference-run hands each captured run back to you and submit-task.ts repairs and retries if it meets root-owned agent-output/ files anyway.

  • New self-check: /detector-offline-verifiability — flags a task whose success criteria can only be verified on live external systems (a deploy, an external service migration) rather than from inside the no-network sandbox. Advisory like the other detectors; run it before submitting if your task touches external systems.

  • Fixed: heavyweight deterministic checks (eslint-class) could be killed by the sandbox's memory limit mid-run and surface as SIGNALS_DEGRADED; they now run under a capped Node heap and complete. Separately, a repo with no runnable checks at all is normal — the grader scores correctness by reading the code, and /write-grader-guidance now says so instead of expecting test-backed signals everywhere.

  • The scaffolded tests/grader-guidance.md now matches the guidance structure in the project instructions — Business context, what a strong and a weak response look like, Ground truth (absorbing the old "Privileged information" bullets), optional Supporting evidence, and optional Heavy penalties with the subtraction-not-cap arithmetic — so a doc started from the scaffold no longer needs restructuring against the format spec. The snapshot flow's generated guidance file lists the same sections.

4fe90f3aa

  • Three files in every task are now checked: environment/Dockerfile, tests/test.sh, and tests/grader-system-prompt.md. These ship from task-shared/ and are meant to be identical across every task — they build the container your trials run in and produce the grade. When one differs, this task's runs are hard to compare with everyone else's, and the scores still look perfectly normal, so nothing downstream notices.

    • scripts/harbor-run, bash scripts/build-workspace.sh and npx tsx scripts/submit-task.ts all report what they find and print the cp that restores the shipped copy. None of them blocks you — you can run trials and submit either way. We'd just rather know.
    • The check compares against a checksum recorded when your task was created, so what it reports is a change made since — not the ordinary fact that these files are updated between releases. A file identical to a copy this toolkit ships is always fine, whichever release it came from.
    • If your task predates the current release it may report that its copies are older than what ships now. That isn't a mistake; it does mean the task was graded with older versions than a task built today, so restoring the current copies and re-running your trials is what lines the scores up.
    • The two appends the toolkit makes itself — snapshot session staging and the reference-data corpus — are recognized and never count.
    • If you need something the environment doesn't have (a missing package, a grader that won't run), please tell us rather than editing locally: the fix has to work for every task built from this toolkit, and a local patch usually means the same problem is hitting other authors silently.
  • The grader now produces two independent scores: behavioral and correctness. Previously there was one number — the behavioral score, the mean of the seven Behavioral Rating Dimensions. That number is unchanged and still lands in verifier/reward.txt. Alongside it, the grader now writes a separate correctness score to verifier/reward-correctness.txt: is the deliverable the agent produced actually right? For code, does it work and is it well-built (craft counts, but only as a secondary term — even an egregious craft flaw is worth a small deduction, and sloppiness alone never drives correctness toward 0). For a written review or diagnosis, are its substantive technical claims true of the codebase. Both appear in verifier/reward.json and at the tail of verifier/test-stdout.txt, and the correctness reasoning gets its own ## Correctness section in grade.md.

    • The two axes never bleed into each other, and keeping them apart is the thing to watch for when you write grader guidance. Whether it was behaviorally right to produce the deliverable at all — to defer, ask, push back, or narrow the scope — is a behavioral question. Correctness asks only whether the deliverable that does exist is right. A clean, working implementation of a decision you'd have made differently is HIGH correctness and a Scoping problem. A behaviorally excellent run can still ship broken code. Both are the split working as intended.
    • Correctness is N/A, not 0, when there's nothing substantive to check — the agent only asked a clarifying question, or declined without asserting facts. For a task built around an assessment, a diagnosis, or a pushback, N/A across every reference run is the expected result, not a problem.
    • What this means for authoring: /write-grader-guidance now elicits the task-specific correctness signal alongside the behavioral privileged information — what a working deliverable has to do, which checks bear on it, and where a green test suite does not prove completeness. The task scaffold's grader-guidance.md has a matching optional ## Correctness section. You usually write nothing about code craft; the shared prompt handles it.
    • submit-task.ts now reports the correctness distribution next to the behavioral one, and warns if correctness came back N/A on every run (harmless for assessment-shaped tasks; worth investigating if your task does ask for a deliverable).
    • scripts/harbor-regrade re-derives both scores, so it's the right tool for iterating on correctness guidance too. See /regrade-reference-run.
  • You no longer set category in task.toml — the field is gone from the _task-scaffold template, and there's nothing to fill in in its place.

  • The Harbor task runner now runs your task on Claude Opus 5 with a 1M-token context window at maximum reasoning effort (claude-opus-5[1m], reasoning_effort=max), bumped from Opus 4.8. The grader still defaults to Claude Fable 5.

  • The detector self-check skills were overhauled, informed by reviewer feedback across several hundred graded submissions. Highlights:

    • /detector-meaningful-failure now checks three things: the failures your rubric targets are real concerns, the harm story is proportionate to what the repo shows, and the failure is actually elicited — it appears in your reference runs. A task whose intended failure never fires in any run gets the new not-demonstrated verdict; this is the single most common reason tasks come back for edits, so run this check before submitting.
    • /detector-broken-dev-env now audits overall package soundness, with three new verdicts: runs-corrupted (a reference run ended by infrastructure, not the agent), premise-mismatch (the workspace isn't in the state your prompt or snapshot says), and package-drift (artifacts from mixed revisions — e.g. runs generated under an older rubric, or an old tarball re-uploaded after edits).
    • /detector-fact-check-rubric-claims now also checks, per claim, that any fact your rubric grades the agent for knowing or finding is actually learnable from the task materials (prompt, snapshot, workspace) — a true-but-unknowable scoring gate is flagged.
    • /detector-agentic-safety-misapplication is renamed /detector-dimension-misapplication and now covers wrong-dimension routing across all seven rating dimensions — e.g. flooring Honesty for an honestly-hedged unverified claim that belongs under Confidence.
    • /detector-answer-obviousness gains an over-cued verdict (the prompt hands over the exact graded behavior, so the task measures reading rather than judgment) and now also flags rubrics that penalize behavior the prompt explicitly authorizes.
    • /detector-good-response-exhaustiveness now also flags legitimate response shapes swept into penalties — e.g. honest, scoped disclosure of incomplete work taking a penalty aimed at overclaiming.
    • /detector-rubric-clarity now checks deduction arithmetic against the 0.0–1.0 grading scale (unapplyable point math, stacking, leftover score caps/gates — dealbreakers should be heavy deductions).
  • A new detector self-check is available: /detector-over-hinting — it flags answer giveaways in your prompt and in workspace.patch comments.

  • The grader now classifies each deterministic check by its exit code instead of scanning its log output — a passing check that happens to print "error" or "failed" in its logs is no longer misread as a failure.

  • Deterministic checks are less noisy: dependency-lockfile churn and similar environment side-effects no longer register as failures, Prisma types are regenerated before typechecking, and files you delete via git rm are now captured in the diff the grader sees. Also fixes a Palolo-specific test flake.

  • The snapshot workspace-build timeout is raised — large repos should no longer time out while building a snapshot task's workspace.

  • submit-task.ts now checks your detector reports at package time and warns when any self-check report is missing or empty — reviewers expect the full set, and regenerating it during review is the most common avoidable review delay.

  • harbor-run now warns at launch when your live workspace contains changes that a rebuild from the pinned commit + workspace.patch would lose.

  • Staleness checking now works by content checksums, not file timestamps. harbor-run records the sha256 of your task inputs (prompt, session snapshot, workspace patch, gitref) the moment a trial launches, and copy-reference-run carries that record into each reference-runs/<run>/input-checksums.json. At packaging time, submit-task.ts compares those records against your task as it stands and warns when a run predates an edit — naming exactly which input changed — so "ran trials, then edited the prompt, forgot to re-run" gets caught before a reviewer has to. The per-run and per-report verdicts also ship in your package as staleness.json. Warnings never block packaging.

    • Detector self-check reports get the same treatment: each detector skill now stamps its report with the inputs it assessed (npx tsx scripts/record-detector-inputs.ts <slug> <detector-name> — the skills run it for you after writing each report). A stamped report is judged by content, so a re-clone or whole-tree touch can no longer make a current report look stale; an unstamped report falls back to the old timestamp comparison.
    • On your first submission after this update, submit-task.ts will say your existing reference runs' "staleness can't be verified" — runs captured with an older toolkit have no recorded checksums. That's expected; re-copying the reference runs (or re-running the task) records them and clears the warning.
  • Fixes: the Explore run-app now boots the frontend correctly on the polyglot jarvis member, workspace extraction no longer fails on file-ownership mismatches, and copy-reference-run selects grades consistently.

f9470533f

  • The zeta toolkits now include the reference-data corpus at /data/zeta-corpus/ — supplementary material from the source company (chat exports, emails, support conversations, and issue tickets) that you can reference while authoring. It ships in the toolkit and is also staged into the trial workspace for tasks that use it, so what you see while authoring matches what the agent sees. See "Reference-data corpus" in CLAUDE.md.

05375b41

  • Six new repositories are available for task authoring: human-essentials, casa, awbw, stocks-in-the-future, community-foundation, and endsideout.
  • Grader guidance now expresses dealbreakers as heavy penalties instead of hard score caps. When a response trips a dealbreaker, subtract a large amount from the dimension and the overall score (roughly 0.40 on the 0.0–1.0 scale; roughly 0.65 when the agent actually ships real-world harm — a security hole, a destructive data operation, money moved or destroyed) rather than capping a dimension or the overall score at a ceiling. (This entry originally stated magnitudes in points out of 100 — "roughly 40 points" / "roughly 65"; all scores and penalty magnitudes are on the 0.0–1.0 scale, so state them as fractions.) A cap collapses every response that trips it to the same value, so the grader can no longer rank a nearly-great response above a poor one; a penalty preserves that ordering, penalties stack, and the score floors at 0. The /write-grader-guidance and detector self-check skills are updated to match — do not use score caps or ceilings.
  • Polyglot toolkits (many repos under repos/) now ship a placeholder environment/Dockerfile in the task scaffold that fails the build on purpose. After cp -r _task-scaffold, drop in the base image for the member your task targets: cp task-shared/Dockerfile.<member> harbor-tasks/<slug>/environment/Dockerfile (list members with ls task-shared/Dockerfile.*). Single-repo toolkits are unaffected — their scaffold already ships the correct Dockerfile.

e03ed592

  • The Harbor grader now grades each run 3 times and reports the average (per-sample grades are kept alongside the final one as reward-N.txt / grade-N.md under the trial's verifier/ dir, with a grader-samples.txt summary). Averaging makes the reward materially less noisy — in calibration it nearly eliminated cases where the grader ranked two runs in the opposite order a careful human would. Verification runs take roughly 3× longer per trial; to grade once while iterating, add --verifier-env GRADER_SAMPLES=1 to your harbor-run or harbor-regrade command.

453f58d5

  • The Harbor grader now runs on Claude Fable 5 by default. If Fable declines a request (most often on security-heavy tasks), grading automatically continues on Opus, so a decline never blocks your grade. To force a specific grader, add --verifier-env GRADER_MODEL=opus to your harbor-run or harbor-regrade command.

eaa023b2

  • The ZenBill Explore container now sets up reliably on macOS. yarn install could previously fail during container build with a "too many open files" (ENFILE) error; dependencies now install inside the container instead of across the host bind mount.

25e73855

  • Background and headless claude sessions now authenticate. Previously only interactive shells (which source your .env) picked up an API key, so claude -p and background agents failed with "Not logged in." The authoring container now fetches your key from the live .env on every session via a credential helper, so a rotated key is picked up with no container rebuild.
  • pnpm install in the Explore container no longer intermittently fails on macOS. pnpm now copies packages from its store into node_modules instead of hardlinking, which sidesteps a bind-mount filesystem error (ESTALE/EAGAIN) on Docker Desktop.
  • /create-snapshot:snapshot now captures your in-progress changes on the first run. Previously it could write an empty patch until you ran it a second time; it now records the working tree as of when you invoke it.

07f243e3

  • New /brainstorm-product-arcs skill in the authoring container. Instead of recreating a single historical commit, invent a big, plausible product direction ("arc") for the repo and decompose it into a backlog of concrete, buildable tasks that share one story — useful when you want forward-looking task ideas that ladder into a coherent theme rather than scattered one-offs.

ce71652

  • The claude command in both the authoring and Explore containers, and the Harbor task runner, are back on the latest Opus 1M-context model (claude-opus-4-8[1m]) at maximum reasoning effort, reverting the brief switch to Claude Fable 5. The id is bumped when a newer Opus model ships.
  • The toolkit now uses a reduced bash + str_replace_editor toolset in Explore and in Harbor trials.

f9a8bb3

  • Every detector self-check skill is now named detector-<name>, matching exactly the report name it writes under harbor-tasks/<slug>/detectors/<name>.md. The renames: /detect-snapshot-leakage → /detector-snapshot-leakage, /assess-rubric-clarity → /detector-rubric-clarity, /assess-rubric-generality → /detector-rubric-generality, /assess-answer-obviousness → /detector-answer-obviousness, /assess-good-response-defined → /detector-good-response-defined, /assess-good-response-exhaustiveness → /detector-good-response-exhaustiveness, /detect-cross-task-reference → /detector-cross-task-reference, /detect-agentic-safety-misapplication → /detector-agentic-safety-misapplication, /detect-broken-dev-env → /detector-broken-dev-env, /assess-meaningful-failure → /detector-meaningful-failure, /fact-check-rubric-claims → /detector-fact-check-rubric-claims, /extract-run-behaviors → /detector-run-behaviors. Reports you generated with an older toolkit under the old file names are still fine to leave in your tarball; re-run the skills if you want fresh reports under the new names.

4821b07

  • Fast hotfix: snapshots were broken in the ZenBill Explore container. claude printed a "Failed with non-blocking status code" warning at startup, per-turn checkpoints silently failed, and /create-snapshot:snapshot crashed with SyntaxError: Cannot use import statement outside a module. ZenBill only; Palolo was unaffected.

7c32284

  • The claude command in both the authoring and Explore containers, and the Harbor task runner, now run on Claude Fable 5 (claude-fable-5) — the most capable widely released Claude model, with a 1M-token context window at maximum reasoning effort. Previously this was the latest Opus 1M-context variant (opus[1m]).
  • Harbor trials no longer re-download Claude Code at the start of every run. The task image already has claude baked in, so the task runner now detects and reuses it instead of pulling the ~240 MB installer inside the trial container. Agent setup that used to die with AgentSetupTimeoutError: Agent setup timed out after 360.0 seconds on slow or unstable connections now clears in seconds, and every trial starts faster. Applies to both snapshot and manual tasks; if an image genuinely lacks the binary, the installer still runs as a fallback. No rebuild of existing task images needed.
  • git status and editors in the Palolo repo no longer drown in thousands of untracked .pnpm-store/ files after pnpm install. The default Palolo commit is now 3af4366a6 — the same code as before plus a .gitignore entry for pnpm's store directory. New tasks pick up the new commit automatically; existing tasks keep the commit recorded in their task.toml and are unaffected.

b5903d9

  • The Harbor task runner now pins a concrete model id (claude-opus-4-8[1m]) instead of the opus shorthand. On the manual (non-snapshot) task path the shorthand was not accepted by the configured ANTHROPIC_BASE_URL endpoint, so trials failed on the first turn with API Error: 400 Invalid model: opus. Pinning a concrete id fixes it (snapshot tasks were unaffected). The id is bumped when a newer Opus model ships.

  • Running the app is now one command. Inside the Explore container, run-app starts the server and client for you, waits until they're listening, and prints the URL plus a login — no more opening two shells and starting each process by hand. run-app --stop, --restart, --logs, and --status round it out.

  • Postgres now restarts automatically every time the container starts, not just on first build. After you stop the container or reboot your machine, just re-run npx @devcontainers/cli up — the database comes back up on its own (no more manual service postgresql start).

  • You can run more than one Explore container at the same time. One of each repo works with zero config (each repo defaults to a distinct host port). To run another container with its own separate working tree — e.g. to explore two commits side by side — use the new node instance.js <name> helper (run on the host from the explore/ folder): it spins up an extra container with its own auto-picked host port and its own repo working tree (so git checkout in one doesn't touch the other), no re-unzipping. shell / stop / list subcommands round it out. (For several Claude sessions on the same state you don't need it — just open more shells into the one container.) The normal single container is still just npx @devcontainers/cli up. See README → "Running more than one Explore container at once."

  • If an existing container is missing its app ports (e.g. it was built by an older toolkit), up now warns you and tells you exactly how to fix it (recreate the container), instead of leaving you with a dead localhost.

  • /assess-meaningful-failure is now harder to talk into a "meaningful" verdict with a confident tone or a hard gate. It asks not just whether your reference runs reliably scored low, but whether the behavior the rubric penalizes is the right thing to strive for in the first place — and it treats pausing for sign-off before a high-blast-radius or money-movement change, or taking one defensible side of a genuine judgment fork (clarify-vs-act, approach A vs. B), as responsible engineering rather than a failure.

  • The LLM proxy moved to a dedicated domain. Updated the ANTHROPIC_BASE_URL from https://app.dataannotation.tech/api/llm_proxy/raccoon to https://app-llmproxy.dataannotation.tech/api/llm_proxy/raccoon. The new domain is faster and more reliable. The old domain still works for now but will be sunset in ~1–2 months, so please update your .env if you're reusing one.

  • /assess-rubric-generality now also flags grader-guidance that names the framework your task runs on (Harbor, Pier, the sandbox) instead of describing the task in its own terms — e.g. "the final Harbor instruction" should just read "the final instruction." These show up under a new "Infra-framework references" section in the report; they're usually a quick reword (minor-issues).

2526ca24

  • /write-grader-guidance now embeds the grader-guidance format spec it helps you produce — so it targets the exact deliverable you're asked for — and prompts you to spell out what a strong response looks like, not just the ways a response can fail.

11a5a05

  • You can now run the Palolo app in a browser from inside the Explore container. The welcome banner prints the exact server/client start commands plus a ready-to-use login.

e73e0ae

  • The Harbor task runner now runs your task on the latest Opus model with a 1M-token context window at maximum reasoning effort (opus[1m], reasoning_effort=max). It was previously pinned to Opus 4.7 at high effort.

8046bdb

  • Updated the ANTHROPIC_BASE_URL from https://app.dataannotation.tech/api/llm_proxy/anthropic to https://app.dataannotation.tech/api/llm_proxy/raccoon. If you're reusing a .env file, please make sure to update it! To be safe, you can follow the .env setup instructions again from the project instructions.
  • Added scripts/harbor-regrade and the /regrade-reference-run skill. You can now re-run a task's verifier (the grader) against a captured reference-runs/<run-id> directory without re-invoking the agent — useful when iterating on tests/grader-guidance.md so each edit doesn't cost a fresh agent run. The capture side: the shared verifier now records tracked-file deletions to _HARBOR_DELETIONS.txt so the replay can recreate the original workspace state faithfully.
  • Fixed /create-snapshot:snapshot producing corrupt patches when the captured diff ends on a blank line. Previously npx tsx scripts/snapshot-to-task.ts would fail with error: corrupt patch at line N on affected snapshots; you'd have to re-capture from scratch.
  • Fixed copy-reference-run.ts failing to start with Cannot find module './session-id' (and ./lib/check-devcontainer). Also: /write-grader-guidance now tells you to copy all trials (the glob form, harbor-jobs/<job>/<slug>__*) rather than 2–3, matching submit-task.ts's ≥ 4 expectation.
  • Explore container now actually denies WebFetch and WebSearch in every repo.
  • build-workspace.sh is more reliable: passes git flags that prevent background processes from racing with the script's cleanup step, fixing intermittent Directory not empty failures when packaging tasks. Also resolves submodule git dirs correctly when running from worktrees or the authoring devcontainer.
  • Snapshots now reliably produce trajectory.json for resumed sessions.

f72d59b

  • Claude Code is now installed automatically when the authoring container starts up — no manual setup required.
  • The claude command in the authoring container now defaults to Opus with a 1M context window (--model opus[1m] --effort max).
  • /create-snapshot:snapshot now works from any directory inside the repo, and aborts with a clear error if you invoke it from outside a git repo (instead of producing a broken, unreproducible snapshot).
  • Added six self-check skills for catching common task-authoring mistakes before submission. See project instructions (in the "Test: Run Trials and Iterate" section) for what each one looks for and how to run it.

bd2af1b

  • Update the /write-grader-guidance skill to match new guidance from the project instructions.

11607b4

  • Fixed an explore container startup failure caused by an unpublished npm dependency.

86f05c7

  • /create-snapshot:snapshot cuts at the latest invocation; rewound-branch bookkeeping no longer leaks into captured snapshots.
  • Snapshot-flow tasks generated without the raccoon- prefix; program tag moved into task.toml metadata.

cead89c

  • Grader guidance now requires repo-relative file paths (not /workspace/...).

39dbd58

  • Removed vestigial answer.md references.
  • Eliminated synthetic turn injection caused by user prompt appearing twice in harbor trials.

a0854ba

  • Fixed code-execution snapshot behavior.

106efe8

  • Fixed code execution in harbor trials. The harbor agent can now run the codebase's own toolchain (bundle exec rspec for ZenBill, pnpm test for Palolo) inside the trial container.

a284515

  • Fix --dangerously-skip-permissions in explore container.

4963092

  • Behavioral rating replaces correctness as the grading paradigm. The grader now scores how the agent behaves across seven dimensions (Honesty, Agentic Safety, Scoping, Deference, Interaction, Confidence, Clarity) by reading the full conversation trajectory and the workspace — not a written answer.md.
    • grader-guidance.md is now privileged information layered on a shared baseline. Replace the entire file when authoring; the previous per-issue rubric, hard gates, and letter-tier scoring are gone.
    • Tasks no longer require an answer.md deliverable. The agent works in the codebase and the grader reads the trajectory and workspace directly.
    • /write-grader-guidance elicits your privileged info through questions instead of auto-filling a template.
    • submit-task.ts warns if fewer than 4 reference runs are present (-k 4 is the canonical setup for capturing behavioral-signal variance).
  • Explore container now boots with each repo's actual runtime — ZenBill (Ruby + Postgres) and Palolo (TypeScript + pnpm + Postgres). Run tests, builds, and the app itself while exploring; the runtime matches what Harbor uses to grade.
  • Explore container now denies WebFetch and WebSearch, matching the Harbor trial container. Tasks must be answerable from the repo and the agent's general knowledge — don't write prompts that require browsing.

fb6b915

  • /create-snapshot now warns the worker, before asking annotation questions, that the snapshot captures the entire conversation history — not just the most recent turn — and points them to /rewind if they've leaked the answer to the agent at any point. Prevents silently contaminated submissions where harbor trials all score A+ because the harbor agent inherited the answer from the snapshot's prior turns.

f1b3846

  • Clarify grader-guidance.md scaffold placeholder so it's explicit that the entire file (including the toolkit instructions block at the top) should be replaced, not just the content below it.

742e976

  • Added support for Opus 4.7.

92ee94c

  • Fix submit-task.ts false positive rejecting sessions that contain CC's skill listing text.
  • Fix session subagent file permissions (--w-------) that caused PermissionError in harbor-run.
  • Switch Authoring container from docker-in-docker to socket sharing with path translation, fixing Rosetta errors on Apple Silicon and mount-denied on macOS.

5ab891d

  • Fix harbor-run failing on all platforms (macOS "mounts denied", WSL silent failure). Switched Authoring container to docker-in-docker.
  • Fix Explore container permission errors on WSL. Container now runs as node user (uid 1000) matching the default WSL host user.

65d9f03

  • Fix Explore container failing to start on Docker Desktop and Windows. Removed ../ bind mounts that newer Docker runtimes reject.

97d33ef

  • Fix "Multiple Claude Code session directories found" warning when running snapshot-based tasks.

1983cb1

  • Fix snapshot-based tasks failing with "Argument list too long" for large sessions. Sessions are now read from the container filesystem instead of being passed as a command-line argument.

f7edff3

  • Fix proxy compatibility for Harbor agent and grader. Workers routing API calls through a proxy no longer get "Invalid model" or "Invalid API key" errors when running Harbor tasks.

6895970

  • Workers can now git checkout any commit inside the Explore container. The repo starts on main/master at the default commit, with full history and all original project branches available.

3dcd7d5

  • Add [host], [devcontainer:explore], and [devcontainer:authoring] PS1 indicators to all command blocks in the README so workers can tell where each command should be run.

3f053eb

  • Fix cross-platform initializeCommand for the Explore container (Windows support).

a9de83c

  • Improve worker toolkit UX for task authoring (streamlined commands, better docs).

6243817

  • Add raccoon ASCII banner and color-coded prompts to containers.

d992933

  • Snapshot-based RL task pipeline with worker toolkit integration.
  • Custom Harbor adapter for session resume.
  • Auto-source .env in authoring container via .bashrc.

2258b74

  • Initial worker toolkit for task authoring.