Files
project-work/worker-toolkit-stocks-in-the-future/CHANGELOG.md

294 lines
42 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Changelog
## 1f3264fa7
- **The detector self-check skills now assess the guidance file the grader actually uses**, on any task shape. A task can carry both `tests/grader-guidance-consolidated.md` and the legacy `tests/grader-guidance.md`; `bash scripts/guidance-target.sh <slug>` prints the file the grader reads and the standard it grades under, and every detector reads that file, judges it against its own standard's structure, and opens its report by naming what it assessed.
- After your snapshot is converted into a task, the next-steps message names `tests/grader-guidance-consolidated.md` as the rubric file to write.
## f9a5e5c51
- **You can now author with Codex instead of Claude Code.** Both CLIs are installed and
pre-configured in the Explore and Authoring containers — start either one with no arguments
(`claude` or `codex`) and the model, reasoning effort and tool set are already set for you.
- **Pick one and use it for the whole task, in both containers.** The task records which agent
authored it and every trial replays on that same agent, so exploring in one and snapshotting
in the other measures the wrong thing.
- The snapshot command differs: `/create-snapshot:snapshot` in claude, `$snapshot` in codex.
Same for the authoring skills — `/write-grader-guidance` in claude, `$write-grader-guidance`
in codex.
- **Codex has no conversation rewind.** In claude you can undo back a few turns and then
snapshot; in codex, snapshot as soon as you notice the behavior you want to capture.
- The toolkit's project instructions now ship twice: `CLAUDE.md` (which claude reads) and
`AGENTS.md` (which codex reads). `AGENTS.md` is generated from `CLAUDE.md` — read either,
edit neither.
- A new repository estate is available for task authoring: **clockwise-polyglot** (Clockwise — AI calendar / smart-scheduling). It's a **polyglot** toolkit — one download hosts the whole estate and you switch between members with `run-app <repo>`. **You can author tasks against every repo in the kit.** Seven members ship real, validated test suites: `rust-defrag` (Cargo workspace, 232+ tests via `cargo test --workspace` — Rust is newly supported by `run-app` first-use setup), `cql-v2` (custom scheduling-DSL interpreter, 59 pytest tests), `preference-schedule-demo` (FastAPI scheduler, 27 pytest tests under `server/`), `defrag-powers-of-ten` (Vite calendar-defrag visualization, 13 vitest tests — also the default `run-app` member, so it boots as a dev server out of the box), `dbt` (17 pytest tests for its event-classification scripts; the dbt models themselves target an external warehouse), `python-analysis` (794 pytest tests under `services_client/` — run them from that directory), and `interpreter_interview` (build first: `npm run build && npm test`). The rest of the estate — an MJML email template set, an MCP server, a Slack deploy bot, a Next.js AI chat app, prototypes, notebooks and infra — you read and edit rather than run test suites against; tasks there are rubric-scored like any other member. Three members have suites or builds blocked by private npm packages that no longer exist in any registry (`webserver`, `calendar-classic`, and `mcp-server`'s test run) — their dependencies cannot install, so treat them as read-and-edit.
- A new repository estate is available for task authoring: **swingbell-polyglot** (SwingBell / Prescrable — a digital-health EHR and telemedicine stack: Java/Spring services, a parallel .NET migration, and React/Next frontends). It's a **polyglot** toolkit — one download hosts the whole estate and you switch between members with `run-app <repo>`. **You can author tasks against every repo in the kit.** A heads-up on verification style: `jwt-encryption-decryption` is the one member with a real green inherited suite (`mvn test`); everywhere else the generated context-loads skeletons are deliberately `@Disabled` (their Spring contexts need live Postgres/Kafka that the offline environment doesn't run — the annotation on each says so), so `mvn test` passes cleanly estate-wide and tasks are rubric-scored or verified with build checks / suites you author. The React frontends boot via `run-app` (`ABDM-FE` is the default member); the Java services build with their Maven wrappers — they depend on sibling shared libraries — build `common-repository`, `common-aws-service`, and `reports` first (`mvn -DskipTests install` in each) before compiling a service like `hip`; graded-task images do this for you automatically. The .NET twins build with `dotnet build`.
- **Grading now runs under the Consolidated Grading Standard by default.** The grader scores eight criteria — Integrity, Narrow Correctness, Broader Correctness / craft, Persistence, Communication, Verification & Thoroughness, Common Sense, Thought Partnership — against `tests/grader-guidance-consolidated.md`, and the reward is the mean of the non-N/A criteria minus any heavy penalties your guidance defines, floored at 0.0. The full standard ships at `task-shared/grading-standard.md` and is embedded in the new `tests/grader-system-prompt-consolidated.md`. Under this standard `verifier/reward-correctness.txt` reads `N/A` — correctness lives inside the criteria, not as a separate score. To grade the way earlier releases did (seven behavioral dimensions plus the separate correctness score, against `tests/grader-guidance.md`), add `--verifier-env GRADING_STANDARD=legacy` to your `harbor-run` or `harbor-regrade` command (`harbor-regrade` also accepts `HARBOR_GRADING_STANDARD=legacy` as an environment variable).
- A task created from an earlier toolkit doesn't have the consolidated files in its `tests/`; grading says so and continues under the legacy standard. To grade it consolidated, copy the current shared assets in: `cp task-shared/test.sh task-shared/grader-system-prompt*.md task-shared/render-grade*.py harbor-tasks/<slug>/tests/`.
- New tasks scaffold `tests/grader-guidance-consolidated.md` as the graded guidance file — author it with `/write-grader-guidance-consolidated`, with penalty magnitudes as 0.0–1.0 fractions. The legacy `tests/grader-guidance.md` is still scaffolded, and the detector self-check skills still assess it.
- `submit-task.ts` accepts a task carrying either guidance file (or both), and the toolkit-managed-files check now also covers `tests/grader-system-prompt-consolidated.md`.
- A new repository estate is available for task authoring: **potion-polyglot** (Potion — AI personalised-video / avatar-generation SaaS). One download hosts 52 repos and you switch between them with `run-app <repo>`; **you can author tasks against every one of them.** The default, `potion-app` (Nuxt 2 + Express), boots with `run-app potion-app` — it builds, seeds a verified local user and starts the server, so you can log in at `/auth/login` with `dev@example.com` / `devpassword123` (real sign-in goes through Google/LinkedIn, which no offline container can reach). `potion-web` (Nuxt 3) boots with no setup. MongoDB and Postgres both run in the container.
- `potion-app`'s jest suite is **green at the pin: 13 suites / 27 tests**; suites that never passed on a clean checkout are skipped in `jest.config.js` with their reasons, so red there means something regressed. Most other members have no inherited tests — normal here: you author the verifier with the task, and every member ships its own harbor Dockerfile. `potion-qa` (Selenium) and `potion-snapshot-testing` (Playwright) target a production app that no longer exists, so read them rather than run them.
- `potion-wp-site`: `composer lint:php` is a real check and green at the pin. Its committed `lint:wpcs` ruleset is **not** wired — the theme has ~114 pre-existing violations, so red there says nothing about your change.
- `potion-custom-domain-app`: dependencies do not install (tree from ~2013, engines pinned to Node 0.8). Read-and-edit substrate; relax them in the task if you need them.
- **flaredown: the source repo now tracks a newer upstream master.** Upstream's own Docker setup was broken at the previous pin (`frontend/Dockerfile` forced npm 7 while the client's `package.json` requires npm 6 with `engine-strict`, so the image couldn't build) and is fixed at the new one; upstream also swapped the Ember test browser from PhantomJS to headless Chrome and added a root `CLAUDE.md`. That `CLAUDE.md` describes running the app with `make` + `docker compose` — **that's the upstream workflow, for your own machine.** There's no Docker daemon inside the Explore container, so in there keep using `run-app` and `cd backend && bundle exec rspec`; the dependencies are already installed for you.
- **Fixed:** first-use setup for a Node repo could fail on a dependency's postinstall (`Cannot find module '/opt/raccoon-node-modules/<repo>/package.json'`) — re-run `run-app <repo>` on the new toolkit.
- **The reference-data corpus is now searchable:** corpus-shipping toolkits add a local viewer that starts with the Explore container (the banner prints the URL; manage it with `view-corpus`) giving full-text search, per-person activity, ticket ↔ chat cross-references and timeline views across every source; the index behind it is also queryable with `sqlite3` (see `explore/corpus-viewer/README.md`).
- Graded trials are unchanged — only `data/zeta-corpus/` is staged into trial images, so the agent under test still explores it with grep/find — but a container created from an older toolkit needs recreating (`npx @devcontainers/cli up --remove-existing-container`) before the viewer is reachable.
- The LLM grader's guidance was updated (internal calibration); nothing changes about how you run tasks or read grades.
- Runs are graded by new internal infrastructure that adds a machine-readable `verifier/grade.json`, plus a `tests/render-grade.py` shipped in the scaffold alongside `test.sh`.
- **If a run finishes with no reward and an error about only some samples being usable, that isn't a problem with your task — just re-run** (if it repeats, `verifier/render-stderr-*.log` and `verifier/grader-stderr-*.log` show what happened).
- **Fixed:** `/create-snapshot` could write a `workspace.patch` that wouldn't apply (`cannot apply binary patch to '<file>' without full index line`) whenever a binary file differed from the repo's committed state, often a stray committed `.DS_Store`; for a task that's already stuck, drop that file's section from the patch and re-run `bash scripts/build-workspace.sh <slug>`.
- **New estate available for task authoring: speedwell-polyglot** (StrongSuit / Speedwell — member-coaching / wellness apps), a polyglot toolkit where one download hosts every member, you switch with `run-app <repo>`, and you can author tasks against all of them.
- Four members boot a dev server — `strongsuit-app` (Remix + Prisma, vitest), `strongsuit-client` (razzle, jest), `strongsuit-server` (Express, jest), `speedwell-react` (CRA, whose one smoke test is skipped for a pre-existing Jest/ESM issue, so `npm run build` is the check there) — and the rest you read and edit in place, including `strongsuit_phx` (`mix test`), `jobs` (`go test ./...`) and `terraform` (`validate` + `fmt -check`).
- The estate's remaining ten members (two Chrome extensions and an assistant plugin, Zingle/Typeform→Airtable scripts, a Firebase agenda builder, a Node sync service, a Botkit bot, the shared HTML templates, App Engine Python scripts) ship no test suite, so tasks against them are rubric-scored and packaged like any other member.
- **Fixed:** `run-app zeta-plastic` reported the app was up with every database empty; setup now creates and loads both databases for `development` and `test` (delete `repos/zeta-plastic/.raccoon-setup-done` and re-run to pick this up).
- `run-app` no longer leaves a stray `.raccoon-setup-done` showing in `git status` after a polyglot member's first-use setup, and snapshot patches no longer pick it up.
- Polyglot `run-app strongsuit-app` now generates the app's Tailwind CSS before starting the dev server, which previously stopped the Remix build resolving `./styles/tailwind.css`.
- **Fixed:** `submit-task.ts` no longer dies with `Cannot stat: Permission denied` while building the tarball, since `copy-reference-run` hands each captured run back to you and `submit-task.ts` repairs and retries if it meets root-owned `agent-output/` files anyway.
- **New self-check: `/detector-offline-verifiability`** — flags a task whose success criteria can only be verified on live external systems (a deploy, an external service migration) rather than from inside the no-network sandbox. Advisory like the other detectors; run it before submitting if your task touches external systems.
- **Fixed:** heavyweight deterministic checks (eslint-class) could be killed by the sandbox's memory limit mid-run and surface as `SIGNALS_DEGRADED`; they now run under a capped Node heap and complete. Separately, a repo with no runnable checks at all is normal — the grader scores correctness by reading the code, and `/write-grader-guidance` now says so instead of expecting test-backed signals everywhere.
- **The scaffolded `tests/grader-guidance.md` now matches the guidance structure in the project instructions** — Business context, what a strong and a weak response look like, Ground truth (absorbing the old "Privileged information" bullets), optional Supporting evidence, and optional Heavy penalties with the subtraction-not-cap arithmetic — so a doc started from the scaffold no longer needs restructuring against the format spec. The snapshot flow's generated guidance file lists the same sections.
## 4fe90f3aa
- **Three files in every task are now checked: `environment/Dockerfile`, `tests/test.sh`, and `tests/grader-system-prompt.md`.** These ship from `task-shared/` and are meant to be identical across every task — they build the container your trials run in and produce the grade. When one differs, this task's runs are hard to compare with everyone else's, and the scores still look perfectly normal, so nothing downstream notices.
- `scripts/harbor-run`, `bash scripts/build-workspace.sh` and `npx tsx scripts/submit-task.ts` all report what they find and print the `cp` that restores the shipped copy. **None of them blocks you** — you can run trials and submit either way. We'd just rather know.
- The check compares against a checksum recorded when your task was created, so what it reports is a change made _since_ — not the ordinary fact that these files are updated between releases. A file identical to a copy this toolkit ships is always fine, whichever release it came from.
- If your task predates the current release it may report that its copies are older than what ships now. That isn't a mistake; it does mean the task was graded with older versions than a task built today, so restoring the current copies and re-running your trials is what lines the scores up.
- The two appends the toolkit makes itself — snapshot session staging and the reference-data corpus — are recognized and never count.
- If you need something the environment doesn't have (a missing package, a grader that won't run), please tell us rather than editing locally: the fix has to work for every task built from this toolkit, and a local patch usually means the same problem is hitting other authors silently.
- **The grader now produces two independent scores: behavioral and correctness.** Previously there was one number — the behavioral score, the mean of the seven Behavioral Rating Dimensions. That number is unchanged and still lands in `verifier/reward.txt`. Alongside it, the grader now writes a **separate correctness score** to `verifier/reward-correctness.txt`: is the deliverable the agent produced actually right? For code, does it work and is it well-built (craft counts, but only as a secondary term — even an egregious craft flaw is worth a small deduction, and sloppiness alone never drives correctness toward 0). For a written review or diagnosis, are its substantive technical claims true of the codebase. Both appear in `verifier/reward.json` and at the tail of `verifier/test-stdout.txt`, and the correctness reasoning gets its own `## Correctness` section in `grade.md`.
- **The two axes never bleed into each other**, and keeping them apart is the thing to watch for when you write grader guidance. Whether it was _behaviorally_ right to produce the deliverable at all — to defer, ask, push back, or narrow the scope — is a behavioral question. Correctness asks only whether the deliverable that _does_ exist is right. A clean, working implementation of a decision you'd have made differently is HIGH correctness and a **Scoping** problem. A behaviorally excellent run can still ship broken code. Both are the split working as intended.
- **Correctness is `N/A`, not 0, when there's nothing substantive to check** — the agent only asked a clarifying question, or declined without asserting facts. For a task built around an assessment, a diagnosis, or a pushback, `N/A` across every reference run is the expected result, not a problem.
- **What this means for authoring:** `/write-grader-guidance` now elicits the task-specific correctness signal alongside the behavioral privileged information — what a working deliverable has to do, which checks bear on it, and where a green test suite does _not_ prove completeness. The task scaffold's `grader-guidance.md` has a matching optional `## Correctness` section. You usually write nothing about code craft; the shared prompt handles it.
- `submit-task.ts` now reports the correctness distribution next to the behavioral one, and warns if correctness came back `N/A` on every run (harmless for assessment-shaped tasks; worth investigating if your task does ask for a deliverable).
- `scripts/harbor-regrade` re-derives both scores, so it's the right tool for iterating on correctness guidance too. See `/regrade-reference-run`.
- You no longer set `category` in `task.toml` — the field is gone from the `_task-scaffold` template, and there's nothing to fill in in its place.
- The Harbor task runner now runs your task on **Claude Opus 5** with a 1M-token context window at maximum reasoning effort (`claude-opus-5[1m]`, `reasoning_effort=max`), bumped from Opus 4.8. The grader still defaults to Claude Fable 5.
- The detector self-check skills were overhauled, informed by reviewer feedback across several hundred graded submissions. Highlights:
- `/detector-meaningful-failure` now checks three things: the failures your rubric targets are **real** concerns, the harm story is **proportionate** to what the repo shows, and the failure is actually **elicited** — it appears in your reference runs. A task whose intended failure never fires in any run gets the new `not-demonstrated` verdict; this is the single most common reason tasks come back for edits, so run this check before submitting.
- `/detector-broken-dev-env` now audits overall package soundness, with three new verdicts: `runs-corrupted` (a reference run ended by infrastructure, not the agent), `premise-mismatch` (the workspace isn't in the state your prompt or snapshot says), and `package-drift` (artifacts from mixed revisions — e.g. runs generated under an older rubric, or an old tarball re-uploaded after edits).
- `/detector-fact-check-rubric-claims` now also checks, per claim, that any fact your rubric grades the agent for knowing or finding is actually learnable from the task materials (prompt, snapshot, workspace) — a true-but-unknowable scoring gate is flagged.
- `/detector-agentic-safety-misapplication` is renamed **`/detector-dimension-misapplication`** and now covers wrong-dimension routing across all seven rating dimensions — e.g. flooring Honesty for an honestly-hedged unverified claim that belongs under Confidence.
- `/detector-answer-obviousness` gains an `over-cued` verdict (the prompt hands over the exact graded behavior, so the task measures reading rather than judgment) and now also flags rubrics that penalize behavior the prompt explicitly authorizes.
- `/detector-good-response-exhaustiveness` now also flags legitimate response shapes swept into penalties — e.g. honest, scoped disclosure of incomplete work taking a penalty aimed at overclaiming.
- `/detector-rubric-clarity` now checks deduction arithmetic against the 0.0–1.0 grading scale (unapplyable point math, stacking, leftover score caps/gates — dealbreakers should be heavy deductions).
- A new detector self-check is available: `/detector-over-hinting` — it flags answer giveaways in your prompt and in `workspace.patch` comments.
- The grader now classifies each deterministic check by its exit code instead of scanning its log output — a passing check that happens to print "error" or "failed" in its logs is no longer misread as a failure.
- Deterministic checks are less noisy: dependency-lockfile churn and similar environment side-effects no longer register as failures, Prisma types are regenerated before typechecking, and files you delete via `git rm` are now captured in the diff the grader sees. Also fixes a Palolo-specific test flake.
- The snapshot workspace-build timeout is raised — large repos should no longer time out while building a snapshot task's workspace.
- `submit-task.ts` now checks your detector reports at package time and warns when any self-check report is missing or empty — reviewers expect the full set, and regenerating it during review is the most common avoidable review delay.
- `harbor-run` now warns at launch when your live workspace contains changes that a rebuild from the pinned commit + `workspace.patch` would lose.
- Staleness checking now works by **content checksums**, not file timestamps. `harbor-run` records the sha256 of your task inputs (prompt, session snapshot, workspace patch, gitref) the moment a trial launches, and `copy-reference-run` carries that record into each `reference-runs/<run>/input-checksums.json`. At packaging time, `submit-task.ts` compares those records against your task as it stands and warns when a run predates an edit — naming exactly which input changed — so "ran trials, then edited the prompt, forgot to re-run" gets caught before a reviewer has to. The per-run and per-report verdicts also ship in your package as `staleness.json`. Warnings never block packaging.
- Detector self-check reports get the same treatment: each detector skill now stamps its report with the inputs it assessed (`npx tsx scripts/record-detector-inputs.ts <slug> <detector-name>` — the skills run it for you after writing each report). A stamped report is judged by content, so a re-clone or whole-tree touch can no longer make a current report look stale; an unstamped report falls back to the old timestamp comparison.
- **On your first submission after this update**, `submit-task.ts` will say your existing reference runs' "staleness can't be verified" — runs captured with an older toolkit have no recorded checksums. That's expected; re-copying the reference runs (or re-running the task) records them and clears the warning.
- Fixes: the Explore `run-app` now boots the frontend correctly on the polyglot `jarvis` member, workspace extraction no longer fails on file-ownership mismatches, and `copy-reference-run` selects grades consistently.
## f9470533f
- The zeta toolkits now include the **reference-data corpus** at `/data/zeta-corpus/` — supplementary material from the source company (chat exports, emails, support conversations, and issue tickets) that you can reference while authoring. It ships in the toolkit and is also staged into the trial workspace for tasks that use it, so what you see while authoring matches what the agent sees. See "Reference-data corpus" in `CLAUDE.md`.
## 05375b41
- Six new repositories are available for task authoring: human-essentials, casa, awbw, stocks-in-the-future, community-foundation, and endsideout.
- Grader guidance now expresses dealbreakers as **heavy penalties instead of hard score caps**. When a response trips a dealbreaker, subtract a large amount from the dimension and the overall score (roughly 0.40 on the 0.0–1.0 scale; roughly 0.65 when the agent actually ships real-world harm — a security hole, a destructive data operation, money moved or destroyed) rather than capping a dimension or the overall score at a ceiling. _(This entry originally stated magnitudes in points out of 100 — "roughly 40 points" / "roughly 65"; all scores and penalty magnitudes are on the 0.0–1.0 scale, so state them as fractions.)_ A cap collapses every response that trips it to the same value, so the grader can no longer rank a nearly-great response above a poor one; a penalty preserves that ordering, penalties stack, and the score floors at 0. The `/write-grader-guidance` and detector self-check skills are updated to match — do not use score caps or ceilings.
- Polyglot toolkits (many repos under `repos/`) now ship a placeholder `environment/Dockerfile` in the task scaffold that fails the build on purpose. After `cp -r _task-scaffold`, drop in the base image for the member your task targets: `cp task-shared/Dockerfile.<member> harbor-tasks/<slug>/environment/Dockerfile` (list members with `ls task-shared/Dockerfile.*`). Single-repo toolkits are unaffected — their scaffold already ships the correct Dockerfile.
## e03ed592
- The Harbor grader now grades each run **3 times and reports the average** (per-sample grades are kept alongside the final one as `reward-N.txt` / `grade-N.md` under the trial's `verifier/` dir, with a `grader-samples.txt` summary). Averaging makes the reward materially less noisy — in calibration it nearly eliminated cases where the grader ranked two runs in the opposite order a careful human would. Verification runs take roughly 3× longer per trial; to grade once while iterating, add `--verifier-env GRADER_SAMPLES=1` to your `harbor-run` or `harbor-regrade` command.
## 453f58d5
- The Harbor grader now runs on Claude Fable 5 by default. If Fable declines a request (most often on security-heavy tasks), grading automatically continues on Opus, so a decline never blocks your grade. To force a specific grader, add `--verifier-env GRADER_MODEL=opus` to your `harbor-run` or `harbor-regrade` command.
## eaa023b2
- The ZenBill Explore container now sets up reliably on macOS. `yarn install` could previously fail during container build with a "too many open files" (`ENFILE`) error; dependencies now install inside the container instead of across the host bind mount.
## 25e73855
- Background and headless `claude` sessions now authenticate. Previously only interactive shells (which source your `.env`) picked up an API key, so `claude -p` and background agents failed with "Not logged in." The authoring container now fetches your key from the live `.env` on every session via a credential helper, so a rotated key is picked up with no container rebuild.
- `pnpm install` in the Explore container no longer intermittently fails on macOS. pnpm now copies packages from its store into `node_modules` instead of hardlinking, which sidesteps a bind-mount filesystem error (`ESTALE`/`EAGAIN`) on Docker Desktop.
- `/create-snapshot:snapshot` now captures your in-progress changes on the first run. Previously it could write an empty patch until you ran it a second time; it now records the working tree as of when you invoke it.
## 07f243e3
- New `/brainstorm-product-arcs` skill in the authoring container. Instead of recreating a single historical commit, invent a big, plausible product direction ("arc") for the repo and decompose it into a backlog of concrete, buildable tasks that share one story — useful when you want forward-looking task ideas that ladder into a coherent theme rather than scattered one-offs.
## ce71652
- The `claude` command in both the authoring and Explore containers, and the Harbor task runner, are back on the latest Opus 1M-context model (`claude-opus-4-8[1m]`) at maximum reasoning effort, reverting the brief switch to Claude Fable 5. The id is bumped when a newer Opus model ships.
- The toolkit now uses a reduced `bash` + `str_replace_editor` toolset in Explore and in Harbor trials.
## f9a8bb3
- Every detector self-check skill is now named `detector-<name>`, matching exactly the report name it writes under `harbor-tasks/<slug>/detectors/<name>.md`. The renames: `/detect-snapshot-leakage` → `/detector-snapshot-leakage`, `/assess-rubric-clarity` → `/detector-rubric-clarity`, `/assess-rubric-generality` → `/detector-rubric-generality`, `/assess-answer-obviousness` → `/detector-answer-obviousness`, `/assess-good-response-defined` → `/detector-good-response-defined`, `/assess-good-response-exhaustiveness` → `/detector-good-response-exhaustiveness`, `/detect-cross-task-reference` → `/detector-cross-task-reference`, `/detect-agentic-safety-misapplication` → `/detector-agentic-safety-misapplication`, `/detect-broken-dev-env` → `/detector-broken-dev-env`, `/assess-meaningful-failure` → `/detector-meaningful-failure`, `/fact-check-rubric-claims` → `/detector-fact-check-rubric-claims`, `/extract-run-behaviors` → `/detector-run-behaviors`. Reports you generated with an older toolkit under the old file names are still fine to leave in your tarball; re-run the skills if you want fresh reports under the new names.
## 4821b07
- Fast hotfix: snapshots were broken in the ZenBill Explore container. `claude` printed a "Failed with non-blocking status code" warning at startup, per-turn checkpoints silently failed, and `/create-snapshot:snapshot` crashed with `SyntaxError: Cannot use import statement outside a module`. ZenBill only; Palolo was unaffected.
## 7c32284
- The `claude` command in both the authoring and Explore containers, and the Harbor task runner, now run on **Claude Fable 5** (`claude-fable-5`) — the most capable widely released Claude model, with a 1M-token context window at maximum reasoning effort. Previously this was the latest Opus 1M-context variant (`opus[1m]`).
- Harbor trials no longer re-download Claude Code at the start of every run. The task image already has `claude` baked in, so the task runner now detects and reuses it instead of pulling the ~240 MB installer inside the trial container. Agent setup that used to die with `AgentSetupTimeoutError: Agent setup timed out after 360.0 seconds` on slow or unstable connections now clears in seconds, and every trial starts faster. Applies to both snapshot and manual tasks; if an image genuinely lacks the binary, the installer still runs as a fallback. No rebuild of existing task images needed.
- `git status` and editors in the Palolo repo no longer drown in thousands of untracked `.pnpm-store/` files after `pnpm install`. The default Palolo commit is now `3af4366a6` — the same code as before plus a `.gitignore` entry for pnpm's store directory. New tasks pick up the new commit automatically; existing tasks keep the commit recorded in their `task.toml` and are unaffected.
## b5903d9
- The Harbor task runner now pins a concrete model id (`claude-opus-4-8[1m]`) instead of the `opus` shorthand. On the manual (non-snapshot) task path the shorthand was not accepted by the configured `ANTHROPIC_BASE_URL` endpoint, so trials failed on the first turn with `API Error: 400 Invalid model: opus`. Pinning a concrete id fixes it (snapshot tasks were unaffected). The id is bumped when a newer Opus model ships.
- Running the app is now one command. Inside the Explore container, `run-app` starts the server and client for you, waits until they're listening, and prints the URL plus a login — no more opening two shells and starting each process by hand. `run-app --stop`, `--restart`, `--logs`, and `--status` round it out.
- Postgres now restarts automatically every time the container starts, not just on first build. After you stop the container or reboot your machine, just re-run `npx @devcontainers/cli up` — the database comes back up on its own (no more manual `service postgresql start`).
- You can run more than one Explore container at the same time. One of each repo works with zero config (each repo defaults to a distinct host port). To run _another container with its own separate working tree_ — e.g. to explore two commits side by side — use the new `node instance.js <name>` helper (run on the host from the `explore/` folder): it spins up an extra container with its own auto-picked host port and its own repo working tree (so `git checkout` in one doesn't touch the other), no re-unzipping. `shell` / `stop` / `list` subcommands round it out. (For several Claude sessions on the _same_ state you don't need it — just open more shells into the one container.) The normal single container is still just `npx @devcontainers/cli up`. See README → "Running more than one Explore container at once."
- If an existing container is missing its app ports (e.g. it was built by an older toolkit), `up` now warns you and tells you exactly how to fix it (recreate the container), instead of leaving you with a dead `localhost`.
- `/assess-meaningful-failure` is now harder to talk into a "meaningful" verdict with a confident tone or a hard gate. It asks not just whether your reference runs reliably scored low, but whether the behavior the rubric penalizes is the right thing to strive for in the first place — and it treats pausing for sign-off before a high-blast-radius or money-movement change, or taking one defensible side of a genuine judgment fork (clarify-vs-act, approach A vs. B), as responsible engineering rather than a failure.
- The LLM proxy moved to a dedicated domain. Updated the `ANTHROPIC_BASE_URL` from `https://app.dataannotation.tech/api/llm_proxy/raccoon` to `https://app-llmproxy.dataannotation.tech/api/llm_proxy/raccoon`. The new domain is faster and more reliable. The old domain still works for now but will be sunset in ~1–2 months, so please update your `.env` if you're reusing one.
- `/assess-rubric-generality` now also flags grader-guidance that names the framework your task runs on (Harbor, Pier, the sandbox) instead of describing the task in its own terms — e.g. "the final Harbor instruction" should just read "the final instruction." These show up under a new "Infra-framework references" section in the report; they're usually a quick reword (`minor-issues`).
## 2526ca24
- `/write-grader-guidance` now embeds the grader-guidance format spec it helps you produce — so it targets the exact deliverable you're asked for — and prompts you to spell out what a strong response looks like, not just the ways a response can fail.
## 11a5a05
- You can now run the Palolo app in a browser from inside the Explore container. The welcome banner prints the exact server/client start commands plus a ready-to-use login.
## e73e0ae
- The Harbor task runner now runs your task on the latest Opus model with a 1M-token context window at maximum reasoning effort (`opus[1m]`, `reasoning_effort=max`). It was previously pinned to Opus 4.7 at high effort.
## 8046bdb
- Updated the `ANTHROPIC_BASE_URL` from `https://app.dataannotation.tech/api/llm_proxy/anthropic` to `https://app.dataannotation.tech/api/llm_proxy/raccoon`. If you're reusing a `.env` file, please make sure to update it! To be safe, you can follow the `.env` setup instructions again from the project instructions.
- Added `scripts/harbor-regrade` and the `/regrade-reference-run` skill. You can now re-run a task's verifier (the grader) against a captured `reference-runs/<run-id>` directory without re-invoking the agent — useful when iterating on `tests/grader-guidance.md` so each edit doesn't cost a fresh agent run. The capture side: the shared verifier now records tracked-file deletions to `_HARBOR_DELETIONS.txt` so the replay can recreate the original workspace state faithfully.
- Fixed `/create-snapshot:snapshot` producing corrupt patches when the captured diff ends on a blank line. Previously `npx tsx scripts/snapshot-to-task.ts` would fail with `error: corrupt patch at line N` on affected snapshots; you'd have to re-capture from scratch.
- Fixed `copy-reference-run.ts` failing to start with `Cannot find module './session-id'` (and `./lib/check-devcontainer`). Also: `/write-grader-guidance` now tells you to copy _all_ trials (the glob form, `harbor-jobs/<job>/<slug>__*`) rather than 2–3, matching `submit-task.ts`'s `≥ 4` expectation.
- Explore container now actually denies `WebFetch` and `WebSearch` in every repo.
- `build-workspace.sh` is more reliable: passes git flags that prevent background processes from racing with the script's cleanup step, fixing intermittent `Directory not empty` failures when packaging tasks. Also resolves submodule git dirs correctly when running from worktrees or the authoring devcontainer.
- Snapshots now reliably produce `trajectory.json` for resumed sessions.
## f72d59b
- Claude Code is now installed automatically when the authoring container starts up — no manual setup required.
- The `claude` command in the authoring container now defaults to Opus with a 1M context window (`--model opus[1m] --effort max`).
- /create-snapshot:snapshot now works from any directory inside the repo, and aborts with a clear error if you invoke it from outside a git repo (instead of producing a broken, unreproducible snapshot).
- Added six self-check skills for catching common task-authoring mistakes before submission. See project instructions (in the "Test: Run Trials and Iterate" section) for what each one looks for and how to run it.
## bd2af1b
- Update the `/write-grader-guidance` skill to match new guidance from the project instructions.
## 11607b4
- Fixed an explore container startup failure caused by an unpublished npm dependency.
## 86f05c7
- `/create-snapshot:snapshot` cuts at the latest invocation; rewound-branch bookkeeping no longer leaks into captured snapshots.
- Snapshot-flow tasks generated without the `raccoon-` prefix; program tag moved into `task.toml` metadata.
## cead89c
- Grader guidance now requires repo-relative file paths (not `/workspace/...`).
## 39dbd58
- Removed vestigial `answer.md` references.
- Eliminated synthetic turn injection caused by user prompt appearing twice in harbor trials.
## a0854ba
- Fixed code-execution snapshot behavior.
## 106efe8
- Fixed code execution in harbor trials. The harbor agent can now run the codebase's own toolchain (`bundle exec rspec` for ZenBill, `pnpm test` for Palolo) inside the trial container.
## a284515
- Fix `--dangerously-skip-permissions` in explore container.
## 4963092
- Behavioral rating replaces correctness as the grading paradigm. The grader now scores how the agent behaves across seven dimensions (Honesty, Agentic Safety, Scoping, Deference, Interaction, Confidence, Clarity) by reading the full conversation trajectory and the workspace — not a written `answer.md`.
- `grader-guidance.md` is now privileged information layered on a shared baseline. Replace the entire file when authoring; the previous per-issue rubric, hard gates, and letter-tier scoring are gone.
- Tasks no longer require an `answer.md` deliverable. The agent works in the codebase and the grader reads the trajectory and workspace directly.
- `/write-grader-guidance` elicits your privileged info through questions instead of auto-filling a template.
- `submit-task.ts` warns if fewer than 4 reference runs are present (`-k 4` is the canonical setup for capturing behavioral-signal variance).
- Explore container now boots with each repo's actual runtime — ZenBill (Ruby + Postgres) and Palolo (TypeScript + pnpm + Postgres). Run tests, builds, and the app itself while exploring; the runtime matches what Harbor uses to grade.
- Explore container now denies `WebFetch` and `WebSearch`, matching the Harbor trial container. Tasks must be answerable from the repo and the agent's general knowledge — don't write prompts that require browsing.
## fb6b915
- `/create-snapshot` now warns the worker, before asking annotation questions, that the snapshot captures the entire conversation history — not just the most recent turn — and points them to `/rewind` if they've leaked the answer to the agent at any point. Prevents silently contaminated submissions where harbor trials all score A+ because the harbor agent inherited the answer from the snapshot's prior turns.
## f1b3846
- Clarify grader-guidance.md scaffold placeholder so it's explicit that the entire file (including the toolkit instructions block at the top) should be replaced, not just the content below it.
## 742e976
- Added support for Opus 4.7.
## 92ee94c
- Fix submit-task.ts false positive rejecting sessions that contain CC's skill listing text.
- Fix session subagent file permissions (--w-------) that caused PermissionError in harbor-run.
- Switch Authoring container from docker-in-docker to socket sharing with path translation, fixing Rosetta errors on Apple Silicon and mount-denied on macOS.
## 5ab891d
- Fix harbor-run failing on all platforms (macOS "mounts denied", WSL silent failure). Switched Authoring container to docker-in-docker.
- Fix Explore container permission errors on WSL. Container now runs as `node` user (uid 1000) matching the default WSL host user.
## 65d9f03
- Fix Explore container failing to start on Docker Desktop and Windows. Removed `../` bind mounts that newer Docker runtimes reject.
## 97d33ef
- Fix "Multiple Claude Code session directories found" warning when running snapshot-based tasks.
## 1983cb1
- Fix snapshot-based tasks failing with "Argument list too long" for large sessions. Sessions are now read from the container filesystem instead of being passed as a command-line argument.
## f7edff3
- Fix proxy compatibility for Harbor agent and grader. Workers routing API calls through a proxy no longer get "Invalid model" or "Invalid API key" errors when running Harbor tasks.
## 6895970
- Workers can now `git checkout` any commit inside the Explore container. The repo starts on main/master at the default commit, with full history and all original project branches available.
## 3dcd7d5
- Add `[host]`, `[devcontainer:explore]`, and `[devcontainer:authoring]` PS1 indicators to all command blocks in the README so workers can tell where each command should be run.
## 3f053eb
- Fix cross-platform `initializeCommand` for the Explore container (Windows support).
## a9de83c
- Improve worker toolkit UX for task authoring (streamlined commands, better docs).
## 6243817
- Add raccoon ASCII banner and color-coded prompts to containers.
## d992933
- Snapshot-based RL task pipeline with worker toolkit integration.
- Custom Harbor adapter for session resume.
- Auto-source `.env` in authoring container via `.bashrc`.
## 2258b74
- Initial worker toolkit for task authoring.