ren worker folder adding orig, mv new one into root
This commit is contained in:
@@ -1,5 +1,33 @@
|
||||
# Changelog
|
||||
|
||||
## 4ead3c97eb
|
||||
|
||||
- **An atomic rubric carries as many criteria as its holistic rubric needs.** The `/write-atomic-rubric` skill, the `/detector-rubric-form` self-check and the review validator no longer bound the criteria count; write one criterion per scoring-relevant rule and let the count follow the source.
|
||||
- **A task built from a snapshot now scaffolds the full holistic-rubric template.** It used to hand you an empty file with none of the section headings and a stale instruction to write penalties as numeric subtractions; you now get the same template the manual scaffold uses, with the current qualitative penalty phrasing.
|
||||
- **Fixed: writing your atomic rubric no longer marks every earlier detector report stale.** Only `/detector-rubric-coverage` and `/detector-rubric-form` read that file, so the other fifteen reports stay fresh, and the staleness warning now names each report against the input it actually reads.
|
||||
- **New tasks now grade once rather than averaging three samples, so a run finishes sooner and its reward moves around more between runs.** Set `GRADER_SAMPLES=3` — a prefix on `harbor-run`/`harbor-regrade`, or a line in `.env` — to average again; tasks built on an earlier release keep the sample count they were created with.
|
||||
- **`harbor-regrade` can now re-grade all your reference runs in one go.** Pass `--all` instead of a single run directory and it works through every run you have captured, 2 at a time (`--jobs N` to change), which is the usual thing to want after editing your rubric.
|
||||
- **A regrade now files itself, and `--replace` adopts one that would overwrite an existing grade.** An atomic regrade lands in your task's `rubric-regrades/<run>/` with no copying or naming on your part. When a grade is already filed for that run — always the case for a holistic regrade, since the run has one — the result is left in its job directory so you can compare it first, and `--replace` writes it back.
|
||||
- **Fixed: your task images no longer report their own setup as changes your agent made.** The image finished preparing dependencies after taking its baseline snapshot, so files like `Gemfile.lock` and `poetry.lock` reached the grader as part of your agent's work in every run — including runs where the agent changed nothing at all.
|
||||
- **The grader now sees when a task's setup step failed before its checks ran.** Some toolkits run a preparation command, such as a dependency install or a code-generation step, before the test, lint and type-check commands in `tests/test-commands.sh`. When that command failed, the failure reached only the verifier log, so the grader could not tell whether a red check followed from it or from your agent's work, and a holistic rubric had to hedge with an instruction like "if code generation failed, treat the type-check failures as environmental". The failed command, its exit code and the end of its output now appear with the check results the grader reads, together with a note that failures which follow from it are not the response's doing. Nothing is added when setup succeeds. A task you created on an earlier release keeps its own `tests/test.sh`, so this reaches the tasks you create on this toolkit.
|
||||
- **A new task's `task.toml` now says to leave `allow_internet = true` as it is.** Setting it false, or adding `network_mode` or `allowed_hosts`, cuts the grader off from the model API and the trial comes back with no reward.
|
||||
- **Fixed: re-running `run-app` after a failed setup now works on the Python members.** A failed first attempt left a half-built environment that every later run refused to replace, so the member stayed broken until you rebuilt the container. Setup also installs `pkg_resources` now, which some older Python libraries import at start-up.
|
||||
- **A Python member whose dependencies cannot install no longer fails the whole `run-app`.** You get a line saying that member is explore-only in this container, and how to retry it, instead of a wall of installer errors and a bare "setup failed".
|
||||
- **`run-app` with no repo now says which member it picked.** On a polyglot toolkit it quietly started the default member; it now names it and lists the others. The README also uses the `run-app <repo>` form and no longer describes a single-repo layout.
|
||||
- **`run-app` now points you at the error on screen when setup fails.** It used to send you to `run-app --logs`, which is empty at that point because the app has not started yet.
|
||||
- **On the Palolo toolkit, the test suite runs against its own database again, and a bare `pnpm test` needs `--run`.** The image set `CI=true` so plain `vitest` would not sit in watch mode forever, but the app reads that variable in its own code: it was pointing the suite at the main development database and switching off one of the app's environment guards.
|
||||
- **Fixed: on the ZenBill toolkit, the jest check no longer runs out of memory partway through the suite.** Node sizes its heap from the machine's memory rather than from the container's limit, so the run kept growing until the container killed it, and roughly one graded run in ten ended with no frontend test evidence at all. The check now runs with an explicit heap cap, which makes Node collect garbage in time to finish all 134 suites inside the limit. A task you created on an earlier release keeps its own `tests/test-commands.sh`; the new jest line is in `task-shared/test-commands.sh`.
|
||||
- **On the human-essentials toolkit, one flaky spec no longer counts towards your task's test signal.** A request spec asserts two names in an order the app does not guarantee, so the suite went red on roughly one run in ten with nothing actually wrong; it is now skipped.
|
||||
- **On the human-essentials toolkit, a task container's development database now comes up with the app's demo data loaded.** It was created empty, so the app had no account you could sign in with and anything reading it saw no organizations, users or requests.
|
||||
- **On the endsideout toolkit, the minitest check now installs any gems your agent added before it runs.** A response that changed the `Gemfile` left the lockfile needing an install, and the check errored outright instead of running, so a task where adding a gem is the right answer had no test signal at all, and one graded run was marked down for a lockfile the verifier's own setup could not satisfy. The check now runs `bundle check` first and installs only when the lockfile calls for it, so nothing changes when the lockfile is untouched. A task you created on an earlier release keeps its own `tests/test-commands.sh`; the new setup line is in `task-shared/test-commands.sh`.
|
||||
- **On the stocks-in-the-future toolkit, the minitest check now runs with a single test worker.** Parallel workers each restarted the same FactoryBot `level` sequence, which collided with the grade levels several tests hard-code, so tests could fail with "Level has already been taken" on nothing your agent did. The baseline entry that excused this failure now matches on that error message rather than on a whole test class, so a real failure in `ClassroomsControllerTest` counts again. A task you created on an earlier release keeps its own `tests/test-commands.sh`; the new command is in `task-shared/test-commands.sh`.
|
||||
- **On the flaredown toolkit, the rspec check now prepares its own test database and skips two specs that fail on their own.** The container's start-up script creates the test database after its own readiness wait, so a check that started first found no database and could not create one; the check now runs `db:test:prepare`, which creates the database when it is missing. Two specs are excluded because they fail intermittently on an unmodified checkout: a food search spec whose two sample records can produce identical search vectors with nothing to break the tie, and a collection retriever spec whose random sample can draw counts that violate its strict inequality. A red rspec result is now the response's doing. A task you created on an earlier release keeps its own `tests/test-commands.sh`; the new lines are in `task-shared/test-commands.sh`.
|
||||
- **On the zeta-polyglot toolkit, a task image for the zeta-wasabi-platform member no longer switches on two integrations the app's own test setup leaves off.** The image builds the app's `.env` from the repository's `.env.example`, which sets `HT_ENABLED` to true, so any spec that created an internal account sent a request to a service that was not running and raised an unhandled request error. The image now sets `HT_ENABLED=false` and `SAM_ONLINE=false`, which is how the app's own test suite expects to run. A task you created on an earlier release keeps its own `environment/Dockerfile`; the updated image is `task-shared/Dockerfile.zeta-wasabi-platform`.
|
||||
- **Fixed: on the zeta-polyglot toolkit, setting up a second Python member no longer breaks Poetry.** Installing one member's dependencies downgraded a package Poetry itself needs, so every later `run-app` on a Python member failed with `No module named 'packaging.licenses'` and only a container rebuild cleared it.
|
||||
- **Fixed: on the potion-polyglot toolkit, `yarn test` and `npm test` on the potion-app member now run the same suites as the graded jest check.** The member's `package.json` test script passed a path-ignore flag on the command line, and jest treats that flag as a replacement for the ignore list in `jest.config.js` rather than an addition to it, so `yarn test` ran thirty-one suites the config skips, went red, and took far longer than `npx jest`, which stayed green. Agents under test hit this, spent turns untangling it, and some abandoned the command. The script no longer passes the flag, and the member is pinned to that commit. A task you created on an earlier release keeps the potion-app commit it was pinned to. Reported by a worker.
|
||||
- **Fixed: on the potion-polyglot toolkit, seven members' task images now build and come with their dependencies installed.** The `potion-video-processing-devops` image did not build at all, because a Terraform module it uses was not pinned to a version and the module's current release needs a newer provider than the one the repository locks. Six more members built but were allowed to skip a failed dependency install with only a warning, so a task on any of them met a missing-module error at the first import: `browser-extensions`, `lambda-datadog-forwarder`, `potion-voice-utils`, `sentence-split-service`, `microservice-dynamic-screen-recording` and `potion-voice-dataset`. Each now installs what it needs, and a failed install on those members fails the image build so the problem is visible when it happens.
|
||||
- **Fixed: on the potion-polyglot toolkit, unzipping the toolkit on macOS or Windows no longer stops to ask about duplicate files or quietly leaves eight of them out.** The `potion-website` member tracked four poster and video pairs under two spellings of the same filename that differ only in letter case. A filesystem that ignores case cannot hold both, so `unzip` either paused mid-extraction to ask which copy to keep, which reads as a hang on an archive this large, or, when told to overwrite, dropped the second copy without saying so. The duplicate entries are removed and the files themselves are unchanged.
|
||||
|
||||
## 7b6b67ea3d
|
||||
|
||||
- **Fixed: the breezy-complete and zeta toolkits build their containers again.** The Debian release they are built on left long-term support and its package mirror is being retired, so building an Explore container or a task image failed part-way with a "404 Not Found" on a system package; those packages now come from Debian's archive instead.
|
||||
|
||||
Reference in New Issue
Block a user