Project-2 baseline

This commit is contained in:
2026-10-04 21:19:23 -04:00
parent 0f04889edf
commit 213eb3c403
861 changed files with 1710 additions and 3363322 deletions

View File

@@ -1,6 +1,6 @@
---
name: brainstorm-product-arcs
description: Invent big, plausible product directions ("arcs") for a source repo and decompose each into a backlog of concrete tasks that can actually be built and verified with no network access. Use when you want task ideas that ladder into a coherent product story instead of one-off commits.
description: Invent big, plausible product directions ("arcs") for a source repo and decompose each into a backlog of concrete tasks that can actually be built and verified without depending on the network. Use when you want task ideas that ladder into a coherent product story instead of one-off commits.
allowed-tools: Read, Glob, Grep, Bash, Write, Edit, WebSearch, WebFetch, Task
---
@@ -24,7 +24,7 @@ Every arc has to pass all three. Most ideas die on lens 3.
1. **Plausible** — obviously something _this_ company would do. The test: is it an expansion of what they already do, or a pivot "into making printers"? Ground it in the real product, not the brand.
2. **Differentiated** — a sharp, concrete delta against both (a) what the product does _today_ and (b) the _workaround_ a user reaches for now (a named competitor or a manual process).
3. **Buildable with no network** — the substance has to be exercisable by a test suite in a sandbox with no internet. This is the gate, and it's the heart of this skill (Step 4).
3. **Buildable without the network** — the substance has to be exercisable by a test suite that never leaves the sandbox. This is the gate, and it's the heart of this skill (Step 4).
## Step 1 — Map the product surface first (go deep; don't guess)
@@ -67,7 +67,7 @@ Name the real external services and link them. They're load-bearing twice over:
## Step 4 — The buildability filter ("simulate the protocol, not the product")
The sandbox that runs a finished task has **no outbound network access** — you can confirm this yourself by running the task under `harbor-run`. So any external service the feature depends on must be faked locally; there's no calling the real API at grade time. The question is never "does it touch the network" — it's whether a _faithful_ local mock is possible.
**Don't invent an arc whose tasks use the internet.** The sandbox that runs a finished task does reach the network — that can't be switched off — but a task must never *depend* on it. Its correctness can't ride on a third party being up, unchanged, and reachable on grading day, and reaching a real service needs credentials and live state that don't exist here anyway. So any external service the feature depends on must be faked locally. The question is never "does it touch the network" — it's whether a _faithful_ local mock is possible.
Grade every feature into one of three buckets:

View File

@@ -104,10 +104,11 @@ Read whatever you need from `harbor-tasks/<slug>/`. The load-bearing artifacts:
the shipped `instruction.md` and the resolved guidance file — see the shape
section below.
- `environment/Dockerfile` plus the workspace's manifests and lockfiles — the
static view of what the shipped image can actually do. The execution
environment has no network access, so a tool, package, or runtime the ask or
its verification depends on must already be present; check for it here even
when the runs look quiet.
static view of what the shipped image can actually do. A tool, package, or
runtime the ask or its verification depends on must already be present — the
sandbox does have network access, but a task whose correctness rides on a
mid-run fetch isn't reproducible; check for it here even when the runs look
quiet.
- The snapshot session (`environment/session*`, when the task has one) and the
**shipped workspace state** (the declared repo+commit plus
`environment/workspace.patch`) — the two halves of the premise check.
@@ -167,12 +168,13 @@ task. The agent burns turns reaching a runnable baseline (dependency versions,
missing files, broken config, unset env). The prompt never asked for any of it.
Shape 1 also fires **statically**, even when no reference run visibly fights
it: the shipped image can't support what the prompt or rubric requires. The
execution environment has no network access, so anything the ask or its
verification depends on must already be in the image and lockfiles — a browser
the rubric's top tier expects the agent to verify in, a package absent from
every manifest and lockfile, a binary that can only be installed from the
network. The tell isn't a fight in the runs; it's verification that silently
it: the shipped image can't support what the prompt or rubric requires.
Anything the ask or its verification depends on should already be in the image
and lockfiles — a browser the rubric's top tier expects the agent to verify in,
a package absent from every manifest and lockfile, a binary that only a
download would provide. The sandbox has network access, so the agent may well
fetch what's missing; that it can does not make the image adequate, since the
grade then depends on a fetch nothing pins. The tell isn't a fight in the runs; it's verification that silently
never happens. Scope this check to capabilities the prompt or rubric actually
require or score — not to any tool the agent might conceivably reach for.

View File

@@ -29,8 +29,8 @@ USER_ID=6428…[redacted]
— the author's own API key, proxy endpoint, and user identity, swept out of
their authoring container and checked into the task. Nothing about the task
needs these; the agent under test can't use them (no network); and the key is
now distributed to every downstream consumer. The same sweep brings in a `.env`
needs these; the sandbox has network access, so the agent under test could
use them; and the key is now distributed to every downstream consumer. The same sweep brings in a `.env`
symlink into the author's home directory, an `.env.bak-*` full of real
third-party secrets, or a captured HTTP request with a live bearer token.

View File

@@ -63,7 +63,7 @@ The reduction is checked in order. The reduction is a simple computation over th
For per-claim verification, **the canonical source is the patched workspace at `harbor-tasks/<slug>/environment/workspace/`**, not `git show <commit>:<path>` against the baseline commit. The test agent sees `git archive <commit>` followed by `environment/workspace.patch` applied — when the patch adds, modifies, or deletes files, the workspace differs from the bare commit. The rubric describes the workspace state (what the test agent reads), so fact-checking must too. Reading the baseline alone produces false `fail` verdicts on every file the patch creates, and false `pass` verdicts on every file the patch modifies.
The workspace is gitignored. If `harbor-tasks/<slug>/environment/workspace/` is missing, build it with `bash scripts/build-workspace.sh <slug>` before checking claims (in a repo checkout, `harbor-tasks/raccoon-shared/build-workspace.sh <slug> <repo from task.toml> <commit from task.toml>`). The build is idempotent (it `rm -rf`s the workspace before re-exporting), takes seconds, and applies any `workspace.patch` it finds.
The workspace is gitignored. If `harbor-tasks/<slug>/environment/workspace/` is missing, build it with `bash scripts/build-workspace.sh <slug>` before checking claims (in a repo checkout, `harbor-tasks/raccoon-private/build-workspace.sh <slug> <repo from task.toml> <commit from task.toml>`). The build is idempotent (it `rm -rf`s the workspace before re-exporting), takes seconds, and applies any `workspace.patch` it finds.
Read patterns:
@@ -144,7 +144,7 @@ per-claim record (see schema below) does NOT carry the `claimType` field.
## Failure modes to handle
- **Workspace not built and source repo unavailable.** `harbor-tasks/<slug>/environment/workspace/` is missing AND the build script (`scripts/build-workspace.sh` in the toolkit; `harbor-tasks/raccoon-shared/build-workspace.sh` in a repo checkout) can't build it (no source checkout at the toolkit's `repo/` or the repo checkout's `repos/<RepoName>/repo`, and no other local clone with the declared commit). Per-claim verdict for any claim whose cited file lives in that workspace: `unclear` (sub-case: source unavailable). If every claim is `unclear`, the top-level verdict is `not-applicable`. Note the build failure in the body's "Source" line.
- **Workspace not built and source repo unavailable.** `harbor-tasks/<slug>/environment/workspace/` is missing AND the build script (`scripts/build-workspace.sh` in the toolkit; `harbor-tasks/raccoon-private/build-workspace.sh` in a repo checkout) can't build it (no source checkout at the toolkit's `repo/` or the repo checkout's `repos/<RepoName>/repo`, and no other local clone with the declared commit). Per-claim verdict for any claim whose cited file lives in that workspace: `unclear` (sub-case: source unavailable). If every claim is `unclear`, the top-level verdict is `not-applicable`. Note the build failure in the body's "Source" line.
- **Workspace missing but buildable.** `environment/workspace/` is absent but the source repo and `workspace.patch` are present. Build the workspace before fact-checking — don't return `unclear`, you have everything you need.
- **Rubric is empty / template.** Extract step emits `[]`. The save step records `claims: []` and `verdict: not-applicable`.
- **`task.toml` missing or unreadable.** Treat as `not-applicable` with an explanatory note in the body.

View File

@@ -1,14 +1,15 @@
---
name: detector-offline-verifiability
description: |
Self-check whether your task makes sense in the no-network sandbox it runs
in. The test agent's environment is initialized up front — repo checked
out, packages installed — and then runs with no outbound network access, so
a good task is offline-completable and offline-verifiable: a competent SWE
could do the work AND trust their verification of it entirely from within
the repo. Flags tasks whose success criteria live materially outside the
sandbox — "speed up our CI/CD pipeline" (verifying needs the live
pipeline), "redeploy to prod" (prod doesn't exist in the sandbox),
Self-check whether your task uses the internet — it must not. The test
agent's environment is initialized up front (repo checked out, packages
installed) and a good task is offline-completable and offline-verifiable: a
competent SWE could do the work AND trust their verification of it entirely
from within the repo. The trial does reach the network and that can't be
changed, so this reads your task, never what an agent did in a run.
Flags tasks whose success criteria live materially outside the sandbox —
"speed up our CI/CD pipeline" (verifying needs the live pipeline),
"redeploy to prod" (prod doesn't exist in the sandbox),
"migrate from Zendesk to Intercom" (neither service is reachable, so
mocks are guesses that likely won't survive real integration), "check the
dashboard," published-package behavior. External services as scenario
@@ -24,11 +25,26 @@ allowed-tools: Bash, Read, Write
This skill checks one of your tasks for **offline-verifiability** — whether
the ask still makes sense inside the sandbox the test agent actually gets.
That sandbox is initialized before the task starts (repo checked out,
dependencies installed) and then has **no outbound network access**. So the
question is: could a competent SWE complete AND verify your task entirely
from within the initialized repo — and would their "it works" actually be
trustworthy?
**The rule is: don't create a task that uses the internet.** It has to be
solvable and checkable without one, and the network must never be a central
component of the work. The sandbox is initialized before the task starts (repo
checked out, dependencies installed), and from there everything that decides
the grade should live in the repo. So the question is: could a competent SWE
complete AND verify your task entirely from within the initialized repo — and
would their "it works" actually be trustworthy?
**What that rule is not.** The trial does reach the network, and you can't
change that — it needs network access to reach the model, so leave
`allow_internet` at its default and don't add `network_mode` or
`allowed_hosts`. If an agent goes and reads something on the web during one of
your runs, that's outside your control and it's fine: the run and the task
still stand, provided the task works without the internet and its outcome
doesn't rest on what the agent found. And don't write the restriction into the
task — a prompt telling the agent it has no internet access, a justification
invented for it ("the security team has blocked outbound traffic"), or a rubric
that deducts for a lookup are unrealistic constraints that make the task
worse. This skill reads your task, never your runs.
The failure shape to catch: tasks whose *success criteria* live outside the
sandbox. "Speed up our CI/CD pipeline" — the pipeline the work would be
@@ -39,13 +55,24 @@ services almost certainly won't work at integration time. When a task has
this shape, the grade measures how convincingly the agent pantomimes the
work, not whether the work is right — and an agent that honestly says "I
can't verify this from here" can end up scoring worse than one that
confidently fakes it.
confidently fakes it. Egress doesn't rescue any of these: reaching a live
pipeline or a real SaaS tenant needs credentials and real state, not just a
route out.
What *doesn't* trip this check: external services as scenario dressing (a
prompt set at a company that uses Stripe is realism, as long as the graded
work and its verification are local), and integrations scoped to a documented
protocol slice with a faithful local fake — ideally wired through the fake
providers your repo already ships (see `/brainstorm-product-arcs` for the
The other shape to watch is an ask whose first step is a fetch — "migrate the
cache to Redis" in an app that ships no Redis client, "add TOTP" with no OTP
library in any manifest. The install will probably work, because the network
is there. That's the problem: your task's correctness then rides on what a
registry serves on grading day. Put what the task needs into
`environment/workspace.patch` instead, where it's pinned and identical on
every run.
What *doesn't* trip this check: a run in which the agent went online (that's
not something your task did), external services as
scenario dressing (a prompt set at a company that uses Stripe is realism, as
long as the graded work and its verification are local), and integrations
scoped to a documented protocol slice with a faithful local fake — ideally
wired through the fake providers your repo ships (see `/brainstorm-product-arcs` for the
"simulate the protocol, not the product" filter this mirrors).
**This check is advisory.** Where the line falls is a judgment call — a task

View File

@@ -2,49 +2,73 @@
This file is the canonical, context-neutral content for the
detector-offline-verifiability detector. It defines the signal (does the task
make sense in a no-network sandbox?), the controlling test, the external-
dependency shapes to recognize, the verdict enums, and the output schema. It is
read in two contexts — the base repo's review pipeline and the worker toolkit's
self-check — so nothing here should reference how the report is stored
downstream.
depend on the network to be done right or graded right?), the controlling
test, the external-dependency shapes to recognize, the verdict enums, and the
output schema. It is read in two contexts — the base repo's review pipeline
and the worker toolkit's self-check — so nothing here should reference how the
report is stored downstream.
## What this detector is for
Every task runs in a sandbox that is initialized up front — the repo checked
out, dependencies installed — and then executes with **no outbound network
access**. The agent under test can read, build, run, and test everything inside
the workspace, and nothing outside it. A task fits that world when everything
is totally verifiable from within the repo: offline-completable and
offline-verifiable, because the setup happened before the network went away.
**The rule is that a task must not use the internet.** It has to be solvable
and checkable without one, and the network must never be a central component of
the work: the deliverable is code, config, tests or analysis over what is in
the repo, and every load-bearing success criterion is checkable against the
repo. That is what "offline-completable and offline-verifiable" mean here, and
it is the only thing this detector measures.
**The rule is about the task, not about a run, and conflating the two is the
main way this detector goes wrong.** The sandbox does reach the network — the
egress allowlist harbor would need to switch that off does not work on the
machines these tasks are built and run on, so every task runs with network
access whether or not its author wanted it. An agent that opens a doc page,
checks a changelog, or installs something mid-run is therefore doing something
the author could not have prevented. That is never a finding here. The
questions are whether the task is still solvable without the internet and
whether the network is central to it; if it is solvable and the network is not
central, the run is fine and goes unremarked.
The inverse *is* a finding. A prompt that tells the agent it has no internet
access, a justification invented for that absence ("the security team has
blocked outbound traffic"), or a rubric that deducts for a lookup, are
unrealistic constraints the author wrote into the task, and they make it
worse.
Setup installs what the repo's own manifests and lockfiles declare **after
`environment/workspace.patch` has been applied** — the image copies the patched
workspace in and only *then* runs the dependency install. So a package the
author added, upgraded, downgraded or re-pinned in the patch is present in the
sandbox, and is never a completability finding; judge the manifests as the
patch leaves them, not as the pinned commit left them. What is never installed
is a library the ask requires the *agent* to add: acquiring that means `bundle
add`, `npm install <pkg>`, `pip install` — a registry fetch, mid-task, after
the network is gone.
patch leaves them, not as the pinned commit left them. What setup never
installs is a library the ask requires the *agent* to add: acquiring that means
`bundle add`, `npm install <pkg>`, `pip install` — a registry fetch the task's
happy path now hangs on.
**Do not consider the task's network policy. At all.** `task.toml`'s
`allow_internet` / `network_mode` / `allowed_hosts` fields are not about the
agent — `allow_internet = true` is scaffold boilerplate carried by essentially
every task so the *grading harness* can call its own API. It is not a grant of
registry access to the task, and it is out of scope for this detector: do not
read those fields, do not mention them in the report, and do not let them move
the verdict.
`allow_internet` / `network_mode` / `allowed_hosts` fields settle nothing here.
`allow_internet = true` is the default every task carries so the *grading
harness* can call its own API, and `allow_internet = false` is not an
enforceable design tool — the allowlist it would need cannot run on our
machines, so a task that sets it gets the whole internet anyway. Leave the
field at its default, do not read it, do not mention it in the report, and do
not let it move the verdict.
The corollary matters just as much: **a mid-run install that succeeded is not a
clearance.** If the reference runs show the agent fetching the package from a
registry, that is evidence the dependency was missing and needed — cite it as
support for the finding, never as a reason to soften it. "The runs prove it
worked, so this isn't a failure" is the wrong question, answered.
clearance.** The network was on, so of course it worked. If the reference runs
show the agent fetching the package from a registry, that is the dependency
demonstrated, not excused — cite it as support for the finding, never as a
reason to soften it. "The runs prove it worked, so this isn't a failure" is the
wrong question, answered. Note the asymmetry with the paragraph above, because
it is easy to get backwards: run evidence can *corroborate* a finding the
manifests already establish, but it can never *create* one. A run that fetched
something the ask never required stays unremarked.
Some task ideas don't really make sense in that world, because a human SWE
would need internet access — or access to live systems that only exist outside
the sandbox — to really do the task well or to verify the result. The
canonical examples:
The line an open network does *not* move is where live systems sit. Public
documentation is reachable; your CI pipeline, your prod, your customer's SaaS
tenant, your dashboard are not, because reaching them needs credentials and
accumulated state that exist only outside this sandbox. So a task still
doesn't make sense when doing it well, or verifying it, means touching one of
those. The canonical examples:
> Speed up our CI/CD pipeline
@@ -89,8 +113,8 @@ one that isn't there cannot be carried out here at all.
For the task as a whole, ask:
**Could a competent SWE complete AND verify this task entirely from within the
initialized repo — packages already installed, no network — and would their
"it works" claim actually be trustworthy?**
initialized repo — packages already installed, nothing fetched — and would
their "it works" claim actually be trustworthy?**
Break that into the two halves:
@@ -99,6 +123,9 @@ Break that into the two halves:
something outside — a live pipeline, a running production system, a
third-party API, a package registry, data that isn't in the repo?
The agent is free to *consult* the network while doing it; the test is
whether the work can be done without it.
**This half has a mechanical check, and it is not optional.** List every
library, framework, runner, or binary the ask or the rubric's criteria
name, then check each against every manifest and lockfile in the repo **as
@@ -117,23 +144,27 @@ Break that into the two halves:
itself ("the provider is installed and pinned compatibly"). Rebuilding
the image wouldn't help, because the dependency was never the repo's.
This is the completability failure — flag it, and cite the manifests you
read plus the runs that installed the package mid-session.
read plus the runs that installed the package mid-session. That those
installs succeeded is not a defence: the sandbox has egress, so the fetch
was always going to work. The defect is that the ask needs one.
- **No — consequential.** The image simply forgot something the repo
already depends on: a runner, linter or type checker its own config
expects, or a sub-package the build skipped. That is an image-packaging
bug on our side, not a defect in the task's design. Do not flag the task
for it; record what is missing so the image can be fixed.
2. **Offline-verifiable.** Where do the success criteria live? If the honest
check for "did this work?" is *outside* the sandbox — watch the pipeline
get faster, see the dashboard update, confirm the third-party service
accepts the calls, install the published package — then the sandbox can
only verify a proxy, and the question is whether that proxy is faithful
enough to carry the grade.
check for "did this work?" is on a *live system* — watch the pipeline get
faster, see the dashboard update, confirm the third-party service accepts
the calls, install the published package — then the sandbox can only verify
a proxy, and the question is whether that proxy is faithful enough to carry
the grade. Egress doesn't help here: these systems need credentials and
real state, not just a route out.
A task passes when both halves stay inside the workspace: the deliverable is
code, config, tests, or analysis over what's in the repo, and the rubric's
success criteria are checkable against the repo (its test suite, its local
mocks and fakes, its own artifacts). A task gets flagged when the success
mocks and fakes, its own artifacts). Whether the agent happened to browse the
web along the way is irrelevant to that. A task gets flagged when the success
criteria live materially outside — external services, live pipelines, prod
deploys, third-party SaaS integration, "check the dashboard," published-package
behavior — even when the environment itself is perfectly healthy.
@@ -201,13 +232,15 @@ Read from `harbor-tasks/<slug>/`:
- **Published-artifact behavior.** Release the package and verify it installs
from the registry, publish the image, ship the SDK update to consumers —
the verifying step is inherently on the other side of the network boundary.
- **Missing-at-runtime acquisitions.** The task's happy path requires
fetching something after the network is gone: installing a dependency that
isn't pre-installed or vendored, pulling a dataset from a URL, cloning
another repo, calling a real API for live data. (Setup-time installation is
fine only for what a manifest already declares — that got installed before
the shutoff. A package the ask tells the agent to add is not setup-time; it
is a runtime acquisition, and by then the network is gone.)
- **Missing-at-runtime acquisitions.** The task's happy path requires fetching
something mid-run: installing a dependency that isn't pre-installed or
vendored, pulling a dataset from a URL, cloning another repo, calling a real
API for live data. The fetch will probably succeed — that is not the point.
The task's correctness then rides on a registry, a URL, or a remote service
behaving a particular way on the day it is graded, none of which is pinned,
reproducible, or ours. Setup-time installation is the sound version: what a
manifest already declares is installed once, into the image, and is the same
on every run.
- **An uninstallable dependency as the deliverable.** The ask names a
technology the repo does not carry — migrate the cache to Redis in an app
@@ -253,6 +286,11 @@ Read from `harbor-tasks/<slug>/`:
- **Hard-but-local work.** Big refactors, gnarly debugging, performance work
measured by local benchmarks — difficulty is not an offline-verifiability
problem. This detector is orthogonal to how hard the task is.
- **A run in which the agent used the internet.** Reading documentation,
checking a changelog, searching an error message, even installing something
the ask never required — the author cannot switch the network off, so none of
this is theirs to answer for. Flag what the *task* needs, never what a run
happened to do.
## Verdict definitions
@@ -276,9 +314,9 @@ Read from `harbor-tasks/<slug>/`:
neither doing the work well nor verifying it can happen in the workspace.
*Completability form:* the ask names a technology the repo carries no
library for, so step one is a registry fetch that rebuilding the image
correctly would not remove. Whether the sandbox happened to permit that
fetch is irrelevant and plays no part in the verdict. A human SWE handed this task in this environment would say "I
can't actually do or check this from here."
correctly would not remove. That the sandbox permits the fetch is irrelevant
and plays no part in the verdict — the task's correctness is not supposed to
hang on what a registry serves that day.
- **`not-applicable`** — nothing to assess: `instruction.md` is missing,
empty, or only template/placeholder content, and there is no session
history to read an ask from. Re-run once the prompt lands.
@@ -351,9 +389,18 @@ integral, a library the image forgot to install is ours to fix.
manifests, so state it rather than softening it into a consideration.
- **Don't consult the task's network policy.** `allow_internet`,
`network_mode` and `allowed_hosts` exist for the grading harness, not the
agent. Reading them can only mislead you here: nearly every task allows
egress, so weighing it would clear every missing-dependency finding in the
corpus. Judge the repo's manifests against the ask and nothing else.
agent, and none of them actually closes the sandbox. Reading them can only
mislead you here: every task has egress, so weighing it would clear every
missing-dependency finding in the corpus. Judge the repo's manifests against
the ask and nothing else.
- **Don't fault a run for going online, and do flag a task that faults it for
you.** A run reaching the web is not grounds for any finding, and never
grounds to return a submission — ask only whether the task is solvable
without the internet and whether the network is central to it. A prompt or
rubric asserting the environment has no internet, inventing a reason for that
("the security team has blocked outbound traffic"), or deducting for a
lookup, is an unrealistic constraint the author added: say so as a finding on
the authored text.
- **Don't clear a missing dependency because the framework supports it.**
"Rails ships `:redis_cache_store`", "pytest has a coverage plugin" — an
adapter existing upstream says nothing about whether the gem or package is

View File

@@ -52,7 +52,7 @@ The five shapes can co-occur, and any one of them gets verdicted as a leak. Shap
- **`not-applicable`** — There is no way to decide leakage from this submission. Three triggers:
- **No snapshot**: `harbor-tasks/<slug>/environment/session.jsonl` does not exist. The task isn't a snapshot task; there's nothing for the snapshot to leak. Before concluding this, confirm `environment/` truly ships nothing else — no `session/` directory, no packaging-added artifacts.
- **Empty snapshot**: `environment/session.jsonl` exists but is empty (zero bytes or whitespace-only). This is `snapshot-to-task`'s designed fallback for one-shot snapshots — when the worker's session held no completed exchange before the prompt (each harness's reader decides what counts as completed), the truncation algorithm has nothing to keep, writes an empty file, and the agent then skips resuming entirely and runs the trial cold from `instruction.md`. An empty `session.jsonl` makes *conversation-text* leakage impossible, but it does NOT make the detector not-applicable on its own — check the rest of the bundle first: subagent sidechain JSONLs under `environment/session/subagents/` and workspace files added by the packaging (`workspace.patch`, `results/` dirs) can carry the answer even when the seeded conversation is empty (Shape 4). Return `not-applicable` only when the session is empty AND no bundled artifact states the rubric's scored answer. (See `plugins/create-snapshot/snapshot-to-task.ts` lines 309–394, and the unit test `truncates one-shot snapshot to empty session` in `snapshot-to-task.test.ts`.) A non-empty `session-full.jsonl` at the slug root in this state is expected and not a sign of over-truncation — it's the unredacted reference copy preserved for human review; the test agent does not see it.
- **Empty snapshot**: `environment/session.jsonl` exists but is empty (zero bytes or whitespace-only). This is `snapshot-to-task`'s designed fallback for one-shot snapshots — when the worker's session held no completed exchange before the prompt (each harness's reader decides what counts as completed), the truncation algorithm has nothing to keep, writes an empty file, and the agent then skips resuming entirely and runs the trial cold from `instruction.md`. An empty `session.jsonl` makes *conversation-text* leakage impossible, but it does NOT make the detector not-applicable on its own — check the rest of the bundle first: subagent sidechain JSONLs under `environment/session/subagents/` and workspace files added by the packaging (`workspace.patch`, `results/` dirs) can carry the answer even when the seeded conversation is empty (Shape 4). Return `not-applicable` only when the session is empty AND no bundled artifact states the rubric's scored answer. (See `toolkit/plugins/create-snapshot/snapshot-to-task.ts` lines 309–394, and the unit test `truncates one-shot snapshot to empty session` in `snapshot-to-task.test.ts`.) A non-empty `session-full.jsonl` at the slug root in this state is expected and not a sign of over-truncation — it's the unredacted reference copy preserved for human review; the test agent does not see it.
- **No rubric to leak against**: the resolved guidance file is missing, empty, or only contains template/placeholder content (header scaffolding without scored issues, all-TODO stubs, the unmodified default that ships with the task harness). Leakage is *relative* to the rubric's load-bearing claim — if the rubric doesn't yet name what the canonical answer is, the snapshot can't be shown to leak it. We don't try to reverse-engineer the answer from reference runs; that would let us "find" leakage in any thorough snapshot. Wait for the rubric to land, then re-run.
- **`clear-leak`** — Shape 1, strong Shape 2, Shape 3, strong Shape 4, or strong Shape 5. Any of:
- The snapshot contains explicit content that is the rubric's scored answer. Rubric scores X being identified, snapshot's prior conversation already identifies X. Rubric scores calibrated hedging (the agent should state its uncertainty plainly), snapshot ends with the calibrated hedge. Rubric grades "agent should refuse to close the ticket as expected", snapshot ends with the assistant saying "actually I should keep this open because Y" where Y is the rubric's exact reasoning.

View File

@@ -56,8 +56,9 @@ Each criterion carries:
descriptive enough to be quoted on its own ("names-the-injected-config-key").
- **`category`** — one of three values. `primary_intent` marks a requirement at the
heart of what the task asks for. `extra_credit` marks a valuable behavior beyond the
task's requirements; it can only raise the score, and a response that does not earn
it loses nothing. `dodged_bullet` marks a specific failure the response must avoid; a
task's requirements; it can only raise the score. A response that earns it partly
gets half the effect of a full pass, and a response that does not earn it loses
nothing. `dodged_bullet` marks a specific failure the response must avoid; a
response that avoids it passes the criterion.
- **`severity`** — how heavily a failed criterion weighs in the score: `crux`,
`certain_dealbreaker`, `possible_dealbreaker`, or `unlikely_dealbreaker` (displayed
@@ -82,7 +83,9 @@ Each criterion carries:
- **Phrase requirements positively.** Write "The response should ..." or "The response
should avoid ..."; never write "should not". Factual criteria carry their answer key
inline, in bold, so the criterion is judgeable without opening another document.
- **Keep each criterion self-contained.** Never reference one criterion from another.
- **Keep each criterion self-contained.** Its verdict must never depend on another
criterion's verdict. An elaboration may name the criterion that primarily assesses
a concern, so the same failure is not charged twice; that scope note is fine.
A criterion may briefly restate a fact that also lives in `grader-context.md` so
that it stands alone; that duplication is intended, and it is the one exception to
the source's say-each-thing-once rule.

View File

@@ -44,6 +44,9 @@ section — the examples there are normative for how criteria interact.
the full Task context, Business context, and Ground truth the grader needs.
- All eight criterion sections are present, in the standard's order, even when a
criterion has no task-specific content (see placeholder discipline below).
- `<task-slug>` is the task's own name, with no worker-id prefix. If your task
directory is `2QTCWAWMJNJJ-late-fee-rounding`, the title is
`# Holistic Rubric — late-fee-rounding`: the title names the task, not its author.
## The doc must stand alone