Project-2 baseline
This commit is contained in:
@@ -1,6 +1,6 @@
|
||||
---
|
||||
name: brainstorm-product-arcs
|
||||
description: Invent big, plausible product directions ("arcs") for a source repo and decompose each into a backlog of concrete tasks that can actually be built and verified with no network access. Use when you want task ideas that ladder into a coherent product story instead of one-off commits.
|
||||
description: Invent big, plausible product directions ("arcs") for a source repo and decompose each into a backlog of concrete tasks that can actually be built and verified without depending on the network. Use when you want task ideas that ladder into a coherent product story instead of one-off commits.
|
||||
allowed-tools: Read, Glob, Grep, Bash, Write, Edit, WebSearch, WebFetch, Task
|
||||
---
|
||||
|
||||
@@ -24,7 +24,7 @@ Every arc has to pass all three. Most ideas die on lens 3.
|
||||
|
||||
1. **Plausible** — obviously something _this_ company would do. The test: is it an expansion of what they already do, or a pivot "into making printers"? Ground it in the real product, not the brand.
|
||||
2. **Differentiated** — a sharp, concrete delta against both (a) what the product does _today_ and (b) the _workaround_ a user reaches for now (a named competitor or a manual process).
|
||||
3. **Buildable with no network** — the substance has to be exercisable by a test suite in a sandbox with no internet. This is the gate, and it's the heart of this skill (Step 4).
|
||||
3. **Buildable without the network** — the substance has to be exercisable by a test suite that never leaves the sandbox. This is the gate, and it's the heart of this skill (Step 4).
|
||||
|
||||
## Step 1 — Map the product surface first (go deep; don't guess)
|
||||
|
||||
@@ -67,7 +67,7 @@ Name the real external services and link them. They're load-bearing twice over:
|
||||
|
||||
## Step 4 — The buildability filter ("simulate the protocol, not the product")
|
||||
|
||||
The sandbox that runs a finished task has **no outbound network access** — you can confirm this yourself by running the task under `harbor-run`. So any external service the feature depends on must be faked locally; there's no calling the real API at grade time. The question is never "does it touch the network" — it's whether a _faithful_ local mock is possible.
|
||||
**Don't invent an arc whose tasks use the internet.** The sandbox that runs a finished task does reach the network — that can't be switched off — but a task must never *depend* on it. Its correctness can't ride on a third party being up, unchanged, and reachable on grading day, and reaching a real service needs credentials and live state that don't exist here anyway. So any external service the feature depends on must be faked locally. The question is never "does it touch the network" — it's whether a _faithful_ local mock is possible.
|
||||
|
||||
Grade every feature into one of three buckets:
|
||||
|
||||
|
||||
@@ -104,10 +104,11 @@ Read whatever you need from `harbor-tasks/<slug>/`. The load-bearing artifacts:
|
||||
the shipped `instruction.md` and the resolved guidance file — see the shape
|
||||
section below.
|
||||
- `environment/Dockerfile` plus the workspace's manifests and lockfiles — the
|
||||
static view of what the shipped image can actually do. The execution
|
||||
environment has no network access, so a tool, package, or runtime the ask or
|
||||
its verification depends on must already be present; check for it here even
|
||||
when the runs look quiet.
|
||||
static view of what the shipped image can actually do. A tool, package, or
|
||||
runtime the ask or its verification depends on must already be present — the
|
||||
sandbox does have network access, but a task whose correctness rides on a
|
||||
mid-run fetch isn't reproducible; check for it here even when the runs look
|
||||
quiet.
|
||||
- The snapshot session (`environment/session*`, when the task has one) and the
|
||||
**shipped workspace state** (the declared repo+commit plus
|
||||
`environment/workspace.patch`) — the two halves of the premise check.
|
||||
@@ -167,12 +168,13 @@ task. The agent burns turns reaching a runnable baseline (dependency versions,
|
||||
missing files, broken config, unset env). The prompt never asked for any of it.
|
||||
|
||||
Shape 1 also fires **statically**, even when no reference run visibly fights
|
||||
it: the shipped image can't support what the prompt or rubric requires. The
|
||||
execution environment has no network access, so anything the ask or its
|
||||
verification depends on must already be in the image and lockfiles — a browser
|
||||
the rubric's top tier expects the agent to verify in, a package absent from
|
||||
every manifest and lockfile, a binary that can only be installed from the
|
||||
network. The tell isn't a fight in the runs; it's verification that silently
|
||||
it: the shipped image can't support what the prompt or rubric requires.
|
||||
Anything the ask or its verification depends on should already be in the image
|
||||
and lockfiles — a browser the rubric's top tier expects the agent to verify in,
|
||||
a package absent from every manifest and lockfile, a binary that only a
|
||||
download would provide. The sandbox has network access, so the agent may well
|
||||
fetch what's missing; that it can does not make the image adequate, since the
|
||||
grade then depends on a fetch nothing pins. The tell isn't a fight in the runs; it's verification that silently
|
||||
never happens. Scope this check to capabilities the prompt or rubric actually
|
||||
require or score — not to any tool the agent might conceivably reach for.
|
||||
|
||||
|
||||
@@ -29,8 +29,8 @@ USER_ID=6428…[redacted]
|
||||
|
||||
— the author's own API key, proxy endpoint, and user identity, swept out of
|
||||
their authoring container and checked into the task. Nothing about the task
|
||||
needs these; the agent under test can't use them (no network); and the key is
|
||||
now distributed to every downstream consumer. The same sweep brings in a `.env`
|
||||
needs these; the sandbox has network access, so the agent under test could
|
||||
use them; and the key is now distributed to every downstream consumer. The same sweep brings in a `.env`
|
||||
symlink into the author's home directory, an `.env.bak-*` full of real
|
||||
third-party secrets, or a captured HTTP request with a live bearer token.
|
||||
|
||||
|
||||
@@ -63,7 +63,7 @@ The reduction is checked in order. The reduction is a simple computation over th
|
||||
|
||||
For per-claim verification, **the canonical source is the patched workspace at `harbor-tasks/<slug>/environment/workspace/`**, not `git show <commit>:<path>` against the baseline commit. The test agent sees `git archive <commit>` followed by `environment/workspace.patch` applied — when the patch adds, modifies, or deletes files, the workspace differs from the bare commit. The rubric describes the workspace state (what the test agent reads), so fact-checking must too. Reading the baseline alone produces false `fail` verdicts on every file the patch creates, and false `pass` verdicts on every file the patch modifies.
|
||||
|
||||
The workspace is gitignored. If `harbor-tasks/<slug>/environment/workspace/` is missing, build it with `bash scripts/build-workspace.sh <slug>` before checking claims (in a repo checkout, `harbor-tasks/raccoon-shared/build-workspace.sh <slug> <repo from task.toml> <commit from task.toml>`). The build is idempotent (it `rm -rf`s the workspace before re-exporting), takes seconds, and applies any `workspace.patch` it finds.
|
||||
The workspace is gitignored. If `harbor-tasks/<slug>/environment/workspace/` is missing, build it with `bash scripts/build-workspace.sh <slug>` before checking claims (in a repo checkout, `harbor-tasks/raccoon-private/build-workspace.sh <slug> <repo from task.toml> <commit from task.toml>`). The build is idempotent (it `rm -rf`s the workspace before re-exporting), takes seconds, and applies any `workspace.patch` it finds.
|
||||
|
||||
Read patterns:
|
||||
|
||||
@@ -144,7 +144,7 @@ per-claim record (see schema below) does NOT carry the `claimType` field.
|
||||
|
||||
## Failure modes to handle
|
||||
|
||||
- **Workspace not built and source repo unavailable.** `harbor-tasks/<slug>/environment/workspace/` is missing AND the build script (`scripts/build-workspace.sh` in the toolkit; `harbor-tasks/raccoon-shared/build-workspace.sh` in a repo checkout) can't build it (no source checkout at the toolkit's `repo/` or the repo checkout's `repos/<RepoName>/repo`, and no other local clone with the declared commit). Per-claim verdict for any claim whose cited file lives in that workspace: `unclear` (sub-case: source unavailable). If every claim is `unclear`, the top-level verdict is `not-applicable`. Note the build failure in the body's "Source" line.
|
||||
- **Workspace not built and source repo unavailable.** `harbor-tasks/<slug>/environment/workspace/` is missing AND the build script (`scripts/build-workspace.sh` in the toolkit; `harbor-tasks/raccoon-private/build-workspace.sh` in a repo checkout) can't build it (no source checkout at the toolkit's `repo/` or the repo checkout's `repos/<RepoName>/repo`, and no other local clone with the declared commit). Per-claim verdict for any claim whose cited file lives in that workspace: `unclear` (sub-case: source unavailable). If every claim is `unclear`, the top-level verdict is `not-applicable`. Note the build failure in the body's "Source" line.
|
||||
- **Workspace missing but buildable.** `environment/workspace/` is absent but the source repo and `workspace.patch` are present. Build the workspace before fact-checking — don't return `unclear`, you have everything you need.
|
||||
- **Rubric is empty / template.** Extract step emits `[]`. The save step records `claims: []` and `verdict: not-applicable`.
|
||||
- **`task.toml` missing or unreadable.** Treat as `not-applicable` with an explanatory note in the body.
|
||||
|
||||
@@ -1,14 +1,15 @@
|
||||
---
|
||||
name: detector-offline-verifiability
|
||||
description: |
|
||||
Self-check whether your task makes sense in the no-network sandbox it runs
|
||||
in. The test agent's environment is initialized up front — repo checked
|
||||
out, packages installed — and then runs with no outbound network access, so
|
||||
a good task is offline-completable and offline-verifiable: a competent SWE
|
||||
could do the work AND trust their verification of it entirely from within
|
||||
the repo. Flags tasks whose success criteria live materially outside the
|
||||
sandbox — "speed up our CI/CD pipeline" (verifying needs the live
|
||||
pipeline), "redeploy to prod" (prod doesn't exist in the sandbox),
|
||||
Self-check whether your task uses the internet — it must not. The test
|
||||
agent's environment is initialized up front (repo checked out, packages
|
||||
installed) and a good task is offline-completable and offline-verifiable: a
|
||||
competent SWE could do the work AND trust their verification of it entirely
|
||||
from within the repo. The trial does reach the network and that can't be
|
||||
changed, so this reads your task, never what an agent did in a run.
|
||||
Flags tasks whose success criteria live materially outside the sandbox —
|
||||
"speed up our CI/CD pipeline" (verifying needs the live pipeline),
|
||||
"redeploy to prod" (prod doesn't exist in the sandbox),
|
||||
"migrate from Zendesk to Intercom" (neither service is reachable, so
|
||||
mocks are guesses that likely won't survive real integration), "check the
|
||||
dashboard," published-package behavior. External services as scenario
|
||||
@@ -24,11 +25,26 @@ allowed-tools: Bash, Read, Write
|
||||
|
||||
This skill checks one of your tasks for **offline-verifiability** — whether
|
||||
the ask still makes sense inside the sandbox the test agent actually gets.
|
||||
That sandbox is initialized before the task starts (repo checked out,
|
||||
dependencies installed) and then has **no outbound network access**. So the
|
||||
question is: could a competent SWE complete AND verify your task entirely
|
||||
from within the initialized repo — and would their "it works" actually be
|
||||
trustworthy?
|
||||
|
||||
**The rule is: don't create a task that uses the internet.** It has to be
|
||||
solvable and checkable without one, and the network must never be a central
|
||||
component of the work. The sandbox is initialized before the task starts (repo
|
||||
checked out, dependencies installed), and from there everything that decides
|
||||
the grade should live in the repo. So the question is: could a competent SWE
|
||||
complete AND verify your task entirely from within the initialized repo — and
|
||||
would their "it works" actually be trustworthy?
|
||||
|
||||
**What that rule is not.** The trial does reach the network, and you can't
|
||||
change that — it needs network access to reach the model, so leave
|
||||
`allow_internet` at its default and don't add `network_mode` or
|
||||
`allowed_hosts`. If an agent goes and reads something on the web during one of
|
||||
your runs, that's outside your control and it's fine: the run and the task
|
||||
still stand, provided the task works without the internet and its outcome
|
||||
doesn't rest on what the agent found. And don't write the restriction into the
|
||||
task — a prompt telling the agent it has no internet access, a justification
|
||||
invented for it ("the security team has blocked outbound traffic"), or a rubric
|
||||
that deducts for a lookup are unrealistic constraints that make the task
|
||||
worse. This skill reads your task, never your runs.
|
||||
|
||||
The failure shape to catch: tasks whose *success criteria* live outside the
|
||||
sandbox. "Speed up our CI/CD pipeline" — the pipeline the work would be
|
||||
@@ -39,13 +55,24 @@ services almost certainly won't work at integration time. When a task has
|
||||
this shape, the grade measures how convincingly the agent pantomimes the
|
||||
work, not whether the work is right — and an agent that honestly says "I
|
||||
can't verify this from here" can end up scoring worse than one that
|
||||
confidently fakes it.
|
||||
confidently fakes it. Egress doesn't rescue any of these: reaching a live
|
||||
pipeline or a real SaaS tenant needs credentials and real state, not just a
|
||||
route out.
|
||||
|
||||
What *doesn't* trip this check: external services as scenario dressing (a
|
||||
prompt set at a company that uses Stripe is realism, as long as the graded
|
||||
work and its verification are local), and integrations scoped to a documented
|
||||
protocol slice with a faithful local fake — ideally wired through the fake
|
||||
providers your repo already ships (see `/brainstorm-product-arcs` for the
|
||||
The other shape to watch is an ask whose first step is a fetch — "migrate the
|
||||
cache to Redis" in an app that ships no Redis client, "add TOTP" with no OTP
|
||||
library in any manifest. The install will probably work, because the network
|
||||
is there. That's the problem: your task's correctness then rides on what a
|
||||
registry serves on grading day. Put what the task needs into
|
||||
`environment/workspace.patch` instead, where it's pinned and identical on
|
||||
every run.
|
||||
|
||||
What *doesn't* trip this check: a run in which the agent went online (that's
|
||||
not something your task did), external services as
|
||||
scenario dressing (a prompt set at a company that uses Stripe is realism, as
|
||||
long as the graded work and its verification are local), and integrations
|
||||
scoped to a documented protocol slice with a faithful local fake — ideally
|
||||
wired through the fake providers your repo ships (see `/brainstorm-product-arcs` for the
|
||||
"simulate the protocol, not the product" filter this mirrors).
|
||||
|
||||
**This check is advisory.** Where the line falls is a judgment call — a task
|
||||
|
||||
@@ -2,49 +2,73 @@
|
||||
|
||||
This file is the canonical, context-neutral content for the
|
||||
detector-offline-verifiability detector. It defines the signal (does the task
|
||||
make sense in a no-network sandbox?), the controlling test, the external-
|
||||
dependency shapes to recognize, the verdict enums, and the output schema. It is
|
||||
read in two contexts — the base repo's review pipeline and the worker toolkit's
|
||||
self-check — so nothing here should reference how the report is stored
|
||||
downstream.
|
||||
depend on the network to be done right or graded right?), the controlling
|
||||
test, the external-dependency shapes to recognize, the verdict enums, and the
|
||||
output schema. It is read in two contexts — the base repo's review pipeline
|
||||
and the worker toolkit's self-check — so nothing here should reference how the
|
||||
report is stored downstream.
|
||||
|
||||
## What this detector is for
|
||||
|
||||
Every task runs in a sandbox that is initialized up front — the repo checked
|
||||
out, dependencies installed — and then executes with **no outbound network
|
||||
access**. The agent under test can read, build, run, and test everything inside
|
||||
the workspace, and nothing outside it. A task fits that world when everything
|
||||
is totally verifiable from within the repo: offline-completable and
|
||||
offline-verifiable, because the setup happened before the network went away.
|
||||
**The rule is that a task must not use the internet.** It has to be solvable
|
||||
and checkable without one, and the network must never be a central component of
|
||||
the work: the deliverable is code, config, tests or analysis over what is in
|
||||
the repo, and every load-bearing success criterion is checkable against the
|
||||
repo. That is what "offline-completable and offline-verifiable" mean here, and
|
||||
it is the only thing this detector measures.
|
||||
|
||||
**The rule is about the task, not about a run, and conflating the two is the
|
||||
main way this detector goes wrong.** The sandbox does reach the network — the
|
||||
egress allowlist harbor would need to switch that off does not work on the
|
||||
machines these tasks are built and run on, so every task runs with network
|
||||
access whether or not its author wanted it. An agent that opens a doc page,
|
||||
checks a changelog, or installs something mid-run is therefore doing something
|
||||
the author could not have prevented. That is never a finding here. The
|
||||
questions are whether the task is still solvable without the internet and
|
||||
whether the network is central to it; if it is solvable and the network is not
|
||||
central, the run is fine and goes unremarked.
|
||||
|
||||
The inverse *is* a finding. A prompt that tells the agent it has no internet
|
||||
access, a justification invented for that absence ("the security team has
|
||||
blocked outbound traffic"), or a rubric that deducts for a lookup, are
|
||||
unrealistic constraints the author wrote into the task, and they make it
|
||||
worse.
|
||||
|
||||
Setup installs what the repo's own manifests and lockfiles declare **after
|
||||
`environment/workspace.patch` has been applied** — the image copies the patched
|
||||
workspace in and only *then* runs the dependency install. So a package the
|
||||
author added, upgraded, downgraded or re-pinned in the patch is present in the
|
||||
sandbox, and is never a completability finding; judge the manifests as the
|
||||
patch leaves them, not as the pinned commit left them. What is never installed
|
||||
is a library the ask requires the *agent* to add: acquiring that means `bundle
|
||||
add`, `npm install <pkg>`, `pip install` — a registry fetch, mid-task, after
|
||||
the network is gone.
|
||||
patch leaves them, not as the pinned commit left them. What setup never
|
||||
installs is a library the ask requires the *agent* to add: acquiring that means
|
||||
`bundle add`, `npm install <pkg>`, `pip install` — a registry fetch the task's
|
||||
happy path now hangs on.
|
||||
|
||||
**Do not consider the task's network policy. At all.** `task.toml`'s
|
||||
`allow_internet` / `network_mode` / `allowed_hosts` fields are not about the
|
||||
agent — `allow_internet = true` is scaffold boilerplate carried by essentially
|
||||
every task so the *grading harness* can call its own API. It is not a grant of
|
||||
registry access to the task, and it is out of scope for this detector: do not
|
||||
read those fields, do not mention them in the report, and do not let them move
|
||||
the verdict.
|
||||
`allow_internet` / `network_mode` / `allowed_hosts` fields settle nothing here.
|
||||
`allow_internet = true` is the default every task carries so the *grading
|
||||
harness* can call its own API, and `allow_internet = false` is not an
|
||||
enforceable design tool — the allowlist it would need cannot run on our
|
||||
machines, so a task that sets it gets the whole internet anyway. Leave the
|
||||
field at its default, do not read it, do not mention it in the report, and do
|
||||
not let it move the verdict.
|
||||
|
||||
The corollary matters just as much: **a mid-run install that succeeded is not a
|
||||
clearance.** If the reference runs show the agent fetching the package from a
|
||||
registry, that is evidence the dependency was missing and needed — cite it as
|
||||
support for the finding, never as a reason to soften it. "The runs prove it
|
||||
worked, so this isn't a failure" is the wrong question, answered.
|
||||
clearance.** The network was on, so of course it worked. If the reference runs
|
||||
show the agent fetching the package from a registry, that is the dependency
|
||||
demonstrated, not excused — cite it as support for the finding, never as a
|
||||
reason to soften it. "The runs prove it worked, so this isn't a failure" is the
|
||||
wrong question, answered. Note the asymmetry with the paragraph above, because
|
||||
it is easy to get backwards: run evidence can *corroborate* a finding the
|
||||
manifests already establish, but it can never *create* one. A run that fetched
|
||||
something the ask never required stays unremarked.
|
||||
|
||||
Some task ideas don't really make sense in that world, because a human SWE
|
||||
would need internet access — or access to live systems that only exist outside
|
||||
the sandbox — to really do the task well or to verify the result. The
|
||||
canonical examples:
|
||||
The line an open network does *not* move is where live systems sit. Public
|
||||
documentation is reachable; your CI pipeline, your prod, your customer's SaaS
|
||||
tenant, your dashboard are not, because reaching them needs credentials and
|
||||
accumulated state that exist only outside this sandbox. So a task still
|
||||
doesn't make sense when doing it well, or verifying it, means touching one of
|
||||
those. The canonical examples:
|
||||
|
||||
> Speed up our CI/CD pipeline
|
||||
|
||||
@@ -89,8 +113,8 @@ one that isn't there cannot be carried out here at all.
|
||||
For the task as a whole, ask:
|
||||
|
||||
**Could a competent SWE complete AND verify this task entirely from within the
|
||||
initialized repo — packages already installed, no network — and would their
|
||||
"it works" claim actually be trustworthy?**
|
||||
initialized repo — packages already installed, nothing fetched — and would
|
||||
their "it works" claim actually be trustworthy?**
|
||||
|
||||
Break that into the two halves:
|
||||
|
||||
@@ -99,6 +123,9 @@ Break that into the two halves:
|
||||
something outside — a live pipeline, a running production system, a
|
||||
third-party API, a package registry, data that isn't in the repo?
|
||||
|
||||
The agent is free to *consult* the network while doing it; the test is
|
||||
whether the work can be done without it.
|
||||
|
||||
**This half has a mechanical check, and it is not optional.** List every
|
||||
library, framework, runner, or binary the ask or the rubric's criteria
|
||||
name, then check each against every manifest and lockfile in the repo **as
|
||||
@@ -117,23 +144,27 @@ Break that into the two halves:
|
||||
itself ("the provider is installed and pinned compatibly"). Rebuilding
|
||||
the image wouldn't help, because the dependency was never the repo's.
|
||||
This is the completability failure — flag it, and cite the manifests you
|
||||
read plus the runs that installed the package mid-session.
|
||||
read plus the runs that installed the package mid-session. That those
|
||||
installs succeeded is not a defence: the sandbox has egress, so the fetch
|
||||
was always going to work. The defect is that the ask needs one.
|
||||
- **No — consequential.** The image simply forgot something the repo
|
||||
already depends on: a runner, linter or type checker its own config
|
||||
expects, or a sub-package the build skipped. That is an image-packaging
|
||||
bug on our side, not a defect in the task's design. Do not flag the task
|
||||
for it; record what is missing so the image can be fixed.
|
||||
2. **Offline-verifiable.** Where do the success criteria live? If the honest
|
||||
check for "did this work?" is *outside* the sandbox — watch the pipeline
|
||||
get faster, see the dashboard update, confirm the third-party service
|
||||
accepts the calls, install the published package — then the sandbox can
|
||||
only verify a proxy, and the question is whether that proxy is faithful
|
||||
enough to carry the grade.
|
||||
check for "did this work?" is on a *live system* — watch the pipeline get
|
||||
faster, see the dashboard update, confirm the third-party service accepts
|
||||
the calls, install the published package — then the sandbox can only verify
|
||||
a proxy, and the question is whether that proxy is faithful enough to carry
|
||||
the grade. Egress doesn't help here: these systems need credentials and
|
||||
real state, not just a route out.
|
||||
|
||||
A task passes when both halves stay inside the workspace: the deliverable is
|
||||
code, config, tests, or analysis over what's in the repo, and the rubric's
|
||||
success criteria are checkable against the repo (its test suite, its local
|
||||
mocks and fakes, its own artifacts). A task gets flagged when the success
|
||||
mocks and fakes, its own artifacts). Whether the agent happened to browse the
|
||||
web along the way is irrelevant to that. A task gets flagged when the success
|
||||
criteria live materially outside — external services, live pipelines, prod
|
||||
deploys, third-party SaaS integration, "check the dashboard," published-package
|
||||
behavior — even when the environment itself is perfectly healthy.
|
||||
@@ -201,13 +232,15 @@ Read from `harbor-tasks/<slug>/`:
|
||||
- **Published-artifact behavior.** Release the package and verify it installs
|
||||
from the registry, publish the image, ship the SDK update to consumers —
|
||||
the verifying step is inherently on the other side of the network boundary.
|
||||
- **Missing-at-runtime acquisitions.** The task's happy path requires
|
||||
fetching something after the network is gone: installing a dependency that
|
||||
isn't pre-installed or vendored, pulling a dataset from a URL, cloning
|
||||
another repo, calling a real API for live data. (Setup-time installation is
|
||||
fine only for what a manifest already declares — that got installed before
|
||||
the shutoff. A package the ask tells the agent to add is not setup-time; it
|
||||
is a runtime acquisition, and by then the network is gone.)
|
||||
- **Missing-at-runtime acquisitions.** The task's happy path requires fetching
|
||||
something mid-run: installing a dependency that isn't pre-installed or
|
||||
vendored, pulling a dataset from a URL, cloning another repo, calling a real
|
||||
API for live data. The fetch will probably succeed — that is not the point.
|
||||
The task's correctness then rides on a registry, a URL, or a remote service
|
||||
behaving a particular way on the day it is graded, none of which is pinned,
|
||||
reproducible, or ours. Setup-time installation is the sound version: what a
|
||||
manifest already declares is installed once, into the image, and is the same
|
||||
on every run.
|
||||
|
||||
- **An uninstallable dependency as the deliverable.** The ask names a
|
||||
technology the repo does not carry — migrate the cache to Redis in an app
|
||||
@@ -253,6 +286,11 @@ Read from `harbor-tasks/<slug>/`:
|
||||
- **Hard-but-local work.** Big refactors, gnarly debugging, performance work
|
||||
measured by local benchmarks — difficulty is not an offline-verifiability
|
||||
problem. This detector is orthogonal to how hard the task is.
|
||||
- **A run in which the agent used the internet.** Reading documentation,
|
||||
checking a changelog, searching an error message, even installing something
|
||||
the ask never required — the author cannot switch the network off, so none of
|
||||
this is theirs to answer for. Flag what the *task* needs, never what a run
|
||||
happened to do.
|
||||
|
||||
## Verdict definitions
|
||||
|
||||
@@ -276,9 +314,9 @@ Read from `harbor-tasks/<slug>/`:
|
||||
neither doing the work well nor verifying it can happen in the workspace.
|
||||
*Completability form:* the ask names a technology the repo carries no
|
||||
library for, so step one is a registry fetch that rebuilding the image
|
||||
correctly would not remove. Whether the sandbox happened to permit that
|
||||
fetch is irrelevant and plays no part in the verdict. A human SWE handed this task in this environment would say "I
|
||||
can't actually do or check this from here."
|
||||
correctly would not remove. That the sandbox permits the fetch is irrelevant
|
||||
and plays no part in the verdict — the task's correctness is not supposed to
|
||||
hang on what a registry serves that day.
|
||||
- **`not-applicable`** — nothing to assess: `instruction.md` is missing,
|
||||
empty, or only template/placeholder content, and there is no session
|
||||
history to read an ask from. Re-run once the prompt lands.
|
||||
@@ -351,9 +389,18 @@ integral, a library the image forgot to install is ours to fix.
|
||||
manifests, so state it rather than softening it into a consideration.
|
||||
- **Don't consult the task's network policy.** `allow_internet`,
|
||||
`network_mode` and `allowed_hosts` exist for the grading harness, not the
|
||||
agent. Reading them can only mislead you here: nearly every task allows
|
||||
egress, so weighing it would clear every missing-dependency finding in the
|
||||
corpus. Judge the repo's manifests against the ask and nothing else.
|
||||
agent, and none of them actually closes the sandbox. Reading them can only
|
||||
mislead you here: every task has egress, so weighing it would clear every
|
||||
missing-dependency finding in the corpus. Judge the repo's manifests against
|
||||
the ask and nothing else.
|
||||
- **Don't fault a run for going online, and do flag a task that faults it for
|
||||
you.** A run reaching the web is not grounds for any finding, and never
|
||||
grounds to return a submission — ask only whether the task is solvable
|
||||
without the internet and whether the network is central to it. A prompt or
|
||||
rubric asserting the environment has no internet, inventing a reason for that
|
||||
("the security team has blocked outbound traffic"), or deducting for a
|
||||
lookup, is an unrealistic constraint the author added: say so as a finding on
|
||||
the authored text.
|
||||
- **Don't clear a missing dependency because the framework supports it.**
|
||||
"Rails ships `:redis_cache_store`", "pytest has a coverage plugin" — an
|
||||
adapter existing upstream says nothing about whether the gem or package is
|
||||
|
||||
@@ -52,7 +52,7 @@ The five shapes can co-occur, and any one of them gets verdicted as a leak. Shap
|
||||
|
||||
- **`not-applicable`** — There is no way to decide leakage from this submission. Three triggers:
|
||||
- **No snapshot**: `harbor-tasks/<slug>/environment/session.jsonl` does not exist. The task isn't a snapshot task; there's nothing for the snapshot to leak. Before concluding this, confirm `environment/` truly ships nothing else — no `session/` directory, no packaging-added artifacts.
|
||||
- **Empty snapshot**: `environment/session.jsonl` exists but is empty (zero bytes or whitespace-only). This is `snapshot-to-task`'s designed fallback for one-shot snapshots — when the worker's session held no completed exchange before the prompt (each harness's reader decides what counts as completed), the truncation algorithm has nothing to keep, writes an empty file, and the agent then skips resuming entirely and runs the trial cold from `instruction.md`. An empty `session.jsonl` makes *conversation-text* leakage impossible, but it does NOT make the detector not-applicable on its own — check the rest of the bundle first: subagent sidechain JSONLs under `environment/session/subagents/` and workspace files added by the packaging (`workspace.patch`, `results/` dirs) can carry the answer even when the seeded conversation is empty (Shape 4). Return `not-applicable` only when the session is empty AND no bundled artifact states the rubric's scored answer. (See `plugins/create-snapshot/snapshot-to-task.ts` lines 309–394, and the unit test `truncates one-shot snapshot to empty session` in `snapshot-to-task.test.ts`.) A non-empty `session-full.jsonl` at the slug root in this state is expected and not a sign of over-truncation — it's the unredacted reference copy preserved for human review; the test agent does not see it.
|
||||
- **Empty snapshot**: `environment/session.jsonl` exists but is empty (zero bytes or whitespace-only). This is `snapshot-to-task`'s designed fallback for one-shot snapshots — when the worker's session held no completed exchange before the prompt (each harness's reader decides what counts as completed), the truncation algorithm has nothing to keep, writes an empty file, and the agent then skips resuming entirely and runs the trial cold from `instruction.md`. An empty `session.jsonl` makes *conversation-text* leakage impossible, but it does NOT make the detector not-applicable on its own — check the rest of the bundle first: subagent sidechain JSONLs under `environment/session/subagents/` and workspace files added by the packaging (`workspace.patch`, `results/` dirs) can carry the answer even when the seeded conversation is empty (Shape 4). Return `not-applicable` only when the session is empty AND no bundled artifact states the rubric's scored answer. (See `toolkit/plugins/create-snapshot/snapshot-to-task.ts` lines 309–394, and the unit test `truncates one-shot snapshot to empty session` in `snapshot-to-task.test.ts`.) A non-empty `session-full.jsonl` at the slug root in this state is expected and not a sign of over-truncation — it's the unredacted reference copy preserved for human review; the test agent does not see it.
|
||||
- **No rubric to leak against**: the resolved guidance file is missing, empty, or only contains template/placeholder content (header scaffolding without scored issues, all-TODO stubs, the unmodified default that ships with the task harness). Leakage is *relative* to the rubric's load-bearing claim — if the rubric doesn't yet name what the canonical answer is, the snapshot can't be shown to leak it. We don't try to reverse-engineer the answer from reference runs; that would let us "find" leakage in any thorough snapshot. Wait for the rubric to land, then re-run.
|
||||
- **`clear-leak`** — Shape 1, strong Shape 2, Shape 3, strong Shape 4, or strong Shape 5. Any of:
|
||||
- The snapshot contains explicit content that is the rubric's scored answer. Rubric scores X being identified, snapshot's prior conversation already identifies X. Rubric scores calibrated hedging (the agent should state its uncertainty plainly), snapshot ends with the calibrated hedge. Rubric grades "agent should refuse to close the ticket as expected", snapshot ends with the assistant saying "actually I should keep this open because Y" where Y is the rubric's exact reasoning.
|
||||
|
||||
@@ -56,8 +56,9 @@ Each criterion carries:
|
||||
descriptive enough to be quoted on its own ("names-the-injected-config-key").
|
||||
- **`category`** — one of three values. `primary_intent` marks a requirement at the
|
||||
heart of what the task asks for. `extra_credit` marks a valuable behavior beyond the
|
||||
task's requirements; it can only raise the score, and a response that does not earn
|
||||
it loses nothing. `dodged_bullet` marks a specific failure the response must avoid; a
|
||||
task's requirements; it can only raise the score. A response that earns it partly
|
||||
gets half the effect of a full pass, and a response that does not earn it loses
|
||||
nothing. `dodged_bullet` marks a specific failure the response must avoid; a
|
||||
response that avoids it passes the criterion.
|
||||
- **`severity`** — how heavily a failed criterion weighs in the score: `crux`,
|
||||
`certain_dealbreaker`, `possible_dealbreaker`, or `unlikely_dealbreaker` (displayed
|
||||
@@ -82,7 +83,9 @@ Each criterion carries:
|
||||
- **Phrase requirements positively.** Write "The response should ..." or "The response
|
||||
should avoid ..."; never write "should not". Factual criteria carry their answer key
|
||||
inline, in bold, so the criterion is judgeable without opening another document.
|
||||
- **Keep each criterion self-contained.** Never reference one criterion from another.
|
||||
- **Keep each criterion self-contained.** Its verdict must never depend on another
|
||||
criterion's verdict. An elaboration may name the criterion that primarily assesses
|
||||
a concern, so the same failure is not charged twice; that scope note is fine.
|
||||
A criterion may briefly restate a fact that also lives in `grader-context.md` so
|
||||
that it stands alone; that duplication is intended, and it is the one exception to
|
||||
the source's say-each-thing-once rule.
|
||||
|
||||
@@ -44,6 +44,9 @@ section — the examples there are normative for how criteria interact.
|
||||
the full Task context, Business context, and Ground truth the grader needs.
|
||||
- All eight criterion sections are present, in the standard's order, even when a
|
||||
criterion has no task-specific content (see placeholder discipline below).
|
||||
- `<task-slug>` is the task's own name, with no worker-id prefix. If your task
|
||||
directory is `2QTCWAWMJNJJ-late-fee-rounding`, the title is
|
||||
`# Holistic Rubric — late-fee-rounding`: the title names the task, not its author.
|
||||
|
||||
## The doc must stand alone
|
||||
|
||||
|
||||
Reference in New Issue
Block a user