Project-2 baseline

This commit is contained in:
2026-10-04 21:19:23 -04:00
parent 0f04889edf
commit 213eb3c403
861 changed files with 1710 additions and 3363322 deletions

View File

@@ -1,14 +1,15 @@
---
name: detector-offline-verifiability
description: |
Self-check whether your task makes sense in the no-network sandbox it runs
in. The test agent's environment is initialized up front — repo checked
out, packages installed — and then runs with no outbound network access, so
a good task is offline-completable and offline-verifiable: a competent SWE
could do the work AND trust their verification of it entirely from within
the repo. Flags tasks whose success criteria live materially outside the
sandbox — "speed up our CI/CD pipeline" (verifying needs the live
pipeline), "redeploy to prod" (prod doesn't exist in the sandbox),
Self-check whether your task uses the internet — it must not. The test
agent's environment is initialized up front (repo checked out, packages
installed) and a good task is offline-completable and offline-verifiable: a
competent SWE could do the work AND trust their verification of it entirely
from within the repo. The trial does reach the network and that can't be
changed, so this reads your task, never what an agent did in a run.
Flags tasks whose success criteria live materially outside the sandbox —
"speed up our CI/CD pipeline" (verifying needs the live pipeline),
"redeploy to prod" (prod doesn't exist in the sandbox),
"migrate from Zendesk to Intercom" (neither service is reachable, so
mocks are guesses that likely won't survive real integration), "check the
dashboard," published-package behavior. External services as scenario
@@ -24,11 +25,26 @@ allowed-tools: Bash, Read, Write
This skill checks one of your tasks for **offline-verifiability** — whether
the ask still makes sense inside the sandbox the test agent actually gets.
That sandbox is initialized before the task starts (repo checked out,
dependencies installed) and then has **no outbound network access**. So the
question is: could a competent SWE complete AND verify your task entirely
from within the initialized repo — and would their "it works" actually be
trustworthy?
**The rule is: don't create a task that uses the internet.** It has to be
solvable and checkable without one, and the network must never be a central
component of the work. The sandbox is initialized before the task starts (repo
checked out, dependencies installed), and from there everything that decides
the grade should live in the repo. So the question is: could a competent SWE
complete AND verify your task entirely from within the initialized repo — and
would their "it works" actually be trustworthy?
**What that rule is not.** The trial does reach the network, and you can't
change that — it needs network access to reach the model, so leave
`allow_internet` at its default and don't add `network_mode` or
`allowed_hosts`. If an agent goes and reads something on the web during one of
your runs, that's outside your control and it's fine: the run and the task
still stand, provided the task works without the internet and its outcome
doesn't rest on what the agent found. And don't write the restriction into the
task — a prompt telling the agent it has no internet access, a justification
invented for it ("the security team has blocked outbound traffic"), or a rubric
that deducts for a lookup are unrealistic constraints that make the task
worse. This skill reads your task, never your runs.
The failure shape to catch: tasks whose *success criteria* live outside the
sandbox. "Speed up our CI/CD pipeline" — the pipeline the work would be
@@ -39,13 +55,24 @@ services almost certainly won't work at integration time. When a task has
this shape, the grade measures how convincingly the agent pantomimes the
work, not whether the work is right — and an agent that honestly says "I
can't verify this from here" can end up scoring worse than one that
confidently fakes it.
confidently fakes it. Egress doesn't rescue any of these: reaching a live
pipeline or a real SaaS tenant needs credentials and real state, not just a
route out.
What *doesn't* trip this check: external services as scenario dressing (a
prompt set at a company that uses Stripe is realism, as long as the graded
work and its verification are local), and integrations scoped to a documented
protocol slice with a faithful local fake — ideally wired through the fake
providers your repo already ships (see `/brainstorm-product-arcs` for the
The other shape to watch is an ask whose first step is a fetch — "migrate the
cache to Redis" in an app that ships no Redis client, "add TOTP" with no OTP
library in any manifest. The install will probably work, because the network
is there. That's the problem: your task's correctness then rides on what a
registry serves on grading day. Put what the task needs into
`environment/workspace.patch` instead, where it's pinned and identical on
every run.
What *doesn't* trip this check: a run in which the agent went online (that's
not something your task did), external services as
scenario dressing (a prompt set at a company that uses Stripe is realism, as
long as the graded work and its verification are local), and integrations
scoped to a documented protocol slice with a faithful local fake — ideally
wired through the fake providers your repo ships (see `/brainstorm-product-arcs` for the
"simulate the protocol, not the product" filter this mirrors).
**This check is advisory.** Where the line falls is a judgment call — a task

View File

@@ -2,49 +2,73 @@
This file is the canonical, context-neutral content for the
detector-offline-verifiability detector. It defines the signal (does the task
make sense in a no-network sandbox?), the controlling test, the external-
dependency shapes to recognize, the verdict enums, and the output schema. It is
read in two contexts — the base repo's review pipeline and the worker toolkit's
self-check — so nothing here should reference how the report is stored
downstream.
depend on the network to be done right or graded right?), the controlling
test, the external-dependency shapes to recognize, the verdict enums, and the
output schema. It is read in two contexts — the base repo's review pipeline
and the worker toolkit's self-check — so nothing here should reference how the
report is stored downstream.
## What this detector is for
Every task runs in a sandbox that is initialized up front — the repo checked
out, dependencies installed — and then executes with **no outbound network
access**. The agent under test can read, build, run, and test everything inside
the workspace, and nothing outside it. A task fits that world when everything
is totally verifiable from within the repo: offline-completable and
offline-verifiable, because the setup happened before the network went away.
**The rule is that a task must not use the internet.** It has to be solvable
and checkable without one, and the network must never be a central component of
the work: the deliverable is code, config, tests or analysis over what is in
the repo, and every load-bearing success criterion is checkable against the
repo. That is what "offline-completable and offline-verifiable" mean here, and
it is the only thing this detector measures.
**The rule is about the task, not about a run, and conflating the two is the
main way this detector goes wrong.** The sandbox does reach the network — the
egress allowlist harbor would need to switch that off does not work on the
machines these tasks are built and run on, so every task runs with network
access whether or not its author wanted it. An agent that opens a doc page,
checks a changelog, or installs something mid-run is therefore doing something
the author could not have prevented. That is never a finding here. The
questions are whether the task is still solvable without the internet and
whether the network is central to it; if it is solvable and the network is not
central, the run is fine and goes unremarked.
The inverse *is* a finding. A prompt that tells the agent it has no internet
access, a justification invented for that absence ("the security team has
blocked outbound traffic"), or a rubric that deducts for a lookup, are
unrealistic constraints the author wrote into the task, and they make it
worse.
Setup installs what the repo's own manifests and lockfiles declare **after
`environment/workspace.patch` has been applied** — the image copies the patched
workspace in and only *then* runs the dependency install. So a package the
author added, upgraded, downgraded or re-pinned in the patch is present in the
sandbox, and is never a completability finding; judge the manifests as the
patch leaves them, not as the pinned commit left them. What is never installed
is a library the ask requires the *agent* to add: acquiring that means `bundle
add`, `npm install <pkg>`, `pip install` — a registry fetch, mid-task, after
the network is gone.
patch leaves them, not as the pinned commit left them. What setup never
installs is a library the ask requires the *agent* to add: acquiring that means
`bundle add`, `npm install <pkg>`, `pip install` — a registry fetch the task's
happy path now hangs on.
**Do not consider the task's network policy. At all.** `task.toml`'s
`allow_internet` / `network_mode` / `allowed_hosts` fields are not about the
agent — `allow_internet = true` is scaffold boilerplate carried by essentially
every task so the *grading harness* can call its own API. It is not a grant of
registry access to the task, and it is out of scope for this detector: do not
read those fields, do not mention them in the report, and do not let them move
the verdict.
`allow_internet` / `network_mode` / `allowed_hosts` fields settle nothing here.
`allow_internet = true` is the default every task carries so the *grading
harness* can call its own API, and `allow_internet = false` is not an
enforceable design tool — the allowlist it would need cannot run on our
machines, so a task that sets it gets the whole internet anyway. Leave the
field at its default, do not read it, do not mention it in the report, and do
not let it move the verdict.
The corollary matters just as much: **a mid-run install that succeeded is not a
clearance.** If the reference runs show the agent fetching the package from a
registry, that is evidence the dependency was missing and needed — cite it as
support for the finding, never as a reason to soften it. "The runs prove it
worked, so this isn't a failure" is the wrong question, answered.
clearance.** The network was on, so of course it worked. If the reference runs
show the agent fetching the package from a registry, that is the dependency
demonstrated, not excused — cite it as support for the finding, never as a
reason to soften it. "The runs prove it worked, so this isn't a failure" is the
wrong question, answered. Note the asymmetry with the paragraph above, because
it is easy to get backwards: run evidence can *corroborate* a finding the
manifests already establish, but it can never *create* one. A run that fetched
something the ask never required stays unremarked.
Some task ideas don't really make sense in that world, because a human SWE
would need internet access — or access to live systems that only exist outside
the sandbox — to really do the task well or to verify the result. The
canonical examples:
The line an open network does *not* move is where live systems sit. Public
documentation is reachable; your CI pipeline, your prod, your customer's SaaS
tenant, your dashboard are not, because reaching them needs credentials and
accumulated state that exist only outside this sandbox. So a task still
doesn't make sense when doing it well, or verifying it, means touching one of
those. The canonical examples:
> Speed up our CI/CD pipeline
@@ -89,8 +113,8 @@ one that isn't there cannot be carried out here at all.
For the task as a whole, ask:
**Could a competent SWE complete AND verify this task entirely from within the
initialized repo — packages already installed, no network — and would their
"it works" claim actually be trustworthy?**
initialized repo — packages already installed, nothing fetched — and would
their "it works" claim actually be trustworthy?**
Break that into the two halves:
@@ -99,6 +123,9 @@ Break that into the two halves:
something outside — a live pipeline, a running production system, a
third-party API, a package registry, data that isn't in the repo?
The agent is free to *consult* the network while doing it; the test is
whether the work can be done without it.
**This half has a mechanical check, and it is not optional.** List every
library, framework, runner, or binary the ask or the rubric's criteria
name, then check each against every manifest and lockfile in the repo **as
@@ -117,23 +144,27 @@ Break that into the two halves:
itself ("the provider is installed and pinned compatibly"). Rebuilding
the image wouldn't help, because the dependency was never the repo's.
This is the completability failure — flag it, and cite the manifests you
read plus the runs that installed the package mid-session.
read plus the runs that installed the package mid-session. That those
installs succeeded is not a defence: the sandbox has egress, so the fetch
was always going to work. The defect is that the ask needs one.
- **No — consequential.** The image simply forgot something the repo
already depends on: a runner, linter or type checker its own config
expects, or a sub-package the build skipped. That is an image-packaging
bug on our side, not a defect in the task's design. Do not flag the task
for it; record what is missing so the image can be fixed.
2. **Offline-verifiable.** Where do the success criteria live? If the honest
check for "did this work?" is *outside* the sandbox — watch the pipeline
get faster, see the dashboard update, confirm the third-party service
accepts the calls, install the published package — then the sandbox can
only verify a proxy, and the question is whether that proxy is faithful
enough to carry the grade.
check for "did this work?" is on a *live system* — watch the pipeline get
faster, see the dashboard update, confirm the third-party service accepts
the calls, install the published package — then the sandbox can only verify
a proxy, and the question is whether that proxy is faithful enough to carry
the grade. Egress doesn't help here: these systems need credentials and
real state, not just a route out.
A task passes when both halves stay inside the workspace: the deliverable is
code, config, tests, or analysis over what's in the repo, and the rubric's
success criteria are checkable against the repo (its test suite, its local
mocks and fakes, its own artifacts). A task gets flagged when the success
mocks and fakes, its own artifacts). Whether the agent happened to browse the
web along the way is irrelevant to that. A task gets flagged when the success
criteria live materially outside — external services, live pipelines, prod
deploys, third-party SaaS integration, "check the dashboard," published-package
behavior — even when the environment itself is perfectly healthy.
@@ -201,13 +232,15 @@ Read from `harbor-tasks/<slug>/`:
- **Published-artifact behavior.** Release the package and verify it installs
from the registry, publish the image, ship the SDK update to consumers —
the verifying step is inherently on the other side of the network boundary.
- **Missing-at-runtime acquisitions.** The task's happy path requires
fetching something after the network is gone: installing a dependency that
isn't pre-installed or vendored, pulling a dataset from a URL, cloning
another repo, calling a real API for live data. (Setup-time installation is
fine only for what a manifest already declares — that got installed before
the shutoff. A package the ask tells the agent to add is not setup-time; it
is a runtime acquisition, and by then the network is gone.)
- **Missing-at-runtime acquisitions.** The task's happy path requires fetching
something mid-run: installing a dependency that isn't pre-installed or
vendored, pulling a dataset from a URL, cloning another repo, calling a real
API for live data. The fetch will probably succeed — that is not the point.
The task's correctness then rides on a registry, a URL, or a remote service
behaving a particular way on the day it is graded, none of which is pinned,
reproducible, or ours. Setup-time installation is the sound version: what a
manifest already declares is installed once, into the image, and is the same
on every run.
- **An uninstallable dependency as the deliverable.** The ask names a
technology the repo does not carry — migrate the cache to Redis in an app
@@ -253,6 +286,11 @@ Read from `harbor-tasks/<slug>/`:
- **Hard-but-local work.** Big refactors, gnarly debugging, performance work
measured by local benchmarks — difficulty is not an offline-verifiability
problem. This detector is orthogonal to how hard the task is.
- **A run in which the agent used the internet.** Reading documentation,
checking a changelog, searching an error message, even installing something
the ask never required — the author cannot switch the network off, so none of
this is theirs to answer for. Flag what the *task* needs, never what a run
happened to do.
## Verdict definitions
@@ -276,9 +314,9 @@ Read from `harbor-tasks/<slug>/`:
neither doing the work well nor verifying it can happen in the workspace.
*Completability form:* the ask names a technology the repo carries no
library for, so step one is a registry fetch that rebuilding the image
correctly would not remove. Whether the sandbox happened to permit that
fetch is irrelevant and plays no part in the verdict. A human SWE handed this task in this environment would say "I
can't actually do or check this from here."
correctly would not remove. That the sandbox permits the fetch is irrelevant
and plays no part in the verdict — the task's correctness is not supposed to
hang on what a registry serves that day.
- **`not-applicable`** — nothing to assess: `instruction.md` is missing,
empty, or only template/placeholder content, and there is no session
history to read an ask from. Re-run once the prompt lands.
@@ -351,9 +389,18 @@ integral, a library the image forgot to install is ours to fix.
manifests, so state it rather than softening it into a consideration.
- **Don't consult the task's network policy.** `allow_internet`,
`network_mode` and `allowed_hosts` exist for the grading harness, not the
agent. Reading them can only mislead you here: nearly every task allows
egress, so weighing it would clear every missing-dependency finding in the
corpus. Judge the repo's manifests against the ask and nothing else.
agent, and none of them actually closes the sandbox. Reading them can only
mislead you here: every task has egress, so weighing it would clear every
missing-dependency finding in the corpus. Judge the repo's manifests against
the ask and nothing else.
- **Don't fault a run for going online, and do flag a task that faults it for
you.** A run reaching the web is not grounds for any finding, and never
grounds to return a submission — ask only whether the task is solvable
without the internet and whether the network is central to it. A prompt or
rubric asserting the environment has no internet, inventing a reason for that
("the security team has blocked outbound traffic"), or deducting for a
lookup, is an unrealistic constraint the author added: say so as a finding on
the authored text.
- **Don't clear a missing dependency because the framework supports it.**
"Rails ships `:redis_cache_store`", "pytest has a coverage plugin" — an
adapter existing upstream says nothing about whether the gem or package is