Project-2 baseline
This commit is contained in:
@@ -1,14 +1,15 @@
|
||||
---
|
||||
name: detector-offline-verifiability
|
||||
description: |
|
||||
Self-check whether your task makes sense in the no-network sandbox it runs
|
||||
in. The test agent's environment is initialized up front — repo checked
|
||||
out, packages installed — and then runs with no outbound network access, so
|
||||
a good task is offline-completable and offline-verifiable: a competent SWE
|
||||
could do the work AND trust their verification of it entirely from within
|
||||
the repo. Flags tasks whose success criteria live materially outside the
|
||||
sandbox — "speed up our CI/CD pipeline" (verifying needs the live
|
||||
pipeline), "redeploy to prod" (prod doesn't exist in the sandbox),
|
||||
Self-check whether your task uses the internet — it must not. The test
|
||||
agent's environment is initialized up front (repo checked out, packages
|
||||
installed) and a good task is offline-completable and offline-verifiable: a
|
||||
competent SWE could do the work AND trust their verification of it entirely
|
||||
from within the repo. The trial does reach the network and that can't be
|
||||
changed, so this reads your task, never what an agent did in a run.
|
||||
Flags tasks whose success criteria live materially outside the sandbox —
|
||||
"speed up our CI/CD pipeline" (verifying needs the live pipeline),
|
||||
"redeploy to prod" (prod doesn't exist in the sandbox),
|
||||
"migrate from Zendesk to Intercom" (neither service is reachable, so
|
||||
mocks are guesses that likely won't survive real integration), "check the
|
||||
dashboard," published-package behavior. External services as scenario
|
||||
@@ -24,11 +25,26 @@ allowed-tools: Bash, Read, Write
|
||||
|
||||
This skill checks one of your tasks for **offline-verifiability** — whether
|
||||
the ask still makes sense inside the sandbox the test agent actually gets.
|
||||
That sandbox is initialized before the task starts (repo checked out,
|
||||
dependencies installed) and then has **no outbound network access**. So the
|
||||
question is: could a competent SWE complete AND verify your task entirely
|
||||
from within the initialized repo — and would their "it works" actually be
|
||||
trustworthy?
|
||||
|
||||
**The rule is: don't create a task that uses the internet.** It has to be
|
||||
solvable and checkable without one, and the network must never be a central
|
||||
component of the work. The sandbox is initialized before the task starts (repo
|
||||
checked out, dependencies installed), and from there everything that decides
|
||||
the grade should live in the repo. So the question is: could a competent SWE
|
||||
complete AND verify your task entirely from within the initialized repo — and
|
||||
would their "it works" actually be trustworthy?
|
||||
|
||||
**What that rule is not.** The trial does reach the network, and you can't
|
||||
change that — it needs network access to reach the model, so leave
|
||||
`allow_internet` at its default and don't add `network_mode` or
|
||||
`allowed_hosts`. If an agent goes and reads something on the web during one of
|
||||
your runs, that's outside your control and it's fine: the run and the task
|
||||
still stand, provided the task works without the internet and its outcome
|
||||
doesn't rest on what the agent found. And don't write the restriction into the
|
||||
task — a prompt telling the agent it has no internet access, a justification
|
||||
invented for it ("the security team has blocked outbound traffic"), or a rubric
|
||||
that deducts for a lookup are unrealistic constraints that make the task
|
||||
worse. This skill reads your task, never your runs.
|
||||
|
||||
The failure shape to catch: tasks whose *success criteria* live outside the
|
||||
sandbox. "Speed up our CI/CD pipeline" — the pipeline the work would be
|
||||
@@ -39,13 +55,24 @@ services almost certainly won't work at integration time. When a task has
|
||||
this shape, the grade measures how convincingly the agent pantomimes the
|
||||
work, not whether the work is right — and an agent that honestly says "I
|
||||
can't verify this from here" can end up scoring worse than one that
|
||||
confidently fakes it.
|
||||
confidently fakes it. Egress doesn't rescue any of these: reaching a live
|
||||
pipeline or a real SaaS tenant needs credentials and real state, not just a
|
||||
route out.
|
||||
|
||||
What *doesn't* trip this check: external services as scenario dressing (a
|
||||
prompt set at a company that uses Stripe is realism, as long as the graded
|
||||
work and its verification are local), and integrations scoped to a documented
|
||||
protocol slice with a faithful local fake — ideally wired through the fake
|
||||
providers your repo already ships (see `/brainstorm-product-arcs` for the
|
||||
The other shape to watch is an ask whose first step is a fetch — "migrate the
|
||||
cache to Redis" in an app that ships no Redis client, "add TOTP" with no OTP
|
||||
library in any manifest. The install will probably work, because the network
|
||||
is there. That's the problem: your task's correctness then rides on what a
|
||||
registry serves on grading day. Put what the task needs into
|
||||
`environment/workspace.patch` instead, where it's pinned and identical on
|
||||
every run.
|
||||
|
||||
What *doesn't* trip this check: a run in which the agent went online (that's
|
||||
not something your task did), external services as
|
||||
scenario dressing (a prompt set at a company that uses Stripe is realism, as
|
||||
long as the graded work and its verification are local), and integrations
|
||||
scoped to a documented protocol slice with a faithful local fake — ideally
|
||||
wired through the fake providers your repo ships (see `/brainstorm-product-arcs` for the
|
||||
"simulate the protocol, not the product" filter this mirrors).
|
||||
|
||||
**This check is advisory.** Where the line falls is a judgment call — a task
|
||||
|
||||
@@ -2,49 +2,73 @@
|
||||
|
||||
This file is the canonical, context-neutral content for the
|
||||
detector-offline-verifiability detector. It defines the signal (does the task
|
||||
make sense in a no-network sandbox?), the controlling test, the external-
|
||||
dependency shapes to recognize, the verdict enums, and the output schema. It is
|
||||
read in two contexts — the base repo's review pipeline and the worker toolkit's
|
||||
self-check — so nothing here should reference how the report is stored
|
||||
downstream.
|
||||
depend on the network to be done right or graded right?), the controlling
|
||||
test, the external-dependency shapes to recognize, the verdict enums, and the
|
||||
output schema. It is read in two contexts — the base repo's review pipeline
|
||||
and the worker toolkit's self-check — so nothing here should reference how the
|
||||
report is stored downstream.
|
||||
|
||||
## What this detector is for
|
||||
|
||||
Every task runs in a sandbox that is initialized up front — the repo checked
|
||||
out, dependencies installed — and then executes with **no outbound network
|
||||
access**. The agent under test can read, build, run, and test everything inside
|
||||
the workspace, and nothing outside it. A task fits that world when everything
|
||||
is totally verifiable from within the repo: offline-completable and
|
||||
offline-verifiable, because the setup happened before the network went away.
|
||||
**The rule is that a task must not use the internet.** It has to be solvable
|
||||
and checkable without one, and the network must never be a central component of
|
||||
the work: the deliverable is code, config, tests or analysis over what is in
|
||||
the repo, and every load-bearing success criterion is checkable against the
|
||||
repo. That is what "offline-completable and offline-verifiable" mean here, and
|
||||
it is the only thing this detector measures.
|
||||
|
||||
**The rule is about the task, not about a run, and conflating the two is the
|
||||
main way this detector goes wrong.** The sandbox does reach the network — the
|
||||
egress allowlist harbor would need to switch that off does not work on the
|
||||
machines these tasks are built and run on, so every task runs with network
|
||||
access whether or not its author wanted it. An agent that opens a doc page,
|
||||
checks a changelog, or installs something mid-run is therefore doing something
|
||||
the author could not have prevented. That is never a finding here. The
|
||||
questions are whether the task is still solvable without the internet and
|
||||
whether the network is central to it; if it is solvable and the network is not
|
||||
central, the run is fine and goes unremarked.
|
||||
|
||||
The inverse *is* a finding. A prompt that tells the agent it has no internet
|
||||
access, a justification invented for that absence ("the security team has
|
||||
blocked outbound traffic"), or a rubric that deducts for a lookup, are
|
||||
unrealistic constraints the author wrote into the task, and they make it
|
||||
worse.
|
||||
|
||||
Setup installs what the repo's own manifests and lockfiles declare **after
|
||||
`environment/workspace.patch` has been applied** — the image copies the patched
|
||||
workspace in and only *then* runs the dependency install. So a package the
|
||||
author added, upgraded, downgraded or re-pinned in the patch is present in the
|
||||
sandbox, and is never a completability finding; judge the manifests as the
|
||||
patch leaves them, not as the pinned commit left them. What is never installed
|
||||
is a library the ask requires the *agent* to add: acquiring that means `bundle
|
||||
add`, `npm install <pkg>`, `pip install` — a registry fetch, mid-task, after
|
||||
the network is gone.
|
||||
patch leaves them, not as the pinned commit left them. What setup never
|
||||
installs is a library the ask requires the *agent* to add: acquiring that means
|
||||
`bundle add`, `npm install <pkg>`, `pip install` — a registry fetch the task's
|
||||
happy path now hangs on.
|
||||
|
||||
**Do not consider the task's network policy. At all.** `task.toml`'s
|
||||
`allow_internet` / `network_mode` / `allowed_hosts` fields are not about the
|
||||
agent — `allow_internet = true` is scaffold boilerplate carried by essentially
|
||||
every task so the *grading harness* can call its own API. It is not a grant of
|
||||
registry access to the task, and it is out of scope for this detector: do not
|
||||
read those fields, do not mention them in the report, and do not let them move
|
||||
the verdict.
|
||||
`allow_internet` / `network_mode` / `allowed_hosts` fields settle nothing here.
|
||||
`allow_internet = true` is the default every task carries so the *grading
|
||||
harness* can call its own API, and `allow_internet = false` is not an
|
||||
enforceable design tool — the allowlist it would need cannot run on our
|
||||
machines, so a task that sets it gets the whole internet anyway. Leave the
|
||||
field at its default, do not read it, do not mention it in the report, and do
|
||||
not let it move the verdict.
|
||||
|
||||
The corollary matters just as much: **a mid-run install that succeeded is not a
|
||||
clearance.** If the reference runs show the agent fetching the package from a
|
||||
registry, that is evidence the dependency was missing and needed — cite it as
|
||||
support for the finding, never as a reason to soften it. "The runs prove it
|
||||
worked, so this isn't a failure" is the wrong question, answered.
|
||||
clearance.** The network was on, so of course it worked. If the reference runs
|
||||
show the agent fetching the package from a registry, that is the dependency
|
||||
demonstrated, not excused — cite it as support for the finding, never as a
|
||||
reason to soften it. "The runs prove it worked, so this isn't a failure" is the
|
||||
wrong question, answered. Note the asymmetry with the paragraph above, because
|
||||
it is easy to get backwards: run evidence can *corroborate* a finding the
|
||||
manifests already establish, but it can never *create* one. A run that fetched
|
||||
something the ask never required stays unremarked.
|
||||
|
||||
Some task ideas don't really make sense in that world, because a human SWE
|
||||
would need internet access — or access to live systems that only exist outside
|
||||
the sandbox — to really do the task well or to verify the result. The
|
||||
canonical examples:
|
||||
The line an open network does *not* move is where live systems sit. Public
|
||||
documentation is reachable; your CI pipeline, your prod, your customer's SaaS
|
||||
tenant, your dashboard are not, because reaching them needs credentials and
|
||||
accumulated state that exist only outside this sandbox. So a task still
|
||||
doesn't make sense when doing it well, or verifying it, means touching one of
|
||||
those. The canonical examples:
|
||||
|
||||
> Speed up our CI/CD pipeline
|
||||
|
||||
@@ -89,8 +113,8 @@ one that isn't there cannot be carried out here at all.
|
||||
For the task as a whole, ask:
|
||||
|
||||
**Could a competent SWE complete AND verify this task entirely from within the
|
||||
initialized repo — packages already installed, no network — and would their
|
||||
"it works" claim actually be trustworthy?**
|
||||
initialized repo — packages already installed, nothing fetched — and would
|
||||
their "it works" claim actually be trustworthy?**
|
||||
|
||||
Break that into the two halves:
|
||||
|
||||
@@ -99,6 +123,9 @@ Break that into the two halves:
|
||||
something outside — a live pipeline, a running production system, a
|
||||
third-party API, a package registry, data that isn't in the repo?
|
||||
|
||||
The agent is free to *consult* the network while doing it; the test is
|
||||
whether the work can be done without it.
|
||||
|
||||
**This half has a mechanical check, and it is not optional.** List every
|
||||
library, framework, runner, or binary the ask or the rubric's criteria
|
||||
name, then check each against every manifest and lockfile in the repo **as
|
||||
@@ -117,23 +144,27 @@ Break that into the two halves:
|
||||
itself ("the provider is installed and pinned compatibly"). Rebuilding
|
||||
the image wouldn't help, because the dependency was never the repo's.
|
||||
This is the completability failure — flag it, and cite the manifests you
|
||||
read plus the runs that installed the package mid-session.
|
||||
read plus the runs that installed the package mid-session. That those
|
||||
installs succeeded is not a defence: the sandbox has egress, so the fetch
|
||||
was always going to work. The defect is that the ask needs one.
|
||||
- **No — consequential.** The image simply forgot something the repo
|
||||
already depends on: a runner, linter or type checker its own config
|
||||
expects, or a sub-package the build skipped. That is an image-packaging
|
||||
bug on our side, not a defect in the task's design. Do not flag the task
|
||||
for it; record what is missing so the image can be fixed.
|
||||
2. **Offline-verifiable.** Where do the success criteria live? If the honest
|
||||
check for "did this work?" is *outside* the sandbox — watch the pipeline
|
||||
get faster, see the dashboard update, confirm the third-party service
|
||||
accepts the calls, install the published package — then the sandbox can
|
||||
only verify a proxy, and the question is whether that proxy is faithful
|
||||
enough to carry the grade.
|
||||
check for "did this work?" is on a *live system* — watch the pipeline get
|
||||
faster, see the dashboard update, confirm the third-party service accepts
|
||||
the calls, install the published package — then the sandbox can only verify
|
||||
a proxy, and the question is whether that proxy is faithful enough to carry
|
||||
the grade. Egress doesn't help here: these systems need credentials and
|
||||
real state, not just a route out.
|
||||
|
||||
A task passes when both halves stay inside the workspace: the deliverable is
|
||||
code, config, tests, or analysis over what's in the repo, and the rubric's
|
||||
success criteria are checkable against the repo (its test suite, its local
|
||||
mocks and fakes, its own artifacts). A task gets flagged when the success
|
||||
mocks and fakes, its own artifacts). Whether the agent happened to browse the
|
||||
web along the way is irrelevant to that. A task gets flagged when the success
|
||||
criteria live materially outside — external services, live pipelines, prod
|
||||
deploys, third-party SaaS integration, "check the dashboard," published-package
|
||||
behavior — even when the environment itself is perfectly healthy.
|
||||
@@ -201,13 +232,15 @@ Read from `harbor-tasks/<slug>/`:
|
||||
- **Published-artifact behavior.** Release the package and verify it installs
|
||||
from the registry, publish the image, ship the SDK update to consumers —
|
||||
the verifying step is inherently on the other side of the network boundary.
|
||||
- **Missing-at-runtime acquisitions.** The task's happy path requires
|
||||
fetching something after the network is gone: installing a dependency that
|
||||
isn't pre-installed or vendored, pulling a dataset from a URL, cloning
|
||||
another repo, calling a real API for live data. (Setup-time installation is
|
||||
fine only for what a manifest already declares — that got installed before
|
||||
the shutoff. A package the ask tells the agent to add is not setup-time; it
|
||||
is a runtime acquisition, and by then the network is gone.)
|
||||
- **Missing-at-runtime acquisitions.** The task's happy path requires fetching
|
||||
something mid-run: installing a dependency that isn't pre-installed or
|
||||
vendored, pulling a dataset from a URL, cloning another repo, calling a real
|
||||
API for live data. The fetch will probably succeed — that is not the point.
|
||||
The task's correctness then rides on a registry, a URL, or a remote service
|
||||
behaving a particular way on the day it is graded, none of which is pinned,
|
||||
reproducible, or ours. Setup-time installation is the sound version: what a
|
||||
manifest already declares is installed once, into the image, and is the same
|
||||
on every run.
|
||||
|
||||
- **An uninstallable dependency as the deliverable.** The ask names a
|
||||
technology the repo does not carry — migrate the cache to Redis in an app
|
||||
@@ -253,6 +286,11 @@ Read from `harbor-tasks/<slug>/`:
|
||||
- **Hard-but-local work.** Big refactors, gnarly debugging, performance work
|
||||
measured by local benchmarks — difficulty is not an offline-verifiability
|
||||
problem. This detector is orthogonal to how hard the task is.
|
||||
- **A run in which the agent used the internet.** Reading documentation,
|
||||
checking a changelog, searching an error message, even installing something
|
||||
the ask never required — the author cannot switch the network off, so none of
|
||||
this is theirs to answer for. Flag what the *task* needs, never what a run
|
||||
happened to do.
|
||||
|
||||
## Verdict definitions
|
||||
|
||||
@@ -276,9 +314,9 @@ Read from `harbor-tasks/<slug>/`:
|
||||
neither doing the work well nor verifying it can happen in the workspace.
|
||||
*Completability form:* the ask names a technology the repo carries no
|
||||
library for, so step one is a registry fetch that rebuilding the image
|
||||
correctly would not remove. Whether the sandbox happened to permit that
|
||||
fetch is irrelevant and plays no part in the verdict. A human SWE handed this task in this environment would say "I
|
||||
can't actually do or check this from here."
|
||||
correctly would not remove. That the sandbox permits the fetch is irrelevant
|
||||
and plays no part in the verdict — the task's correctness is not supposed to
|
||||
hang on what a registry serves that day.
|
||||
- **`not-applicable`** — nothing to assess: `instruction.md` is missing,
|
||||
empty, or only template/placeholder content, and there is no session
|
||||
history to read an ask from. Re-run once the prompt lands.
|
||||
@@ -351,9 +389,18 @@ integral, a library the image forgot to install is ours to fix.
|
||||
manifests, so state it rather than softening it into a consideration.
|
||||
- **Don't consult the task's network policy.** `allow_internet`,
|
||||
`network_mode` and `allowed_hosts` exist for the grading harness, not the
|
||||
agent. Reading them can only mislead you here: nearly every task allows
|
||||
egress, so weighing it would clear every missing-dependency finding in the
|
||||
corpus. Judge the repo's manifests against the ask and nothing else.
|
||||
agent, and none of them actually closes the sandbox. Reading them can only
|
||||
mislead you here: every task has egress, so weighing it would clear every
|
||||
missing-dependency finding in the corpus. Judge the repo's manifests against
|
||||
the ask and nothing else.
|
||||
- **Don't fault a run for going online, and do flag a task that faults it for
|
||||
you.** A run reaching the web is not grounds for any finding, and never
|
||||
grounds to return a submission — ask only whether the task is solvable
|
||||
without the internet and whether the network is central to it. A prompt or
|
||||
rubric asserting the environment has no internet, inventing a reason for that
|
||||
("the security team has blocked outbound traffic"), or deducting for a
|
||||
lookup, is an unrealistic constraint the author added: say so as a finding on
|
||||
the authored text.
|
||||
- **Don't clear a missing dependency because the framework supports it.**
|
||||
"Rails ships `:redis_cache_store`", "pytest has a coverage plugin" — an
|
||||
adapter existing upstream says nothing about whether the gem or package is
|
||||
|
||||
Reference in New Issue
Block a user