restored entire zip and config'd
This commit is contained in:
@@ -0,0 +1,114 @@
|
||||
---
|
||||
name: detector-offline-verifiability
|
||||
description: |
|
||||
Self-check whether your task uses the internet — it must not. The test
|
||||
agent's environment is initialized up front (repo checked out, packages
|
||||
installed) and a good task is offline-completable and offline-verifiable: a
|
||||
competent SWE could do the work AND trust their verification of it entirely
|
||||
from within the repo. The trial does reach the network and that can't be
|
||||
changed, so this reads your task, never what an agent did in a run.
|
||||
Flags tasks whose success criteria live materially outside the sandbox —
|
||||
"speed up our CI/CD pipeline" (verifying needs the live pipeline),
|
||||
"redeploy to prod" (prod doesn't exist in the sandbox),
|
||||
"migrate from Zendesk to Intercom" (neither service is reachable, so
|
||||
mocks are guesses that likely won't survive real integration), "check the
|
||||
dashboard," published-package behavior. External services as scenario
|
||||
dressing are fine; protocol-slice integrations against a faithful local
|
||||
fake are fine. Explicitly advisory: every finding is something to
|
||||
consider, never a failure, and it blocks nothing. Reads instruction.md +
|
||||
the resolved holistic rubric (+ the workspace for local fakes); runs
|
||||
before or after reference runs exist.
|
||||
allowed-tools: Bash, Read, Write
|
||||
---
|
||||
|
||||
# Offline-verifiability detector
|
||||
|
||||
This skill checks one of your tasks for **offline-verifiability** — whether
|
||||
the ask still makes sense inside the sandbox the test agent actually gets.
|
||||
|
||||
**The rule is: don't create a task that uses the internet.** It has to be
|
||||
solvable and checkable without one, and the network must never be a central
|
||||
component of the work. The sandbox is initialized before the task starts (repo
|
||||
checked out, dependencies installed), and from there everything that decides
|
||||
the grade should live in the repo. So the question is: could a competent SWE
|
||||
complete AND verify your task entirely from within the initialized repo — and
|
||||
would their "it works" actually be trustworthy?
|
||||
|
||||
**What that rule is not.** The trial does reach the network, and you can't
|
||||
change that — it needs network access to reach the model, so leave
|
||||
`allow_internet` at its default and don't add `network_mode` or
|
||||
`allowed_hosts`. If an agent goes and reads something on the web during one of
|
||||
your runs, that's outside your control and it's fine: the run and the task
|
||||
still stand, provided the task works without the internet and its outcome
|
||||
doesn't rest on what the agent found. And don't write the restriction into the
|
||||
task — a prompt telling the agent it has no internet access, a justification
|
||||
invented for it ("the security team has blocked outbound traffic"), or a rubric
|
||||
that deducts for a lookup are unrealistic constraints that make the task
|
||||
worse. This skill reads your task, never your runs.
|
||||
|
||||
The failure shape to catch: tasks whose *success criteria* live outside the
|
||||
sandbox. "Speed up our CI/CD pipeline" — the pipeline the work would be
|
||||
verified against isn't there. "Redeploy to prod" — there is no prod. "Migrate
|
||||
from Zendesk to Intercom" — the agent can't interact with either service, so
|
||||
it can only mock both ends, and mocks written without ever touching the real
|
||||
services almost certainly won't work at integration time. When a task has
|
||||
this shape, the grade measures how convincingly the agent pantomimes the
|
||||
work, not whether the work is right — and an agent that honestly says "I
|
||||
can't verify this from here" can end up scoring worse than one that
|
||||
confidently fakes it. Egress doesn't rescue any of these: reaching a live
|
||||
pipeline or a real SaaS tenant needs credentials and real state, not just a
|
||||
route out.
|
||||
|
||||
The other shape to watch is an ask whose first step is a fetch — "migrate the
|
||||
cache to Redis" in an app that ships no Redis client, "add TOTP" with no OTP
|
||||
library in any manifest. The install will probably work, because the network
|
||||
is there. That's the problem: your task's correctness then rides on what a
|
||||
registry serves on grading day. Put what the task needs into
|
||||
`environment/workspace.patch` instead, where it's pinned and identical on
|
||||
every run.
|
||||
|
||||
What *doesn't* trip this check: a run in which the agent went online (that's
|
||||
not something your task did), external services as
|
||||
scenario dressing (a prompt set at a company that uses Stripe is realism, as
|
||||
long as the graded work and its verification are local), and integrations
|
||||
scoped to a documented protocol slice with a faithful local fake — ideally
|
||||
wired through the fake providers your repo ships (see `/brainstorm-product-arcs` for the
|
||||
"simulate the protocol, not the product" filter this mirrors).
|
||||
|
||||
**This check is advisory.** Where the line falls is a judgment call — a task
|
||||
can even be deliberately built around recognizing the sandbox's limits, with
|
||||
a holistic rubric that credits saying so. The report exists so you can *consider*
|
||||
where your success criteria live: each finding quotes the passage, says what a
|
||||
human SWE would need the network or a live system for, and offers a rescoping
|
||||
option, so the decision stays yours. Nothing here blocks your submission.
|
||||
|
||||
Read these before deciding:
|
||||
|
||||
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
|
||||
2. `.claude/skills/detector-offline-verifiability/core.md` — the controlling test (offline-completable + offline-verifiable), the external-dependency shapes, the mock-fidelity boundary, what is NOT a finding, verdict enums, and the body schema.
|
||||
|
||||
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
|
||||
|
||||
## Acting on the verdict
|
||||
|
||||
- **`offline-verifiable`** — the work and its verification both live inside
|
||||
the workspace; any external names are scenario context or faithfully
|
||||
faked. Good. Move on.
|
||||
- **`partial`** — the core of your task is offline-completable, but some
|
||||
success criteria lean outside: a "works in prod"-shaped expectation, a
|
||||
local proxy (config parses, unit tests pass) standing in for an external
|
||||
outcome (the pipeline gets faster), or a mock whose fidelity is carrying a
|
||||
lot of the grade. Read each finding and decide: tighten the prompt so it
|
||||
asks for the local slice, point the criterion at your repo's fake provider,
|
||||
or keep the framing deliberately and make sure your holistic rubric grades
|
||||
only what the sandbox can check (crediting honest disclosure of the rest).
|
||||
- **`not-offline-verifiable`** — the system your task operates on (pipeline,
|
||||
prod, third-party service) isn't in the sandbox and can't be faithfully
|
||||
faked, so neither doing the work well nor verifying it can happen there.
|
||||
Consider the rescoping option in each finding: extract the protocol slice
|
||||
and build an adversarial local mock for it, reframe the ask as an
|
||||
assessment or plan graded on repo evidence, or pick a different behavior to
|
||||
test. If you believe the task works as-is, that's your call — but make sure
|
||||
the holistic rubric never asks the grader (or the agent) for a verification
|
||||
the sandbox cannot perform.
|
||||
- **`not-applicable`** — there's no prompt to assess yet. Draft it first.
|
||||
@@ -0,0 +1,472 @@
|
||||
# Offline-verifiability detector — core
|
||||
|
||||
This file is the canonical, context-neutral content for the
|
||||
detector-offline-verifiability detector. It defines the signal (does the task
|
||||
depend on the network to be done right or graded right?), the controlling
|
||||
test, the external-dependency shapes to recognize, the verdict enums, and the
|
||||
output schema. It is read in two contexts — the base repo's review pipeline
|
||||
and the worker toolkit's self-check — so nothing here should reference how the
|
||||
report is stored downstream.
|
||||
|
||||
## What this detector is for
|
||||
|
||||
**The rule is that a task must not use the internet.** It has to be solvable
|
||||
and checkable without one, and the network must never be a central component of
|
||||
the work: the deliverable is code, config, tests or analysis over what is in
|
||||
the repo, and every load-bearing success criterion is checkable against the
|
||||
repo. That is what "offline-completable and offline-verifiable" mean here, and
|
||||
it is the only thing this detector measures.
|
||||
|
||||
**The rule is about the task, not about a run, and conflating the two is the
|
||||
main way this detector goes wrong.** The sandbox does reach the network — the
|
||||
egress allowlist harbor would need to switch that off does not work on the
|
||||
machines these tasks are built and run on, so every task runs with network
|
||||
access whether or not its author wanted it. An agent that opens a doc page,
|
||||
checks a changelog, or installs something mid-run is therefore doing something
|
||||
the author could not have prevented. That is never a finding here. The
|
||||
questions are whether the task is still solvable without the internet and
|
||||
whether the network is central to it; if it is solvable and the network is not
|
||||
central, the run is fine and goes unremarked.
|
||||
|
||||
The inverse *is* a finding. A prompt that tells the agent it has no internet
|
||||
access, a justification invented for that absence ("the security team has
|
||||
blocked outbound traffic"), or a rubric that deducts for a lookup, are
|
||||
unrealistic constraints the author wrote into the task, and they make it
|
||||
worse.
|
||||
|
||||
Setup installs what the repo's own manifests and lockfiles declare **after
|
||||
`environment/workspace.patch` has been applied** — the image copies the patched
|
||||
workspace in and only *then* runs the dependency install. So a package the
|
||||
author added, upgraded, downgraded or re-pinned in the patch is present in the
|
||||
sandbox, and is never a completability finding; judge the manifests as the
|
||||
patch leaves them, not as the pinned commit left them. What setup never
|
||||
installs is a library the ask requires the *agent* to add: acquiring that means
|
||||
`bundle add`, `npm install <pkg>`, `pip install` — a registry fetch the task's
|
||||
happy path now hangs on.
|
||||
|
||||
**Do not consider the task's network policy. At all.** `task.toml`'s
|
||||
`allow_internet` / `network_mode` / `allowed_hosts` fields settle nothing here.
|
||||
`allow_internet = true` is the default every task carries so the *grading
|
||||
harness* can call its own API, and `allow_internet = false` is not an
|
||||
enforceable design tool — the allowlist it would need cannot run on our
|
||||
machines, so a task that sets it gets the whole internet anyway. Leave the
|
||||
field at its default, do not read it, do not mention it in the report, and do
|
||||
not let it move the verdict.
|
||||
|
||||
The corollary matters just as much: **a mid-run install that succeeded is not a
|
||||
clearance.** The network was on, so of course it worked. If the reference runs
|
||||
show the agent fetching the package from a registry, that is the dependency
|
||||
demonstrated, not excused — cite it as support for the finding, never as a
|
||||
reason to soften it. "The runs prove it worked, so this isn't a failure" is the
|
||||
wrong question, answered. Note the asymmetry with the paragraph above, because
|
||||
it is easy to get backwards: run evidence can *corroborate* a finding the
|
||||
manifests already establish, but it can never *create* one. A run that fetched
|
||||
something the ask never required stays unremarked.
|
||||
|
||||
The line an open network does *not* move is where live systems sit. Public
|
||||
documentation is reachable; your CI pipeline, your prod, your customer's SaaS
|
||||
tenant, your dashboard are not, because reaching them needs credentials and
|
||||
accumulated state that exist only outside this sandbox. So a task still
|
||||
doesn't make sense when doing it well, or verifying it, means touching one of
|
||||
those. The canonical examples:
|
||||
|
||||
> Speed up our CI/CD pipeline
|
||||
|
||||
— you need access to that pipeline to verify your work. The pipeline's actual
|
||||
runtime, caching behavior, and bottlenecks live on an external system the
|
||||
sandbox doesn't have; the agent can edit config files but never observe whether
|
||||
anything got faster.
|
||||
|
||||
> Redeploy to prod
|
||||
|
||||
— prod doesn't exist in the sandbox environment. There is nothing to deploy
|
||||
to, so the "success" the prompt asks for cannot occur, let alone be checked.
|
||||
|
||||
> Migrate from Zendesk to Intercom
|
||||
|
||||
— the agent can't interact with either service, so it can't test the
|
||||
migration end-to-end; it's just mocking things out, and mocks written without
|
||||
ever touching the real services almost certainly won't work at integration
|
||||
time. The part that makes the task hard — does it actually work against the
|
||||
real thing? — is exactly the part the sandbox can't answer.
|
||||
|
||||
When a task has this shape, the reference runs and the grade measure how
|
||||
convincingly the agent *pantomimes* the work, not whether the work is right.
|
||||
The verifier can't check the thing that matters, the rubric drifts toward
|
||||
style points, and an agent that (correctly) says "I can't verify this from
|
||||
here" may score worse than one that confidently fakes it.
|
||||
|
||||
**The verifiability half of this detector is advisory.** Whether a task's
|
||||
success criteria live too far outside the sandbox is a judgment call — most real tasks mention external services *somewhere*, and a scenario
|
||||
can legitimately be about recognizing the limits of what's verifiable. A
|
||||
flagged verdict means "here is something to consider about where this task's
|
||||
success criteria live," never "this task is invalid." The author may have
|
||||
deliberately scoped the graded substance to the local slice, and the flag is
|
||||
the prompt to confirm that scoping is real.
|
||||
|
||||
The completability half is not a judgment call. Whether a library the ask
|
||||
requires appears in any manifest is a fact you check, and a task that needs
|
||||
one that isn't there cannot be carried out here at all.
|
||||
|
||||
## The controlling test
|
||||
|
||||
For the task as a whole, ask:
|
||||
|
||||
**Could a competent SWE complete AND verify this task entirely from within the
|
||||
initialized repo — packages already installed, nothing fetched — and would
|
||||
their "it works" claim actually be trustworthy?**
|
||||
|
||||
Break that into the two halves:
|
||||
|
||||
1. **Offline-completable.** Is everything the prompt asks for buildable from
|
||||
what's in the workspace? Or does doing the work well require reaching
|
||||
something outside — a live pipeline, a running production system, a
|
||||
third-party API, a package registry, data that isn't in the repo?
|
||||
|
||||
The agent is free to *consult* the network while doing it; the test is
|
||||
whether the work can be done without it.
|
||||
|
||||
**This half has a mechanical check, and it is not optional.** List every
|
||||
library, framework, runner, or binary the ask or the rubric's criteria
|
||||
name, then check each against every manifest and lockfile in the repo **as
|
||||
the workspace patch leaves it** (`Gemfile`/`Gemfile.lock`, `package.json` + its lockfile,
|
||||
`pyproject.toml`/`requirements*.txt`/`uv.lock`, `go.mod`, the Dockerfile).
|
||||
Read the files — never settle this from knowledge of what the framework
|
||||
supports. When a name is absent from all of them, the deciding question is
|
||||
**integral or consequential**:
|
||||
|
||||
> If we rebuilt the image correctly, would this task still need the
|
||||
> network?
|
||||
|
||||
- **Yes — integral.** The repo has no library for the thing the ask names:
|
||||
migrate to Redis Cluster with no Redis client, add TOTP with no OTP gem,
|
||||
write BDD features with no BDD runner, or a rubric that grades the fetch
|
||||
itself ("the provider is installed and pinned compatibly"). Rebuilding
|
||||
the image wouldn't help, because the dependency was never the repo's.
|
||||
This is the completability failure — flag it, and cite the manifests you
|
||||
read plus the runs that installed the package mid-session. That those
|
||||
installs succeeded is not a defence: the sandbox has egress, so the fetch
|
||||
was always going to work. The defect is that the ask needs one.
|
||||
- **No — consequential.** The image simply forgot something the repo
|
||||
already depends on: a runner, linter or type checker its own config
|
||||
expects, or a sub-package the build skipped. That is an image-packaging
|
||||
bug on our side, not a defect in the task's design. Do not flag the task
|
||||
for it; record what is missing so the image can be fixed.
|
||||
2. **Offline-verifiable.** Where do the success criteria live? If the honest
|
||||
check for "did this work?" is on a *live system* — watch the pipeline get
|
||||
faster, see the dashboard update, confirm the third-party service accepts
|
||||
the calls, install the published package — then the sandbox can only verify
|
||||
a proxy, and the question is whether that proxy is faithful enough to carry
|
||||
the grade. Egress doesn't help here: these systems need credentials and
|
||||
real state, not just a route out.
|
||||
|
||||
A task passes when both halves stay inside the workspace: the deliverable is
|
||||
code, config, tests, or analysis over what's in the repo, and the rubric's
|
||||
success criteria are checkable against the repo (its test suite, its local
|
||||
mocks and fakes, its own artifacts). Whether the agent happened to browse the
|
||||
web along the way is irrelevant to that. A task gets flagged when the success
|
||||
criteria live materially outside — external services, live pipelines, prod
|
||||
deploys, third-party SaaS integration, "check the dashboard," published-package
|
||||
behavior — even when the environment itself is perfectly healthy.
|
||||
|
||||
**Mocks are the boundary case, and fidelity is the question.** External
|
||||
dependencies faked through a faithful local mock — a documented protocol
|
||||
(file formats, webhook signatures, return codes) simulated the way the repo
|
||||
already fakes its providers — keep a task offline-verifiable: the hard work is
|
||||
on the repo's side and the mock exercises it honestly. The flag condition is a
|
||||
mock that has to *invent* the external side because nobody can check it: an
|
||||
undocumented or proprietary behavior, a product rather than a protocol, or an
|
||||
integration whose entire difficulty is "does the real service accept this?"
|
||||
A useful rule of thumb: if the mock's spec could be written straight from
|
||||
public documentation and a correct integration against the mock would also be
|
||||
correct against the real service, the mock carries the verification; if the
|
||||
mock is a guess about the real thing, it doesn't.
|
||||
|
||||
## Inputs
|
||||
|
||||
Read from `harbor-tasks/<slug>/`:
|
||||
|
||||
- `instruction.md` — the prompt the agent under test receives. The primary
|
||||
surface: what is the agent actually being asked to deliver, and what would
|
||||
"done, and correct" mean for that ask? For a snapshot / multi-turn task,
|
||||
also read the standing user turns in the session history
|
||||
(`environment/session.jsonl` or `session-full.jsonl`) — an ask that arrives
|
||||
in a prior turn binds the agent the same way.
|
||||
- The grader guidance — context for what is actually verified. Resolve which
|
||||
guidance file the grader actually reads (`bash scripts/guidance-target.sh
|
||||
<slug>` — the worker shell's guidance-target resolution) and read that
|
||||
file, never its sibling. This is
|
||||
where the flag is confirmed or cleared: a prompt that *mentions* deployment
|
||||
can still be graded entirely on local substance, and a local-sounding prompt
|
||||
can hide a rubric criterion that only a live system could check ("the
|
||||
webhook must be accepted by the provider"). Ask of each load-bearing
|
||||
criterion: what would the grader look at, and is it in the workspace?
|
||||
- `environment/workspace.patch` and the workspace — context for whether the
|
||||
external side is actually represented locally: an existing fake provider,
|
||||
fixtures, a stub server, seeded data. A prompt naming a third-party service
|
||||
reads very differently when the repo ships a faithful fake of it.
|
||||
- `reference-runs/*/grade.md` — not required, but a useful cross-check when
|
||||
present: runs where the agent had to invent mock behavior wholesale, spent
|
||||
its effort simulating an absent system, or was penalized for saying it
|
||||
couldn't verify something the sandbox genuinely can't verify, all
|
||||
corroborate the flag.
|
||||
|
||||
## External-dependency shapes to look for
|
||||
|
||||
- **Live infrastructure as the subject.** The deliverable is an operation on
|
||||
a system that exists only outside the sandbox: speed up the CI/CD pipeline,
|
||||
redeploy to prod, rotate the certs, fix the DNS, tune the production
|
||||
database. The workspace may contain the *config* for these systems, but the
|
||||
success criteria — the pipeline runs faster, the deploy succeeds — are
|
||||
observable only on the real thing.
|
||||
- **Third-party SaaS integration as the deliverable.** Migrate from Zendesk
|
||||
to Intercom, integrate the new payment provider, sync with the CRM — where
|
||||
the graded outcome is end-to-end behavior against services the agent can't
|
||||
reach, and no faithful local fake exists or could exist. (A protocol-slice
|
||||
task against a documented contract with a faithful adversarial mock is the
|
||||
acceptable version — see the controlling test.)
|
||||
- **Success criteria that name an external observation.** "Check the
|
||||
dashboard," "confirm the metrics improve," "verify the alert fires in
|
||||
PagerDuty," "make sure the docs site renders" — the rubric or prompt defines
|
||||
done-ness as something seen on a system that isn't in the workspace.
|
||||
- **Published-artifact behavior.** Release the package and verify it installs
|
||||
from the registry, publish the image, ship the SDK update to consumers —
|
||||
the verifying step is inherently on the other side of the network boundary.
|
||||
- **Missing-at-runtime acquisitions.** The task's happy path requires fetching
|
||||
something mid-run: installing a dependency that isn't pre-installed or
|
||||
vendored, pulling a dataset from a URL, cloning another repo, calling a real
|
||||
API for live data. The fetch will probably succeed — that is not the point.
|
||||
The task's correctness then rides on a registry, a URL, or a remote service
|
||||
behaving a particular way on the day it is graded, none of which is pinned,
|
||||
reproducible, or ours. Setup-time installation is the sound version: what a
|
||||
manifest already declares is installed once, into the image, and is the same
|
||||
on every run.
|
||||
|
||||
- **An uninstallable dependency as the deliverable.** The ask names a
|
||||
technology the repo does not carry — migrate the cache to Redis in an app
|
||||
whose only cache gem is `solid_cache`, add TOTP and lockout to an app
|
||||
shipping no auth library, add coverage or BDD tooling that appears in no
|
||||
manifest — and the rubric grades the result as installed and working. The
|
||||
graded substance can look entirely local (config, key shapes, call sites)
|
||||
while step one is an impossible `bundle add`. A first-party framework
|
||||
adapter still needs its gem: "documented upstream" is not "present here".
|
||||
- **External knowledge as the graded substance.** The rubric's success hinges
|
||||
on looking up volatile external state — current API behavior of a live
|
||||
provider, today's prices, the latest version of a service's schema — that
|
||||
isn't captured in the workspace and can't be derived from it.
|
||||
|
||||
## What is NOT a finding
|
||||
|
||||
- **External services as scenario dressing.** A prompt set at a company that
|
||||
uses Stripe, Zendesk, and AWS is realism. The question is where the *graded
|
||||
work and its verification* happen — if the deliverable is repo code and the
|
||||
rubric checks repo behavior, the named services are backdrop, not
|
||||
dependencies.
|
||||
- **Protocol-slice integrations with a faithful local fake.** Build the
|
||||
webhook verifier, parse the provider's documented file format, reconcile
|
||||
against the seeded fixture service — especially when the repo already fakes
|
||||
that provider and the task extends the existing seam. That is the sanctioned
|
||||
way to do external-facing work offline.
|
||||
- **Deploy/CI config work graded on local substance.** Editing a CI config or
|
||||
a deploy manifest where the rubric checks properties verifiable in the
|
||||
workspace — the config parses, the referenced scripts exist and run, the
|
||||
documented invariants hold — is bounded. It may still merit `partial` when
|
||||
the *real* success criterion (the pipeline actually gets faster) is external
|
||||
and the local checks are a thin proxy; say which.
|
||||
- **Assessments and plans about external systems, graded on repo evidence.**
|
||||
"Review our migration plan," "assess what moving to Intercom would take" —
|
||||
where the deliverable is analysis whose load-bearing claims are checkable
|
||||
against the repo. A *plan* for external work is offline-verifiable; only
|
||||
*executing and confirming* the external work isn't.
|
||||
- **Tasks deliberately about recognizing the limit.** A scenario can be built
|
||||
so that the right behavior is to say "this part can't be verified from
|
||||
here" — and the rubric credits exactly that. If the grader guidance treats
|
||||
the boundary honestly (credits disclosure, doesn't demand the impossible
|
||||
verification), the external dependency is the task working as designed.
|
||||
- **Hard-but-local work.** Big refactors, gnarly debugging, performance work
|
||||
measured by local benchmarks — difficulty is not an offline-verifiability
|
||||
problem. This detector is orthogonal to how hard the task is.
|
||||
- **A run in which the agent used the internet.** Reading documentation,
|
||||
checking a changelog, searching an error message, even installing something
|
||||
the ask never required — the author cannot switch the network off, so none of
|
||||
this is theirs to answer for. Flag what the *task* needs, never what a run
|
||||
happened to do.
|
||||
|
||||
## Verdict definitions
|
||||
|
||||
- **`offline-verifiable`** — the controlling test passes: a competent SWE
|
||||
could complete the ask and trust their own verification of it entirely
|
||||
within the initialized workspace. External services, if named, are scenario
|
||||
context or are represented by faithful local fakes; every load-bearing
|
||||
rubric criterion is checkable against the repo.
|
||||
- **`partial`** — the core of the task is offline-completable and the rubric
|
||||
mostly grades local substance, but some of the success criteria lean
|
||||
outside the sandbox: a secondary "and it works in prod"-shaped expectation,
|
||||
a mock whose fidelity is doing a lot of load-bearing work, a local proxy
|
||||
(config parses, unit tests pass) standing in for an external outcome (the
|
||||
pipeline is faster), or a prompt whose natural reading promises more
|
||||
end-to-end confidence than the sandbox can deliver. The task works; the
|
||||
author should look at each finding and decide whether to rescope, reword,
|
||||
or accept the gap knowingly.
|
||||
- **`not-offline-verifiable`** — either half of the controlling test fails
|
||||
outright. *Success-criteria form:* the system being operated on (pipeline,
|
||||
prod, third-party service) isn't there and can't be faithfully faked, so
|
||||
neither doing the work well nor verifying it can happen in the workspace.
|
||||
*Completability form:* the ask names a technology the repo carries no
|
||||
library for, so step one is a registry fetch that rebuilding the image
|
||||
correctly would not remove. That the sandbox permits the fetch is irrelevant
|
||||
and plays no part in the verdict — the task's correctness is not supposed to
|
||||
hang on what a registry serves that day.
|
||||
- **`not-applicable`** — nothing to assess: `instruction.md` is missing,
|
||||
empty, or only template/placeholder content, and there is no session
|
||||
history to read an ask from. Re-run once the prompt lands.
|
||||
|
||||
`not-offline-verifiable` and `partial` are the flagged outcomes. Findings on
|
||||
the *verifiability* half stay advisory: where success criteria live is a
|
||||
judgment call, and the finding is a consideration for the author. A finding on
|
||||
the *completability* half is not — whether a named dependency appears in any
|
||||
manifest is a checked fact. Report it plainly and say which manifests you
|
||||
read. Only the integral-or-consequential call stands between that fact and
|
||||
the verdict, and the ask itself settles it: a library the repo never had is
|
||||
integral, a library the image forgot to install is ours to fix.
|
||||
|
||||
## Confidence
|
||||
|
||||
- **HIGH** — the call is unambiguous: the success criteria plainly live
|
||||
outside the sandbox (or plainly don't), and the rubric confirms the
|
||||
reading.
|
||||
- **MEDIUM** — at least one finding is genuinely two-sided: a mock whose
|
||||
fidelity a reasonable reviewer might judge either way, or a prompt that
|
||||
reads external but a rubric that grades local.
|
||||
- **LOW** — limited information: the rubric is thin or absent so you can't
|
||||
tell what's actually verified, or the workspace's representation of the
|
||||
external side couldn't be assessed.
|
||||
|
||||
## Relationship to other detectors
|
||||
|
||||
- **vs. detector-broken-dev-env.** That detector owns the *environment being
|
||||
broken*: the workspace doesn't build, tests flake, artifacts contradict the
|
||||
premise. This detector fires even when the environment is perfectly healthy
|
||||
— the defect is that the TASK's success criteria live outside the sandbox.
|
||||
"The tests won't run" is broken-dev-env; "no test that could run here can
|
||||
tell you whether this worked" is this detector. A dependency the *ask*
|
||||
requires but no manifest declares is this detector's (the env is fine, the
|
||||
ask isn't completable); a dependency the *existing code* imports but no
|
||||
manifest declares is broken-dev-env's (the env is broken).
|
||||
- **vs. detector-fact-check-rubric-claims.** Its reachability axis asks
|
||||
whether a specific *fact* the rubric grades the response for knowing is
|
||||
reachable from the package. This detector asks the structural version:
|
||||
whether the task's *success criteria as a whole* are checkable from inside
|
||||
the sandbox. A rubric criterion "the provider accepts the payload" can
|
||||
surface in both — as an unreachable/unverifiable claim there, and as an
|
||||
offline-verifiability finding here.
|
||||
- **vs. detector-meaningful-failure.** That detector asks whether the graded
|
||||
failure is real, proportionate, and elicited. A not-offline-verifiable task
|
||||
often *also* fails to elicit meaningfully (the runs are all pantomime), but
|
||||
the diagnosis differs: meaningful-failure says "this failure isn't worth
|
||||
grading"; this detector says "no one inside the sandbox can check the thing
|
||||
being graded."
|
||||
- **vs. detector-answer-obviousness.** Unrelated axis (is the expected answer
|
||||
inferable from the prompt?). No overlap expected; neither subsumes the
|
||||
other.
|
||||
|
||||
## Anti-patterns: do not do these
|
||||
|
||||
- **Don't flag every mention of an external service.** Scenario realism
|
||||
requires them. Trace the graded success criteria; flag only when *they*
|
||||
live outside.
|
||||
- **Don't demand hermetic purity.** Nearly every repo talks to something.
|
||||
The bar is the controlling test — complete AND verify from within the
|
||||
initialized workspace — not "the prompt never says the word 'deploy'."
|
||||
- **Don't punish tasks that are honest about the boundary.** A rubric that
|
||||
credits the agent for saying "this can't be verified from here" has priced
|
||||
the sandbox in; that's a strength, not a finding.
|
||||
- **Don't treat a verifiability flag as a verdict on the author or the
|
||||
task's worth.** That output is something to consider — a pointer at where
|
||||
the success criteria live — phrased so the author can decide. Never assert
|
||||
the task is invalid; never frame the finding as a failure. A
|
||||
missing-dependency finding is the exception: it is a fact about the
|
||||
manifests, so state it rather than softening it into a consideration.
|
||||
- **Don't consult the task's network policy.** `allow_internet`,
|
||||
`network_mode` and `allowed_hosts` exist for the grading harness, not the
|
||||
agent, and none of them actually closes the sandbox. Reading them can only
|
||||
mislead you here: every task has egress, so weighing it would clear every
|
||||
missing-dependency finding in the corpus. Judge the repo's manifests against
|
||||
the ask and nothing else.
|
||||
- **Don't fault a run for going online, and do flag a task that faults it for
|
||||
you.** A run reaching the web is not grounds for any finding, and never
|
||||
grounds to return a submission — ask only whether the task is solvable
|
||||
without the internet and whether the network is central to it. A prompt or
|
||||
rubric asserting the environment has no internet, inventing a reason for that
|
||||
("the security team has blocked outbound traffic"), or deducting for a
|
||||
lookup, is an unrealistic constraint the author added: say so as a finding on
|
||||
the authored text.
|
||||
- **Don't clear a missing dependency because the framework supports it.**
|
||||
"Rails ships `:redis_cache_store`", "pytest has a coverage plugin" — an
|
||||
adapter existing upstream says nothing about whether the gem or package is
|
||||
in this repo's lockfile. Open the manifest.
|
||||
- **Don't re-litigate env health.** Whether the workspace builds and the
|
||||
suite passes belongs to detector-broken-dev-env. Assume a healthy env and
|
||||
ask where the success criteria live — a healthy env does not imply the ask's
|
||||
own dependencies are present, which is the completability check above.
|
||||
- **Don't cite evidence you haven't verified in the submitted package.**
|
||||
Quote the prompt, rubric, and workspace as they exist in the actual
|
||||
submission — not as you remember or infer them.
|
||||
|
||||
## Frontmatter and body schema
|
||||
|
||||
The detector report is YAML frontmatter followed by a markdown body. Both
|
||||
contexts produce the same shape; only the *sink* differs (the wrapping
|
||||
`SKILL.md` tells you where to send the report).
|
||||
|
||||
**Frontmatter** — exactly these keys, exactly these enum values:
|
||||
|
||||
```yaml
|
||||
---
|
||||
detector: detector-offline-verifiability
|
||||
verdict: offline-verifiable | partial | not-offline-verifiable | not-applicable
|
||||
confidence: HIGH | MEDIUM | LOW
|
||||
---
|
||||
```
|
||||
|
||||
**Body sections**, in this order:
|
||||
|
||||
```markdown
|
||||
# Offline-verifiability check: <slug>
|
||||
|
||||
## Findings
|
||||
|
||||
One block per finding, strongest first:
|
||||
|
||||
### <short label> — <external-dependency shape> (<clear | partial>)
|
||||
|
||||
- **Where:** the file (and line/section, or the turn for a session message)
|
||||
where the ask or success criterion appears.
|
||||
- **Quote:** the passage verbatim, as a blockquote — never a paraphrase.
|
||||
- **Why it lives outside:** one or two sentences — what a human SWE would
|
||||
need the network or a live system for, in doing or verifying this, and
|
||||
what the sandbox can actually check instead.
|
||||
- **Something to consider:** a concrete rescoping option — grade the local
|
||||
protocol slice, reword the ask as a plan/assessment, point the criterion
|
||||
at the repo's fake provider, credit honest disclosure of the boundary —
|
||||
worded so the author can decide whether to take it.
|
||||
|
||||
For `offline-verifiable`, quote the strongest near-miss (the most
|
||||
external-sounding passage) and say why it was cleared. For `not-applicable`,
|
||||
name the missing artifacts.
|
||||
|
||||
## Overall verdict
|
||||
|
||||
2–3 paragraphs reducing the findings to the chosen verdict: where the
|
||||
task's success criteria live, whether the workspace (including any local
|
||||
fakes) can honestly check them, and — for verifiability findings, which are
|
||||
advisory — what a rescoping pass would consider first. For a
|
||||
missing-dependency finding, drop the hedging: name the package, name every
|
||||
manifest and lockfile you checked, and say the ask can't be completed offline
|
||||
as shipped. For `offline-verifiable`, why
|
||||
the near-misses are scenario context or faithfully mocked rather than live
|
||||
dependencies.
|
||||
```
|
||||
|
||||
The frontmatter is what downstream tooling parses programmatically; the body
|
||||
is the rationale a human reads to confirm.
|
||||
Reference in New Issue
Block a user