6.9 KiB
name, description, allowed-tools
| name | description | allowed-tools |
|---|---|---|
| detector-offline-verifiability | Self-check whether your task uses the internet — it must not. The test agent's environment is initialized up front (repo checked out, packages installed) and a good task is offline-completable and offline-verifiable: a competent SWE could do the work AND trust their verification of it entirely from within the repo. The trial does reach the network and that can't be changed, so this reads your task, never what an agent did in a run. Flags tasks whose success criteria live materially outside the sandbox — "speed up our CI/CD pipeline" (verifying needs the live pipeline), "redeploy to prod" (prod doesn't exist in the sandbox), "migrate from Zendesk to Intercom" (neither service is reachable, so mocks are guesses that likely won't survive real integration), "check the dashboard," published-package behavior. External services as scenario dressing are fine; protocol-slice integrations against a faithful local fake are fine. Explicitly advisory: every finding is something to consider, never a failure, and it blocks nothing. Reads instruction.md + the resolved holistic rubric (+ the workspace for local fakes); runs before or after reference runs exist. | Bash, Read, Write |
Offline-verifiability detector
This skill checks one of your tasks for offline-verifiability — whether the ask still makes sense inside the sandbox the test agent actually gets.
The rule is: don't create a task that uses the internet. It has to be solvable and checkable without one, and the network must never be a central component of the work. The sandbox is initialized before the task starts (repo checked out, dependencies installed), and from there everything that decides the grade should live in the repo. So the question is: could a competent SWE complete AND verify your task entirely from within the initialized repo — and would their "it works" actually be trustworthy?
What that rule is not. The trial does reach the network, and you can't
change that — it needs network access to reach the model, so leave
allow_internet at its default and don't add network_mode or
allowed_hosts. If an agent goes and reads something on the web during one of
your runs, that's outside your control and it's fine: the run and the task
still stand, provided the task works without the internet and its outcome
doesn't rest on what the agent found. And don't write the restriction into the
task — a prompt telling the agent it has no internet access, a justification
invented for it ("the security team has blocked outbound traffic"), or a rubric
that deducts for a lookup are unrealistic constraints that make the task
worse. This skill reads your task, never your runs.
The failure shape to catch: tasks whose success criteria live outside the sandbox. "Speed up our CI/CD pipeline" — the pipeline the work would be verified against isn't there. "Redeploy to prod" — there is no prod. "Migrate from Zendesk to Intercom" — the agent can't interact with either service, so it can only mock both ends, and mocks written without ever touching the real services almost certainly won't work at integration time. When a task has this shape, the grade measures how convincingly the agent pantomimes the work, not whether the work is right — and an agent that honestly says "I can't verify this from here" can end up scoring worse than one that confidently fakes it. Egress doesn't rescue any of these: reaching a live pipeline or a real SaaS tenant needs credentials and real state, not just a route out.
The other shape to watch is an ask whose first step is a fetch — "migrate the
cache to Redis" in an app that ships no Redis client, "add TOTP" with no OTP
library in any manifest. The install will probably work, because the network
is there. That's the problem: your task's correctness then rides on what a
registry serves on grading day. Put what the task needs into
environment/workspace.patch instead, where it's pinned and identical on
every run.
What doesn't trip this check: a run in which the agent went online (that's
not something your task did), external services as
scenario dressing (a prompt set at a company that uses Stripe is realism, as
long as the graded work and its verification are local), and integrations
scoped to a documented protocol slice with a faithful local fake — ideally
wired through the fake providers your repo ships (see /brainstorm-product-arcs for the
"simulate the protocol, not the product" filter this mirrors).
This check is advisory. Where the line falls is a judgment call — a task can even be deliberately built around recognizing the sandbox's limits, with a holistic rubric that credits saying so. The report exists so you can consider where your success criteria live: each finding quotes the passage, says what a human SWE would need the network or a live system for, and offers a rescoping option, so the decision stays yours. Nothing here blocks your submission.
Read these before deciding:
.claude/skills/_detector-worker-shell.md— where to write the report and how to handle re-runs..claude/skills/detector-offline-verifiability/core.md— the controlling test (offline-completable + offline-verifiable), the external-dependency shapes, the mock-fidelity boundary, what is NOT a finding, verdict enums, and the body schema.
Compose the report per the schema in core.md and write it per _detector-worker-shell.md.
Acting on the verdict
offline-verifiable— the work and its verification both live inside the workspace; any external names are scenario context or faithfully faked. Good. Move on.partial— the core of your task is offline-completable, but some success criteria lean outside: a "works in prod"-shaped expectation, a local proxy (config parses, unit tests pass) standing in for an external outcome (the pipeline gets faster), or a mock whose fidelity is carrying a lot of the grade. Read each finding and decide: tighten the prompt so it asks for the local slice, point the criterion at your repo's fake provider, or keep the framing deliberately and make sure your holistic rubric grades only what the sandbox can check (crediting honest disclosure of the rest).not-offline-verifiable— the system your task operates on (pipeline, prod, third-party service) isn't in the sandbox and can't be faithfully faked, so neither doing the work well nor verifying it can happen there. Consider the rescoping option in each finding: extract the protocol slice and build an adversarial local mock for it, reframe the ask as an assessment or plan graded on repo evidence, or pick a different behavior to test. If you believe the task works as-is, that's your call — but make sure the holistic rubric never asks the grader (or the agent) for a verification the sandbox cannot perform.not-applicable— there's no prompt to assess yet. Draft it first.