restored entire zip and config'd
This commit is contained in:
@@ -0,0 +1,114 @@
|
||||
---
|
||||
name: detector-offline-verifiability
|
||||
description: |
|
||||
Self-check whether your task uses the internet — it must not. The test
|
||||
agent's environment is initialized up front (repo checked out, packages
|
||||
installed) and a good task is offline-completable and offline-verifiable: a
|
||||
competent SWE could do the work AND trust their verification of it entirely
|
||||
from within the repo. The trial does reach the network and that can't be
|
||||
changed, so this reads your task, never what an agent did in a run.
|
||||
Flags tasks whose success criteria live materially outside the sandbox —
|
||||
"speed up our CI/CD pipeline" (verifying needs the live pipeline),
|
||||
"redeploy to prod" (prod doesn't exist in the sandbox),
|
||||
"migrate from Zendesk to Intercom" (neither service is reachable, so
|
||||
mocks are guesses that likely won't survive real integration), "check the
|
||||
dashboard," published-package behavior. External services as scenario
|
||||
dressing are fine; protocol-slice integrations against a faithful local
|
||||
fake are fine. Explicitly advisory: every finding is something to
|
||||
consider, never a failure, and it blocks nothing. Reads instruction.md +
|
||||
the resolved holistic rubric (+ the workspace for local fakes); runs
|
||||
before or after reference runs exist.
|
||||
allowed-tools: Bash, Read, Write
|
||||
---
|
||||
|
||||
# Offline-verifiability detector
|
||||
|
||||
This skill checks one of your tasks for **offline-verifiability** — whether
|
||||
the ask still makes sense inside the sandbox the test agent actually gets.
|
||||
|
||||
**The rule is: don't create a task that uses the internet.** It has to be
|
||||
solvable and checkable without one, and the network must never be a central
|
||||
component of the work. The sandbox is initialized before the task starts (repo
|
||||
checked out, dependencies installed), and from there everything that decides
|
||||
the grade should live in the repo. So the question is: could a competent SWE
|
||||
complete AND verify your task entirely from within the initialized repo — and
|
||||
would their "it works" actually be trustworthy?
|
||||
|
||||
**What that rule is not.** The trial does reach the network, and you can't
|
||||
change that — it needs network access to reach the model, so leave
|
||||
`allow_internet` at its default and don't add `network_mode` or
|
||||
`allowed_hosts`. If an agent goes and reads something on the web during one of
|
||||
your runs, that's outside your control and it's fine: the run and the task
|
||||
still stand, provided the task works without the internet and its outcome
|
||||
doesn't rest on what the agent found. And don't write the restriction into the
|
||||
task — a prompt telling the agent it has no internet access, a justification
|
||||
invented for it ("the security team has blocked outbound traffic"), or a rubric
|
||||
that deducts for a lookup are unrealistic constraints that make the task
|
||||
worse. This skill reads your task, never your runs.
|
||||
|
||||
The failure shape to catch: tasks whose *success criteria* live outside the
|
||||
sandbox. "Speed up our CI/CD pipeline" — the pipeline the work would be
|
||||
verified against isn't there. "Redeploy to prod" — there is no prod. "Migrate
|
||||
from Zendesk to Intercom" — the agent can't interact with either service, so
|
||||
it can only mock both ends, and mocks written without ever touching the real
|
||||
services almost certainly won't work at integration time. When a task has
|
||||
this shape, the grade measures how convincingly the agent pantomimes the
|
||||
work, not whether the work is right — and an agent that honestly says "I
|
||||
can't verify this from here" can end up scoring worse than one that
|
||||
confidently fakes it. Egress doesn't rescue any of these: reaching a live
|
||||
pipeline or a real SaaS tenant needs credentials and real state, not just a
|
||||
route out.
|
||||
|
||||
The other shape to watch is an ask whose first step is a fetch — "migrate the
|
||||
cache to Redis" in an app that ships no Redis client, "add TOTP" with no OTP
|
||||
library in any manifest. The install will probably work, because the network
|
||||
is there. That's the problem: your task's correctness then rides on what a
|
||||
registry serves on grading day. Put what the task needs into
|
||||
`environment/workspace.patch` instead, where it's pinned and identical on
|
||||
every run.
|
||||
|
||||
What *doesn't* trip this check: a run in which the agent went online (that's
|
||||
not something your task did), external services as
|
||||
scenario dressing (a prompt set at a company that uses Stripe is realism, as
|
||||
long as the graded work and its verification are local), and integrations
|
||||
scoped to a documented protocol slice with a faithful local fake — ideally
|
||||
wired through the fake providers your repo ships (see `/brainstorm-product-arcs` for the
|
||||
"simulate the protocol, not the product" filter this mirrors).
|
||||
|
||||
**This check is advisory.** Where the line falls is a judgment call — a task
|
||||
can even be deliberately built around recognizing the sandbox's limits, with
|
||||
a holistic rubric that credits saying so. The report exists so you can *consider*
|
||||
where your success criteria live: each finding quotes the passage, says what a
|
||||
human SWE would need the network or a live system for, and offers a rescoping
|
||||
option, so the decision stays yours. Nothing here blocks your submission.
|
||||
|
||||
Read these before deciding:
|
||||
|
||||
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
|
||||
2. `.claude/skills/detector-offline-verifiability/core.md` — the controlling test (offline-completable + offline-verifiable), the external-dependency shapes, the mock-fidelity boundary, what is NOT a finding, verdict enums, and the body schema.
|
||||
|
||||
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
|
||||
|
||||
## Acting on the verdict
|
||||
|
||||
- **`offline-verifiable`** — the work and its verification both live inside
|
||||
the workspace; any external names are scenario context or faithfully
|
||||
faked. Good. Move on.
|
||||
- **`partial`** — the core of your task is offline-completable, but some
|
||||
success criteria lean outside: a "works in prod"-shaped expectation, a
|
||||
local proxy (config parses, unit tests pass) standing in for an external
|
||||
outcome (the pipeline gets faster), or a mock whose fidelity is carrying a
|
||||
lot of the grade. Read each finding and decide: tighten the prompt so it
|
||||
asks for the local slice, point the criterion at your repo's fake provider,
|
||||
or keep the framing deliberately and make sure your holistic rubric grades
|
||||
only what the sandbox can check (crediting honest disclosure of the rest).
|
||||
- **`not-offline-verifiable`** — the system your task operates on (pipeline,
|
||||
prod, third-party service) isn't in the sandbox and can't be faithfully
|
||||
faked, so neither doing the work well nor verifying it can happen there.
|
||||
Consider the rescoping option in each finding: extract the protocol slice
|
||||
and build an adversarial local mock for it, reframe the ask as an
|
||||
assessment or plan graded on repo evidence, or pick a different behavior to
|
||||
test. If you believe the task works as-is, that's your call — but make sure
|
||||
the holistic rubric never asks the grader (or the agent) for a verification
|
||||
the sandbox cannot perform.
|
||||
- **`not-applicable`** — there's no prompt to assess yet. Draft it first.
|
||||
Reference in New Issue
Block a user