Project-2 baseline

This commit is contained in:
2026-10-04 21:19:23 -04:00
parent 0f04889edf
commit 213eb3c403
861 changed files with 1710 additions and 3363322 deletions

View File

@@ -1,14 +1,15 @@
---
name: detector-offline-verifiability
description: |
Self-check whether your task makes sense in the no-network sandbox it runs
in. The test agent's environment is initialized up front — repo checked
out, packages installed — and then runs with no outbound network access, so
a good task is offline-completable and offline-verifiable: a competent SWE
could do the work AND trust their verification of it entirely from within
the repo. Flags tasks whose success criteria live materially outside the
sandbox — "speed up our CI/CD pipeline" (verifying needs the live
pipeline), "redeploy to prod" (prod doesn't exist in the sandbox),
Self-check whether your task uses the internet — it must not. The test
agent's environment is initialized up front (repo checked out, packages
installed) and a good task is offline-completable and offline-verifiable: a
competent SWE could do the work AND trust their verification of it entirely
from within the repo. The trial does reach the network and that can't be
changed, so this reads your task, never what an agent did in a run.
Flags tasks whose success criteria live materially outside the sandbox —
"speed up our CI/CD pipeline" (verifying needs the live pipeline),
"redeploy to prod" (prod doesn't exist in the sandbox),
"migrate from Zendesk to Intercom" (neither service is reachable, so
mocks are guesses that likely won't survive real integration), "check the
dashboard," published-package behavior. External services as scenario
@@ -24,11 +25,26 @@ allowed-tools: Bash, Read, Write
This skill checks one of your tasks for **offline-verifiability** — whether
the ask still makes sense inside the sandbox the test agent actually gets.
That sandbox is initialized before the task starts (repo checked out,
dependencies installed) and then has **no outbound network access**. So the
question is: could a competent SWE complete AND verify your task entirely
from within the initialized repo — and would their "it works" actually be
trustworthy?
**The rule is: don't create a task that uses the internet.** It has to be
solvable and checkable without one, and the network must never be a central
component of the work. The sandbox is initialized before the task starts (repo
checked out, dependencies installed), and from there everything that decides
the grade should live in the repo. So the question is: could a competent SWE
complete AND verify your task entirely from within the initialized repo — and
would their "it works" actually be trustworthy?
**What that rule is not.** The trial does reach the network, and you can't
change that — it needs network access to reach the model, so leave
`allow_internet` at its default and don't add `network_mode` or
`allowed_hosts`. If an agent goes and reads something on the web during one of
your runs, that's outside your control and it's fine: the run and the task
still stand, provided the task works without the internet and its outcome
doesn't rest on what the agent found. And don't write the restriction into the
task — a prompt telling the agent it has no internet access, a justification
invented for it ("the security team has blocked outbound traffic"), or a rubric
that deducts for a lookup are unrealistic constraints that make the task
worse. This skill reads your task, never your runs.
The failure shape to catch: tasks whose *success criteria* live outside the
sandbox. "Speed up our CI/CD pipeline" — the pipeline the work would be
@@ -39,13 +55,24 @@ services almost certainly won't work at integration time. When a task has
this shape, the grade measures how convincingly the agent pantomimes the
work, not whether the work is right — and an agent that honestly says "I
can't verify this from here" can end up scoring worse than one that
confidently fakes it.
confidently fakes it. Egress doesn't rescue any of these: reaching a live
pipeline or a real SaaS tenant needs credentials and real state, not just a
route out.
What *doesn't* trip this check: external services as scenario dressing (a
prompt set at a company that uses Stripe is realism, as long as the graded
work and its verification are local), and integrations scoped to a documented
protocol slice with a faithful local fake — ideally wired through the fake
providers your repo already ships (see `/brainstorm-product-arcs` for the
The other shape to watch is an ask whose first step is a fetch — "migrate the
cache to Redis" in an app that ships no Redis client, "add TOTP" with no OTP
library in any manifest. The install will probably work, because the network
is there. That's the problem: your task's correctness then rides on what a
registry serves on grading day. Put what the task needs into
`environment/workspace.patch` instead, where it's pinned and identical on
every run.
What *doesn't* trip this check: a run in which the agent went online (that's
not something your task did), external services as
scenario dressing (a prompt set at a company that uses Stripe is realism, as
long as the graded work and its verification are local), and integrations
scoped to a documented protocol slice with a faithful local fake — ideally
wired through the fake providers your repo ships (see `/brainstorm-product-arcs` for the
"simulate the protocol, not the product" filter this mirrors).
**This check is advisory.** Where the line falls is a judgment call — a task