Project-2 baseline
This commit is contained in:
@@ -1,14 +1,15 @@
|
||||
---
|
||||
name: detector-offline-verifiability
|
||||
description: |
|
||||
Self-check whether your task makes sense in the no-network sandbox it runs
|
||||
in. The test agent's environment is initialized up front — repo checked
|
||||
out, packages installed — and then runs with no outbound network access, so
|
||||
a good task is offline-completable and offline-verifiable: a competent SWE
|
||||
could do the work AND trust their verification of it entirely from within
|
||||
the repo. Flags tasks whose success criteria live materially outside the
|
||||
sandbox — "speed up our CI/CD pipeline" (verifying needs the live
|
||||
pipeline), "redeploy to prod" (prod doesn't exist in the sandbox),
|
||||
Self-check whether your task uses the internet — it must not. The test
|
||||
agent's environment is initialized up front (repo checked out, packages
|
||||
installed) and a good task is offline-completable and offline-verifiable: a
|
||||
competent SWE could do the work AND trust their verification of it entirely
|
||||
from within the repo. The trial does reach the network and that can't be
|
||||
changed, so this reads your task, never what an agent did in a run.
|
||||
Flags tasks whose success criteria live materially outside the sandbox —
|
||||
"speed up our CI/CD pipeline" (verifying needs the live pipeline),
|
||||
"redeploy to prod" (prod doesn't exist in the sandbox),
|
||||
"migrate from Zendesk to Intercom" (neither service is reachable, so
|
||||
mocks are guesses that likely won't survive real integration), "check the
|
||||
dashboard," published-package behavior. External services as scenario
|
||||
@@ -24,11 +25,26 @@ allowed-tools: Bash, Read, Write
|
||||
|
||||
This skill checks one of your tasks for **offline-verifiability** — whether
|
||||
the ask still makes sense inside the sandbox the test agent actually gets.
|
||||
That sandbox is initialized before the task starts (repo checked out,
|
||||
dependencies installed) and then has **no outbound network access**. So the
|
||||
question is: could a competent SWE complete AND verify your task entirely
|
||||
from within the initialized repo — and would their "it works" actually be
|
||||
trustworthy?
|
||||
|
||||
**The rule is: don't create a task that uses the internet.** It has to be
|
||||
solvable and checkable without one, and the network must never be a central
|
||||
component of the work. The sandbox is initialized before the task starts (repo
|
||||
checked out, dependencies installed), and from there everything that decides
|
||||
the grade should live in the repo. So the question is: could a competent SWE
|
||||
complete AND verify your task entirely from within the initialized repo — and
|
||||
would their "it works" actually be trustworthy?
|
||||
|
||||
**What that rule is not.** The trial does reach the network, and you can't
|
||||
change that — it needs network access to reach the model, so leave
|
||||
`allow_internet` at its default and don't add `network_mode` or
|
||||
`allowed_hosts`. If an agent goes and reads something on the web during one of
|
||||
your runs, that's outside your control and it's fine: the run and the task
|
||||
still stand, provided the task works without the internet and its outcome
|
||||
doesn't rest on what the agent found. And don't write the restriction into the
|
||||
task — a prompt telling the agent it has no internet access, a justification
|
||||
invented for it ("the security team has blocked outbound traffic"), or a rubric
|
||||
that deducts for a lookup are unrealistic constraints that make the task
|
||||
worse. This skill reads your task, never your runs.
|
||||
|
||||
The failure shape to catch: tasks whose *success criteria* live outside the
|
||||
sandbox. "Speed up our CI/CD pipeline" — the pipeline the work would be
|
||||
@@ -39,13 +55,24 @@ services almost certainly won't work at integration time. When a task has
|
||||
this shape, the grade measures how convincingly the agent pantomimes the
|
||||
work, not whether the work is right — and an agent that honestly says "I
|
||||
can't verify this from here" can end up scoring worse than one that
|
||||
confidently fakes it.
|
||||
confidently fakes it. Egress doesn't rescue any of these: reaching a live
|
||||
pipeline or a real SaaS tenant needs credentials and real state, not just a
|
||||
route out.
|
||||
|
||||
What *doesn't* trip this check: external services as scenario dressing (a
|
||||
prompt set at a company that uses Stripe is realism, as long as the graded
|
||||
work and its verification are local), and integrations scoped to a documented
|
||||
protocol slice with a faithful local fake — ideally wired through the fake
|
||||
providers your repo already ships (see `/brainstorm-product-arcs` for the
|
||||
The other shape to watch is an ask whose first step is a fetch — "migrate the
|
||||
cache to Redis" in an app that ships no Redis client, "add TOTP" with no OTP
|
||||
library in any manifest. The install will probably work, because the network
|
||||
is there. That's the problem: your task's correctness then rides on what a
|
||||
registry serves on grading day. Put what the task needs into
|
||||
`environment/workspace.patch` instead, where it's pinned and identical on
|
||||
every run.
|
||||
|
||||
What *doesn't* trip this check: a run in which the agent went online (that's
|
||||
not something your task did), external services as
|
||||
scenario dressing (a prompt set at a company that uses Stripe is realism, as
|
||||
long as the graded work and its verification are local), and integrations
|
||||
scoped to a documented protocol slice with a faithful local fake — ideally
|
||||
wired through the fake providers your repo ships (see `/brainstorm-product-arcs` for the
|
||||
"simulate the protocol, not the product" filter this mirrors).
|
||||
|
||||
**This check is advisory.** Where the line falls is a judgment call — a task
|
||||
|
||||
Reference in New Issue
Block a user