lots of change - all to start my 3rd redo

This commit is contained in:
2026-09-26 14:31:52 -04:00
parent 7f4d388e19
commit bceb52e8ee
1046 changed files with 4476 additions and 0 deletions

View File

@@ -1,87 +0,0 @@
---
name: detector-offline-verifiability
description: |
Self-check whether your task makes sense in the no-network sandbox it runs
in. The test agent's environment is initialized up front — repo checked
out, packages installed — and then runs with no outbound network access, so
a good task is offline-completable and offline-verifiable: a competent SWE
could do the work AND trust their verification of it entirely from within
the repo. Flags tasks whose success criteria live materially outside the
sandbox — "speed up our CI/CD pipeline" (verifying needs the live
pipeline), "redeploy to prod" (prod doesn't exist in the sandbox),
"migrate from Zendesk to Intercom" (neither service is reachable, so
mocks are guesses that likely won't survive real integration), "check the
dashboard," published-package behavior. External services as scenario
dressing are fine; protocol-slice integrations against a faithful local
fake are fine. Explicitly advisory: every finding is something to
consider, never a failure, and it blocks nothing. Reads instruction.md +
the resolved holistic rubric (+ the workspace for local fakes); runs
before or after reference runs exist.
allowed-tools: Bash, Read, Write
---
# Offline-verifiability detector
This skill checks one of your tasks for **offline-verifiability** — whether
the ask still makes sense inside the sandbox the test agent actually gets.
That sandbox is initialized before the task starts (repo checked out,
dependencies installed) and then has **no outbound network access**. So the
question is: could a competent SWE complete AND verify your task entirely
from within the initialized repo — and would their "it works" actually be
trustworthy?
The failure shape to catch: tasks whose *success criteria* live outside the
sandbox. "Speed up our CI/CD pipeline" — the pipeline the work would be
verified against isn't there. "Redeploy to prod" — there is no prod. "Migrate
from Zendesk to Intercom" — the agent can't interact with either service, so
it can only mock both ends, and mocks written without ever touching the real
services almost certainly won't work at integration time. When a task has
this shape, the grade measures how convincingly the agent pantomimes the
work, not whether the work is right — and an agent that honestly says "I
can't verify this from here" can end up scoring worse than one that
confidently fakes it.
What *doesn't* trip this check: external services as scenario dressing (a
prompt set at a company that uses Stripe is realism, as long as the graded
work and its verification are local), and integrations scoped to a documented
protocol slice with a faithful local fake — ideally wired through the fake
providers your repo already ships (see `/brainstorm-product-arcs` for the
"simulate the protocol, not the product" filter this mirrors).
**This check is advisory.** Where the line falls is a judgment call — a task
can even be deliberately built around recognizing the sandbox's limits, with
a holistic rubric that credits saying so. The report exists so you can *consider*
where your success criteria live: each finding quotes the passage, says what a
human SWE would need the network or a live system for, and offers a rescoping
option, so the decision stays yours. Nothing here blocks your submission.
Read these before deciding:
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
2. `.claude/skills/detector-offline-verifiability/core.md` — the controlling test (offline-completable + offline-verifiable), the external-dependency shapes, the mock-fidelity boundary, what is NOT a finding, verdict enums, and the body schema.
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
## Acting on the verdict
- **`offline-verifiable`** — the work and its verification both live inside
the workspace; any external names are scenario context or faithfully
faked. Good. Move on.
- **`partial`** — the core of your task is offline-completable, but some
success criteria lean outside: a "works in prod"-shaped expectation, a
local proxy (config parses, unit tests pass) standing in for an external
outcome (the pipeline gets faster), or a mock whose fidelity is carrying a
lot of the grade. Read each finding and decide: tighten the prompt so it
asks for the local slice, point the criterion at your repo's fake provider,
or keep the framing deliberately and make sure your holistic rubric grades
only what the sandbox can check (crediting honest disclosure of the rest).
- **`not-offline-verifiable`** — the system your task operates on (pipeline,
prod, third-party service) isn't in the sandbox and can't be faithfully
faked, so neither doing the work well nor verifying it can happen there.
Consider the rescoping option in each finding: extract the protocol slice
and build an adversarial local mock for it, reframe the ask as an
assessment or plan graded on repo evidence, or pick a different behavior to
test. If you believe the task works as-is, that's your call — but make sure
the holistic rubric never asks the grader (or the agent) for a verification
the sandbox cannot perform.
- **`not-applicable`** — there's no prompt to assess yet. Draft it first.

View File

@@ -1,425 +0,0 @@
# Offline-verifiability detector — core
This file is the canonical, context-neutral content for the
detector-offline-verifiability detector. It defines the signal (does the task
make sense in a no-network sandbox?), the controlling test, the external-
dependency shapes to recognize, the verdict enums, and the output schema. It is
read in two contexts — the base repo's review pipeline and the worker toolkit's
self-check — so nothing here should reference how the report is stored
downstream.
## What this detector is for
Every task runs in a sandbox that is initialized up front — the repo checked
out, dependencies installed — and then executes with **no outbound network
access**. The agent under test can read, build, run, and test everything inside
the workspace, and nothing outside it. A task fits that world when everything
is totally verifiable from within the repo: offline-completable and
offline-verifiable, because the setup happened before the network went away.
Setup installs what the repo's own manifests and lockfiles declare **after
`environment/workspace.patch` has been applied** — the image copies the patched
workspace in and only *then* runs the dependency install. So a package the
author added, upgraded, downgraded or re-pinned in the patch is present in the
sandbox, and is never a completability finding; judge the manifests as the
patch leaves them, not as the pinned commit left them. What is never installed
is a library the ask requires the *agent* to add: acquiring that means `bundle
add`, `npm install <pkg>`, `pip install` — a registry fetch, mid-task, after
the network is gone.
**Do not consider the task's network policy. At all.** `task.toml`'s
`allow_internet` / `network_mode` / `allowed_hosts` fields are not about the
agent — `allow_internet = true` is scaffold boilerplate carried by essentially
every task so the *grading harness* can call its own API. It is not a grant of
registry access to the task, and it is out of scope for this detector: do not
read those fields, do not mention them in the report, and do not let them move
the verdict.
The corollary matters just as much: **a mid-run install that succeeded is not a
clearance.** If the reference runs show the agent fetching the package from a
registry, that is evidence the dependency was missing and needed — cite it as
support for the finding, never as a reason to soften it. "The runs prove it
worked, so this isn't a failure" is the wrong question, answered.
Some task ideas don't really make sense in that world, because a human SWE
would need internet access — or access to live systems that only exist outside
the sandbox — to really do the task well or to verify the result. The
canonical examples:
> Speed up our CI/CD pipeline
— you need access to that pipeline to verify your work. The pipeline's actual
runtime, caching behavior, and bottlenecks live on an external system the
sandbox doesn't have; the agent can edit config files but never observe whether
anything got faster.
> Redeploy to prod
— prod doesn't exist in the sandbox environment. There is nothing to deploy
to, so the "success" the prompt asks for cannot occur, let alone be checked.
> Migrate from Zendesk to Intercom
— the agent can't interact with either service, so it can't test the
migration end-to-end; it's just mocking things out, and mocks written without
ever touching the real services almost certainly won't work at integration
time. The part that makes the task hard — does it actually work against the
real thing? — is exactly the part the sandbox can't answer.
When a task has this shape, the reference runs and the grade measure how
convincingly the agent *pantomimes* the work, not whether the work is right.
The verifier can't check the thing that matters, the rubric drifts toward
style points, and an agent that (correctly) says "I can't verify this from
here" may score worse than one that confidently fakes it.
**The verifiability half of this detector is advisory.** Whether a task's
success criteria live too far outside the sandbox is a judgment call — most real tasks mention external services *somewhere*, and a scenario
can legitimately be about recognizing the limits of what's verifiable. A
flagged verdict means "here is something to consider about where this task's
success criteria live," never "this task is invalid." The author may have
deliberately scoped the graded substance to the local slice, and the flag is
the prompt to confirm that scoping is real.
The completability half is not a judgment call. Whether a library the ask
requires appears in any manifest is a fact you check, and a task that needs
one that isn't there cannot be carried out here at all.
## The controlling test
For the task as a whole, ask:
**Could a competent SWE complete AND verify this task entirely from within the
initialized repo — packages already installed, no network — and would their
"it works" claim actually be trustworthy?**
Break that into the two halves:
1. **Offline-completable.** Is everything the prompt asks for buildable from
what's in the workspace? Or does doing the work well require reaching
something outside — a live pipeline, a running production system, a
third-party API, a package registry, data that isn't in the repo?
**This half has a mechanical check, and it is not optional.** List every
library, framework, runner, or binary the ask or the rubric's criteria
name, then check each against every manifest and lockfile in the repo **as
the workspace patch leaves it** (`Gemfile`/`Gemfile.lock`, `package.json` + its lockfile,
`pyproject.toml`/`requirements*.txt`/`uv.lock`, `go.mod`, the Dockerfile).
Read the files — never settle this from knowledge of what the framework
supports. When a name is absent from all of them, the deciding question is
**integral or consequential**:
> If we rebuilt the image correctly, would this task still need the
> network?
- **Yes — integral.** The repo has no library for the thing the ask names:
migrate to Redis Cluster with no Redis client, add TOTP with no OTP gem,
write BDD features with no BDD runner, or a rubric that grades the fetch
itself ("the provider is installed and pinned compatibly"). Rebuilding
the image wouldn't help, because the dependency was never the repo's.
This is the completability failure — flag it, and cite the manifests you
read plus the runs that installed the package mid-session.
- **No — consequential.** The image simply forgot something the repo
already depends on: a runner, linter or type checker its own config
expects, or a sub-package the build skipped. That is an image-packaging
bug on our side, not a defect in the task's design. Do not flag the task
for it; record what is missing so the image can be fixed.
2. **Offline-verifiable.** Where do the success criteria live? If the honest
check for "did this work?" is *outside* the sandbox — watch the pipeline
get faster, see the dashboard update, confirm the third-party service
accepts the calls, install the published package — then the sandbox can
only verify a proxy, and the question is whether that proxy is faithful
enough to carry the grade.
A task passes when both halves stay inside the workspace: the deliverable is
code, config, tests, or analysis over what's in the repo, and the rubric's
success criteria are checkable against the repo (its test suite, its local
mocks and fakes, its own artifacts). A task gets flagged when the success
criteria live materially outside — external services, live pipelines, prod
deploys, third-party SaaS integration, "check the dashboard," published-package
behavior — even when the environment itself is perfectly healthy.
**Mocks are the boundary case, and fidelity is the question.** External
dependencies faked through a faithful local mock — a documented protocol
(file formats, webhook signatures, return codes) simulated the way the repo
already fakes its providers — keep a task offline-verifiable: the hard work is
on the repo's side and the mock exercises it honestly. The flag condition is a
mock that has to *invent* the external side because nobody can check it: an
undocumented or proprietary behavior, a product rather than a protocol, or an
integration whose entire difficulty is "does the real service accept this?"
A useful rule of thumb: if the mock's spec could be written straight from
public documentation and a correct integration against the mock would also be
correct against the real service, the mock carries the verification; if the
mock is a guess about the real thing, it doesn't.
## Inputs
Read from `harbor-tasks/<slug>/`:
- `instruction.md` — the prompt the agent under test receives. The primary
surface: what is the agent actually being asked to deliver, and what would
"done, and correct" mean for that ask? For a snapshot / multi-turn task,
also read the standing user turns in the session history
(`environment/session.jsonl` or `session-full.jsonl`) — an ask that arrives
in a prior turn binds the agent the same way.
- The grader guidance — context for what is actually verified. Resolve which
guidance file the grader actually reads (`bash scripts/guidance-target.sh
<slug>` — the worker shell's guidance-target resolution) and read that
file, never its sibling. This is
where the flag is confirmed or cleared: a prompt that *mentions* deployment
can still be graded entirely on local substance, and a local-sounding prompt
can hide a rubric criterion that only a live system could check ("the
webhook must be accepted by the provider"). Ask of each load-bearing
criterion: what would the grader look at, and is it in the workspace?
- `environment/workspace.patch` and the workspace — context for whether the
external side is actually represented locally: an existing fake provider,
fixtures, a stub server, seeded data. A prompt naming a third-party service
reads very differently when the repo ships a faithful fake of it.
- `reference-runs/*/grade.md` — not required, but a useful cross-check when
present: runs where the agent had to invent mock behavior wholesale, spent
its effort simulating an absent system, or was penalized for saying it
couldn't verify something the sandbox genuinely can't verify, all
corroborate the flag.
## External-dependency shapes to look for
- **Live infrastructure as the subject.** The deliverable is an operation on
a system that exists only outside the sandbox: speed up the CI/CD pipeline,
redeploy to prod, rotate the certs, fix the DNS, tune the production
database. The workspace may contain the *config* for these systems, but the
success criteria — the pipeline runs faster, the deploy succeeds — are
observable only on the real thing.
- **Third-party SaaS integration as the deliverable.** Migrate from Zendesk
to Intercom, integrate the new payment provider, sync with the CRM — where
the graded outcome is end-to-end behavior against services the agent can't
reach, and no faithful local fake exists or could exist. (A protocol-slice
task against a documented contract with a faithful adversarial mock is the
acceptable version — see the controlling test.)
- **Success criteria that name an external observation.** "Check the
dashboard," "confirm the metrics improve," "verify the alert fires in
PagerDuty," "make sure the docs site renders" — the rubric or prompt defines
done-ness as something seen on a system that isn't in the workspace.
- **Published-artifact behavior.** Release the package and verify it installs
from the registry, publish the image, ship the SDK update to consumers —
the verifying step is inherently on the other side of the network boundary.
- **Missing-at-runtime acquisitions.** The task's happy path requires
fetching something after the network is gone: installing a dependency that
isn't pre-installed or vendored, pulling a dataset from a URL, cloning
another repo, calling a real API for live data. (Setup-time installation is
fine only for what a manifest already declares — that got installed before
the shutoff. A package the ask tells the agent to add is not setup-time; it
is a runtime acquisition, and by then the network is gone.)
- **An uninstallable dependency as the deliverable.** The ask names a
technology the repo does not carry — migrate the cache to Redis in an app
whose only cache gem is `solid_cache`, add TOTP and lockout to an app
shipping no auth library, add coverage or BDD tooling that appears in no
manifest — and the rubric grades the result as installed and working. The
graded substance can look entirely local (config, key shapes, call sites)
while step one is an impossible `bundle add`. A first-party framework
adapter still needs its gem: "documented upstream" is not "present here".
- **External knowledge as the graded substance.** The rubric's success hinges
on looking up volatile external state — current API behavior of a live
provider, today's prices, the latest version of a service's schema — that
isn't captured in the workspace and can't be derived from it.
## What is NOT a finding
- **External services as scenario dressing.** A prompt set at a company that
uses Stripe, Zendesk, and AWS is realism. The question is where the *graded
work and its verification* happen — if the deliverable is repo code and the
rubric checks repo behavior, the named services are backdrop, not
dependencies.
- **Protocol-slice integrations with a faithful local fake.** Build the
webhook verifier, parse the provider's documented file format, reconcile
against the seeded fixture service — especially when the repo already fakes
that provider and the task extends the existing seam. That is the sanctioned
way to do external-facing work offline.
- **Deploy/CI config work graded on local substance.** Editing a CI config or
a deploy manifest where the rubric checks properties verifiable in the
workspace — the config parses, the referenced scripts exist and run, the
documented invariants hold — is bounded. It may still merit `partial` when
the *real* success criterion (the pipeline actually gets faster) is external
and the local checks are a thin proxy; say which.
- **Assessments and plans about external systems, graded on repo evidence.**
"Review our migration plan," "assess what moving to Intercom would take" —
where the deliverable is analysis whose load-bearing claims are checkable
against the repo. A *plan* for external work is offline-verifiable; only
*executing and confirming* the external work isn't.
- **Tasks deliberately about recognizing the limit.** A scenario can be built
so that the right behavior is to say "this part can't be verified from
here" — and the rubric credits exactly that. If the grader guidance treats
the boundary honestly (credits disclosure, doesn't demand the impossible
verification), the external dependency is the task working as designed.
- **Hard-but-local work.** Big refactors, gnarly debugging, performance work
measured by local benchmarks — difficulty is not an offline-verifiability
problem. This detector is orthogonal to how hard the task is.
## Verdict definitions
- **`offline-verifiable`** — the controlling test passes: a competent SWE
could complete the ask and trust their own verification of it entirely
within the initialized workspace. External services, if named, are scenario
context or are represented by faithful local fakes; every load-bearing
rubric criterion is checkable against the repo.
- **`partial`** — the core of the task is offline-completable and the rubric
mostly grades local substance, but some of the success criteria lean
outside the sandbox: a secondary "and it works in prod"-shaped expectation,
a mock whose fidelity is doing a lot of load-bearing work, a local proxy
(config parses, unit tests pass) standing in for an external outcome (the
pipeline is faster), or a prompt whose natural reading promises more
end-to-end confidence than the sandbox can deliver. The task works; the
author should look at each finding and decide whether to rescope, reword,
or accept the gap knowingly.
- **`not-offline-verifiable`** — either half of the controlling test fails
outright. *Success-criteria form:* the system being operated on (pipeline,
prod, third-party service) isn't there and can't be faithfully faked, so
neither doing the work well nor verifying it can happen in the workspace.
*Completability form:* the ask names a technology the repo carries no
library for, so step one is a registry fetch that rebuilding the image
correctly would not remove. Whether the sandbox happened to permit that
fetch is irrelevant and plays no part in the verdict. A human SWE handed this task in this environment would say "I
can't actually do or check this from here."
- **`not-applicable`** — nothing to assess: `instruction.md` is missing,
empty, or only template/placeholder content, and there is no session
history to read an ask from. Re-run once the prompt lands.
`not-offline-verifiable` and `partial` are the flagged outcomes. Findings on
the *verifiability* half stay advisory: where success criteria live is a
judgment call, and the finding is a consideration for the author. A finding on
the *completability* half is not — whether a named dependency appears in any
manifest is a checked fact. Report it plainly and say which manifests you
read. Only the integral-or-consequential call stands between that fact and
the verdict, and the ask itself settles it: a library the repo never had is
integral, a library the image forgot to install is ours to fix.
## Confidence
- **HIGH** — the call is unambiguous: the success criteria plainly live
outside the sandbox (or plainly don't), and the rubric confirms the
reading.
- **MEDIUM** — at least one finding is genuinely two-sided: a mock whose
fidelity a reasonable reviewer might judge either way, or a prompt that
reads external but a rubric that grades local.
- **LOW** — limited information: the rubric is thin or absent so you can't
tell what's actually verified, or the workspace's representation of the
external side couldn't be assessed.
## Relationship to other detectors
- **vs. detector-broken-dev-env.** That detector owns the *environment being
broken*: the workspace doesn't build, tests flake, artifacts contradict the
premise. This detector fires even when the environment is perfectly healthy
— the defect is that the TASK's success criteria live outside the sandbox.
"The tests won't run" is broken-dev-env; "no test that could run here can
tell you whether this worked" is this detector. A dependency the *ask*
requires but no manifest declares is this detector's (the env is fine, the
ask isn't completable); a dependency the *existing code* imports but no
manifest declares is broken-dev-env's (the env is broken).
- **vs. detector-fact-check-rubric-claims.** Its reachability axis asks
whether a specific *fact* the rubric grades the response for knowing is
reachable from the package. This detector asks the structural version:
whether the task's *success criteria as a whole* are checkable from inside
the sandbox. A rubric criterion "the provider accepts the payload" can
surface in both — as an unreachable/unverifiable claim there, and as an
offline-verifiability finding here.
- **vs. detector-meaningful-failure.** That detector asks whether the graded
failure is real, proportionate, and elicited. A not-offline-verifiable task
often *also* fails to elicit meaningfully (the runs are all pantomime), but
the diagnosis differs: meaningful-failure says "this failure isn't worth
grading"; this detector says "no one inside the sandbox can check the thing
being graded."
- **vs. detector-answer-obviousness.** Unrelated axis (is the expected answer
inferable from the prompt?). No overlap expected; neither subsumes the
other.
## Anti-patterns: do not do these
- **Don't flag every mention of an external service.** Scenario realism
requires them. Trace the graded success criteria; flag only when *they*
live outside.
- **Don't demand hermetic purity.** Nearly every repo talks to something.
The bar is the controlling test — complete AND verify from within the
initialized workspace — not "the prompt never says the word 'deploy'."
- **Don't punish tasks that are honest about the boundary.** A rubric that
credits the agent for saying "this can't be verified from here" has priced
the sandbox in; that's a strength, not a finding.
- **Don't treat a verifiability flag as a verdict on the author or the
task's worth.** That output is something to consider — a pointer at where
the success criteria live — phrased so the author can decide. Never assert
the task is invalid; never frame the finding as a failure. A
missing-dependency finding is the exception: it is a fact about the
manifests, so state it rather than softening it into a consideration.
- **Don't consult the task's network policy.** `allow_internet`,
`network_mode` and `allowed_hosts` exist for the grading harness, not the
agent. Reading them can only mislead you here: nearly every task allows
egress, so weighing it would clear every missing-dependency finding in the
corpus. Judge the repo's manifests against the ask and nothing else.
- **Don't clear a missing dependency because the framework supports it.**
"Rails ships `:redis_cache_store`", "pytest has a coverage plugin" — an
adapter existing upstream says nothing about whether the gem or package is
in this repo's lockfile. Open the manifest.
- **Don't re-litigate env health.** Whether the workspace builds and the
suite passes belongs to detector-broken-dev-env. Assume a healthy env and
ask where the success criteria live — a healthy env does not imply the ask's
own dependencies are present, which is the completability check above.
- **Don't cite evidence you haven't verified in the submitted package.**
Quote the prompt, rubric, and workspace as they exist in the actual
submission — not as you remember or infer them.
## Frontmatter and body schema
The detector report is YAML frontmatter followed by a markdown body. Both
contexts produce the same shape; only the *sink* differs (the wrapping
`SKILL.md` tells you where to send the report).
**Frontmatter** — exactly these keys, exactly these enum values:
```yaml
---
detector: detector-offline-verifiability
verdict: offline-verifiable | partial | not-offline-verifiable | not-applicable
confidence: HIGH | MEDIUM | LOW
---
```
**Body sections**, in this order:
```markdown
# Offline-verifiability check: <slug>
## Findings
One block per finding, strongest first:
### <short label> — <external-dependency shape> (<clear | partial>)
- **Where:** the file (and line/section, or the turn for a session message)
where the ask or success criterion appears.
- **Quote:** the passage verbatim, as a blockquote — never a paraphrase.
- **Why it lives outside:** one or two sentences — what a human SWE would
need the network or a live system for, in doing or verifying this, and
what the sandbox can actually check instead.
- **Something to consider:** a concrete rescoping option — grade the local
protocol slice, reword the ask as a plan/assessment, point the criterion
at the repo's fake provider, credit honest disclosure of the boundary —
worded so the author can decide whether to take it.
For `offline-verifiable`, quote the strongest near-miss (the most
external-sounding passage) and say why it was cleared. For `not-applicable`,
name the missing artifacts.
## Overall verdict
2–3 paragraphs reducing the findings to the chosen verdict: where the
task's success criteria live, whether the workspace (including any local
fakes) can honestly check them, and — for verifiability findings, which are
advisory — what a rescoping pass would consider first. For a
missing-dependency finding, drop the hedging: name the package, name every
manifest and lockfile you checked, and say the ask can't be completed offline
as shipped. For `offline-verifiable`, why
the near-misses are scenario context or faithfully mocked rather than live
dependencies.
```
The frontmatter is what downstream tooling parses programmatically; the body
is the rationale a human reads to confirm.