added stocks app codebase and md

This commit is contained in:
2026-08-10 21:38:03 -04:00
parent 87f070f033
commit 35aa848168
143 changed files with 33558 additions and 0 deletions

View File

@@ -0,0 +1,79 @@
---
name: detector-credential-leakage
description: |
Self-check whether your submission ships credentials or other content from
your authoring environment inside its authored surfaces — above all
`environment/workspace.patch`. Two tiers. (1) **Known-credential tier
(deterministic):** hard-flags your authoring environment's own env vars
(`ANTHROPIC_API_KEY`, `ANTHROPIC_BASE_URL`, `USER_ID` as an env
assignment) and well-known secret shapes (`sk-ant-…`, AWS `AKIA…`, GitHub
`ghp_…`, Google `AIza…`, Stripe secret keys, bearer tokens, private-key
blocks, URL-embedded passwords) on lines your patch adds. The canonical
incident: your toolkit `.env` — your personal API key, proxy URL, and user
id — swept into the workspace as a new `.env` file. (2) **Task-relevance
tier (judgment):** content your patch adds that doesn't appear to serve
the task — `.env`-style files, env-file symlinks into your home directory,
`export FOO=` lines, credential-shaped assignments with real values. A
`credential-leak` finding must be acted on before submitting (remove the
material AND report the key as compromised so it can be rotated);
`suspicious-content` findings are advisory. The report never reproduces
secret values. Reads workspace.patch (+ Dockerfile, instruction.md,
tests/*.md); runs before or after reference runs exist.
allowed-tools: Bash, Read, Write
---
# Credential-leakage detector
This skill checks one of your tasks for **credential leakage** — whether
anything from your own authoring environment (or any other secret) has been
swept into the submission's authored surfaces, above all
`environment/workspace.patch`. Everything your patch adds ships to everyone
downstream, so a leaked key is compromised the moment you submit: deleting the
line later does not un-ship it.
The failure shape to catch: your toolkit's `.env` — the file holding your
personal `ANTHROPIC_API_KEY`, `ANTHROPIC_BASE_URL`, and `USER_ID` — landing in
the workspace as a new `.env` file (or a `.env.bak-*` backup, or a symlink to
`/home/<you>/.env`). It happens easily: a stray `git add`, a working-tree
backup, a captured terminal snippet. None of it serves the task; the test
agent has no network to use a key with; and the key is now distributed.
What *doesn't* trip this check: placeholder and example values
(`.env.example` with empty or dummy entries, `sk-ant-...` as a literal
template, `changeme`), dev-infrastructure defaults (`POSTGRES_PASSWORD=postgres`
in a local docker-compose), code identifiers (`USER_ID = 4958` as a test
constant), and env vars your task's scenario genuinely needs documented.
**Severity differs by tier.** Unlike most self-checks, a `credential-leak`
finding is not a consideration: remove the material from the patch, rebuild it
(`bash scripts/check-workspace-sync.sh --update-patch harbor-tasks/<slug>`),
and report the leaked credential through your support channel so it can be
rotated — treat it as compromised even after you scrub it.
`suspicious-content` findings are the usual advisory kind: read each one and
fix or justify it.
Read these before deciding:
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
2. `.claude/skills/detector-credential-leakage/core.md` — the two tiers, the deterministic pattern checks to run, the placeholder test, the redaction rule (never quote a secret value), what is NOT a finding, verdict enums, and the body schema.
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
## Acting on the verdict
- **`clean`** — nothing your patch adds looks like a credential or foreign
content. Good. Move on.
- **`suspicious-content`** — no confirmed credential, but something your patch
adds doesn't look like it belongs to the task: an env-file symlink into your
home directory, a captured request with a real (if low-sensitivity) token, a
config file of credential-shaped values. Fix each finding (replace tokens
with placeholders, drop the file, or make its task relevance explicit) or
satisfy yourself it's genuinely scenario material.
- **`credential-leak`** — a real credential or your authoring environment's
own env vars are in the patch. Act before submitting: (1) remove the
material and regenerate `workspace.patch`; (2) re-run this detector to
confirm it's gone; (3) report the leaked value as compromised so it can be
rotated — scrubbing the patch does not un-ship a key that already left your
machine in an earlier submission.
- **`not-applicable`** — there's no workspace patch to assess yet. Build the
workspace first.

View File

@@ -0,0 +1,303 @@
# Credential-leakage detector — core
This file is the canonical, context-neutral content for the
detector-credential-leakage detector. It defines the signal (authoring-environment
credentials or other task-irrelevant content shipped inside the submission's
authored surfaces), the two tiers of the check, the deterministic patterns, the
verdict enums, and the output schema. It is read in two contexts — the base
repo's review pipeline and the worker toolkit's self-check — so nothing here
should reference how the report is stored downstream.
## What this detector is for
Everything a task adds to the workspace ships to everyone downstream: the test
agent reads it, graders read it, and the patch text itself travels with the
submission. The task author's *authoring environment*, though, contains things
that must never make that trip — most importantly the author's own credentials.
The canonical incident: a `workspace.patch` that adds a `.env` file containing
```
ANTHROPIC_API_KEY=DKRY…[redacted]
ANTHROPIC_BASE_URL=https://…/llm_proxy/…
USER_ID=6428…[redacted]
```
— the author's personal LLM-proxy API key, proxy endpoint, and user identity,
swept out of their authoring container and checked into the task. Nothing about
the task needs these; the agent under test can't use them (no network); and the
key is now distributed to every downstream consumer of the task. The same
mechanism sweeps in other authoring-environment artifacts: a `.env` symlink
pointing at the author's home directory, a backup copy of a modified env file
(`.env.bak-*`) full of real third-party secrets, a captured HTTP request with a
live bearer token.
Two tiers, one report:
1. **Known-credential tier (deterministic).** Specific, unambiguous signatures
of authoring-environment credentials and well-known secret shapes,
detected by running fixed pattern checks — not judgment. Any hit here is a
`credential-leak`.
2. **Task-relevance tier (judgment).** Content added by `workspace.patch` that
doesn't appear to serve the task — particularly env-var or credential-shaped
content: new `.env`-style files, `export FOO=` lines in scripts the task
never uses, credential assignments in config files, absolute paths into
somebody's home directory. Judged against what the task is actually about.
**Severity differs by tier.** A `credential-leak` finding is not a style
consideration: the leaked material must be removed from the submission, and any
real credential in it treated as compromised (reported so it can be rotated) —
deleting the line later does not un-ship the key. The `suspicious-content` tier
is advisory in the usual way: each finding is something for the author to look
at and decide, since plenty of env-var-shaped content is legitimate task
material.
## NEVER quote secret values — redact
This detector's report is itself distributed. Reproducing a leaked value in the
report would spread the leak further. **Never copy a candidate secret value into
the report.** Quote the variable name, the file path, and at most the first 4
characters of the value followed by `…[redacted]`:
> `ANTHROPIC_API_KEY=DKRY…[redacted]` in `.env` (new file, line 1)
This overrides the sibling detectors' quote-verbatim convention — for this
detector, redaction wins.
## Inputs
Read from `harbor-tasks/<slug>/`:
- `environment/workspace.patch` — the primary surface: the diff of files the
task adds to or edits in the workspace. **Added lines and newly added files
are the authored surface.** Also scan the *whole* patch text for secret
shapes: a secret on a context or removed line is pre-existing repo content
(see "What is NOT a finding"), but it still ships in the patch, so it is
worth an informational note.
- `environment/Dockerfile` — task-owned build steps can carry `ENV`/`ARG`
credentials the same way.
- `instruction.md` and `tests/*.md` — secondary authored surfaces; a pasted
terminal capture or setup snippet can carry the same leak.
- `task.toml` — context only: what the task is about, which informs the
relevance judgment in tier 2.
## Tier 1 — known credentials and secret shapes (deterministic)
Run these checks over the patch. Treat the pattern list as the contract: a hit
on an **added** line (or a newly added file) is a `credential-leak` finding
unless it is an explicit placeholder (see the placeholder test below).
```bash
# Authoring-environment env vars, on added lines:
grep -nE '^\+' environment/workspace.patch \
| grep -E 'ANTHROPIC_[A-Z_]+[[:space:]]*[=:]|(^|[^A-Za-z0-9_.])USER_ID[[:space:]]*='
# Well-known secret shapes, over the WHOLE patch (added hits are findings;
# context/removed hits are informational notes):
grep -nE 'sk-ant-[A-Za-z0-9_-]{8,}|AKIA[0-9A-Z]{16}|(ghp|gho|ghu|ghs|ghr)_[A-Za-z0-9]{20,}|github_pat_[A-Za-z0-9_]{20,}|xox[baprs]-[A-Za-z0-9-]{10,}|AIza[0-9A-Za-z_-]{35}|sk_(live|test)_[A-Za-z0-9]{16,}|-----BEGIN [A-Z ]*PRIVATE KEY-----|[Aa]uthorization[^A-Za-z0-9]{0,3}Bearer [A-Za-z0-9._~+/=-]{20,}|[a-z][a-z0-9+.-]*://[^/:@[:space:]]{3,}:[^@[:space:]]{8,}@' \
environment/workspace.patch
# LLM-proxy endpoints from the authoring environment:
grep -nE '^\+' environment/workspace.patch | grep -iE 'llm[_-]?proxy|dataannotation\.tech'
```
The named env vars to hard-flag on added lines:
- **`ANTHROPIC_API_KEY`** (or any `ANTHROPIC_*` var carrying a value) — the
author's personal API credential.
- **`ANTHROPIC_BASE_URL`** — the authoring environment's LLM-proxy endpoint;
not a secret by itself, but pure authoring-environment plumbing that has no
business in a task workspace, and its presence marks the leak.
- **`USER_ID`** *as an env-var assignment* (a `.env` line, `export USER_ID=`,
`ENV USER_ID=`, especially with a UUID value) — the author's platform
identity. `USER_ID` / `user_id` as a *code identifier* (a column, a variable,
a test constant like `USER_ID = 4958`) is normal code, not a leak — the flag
is the env-assignment shape.
**The placeholder test.** A pattern hit whose value is plainly not real is not
a leak: empty (`QBO_SECRET=`), a template marker (`sk-ant-...`, `<your-key>`,
`${STRIPE_KEY}`, `changeme`, `your-key-here`), a documented dummy the repo
already uses in fixtures, or a commented-out no-value line in an
`.env.example`. When in doubt — the value looks high-entropy and real — flag
it; a false "compromised" alarm is far cheaper than a shipped key.
## Tier 2 — task-relevance judgment (advisory)
For everything else the patch adds, ask: **does this content serve the task,
or did it fall in from the author's environment?** Shapes to look at:
- **New env-style files** — `.env`, `.env.local`, `.env.bak*`, or a config
file of credentials the task never references. A new `.env.example` with
placeholder values that documents setup the task genuinely needs is fine;
a backup of somebody's real env file is not.
- **Env files added as symlinks** — a patch adding `.env -> /home/<user>/.env`
ships an absolute path into the author's machine and dangles in the
sandbox. Pure authoring artifact.
- **`export FOO=value` lines** added to scripts, docs, or shell profiles —
legitimate when the task's setup genuinely needs them (`export
PATH=…`, parameterized `${VAR:-default}` deploy scripts), suspicious when
they set credentials or author-specific values.
- **Credential-shaped assignments in added code/config/fixtures** —
`*_API_KEY`, `*_SECRET`, `*_TOKEN`, `*_PASSWORD` set to real-looking
(high-entropy, non-placeholder) values: recorded HTTP fixtures carrying live
`Authorization` headers, a pasted curl with a real bearer token, a
docker-compose with a non-dev password.
- **Other authoring-environment artifacts** — absolute paths into a home
directory, editor/agent config files (`.claude/`, `.vscode/` state), shell
history, tool caches: content whose only plausible origin is the author's
working environment rather than the task's scenario.
The controlling question is relevance, not vocabulary. A task about payment
webhooks legitimately adds webhook-secret *placeholders*; a task about i18n
that adds a translation script legitimately documents the env var the script
reads. The finding is content whose presence the task cannot explain.
## What is NOT a finding
- **Placeholder and example values.** `.env.example` / `.env.sample` /
`.env.test` files with empty or dummy values, `sk_test`-style fixture
strings the repo's test suite already uses as fakes, `changeme`,
`dev-insecure-session-secret-change-me`, `${VAR:-default}` expansions.
- **Dev-infrastructure defaults.** `POSTGRES_PASSWORD=postgres` in a local
docker-compose, `SESSION_SECRET: dev-…` in a dev config, a `bin/dev`
exporting `BINDING=0.0.0.0` — local-only, value-free-by-convention.
- **Code identifiers.** `USER_ID` as a constant, column, or variable in
code or tests; `ANTHROPIC_VERSION`-style constants in an app that
genuinely integrates an LLM API as its product feature.
- **Env vars the task's own scenario needs.** If the repo's product calls an
external API and the task is about that integration, documenting the env
var (with a placeholder value) is task material.
- **Pre-existing repo content.** Secrets on *context or removed* lines of the
patch were committed by the source repo, not the author — the author
removing one is good hygiene, not a leak. Don't flag the author; DO add an
informational note (the secret still ships inside the patch text, and the
repo owner should hear about it).
- **A task whose subject IS a leaked credential.** A scenario can plant a
fake "leaked key" for the agent to find. The planted value should still be
fake; flag only if it's real.
## Verdict definitions
- **`clean`** — no tier-1 hit survives the placeholder test, and nothing the
patch adds looks foreign to the task. Placeholder env files, dev defaults,
and scenario-relevant env vars are all clean (see the list above).
- **`suspicious-content`** — no confirmed credential, but the patch carries
content that doesn't look like it belongs to the task: an env-file symlink
into a home directory, a real-looking-but-low-sensitivity token (a
public-by-design client token, a locally-signed dev JWT), an unexplained
env/config addition. Advisory: each finding is for the author to resolve
or justify.
- **`credential-leak`** — a tier-1 pattern hit on added content survives the
placeholder test: a named authoring-environment variable carrying a value,
or a known secret shape. This is the strong form: the material must be
removed from the submission and any real credential in it treated as
compromised and reported for rotation. Scrubbing the patch alone is not
sufficient remediation for the key itself.
- **`not-applicable`** — nothing to assess: no `environment/workspace.patch`
(and no authored Dockerfile/doc surfaces) exists yet. Re-run once the
workspace lands.
`credential-leak` and `suspicious-content` are the flagged outcomes.
`suspicious-content` is advisory in the usual way; `credential-leak` is the
one finding in this detector that is not a judgment call to sit on — it
should be acted on before the task ships.
## Confidence
- **HIGH** — a tier-1 hit with a real-looking value (or plainly nothing
anywhere): the deterministic tier makes most calls HIGH by construction.
- **MEDIUM** — the call rests on tier-2 judgment a reasonable reviewer could
make either way: a token that may be public-by-design, an env file whose
values might all be dummies, content whose task relevance is arguable.
- **LOW** — limited information: the patch is enormous and only sampled, or
the task's subject couldn't be established well enough to judge relevance.
## Relationship to other detectors
- **vs. detector-over-hinting.** Same primary surface (`workspace.patch`
additions), different defect: over-hinting reads authored *comments* for
content that does the agent's thinking; this detector reads authored
content for material that belongs to the author's environment, not the
task. A file can trip both; the verdicts are independent.
- **vs. detector-snapshot-leakage.** "Leakage" there means the *answer*
leaking to the test agent through the inherited session. Here it means the
*author's credentials* leaking into the shipped workspace. No overlap in
substance; the shared word is coincidence.
- **vs. detector-broken-dev-env.** A dangling `.env` symlink or a bogus env
file can also break the workspace at runtime — that detector owns the
build/run consequences; this one owns the provenance/exposure question.
Expect both to fire on the same artifact occasionally, each with its own
rationale.
## Anti-patterns: do not do these
- **Never reproduce a secret value in the report.** Redact to a 4-character
stub. This is the detector's own hygiene bar; failing it is worse than a
missed finding.
- **Don't flag vocabulary.** `SECRET`, `TOKEN`, `PASSWORD` in a variable
name is not a finding; a real-looking *value* is. Run the placeholder test
before flagging anything.
- **Don't flag pre-existing repo secrets as author leaks.** Context and
removed lines belong to the source repo. Note them informationally;
attribute them correctly.
- **Don't soften a tier-1 hit into advice.** A real key in the patch is not
"something to consider" — say plainly that it must be removed and the
credential rotated.
- **Don't skip the deterministic tier because the patch "looks clean".**
Run the pattern checks; the canonical incident sat in plain sight at the
top of the patch.
- **Don't cite evidence you haven't verified in the submitted package.**
Point at the actual file and line in the actual patch — not at what you
remember or infer.
## Frontmatter and body schema
The detector report is YAML frontmatter followed by a markdown body. Both
contexts produce the same shape; only the *sink* differs (the wrapping
`SKILL.md` tells you where to send the report).
**Frontmatter** — exactly these keys, exactly these enum values:
```yaml
---
detector: detector-credential-leakage
verdict: credential-leak | suspicious-content | clean | not-applicable
confidence: HIGH | MEDIUM | LOW
---
```
**Body sections**, in this order:
```markdown
# Credential-leakage check: <slug>
## Findings
One block per finding, strongest first:
### <short label> — <known-credential | task-relevance> (<leak | suspicious | informational>)
- **Where:** the file and line (patch hunk) where the content appears, and
whether the line is added, context, or removed.
- **What:** the variable name(s) / content shape, with every value REDACTED
to at most 4 characters + `…[redacted]`. Never the full value.
- **Why it doesn't belong:** one or two sentences — what marks this as
authoring-environment material or task-irrelevant, and (for tier 1) which
pattern hit.
- **Action:** for a leak — remove the material from the patch AND treat the
credential as compromised (report it for rotation). For suspicious
content — the concrete fix or the justification that would clear it.
For `clean`, name the strongest near-miss (a placeholder env file, a dev
default) and say why the placeholder test cleared it. For `not-applicable`,
name the missing artifacts.
## Overall verdict
2–3 paragraphs reducing the findings to the chosen verdict: what the patch
adds that shouldn't ship, which tier the strongest finding sits in, and what
remediation looks like — including, for any real credential, that removal
from the patch does not un-ship it and rotation is the actual fix.
```
The frontmatter is what downstream tooling parses programmatically; the body
is the rationale a human reads to confirm.