remove folder - untrustworthy
This commit is contained in:
@@ -1,79 +0,0 @@
|
||||
---
|
||||
name: detector-credential-leakage
|
||||
description: |
|
||||
Self-check whether your submission ships credentials or other content from
|
||||
your authoring environment inside its authored surfaces — above all
|
||||
`environment/workspace.patch`. Two tiers. (1) **Known-credential tier
|
||||
(deterministic):** hard-flags your authoring environment's own env vars
|
||||
(`ANTHROPIC_API_KEY`, `ANTHROPIC_BASE_URL`, `USER_ID` as an env
|
||||
assignment) and well-known secret shapes (`sk-ant-…`, AWS `AKIA…`, GitHub
|
||||
`ghp_…`, Google `AIza…`, Stripe secret keys, bearer tokens, private-key
|
||||
blocks, URL-embedded passwords) on lines your patch adds. The canonical
|
||||
incident: your toolkit `.env` — your personal API key, proxy URL, and user
|
||||
id — swept into the workspace as a new `.env` file. (2) **Task-relevance
|
||||
tier (judgment):** content your patch adds that doesn't appear to serve
|
||||
the task — `.env`-style files, env-file symlinks into your home directory,
|
||||
`export FOO=` lines, credential-shaped assignments with real values. A
|
||||
`credential-leak` finding must be acted on before submitting (remove the
|
||||
material AND report the key as compromised so it can be rotated);
|
||||
`suspicious-content` findings are advisory. The report never reproduces
|
||||
secret values. Reads workspace.patch (+ Dockerfile, instruction.md,
|
||||
tests/*.md); runs before or after reference runs exist.
|
||||
allowed-tools: Bash, Read, Write
|
||||
---
|
||||
|
||||
# Credential-leakage detector
|
||||
|
||||
This skill checks one of your tasks for **credential leakage** — whether
|
||||
anything from your own authoring environment (or any other secret) has been
|
||||
swept into the submission's authored surfaces, above all
|
||||
`environment/workspace.patch`. Everything your patch adds ships to everyone
|
||||
downstream, so a leaked key is compromised the moment you submit: deleting the
|
||||
line later does not un-ship it.
|
||||
|
||||
The failure shape to catch: your toolkit's `.env` — the file holding your
|
||||
personal `ANTHROPIC_API_KEY`, `ANTHROPIC_BASE_URL`, and `USER_ID` — landing in
|
||||
the workspace as a new `.env` file (or a `.env.bak-*` backup, or a symlink to
|
||||
`/home/<you>/.env`). It happens easily: a stray `git add`, a working-tree
|
||||
backup, a captured terminal snippet. None of it serves the task; the test
|
||||
agent has no network to use a key with; and the key is now distributed.
|
||||
|
||||
What *doesn't* trip this check: placeholder and example values
|
||||
(`.env.example` with empty or dummy entries, `sk-ant-...` as a literal
|
||||
template, `changeme`), dev-infrastructure defaults (`POSTGRES_PASSWORD=postgres`
|
||||
in a local docker-compose), code identifiers (`USER_ID = 4958` as a test
|
||||
constant), and env vars your task's scenario genuinely needs documented.
|
||||
|
||||
**Severity differs by tier.** Unlike most self-checks, a `credential-leak`
|
||||
finding is not a consideration: remove the material from the patch, rebuild it
|
||||
(`bash scripts/check-workspace-sync.sh --update-patch harbor-tasks/<slug>`),
|
||||
and report the leaked credential through your support channel so it can be
|
||||
rotated — treat it as compromised even after you scrub it.
|
||||
`suspicious-content` findings are the usual advisory kind: read each one and
|
||||
fix or justify it.
|
||||
|
||||
Read these before deciding:
|
||||
|
||||
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
|
||||
2. `.claude/skills/detector-credential-leakage/core.md` — the two tiers, the deterministic pattern checks to run, the placeholder test, the redaction rule (never quote a secret value), what is NOT a finding, verdict enums, and the body schema.
|
||||
|
||||
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
|
||||
|
||||
## Acting on the verdict
|
||||
|
||||
- **`clean`** — nothing your patch adds looks like a credential or foreign
|
||||
content. Good. Move on.
|
||||
- **`suspicious-content`** — no confirmed credential, but something your patch
|
||||
adds doesn't look like it belongs to the task: an env-file symlink into your
|
||||
home directory, a captured request with a real (if low-sensitivity) token, a
|
||||
config file of credential-shaped values. Fix each finding (replace tokens
|
||||
with placeholders, drop the file, or make its task relevance explicit) or
|
||||
satisfy yourself it's genuinely scenario material.
|
||||
- **`credential-leak`** — a real credential or your authoring environment's
|
||||
own env vars are in the patch. Act before submitting: (1) remove the
|
||||
material and regenerate `workspace.patch`; (2) re-run this detector to
|
||||
confirm it's gone; (3) report the leaked value as compromised so it can be
|
||||
rotated — scrubbing the patch does not un-ship a key that already left your
|
||||
machine in an earlier submission.
|
||||
- **`not-applicable`** — there's no workspace patch to assess yet. Build the
|
||||
workspace first.
|
||||
@@ -1,303 +0,0 @@
|
||||
# Credential-leakage detector — core
|
||||
|
||||
This file is the canonical, context-neutral content for the
|
||||
detector-credential-leakage detector. It defines the signal (authoring-environment
|
||||
credentials or other task-irrelevant content shipped inside the submission's
|
||||
authored surfaces), the two tiers of the check, the deterministic patterns, the
|
||||
verdict enums, and the output schema. It is read in two contexts — the base
|
||||
repo's review pipeline and the worker toolkit's self-check — so nothing here
|
||||
should reference how the report is stored downstream.
|
||||
|
||||
## What this detector is for
|
||||
|
||||
Everything a task adds to the workspace ships to everyone downstream: the test
|
||||
agent reads it, graders read it, and the patch text itself travels with the
|
||||
submission. The task author's *authoring environment*, though, contains things
|
||||
that must never make that trip — most importantly the author's own credentials.
|
||||
The canonical incident: a `workspace.patch` that adds a `.env` file containing
|
||||
|
||||
```
|
||||
ANTHROPIC_API_KEY=DKRY…[redacted]
|
||||
ANTHROPIC_BASE_URL=https://…/llm_proxy/…
|
||||
USER_ID=6428…[redacted]
|
||||
```
|
||||
|
||||
— the author's personal LLM-proxy API key, proxy endpoint, and user identity,
|
||||
swept out of their authoring container and checked into the task. Nothing about
|
||||
the task needs these; the agent under test can't use them (no network); and the
|
||||
key is now distributed to every downstream consumer of the task. The same
|
||||
mechanism sweeps in other authoring-environment artifacts: a `.env` symlink
|
||||
pointing at the author's home directory, a backup copy of a modified env file
|
||||
(`.env.bak-*`) full of real third-party secrets, a captured HTTP request with a
|
||||
live bearer token.
|
||||
|
||||
Two tiers, one report:
|
||||
|
||||
1. **Known-credential tier (deterministic).** Specific, unambiguous signatures
|
||||
of authoring-environment credentials and well-known secret shapes,
|
||||
detected by running fixed pattern checks — not judgment. Any hit here is a
|
||||
`credential-leak`.
|
||||
2. **Task-relevance tier (judgment).** Content added by `workspace.patch` that
|
||||
doesn't appear to serve the task — particularly env-var or credential-shaped
|
||||
content: new `.env`-style files, `export FOO=` lines in scripts the task
|
||||
never uses, credential assignments in config files, absolute paths into
|
||||
somebody's home directory. Judged against what the task is actually about.
|
||||
|
||||
**Severity differs by tier.** A `credential-leak` finding is not a style
|
||||
consideration: the leaked material must be removed from the submission, and any
|
||||
real credential in it treated as compromised (reported so it can be rotated) —
|
||||
deleting the line later does not un-ship the key. The `suspicious-content` tier
|
||||
is advisory in the usual way: each finding is something for the author to look
|
||||
at and decide, since plenty of env-var-shaped content is legitimate task
|
||||
material.
|
||||
|
||||
## NEVER quote secret values — redact
|
||||
|
||||
This detector's report is itself distributed. Reproducing a leaked value in the
|
||||
report would spread the leak further. **Never copy a candidate secret value into
|
||||
the report.** Quote the variable name, the file path, and at most the first 4
|
||||
characters of the value followed by `…[redacted]`:
|
||||
|
||||
> `ANTHROPIC_API_KEY=DKRY…[redacted]` in `.env` (new file, line 1)
|
||||
|
||||
This overrides the sibling detectors' quote-verbatim convention — for this
|
||||
detector, redaction wins.
|
||||
|
||||
## Inputs
|
||||
|
||||
Read from `harbor-tasks/<slug>/`:
|
||||
|
||||
- `environment/workspace.patch` — the primary surface: the diff of files the
|
||||
task adds to or edits in the workspace. **Added lines and newly added files
|
||||
are the authored surface.** Also scan the *whole* patch text for secret
|
||||
shapes: a secret on a context or removed line is pre-existing repo content
|
||||
(see "What is NOT a finding"), but it still ships in the patch, so it is
|
||||
worth an informational note.
|
||||
- `environment/Dockerfile` — task-owned build steps can carry `ENV`/`ARG`
|
||||
credentials the same way.
|
||||
- `instruction.md` and `tests/*.md` — secondary authored surfaces; a pasted
|
||||
terminal capture or setup snippet can carry the same leak.
|
||||
- `task.toml` — context only: what the task is about, which informs the
|
||||
relevance judgment in tier 2.
|
||||
|
||||
## Tier 1 — known credentials and secret shapes (deterministic)
|
||||
|
||||
Run these checks over the patch. Treat the pattern list as the contract: a hit
|
||||
on an **added** line (or a newly added file) is a `credential-leak` finding
|
||||
unless it is an explicit placeholder (see the placeholder test below).
|
||||
|
||||
```bash
|
||||
# Authoring-environment env vars, on added lines:
|
||||
grep -nE '^\+' environment/workspace.patch \
|
||||
| grep -E 'ANTHROPIC_[A-Z_]+[[:space:]]*[=:]|(^|[^A-Za-z0-9_.])USER_ID[[:space:]]*='
|
||||
|
||||
# Well-known secret shapes, over the WHOLE patch (added hits are findings;
|
||||
# context/removed hits are informational notes):
|
||||
grep -nE 'sk-ant-[A-Za-z0-9_-]{8,}|AKIA[0-9A-Z]{16}|(ghp|gho|ghu|ghs|ghr)_[A-Za-z0-9]{20,}|github_pat_[A-Za-z0-9_]{20,}|xox[baprs]-[A-Za-z0-9-]{10,}|AIza[0-9A-Za-z_-]{35}|sk_(live|test)_[A-Za-z0-9]{16,}|-----BEGIN [A-Z ]*PRIVATE KEY-----|[Aa]uthorization[^A-Za-z0-9]{0,3}Bearer [A-Za-z0-9._~+/=-]{20,}|[a-z][a-z0-9+.-]*://[^/:@[:space:]]{3,}:[^@[:space:]]{8,}@' \
|
||||
environment/workspace.patch
|
||||
|
||||
# LLM-proxy endpoints from the authoring environment:
|
||||
grep -nE '^\+' environment/workspace.patch | grep -iE 'llm[_-]?proxy|dataannotation\.tech'
|
||||
```
|
||||
|
||||
The named env vars to hard-flag on added lines:
|
||||
|
||||
- **`ANTHROPIC_API_KEY`** (or any `ANTHROPIC_*` var carrying a value) — the
|
||||
author's personal API credential.
|
||||
- **`ANTHROPIC_BASE_URL`** — the authoring environment's LLM-proxy endpoint;
|
||||
not a secret by itself, but pure authoring-environment plumbing that has no
|
||||
business in a task workspace, and its presence marks the leak.
|
||||
- **`USER_ID`** *as an env-var assignment* (a `.env` line, `export USER_ID=`,
|
||||
`ENV USER_ID=`, especially with a UUID value) — the author's platform
|
||||
identity. `USER_ID` / `user_id` as a *code identifier* (a column, a variable,
|
||||
a test constant like `USER_ID = 4958`) is normal code, not a leak — the flag
|
||||
is the env-assignment shape.
|
||||
|
||||
**The placeholder test.** A pattern hit whose value is plainly not real is not
|
||||
a leak: empty (`QBO_SECRET=`), a template marker (`sk-ant-...`, `<your-key>`,
|
||||
`${STRIPE_KEY}`, `changeme`, `your-key-here`), a documented dummy the repo
|
||||
already uses in fixtures, or a commented-out no-value line in an
|
||||
`.env.example`. When in doubt — the value looks high-entropy and real — flag
|
||||
it; a false "compromised" alarm is far cheaper than a shipped key.
|
||||
|
||||
## Tier 2 — task-relevance judgment (advisory)
|
||||
|
||||
For everything else the patch adds, ask: **does this content serve the task,
|
||||
or did it fall in from the author's environment?** Shapes to look at:
|
||||
|
||||
- **New env-style files** — `.env`, `.env.local`, `.env.bak*`, or a config
|
||||
file of credentials the task never references. A new `.env.example` with
|
||||
placeholder values that documents setup the task genuinely needs is fine;
|
||||
a backup of somebody's real env file is not.
|
||||
- **Env files added as symlinks** — a patch adding `.env -> /home/<user>/.env`
|
||||
ships an absolute path into the author's machine and dangles in the
|
||||
sandbox. Pure authoring artifact.
|
||||
- **`export FOO=value` lines** added to scripts, docs, or shell profiles —
|
||||
legitimate when the task's setup genuinely needs them (`export
|
||||
PATH=…`, parameterized `${VAR:-default}` deploy scripts), suspicious when
|
||||
they set credentials or author-specific values.
|
||||
- **Credential-shaped assignments in added code/config/fixtures** —
|
||||
`*_API_KEY`, `*_SECRET`, `*_TOKEN`, `*_PASSWORD` set to real-looking
|
||||
(high-entropy, non-placeholder) values: recorded HTTP fixtures carrying live
|
||||
`Authorization` headers, a pasted curl with a real bearer token, a
|
||||
docker-compose with a non-dev password.
|
||||
- **Other authoring-environment artifacts** — absolute paths into a home
|
||||
directory, editor/agent config files (`.claude/`, `.vscode/` state), shell
|
||||
history, tool caches: content whose only plausible origin is the author's
|
||||
working environment rather than the task's scenario.
|
||||
|
||||
The controlling question is relevance, not vocabulary. A task about payment
|
||||
webhooks legitimately adds webhook-secret *placeholders*; a task about i18n
|
||||
that adds a translation script legitimately documents the env var the script
|
||||
reads. The finding is content whose presence the task cannot explain.
|
||||
|
||||
## What is NOT a finding
|
||||
|
||||
- **Placeholder and example values.** `.env.example` / `.env.sample` /
|
||||
`.env.test` files with empty or dummy values, `sk_test`-style fixture
|
||||
strings the repo's test suite already uses as fakes, `changeme`,
|
||||
`dev-insecure-session-secret-change-me`, `${VAR:-default}` expansions.
|
||||
- **Dev-infrastructure defaults.** `POSTGRES_PASSWORD=postgres` in a local
|
||||
docker-compose, `SESSION_SECRET: dev-…` in a dev config, a `bin/dev`
|
||||
exporting `BINDING=0.0.0.0` — local-only, value-free-by-convention.
|
||||
- **Code identifiers.** `USER_ID` as a constant, column, or variable in
|
||||
code or tests; `ANTHROPIC_VERSION`-style constants in an app that
|
||||
genuinely integrates an LLM API as its product feature.
|
||||
- **Env vars the task's own scenario needs.** If the repo's product calls an
|
||||
external API and the task is about that integration, documenting the env
|
||||
var (with a placeholder value) is task material.
|
||||
- **Pre-existing repo content.** Secrets on *context or removed* lines of the
|
||||
patch were committed by the source repo, not the author — the author
|
||||
removing one is good hygiene, not a leak. Don't flag the author; DO add an
|
||||
informational note (the secret still ships inside the patch text, and the
|
||||
repo owner should hear about it).
|
||||
- **A task whose subject IS a leaked credential.** A scenario can plant a
|
||||
fake "leaked key" for the agent to find. The planted value should still be
|
||||
fake; flag only if it's real.
|
||||
|
||||
## Verdict definitions
|
||||
|
||||
- **`clean`** — no tier-1 hit survives the placeholder test, and nothing the
|
||||
patch adds looks foreign to the task. Placeholder env files, dev defaults,
|
||||
and scenario-relevant env vars are all clean (see the list above).
|
||||
- **`suspicious-content`** — no confirmed credential, but the patch carries
|
||||
content that doesn't look like it belongs to the task: an env-file symlink
|
||||
into a home directory, a real-looking-but-low-sensitivity token (a
|
||||
public-by-design client token, a locally-signed dev JWT), an unexplained
|
||||
env/config addition. Advisory: each finding is for the author to resolve
|
||||
or justify.
|
||||
- **`credential-leak`** — a tier-1 pattern hit on added content survives the
|
||||
placeholder test: a named authoring-environment variable carrying a value,
|
||||
or a known secret shape. This is the strong form: the material must be
|
||||
removed from the submission and any real credential in it treated as
|
||||
compromised and reported for rotation. Scrubbing the patch alone is not
|
||||
sufficient remediation for the key itself.
|
||||
- **`not-applicable`** — nothing to assess: no `environment/workspace.patch`
|
||||
(and no authored Dockerfile/doc surfaces) exists yet. Re-run once the
|
||||
workspace lands.
|
||||
|
||||
`credential-leak` and `suspicious-content` are the flagged outcomes.
|
||||
`suspicious-content` is advisory in the usual way; `credential-leak` is the
|
||||
one finding in this detector that is not a judgment call to sit on — it
|
||||
should be acted on before the task ships.
|
||||
|
||||
## Confidence
|
||||
|
||||
- **HIGH** — a tier-1 hit with a real-looking value (or plainly nothing
|
||||
anywhere): the deterministic tier makes most calls HIGH by construction.
|
||||
- **MEDIUM** — the call rests on tier-2 judgment a reasonable reviewer could
|
||||
make either way: a token that may be public-by-design, an env file whose
|
||||
values might all be dummies, content whose task relevance is arguable.
|
||||
- **LOW** — limited information: the patch is enormous and only sampled, or
|
||||
the task's subject couldn't be established well enough to judge relevance.
|
||||
|
||||
## Relationship to other detectors
|
||||
|
||||
- **vs. detector-over-hinting.** Same primary surface (`workspace.patch`
|
||||
additions), different defect: over-hinting reads authored *comments* for
|
||||
content that does the agent's thinking; this detector reads authored
|
||||
content for material that belongs to the author's environment, not the
|
||||
task. A file can trip both; the verdicts are independent.
|
||||
- **vs. detector-snapshot-leakage.** "Leakage" there means the *answer*
|
||||
leaking to the test agent through the inherited session. Here it means the
|
||||
*author's credentials* leaking into the shipped workspace. No overlap in
|
||||
substance; the shared word is coincidence.
|
||||
- **vs. detector-broken-dev-env.** A dangling `.env` symlink or a bogus env
|
||||
file can also break the workspace at runtime — that detector owns the
|
||||
build/run consequences; this one owns the provenance/exposure question.
|
||||
Expect both to fire on the same artifact occasionally, each with its own
|
||||
rationale.
|
||||
|
||||
## Anti-patterns: do not do these
|
||||
|
||||
- **Never reproduce a secret value in the report.** Redact to a 4-character
|
||||
stub. This is the detector's own hygiene bar; failing it is worse than a
|
||||
missed finding.
|
||||
- **Don't flag vocabulary.** `SECRET`, `TOKEN`, `PASSWORD` in a variable
|
||||
name is not a finding; a real-looking *value* is. Run the placeholder test
|
||||
before flagging anything.
|
||||
- **Don't flag pre-existing repo secrets as author leaks.** Context and
|
||||
removed lines belong to the source repo. Note them informationally;
|
||||
attribute them correctly.
|
||||
- **Don't soften a tier-1 hit into advice.** A real key in the patch is not
|
||||
"something to consider" — say plainly that it must be removed and the
|
||||
credential rotated.
|
||||
- **Don't skip the deterministic tier because the patch "looks clean".**
|
||||
Run the pattern checks; the canonical incident sat in plain sight at the
|
||||
top of the patch.
|
||||
- **Don't cite evidence you haven't verified in the submitted package.**
|
||||
Point at the actual file and line in the actual patch — not at what you
|
||||
remember or infer.
|
||||
|
||||
## Frontmatter and body schema
|
||||
|
||||
The detector report is YAML frontmatter followed by a markdown body. Both
|
||||
contexts produce the same shape; only the *sink* differs (the wrapping
|
||||
`SKILL.md` tells you where to send the report).
|
||||
|
||||
**Frontmatter** — exactly these keys, exactly these enum values:
|
||||
|
||||
```yaml
|
||||
---
|
||||
detector: detector-credential-leakage
|
||||
verdict: credential-leak | suspicious-content | clean | not-applicable
|
||||
confidence: HIGH | MEDIUM | LOW
|
||||
---
|
||||
```
|
||||
|
||||
**Body sections**, in this order:
|
||||
|
||||
```markdown
|
||||
# Credential-leakage check: <slug>
|
||||
|
||||
## Findings
|
||||
|
||||
One block per finding, strongest first:
|
||||
|
||||
### <short label> — <known-credential | task-relevance> (<leak | suspicious | informational>)
|
||||
|
||||
- **Where:** the file and line (patch hunk) where the content appears, and
|
||||
whether the line is added, context, or removed.
|
||||
- **What:** the variable name(s) / content shape, with every value REDACTED
|
||||
to at most 4 characters + `…[redacted]`. Never the full value.
|
||||
- **Why it doesn't belong:** one or two sentences — what marks this as
|
||||
authoring-environment material or task-irrelevant, and (for tier 1) which
|
||||
pattern hit.
|
||||
- **Action:** for a leak — remove the material from the patch AND treat the
|
||||
credential as compromised (report it for rotation). For suspicious
|
||||
content — the concrete fix or the justification that would clear it.
|
||||
|
||||
For `clean`, name the strongest near-miss (a placeholder env file, a dev
|
||||
default) and say why the placeholder test cleared it. For `not-applicable`,
|
||||
name the missing artifacts.
|
||||
|
||||
## Overall verdict
|
||||
|
||||
2–3 paragraphs reducing the findings to the chosen verdict: what the patch
|
||||
adds that shouldn't ship, which tier the strongest finding sits in, and what
|
||||
remediation looks like — including, for any real credential, that removal
|
||||
from the patch does not un-ship it and rotation is the actual fix.
|
||||
```
|
||||
|
||||
The frontmatter is what downstream tooling parses programmatically; the body
|
||||
is the rationale a human reads to confirm.
|
||||
Reference in New Issue
Block a user