after moving all to cipher
This commit is contained in:
@@ -1,79 +1,106 @@
|
||||
---
|
||||
name: detector-credential-leakage
|
||||
description: |
|
||||
Self-check whether your submission ships credentials or other content from
|
||||
your authoring environment inside its authored surfaces — above all
|
||||
`environment/workspace.patch`. Two tiers. (1) **Known-credential tier
|
||||
(deterministic):** hard-flags your authoring environment's own env vars
|
||||
(`ANTHROPIC_API_KEY`, `ANTHROPIC_BASE_URL`, `USER_ID` as an env
|
||||
assignment) and well-known secret shapes (`sk-ant-…`, AWS `AKIA…`, GitHub
|
||||
`ghp_…`, Google `AIza…`, Stripe secret keys, bearer tokens, private-key
|
||||
blocks, URL-embedded passwords) on lines your patch adds. The canonical
|
||||
incident: your toolkit `.env` — your personal API key, proxy URL, and user
|
||||
id — swept into the workspace as a new `.env` file. (2) **Task-relevance
|
||||
tier (judgment):** content your patch adds that doesn't appear to serve
|
||||
the task — `.env`-style files, env-file symlinks into your home directory,
|
||||
`export FOO=` lines, credential-shaped assignments with real values. A
|
||||
`credential-leak` finding must be acted on before submitting (remove the
|
||||
material AND report the key as compromised so it can be rotated);
|
||||
`suspicious-content` findings are advisory. The report never reproduces
|
||||
secret values. Reads workspace.patch (+ Dockerfile, instruction.md,
|
||||
tests/*.md); runs before or after reference runs exist.
|
||||
Self-check whether your submission ships credentials, internal
|
||||
information, or other content from your authoring environment inside its
|
||||
authored surfaces — above all `environment/workspace.patch`. Three tiers.
|
||||
(1) **Known-credential tier (deterministic):** hard-flags your authoring
|
||||
environment's own env vars (`ANTHROPIC_API_KEY`, `ANTHROPIC_BASE_URL`,
|
||||
`USER_ID` as an env assignment) and well-known secret shapes (`sk-ant-…`,
|
||||
AWS `AKIA…`, GitHub `ghp_…`, Google `AIza…`, Stripe secret keys, bearer
|
||||
tokens, private-key blocks, URL-embedded passwords) on lines your patch
|
||||
adds. (2) **Internal-leakage tier (deterministic candidates):** the
|
||||
project name or being-evaluated framing in workspace content (a CLAUDE.md
|
||||
that tells the agent "this is an assessment"), your identity (home-dir
|
||||
paths, agency names, toolkit checkout paths — HTML-escaped copies
|
||||
included), and authoring artifacts (`.raccoon-setup-done`,
|
||||
`.claude/settings.local.json`, session-export dumps, stray logs).
|
||||
(3) **Task-relevance tier (judgment):** patch content that doesn't serve
|
||||
the task — CLAUDE.md/.claude additions judged on their content against
|
||||
your instruction.md + grader guidance (a task-relevant CLAUDE.md is
|
||||
fine), env files, unexplained config. `credential-leak` and
|
||||
`internal-leak` findings must be fixed before submitting;
|
||||
`suspicious-content` is advisory. The report never reproduces secret
|
||||
values. Reads workspace.patch (+ Dockerfile, instruction.md, tests/*.md);
|
||||
runs before or after reference runs exist.
|
||||
allowed-tools: Bash, Read, Write
|
||||
---
|
||||
|
||||
# Credential-leakage detector
|
||||
|
||||
This skill checks one of your tasks for **credential leakage** — whether
|
||||
anything from your own authoring environment (or any other secret) has been
|
||||
swept into the submission's authored surfaces, above all
|
||||
`environment/workspace.patch`. Everything your patch adds ships to everyone
|
||||
downstream, so a leaked key is compromised the moment you submit: deleting the
|
||||
line later does not un-ship it.
|
||||
This skill checks one of your tasks for **leakage from your authoring
|
||||
environment** — credentials, internal information, your own identity, or
|
||||
plain authoring-machine cruft swept into the submission's authored surfaces,
|
||||
above all `environment/workspace.patch`. Everything your patch adds ships to
|
||||
everyone downstream, and the test agent reads the built workspace: a leaked
|
||||
key is compromised the moment you submit, and a workspace file that names
|
||||
the project or says the agent is being assessed invalidates the task itself.
|
||||
|
||||
The failure shape to catch: your toolkit's `.env` — the file holding your
|
||||
personal `ANTHROPIC_API_KEY`, `ANTHROPIC_BASE_URL`, and `USER_ID` — landing in
|
||||
the workspace as a new `.env` file (or a `.env.bak-*` backup, or a symlink to
|
||||
`/home/<you>/.env`). It happens easily: a stray `git add`, a working-tree
|
||||
backup, a captured terminal snippet. None of it serves the task; the test
|
||||
agent has no network to use a key with; and the key is now distributed.
|
||||
The failure shapes to catch:
|
||||
|
||||
- **Credentials.** Your toolkit `.env` — your personal `ANTHROPIC_API_KEY`,
|
||||
`ANTHROPIC_BASE_URL`, and `USER_ID` — landing in the workspace as a new
|
||||
`.env` file (or a `.env.bak-*` backup, or a symlink to `/home/<you>/.env`).
|
||||
- **Internal information / being-evaluated framing.** A `CLAUDE.md` (or any
|
||||
workspace file) that names the project ("raccoon"), or tells the agent
|
||||
what this really is ("this is a behavioral assessment", "we capture
|
||||
failures for grading"). The agent under test must experience a plausible
|
||||
real-world scenario, not a labeled exam.
|
||||
- **Your identity.** Home-directory paths (`/home/<you>/…`,
|
||||
`/Users/<you>/…`), your agency or employer's name in those paths, toolkit
|
||||
checkout paths (`worker-toolkit-…`) — including HTML-escaped copies inside
|
||||
exported artifacts. A real incident: an HTML export of an authoring
|
||||
session added to the workspace carried the author's name and agency in
|
||||
~13 escaped `file_path` fields.
|
||||
- **Authoring artifacts.** `.raccoon-setup-done`,
|
||||
`.claude/settings.local.json` (your machine-local permission state),
|
||||
stray build logs, session dumps. Even content-harmless, they are unclean
|
||||
patch content nothing in the task explains.
|
||||
|
||||
What *doesn't* trip this check: placeholder and example values
|
||||
(`.env.example` with empty or dummy entries, `sk-ant-...` as a literal
|
||||
template, `changeme`), dev-infrastructure defaults (`POSTGRES_PASSWORD=postgres`
|
||||
in a local docker-compose), code identifiers (`USER_ID = 4958` as a test
|
||||
constant), and env vars your task's scenario genuinely needs documented.
|
||||
(`.env.example` with dummies, `sk-ant-...` as a literal template), dev
|
||||
defaults (`POSTGRES_PASSWORD=postgres` in a local docker-compose), code
|
||||
identifiers (`USER_ID = 4958` as a test constant), generic service-account
|
||||
paths (`/home/app/`), and — importantly — **a task-relevant `CLAUDE.md`**:
|
||||
one that documents codebase conventions the graded behavior depends on, or
|
||||
sets deliberate in-world constraints, is task authoring, judged on its
|
||||
content, never flagged for existing.
|
||||
|
||||
**Severity differs by tier.** Unlike most self-checks, a `credential-leak`
|
||||
finding is not a consideration: remove the material from the patch, rebuild it
|
||||
(`bash scripts/check-workspace-sync.sh --update-patch harbor-tasks/<slug>`),
|
||||
and report the leaked credential through your support channel so it can be
|
||||
rotated — treat it as compromised even after you scrub it.
|
||||
`suspicious-content` findings are the usual advisory kind: read each one and
|
||||
fix or justify it.
|
||||
**Severity differs by tier.** `credential-leak` and `internal-leak` findings
|
||||
are not considerations — fix them before submitting. Remove the material,
|
||||
rebuild the patch (`bash scripts/check-workspace-sync.sh --update-patch
|
||||
harbor-tasks/<slug>`), and for a real credential also report it through your
|
||||
support channel so it can be rotated — it's compromised even after you scrub
|
||||
it. `suspicious-content` findings are the usual advisory kind: read each one
|
||||
and fix or justify it.
|
||||
|
||||
Read these before deciding:
|
||||
|
||||
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
|
||||
2. `.claude/skills/detector-credential-leakage/core.md` — the two tiers, the deterministic pattern checks to run, the placeholder test, the redaction rule (never quote a secret value), what is NOT a finding, verdict enums, and the body schema.
|
||||
2. `.claude/skills/detector-credential-leakage/core.md` — the three tiers, the deterministic pattern checks to run (including entity-decoding), the placeholder test, the confirmation rules, the redaction rule (never quote a secret value), what is NOT a finding, verdict enums, and the body schema.
|
||||
|
||||
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
|
||||
|
||||
## Acting on the verdict
|
||||
|
||||
- **`clean`** — nothing your patch adds looks like a credential or foreign
|
||||
content. Good. Move on.
|
||||
- **`suspicious-content`** — no confirmed credential, but something your patch
|
||||
adds doesn't look like it belongs to the task: an env-file symlink into your
|
||||
home directory, a captured request with a real (if low-sensitivity) token, a
|
||||
config file of credential-shaped values. Fix each finding (replace tokens
|
||||
with placeholders, drop the file, or make its task relevance explicit) or
|
||||
- **`clean`** — nothing your patch adds looks like a credential, internal
|
||||
information, or foreign content. Good. Move on.
|
||||
- **`suspicious-content`** — no confirmed leak, but something your patch adds
|
||||
couldn't be tied to the task: a captured request with a real (if
|
||||
low-sensitivity) token, a config file of credential-shaped values, an
|
||||
addition whose purpose isn't clear. Fix each finding (replace tokens with
|
||||
placeholders, drop the file, or make its task relevance explicit) or
|
||||
satisfy yourself it's genuinely scenario material.
|
||||
- **`credential-leak`** — a real credential or your authoring environment's
|
||||
own env vars are in the patch. Act before submitting: (1) remove the
|
||||
material and regenerate `workspace.patch`; (2) re-run this detector to
|
||||
confirm it's gone; (3) report the leaked value as compromised so it can be
|
||||
rotated — scrubbing the patch does not un-ship a key that already left your
|
||||
machine in an earlier submission.
|
||||
- **`internal-leak`** — your patch carries internal information, your
|
||||
identity, or authoring-machine artifacts. Fix before submitting: delete
|
||||
the file or passage (`.raccoon-setup-done`, `settings.local.json`, the
|
||||
session export, the project-naming paragraph), regenerate
|
||||
`workspace.patch`, and re-run this detector to confirm it's gone.
|
||||
- **`credential-leak`** — a real credential (or your authoring env vars) is
|
||||
in the patch. Act before submitting: (1) remove the material and
|
||||
regenerate `workspace.patch`; (2) re-run this detector to confirm;
|
||||
(3) report the leaked value as compromised so it can be rotated —
|
||||
scrubbing the patch does not un-ship a key that already left your machine
|
||||
in an earlier submission.
|
||||
- **`not-applicable`** — there's no workspace patch to assess yet. Build the
|
||||
workspace first.
|
||||
|
||||
Reference in New Issue
Block a user