Files
project-work/worker-toolkit-potion-polyglot/.claude/skills/detector-snapshot-leakage/SKILL.md

1.8 KiB

name, description, allowed-tools
name description allowed-tools
detector-snapshot-leakage Self-check a snapshot-based task for whether the snapshot session leaks the rubric's intended answer to the test agent. `/create-snapshot` is meant to capture a failure mode the task tests recovery from — not extra context that hands the test agent a roadmap to the answer the rubric scores. Run this skill on your task before submission to catch leaks while you can still fix them. Bash, Read, Write

Snapshot-leakage detector

This skill checks one of your tasks for snapshot leakage — the most common failure mode for snapshot-based tasks, where the prior conversation in session.jsonl already contains the answer the rubric is testing for, so the test agent gets full credit by repeating something the snapshot handed them.

Read these before deciding:

  1. .claude/skills/_detector-worker-shell.md — where to write the report and how to handle re-runs.
  2. .claude/skills/detector-snapshot-leakage/core.md — what this detector looks for, the three shapes a leak can take, verdict enums, frontmatter/body schema.

Compose the report per the schema in core.md and write it per _detector-worker-shell.md.

Acting on the verdict

  • clear-leak or partial-leak — your snapshot is doing work the rubric expects the agent to do. The fix is usually to trim the snapshot (cut the assistant turns that articulate the answer) and replace them with prior conversation that sets up the failure mode without resolving it. Re-run this skill after editing to confirm the verdict moved to clean.
  • clean — the snapshot stops short of giving the answer. Good.
  • not-applicable — the snapshot or rubric is missing/empty. Either this isn't a snapshot task, or the rubric isn't drafted yet. Come back to this skill once both artifacts exist.