# Task Creation Overview ## Purpose Create a task package capturing **meaningful failures** of Claude Code in a real repository. A failure must be: - Recognizable (~80% of senior engineers would flag it) - Real‑world impactful (security, data, permissions, etc.) - Free of artificial or contrived setup ## End‑to‑End Workflow 1. **Explore & Identify Failure** - Use the *Explore* container to locate a genuine defect in a repository. - Verify the defect’s significance against the “Meaningful Failure” criteria. 2. **Build the Task** - **instruction.md** – Write a realistic, self‑contained prompt: * Base it on actual repository state (including any workspace.patch changes). * Avoid hints, AI commentary, or external dependencies. * Do not manufacture breakage; use pre‑existing flaws. - **grader‑guidance‑consolidated.md** – Tailor the grader’s evaluation: * Provide concise task context & optional business context. * Supply privileged ground‑truth facts (file:line references, correct fix). * For each of the 8 grading criteria, describe strong vs. weak responses specific to the task. * Add heavy penalties only for deal‑breakers, targeting a named criterion and stated behavior. 3. **Generate Reference Runs** - Run trials with 'harbor-run' (or copy from existing job) to produce recorded attempts. - Ensure **≥ 4 accepted reference runs** are stored in 'reference-runs/'. - All runs must use the **same trial agent** (Claude Code or Codex) – do not mix agents. 4. **Run Detectors** - Execute each detector skill ('/detector‑…') to catch: * Off‑hinting, broken dev‑env, cross‑task references, fact‑check failures, etc. - Fix any issues and re‑run detectors until all reports pass. 5. **Validate, Export & Submit** - Run 'npx tsx scripts/submit-task.ts ' to validate and package the task. - Export platform state before submitting. - Upload the generated tarball and fill the feedback‑request field (4‑item format). - Include the Slack thread URL for reviewer access. ## Key Constraints & Policies - **Confidentiality**: All task artifacts must stay on your local machine; never post or share publicly. - **Agent Consistency**: Pick Claude Code **or** Codex and stick with it for the entire task. - **Scoring Model**: - Grader uses the **Consolidated Grading Standard** (8 criteria: Integrity, Narrow Correctness, Broader Correctness, Persistence, Communication, Verification & Thoroughness, Common Sense, Thought Partnership). - Scores are averaged (0‑1) and may be reduced by **qualitative heavy penalties** applied to a named criterion. - No numeric caps or combined‑criterion weighting; penalties preserve relative ranking. - **Workspace & Patch**: - Changes are captured in 'environment/workspace.patch'; avoid editing ignored files, Dockerfiles, or external resources. - Verify that the patch applies cleanly on rebuilds ('check-workspace-sync.sh' helps). - **Failure Significance**: - Must meet the “Meaningful Failure” definition (real impact, engineer consensus, blockable PR, etc.). - If reference runs all score high, revisit prompt difficulty or ground‑truth discrimination. ## Quick Reference Commands | Action | Command | |--------|---------| | Start exploration container | 'claude' | | Snapshot current workspace | '/create-snapshot:snapshot' (Claude) | | Build workspace script | 'bash scripts/build-workspace.sh ' | | Verify patch sync | 'bash scripts/check-workspace-sync.sh' | | Run a trial | 'harbor-run' | | Copy reference runs | 'npx tsx scripts/copy-reference-run.ts harbor-jobs//__*' | | Rerun a stale reference run | '/regrade-reference-run' | | Submit task | 'npx tsx scripts/submit-task.ts ' |