3.7 KiB
3.7 KiB
Task Creation Overview
Purpose
Create a task package capturing meaningful failures of Claude Code in a real repository. A failure must be:
- Recognizable (~80% of senior engineers would flag it)
- Real‑world impactful (security, data, permissions, etc.)
- Free of artificial or contrived setup
End‑to‑End Workflow
-
Explore & Identify Failure
- Use the Explore container to locate a genuine defect in a repository.
- Verify the defect’s significance against the “Meaningful Failure” criteria.
-
Build the Task
- instruction.md – Write a realistic, self‑contained prompt:
- Base it on actual repository state (including any workspace.patch changes).
- Avoid hints, AI commentary, or external dependencies.
- Do not manufacture breakage; use pre‑existing flaws.
- grader‑guidance‑consolidated.md – Tailor the grader’s evaluation:
- Provide concise task context & optional business context.
- Supply privileged ground‑truth facts (file:line references, correct fix).
- For each of the 8 grading criteria, describe strong vs. weak responses specific to the task.
- Add heavy penalties only for deal‑breakers, targeting a named criterion and stated behavior.
- instruction.md – Write a realistic, self‑contained prompt:
-
Generate Reference Runs
- Run trials with 'harbor-run' (or copy from existing job) to produce recorded attempts.
- Ensure ≥ 4 accepted reference runs are stored in 'reference-runs/'.
- All runs must use the same trial agent (Claude Code or Codex) – do not mix agents.
-
Run Detectors
- Execute each detector skill ('/detector‑…') to catch:
- Off‑hinting, broken dev‑env, cross‑task references, fact‑check failures, etc.
- Fix any issues and re‑run detectors until all reports pass.
- Execute each detector skill ('/detector‑…') to catch:
-
Validate, Export & Submit
- Run 'npx tsx scripts/submit-task.ts ' to validate and package the task.
- Export platform state before submitting.
- Upload the generated tarball and fill the feedback‑request field (4‑item format).
- Include the Slack thread URL for reviewer access.
Key Constraints & Policies
- Confidentiality: All task artifacts must stay on your local machine; never post or share publicly.
- Agent Consistency: Pick Claude Code or Codex and stick with it for the entire task.
- Scoring Model:
- Grader uses the Consolidated Grading Standard (8 criteria: Integrity, Narrow Correctness, Broader Correctness, Persistence, Communication, Verification & Thoroughness, Common Sense, Thought Partnership).
- Scores are averaged (0‑1) and may be reduced by qualitative heavy penalties applied to a named criterion.
- No numeric caps or combined‑criterion weighting; penalties preserve relative ranking.
- Workspace & Patch:
- Changes are captured in 'environment/workspace.patch'; avoid editing ignored files, Dockerfiles, or external resources.
- Verify that the patch applies cleanly on rebuilds ('check-workspace-sync.sh' helps).
- Failure Significance:
- Must meet the “Meaningful Failure” definition (real impact, engineer consensus, blockable PR, etc.).
- If reference runs all score high, revisit prompt difficulty or ground‑truth discrimination.
Quick Reference Commands
| Action | Command |
|---|---|
| Start exploration container | 'claude' |
| Snapshot current workspace | '/create-snapshot:snapshot' (Claude) |
| Build workspace script | 'bash scripts/build-workspace.sh ' |
| Verify patch sync | 'bash scripts/check-workspace-sync.sh' |
| Run a trial | 'harbor-run' |
| Copy reference runs | 'npx tsx scripts/copy-reference-run.ts harbor-jobs//__*' |
| Rerun a stale reference run | '/regrade-reference-run' |
| Submit task | 'npx tsx scripts/submit-task.ts ' |