70 lines
3.7 KiB
Markdown
70 lines
3.7 KiB
Markdown
# Task Creation Overview
|
||
|
||
## Purpose
|
||
Create a task package capturing **meaningful failures** of Claude Code in a real repository. A failure must be:
|
||
- Recognizable (~80% of senior engineers would flag it)
|
||
- Real‑world impactful (security, data, permissions, etc.)
|
||
- Free of artificial or contrived setup
|
||
|
||
## End‑to‑End Workflow
|
||
|
||
1. **Explore & Identify Failure**
|
||
- Use the *Explore* container to locate a genuine defect in a repository.
|
||
- Verify the defect’s significance against the “Meaningful Failure” criteria.
|
||
|
||
2. **Build the Task**
|
||
- **instruction.md** – Write a realistic, self‑contained prompt:
|
||
* Base it on actual repository state (including any workspace.patch changes).
|
||
* Avoid hints, AI commentary, or external dependencies.
|
||
* Do not manufacture breakage; use pre‑existing flaws.
|
||
- **grader‑guidance‑consolidated.md** – Tailor the grader’s evaluation:
|
||
* Provide concise task context & optional business context.
|
||
* Supply privileged ground‑truth facts (file:line references, correct fix).
|
||
* For each of the 8 grading criteria, describe strong vs. weak responses specific to the task.
|
||
* Add heavy penalties only for deal‑breakers, targeting a named criterion and stated behavior.
|
||
|
||
3. **Generate Reference Runs**
|
||
- Run trials with 'harbor-run' (or copy from existing job) to produce recorded attempts.
|
||
- Ensure **≥ 4 accepted reference runs** are stored in 'reference-runs/'.
|
||
- All runs must use the **same trial agent** (Claude Code or Codex) – do not mix agents.
|
||
|
||
4. **Run Detectors**
|
||
- Execute each detector skill ('/detector‑…') to catch:
|
||
* Off‑hinting, broken dev‑env, cross‑task references, fact‑check failures, etc.
|
||
- Fix any issues and re‑run detectors until all reports pass.
|
||
|
||
5. **Validate, Export & Submit**
|
||
- Run 'npx tsx scripts/submit-task.ts <slug>' to validate and package the task.
|
||
- Export platform state before submitting.
|
||
- Upload the generated tarball and fill the feedback‑request field (4‑item format).
|
||
- Include the Slack thread URL for reviewer access.
|
||
|
||
## Key Constraints & Policies
|
||
|
||
- **Confidentiality**: All task artifacts must stay on your local machine; never post or share publicly.
|
||
- **Agent Consistency**: Pick Claude Code **or** Codex and stick with it for the entire task.
|
||
- **Scoring Model**:
|
||
- Grader uses the **Consolidated Grading Standard** (8 criteria: Integrity, Narrow Correctness, Broader Correctness, Persistence, Communication, Verification & Thoroughness, Common Sense, Thought Partnership).
|
||
- Scores are averaged (0‑1) and may be reduced by **qualitative heavy penalties** applied to a named criterion.
|
||
- No numeric caps or combined‑criterion weighting; penalties preserve relative ranking.
|
||
- **Workspace & Patch**:
|
||
- Changes are captured in 'environment/workspace.patch'; avoid editing ignored files, Dockerfiles, or external resources.
|
||
- Verify that the patch applies cleanly on rebuilds ('check-workspace-sync.sh' helps).
|
||
- **Failure Significance**:
|
||
- Must meet the “Meaningful Failure” definition (real impact, engineer consensus, blockable PR, etc.).
|
||
- If reference runs all score high, revisit prompt difficulty or ground‑truth discrimination.
|
||
|
||
## Quick Reference Commands
|
||
|
||
| Action | Command |
|
||
|--------|---------|
|
||
| Start exploration container | 'claude' |
|
||
| Snapshot current workspace | '/create-snapshot:snapshot' (Claude) |
|
||
| Build workspace script | 'bash scripts/build-workspace.sh <slug>' |
|
||
| Verify patch sync | 'bash scripts/check-workspace-sync.sh' |
|
||
| Run a trial | 'harbor-run' |
|
||
| Copy reference runs | 'npx tsx scripts/copy-reference-run.ts harbor-jobs/<job>/<slug>__*' |
|
||
| Rerun a stale reference run | '/regrade-reference-run' |
|
||
| Submit task | 'npx tsx scripts/submit-task.ts <slug>' |
|
||
|