chore: pull all docs into sources/

This commit is contained in:
2026-08-19 21:09:38 -04:00
parent 9abade1a81
commit c481a3cf5d
17 changed files with 83 additions and 0 deletions

View File

@@ -0,0 +1,69 @@
# Task Creation Overview
## Purpose
Create a task package capturing **meaningful failures** of Claude Code in a real repository. A failure must be:
- Recognizable (~80% of senior engineers would flag it)
- Real‑world impactful (security, data, permissions, etc.)
- Free of artificial or contrived setup
## End‑to‑End Workflow
1. **Explore & Identify Failure**
- Use the *Explore* container to locate a genuine defect in a repository.
- Verify the defect’s significance against the “Meaningful Failure” criteria.
2. **Build the Task**
- **instruction.md** – Write a realistic, self‑contained prompt:
* Base it on actual repository state (including any workspace.patch changes).
* Avoid hints, AI commentary, or external dependencies.
* Do not manufacture breakage; use pre‑existing flaws.
- **grader‑guidance‑consolidated.md** – Tailor the grader’s evaluation:
* Provide concise task context & optional business context.
* Supply privileged ground‑truth facts (file:line references, correct fix).
* For each of the 8 grading criteria, describe strong vs. weak responses specific to the task.
* Add heavy penalties only for deal‑breakers, targeting a named criterion and stated behavior.
3. **Generate Reference Runs**
- Run trials with 'harbor-run' (or copy from existing job) to produce recorded attempts.
- Ensure **≥ 4 accepted reference runs** are stored in 'reference-runs/'.
- All runs must use the **same trial agent** (Claude Code or Codex) – do not mix agents.
4. **Run Detectors**
- Execute each detector skill ('/detector‑…') to catch:
* Off‑hinting, broken dev‑env, cross‑task references, fact‑check failures, etc.
- Fix any issues and re‑run detectors until all reports pass.
5. **Validate, Export & Submit**
- Run 'npx tsx scripts/submit-task.ts <slug>' to validate and package the task.
- Export platform state before submitting.
- Upload the generated tarball and fill the feedback‑request field (4‑item format).
- Include the Slack thread URL for reviewer access.
## Key Constraints & Policies
- **Confidentiality**: All task artifacts must stay on your local machine; never post or share publicly.
- **Agent Consistency**: Pick Claude Code **or** Codex and stick with it for the entire task.
- **Scoring Model**:
- Grader uses the **Consolidated Grading Standard** (8 criteria: Integrity, Narrow Correctness, Broader Correctness, Persistence, Communication, Verification & Thoroughness, Common Sense, Thought Partnership).
- Scores are averaged (0‑1) and may be reduced by **qualitative heavy penalties** applied to a named criterion.
- No numeric caps or combined‑criterion weighting; penalties preserve relative ranking.
- **Workspace & Patch**:
- Changes are captured in 'environment/workspace.patch'; avoid editing ignored files, Dockerfiles, or external resources.
- Verify that the patch applies cleanly on rebuilds ('check-workspace-sync.sh' helps).
- **Failure Significance**:
- Must meet the “Meaningful Failure” definition (real impact, engineer consensus, blockable PR, etc.).
- If reference runs all score high, revisit prompt difficulty or ground‑truth discrimination.
## Quick Reference Commands
| Action | Command |
|--------|---------|
| Start exploration container | 'claude' |
| Snapshot current workspace | '/create-snapshot:snapshot' (Claude) |
| Build workspace script | 'bash scripts/build-workspace.sh <slug>' |
| Verify patch sync | 'bash scripts/check-workspace-sync.sh' |
| Run a trial | 'harbor-run' |
| Copy reference runs | 'npx tsx scripts/copy-reference-run.ts harbor-jobs/<job>/<slug>__*' |
| Rerun a stale reference run | '/regrade-reference-run' |
| Submit task | 'npx tsx scripts/submit-task.ts <slug>' |