Files
project-work/sources/ai-version-instructions.md

70 lines
3.7 KiB
Markdown
Raw Blame History

This file contains invisible Unicode characters

This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Task Creation Overview
## Purpose
Create a task package capturing **meaningful failures** of Claude Code in a real repository. A failure must be:
- Recognizable (~80% of senior engineers would flag it)
- Real‑world impactful (security, data, permissions, etc.)
- Free of artificial or contrived setup
## End‑to‑End Workflow
1. **Explore & Identify Failure**
- Use the *Explore* container to locate a genuine defect in a repository.
- Verify the defect’s significance against the “Meaningful Failure” criteria.
2. **Build the Task**
- **instruction.md** – Write a realistic, self‑contained prompt:
* Base it on actual repository state (including any workspace.patch changes).
* Avoid hints, AI commentary, or external dependencies.
* Do not manufacture breakage; use pre‑existing flaws.
- **grader‑guidance‑consolidated.md** – Tailor the grader’s evaluation:
* Provide concise task context & optional business context.
* Supply privileged ground‑truth facts (file:line references, correct fix).
* For each of the 8 grading criteria, describe strong vs. weak responses specific to the task.
* Add heavy penalties only for deal‑breakers, targeting a named criterion and stated behavior.
3. **Generate Reference Runs**
- Run trials with 'harbor-run' (or copy from existing job) to produce recorded attempts.
- Ensure **≥ 4 accepted reference runs** are stored in 'reference-runs/'.
- All runs must use the **same trial agent** (Claude Code or Codex) – do not mix agents.
4. **Run Detectors**
- Execute each detector skill ('/detector‑…') to catch:
* Off‑hinting, broken dev‑env, cross‑task references, fact‑check failures, etc.
- Fix any issues and re‑run detectors until all reports pass.
5. **Validate, Export & Submit**
- Run 'npx tsx scripts/submit-task.ts <slug>' to validate and package the task.
- Export platform state before submitting.
- Upload the generated tarball and fill the feedback‑request field (4‑item format).
- Include the Slack thread URL for reviewer access.
## Key Constraints & Policies
- **Confidentiality**: All task artifacts must stay on your local machine; never post or share publicly.
- **Agent Consistency**: Pick Claude Code **or** Codex and stick with it for the entire task.
- **Scoring Model**:
- Grader uses the **Consolidated Grading Standard** (8 criteria: Integrity, Narrow Correctness, Broader Correctness, Persistence, Communication, Verification & Thoroughness, Common Sense, Thought Partnership).
- Scores are averaged (0‑1) and may be reduced by **qualitative heavy penalties** applied to a named criterion.
- No numeric caps or combined‑criterion weighting; penalties preserve relative ranking.
- **Workspace & Patch**:
- Changes are captured in 'environment/workspace.patch'; avoid editing ignored files, Dockerfiles, or external resources.
- Verify that the patch applies cleanly on rebuilds ('check-workspace-sync.sh' helps).
- **Failure Significance**:
- Must meet the “Meaningful Failure” definition (real impact, engineer consensus, blockable PR, etc.).
- If reference runs all score high, revisit prompt difficulty or ground‑truth discrimination.
## Quick Reference Commands
| Action | Command |
|--------|---------|
| Start exploration container | 'claude' |
| Snapshot current workspace | '/create-snapshot:snapshot' (Claude) |
| Build workspace script | 'bash scripts/build-workspace.sh <slug>' |
| Verify patch sync | 'bash scripts/check-workspace-sync.sh' |
| Run a trial | 'harbor-run' |
| Copy reference runs | 'npx tsx scripts/copy-reference-run.ts harbor-jobs/<job>/<slug>__*' |
| Rerun a stale reference run | '/regrade-reference-run' |
| Submit task | 'npx tsx scripts/submit-task.ts <slug>' |