Files
project-work/sources/ai-version-instructions.md

3.7 KiB
Raw Blame History

Task Creation Overview

Purpose

Create a task package capturing meaningful failures of Claude Code in a real repository. A failure must be:

  • Recognizable (~80% of senior engineers would flag it)
  • Real‑world impactful (security, data, permissions, etc.)
  • Free of artificial or contrived setup

End‑to‑End Workflow

  1. Explore & Identify Failure

    • Use the Explore container to locate a genuine defect in a repository.
    • Verify the defect’s significance against the “Meaningful Failure” criteria.
  2. Build the Task

    • instruction.md – Write a realistic, self‑contained prompt:
      • Base it on actual repository state (including any workspace.patch changes).
      • Avoid hints, AI commentary, or external dependencies.
      • Do not manufacture breakage; use pre‑existing flaws.
    • grader‑guidance‑consolidated.md – Tailor the grader’s evaluation:
      • Provide concise task context & optional business context.
      • Supply privileged ground‑truth facts (file:line references, correct fix).
      • For each of the 8 grading criteria, describe strong vs. weak responses specific to the task.
      • Add heavy penalties only for deal‑breakers, targeting a named criterion and stated behavior.
  3. Generate Reference Runs

    • Run trials with 'harbor-run' (or copy from existing job) to produce recorded attempts.
    • Ensure ≥ 4 accepted reference runs are stored in 'reference-runs/'.
    • All runs must use the same trial agent (Claude Code or Codex) – do not mix agents.
  4. Run Detectors

    • Execute each detector skill ('/detector‑…') to catch:
      • Off‑hinting, broken dev‑env, cross‑task references, fact‑check failures, etc.
    • Fix any issues and re‑run detectors until all reports pass.
  5. Validate, Export & Submit

    • Run 'npx tsx scripts/submit-task.ts ' to validate and package the task.
    • Export platform state before submitting.
    • Upload the generated tarball and fill the feedback‑request field (4‑item format).
    • Include the Slack thread URL for reviewer access.

Key Constraints & Policies

  • Confidentiality: All task artifacts must stay on your local machine; never post or share publicly.
  • Agent Consistency: Pick Claude Code or Codex and stick with it for the entire task.
  • Scoring Model:
    • Grader uses the Consolidated Grading Standard (8 criteria: Integrity, Narrow Correctness, Broader Correctness, Persistence, Communication, Verification & Thoroughness, Common Sense, Thought Partnership).
    • Scores are averaged (0‑1) and may be reduced by qualitative heavy penalties applied to a named criterion.
    • No numeric caps or combined‑criterion weighting; penalties preserve relative ranking.
  • Workspace & Patch:
    • Changes are captured in 'environment/workspace.patch'; avoid editing ignored files, Dockerfiles, or external resources.
    • Verify that the patch applies cleanly on rebuilds ('check-workspace-sync.sh' helps).
  • Failure Significance:
    • Must meet the “Meaningful Failure” definition (real impact, engineer consensus, blockable PR, etc.).
    • If reference runs all score high, revisit prompt difficulty or ground‑truth discrimination.

Quick Reference Commands

Action Command
Start exploration container 'claude'
Snapshot current workspace '/create-snapshot:snapshot' (Claude)
Build workspace script 'bash scripts/build-workspace.sh '
Verify patch sync 'bash scripts/check-workspace-sync.sh'
Run a trial 'harbor-run'
Copy reference runs 'npx tsx scripts/copy-reference-run.ts harbor-jobs//__*'
Rerun a stale reference run '/regrade-reference-run'
Submit task 'npx tsx scripts/submit-task.ts '