chore: pull all docs into sources/
This commit is contained in:
3
running-containers.sh
Executable file
3
running-containers.sh
Executable file
@@ -0,0 +1,3 @@
|
|||||||
|
#!/bin/bash
|
||||||
|
#
|
||||||
|
docker ps --format "table {{.ID}}\t{{.Names}}"
|
||||||
69
sources/ai-version-instructions.md
Normal file
69
sources/ai-version-instructions.md
Normal file
@@ -0,0 +1,69 @@
|
|||||||
|
# Task Creation Overview
|
||||||
|
|
||||||
|
## Purpose
|
||||||
|
Create a task package capturing **meaningful failures** of Claude Code in a real repository. A failure must be:
|
||||||
|
- Recognizable (~80% of senior engineers would flag it)
|
||||||
|
- Real‑world impactful (security, data, permissions, etc.)
|
||||||
|
- Free of artificial or contrived setup
|
||||||
|
|
||||||
|
## End‑to‑End Workflow
|
||||||
|
|
||||||
|
1. **Explore & Identify Failure**
|
||||||
|
- Use the *Explore* container to locate a genuine defect in a repository.
|
||||||
|
- Verify the defect’s significance against the “Meaningful Failure” criteria.
|
||||||
|
|
||||||
|
2. **Build the Task**
|
||||||
|
- **instruction.md** – Write a realistic, self‑contained prompt:
|
||||||
|
* Base it on actual repository state (including any workspace.patch changes).
|
||||||
|
* Avoid hints, AI commentary, or external dependencies.
|
||||||
|
* Do not manufacture breakage; use pre‑existing flaws.
|
||||||
|
- **grader‑guidance‑consolidated.md** – Tailor the grader’s evaluation:
|
||||||
|
* Provide concise task context & optional business context.
|
||||||
|
* Supply privileged ground‑truth facts (file:line references, correct fix).
|
||||||
|
* For each of the 8 grading criteria, describe strong vs. weak responses specific to the task.
|
||||||
|
* Add heavy penalties only for deal‑breakers, targeting a named criterion and stated behavior.
|
||||||
|
|
||||||
|
3. **Generate Reference Runs**
|
||||||
|
- Run trials with 'harbor-run' (or copy from existing job) to produce recorded attempts.
|
||||||
|
- Ensure **≥ 4 accepted reference runs** are stored in 'reference-runs/'.
|
||||||
|
- All runs must use the **same trial agent** (Claude Code or Codex) – do not mix agents.
|
||||||
|
|
||||||
|
4. **Run Detectors**
|
||||||
|
- Execute each detector skill ('/detector‑…') to catch:
|
||||||
|
* Off‑hinting, broken dev‑env, cross‑task references, fact‑check failures, etc.
|
||||||
|
- Fix any issues and re‑run detectors until all reports pass.
|
||||||
|
|
||||||
|
5. **Validate, Export & Submit**
|
||||||
|
- Run 'npx tsx scripts/submit-task.ts <slug>' to validate and package the task.
|
||||||
|
- Export platform state before submitting.
|
||||||
|
- Upload the generated tarball and fill the feedback‑request field (4‑item format).
|
||||||
|
- Include the Slack thread URL for reviewer access.
|
||||||
|
|
||||||
|
## Key Constraints & Policies
|
||||||
|
|
||||||
|
- **Confidentiality**: All task artifacts must stay on your local machine; never post or share publicly.
|
||||||
|
- **Agent Consistency**: Pick Claude Code **or** Codex and stick with it for the entire task.
|
||||||
|
- **Scoring Model**:
|
||||||
|
- Grader uses the **Consolidated Grading Standard** (8 criteria: Integrity, Narrow Correctness, Broader Correctness, Persistence, Communication, Verification & Thoroughness, Common Sense, Thought Partnership).
|
||||||
|
- Scores are averaged (0‑1) and may be reduced by **qualitative heavy penalties** applied to a named criterion.
|
||||||
|
- No numeric caps or combined‑criterion weighting; penalties preserve relative ranking.
|
||||||
|
- **Workspace & Patch**:
|
||||||
|
- Changes are captured in 'environment/workspace.patch'; avoid editing ignored files, Dockerfiles, or external resources.
|
||||||
|
- Verify that the patch applies cleanly on rebuilds ('check-workspace-sync.sh' helps).
|
||||||
|
- **Failure Significance**:
|
||||||
|
- Must meet the “Meaningful Failure” definition (real impact, engineer consensus, blockable PR, etc.).
|
||||||
|
- If reference runs all score high, revisit prompt difficulty or ground‑truth discrimination.
|
||||||
|
|
||||||
|
## Quick Reference Commands
|
||||||
|
|
||||||
|
| Action | Command |
|
||||||
|
|--------|---------|
|
||||||
|
| Start exploration container | 'claude' |
|
||||||
|
| Snapshot current workspace | '/create-snapshot:snapshot' (Claude) |
|
||||||
|
| Build workspace script | 'bash scripts/build-workspace.sh <slug>' |
|
||||||
|
| Verify patch sync | 'bash scripts/check-workspace-sync.sh' |
|
||||||
|
| Run a trial | 'harbor-run' |
|
||||||
|
| Copy reference runs | 'npx tsx scripts/copy-reference-run.ts harbor-jobs/<job>/<slug>__*' |
|
||||||
|
| Rerun a stale reference run | '/regrade-reference-run' |
|
||||||
|
| Submit task | 'npx tsx scripts/submit-task.ts <slug>' |
|
||||||
|
|
||||||
|
Before Width: | Height: | Size: 6.9 KiB After Width: | Height: | Size: 6.9 KiB |
|
Before Width: | Height: | Size: 251 KiB After Width: | Height: | Size: 251 KiB |
|
Before Width: | Height: | Size: 3.0 MiB After Width: | Height: | Size: 3.0 MiB |
1
worker-toolkit-stocks-in-the-future/explore/repo
Symbolic link
1
worker-toolkit-stocks-in-the-future/explore/repo
Symbolic link
@@ -0,0 +1 @@
|
|||||||
|
/home/ericbell/workspaces/dataannotation/current-project/worker-toolkit-stocks-in-the-future/repo
|
||||||
10
worker-toolkit-stocks-in-the-future/explore/toolkit.json
Normal file
10
worker-toolkit-stocks-in-the-future/explore/toolkit.json
Normal file
@@ -0,0 +1,10 @@
|
|||||||
|
{
|
||||||
|
"repo": "stocks-in-the-future",
|
||||||
|
"defaultCommit": "63732df2",
|
||||||
|
"version": "ecee90d5cb",
|
||||||
|
"explorePorts": {
|
||||||
|
"clientHost": 3700,
|
||||||
|
"serverHost": null,
|
||||||
|
"corpusHost": null
|
||||||
|
}
|
||||||
|
}
|
||||||
Reference in New Issue
Block a user