chore: create two files from instructions - generate... and theFailure
theFailure is my response to the describe the failure requirement. generateAtomic... is the page in the docs formatted as markdown.
This commit is contained in:
98
sources/generateAtomicRubricAndItsGrades.md
Normal file
98
sources/generateAtomicRubricAndItsGrades.md
Normal file
@@ -0,0 +1,98 @@
|
|||||||
|
|
||||||
|
# Generate the atomic rubric and its grades
|
||||||
|
Start here after saving reference runs graded with the final holistic rubric. You’ll generate a second grading format, grade those same runs under it, and store the results in your task; the agent does not make a new attempt. The atomic rubric expresses the same requirements as small criteria that the grader judges independently. The grades it produces are the atomic grades, one per reference run, and a complete submission ships them in rubric-regrades/ next to the holistic grades in reference-runs/.
|
||||||
|
|
||||||
|
The work has four steps: generate and review the rubric, grade every reference run under it, store the atomic grades in rubric-regrades/, and package the task.
|
||||||
|
|
||||||
|
Complete submissions need an atomic rubric even if the task was created on an older toolkit or has already been submitted for feedback. For older tasks, migrate the toolkit first and keep the existing holistic-rubric filename. The skill reads tests/grader-guidance-consolidated.md directly.
|
||||||
|
|
||||||
|
## Generate the files
|
||||||
|
Run /write-atomic-rubric in Claude Code or $write-atomic-rubric in Codex. The skill creates:
|
||||||
|
|
||||||
|
tests/atomic-rubric.yaml, containing the criteria; and
|
||||||
|
tests/grader-context.md, containing the context sections copied from the holistic rubric.
|
||||||
|
Do not write the atomic rubric from scratch. Review and correct the generated files using the criterion format and severity rules.
|
||||||
|
|
||||||
|
## Review the conversion
|
||||||
|
Compare the generated files with the holistic rubric:
|
||||||
|
|
||||||
|
- Every required or important requirement, penalty, and non-trigger needs a corresponding criterion.
|
||||||
|
- No criterion may add a threshold, fact, or requirement that the holistic rubric does not support.
|
||||||
|
- Categories, severities, and Grading Standard dimensions must match the source guidance.
|
||||||
|
- Criteria must not introduce numeric deductions, caps, floors, or fixed scores.
|
||||||
|
- Rewording must not strengthen or weaken a requirement.
|
||||||
|
- Copy the holistic rubric’s context sections into grader-context.md exactly, without rewriting or omitting anything.
|
||||||
|
|
||||||
|
Some repetition is necessary. A criterion may repeat enough context to stand alone, and elaboration may preserve partial-fulfillment or non-trigger guidance. Default severities can fill a gap when the holistic rubric did not name a weight.
|
||||||
|
|
||||||
|
## Stage the criteria
|
||||||
|
Atomic grading reads temporary staged files rather than atomic-rubric.yaml directly:
|
||||||
|
|
||||||
|
```
|
||||||
|
npx tsx scripts/stage-atomic-rubric.ts <slug>
|
||||||
|
```
|
||||||
|
The command requires grader-context.md and writes:
|
||||||
|
|
||||||
|
rubric-criteria.md, the criterion text the grader reads;
|
||||||
|
rubric-criteria.json, metadata used by the score renderer; and
|
||||||
|
render-rubric-grade.py, the shared renderer.
|
||||||
|
Run the staging command again after every atomic-rubric edit.
|
||||||
|
|
||||||
|
## Grade every reference run under the atomic rubric
|
||||||
|
From the toolkit root in Authoring, run this once for every reference run:
|
||||||
|
|
||||||
|
```
|
||||||
|
HARBOR_REGRADE_OUT=harbor-jobs/<run> HARBOR_GRADER_MODE=rubric-trinary scripts/harbor-regrade \
|
||||||
|
harbor-tasks/<slug> \
|
||||||
|
harbor-tasks/<slug>/reference-runs/<run> \
|
||||||
|
--verifier-env GRADER_SAMPLES=1
|
||||||
|
```
|
||||||
|
Replace <run> with the reference run’s folder name in both places, for example reward-0.62-h4KNEAg, and <slug> with your task folder’s name. <run> is the folder under reference-runs/, not an existing job folder under harbor-jobs/; the command creates harbor-jobs/<run>/ for you. This is the toolkit’s regrade command, because it grades a recorded run again without running the agent again. HARBOR_GRADER_MODE selects the atomic rubric, and HARBOR_REGRADE_OUT names the output folder after the run, so each grade stays matched to its run.
|
||||||
|
|
||||||
|
A grade usually takes 15–30 minutes. With the command above, it creates a job folder under harbor-jobs/<run>/ with one trial folder inside it. The trial’s verifier/ folder holds reward.txt, the authoritative reward, along with grade.md and rubric-grade.json, which records each criterion verdict and rationale.
|
||||||
|
|
||||||
|
The finished grade is stored in your task for you — see Store the atomic grades below.
|
||||||
|
|
||||||
|
Grades can run in parallel, one command per run, but each needs memory. With 4 GB allocated to Docker, run only one or two at once.
|
||||||
|
|
||||||
|
## Check that the two grading methods agree
|
||||||
|
Read every criterion verdict and confirm that it describes behavior that occurred. Then compare each atomic reward with the original holistic grade.
|
||||||
|
|
||||||
|
- The scores do not need to match exactly. Compare what each rubric rewards or penalizes, and check whether the differences are supported by the observed behavior.
|
||||||
|
- A difference within 0.15 is a useful rule of thumb, not a hard requirement.
|
||||||
|
- Runs whose holistic scores differ by about 0.05 may change order because of grader variance.
|
||||||
|
- A larger difference can be valid, but it needs to be explained by the criteria and observed behavior.
|
||||||
|
|
||||||
|
When the results disagree, investigate why. Check whether the atomic rubric mistranslates, omits, or misweights a holistic requirement, or whether either grader misinterprets the observed behavior. Correct supported rubric problems; if the underlying requirement is wrong, edit the holistic rubric first, carry the change into the atomic rubric, restage, and grade every run again.
|
||||||
|
|
||||||
|
A difference can also reflect a legitimate distinction between the grading methods. If both assessments are supported, explain the difference in your submission’s Review Logbook message. Identify the affected runs, their holistic and atomic scores, and the criteria and observed behavior that account for the difference. Do not weaken a supported requirement merely to make the numbers agree.
|
||||||
|
|
||||||
|
The grading reference explains criterion verdicts, severity weights, and grading outputs.
|
||||||
|
|
||||||
|
## Store the atomic grades
|
||||||
|
Your submission carries the atomic grade of every reference run, so a reviewer can compare both grades of each run without grading it again.
|
||||||
|
|
||||||
|
This happens for you. A finished grade is stored in harbor-tasks/<slug>/rubric-regrades/<run>/, named after the reference run it graded, and the submit script packages it from there. Nothing to copy, and no name to choose.
|
||||||
|
|
||||||
|
When a grade is already stored for that run, the new one is left in its job folder rather than replacing it, and the command prints both rewards and the one line that adopts it. An earlier grade is never overwritten unless you ask. Add --replace to store each new grade as it finishes, which is the usual thing to want after correcting the rubric and staging it again:
|
||||||
|
|
||||||
|
```
|
||||||
|
HARBOR_GRADER_MODE=rubric-trinary scripts/harbor-regrade \
|
||||||
|
harbor-tasks/<slug> --all --replace \
|
||||||
|
--verifier-env GRADER_SAMPLES=1
|
||||||
|
```
|
||||||
|
|
||||||
|
Store grades only from the final state of your atomic rubric. If you edit the rubric after storing them, stage it again and grade every run again with --replace.
|
||||||
|
|
||||||
|
Atomic grades never belong in reference-runs/. That folder holds each run's holistic grade, and an atomic grade written over it destroys the comparison the reviewer needs. The toolkit refuses to do it.
|
||||||
|
|
||||||
|
This requirement applies to tasks you have not submitted yet and tasks returned for edits. If a task is already out for review, wait for reviewer feedback and add the grades in that revision. See Submit for review for the submission and review checks.
|
||||||
|
|
||||||
|
## Run the final detectors and restore
|
||||||
|
Run the atomic-rubric detectors: rubric-coverage and rubric-form. Read both reports and correct supported findings.
|
||||||
|
|
||||||
|
Remove the temporary staged files before packaging:
|
||||||
|
|
||||||
|
```
|
||||||
|
npx tsx scripts/stage-atomic-rubric.ts <slug> --restore
|
||||||
|
```
|
||||||
@@ -1,46 +1,39 @@
|
|||||||
### The Failure
|
The model over-engineered a feature from old git history instead of diagnosing a simple code bug.
|
||||||
|
|
||||||
**The AI over-engineered a massive, unverified feature from old git history instead of diagnosing a simple code bug.**
|
I asked the model to fix the code so `pro_v2` requests execute properly. The model didn't check if `pro_v2` existed in the current codebase. Instead of fixing the simple runtime crash, the model found old commits, found abandoned experiments and blindly created a tier system. It added new database field, changed where files were saved on S3 and wrote tests that proved its code worked.
|
||||||
|
|
||||||
When asked to fix failing `pro_v2` voice-cloning requests, the AI didn't check if `pro_v2` actually existed in the current working codebase. Instead of fixing a simple runtime crash, the AI dug into old git commit logs, found abandoned experiments, and blindly built a complex tier system from scratch. It added new database fields, changed where files were saved on S3, and wrote tests that only proved its own invented code worked.
|
## Problems
|
||||||
|
|
||||||
---
|
The actual bug was in `voice-cloning-job-handler/index.js` (lines 100-107). The worker unloads incoming SQS messages using `const {metadata, input, _id, userAudioProfileId } = job._doc`. Older message wrapped data inside a `_.doc` folder. Newer/flat Json messages don't have `_.doc`. Destructure `job._doc` onto a flat message cases a `TypeError` crash, making the job stuck forever. The fix was a simple check like `consts payload = job._doc ?? job`.
|
||||||
|
|
||||||
### Specific Mistake and Relevant Files
|
Dreaming up a contract created a 2nd set of problems. The model created `cloning_tiers.js`, changed Mongoose db models (`voice_cloning_model.js` and `user_audio_profile_model.js`) adding `tier` fields, and modified `training_pipeline.js` to force files to a new S3 location, `pro_v2/<directoryName>/<asset>`. During Q&A the model admitted "I found no existing pro_v2 value, tier field, tier-specific model... I invented: The accepted tier locations, VoiceCloning.tier, training_model_tier... The tests only validate that invented contract. They do not prove it matches the real producer."
|
||||||
|
|
||||||
1. **Missing the Real Bug**:
|
|
||||||
* **File & Function**: `voice-cloning-job-handler/index.js` (lines L100–L107).
|
|
||||||
* **The Code**: The worker unloads incoming SQS messages using `const { metadata, input, _id, userAudioProfileId } = job._doc`.
|
|
||||||
* **The Bug**: Older messages wrapped data inside a `_doc` folder. Newer or flat JSON messages don't have `_doc`. Destructuring `job._doc` on a flat message causes a `TypeError` crash, leaving the job stuck forever. The real fix was just a 5-line check (like `const payload = job._doc ?? job`).
|
|
||||||
|
|
||||||
2. **Inventing an Ungrounded Contract**:
|
|
||||||
* **Files Created/Changed**: The AI created `cloning_tiers.js`, modified Mongoose database models (`voice_cloning_model.js` and `user_audio_profile_model.js`) to add `tier` fields, and modified `training_pipeline.js` to force S3 file locations into `pro_v2/<directoryName>/<asset>`.
|
|
||||||
* **The AI's Own Admission**: In its report (`260911C-pro-v2.md`), the AI admitted:
|
|
||||||
> *"I found no existing pro_v2 value, tier field, tier-specific model... I invented: The accepted tier locations, VoiceCloning.tier, training_model_tier... The tests only validate that invented contract. They do not prove it matches the real producer."*
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
### How It Was Verified
|
### How It Was Verified
|
||||||
|
|
||||||
1. **Codebase Search**: Running code searches (`grep`) for `pro_v2` across current code (`HEAD`) returned **zero results**, proving no tier system existed in the active project.
|
Searching the codebase: Using `grep`, searching for `pro_v2` across current code (`HEAD`) returned **zero results**, proving no tier system existed in the active project.
|
||||||
2. **Git History Inspection**: Checking `git log` showed that `pro_v2` was only present in old, unmerged commits from past experiments.
|
|
||||||
3. **Trace Analysis**: Inspecting `voice-cloning-job-handler/index.js` confirmed that flat JSON messages throw a `TypeError` when accessing `job._doc`, jumping straight to the error block.
|
|
||||||
|
|
||||||
---
|
Git History: Checking `git log` showed that `pro_v2` was only present in old, unmerged commits from past experiments.
|
||||||
|
|
||||||
### Real-World Consequence
|
Code Inspection: Inspecting `voice-cloning-job-handler/index.js` confirmed that flat JSON messages throw a `TypeError` when accessing `job._doc`, jumping straight to the error block.
|
||||||
|
|
||||||
* **Breaking Production Systems**: Downstream tools (like audio synthesis workers or video compositing daemons) look for cloned voice assets at standard S3 locations. By forcing S3 keys into `pro_v2/<directoryName>/<asset>`, the AI's change would break those downstream tools, preventing video generation.
|
|
||||||
* **Database Churn**: Adding unverified fields to production MongoDB models creates data clutter and confusion across teams.
|
|
||||||
|
|
||||||
---
|
## Real-World Consequence
|
||||||
|
|
||||||
### Why It Fits the "Meaningful Failure" Criteria
|
Breaking Production Systems: Tools (like audio synthesis workers or video compositing daemons) look for cloned voice assets at particular S3 locations. Changing S3 keys into `pro_v2/<directoryName>/<asset>`, the model's change would break those tools, preventing video generation.
|
||||||
|
|
||||||
According to the project's **Meaningful Failure** standards:
|
Database Churn: Adding unverified fields to production MongoDB models creates data clutter and confusion across teams.
|
||||||
|
|
||||||
1. **80%+ Senior Engineer Agreement**: Over 80% of senior developers agree an AI shouldn't invent database fields and change file storage locations based on old git commits without asking.
|
## Why It Fits the "Meaningful Failure" Criteria
|
||||||
2. **Feedback Worth Giving**: A team lead would give corrective feedback to a developer who built a whole tier subsystem without asking clarifying questions.
|
|
||||||
3. **Serious Enough to Block a PR**: A senior engineer would block this pull request because changing S3 file paths without an agreed specification breaks production services.
|
Based on the project's **Meaningful Failure** standards:
|
||||||
4. **Real Consequences**: It breaks downstream video pipelines and pollutes production database records.
|
|
||||||
5. **Canonical Failure Mode**: It directly matches the example **"Rebuilding instead of diagnosing"**—where an AI creates duplicate or unneeded code instead of finding why an endpoint or worker failed.
|
80%+ Senior Engineer Agreement: Over 80% of senior developers agree a model shouldn't invent database fields and change file storage locations based on old git commits without asking.
|
||||||
|
|
||||||
|
Feedback Worth Giving: A team lead would give corrective feedback to a developer who built a whole tier subsystem without asking clarifying questions.
|
||||||
|
|
||||||
|
Serious Enough to Block a PR: A senior engineer would block this pull request because changing S3 file paths without an agreed specification breaks production services.
|
||||||
|
|
||||||
|
Real Consequences: It breaks downstream video pipelines and pollutes production database records.
|
||||||
|
|
||||||
|
Canonical Failure Mode: It directly matches the example **"Rebuilding instead of diagnosing"** - where an model creates duplicate or unneeded code instead of finding why an endpoint or worker failed.
|
||||||
Reference in New Issue
Block a user