- Undefined Async Engineering Standard removed from both sites; the foreclosure now rests on a Ground Truth fact instead of a named authority
- Duplicated grounding clause repaired
- c01 unreachable premise reframed across all four sites — :4, :29, :30, :53
- c07 narrowed to tier-specific model checkpoints
- c11 S3 claim softened to what the repo actually supports
Every load-bearing claim the fact-check flagged is now either true against fcd8a9d or no longer gating credit.
Still open on the rubric — one edit, two decisions:
1. Major Penalty (:68) — name the target score, and resolve whether "without flagging X or confirming Y" means neither or both.
2. Where envelope-probing sits — now that the rubric concedes the producer shape is unknowable, is a normalizer handling two or three wrapper shapes reasonable
defensiveness or still over-engineering? The tier fields and S3 namespaces are clearly still penalized; this is about the middle ground. Worth resolving in
the same edit so :68 and :4 don't contradict each other.
3. dimension-misapplication — conditioning the Integrity clause on evidence the agent actually saw, and decoupling the TP 1.0 tier from successfully
implementing the fix.
Once those land, the compute sequence is: re-run the three rubric-only detectors → build-workspace.sh (to settle the missing test-commands.sh) → regrade the
four runs once → copy reference runs → run-dependent detectors → /write-atomic-rubric.
task-instructions - current instructions as of 8/14 - masked
task-instructions-8-11-before-masking - current instructions as generated
original-task-instructions - the instructions before today