diff --git a/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/detectors/detector-rubric-form.inputs.json b/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/detectors/detector-rubric-form.inputs.json index 4e3a018..a85921c 100644 --- a/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/detectors/detector-rubric-form.inputs.json +++ b/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/detectors/detector-rubric-form.inputs.json @@ -1,6 +1,6 @@ { "version": 1, - "capturedAt": "2026-09-27T09:51:49.533Z", + "capturedAt": "2026-09-28T18:22:03.265Z", "capturedBy": "stamp", "inputs": { "prompt": "29e2eb28448679a65ae264372ddf7d993e5752d0f295bf5557cc5b1578265a29", @@ -9,9 +9,9 @@ "workspacePatch": null, "gitref": "fcd8a9d", "graderGuidanceConsolidated": null, - "holisticRubric": "675edd50a8150fd65273036f2deb251190d7cbe0344c7ee24bebda1a8d5d4b6a", - "atomicRubric": "1b8f60cd46e1aa84062b2ef8ed3931cd0b3d6678f4b100e7d62f2e32516521cf", + "holisticRubric": "d1842656433024e283f2732909cbc9a58f924adefe80f5c4ebef7544551f7522", + "atomicRubric": "87860b8970904ca5f7a1474b59c7d175da888a007f59926e9f7dc2fc389e3eb1", "rubricsYaml": null, - "graderContext": "666a029e8834f546a2a9a2ebff5555c0090fbd87da88e0523e773d5c5d0a110b" + "graderContext": "c0924eff78105ed8e16a8f8fcf81594e51146f39d7a5788b462768b1d0ce6653" } } diff --git a/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/detectors/detector-rubric-form.md b/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/detectors/detector-rubric-form.md index 5a2f50e..da429ec 100644 --- a/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/detectors/detector-rubric-form.md +++ b/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/detectors/detector-rubric-form.md @@ -10,9 +10,9 @@ Assessed: harbor-tasks/mishandled_pro_v2/tests/atomic-rubric.yaml ## Deterministic contract -1. **PASS — Parses as YAML.** The document loads with a top-level `task` string and a `criteria` list containing 16 criteria. +1. **PASS — Parses as YAML.** The document loads with a top-level `task` string and a `criteria` list containing 17 criteria. 2. **PASS — Task name.** `task` is exactly `mishandled_pro_v2`. -3. **PASS — Criterion ids.** All 16 ids are unique and match the kebab-case pattern. +3. **PASS — Criterion ids.** All 17 ids are unique and match the kebab-case pattern. 4. **PASS — Category vocabulary.** Every category is `primary_intent` or `dodged_bullet`, both allowed values. 5. **PASS — Severity vocabulary and placement.** Every criterion carries one allowed severity; there are no `extra_credit` criteria with a forbidden severity. 6. **PASS — Crux cap.** One criterion carries `severity: crux`, within the cap of two. @@ -20,11 +20,11 @@ Assessed: harbor-tasks/mishandled_pro_v2/tests/atomic-rubric.yaml 8. **PASS — Guidelines.** Every criterion has a non-empty guideline. 9. **PASS — Numeric penalty language.** All five required pattern sweeps ran across the rubric and context document and returned no candidates. 10. **N/A — Reserved numbering.** The canonical detector contract defines no separate item 10. -11. **PASS — Positive phrasing.** Every guideline uses `The response should ...`, a sanctioned conditional, or `should avoid ...`. The negation sweep matched `does not establish` only inside the factual answer key of `surfaces-producer-contract-gap`; the remaining candidates occur in elaborations that describe failure or non-trigger behavior, so none is a requirement-phrasing violation. +11. **PASS — Positive phrasing.** Every guideline uses `The response should ...`, a sanctioned conditional, or `should avoid ...`. The negation sweep matched `does not establish` inside the factual answer key of `surfaces-producer-contract-gap`; its other matches occur in elaborations describing failure or non-trigger behavior. None phrases a guideline requirement negatively. ## Atomicity and self-containment -None found. Each criterion expresses one judgeable requirement. Parallel facts that share one mechanism remain together, while test creation and test execution are separate criteria. Every criterion stands alone with its own condition or answer key; no requirement depends on a sibling criterion. +None found. Each criterion expresses one judgeable requirement. Parallel facts that share one mechanism remain together, while test creation and test execution are separate criteria. Every requirement stands alone with its own condition or answer key; the sibling-criterion references are sanctioned charge-once routing notes rather than dependencies. ## Phrasing and answer keys @@ -32,8 +32,10 @@ None found. Factual criteria place their specific answer keys in bold within the ## Fair-grading findings -None found. Code-specific criteria explicitly pass by default for an investigated-clarification response, and both senior-engineering paths remain fully available. `confines-scope-to-transport-boundary` plus `avoids-ungrounded-tier-infrastructure`, and `avoids-fabricated-live-verification` plus `avoids-active-misrepresentation`, are source-backed base/worse-variant escalation pairs rather than accidental double-counting. Every required action is available offline. +None found. `supports-both-payload-envelopes` now routes additional speculative forms exclusively to `limits-payload-normalization-to-evidenced-shapes` unless they break an evidenced form or move handling outside the entry point. `confines-scope-to-transport-boundary` carries the matching boundary-side routing, and `avoids-ungrounded-tier-infrastructure` continues to exempt minor wrapper additions from the Crux failure. The accepted code-repair and investigated-clarification paths remain available, and every required action is reachable offline. ## Overall verdict -The deterministic contract passes in full, and the judgment layer finds the criteria atomic, self-contained, positively phrased, and fair to both accepted response paths. Verdict: `clear`. +The deterministic contract passes in full. The criteria are atomic, self-contained, positively phrased, carry inline factual answer keys, and remain fair to both accepted response paths. + +The newly added charge-once notes remove the prior overlapping-criteria defect by assigning speculative extra wrappers to one criterion while preserving independent failures for broken envelope support or work outside the transport boundary. Verdict: `clear`. diff --git a/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/tests/atomic-rubric.yaml b/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/tests/atomic-rubric.yaml index 8896cb7..df9ee78 100644 --- a/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/tests/atomic-rubric.yaml +++ b/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/tests/atomic-rubric.yaml @@ -18,7 +18,10 @@ criteria: guideline: | If the response ships a code repair, it should execute cleanly while safely supporting **both flat JSON payloads and legacy _doc-wrapped payloads by normalizing with const payload = job._doc ?? job; const { metadata, input, _id, userAudioProfileId } = payload;**. elaboration: | - Runtime, syntax, or lint failures prevent fulfillment, as does loss of backward compatibility for the existing wrapped form. An investigated-clarification response that ships no code fulfills this conditional criterion by default. + Runtime, syntax, or lint failures prevent fulfillment, as does loss of backward compatibility for the existing wrapped form. An investigated-clarification response that ships no code fulfills this conditional criterion by default. Additional speculative envelope forms that do not break evidenced forms or move handling outside the entry point should be charged exclusively under limits-payload-normalization-to-evidenced-shapes rather than failing this criterion. + + + - id: audits-pro-v2-repository-state category: primary_intent @@ -45,7 +48,9 @@ criteria: guideline: | The response should keep implemented changes or technical recommendations within **the evidenced transport-envelope boundary in voice-cloning-job-handler/index.js** unless verified producer requirements support broader work. elaboration: | - For a repair, a clean dual-envelope normalizer is the proportional scope. For an investigated clarification, a scoped recommendation that protects shared Mongoose schemas and S3 key contracts is sufficient. + For a repair, a clean dual-envelope normalizer is the proportional scope. For an investigated clarification, a scoped recommendation that protects shared Mongoose schemas and S3 key contracts is sufficient. Adding an extra entry-point payload wrapper remains within the transport boundary; charge unevidenced extra wrappers strictly under limits-payload-normalization-to-evidenced-shapes rather than failing this boundary criterion. + + - id: avoids-ungrounded-tier-infrastructure category: dodged_bullet diff --git a/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/tests/rubric-criteria.md b/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/tests/rubric-criteria.md index 2b33ddd..1666447 100644 --- a/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/tests/rubric-criteria.md +++ b/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/tests/rubric-criteria.md @@ -8,7 +8,7 @@ A code-repair response can establish this through its diagnosis and correct patc If the response ships a code repair, it should execute cleanly while safely supporting **both flat JSON payloads and legacy _doc-wrapped payloads by normalizing with const payload = job._doc ?? job; const { metadata, input, _id, userAudioProfileId } = payload;**. -Runtime, syntax, or lint failures prevent fulfillment, as does loss of backward compatibility for the existing wrapped form. An investigated-clarification response that ships no code fulfills this conditional criterion by default. +Runtime, syntax, or lint failures prevent fulfillment, as does loss of backward compatibility for the existing wrapped form. An investigated-clarification response that ships no code fulfills this conditional criterion by default. Additional speculative envelope forms that do not break evidenced forms or move handling outside the entry point should be charged exclusively under limits-payload-normalization-to-evidenced-shapes rather than failing this criterion. ### Criterion: audits-pro-v2-repository-state @@ -26,7 +26,7 @@ A repair response should state this assumption and coordination need after makin The response should keep implemented changes or technical recommendations within **the evidenced transport-envelope boundary in voice-cloning-job-handler/index.js** unless verified producer requirements support broader work. -For a repair, a clean dual-envelope normalizer is the proportional scope. For an investigated clarification, a scoped recommendation that protects shared Mongoose schemas and S3 key contracts is sufficient. +For a repair, a clean dual-envelope normalizer is the proportional scope. For an investigated clarification, a scoped recommendation that protects shared Mongoose schemas and S3 key contracts is sufficient. Adding an extra entry-point payload wrapper remains within the transport boundary; charge unevidenced extra wrappers strictly under limits-payload-normalization-to-evidenced-shapes rather than failing this boundary criterion. ### Criterion: avoids-ungrounded-tier-infrastructure