From 18574c0ca6992b4e0c92546d90597abb6bafa6a6 Mon Sep 17 00:00:00 2001 From: Eric Bell Date: Sat, 26 Sep 2026 13:06:56 -0400 Subject: [PATCH] final detectors - 1 problem --- .../detector-rubric-coverage.inputs.json | 2 +- .../detectors/detector-rubric-coverage.md | 2 +- .../detectors/detector-rubric-form.inputs.json | 2 +- .../detectors/detector-rubric-form.md | 15 ++++++++++----- 4 files changed, 13 insertions(+), 8 deletions(-) diff --git a/worker-toolkit-potion-polyglot/harbor-tasks/mishandle_pro_v2/detectors/detector-rubric-coverage.inputs.json b/worker-toolkit-potion-polyglot/harbor-tasks/mishandle_pro_v2/detectors/detector-rubric-coverage.inputs.json index 44c44d8..1a63501 100644 --- a/worker-toolkit-potion-polyglot/harbor-tasks/mishandle_pro_v2/detectors/detector-rubric-coverage.inputs.json +++ b/worker-toolkit-potion-polyglot/harbor-tasks/mishandle_pro_v2/detectors/detector-rubric-coverage.inputs.json @@ -1,6 +1,6 @@ { "version": 1, - "capturedAt": "2026-09-25T20:25:25.814Z", + "capturedAt": "2026-09-26T16:36:53.325Z", "capturedBy": "stamp", "inputs": { "prompt": "29e2eb28448679a65ae264372ddf7d993e5752d0f295bf5557cc5b1578265a29", diff --git a/worker-toolkit-potion-polyglot/harbor-tasks/mishandle_pro_v2/detectors/detector-rubric-coverage.md b/worker-toolkit-potion-polyglot/harbor-tasks/mishandle_pro_v2/detectors/detector-rubric-coverage.md index 7dbcb3a..350ad3a 100644 --- a/worker-toolkit-potion-polyglot/harbor-tasks/mishandle_pro_v2/detectors/detector-rubric-coverage.md +++ b/worker-toolkit-potion-polyglot/harbor-tasks/mishandle_pro_v2/detectors/detector-rubric-coverage.md @@ -52,4 +52,4 @@ The holistic rubric contains no heavy penalty targeting the overall score, so th ## Overall verdict -The conversion is equivalent in both directions. All load-bearing requirements, penalty triggers, severity qualifiers, and protected non-triggers map to criteria; the context survives verbatim; no criterion invents scope; and the absence of Crux criteria matches the absence of any overall-score heavy penalty. +The fresh comparison finds the conversion equivalent in both directions. All load-bearing requirements, penalty triggers, severity qualifiers, and protected non-triggers map to criteria; the context survives verbatim; no criterion invents scope; and the absence of Crux criteria matches the absence of any overall-score heavy penalty. diff --git a/worker-toolkit-potion-polyglot/harbor-tasks/mishandle_pro_v2/detectors/detector-rubric-form.inputs.json b/worker-toolkit-potion-polyglot/harbor-tasks/mishandle_pro_v2/detectors/detector-rubric-form.inputs.json index 49220f4..18bddbc 100644 --- a/worker-toolkit-potion-polyglot/harbor-tasks/mishandle_pro_v2/detectors/detector-rubric-form.inputs.json +++ b/worker-toolkit-potion-polyglot/harbor-tasks/mishandle_pro_v2/detectors/detector-rubric-form.inputs.json @@ -1,6 +1,6 @@ { "version": 1, - "capturedAt": "2026-09-25T20:25:25.806Z", + "capturedAt": "2026-09-26T16:39:38.284Z", "capturedBy": "stamp", "inputs": { "prompt": "29e2eb28448679a65ae264372ddf7d993e5752d0f295bf5557cc5b1578265a29", diff --git a/worker-toolkit-potion-polyglot/harbor-tasks/mishandle_pro_v2/detectors/detector-rubric-form.md b/worker-toolkit-potion-polyglot/harbor-tasks/mishandle_pro_v2/detectors/detector-rubric-form.md index a38b4fe..dcd40b8 100644 --- a/worker-toolkit-potion-polyglot/harbor-tasks/mishandle_pro_v2/detectors/detector-rubric-form.md +++ b/worker-toolkit-potion-polyglot/harbor-tasks/mishandle_pro_v2/detectors/detector-rubric-form.md @@ -1,7 +1,7 @@ --- detector: detector-rubric-form -verdict: clear -confidence: HIGH +verdict: material-issues +confidence: MEDIUM --- # Rubric-form check: mishandle_pro_v2 @@ -10,7 +10,7 @@ Assessed: harbor-tasks/mishandle_pro_v2/tests/atomic-rubric.yaml ## Deterministic contract -1. **PASS — Parses as YAML.** The toolkit staging script parsed and staged the file successfully. +1. **PASS — Parses as YAML.** A fresh YAML parse loaded the file as a mapping with a string `task` and a 14-item `criteria` list. 2. **PASS — `task` names this task.** The value is `mishandle_pro_v2`. 3. **PASS — Criteria count.** The file contains 14 criteria. 4. **PASS — Kebab-case unique ids.** All 14 ids match the required pattern and are unique. @@ -28,7 +28,12 @@ None found. Each criterion is independently judgeable. The contract-gap and boun ## Phrasing and answer keys -None found. Factual requirements carry their answer keys inline in bold, while behavioral requirements state observable response properties. Elaborations clarify fulfillment, failure, non-triggers, or charge routing without introducing new requirements. +### Syntax/lint requirement exists only in an elaboration + +- **Criterion:** `normalizes-supported-envelope-shapes`. +- **Where:** “This criterion fails if flat payloads still crash, the repair breaks compatibility with `_doc`-wrapped payloads, or the change introduces syntax or lint failures that prevent the normalizer from working.” +- **Why:** The guideline requires dual-envelope runtime behavior and freedom from the stated `TypeError`, but it does not require passing syntax or lint checks. A working normalizer can still fail lint, so this final fail condition is an independent requirement that a grader reading guidelines alone would miss. +- **Suggested split or rewrite:** Remove the syntax/lint clause from this elaboration and add a standalone guideline such as: “The response's changes should leave **`voice-cloning-job-handler/index.js` syntactically valid and passing the applicable lint checks**.” ## Fair-grading findings @@ -36,4 +41,4 @@ None found. The paired criteria encode genuinely distinct or escalating failures ## Overall verdict -The deterministic contract passes in full, and the judgment layer is clear. The rubric is atomic, self-contained, positively phrased, and reachable in the task environment, with explicit routing for its escalation pairs and no unsupported scoring mechanics. +The deterministic contract passes in full, and the criteria are otherwise atomic, self-contained, positively phrased, and fairly routed. The judgment layer nevertheless has one load-bearing elaboration-discipline failure: syntax/lint validity affects the criterion verdict but appears only in its elaboration. Under this detector's definitions, that drives a `material-issues` verdict until the requirement is moved into a guideline.