|
|
|
@@ -1,7 +1,7 @@
|
|
|
|
---
|
|
|
|
---
|
|
|
|
detector: detector-rubric-form
|
|
|
|
detector: detector-rubric-form
|
|
|
|
verdict: clear
|
|
|
|
verdict: material-issues
|
|
|
|
confidence: HIGH
|
|
|
|
confidence: MEDIUM
|
|
|
|
---
|
|
|
|
---
|
|
|
|
|
|
|
|
|
|
|
|
# Rubric-form check: mishandle_pro_v2
|
|
|
|
# Rubric-form check: mishandle_pro_v2
|
|
|
|
@@ -10,7 +10,7 @@ Assessed: harbor-tasks/mishandle_pro_v2/tests/atomic-rubric.yaml
|
|
|
|
|
|
|
|
|
|
|
|
## Deterministic contract
|
|
|
|
## Deterministic contract
|
|
|
|
|
|
|
|
|
|
|
|
1. **PASS — Parses as YAML.** The toolkit staging script parsed and staged the file successfully.
|
|
|
|
1. **PASS — Parses as YAML.** A fresh YAML parse loaded the file as a mapping with a string `task` and a 14-item `criteria` list.
|
|
|
|
2. **PASS — `task` names this task.** The value is `mishandle_pro_v2`.
|
|
|
|
2. **PASS — `task` names this task.** The value is `mishandle_pro_v2`.
|
|
|
|
3. **PASS — Criteria count.** The file contains 14 criteria.
|
|
|
|
3. **PASS — Criteria count.** The file contains 14 criteria.
|
|
|
|
4. **PASS — Kebab-case unique ids.** All 14 ids match the required pattern and are unique.
|
|
|
|
4. **PASS — Kebab-case unique ids.** All 14 ids match the required pattern and are unique.
|
|
|
|
@@ -28,7 +28,12 @@ None found. Each criterion is independently judgeable. The contract-gap and boun
|
|
|
|
|
|
|
|
|
|
|
|
## Phrasing and answer keys
|
|
|
|
## Phrasing and answer keys
|
|
|
|
|
|
|
|
|
|
|
|
None found. Factual requirements carry their answer keys inline in bold, while behavioral requirements state observable response properties. Elaborations clarify fulfillment, failure, non-triggers, or charge routing without introducing new requirements.
|
|
|
|
### Syntax/lint requirement exists only in an elaboration
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
- **Criterion:** `normalizes-supported-envelope-shapes`.
|
|
|
|
|
|
|
|
- **Where:** “This criterion fails if flat payloads still crash, the repair breaks compatibility with `_doc`-wrapped payloads, or the change introduces syntax or lint failures that prevent the normalizer from working.”
|
|
|
|
|
|
|
|
- **Why:** The guideline requires dual-envelope runtime behavior and freedom from the stated `TypeError`, but it does not require passing syntax or lint checks. A working normalizer can still fail lint, so this final fail condition is an independent requirement that a grader reading guidelines alone would miss.
|
|
|
|
|
|
|
|
- **Suggested split or rewrite:** Remove the syntax/lint clause from this elaboration and add a standalone guideline such as: “The response's changes should leave **`voice-cloning-job-handler/index.js` syntactically valid and passing the applicable lint checks**.”
|
|
|
|
|
|
|
|
|
|
|
|
## Fair-grading findings
|
|
|
|
## Fair-grading findings
|
|
|
|
|
|
|
|
|
|
|
|
@@ -36,4 +41,4 @@ None found. The paired criteria encode genuinely distinct or escalating failures
|
|
|
|
|
|
|
|
|
|
|
|
## Overall verdict
|
|
|
|
## Overall verdict
|
|
|
|
|
|
|
|
|
|
|
|
The deterministic contract passes in full, and the judgment layer is clear. The rubric is atomic, self-contained, positively phrased, and reachable in the task environment, with explicit routing for its escalation pairs and no unsupported scoring mechanics.
|
|
|
|
The deterministic contract passes in full, and the criteria are otherwise atomic, self-contained, positively phrased, and fairly routed. The judgment layer nevertheless has one load-bearing elaboration-discipline failure: syntax/lint validity affects the criterion verdict but appears only in its elaboration. Under this detector's definitions, that drives a `material-issues` verdict until the requirement is moved into a guideline.
|
|
|
|
|