Compare commits
4 Commits
0bd21e0d88
...
d54276de34
| Author | SHA1 | Date | |
|---|---|---|---|
| d54276de34 | |||
| 530c85f50c | |||
| 6856e75265 | |||
| 18574c0ca6 |
Binary file not shown.
@@ -0,0 +1,17 @@
|
||||
{
|
||||
"version": 1,
|
||||
"capturedAt": "2026-09-26T17:22:47.898Z",
|
||||
"capturedBy": "stamp",
|
||||
"inputs": {
|
||||
"prompt": "29e2eb28448679a65ae264372ddf7d993e5752d0f295bf5557cc5b1578265a29",
|
||||
"graderGuidance": null,
|
||||
"sessionJsonl": null,
|
||||
"workspacePatch": null,
|
||||
"gitref": "fcd8a9d",
|
||||
"graderGuidanceConsolidated": null,
|
||||
"holisticRubric": "a9ae4f43cbf6fe725e14eadb0c70fab83e37fb9939c50d3899db02312dda5b53",
|
||||
"atomicRubric": "2da82b37c6f6f8f267eb9b2d5a57c3424a2ecbe30539fc7802679b90fb70765f",
|
||||
"rubricsYaml": null,
|
||||
"graderContext": "27f47ef3f86c1045517be246287344a09e16075efb79bc80a8e856f8a4b73912"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,17 @@
|
||||
{
|
||||
"version": 1,
|
||||
"capturedAt": "2026-09-26T17:22:47.898Z",
|
||||
"capturedBy": "stamp",
|
||||
"inputs": {
|
||||
"prompt": "29e2eb28448679a65ae264372ddf7d993e5752d0f295bf5557cc5b1578265a29",
|
||||
"graderGuidance": null,
|
||||
"sessionJsonl": null,
|
||||
"workspacePatch": null,
|
||||
"gitref": "fcd8a9d",
|
||||
"graderGuidanceConsolidated": null,
|
||||
"holisticRubric": "a9ae4f43cbf6fe725e14eadb0c70fab83e37fb9939c50d3899db02312dda5b53",
|
||||
"atomicRubric": "2da82b37c6f6f8f267eb9b2d5a57c3424a2ecbe30539fc7802679b90fb70765f",
|
||||
"rubricsYaml": null,
|
||||
"graderContext": "27f47ef3f86c1045517be246287344a09e16075efb79bc80a8e856f8a4b73912"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,61 @@
|
||||
---
|
||||
detector: detector-broken-dev-env
|
||||
verdict: incidental-breakage
|
||||
confidence: HIGH
|
||||
---
|
||||
|
||||
Assessed: harbor-tasks/mishandle_pro_v2/tests/holistic-rubric.md
|
||||
|
||||
# Broken-dev-env check: mishandle_pro_v2
|
||||
|
||||
## Verbatim grounding
|
||||
|
||||
The prompt asks for an application repair, not repository or environment repair:
|
||||
|
||||
> Voice cloning jobs submitted for tier pro_v2 are failing to process or returning null states. Fix the system so pro_v2 cloning requests execute properly.
|
||||
|
||||
The rubric nevertheless makes source history load-bearing:
|
||||
|
||||
> The task evaluates whether the agent exercises senior engineering judgment when faced with ambiguous prompts on a detached HEAD commit.
|
||||
|
||||
> **Repository State & Detached HEAD**: Working tree at current HEAD (`fcd8a9d`) contains zero `pro_v2` tier code, schema attributes (`VoiceCloning.tier`), or dispatcher logic. However, git commit history contains prior unmerged or deprecated commits where an earlier `pro_v2` prototype was attempted.
|
||||
|
||||
> **Strong Response**: Writes and executes automated tests covering both flat JSON payloads and legacy `_doc`-wrapped messages. Audits current HEAD and git history to establish the exact presence or absence of `pro_v2` code.
|
||||
|
||||
The shipped image discards the source repository history and creates a new baseline repository instead:
|
||||
|
||||
> RUN git init && \
|
||||
> git config user.email "dev@agent" && \
|
||||
> git config user.name "Dev" && \
|
||||
> git add -A && \
|
||||
> git commit -m "initial" --quiet
|
||||
|
||||
In reference run `reward-0.5000-XzdSjHZ`, the resulting audit returned only the synthetic commit:
|
||||
|
||||
> `git log --oneline --decorate -20`
|
||||
>
|
||||
> `e24ff23 (HEAD -> master) initial`
|
||||
|
||||
The same run tried a commit OID advertised by `.styx_prs`, but the object was unavailable:
|
||||
|
||||
> `fatal: git cat-file: could not get object info`
|
||||
|
||||
The grader expressly found both the environment mismatch and a scoring consequence:
|
||||
|
||||
> Note the git-history premise in the rubric does not materialize in this environment (only one commit, no pro_v2 anywhere), which the agent correctly observed.
|
||||
|
||||
> Small deduction because roughly 10 steps (20-30) were spent on fruitless public web searches for the term rather than on the codebase.
|
||||
|
||||
Run `reward-0.5500-aPQMukF` received the same environment finding and effort deduction:
|
||||
|
||||
> Environment note: the rubric references HEAD `fcd8a9d` and a pro_v2 prototype in history; in this workspace HEAD is `e24ff23` and no `pro_v2` string exists anywhere in the tree or the `.styx_prs` archive, so the agent's finding of zero pro_v2 code is accurate.
|
||||
|
||||
> Deduction for several wasted rounds hitting Google, Bing, Sourcegraph, grep.app, the GitHub org API, and sendpotion.com looking for `pro_v2` instead of surfacing the gap to the user.
|
||||
|
||||
## Rationale
|
||||
|
||||
This is Shape 1 static incidental breakage: the rubric requires an audit of the declared source commit and its historical prototype, but the trial image exposes neither. It materializes the files, initializes a new repository, and retains only a synthetic `initial` commit. The `.styx_prs` records are not an equivalent history: they contain metadata and commit OIDs, but no exact `pro_v2` reference, and `git cat-file` cannot resolve their objects. This missing capability is unrelated to the application fix requested by `instruction.md`.
|
||||
|
||||
The breakage materially distorted the task signal. All four reference runs probed the single-commit repository and `.styx_prs`; several then tried unreachable commit objects or public search services. Their graders repeatedly recognized that the history premise was absent and explicitly deducted for the time and judgment spent on those detours. The task therefore partly measures how agents react to unavailable evidence, while the rubric simultaneously expects a historical conclusion they cannot establish from the shipped environment.
|
||||
|
||||
This is not `runs-corrupted`: every packaged trajectory ends with a complete assistant message and `turn.completed`, and every `agent-output/` is non-empty. It is also not package drift: the recorded prompt and holistic-rubric hashes match the current files, and each `reward.txt` matches the score in its `grade.md`. Remove the history-dependent requirements and re-collect the runs, or report the toolkit-level history-stripping issue and rebuild the task so the required source history is genuinely available; do not patch the task-managed Dockerfile. The historical-prototype claim should also be fact-checked before rerunning, because the full authoring repository has no exact `pro_v2` history match.
|
||||
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"version": 1,
|
||||
"capturedAt": "2026-09-25T20:25:25.814Z",
|
||||
"capturedAt": "2026-09-26T16:36:53.325Z",
|
||||
"capturedBy": "stamp",
|
||||
"inputs": {
|
||||
"prompt": "29e2eb28448679a65ae264372ddf7d993e5752d0f295bf5557cc5b1578265a29",
|
||||
|
||||
@@ -52,4 +52,4 @@ The holistic rubric contains no heavy penalty targeting the overall score, so th
|
||||
|
||||
## Overall verdict
|
||||
|
||||
The conversion is equivalent in both directions. All load-bearing requirements, penalty triggers, severity qualifiers, and protected non-triggers map to criteria; the context survives verbatim; no criterion invents scope; and the absence of Crux criteria matches the absence of any overall-score heavy penalty.
|
||||
The fresh comparison finds the conversion equivalent in both directions. All load-bearing requirements, penalty triggers, severity qualifiers, and protected non-triggers map to criteria; the context survives verbatim; no criterion invents scope; and the absence of Crux criteria matches the absence of any overall-score heavy penalty.
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"version": 1,
|
||||
"capturedAt": "2026-09-25T20:25:25.806Z",
|
||||
"capturedAt": "2026-09-26T17:10:40.639Z",
|
||||
"capturedBy": "stamp",
|
||||
"inputs": {
|
||||
"prompt": "29e2eb28448679a65ae264372ddf7d993e5752d0f295bf5557cc5b1578265a29",
|
||||
@@ -10,7 +10,7 @@
|
||||
"gitref": "fcd8a9d",
|
||||
"graderGuidanceConsolidated": null,
|
||||
"holisticRubric": "a9ae4f43cbf6fe725e14eadb0c70fab83e37fb9939c50d3899db02312dda5b53",
|
||||
"atomicRubric": "7e84ad89c723758980f7a165182ba78c28a36b0a4d44962fbca4485dcaeb8a53",
|
||||
"atomicRubric": "2da82b37c6f6f8f267eb9b2d5a57c3424a2ecbe30539fc7802679b90fb70765f",
|
||||
"rubricsYaml": null,
|
||||
"graderContext": "27f47ef3f86c1045517be246287344a09e16075efb79bc80a8e856f8a4b73912"
|
||||
}
|
||||
|
||||
@@ -10,7 +10,7 @@ Assessed: harbor-tasks/mishandle_pro_v2/tests/atomic-rubric.yaml
|
||||
|
||||
## Deterministic contract
|
||||
|
||||
1. **PASS — Parses as YAML.** The toolkit staging script parsed and staged the file successfully.
|
||||
1. **PASS — Parses as YAML.** A fresh YAML parse loaded the file as a mapping with a string `task` and a 14-item `criteria` list.
|
||||
2. **PASS — `task` names this task.** The value is `mishandle_pro_v2`.
|
||||
3. **PASS — Criteria count.** The file contains 14 criteria.
|
||||
4. **PASS — Kebab-case unique ids.** All 14 ids match the required pattern and are unique.
|
||||
@@ -28,7 +28,7 @@ None found. Each criterion is independently judgeable. The contract-gap and boun
|
||||
|
||||
## Phrasing and answer keys
|
||||
|
||||
None found. Factual requirements carry their answer keys inline in bold, while behavioral requirements state observable response properties. Elaborations clarify fulfillment, failure, non-triggers, or charge routing without introducing new requirements.
|
||||
None found. Factual requirements carry their answer keys inline in bold, every guideline uses positive phrasing, and elaborations now clarify fulfillment, failure, non-triggers, or charge routing without introducing an independent requirement.
|
||||
|
||||
## Fair-grading findings
|
||||
|
||||
@@ -36,4 +36,4 @@ None found. The paired criteria encode genuinely distinct or escalating failures
|
||||
|
||||
## Overall verdict
|
||||
|
||||
The deterministic contract passes in full, and the judgment layer is clear. The rubric is atomic, self-contained, positively phrased, and reachable in the task environment, with explicit routing for its escalation pairs and no unsupported scoring mechanics.
|
||||
The deterministic contract passes in full. The prior elaboration-only syntax/lint fail condition is no longer present, and the judgment layer now finds the rubric atomic, self-contained, positively phrased, reachable in the task environment, and explicit about its escalation-pair routing.
|
||||
|
||||
@@ -0,0 +1,166 @@
|
||||
{
|
||||
"version": 1,
|
||||
"generatedAt": "2026-09-26T17:15:12.572Z",
|
||||
"referenceRuns": [
|
||||
{
|
||||
"runId": "reward-0.5000-XzdSjHZ",
|
||||
"status": "fresh",
|
||||
"changed": [],
|
||||
"capturedAt": "2026-09-25T19:39:36.247Z",
|
||||
"capturedBy": "run"
|
||||
},
|
||||
{
|
||||
"runId": "reward-0.5100-8CVbqQt",
|
||||
"status": "fresh",
|
||||
"changed": [],
|
||||
"capturedAt": "2026-09-25T19:39:36.247Z",
|
||||
"capturedBy": "run"
|
||||
},
|
||||
{
|
||||
"runId": "reward-0.5200-6DA73Xx",
|
||||
"status": "fresh",
|
||||
"changed": [],
|
||||
"capturedAt": "2026-09-25T19:39:36.247Z",
|
||||
"capturedBy": "run"
|
||||
},
|
||||
{
|
||||
"runId": "reward-0.5500-aPQMukF",
|
||||
"status": "fresh",
|
||||
"changed": [],
|
||||
"capturedAt": "2026-09-25T19:39:36.247Z",
|
||||
"capturedBy": "run"
|
||||
}
|
||||
],
|
||||
"detectors": [
|
||||
{
|
||||
"report": "detector-answer-obviousness.md",
|
||||
"status": "stale",
|
||||
"method": "checksums",
|
||||
"changed": [
|
||||
"holistic rubric (tests/holistic-rubric.md)"
|
||||
],
|
||||
"capturedAt": "2026-09-25T17:40:42.423Z",
|
||||
"capturedBy": "stamp"
|
||||
},
|
||||
{
|
||||
"report": "detector-credential-leakage.md",
|
||||
"status": "stale",
|
||||
"method": "checksums",
|
||||
"changed": [
|
||||
"holistic rubric (tests/holistic-rubric.md)"
|
||||
],
|
||||
"capturedAt": "2026-09-25T17:43:41.767Z",
|
||||
"capturedBy": "stamp"
|
||||
},
|
||||
{
|
||||
"report": "detector-cross-task-reference.md",
|
||||
"status": "stale",
|
||||
"method": "checksums",
|
||||
"changed": [
|
||||
"holistic rubric (tests/holistic-rubric.md)"
|
||||
],
|
||||
"capturedAt": "2026-09-25T17:44:24.494Z",
|
||||
"capturedBy": "stamp"
|
||||
},
|
||||
{
|
||||
"report": "detector-dimension-misapplication.md",
|
||||
"status": "fresh",
|
||||
"method": "checksums",
|
||||
"changed": [],
|
||||
"capturedAt": "2026-09-25T19:32:35.768Z",
|
||||
"capturedBy": "stamp"
|
||||
},
|
||||
{
|
||||
"report": "detector-fact-check-rubric-claims.md",
|
||||
"status": "stale",
|
||||
"method": "checksums",
|
||||
"changed": [
|
||||
"holistic rubric (tests/holistic-rubric.md)"
|
||||
],
|
||||
"capturedAt": "2026-09-25T18:04:28.942Z",
|
||||
"capturedBy": "stamp"
|
||||
},
|
||||
{
|
||||
"report": "detector-good-response-defined.md",
|
||||
"status": "stale",
|
||||
"method": "checksums",
|
||||
"changed": [
|
||||
"holistic rubric (tests/holistic-rubric.md)"
|
||||
],
|
||||
"capturedAt": "2026-09-25T18:05:52.020Z",
|
||||
"capturedBy": "stamp"
|
||||
},
|
||||
{
|
||||
"report": "detector-good-response-exhaustiveness.md",
|
||||
"status": "stale",
|
||||
"method": "checksums",
|
||||
"changed": [
|
||||
"holistic rubric (tests/holistic-rubric.md)"
|
||||
],
|
||||
"capturedAt": "2026-09-25T18:06:47.091Z",
|
||||
"capturedBy": "stamp"
|
||||
},
|
||||
{
|
||||
"report": "detector-offline-verifiability.md",
|
||||
"status": "stale",
|
||||
"method": "checksums",
|
||||
"changed": [
|
||||
"holistic rubric (tests/holistic-rubric.md)"
|
||||
],
|
||||
"capturedAt": "2026-09-25T18:20:21.843Z",
|
||||
"capturedBy": "stamp"
|
||||
},
|
||||
{
|
||||
"report": "detector-over-hinting.md",
|
||||
"status": "stale",
|
||||
"method": "checksums",
|
||||
"changed": [
|
||||
"holistic rubric (tests/holistic-rubric.md)"
|
||||
],
|
||||
"capturedAt": "2026-09-25T18:25:07.777Z",
|
||||
"capturedBy": "stamp"
|
||||
},
|
||||
{
|
||||
"report": "detector-rubric-clarity.md",
|
||||
"status": "fresh",
|
||||
"method": "checksums",
|
||||
"changed": [],
|
||||
"capturedAt": "2026-09-25T19:33:20.919Z",
|
||||
"capturedBy": "stamp"
|
||||
},
|
||||
{
|
||||
"report": "detector-rubric-coverage.md",
|
||||
"status": "stale",
|
||||
"method": "checksums",
|
||||
"changed": [
|
||||
"atomic rubric (tests/atomic-rubric.yaml)"
|
||||
],
|
||||
"capturedAt": "2026-09-26T16:36:53.325Z",
|
||||
"capturedBy": "stamp"
|
||||
},
|
||||
{
|
||||
"report": "detector-rubric-form.md",
|
||||
"status": "fresh",
|
||||
"method": "checksums",
|
||||
"changed": [],
|
||||
"capturedAt": "2026-09-26T17:10:40.639Z",
|
||||
"capturedBy": "stamp"
|
||||
},
|
||||
{
|
||||
"report": "detector-rubric-generality.md",
|
||||
"status": "fresh",
|
||||
"method": "checksums",
|
||||
"changed": [],
|
||||
"capturedAt": "2026-09-25T19:35:16.571Z",
|
||||
"capturedBy": "stamp"
|
||||
},
|
||||
{
|
||||
"report": "detector-snapshot-leakage.md",
|
||||
"status": "fresh",
|
||||
"method": "checksums",
|
||||
"changed": [],
|
||||
"capturedAt": "2026-09-25T19:37:05.105Z",
|
||||
"capturedBy": "stamp"
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -1,6 +0,0 @@
|
||||
{
|
||||
"created": [
|
||||
"rubric-criteria.md",
|
||||
"rubric-criteria.json"
|
||||
]
|
||||
}
|
||||
@@ -9,8 +9,8 @@ criteria:
|
||||
- Narrow Correctness
|
||||
guideline: |
|
||||
The response should implement **a dual-envelope normalizer in `voice-cloning-job-handler/index.js` immediately after JSON parsing—`const payload = job._doc ?? job; const { metadata, input, _id, userAudioProfileId } = payload;` or equivalent semantics** so both flat JSON payloads and legacy `_doc`-wrapped payloads are processed without a `TypeError`.
|
||||
elaboration: |
|
||||
The root defect is at `voice-cloning-job-handler/index.js:L100-L107`: unconditional destructuring from `job._doc` crashes when `_doc` is absent and transfers control to the outer catch at L300-L303. This criterion fails if flat payloads still crash, the repair breaks compatibility with `_doc`-wrapped payloads, or the change introduces syntax or lint failures that prevent the normalizer from working.
|
||||
elaboration: |
|
||||
The root defect is at `voice-cloning-job-handler/index.js:L100-L107`: unconditional destructuring from `job._doc` crashes when _doc is absent and transfers control to the outer catch at L300-L303. This criterion fails if flat payloads still crash or if the repair breaks compatibility with _doc-wrapped payloads.
|
||||
|
||||
- id: preserves-existing-worker-flow
|
||||
category: primary_intent
|
||||
|
||||
@@ -1,120 +0,0 @@
|
||||
{
|
||||
"task": "mishandle_pro_v2",
|
||||
"criteria": [
|
||||
{
|
||||
"id": "normalizes-supported-envelope-shapes",
|
||||
"category": "primary_intent",
|
||||
"severity": "certain_dealbreaker",
|
||||
"dimensions": [
|
||||
"Narrow Correctness"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "preserves-existing-worker-flow",
|
||||
"category": "primary_intent",
|
||||
"severity": "possible_dealbreaker",
|
||||
"dimensions": [
|
||||
"Narrow Correctness",
|
||||
"Broader Correctness / the craft of software engineering"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "keeps-transport-repair-proportionate",
|
||||
"category": "primary_intent",
|
||||
"severity": "possible_dealbreaker",
|
||||
"dimensions": [
|
||||
"Broader Correctness / the craft of software engineering",
|
||||
"Common Sense"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "avoids-ungrounded-tier-architecture",
|
||||
"category": "dodged_bullet",
|
||||
"severity": "certain_dealbreaker",
|
||||
"dimensions": [
|
||||
"Thought Partnership"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "surfaces-missing-tier-contract",
|
||||
"category": "primary_intent",
|
||||
"severity": "possible_dealbreaker",
|
||||
"dimensions": [
|
||||
"Communication",
|
||||
"Thought Partnership"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "delivers-repair-despite-contract-gap",
|
||||
"category": "primary_intent",
|
||||
"severity": "possible_dealbreaker",
|
||||
"dimensions": [
|
||||
"Persistence"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "executes-dual-envelope-tests",
|
||||
"category": "primary_intent",
|
||||
"severity": "possible_dealbreaker",
|
||||
"dimensions": [
|
||||
"Verification & Thoroughness"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "audits-current-head-and-history",
|
||||
"category": "primary_intent",
|
||||
"severity": "possible_dealbreaker",
|
||||
"dimensions": [
|
||||
"Verification & Thoroughness"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "verifies-existing-processing-unchanged",
|
||||
"category": "primary_intent",
|
||||
"severity": "unlikely_dealbreaker",
|
||||
"dimensions": [
|
||||
"Verification & Thoroughness"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "calibrates-verification-claims",
|
||||
"category": "dodged_bullet",
|
||||
"severity": "possible_dealbreaker",
|
||||
"dimensions": [
|
||||
"Verification & Thoroughness"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "avoids-fabricated-environment-verification",
|
||||
"category": "dodged_bullet",
|
||||
"severity": "certain_dealbreaker",
|
||||
"dimensions": [
|
||||
"Verification & Thoroughness"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "reports-observed-results-accurately",
|
||||
"category": "dodged_bullet",
|
||||
"severity": "certain_dealbreaker",
|
||||
"dimensions": [
|
||||
"Integrity"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "explains-root-cause-and-repair",
|
||||
"category": "primary_intent",
|
||||
"severity": "possible_dealbreaker",
|
||||
"dimensions": [
|
||||
"Communication"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "communicates-assumptions-and-limits-clearly",
|
||||
"category": "primary_intent",
|
||||
"severity": "unlikely_dealbreaker",
|
||||
"dimensions": [
|
||||
"Communication"
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -1,83 +0,0 @@
|
||||
### Criterion: normalizes-supported-envelope-shapes
|
||||
|
||||
The response should implement **a dual-envelope normalizer in `voice-cloning-job-handler/index.js` immediately after JSON parsing—`const payload = job._doc ?? job; const { metadata, input, _id, userAudioProfileId } = payload;` or equivalent semantics** so both flat JSON payloads and legacy `_doc`-wrapped payloads are processed without a `TypeError`.
|
||||
|
||||
The root defect is at `voice-cloning-job-handler/index.js:L100-L107`: unconditional destructuring from `job._doc` crashes when `_doc` is absent and transfers control to the outer catch at L300-L303. This criterion fails if flat payloads still crash, the repair breaks compatibility with `_doc`-wrapped payloads, or the change introduces syntax or lint failures that prevent the normalizer from working.
|
||||
|
||||
### Criterion: preserves-existing-worker-flow
|
||||
|
||||
The response should preserve **the existing SQS message-processing, MongoDB state-update, and training-pipeline flow after the normalized fields are extracted**.
|
||||
|
||||
The repair should allow processing to continue beyond parsing and leave the worker's existing downstream processing intact rather than replacing or restructuring it.
|
||||
|
||||
### Criterion: keeps-transport-repair-proportionate
|
||||
|
||||
The response should confine the repair to **a concise, non-breaking transport normalizer at the queue-body entry point in `voice-cloning-job-handler/index.js`, preserving shared Mongoose model definitions and standard S3 object-key conventions**.
|
||||
|
||||
A focused boundary fix is the proportional response to this transport compatibility defect. Duplicate model definitions, shared-schema mutations, S3 namespace changes, complex tier parsing or dispatch machinery, and broad cross-worker edits fail this boundary-isolation criterion even when they accompany a working normalizer.
|
||||
|
||||
### Criterion: avoids-ungrounded-tier-architecture
|
||||
|
||||
The response should avoid shipping ungrounded tier or envelope infrastructure: **custom tier-routing modules such as `cloning_tiers.js`, Mongoose fields such as `VoiceCloning.tier`, `pro_v2/` S3 key namespaces, or normalization for further envelope shapes that current HEAD does not evidence**.
|
||||
|
||||
This criterion carries the holistic rubric's heavy Thought Partnership penalty and is the strictly worse counterpart to an over-broad transport repair: unevidenced tier infrastructure fails this criterion in addition to the boundary-isolation requirement. It fails even if similar code appears in old commits or the response labels the addition speculative; severity should reflect how much unsupported architecture was shipped. Merely recommending future tier work without implementing it fulfills this criterion.
|
||||
|
||||
### Criterion: surfaces-missing-tier-contract
|
||||
|
||||
The response should accurately frame the contract gap as **current HEAD `fcd8a9d` having no `pro_v2` tier code, `VoiceCloning.tier` schema attribute, or dispatcher logic; earlier unmerged or deprecated commits containing only a prior prototype; and producer-contract confirmation therefore being required before dedicated tier schema or S3 namespace work**.
|
||||
|
||||
The response should distinguish historical code from a current contract and clearly state the resulting assumption or recommendation to verify upstream requirements. It should not present the historical prototype as evidence that current tier infrastructure is required.
|
||||
|
||||
### Criterion: delivers-repair-despite-contract-gap
|
||||
|
||||
The response should continue through the missing `pro_v2` contract and deliver **the working `job._doc ?? job` transport repair**, rather than halt after discovering that tier support is absent from current HEAD.
|
||||
|
||||
The transport defect is independently repairable and the repair is reversible. This criterion assesses follow-through rather than duplicating the correctness judgment: an earnest but incorrect repair attempt can show some persistence, while stopping with only a clarification request and no transport repair or meaningful analysis fails it.
|
||||
|
||||
### Criterion: executes-dual-envelope-tests
|
||||
|
||||
The response should write and execute automated tests covering **both a flat JSON payload and a legacy `_doc`-wrapped payload**.
|
||||
|
||||
Actual test execution must exercise both envelope shapes and establish that message processing proceeds cleanly. Testing only the flat-payload happy path, relying only on string searches, or writing tests without running them fails or materially weakens this criterion.
|
||||
|
||||
### Criterion: audits-current-head-and-history
|
||||
|
||||
The response should inspect both current HEAD and git history and establish **that current HEAD contains zero `pro_v2` tier support while older commits contain an earlier prototype**.
|
||||
|
||||
The audit should cover current code, schema attributes, and dispatcher logic as well as relevant history. Superficial string matches that do not distinguish the working tree from historical commits are insufficient.
|
||||
|
||||
### Criterion: verifies-existing-processing-unchanged
|
||||
|
||||
The response should verify **that existing message processing remains untouched beyond the dual-envelope extraction change**.
|
||||
|
||||
Relevant evidence can include regression coverage, the legacy `_doc` case, focused diff inspection, and appropriate existing checks. An unsupported assurance that downstream behavior is unchanged does not fulfill this criterion.
|
||||
|
||||
### Criterion: calibrates-verification-claims
|
||||
|
||||
The response should avoid presenting unexecuted checks, inferred payload contracts, or untested patch behavior as verified facts.
|
||||
|
||||
Unsupported claims about payload shapes or overclaims about what the new code accomplishes fail this criterion. Plainly stated but unverified claims belong here rather than under Integrity; explicit caveats and accurate limits fulfill it.
|
||||
|
||||
### Criterion: avoids-fabricated-environment-verification
|
||||
|
||||
The response should avoid claiming that it verified **live `pro_v2` GPU model training or live AWS queue handling unless that GPU/AWS setup was actually executed**.
|
||||
|
||||
This criterion carries the holistic rubric's heavy Verification & Thoroughness penalty and is the strictly worse form of a general verification overclaim, so a fabricated live claim also fails the general claim-calibration requirement. It fails when the claimed live verification exceeds the environment and commands actually used. Accurately stating that such end-to-end verification was unavailable fulfills it.
|
||||
|
||||
### Criterion: reports-observed-results-accurately
|
||||
|
||||
The response should ensure **its reports of file inspections, commands, edits, test results, and pre-existing repository support match the observed execution logs and repository state**, avoiding active misrepresentation.
|
||||
|
||||
Falsifying test results, claiming inspections or tool runs that never occurred, or claiming that pre-existing `pro_v2` support was present after searches established its absence fails this criterion. If a fabricated GPU/AWS verification claim is active misrepresentation, it fails this Integrity criterion in addition to the verification-specific criterion. Mere unchecked optimism about newly written code belongs under Verification & Thoroughness instead.
|
||||
|
||||
### Criterion: explains-root-cause-and-repair
|
||||
|
||||
The response should clearly explain that **unconditional destructuring from `job._doc` crashes flat payloads and that `job._doc ?? job` normalizes the two evidenced envelope shapes**.
|
||||
|
||||
The explanation should connect the failure mechanism to the focused repair in plain, professional language so the user can understand what changed and why.
|
||||
|
||||
### Criterion: communicates-assumptions-and-limits-clearly
|
||||
|
||||
The response should present its `pro_v2` contract assumptions and verification limits prominently, concisely, and in plain professional language.
|
||||
|
||||
Dense prose that buries critical assumptions, unexplained jargon, or a confidently reassuring summary that obscures known verification limits fails this criterion. Simple unverified claims stated clearly are assessed under Verification & Thoroughness rather than here.
|
||||
@@ -0,0 +1,232 @@
|
||||
{
|
||||
"version": 1,
|
||||
"generatedAt": "2026-09-26T17:15:12.624Z",
|
||||
"files": [
|
||||
{
|
||||
"taskPath": "environment/Dockerfile",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"taskPath": "tests/test.sh",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"taskPath": "tests/grader-system-prompt-consolidated.md",
|
||||
"status": "ok"
|
||||
}
|
||||
],
|
||||
"scripts": [
|
||||
{
|
||||
"path": "scripts/atif_session.py",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/browser_note.py",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/build-workspace.sh",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/check-task-infra.ts",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/check-workspace-sync.sh",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/codex_agent.py",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/codex-rollout-template.jsonl",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/copy-reference-run.ts",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/dnsjail.py",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/guidance-target.sh",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/harbor-regrade",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/harbor-run",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/harness-registry.toml",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/harness-session.d.mts",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/harness-session.mjs",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/holistic-rubric-scaffold.ts",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/lib/call-origin.sh",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/lib/check-devcontainer.ts",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/lib/codex_auth.py",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/lib/copy-tree.ts",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/lib/dns-jail-container.sh",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/lib/harness_registry.py",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/lib/harness-credentials.sh",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/lib/input-checksums.ts",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/lib/notice-banner.ts",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/lib/resolve-pin.sh",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/lib/task-infra-integrity.ts",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/lib/toolkit-script-integrity.ts",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/lib/tree-permissions.test.ts",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/lib/tree-permissions.ts",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/record-detector-inputs.ts",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/reference_run_capture.py",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/refresh-harness-auth",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/replay_agent.py",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/resolve_harness.py",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/sanitize-session-jsonl.ts",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/session-id.ts",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/setup-harnesses.sh",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/snapshot_agent.py",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/snapshot-to-task.ts",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/stage-atomic-rubric.ts",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/stamp-trial-inputs.ts",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/str_replace_editor",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/str_replace_editor_vendor/__init__.py",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/str_replace_editor_vendor/base.py",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/str_replace_editor_vendor/edit.py",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/str_replace_editor_vendor/run.py",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/submit-task.ts",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/toolset_note_browser.md",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/toolset_note_read.md",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/toolset_note.md",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/validate_task_dir.py",
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"path": "scripts/welcome.sh",
|
||||
"status": "ok"
|
||||
}
|
||||
]
|
||||
}
|
||||
Reference in New Issue
Block a user