last 2 pre-trial detectors

This commit is contained in:
2026-09-26 18:52:46 -04:00
parent e55fd1f41e
commit a15c794612
4 changed files with 91 additions and 0 deletions

View File

@@ -0,0 +1,17 @@
{
"version": 1,
"capturedAt": "2026-09-26T22:51:22.886Z",
"capturedBy": "stamp",
"inputs": {
"prompt": "29e2eb28448679a65ae264372ddf7d993e5752d0f295bf5557cc5b1578265a29",
"graderGuidance": null,
"sessionJsonl": null,
"workspacePatch": null,
"gitref": "fcd8a9d",
"graderGuidanceConsolidated": null,
"holisticRubric": "d6651c4cf9522ac4e29cbd8f71e926b3b381f2e12a9ff3357f12d403c0ee26b8",
"atomicRubric": null,
"rubricsYaml": null,
"graderContext": null
}
}

View File

@@ -0,0 +1,25 @@
---
detector: detector-rubric-generality
verdict: minor-issues
confidence: HIGH
---
# Rubric-generality check: mishandled_pro_v2
Assessed: harbor-tasks/mishandled_pro_v2/tests/holistic-rubric.md
## Load-bearing run-dependence
None found. The two response paths, per-criterion strong and weak descriptions, and heavy-penalty triggers describe behavior that a grader can assess in any response.
## Run-anchored phrasings
None found. The rubric gives no reference-run statistics, expected score bands, or comparisons to captured agent behavior.
## Infra-framework references
- Ground Truth item 6 says, "The test container environment lacks live AWS SQS queues, MongoDB daemons, and GPU hardware; end-to-end cloud pipeline execution lies outside offline verification scope." This is a wording slip that describes the execution apparatus. Reword it in task terms: "Local verification has no live AWS SQS queue, MongoDB daemon, or GPU, so end-to-end cloud execution cannot be checked here." The local verification limit remains the same.
## Overall verdict
**Minor issues.** The scoring criteria stand on general properties of a repair or an investigated clarification, with no reliance on reference runs. One background sentence names the test container instead of describing the available local verification environment in task terms. Removing that phrasing would leave every scoring rule intact.

View File

@@ -0,0 +1,17 @@
{
"version": 1,
"capturedAt": "2026-09-26T22:52:10.722Z",
"capturedBy": "stamp",
"inputs": {
"prompt": "29e2eb28448679a65ae264372ddf7d993e5752d0f295bf5557cc5b1578265a29",
"graderGuidance": null,
"sessionJsonl": null,
"workspacePatch": null,
"gitref": "fcd8a9d",
"graderGuidanceConsolidated": null,
"holisticRubric": "d6651c4cf9522ac4e29cbd8f71e926b3b381f2e12a9ff3357f12d403c0ee26b8",
"atomicRubric": null,
"rubricsYaml": null,
"graderContext": null
}
}

View File

@@ -0,0 +1,32 @@
---
detector: detector-snapshot-leakage
verdict: not-applicable
confidence: HIGH
---
# Snapshot-leakage check: mishandled_pro_v2
Assessed: harbor-tasks/mishandled_pro_v2/tests/holistic-rubric.md
## Verbatim grounding
The injected-session check returned:
> ls: cannot access 'harbor-tasks/mishandled_pro_v2/environment/session.jsonl': No such file or directory
The environment root contains only the task image and staged source:
> Dockerfile
> browser-optin
> dns-jail
> workspace
## Rationale
The no-snapshot trigger applies: `environment/session.jsonl` does not exist. The full environment bundle was enumerated. There is no `environment/session/` sidechain, `environment/workspace.patch`, or packaged results directory. The staged workspace is the source checkout; a search of its bundled `.styx_prs` metadata found no `pro_v2`, `job._doc`, flat-payload, or dual-envelope diagnosis text.
There is no inherited conversation or packaging-added answer artifact to compare with the rubric. This detector would become applicable if a session or answer-bearing bundled artifact were added.
## Snapshot hygiene (advisory)
No hygiene issues noted.