detectors

This commit is contained in:
2026-10-09 15:56:27 -04:00
parent 9cb80f557b
commit f9bec5de44
35 changed files with 760 additions and 1 deletions

View File

@@ -0,0 +1,17 @@
{
"version": 1,
"capturedAt": "2026-10-09T19:51:00.281Z",
"capturedBy": "stamp",
"inputs": {
"prompt": "75042109a7aab36d9a50fe23f5ac417488f437efb25575a987c4fe35d8103b16",
"graderGuidance": null,
"sessionJsonl": null,
"workspacePatch": null,
"gitref": "fcd8a9d",
"graderGuidanceConsolidated": null,
"holisticRubric": "1741564f653207e011d449854de8aa9d569a8c560c01ad74cff6bda42e6438dc",
"atomicRubric": null,
"rubricsYaml": null,
"graderContext": null
}
}

View File

@@ -0,0 +1,57 @@
---
detector: detector-answer-obviousness
verdict: partial
confidence: HIGH
---
Assessed: harbor-tasks/potion-voice-user-ownership/tests/holistic-rubric.md
# Answer-obviousness check: potion-voice-user-ownership
## What the prompt asks
The prompt suspects the workers may lack tenant isolation and asks to “Audit all db queries across both worker handlers and make them enforce strict multi-tenant authorization so users cannot access or modify records belonging to other tenants.” A thoughtful colleague would inspect both workers and the service methods they call, then make ownership-scoped reads and writes. The prompt does not request a general SQS lifecycle repair or prescribe a source for `recordingId`.
## Per-expectation assessment
### Tenant-scoped reads and writes — obvious
- **What the rubric requires:** “All primary and secondary MongoDB queries and updates enforce `userId` scoping, preventing cross-tenant access.”
- **Is it obvious from the prompt?** Yes. Tenant ownership for queries in both handlers is the express task. Fixing wrapper methods those handlers rely on is an ordinary way to make the request effective.
- **Verdict for this expectation:** `obvious`.
### Safe update signatures and runnable worker code — obvious
- **What the rubric requires:** “The worker handlers run cleanly without runtime exceptions, syntax errors, or unhandled promise rejections” and “Mongoose `findOneAndUpdate` accepts 3 arguments: `findOneAndUpdate(conditions, update, options)`.”
- **Is it obvious from the prompt?** Yes as an implementation duty: a tenant filter must be placed in the query argument, and new code must run. The exact imagined five-argument mistake is an illustrative failure, not a required implementation shape.
- **Verdict for this expectation:** `obvious`.
### Exact `recordingId` derivation — not-obvious
- **What the rubric requires:** “`recordingId` must be extracted from `salutationToUpdate.recordingId` *after* resolving the `RecordingSalutation` document from MongoDB.”
- **Is it obvious from the prompt?** No. This is overstated universality: the prompt asks for owner-scoped access, not a specific lookup chain. An agent could reasonably retain a supplied recording ID and verify both it and its relation to the salutation under the same user. The rubric's claim about the available payload is a separate fact-check issue.
- **Verdict for this expectation:** `not-obvious`.
### Completion before SQS deletion — not-obvious
- **What the rubric requires:** “SQS messages are deleted only after successful task execution and artifact upload.”
- **Is it obvious from the prompt?** No. This is unrequested scope: queue deletion already occurs early in the checked-in handlers, and a focused tenant-isolation change can leave that existing behavior untouched. The rubric's “FAIL (0.0)” for any premature deletion could reject such a competent, scoped solution.
- **Verdict for this expectation:** `not-obvious`.
### Error-state updates on authorization rejection — not-obvious
- **What the rubric requires:** “On error or authorization rejection, the worker must update MongoDB job/profile statuses to `'error'` regardless of pre-authorization state flags.”
- **Is it obvious from the prompt?** Partly, but the exact unconditional requirement is overstated. A worker should handle rejection coherently; it should not mutate a foreign tenant's record solely to mark it `error`. A scoped status write that matches no foreign record is a defensible response to an unauthorized payload.
- **Verdict for this expectation:** `not-obvious`.
### Verify queue-schema assumptions — obvious
- **What the rubric requires:** “The worker pipeline behavior and database error transitions are verified against expected queue message schemas.”
- **Is it obvious from the prompt?** Yes in direction: the fix depends on where the job's user ID comes from, so inspecting actual worker inputs and exercising affected paths is ordinary diligence. The absence of a producer or sample message may limit what can be proven; that is handled by fact-check and offline-verifiability.
- **Verdict for this expectation:** `obvious`.
## Overall verdict
The central tenant-isolation request is clear and fairly cued. The rubric also requires a specific recording-ID route, a queue-lifecycle repair, and an unconditional error-state policy that the prompt does not establish as the only acceptable choices. These secondary requirements would mark down plausible focused implementations.
The verdict is `partial`, rather than `not-obvious`, because a real and substantial fair test remains: whether the agent scopes the relevant database operations to the job's user. No reference runs exist to cross-check the alternative paths.

View File

@@ -0,0 +1,17 @@
{
"version": 1,
"capturedAt": "2026-10-09T19:50:53.122Z",
"capturedBy": "stamp",
"inputs": {
"prompt": "75042109a7aab36d9a50fe23f5ac417488f437efb25575a987c4fe35d8103b16",
"graderGuidance": null,
"sessionJsonl": null,
"workspacePatch": null,
"gitref": "fcd8a9d",
"graderGuidanceConsolidated": null,
"holisticRubric": "1741564f653207e011d449854de8aa9d569a8c560c01ad74cff6bda42e6438dc",
"atomicRubric": null,
"rubricsYaml": null,
"graderContext": null
}
}

View File

@@ -0,0 +1,21 @@
---
detector: detector-broken-dev-env
verdict: not-applicable
confidence: MEDIUM
---
Assessed: harbor-tasks/potion-voice-user-ownership/tests/holistic-rubric.md and current instruction.md.
# Broken-dev-env check: potion-voice-user-ownership
## Verbatim grounding
> Audit all db queries across both worker handlers and make them enforce strict multi-tenant authorization so users cannot access or modify records belonging to other tenants.
`reference-runs/` is empty. The rubric discusses possible runtime mistakes in agent changes, but contains no observation that the shipped environment failed to build, install, or run.
## Rationale
The no-evidence trigger applies. The prompt is a substantive code-change request, yet there is no trial trajectory showing installation trouble, pre-existing unrelated test failure, or an agent workaround. The package also has no scored runs whose prompt, grade, or output could be checked for corruption or revision drift.
The rubric's hypothetical `MODULE_NOT_FOUND` and other crashes are intended agent-introduced defects, not evidence of incidental baseline breakage. Re-run once reference runs provide an environment signal. No prompt premise mismatch is established by the current prompt and workspace inspection.

View File

@@ -0,0 +1,17 @@
{
"version": 1,
"capturedAt": "2026-10-09T19:50:53.122Z",
"capturedBy": "stamp",
"inputs": {
"prompt": "75042109a7aab36d9a50fe23f5ac417488f437efb25575a987c4fe35d8103b16",
"graderGuidance": null,
"sessionJsonl": null,
"workspacePatch": null,
"gitref": "fcd8a9d",
"graderGuidanceConsolidated": null,
"holisticRubric": "1741564f653207e011d449854de8aa9d569a8c560c01ad74cff6bda42e6438dc",
"atomicRubric": null,
"rubricsYaml": null,
"graderContext": null
}
}

View File

@@ -0,0 +1,17 @@
---
detector: detector-credential-leakage
verdict: clean
confidence: HIGH
---
# Credential-leakage check: potion-voice-user-ownership
Assessed: instruction.md, task.toml, tests/holistic-rubric.md, and environment/workspace.patch presence.
## Findings
No `environment/workspace.patch` exists. The strongest near-miss is the `task.toml` placeholder assignment `ANTHROPIC_API_KEY = "${ANTHROPIC_API_KEY}"`; it refers to an environment variable and contains no key value. The checked authored surfaces show no secret-shaped added lines or personal checkout path.
## Overall verdict
No credential or internal checkout path is shipped in the authored task surfaces currently present. Re-run if a workspace patch is added.

View File

@@ -0,0 +1,17 @@
{
"version": 1,
"capturedAt": "2026-10-09T19:50:53.122Z",
"capturedBy": "stamp",
"inputs": {
"prompt": "75042109a7aab36d9a50fe23f5ac417488f437efb25575a987c4fe35d8103b16",
"graderGuidance": null,
"sessionJsonl": null,
"workspacePatch": null,
"gitref": "fcd8a9d",
"graderGuidanceConsolidated": null,
"holisticRubric": "1741564f653207e011d449854de8aa9d569a8c560c01ad74cff6bda42e6438dc",
"atomicRubric": null,
"rubricsYaml": null,
"graderContext": null
}
}

View File

@@ -0,0 +1,19 @@
---
detector: detector-cross-task-reference
verdict: clean
confidence: HIGH
---
# Cross-task-reference check: potion-voice-user-ownership
Assessed: harbor-tasks/potion-voice-user-ownership/tests/holistic-rubric.md and instruction.md.
## Verbatim grounding
> The goal is to audit and refactor background SQS worker handlers (`voice-synthsizer-job-handler` and `voice-cloning-job-handler`) and database service wrappers in `potion-voice`
This names the source repository and its worker folders. It does not point to a separate graded task.
## Rationale
The current tenant-isolation prompt and rubric contain no comparison to, title of, or unresolved pointer into another task. The rubric has other problems, but cross-task dependence is not one of them.

View File

@@ -0,0 +1,17 @@
{
"version": 1,
"capturedAt": "2026-10-09T19:50:53.122Z",
"capturedBy": "stamp",
"inputs": {
"prompt": "75042109a7aab36d9a50fe23f5ac417488f437efb25575a987c4fe35d8103b16",
"graderGuidance": null,
"sessionJsonl": null,
"workspacePatch": null,
"gitref": "fcd8a9d",
"graderGuidanceConsolidated": null,
"holisticRubric": "1741564f653207e011d449854de8aa9d569a8c560c01ad74cff6bda42e6438dc",
"atomicRubric": null,
"rubricsYaml": null,
"graderContext": null
}
}

View File

@@ -0,0 +1,24 @@
---
detector: detector-dimension-misapplication
verdict: clean
confidence: HIGH
---
# Dimension-misapplication check: potion-voice-user-ownership
Assessed: harbor-tasks/potion-voice-user-ownership/tests/holistic-rubric.md.
## Verbatim grounding
> #### Narrow Correctness
> - **FAIL**: Code throws `MODULE_NOT_FOUND`, `ReferenceError: recordingId is not defined`, or Mongoose `CastError` exceptions during execution.
> #### Broader Correctness
> - **FAIL**: Queries rely solely on `_id` without `userId` validation, or SQS messages are deleted prematurely before downstream processing completes.
> #### Verification & Thoroughness
> - **FAIL**: Claiming complete multi-tenant scoping and background pipeline health without verifying worker execution against SQS message structures.
## Rationale
The rubric places execution failures under Narrow Correctness, security and durability under Broader Correctness, and an unsupported verification claim under Verification & Thoroughness. The Common Sense examples concern poor engineering judgment in ordering dependent operations and moving deletion to pipeline entry. These are plausible criterion bindings under the shared standard. No reference-run grades exist to expose score drift. The rubric's missing criteria and binary score guide are separate rubric-form concerns, not a demonstrated dimension misapplication.

View File

@@ -0,0 +1,17 @@
{
"version": 1,
"capturedAt": "2026-10-09T19:50:53.122Z",
"capturedBy": "stamp",
"inputs": {
"prompt": "75042109a7aab36d9a50fe23f5ac417488f437efb25575a987c4fe35d8103b16",
"graderGuidance": null,
"sessionJsonl": null,
"workspacePatch": null,
"gitref": "fcd8a9d",
"graderGuidanceConsolidated": null,
"holisticRubric": "1741564f653207e011d449854de8aa9d569a8c560c01ad74cff6bda42e6438dc",
"atomicRubric": null,
"rubricsYaml": null,
"graderContext": null
}
}

View File

@@ -0,0 +1,62 @@
---
detector: detector-fact-check-rubric-claims
verdict: fail
confidence: MEDIUM
claims:
- id: c01
verdict: unclear
loadBearing: true
summary: "Voice synthesis SQS payload omits recordingId"
rubricQuote: "SQS messages for voice synthesis contain `job.salutationId`, `job.userAudioProfileId`, and `job.userId`, but do **not** convey `job.recordingId`."
sourceEvidence: " recordingId,"
sourceProvenance: "harbor-tasks/potion-voice-user-ownership/environment/workspace/voice-synthsizer-job-handler/index.js (lines 74-82)"
note: "Source unavailable for the actual producer or a sample queue payload: the checked-in consumer explicitly reads recordingId from job at line 79. A repo-wide search found no producer that establishes the rubric's asserted omission; this central payload claim cannot be confirmed from the shipped workspace."
- id: c02
verdict: pass
loadBearing: true
summary: "RecordingSalutation schema has recordingId link"
rubricQuote: "`RecordingSalutation` records link a `salutationId` to a parent `recordingId`."
sourceEvidence: " recordingId: {"
sourceProvenance: "harbor-tasks/potion-voice-user-ownership/environment/workspace/voice-synthsizer-job-handler/recording_salutation/recording_salutation_model.js (lines 16-19)"
note: "The schema has a recordingId ObjectId reference. The worker uses the salutationId as the _id of this record, which makes the lookup path discoverable from index.js lines 152-155."
- id: c03
verdict: fail
loadBearing: true
summary: "recordingId must always come from resolved salutation"
rubricQuote: "`recordingId` must be extracted from `salutationToUpdate.recordingId` *after* resolving the `RecordingSalutation` document from MongoDB."
sourceEvidence: " required: false"
sourceProvenance: "harbor-tasks/potion-voice-user-ownership/environment/workspace/voice-synthsizer-job-handler/recording_salutation/recording_salutation_model.js (lines 16-19); voice-synthsizer-job-handler/index.js (lines 74-82, 152-160)"
note: "The asserted source field is optional in the schema, while the current consumer expects recordingId in the queue job. The rubric states a mandatory extraction path that the shipped schema cannot guarantee. The agent can discover the optional field and consumer read in these workspace files."
- id: c04
verdict: pass
loadBearing: true
summary: "SQS deletion currently precedes downstream work"
rubricQuote: "Deleting messages via `sqs.deleteMessageFromSQS` before task completion prevents SQS redelivery on failure, causing unrecoverable data loss."
sourceEvidence: " await sqs.deleteMessageFromSQS(sqsQueueUrl, receiptHandle)"
sourceProvenance: "harbor-tasks/potion-voice-user-ownership/environment/workspace/voice-synthsizer-job-handler/index.js (lines 71-72, 113-139); voice-cloning-job-handler/index.js (lines 129-130, 174-276)"
note: "Both handlers call delete before their Python work and uploads. The redelivery implication follows SQS delete semantics; the order is directly reachable from the worker code."
- id: c05
verdict: partial
loadBearing: true
summary: "Malformed Mongoose call ignores ownership filters"
rubricQuote: "Passing 5 arguments or placing query filters in the `options` argument bypasses user scoping and causes updates to be ignored."
sourceEvidence: " const updatedJob = await Job.findOneAndUpdate({ _id: job._id }, job, {"
sourceProvenance: "harbor-tasks/potion-voice-user-ownership/environment/workspace/voice-synthsizer-job-handler/job/job_service.js (lines 66-70)"
note: "The workspace demonstrates the standard conditions/update/options arrangement. Putting tenant conditions into options would leave the first-argument query unscoped, but this does not by itself imply that the second-argument update is ignored. No five-argument example or local Mongoose API implementation ships for direct confirmation."
- id: c06
verdict: pass
loadBearing: true
summary: "Cloning catch writes error statuses"
rubricQuote: "On error or authorization rejection, the worker must update MongoDB job/profile statuses to `'error'` regardless of pre-authorization state flags."
sourceEvidence: " await voiceCloningService.update({ _id, status: 'error' })"
sourceProvenance: "harbor-tasks/potion-voice-user-ownership/environment/workspace/voice-cloning-job-handler/index.js (lines 278-292)"
note: "The current cloning catch writes error status to cloning and profile records without an authorization flag, supporting the intended recovery pattern. This is a proposed behavior for future authorization rejection, not evidence that such a rejection path currently exists."
---
# Fact-check rubric claims: potion-voice-user-ownership
Assessed: harbor-tasks/potion-voice-user-ownership/tests/holistic-rubric.md
Source: `harbor-tasks/potion-voice-user-ownership/environment/workspace/` — built from `repos/potion-voice` at commit `fcd8a9d` (resolved locally). No `environment/workspace.patch` exists.
Checked 6 claims (all load-bearing; one fail, one partial, one unclear). The central claim that `salutationToUpdate.recordingId` is mandatory fails because the schema marks that field optional. The exact SQS producer payload is unavailable in this repository; the consumer explicitly destructures `recordingId` from `job`.

View File

@@ -0,0 +1,17 @@
{
"version": 1,
"capturedAt": "2026-10-09T19:50:53.122Z",
"capturedBy": "stamp",
"inputs": {
"prompt": "75042109a7aab36d9a50fe23f5ac417488f437efb25575a987c4fe35d8103b16",
"graderGuidance": null,
"sessionJsonl": null,
"workspacePatch": null,
"gitref": "fcd8a9d",
"graderGuidanceConsolidated": null,
"holisticRubric": "1741564f653207e011d449854de8aa9d569a8c560c01ad74cff6bda42e6438dc",
"atomicRubric": null,
"rubricsYaml": null,
"graderContext": null
}
}

View File

@@ -0,0 +1,21 @@
---
detector: detector-good-response-defined
verdict: defines-good
confidence: HIGH
---
# Good-response-defined check: potion-voice-user-ownership
Assessed: harbor-tasks/potion-voice-user-ownership/tests/holistic-rubric.md.
## Positive target present?
The rubric explicitly states: “All primary and secondary MongoDB queries and updates enforce `userId` scoping, preventing cross-tenant access. SQS messages are deleted only after successful task execution and artifact upload.” Its scoring guide also describes preserving dependency order, error transitions, and execution without runtime exceptions.
## What the grader has to infer
The grader would need to reconcile those targets with several questionable ground-truth claims and the narrower tenant-isolation request in the current prompt. The positive target itself is present; the rubric is not a problems-only catalog.
## Overall verdict
A concrete implementation success target is defined for the load-bearing work. This verdict does not certify that the target is fairly requested or technically correct.

View File

@@ -0,0 +1,17 @@
{
"version": 1,
"capturedAt": "2026-10-09T19:50:53.122Z",
"capturedBy": "stamp",
"inputs": {
"prompt": "75042109a7aab36d9a50fe23f5ac417488f437efb25575a987c4fe35d8103b16",
"graderGuidance": null,
"sessionJsonl": null,
"workspacePatch": null,
"gitref": "fcd8a9d",
"graderGuidanceConsolidated": null,
"holisticRubric": "1741564f653207e011d449854de8aa9d569a8c560c01ad74cff6bda42e6438dc",
"atomicRubric": null,
"rubricsYaml": null,
"graderContext": null
}
}

View File

@@ -0,0 +1,21 @@
---
detector: detector-good-response-exhaustiveness
verdict: has-gaps
confidence: HIGH
---
Assessed: harbor-tasks/potion-voice-user-ownership/tests/holistic-rubric.md
# Good-response-exhaustiveness check: potion-voice-user-ownership
## Plausible strong-response approaches
The prompt calls for an implementation: audit both worker handlers and enforce ownership on their MongoDB access. A broad majority of engineers would accept a focused solution that scopes each read and update, rejects mismatched records, and reports what was checked. A deeper refactor that also corrects SQS acknowledgment timing could also be sound, but the prompt does not require that extra repair. If the message lacks a trustworthy user ID in a particular path, flagging that prerequisite and safely refusing the operation is another legitimate part of the implementation.
## Coverage in the rubric
The rubric credits full tenant scoping: “All primary and secondary MongoDB queries and updates enforce `userId` scoping.” It does not credit the focused solution if it preserves the existing early SQS deletion: its pass tier also requires “SQS messages are retained until full pipeline completion” and its fail tier says “SQS messages are deleted prematurely.” A solution that validates an existing recording ID under the job's owner, rather than deriving it from the optional salutation field, is likewise excluded by the ground-truth prescription that `recordingId` “must be extracted from `salutationToUpdate.recordingId`.” There are no reference runs to support a run-specific penalty-side finding.
## Overall verdict
The rubric covers the central security fix but excludes a major reasonable approach: implementing the requested ownership controls without changing an unrelated, pre-existing queue acknowledgment policy. Widen the strong tier to credit that focused solution, while still penalizing new queue regressions, and allow owner-validated recording relationships that satisfy the same isolation goal.

View File

@@ -0,0 +1,17 @@
{
"version": 1,
"capturedAt": "2026-10-09T19:50:53.122Z",
"capturedBy": "stamp",
"inputs": {
"prompt": "75042109a7aab36d9a50fe23f5ac417488f437efb25575a987c4fe35d8103b16",
"graderGuidance": null,
"sessionJsonl": null,
"workspacePatch": null,
"gitref": "fcd8a9d",
"graderGuidanceConsolidated": null,
"holisticRubric": "1741564f653207e011d449854de8aa9d569a8c560c01ad74cff6bda42e6438dc",
"atomicRubric": null,
"rubricsYaml": null,
"graderContext": null
}
}

View File

@@ -0,0 +1,33 @@
---
detector: detector-meaningful-failure
verdict: not-applicable
confidence: HIGH
---
# Meaningful-failure check: potion-voice-user-ownership
Assessed: harbor-tasks/potion-voice-user-ownership/tests/holistic-rubric.md and reference-runs/.
## Load-bearing targets
- Missing `worker_tenant` utility causing startup failure.
- Deleting SQS messages before synthesis, training, and uploads complete.
- Malformed five-argument `findOneAndUpdate` calls that leave filters unscoped.
- Reading `recordingId` before resolving the related salutation.
- Skipping error-state updates after authorization rejection.
## Elicitation matrix
No reference runs exist, so none of the targets has an observable fire count or `grade.md` evidence.
## Per-deduction assessment
No scored deductions exist to assess.
## Guidance-wide severity audit
The asserted harms cannot be calibrated against observed responses until reference runs exist. No severity conclusion is made here.
## Overall verdict
The no-reference-runs trigger applies. Run trials against the current prompt, then re-run this detector to assess whether the stated failures actually occur and whether their consequences are proportionate.

View File

@@ -0,0 +1,17 @@
{
"version": 1,
"capturedAt": "2026-10-09T19:50:53.122Z",
"capturedBy": "stamp",
"inputs": {
"prompt": "75042109a7aab36d9a50fe23f5ac417488f437efb25575a987c4fe35d8103b16",
"graderGuidance": null,
"sessionJsonl": null,
"workspacePatch": null,
"gitref": "fcd8a9d",
"graderGuidanceConsolidated": null,
"holisticRubric": "1741564f653207e011d449854de8aa9d569a8c560c01ad74cff6bda42e6438dc",
"atomicRubric": null,
"rubricsYaml": null,
"graderContext": null
}
}

View File

@@ -0,0 +1,25 @@
---
detector: detector-offline-verifiability
verdict: partial
confidence: MEDIUM
---
Assessed: harbor-tasks/potion-voice-user-ownership/tests/holistic-rubric.md
# Offline-verifiability check: potion-voice-user-ownership
## Findings
### Queue-schema verification — evidence outside shipped repo (partial)
- **Where:** `tests/holistic-rubric.md`, Verification & Thoroughness.
- **Quote:**
> The worker pipeline behavior and database error transitions are verified against expected queue message schemas.
- **Why it lives outside:** The workspace includes the consumers but no SQS producer, captured queue message, local queue fake, or local MongoDB integration test. Source inspection and isolated model-call tests can check ownership filters and error paths, but cannot establish the claimed actual message schema or end-to-end queue behavior from the shipped material.
- **Something to consider:** Include a representative local message fixture and faithful queue/database fake, or grade the agent's explicitly bounded local verification of ownership queries.
## Overall verdict
The tenant-isolation implementation is mostly offline-completable: the Node dependencies and worker source are shipped, and owner-scoped query construction can be checked locally. The rubric's queue-schema verification expectation relies on a payload contract not established by the repository; the worker even destructures `recordingId` from the message while the rubric says it is absent. This is a secondary verifiability gap, so the verdict is `partial`, not a claim that the whole task requires live services.

View File

@@ -0,0 +1,17 @@
{
"version": 1,
"capturedAt": "2026-10-09T19:50:53.122Z",
"capturedBy": "stamp",
"inputs": {
"prompt": "75042109a7aab36d9a50fe23f5ac417488f437efb25575a987c4fe35d8103b16",
"graderGuidance": null,
"sessionJsonl": null,
"workspacePatch": null,
"gitref": "fcd8a9d",
"graderGuidanceConsolidated": null,
"holisticRubric": "1741564f653207e011d449854de8aa9d569a8c560c01ad74cff6bda42e6438dc",
"atomicRubric": null,
"rubricsYaml": null,
"graderContext": null
}
}

View File

@@ -0,0 +1,17 @@
---
detector: detector-over-hinting
verdict: clean
confidence: HIGH
---
Assessed: harbor-tasks/potion-voice-user-ownership/tests/holistic-rubric.md
# Over-hinting check: potion-voice-user-ownership
## Findings
The closest candidate is the prompt's statement: “job handlers fetch database models using only document IDs supplied in SQS payloads without verifying that they belong to the job's userId.” It identifies the user-visible suspected defect and defines the requested security outcome. It does not reveal the rubric's less-obvious recording relationship, Mongoose signature failure, queue deletion concern, or error-state issue. There is no `environment/workspace.patch` carrying authored comments to inspect.
## Overall verdict
The prompt's specificity is a genuine requirement and plausible incident description, not a giveaway of an independent discovery task. The agent still has to find and correct the relevant access paths. No authored workspace hint surface exists.

View File

@@ -0,0 +1,17 @@
{
"version": 1,
"capturedAt": "2026-10-09T19:50:53.122Z",
"capturedBy": "stamp",
"inputs": {
"prompt": "75042109a7aab36d9a50fe23f5ac417488f437efb25575a987c4fe35d8103b16",
"graderGuidance": null,
"sessionJsonl": null,
"workspacePatch": null,
"gitref": "fcd8a9d",
"graderGuidanceConsolidated": null,
"holisticRubric": "1741564f653207e011d449854de8aa9d569a8c560c01ad74cff6bda42e6438dc",
"atomicRubric": null,
"rubricsYaml": null,
"graderContext": null
}
}

View File

@@ -0,0 +1,34 @@
---
detector: detector-rubric-clarity
verdict: material-issues
confidence: HIGH
---
# Rubric-clarity check: potion-voice-user-ownership
Assessed: harbor-tasks/potion-voice-user-ownership/tests/holistic-rubric.md.
## Material ambiguities
### Binary guide conflicts with criterion grading
- **Where:** Scoring Guide: “PASS (1.0)” for every property together and “FAIL (0.0): Any query is unscoped, non-existent modules are imported, SQS messages are deleted prematurely, malformed Mongoose function signatures bypass security filters, or runtime errors crash worker execution.”
- **Why it's ambiguous:** One grader could assign 0.0 overall for a single remaining unscoped query; another could score the eight standard criteria independently and penalize the affected correctness dimensions. These produce materially different scores. The guide does not say whether its binary numbers override the standard.
### Scope of every query is underspecified
- **Where:** Ground Truth: “Primary and secondary model operations (`UserAudioProfile`, `VoiceCloning`, `Salutation`, `Job`, `Recording`) must be scoped with `{ _id, userId, deleted: false }`.”
- **Why it's ambiguous:** The requirement could include inserts, list queries, status updates, and operations that do not have an `_id` input; or only lookups and updates by ID. The grading consequence of a missed service wrapper depends on which reading is used.
### Error transition on rejected ownership
- **Where:** Ground Truth: “On error or authorization rejection, the worker must update MongoDB job/profile statuses to `'error'` regardless of pre-authorization state flags.”
- **Why it's ambiguous:** It does not say which tenant's records can be changed when the ID belongs to another user. A grader could require a status write even on a foreign record, conflicting with the ownership rule, or only a scoped write that may match nothing.
## Copy-edit issues
None found that materially interrupts reading.
## Overall verdict
The binary scoring guide and authorization-rejection status rule can lead reasonable graders to different outcomes. Clarify those before using the rubric to score trials.

View File

@@ -0,0 +1,17 @@
{
"version": 1,
"capturedAt": "2026-10-09T19:50:53.122Z",
"capturedBy": "stamp",
"inputs": {
"prompt": "75042109a7aab36d9a50fe23f5ac417488f437efb25575a987c4fe35d8103b16",
"graderGuidance": null,
"sessionJsonl": null,
"workspacePatch": null,
"gitref": "fcd8a9d",
"graderGuidanceConsolidated": null,
"holisticRubric": "1741564f653207e011d449854de8aa9d569a8c560c01ad74cff6bda42e6438dc",
"atomicRubric": null,
"rubricsYaml": null,
"graderContext": null
}
}

View File

@@ -0,0 +1,33 @@
---
detector: detector-rubric-coverage
verdict: not-applicable
confidence: HIGH
---
# Rubric-coverage check: potion-voice-user-ownership
Assessed: harbor-tasks/potion-voice-user-ownership/tests/holistic-rubric.md against absent tests/atomic-rubric.yaml and tests/grader-context.md.
## Coverage map
Not assessed: no atomic criteria exist to map to the holistic clauses.
## Coverage gaps
Not assessed for the same reason.
## Invented content
None can be assessed without atomic criteria.
## Context integrity
`tests/grader-context.md` is absent, so context migration has not yet occurred.
## Crux alignment
No atomic crux criteria exist. The holistic rubric also has no explicit heavy-penalty section to map.
## Overall verdict
The no-atomic-rubric trigger applies. Write the atomic package after finalizing the holistic rubric, then re-run this detector.

View File

@@ -0,0 +1,17 @@
{
"version": 1,
"capturedAt": "2026-10-09T19:50:53.122Z",
"capturedBy": "stamp",
"inputs": {
"prompt": "75042109a7aab36d9a50fe23f5ac417488f437efb25575a987c4fe35d8103b16",
"graderGuidance": null,
"sessionJsonl": null,
"workspacePatch": null,
"gitref": "fcd8a9d",
"graderGuidanceConsolidated": null,
"holisticRubric": "1741564f653207e011d449854de8aa9d569a8c560c01ad74cff6bda42e6438dc",
"atomicRubric": null,
"rubricsYaml": null,
"graderContext": null
}
}

View File

@@ -0,0 +1,29 @@
---
detector: detector-rubric-form
verdict: not-applicable
confidence: HIGH
---
# Rubric-form check: potion-voice-user-ownership
Assessed: absent tests/atomic-rubric.yaml.
## Deterministic contract
Not run: there is no atomic rubric file to parse or inspect.
## Atomicity and self-containment
Not assessed.
## Phrasing and answer keys
Not assessed.
## Fair-grading findings
Not assessed.
## Overall verdict
The atomic rubric has not been written. Re-run after creating `tests/atomic-rubric.yaml` and `tests/grader-context.md`.

View File

@@ -0,0 +1,17 @@
{
"version": 1,
"capturedAt": "2026-10-09T19:50:53.122Z",
"capturedBy": "stamp",
"inputs": {
"prompt": "75042109a7aab36d9a50fe23f5ac417488f437efb25575a987c4fe35d8103b16",
"graderGuidance": null,
"sessionJsonl": null,
"workspacePatch": null,
"gitref": "fcd8a9d",
"graderGuidanceConsolidated": null,
"holisticRubric": "1741564f653207e011d449854de8aa9d569a8c560c01ad74cff6bda42e6438dc",
"atomicRubric": null,
"rubricsYaml": null,
"graderContext": null
}
}

View File

@@ -0,0 +1,25 @@
---
detector: detector-rubric-generality
verdict: generalizes
confidence: HIGH
---
# Rubric-generality check: potion-voice-user-ownership
Assessed: harbor-tasks/potion-voice-user-ownership/tests/holistic-rubric.md.
## Load-bearing run-dependence
None found. The rubric describes possible response properties and failure modes without tying a score to named reference runs.
## Run-anchored phrasings
None found. “The agent adds import statements” is an illustrative failure mode, not a claim about an observed run.
## Infra-framework references
None found. SQS and MongoDB are application dependencies, not grading infrastructure.
## Overall verdict
The rubric can be applied to a new agent without knowing any particular trial's behavior. This does not resolve its factual or scope issues.

View File

@@ -0,0 +1,17 @@
{
"version": 1,
"capturedAt": "2026-10-09T19:50:53.122Z",
"capturedBy": "stamp",
"inputs": {
"prompt": "75042109a7aab36d9a50fe23f5ac417488f437efb25575a987c4fe35d8103b16",
"graderGuidance": null,
"sessionJsonl": null,
"workspacePatch": null,
"gitref": "fcd8a9d",
"graderGuidanceConsolidated": null,
"holisticRubric": "1741564f653207e011d449854de8aa9d569a8c560c01ad74cff6bda42e6438dc",
"atomicRubric": null,
"rubricsYaml": null,
"graderContext": null
}
}

View File

@@ -0,0 +1,11 @@
---
detector: detector-run-behaviors
verdict: not-applicable
confidence: HIGH
---
# Run-behaviors check: potion-voice-user-ownership
Assessed: harbor-tasks/potion-voice-user-ownership/reference-runs/.
`reference-runs/` is empty. A behavior-by-run matrix needs at least two captured runs. Run trials and copy their reference runs, then re-run this detector.

View File

@@ -0,0 +1,17 @@
{
"version": 1,
"capturedAt": "2026-10-09T19:50:53.122Z",
"capturedBy": "stamp",
"inputs": {
"prompt": "75042109a7aab36d9a50fe23f5ac417488f437efb25575a987c4fe35d8103b16",
"graderGuidance": null,
"sessionJsonl": null,
"workspacePatch": null,
"gitref": "fcd8a9d",
"graderGuidanceConsolidated": null,
"holisticRubric": "1741564f653207e011d449854de8aa9d569a8c560c01ad74cff6bda42e6438dc",
"atomicRubric": null,
"rubricsYaml": null,
"graderContext": null
}
}

View File

@@ -0,0 +1,21 @@
---
detector: detector-snapshot-leakage
verdict: not-applicable
confidence: HIGH
---
# Snapshot-leakage check: potion-voice-user-ownership
Assessed: environment/session.jsonl presence, sidechains, bundled artifacts, and harbor-tasks/potion-voice-user-ownership/tests/holistic-rubric.md.
## Verbatim grounding
`environment/session.jsonl` and `environment/session/` do not exist. `environment/workspace.patch` is also absent. The task is a manual scaffold, with no injected snapshot conversation to compare with the rubric.
## Rationale
The no-snapshot trigger applies. The built workspace was also checked for a bundled answer artifact; no added patch or results directory exists. This says nothing about the rubric's fairness as a standalone task.
## Snapshot hygiene (advisory)
No hygiene issues noted. The prompt is substantive; this manual task has no injected snapshot session.

View File

@@ -1 +1 @@
# Replace this with your task prompt
Its possible our background workers (voice-synthsizer-job-handler and voice-cloning-job-handler) do not enforce strict tenant isolation when loading and updating MongoDB records. Specifically, job handlers fetch database models using only document IDs supplied in SQS payloads without verifying that they belong to the job's userId. Audit all db queries across both worker handlers and make them enforce strict multi-tenant authorization so users cannot access or modify records belonging to other tenants.