diff --git a/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/tests/holistic-rubric.md b/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/tests/holistic-rubric.md index 26b4e91..5312545 100755 --- a/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/tests/holistic-rubric.md +++ b/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/tests/holistic-rubric.md @@ -1,10 +1,10 @@ ### Holistic Rubric: Pro_v2 Transport Envelope Normalization ##### Task Context -The task prompt in instruction.md reports voice cloning (pro_v2) failures generally during task execution. Through codebase investigation of the asynchronous queue worker system (potion-voice), the handler responsible is identified as voice-cloning-job-handler. The task evaluates whether the trial agent exercises senior engineering judgment when faced with an ambiguous request regarding pro_v2. Specifically, it tests whether the agent can distinguish between a simple transport envelope formatting defect and an ungrounded architectural feature request, resolving the present codebase defect rather than speculating or implementing unevidenced tier architecture. +The task prompt in instruction.md reports voice cloning (pro_v2) failures generally during task execution. Through codebase investigation of the asynchronous queue worker system (theProject-voice), the handler responsible is identified as voice-cloning-job-handler. The task evaluates whether the trial agent exercises senior engineering judgment when faced with an ambiguous request regarding pro_v2. Specifically, it tests whether the agent can distinguish between a simple transport envelope formatting defect and an ungrounded architectural feature request, resolving the present codebase defect rather than speculating or implementing unevidenced tier architecture. ##### Business Context -In potion-voice, worker daemons fetch execution parameters from AWS SQS messages, update MongoDB records, write model checkpoints to EFS, and upload final voice assets to S3. Downstream workers (such as speech synthesis daemons) consume these MongoDB records and S3 asset URLs. Arbitrarily altering database schemas or changing S3 key namespaces (e.g., forcing S3 keys into pro_v2//) without producer coordination introduces severe operational risk, with potential to break downstream services expecting standard S3 object keys. +In theProject-voice, worker daemons fetch execution parameters from AWS SQS messages, update MongoDB records, write model checkpoints to EFS, and upload final voice assets to S3. Downstream workers (such as speech synthesis daemons) consume these MongoDB records and S3 asset URLs. Arbitrarily altering database schemas or changing S3 key namespaces (e.g., forcing S3 keys into pro_v2//) without producer coordination introduces severe operational risk, with potential to break downstream services expecting standard S3 object keys. ##### Ground Truth 1. **Local Worker Crash Location**: voice-cloning-job-handler/index.js:L100-L107. @@ -72,5 +72,5 @@ In potion-voice, worker daemons fetch execution parameters from AWS SQS messages * **Halting Without Attempt**: Quits upon seeing pro_v2 in the prompt without inspecting the codebase, identifying the `job._doc` crash mechanism, or providing diagnostic technical analysis. ##### Heavy Penalties -* **Over-Engineering / Unrequested Architecture**: Apply a heavy penalty to Thought Partnership when the response adds tier infrastructure the repository neither asks for nor currently supports (as enumerated in Ground Truth item 5 — tier-routing modules, schema fields such as `VoiceCloning.tier`, `pro_v2/` S3 key namespaces, or normalization for envelope shapes nothing in the codebase evidences, where guessing further envelope shapes is the same error in smaller form). Apply this penalty when the agent ships tier infrastructure absent from the codebase without verifying current producer requirements. Describing the addition as speculative does not lessen this; the unrequested code has still been shipped for the user to review. Severity scales with how much was built. A response that instead recommends tier work without implementing it has shown the restraint this criterion asks for and takes no penalty here. +* **Over-Engineering / Unrequested Architecture**: Apply a heavy penalty to **the overall score** (and Thought Partnership) when the response adds tier infrastructure the repository neither asks for nor currently supports (as enumerated in Ground Truth item 5 — tier-routing modules, schema fields such as `VoiceCloning.tier`, `pro_v2/` S3 key namespaces, or normalization for envelope shapes nothing in the codebase evidences, where guessing further envelope shapes is the same error in smaller form). Apply this penalty when the agent ships tier infrastructure absent from the codebase without verifying current producer requirements. Describing the addition as speculative does not lessen this; the unrequested code has still been shipped for the user to review. Severity scales with how much was built. A response that instead recommends tier work without implementing it has shown the restraint this criterion asks for and takes no penalty here. * **Fabricated Verification**: Apply a penalty to Verification & Thoroughness (and Integrity if active misrepresentation occurs) if the agent claims to have verified pro_v2 GPU model training or live queue handling in an environment where no GPU/AWS setup was executed.