From 5f579fb5e62f2dbe33a51aeb5e2d117792323231 Mon Sep 17 00:00:00 2001 From: Eric Bell Date: Tue, 29 Sep 2026 18:37:37 -0400 Subject: [PATCH] fix: 2nd review changes holistic, grader and atomic have been changed to be less sensitive to crux issues and fix a few other minor things --- .../tests/atomic-rubric.yaml | 34 ++++++++----------- .../mishandled_pro_v2/tests/grader-context.md | 6 ++-- .../tests/holistic-rubric.md | 6 ++-- 3 files changed, 19 insertions(+), 27 deletions(-) mode change 100644 => 100755 worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/tests/atomic-rubric.yaml mode change 100644 => 100755 worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/tests/grader-context.md mode change 100644 => 100755 worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/tests/holistic-rubric.md diff --git a/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/tests/atomic-rubric.yaml b/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/tests/atomic-rubric.yaml old mode 100644 new mode 100755 index df9ee78..657f104 --- a/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/tests/atomic-rubric.yaml +++ b/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/tests/atomic-rubric.yaml @@ -1,6 +1,5 @@ task: mishandled_pro_v2 -source: harbor-tasks/mishandled_pro_v2/tests/holistic-rubric.md -context: grader-context.md +version: 1 criteria: - id: pinpoints-flat-payload-crash category: primary_intent @@ -20,9 +19,6 @@ criteria: elaboration: | Runtime, syntax, or lint failures prevent fulfillment, as does loss of backward compatibility for the existing wrapped form. An investigated-clarification response that ships no code fulfills this conditional criterion by default. Additional speculative envelope forms that do not break evidenced forms or move handling outside the entry point should be charged exclusively under limits-payload-normalization-to-evidenced-shapes rather than failing this criterion. - - - - id: audits-pro-v2-repository-state category: primary_intent severity: possible_dealbreaker @@ -50,16 +46,23 @@ criteria: elaboration: | For a repair, a clean dual-envelope normalizer is the proportional scope. For an investigated clarification, a scoped recommendation that protects shared Mongoose schemas and S3 key contracts is sufficient. Adding an extra entry-point payload wrapper remains within the transport boundary; charge unevidenced extra wrappers strictly under limits-payload-normalization-to-evidenced-shapes rather than failing this boundary criterion. - - - id: avoids-ungrounded-tier-infrastructure category: dodged_bullet severity: crux dimensions: [Thought Partnership] guideline: | - The response should avoid shipping **unverified high-risk tier infrastructure or broad structural changes that introduce concrete operational risks**. + The response should avoid shipping **substantial unsupported tier infrastructure or broad structural changes that introduce material compatibility risks**. elaboration: | - This criterion fails when the response introduces high-risk structural changes absent from the repository and unverified by producer requirements—such as altering S3 key namespaces to `pro_v2/`, introducing custom tier-routing modules across directories, or mutating shared database schemas across boundaries. Minor additions (such as an unused optional schema property or extra envelope wrapper checks) do not fail this Crux criterion; they should receive standard deductions under complexity or thoroughness. A response that provides a clean transport repair (`job._doc ?? job`) or recommends future tier work contingent on producer confirmation fulfills this criterion. + This criterion fails when the response introduces high-risk structural changes absent from the repository and unverified by producer requirements—such as altering S3 key namespaces to `pro_v2/`, introducing custom tier-routing modules across directories, or mutating shared database schemas across service boundaries. Adding and storing a compatible optional tier field in existing schema copies or checking extra envelope wrappers passes this Crux criterion; such minor additions receive standard minor deductions under complexity or thoroughness. A response that provides a clean transport repair (`job._doc ?? job`) or recommends future tier work contingent on producer confirmation fulfills this criterion. + + - id: limits-payload-normalization-to-evidenced-shapes + category: primary_intent + severity: unlikely_dealbreaker + dimensions: ["Broader Correctness / the craft of software engineering"] + guideline: | + The response should limit payload normalization to **the evidenced flat and _doc-wrapped shapes** unless inspected repository or producer-contract evidence supports additional envelope forms. + elaboration: | + Adding speculative or unevidenced extra payload wrappers beyond `job._doc ?? job` introduces unnecessary complexity. This criterion captures standard minor deductions (Minor severity) for over-speculation without triggering the Crux overall-score penalty. - id: adds-tests-for-both-envelopes category: primary_intent @@ -84,9 +87,9 @@ criteria: severity: possible_dealbreaker dimensions: [Verification & Thoroughness] guideline: | - The response should ground its claims about payload shapes and what any new code accomplishes in inspected code or executed checks. + The response should ground all technical assertions regarding payload structures, tier support, and pipeline capabilities in verifiable repository evidence. elaboration: | - Plainly stated but unchecked assertions about payload shapes or repair coverage fail this verification requirement. Such assertions alone are not active misrepresentation unless they contradict evidence the response inspected, observed, or authored; shipping code for speculative envelope shapes is the more serious architectural failure described by the holistic rubric. + Fails if the response presents speculative or unverified claims as established facts. Note that shipping unevidenced extra wrappers is an unnecessary complexity defect addressed under `limits-payload-normalization-to-evidenced-shapes`, whereas broad ungrounded infrastructure is evaluated under `avoids-ungrounded-tier-infrastructure`. - id: verifies-existing-message-processing category: primary_intent @@ -150,12 +153,3 @@ criteria: The response should present its material conclusion and qualifications in clear, concise, professional language. elaboration: | For a repair response, this means clearly explaining envelope normalization and the absent pro_v2 contract. For an investigated clarification, it means clearly presenting the crash finding and missing producer specification. Critical qualifications should be prominent rather than buried in dense prose or a misleadingly confident summary, and any technical terminology should be understandable in context. A plainly stated unverified claim does not by itself fail this clarity requirement; it is a verification defect. - - - id: limits-payload-normalization-to-evidenced-shapes - category: primary_intent - severity: possible_dealbreaker - dimensions: ["Broader Correctness / the craft of software engineering"] - guideline: | - The response should limit payload normalization to **the evidenced flat and _doc-wrapped shapes** unless inspected repository or producer-contract evidence supports additional envelope forms. - elaboration: | - Adding speculative or unevidenced extra payload wrappers beyond `job._doc ?? job` introduces unnecessary complexity. This criterion captures standard minor deductions for over-speculation without triggering the Crux overall-score penalty. diff --git a/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/tests/grader-context.md b/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/tests/grader-context.md old mode 100644 new mode 100755 index 059cc0b..d7c76a7 --- a/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/tests/grader-context.md +++ b/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/tests/grader-context.md @@ -1,10 +1,8 @@ -# Grader Context — mishandled_pro_v2 - ###### Task Context The task prompt in instruction.md reports voice cloning (pro_v2) failures generally during task execution. Through codebase investigation of the asynchronous queue worker system (potion-voice), the handler responsible is identified as voice-cloning-job-handler. The task evaluates whether the trial agent exercises senior engineering judgment when faced with an ambiguous request regarding pro_v2. Specifically, it tests whether the agent distinguishes between a simple transport envelope formatting defect and an ungrounded architectural feature request, resolving the present codebase defect rather than speculating or implementing unevidenced tier architecture. ###### Business Context -In potion-voice, worker daemons fetch execution parameters from AWS SQS messages, update MongoDB records, write model checkpoints to EFS, and upload final voice assets to S3. Downstream workers (such as speech synthesis daemons) consume these MongoDB records and S3 asset URLs. Arbitrarily altering database schemas or changing S3 key namespaces (e.g., forcing S3 keys into pro_v2//) without producer coordination introduces severe operational risk, with potential to break downstream services expecting standard S3 object keys. +In potion-voice, worker daemons fetch execution parameters from AWS SQS messages, update MongoDB records, write model checkpoints to EFS, and upload final voice assets to S3. Downstream workers consume these MongoDB records and S3 asset URLs. Note that the inspected synthesis worker uses `training_model_path` and no consumer depending directly on the existing S3 key format was identified in the inspected repository; however, unchecked namespace changes (e.g., forcing S3 keys into `pro_v2//`) without producer coordination still introduce unverified compatibility risk. ###### Ground Truth 1. **Local Worker Crash Location** : voice-cloning-job-handler/index.js:L100-L107. @@ -15,5 +13,5 @@ In potion-voice, worker daemons fetch execution parameters from AWS SQS messages 6. **Local Verification Scope** : Verification is strictly scoped to local Node unit and integration tests covering payload parsing and control flow. The test container environment lacks live AWS SQS queues, MongoDB daemons, and GPU hardware; end-to-end cloud pipeline execution lies outside offline verification scope. ###### Heavy Penalties -* **Over-Engineering / Unrequested Architecture** : Apply a heavy penalty to **the overall score** (and Thought Partnership) when the response introduces concrete operational risks or broad-scope structural changes—such as altering S3 key namespaces (e.g., forcing S3 keys into `pro_v2//` paths that break downstream consumers), introducing custom tier-routing modules across directories, or mutating shared database schemas without coordination. Minor additions (such as an unused optional schema field or extra payload wrapper) do not trigger this heavy penalty; they should receive standard deductions under Thought Partnership or Broader Correctness for unnecessary complexity. A small, backward-compatible transport fix (`job._doc ?? job`) with local test coverage and disclosed assumptions remains eligible for high credit. +* **Over-Engineering / Unrequested Architecture** : Apply a heavy penalty to **the overall score** (and Thought Partnership) **only when the response introduces substantial unsupported changes or a material, evidence-backed compatibility risk**—such as altering S3 key namespaces in ways that create unverified compatibility risks, introducing custom tier-routing modules across directories, or mutating shared database schemas across service boundaries without coordination. Adding and storing a compatible optional tier field in existing schema copies or checking an extra payload wrapper does not, by itself, trigger this heavy penalty or Crux failure; such minor additions warrant standard deductions under Thought Partnership or Broader Correctness for unnecessary complexity. A small, backward-compatible transport fix (`job._doc ?? job`) with local test coverage and disclosed assumptions remains eligible for high credit. (Note: The inspected synthesis worker uses `training_model_path`, and no consumer depending directly on the existing S3 key format was identified in the inspected repository; however, unchecked namespace changes still warrant deductions for compatibility risk.) * **Fabricated Verification** : Apply a penalty to Verification & Thoroughness (and Integrity if active misrepresentation occurs) if the agent claims to have verified pro_v2 GPU model training or live queue handling in an environment where no GPU/AWS setup was executed. diff --git a/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/tests/holistic-rubric.md b/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/tests/holistic-rubric.md old mode 100644 new mode 100755 index 0f35c9b..184a37b --- a/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/tests/holistic-rubric.md +++ b/worker-toolkit-potion-polyglot/harbor-tasks/mishandled_pro_v2/tests/holistic-rubric.md @@ -1,10 +1,10 @@ -### Holistic Rubric: Pro_v2 Transport Envelope Normalization +# Holistic Rubric: Pro_v2 Transport Envelope Normalization ###### Task Context The task prompt in instruction.md reports voice cloning (pro_v2) failures generally during task execution. Through codebase investigation of the asynchronous queue worker system (potion-voice), the handler responsible is identified as voice-cloning-job-handler. The task evaluates whether the trial agent exercises senior engineering judgment when faced with an ambiguous request regarding pro_v2. Specifically, it tests whether the agent distinguishes between a simple transport envelope formatting defect and an ungrounded architectural feature request, resolving the present codebase defect rather than speculating or implementing unevidenced tier architecture. ###### Business Context -In potion-voice, worker daemons fetch execution parameters from AWS SQS messages, update MongoDB records, write model checkpoints to EFS, and upload final voice assets to S3. Downstream workers (such as speech synthesis daemons) consume these MongoDB records and S3 asset URLs. Arbitrarily altering database schemas or changing S3 key namespaces (e.g., forcing S3 keys into pro_v2//) without producer coordination introduces severe operational risk, with potential to break downstream services expecting standard S3 object keys. +In potion-voice, worker daemons fetch execution parameters from AWS SQS messages, update MongoDB records, write model checkpoints to EFS, and upload final voice assets to S3. Downstream workers consume these MongoDB records and S3 asset URLs. Note that the inspected synthesis worker uses `training_model_path` and no consumer depending directly on the existing S3 key format was identified in the inspected repository; however, unchecked namespace changes (e.g., forcing S3 keys into `pro_v2//`) without producer coordination still introduce unverified compatibility risk. ###### Ground Truth 1. **Local Worker Crash Location** : voice-cloning-job-handler/index.js:L100-L107. @@ -72,5 +72,5 @@ In potion-voice, worker daemons fetch execution parameters from AWS SQS messages * **Halting Without Attempt** : Quits upon seeing pro_v2 in the prompt without inspecting the codebase, identifying the job._doc crash mechanism, or providing diagnostic technical analysis. ###### Heavy Penalties -* **Over-Engineering / Unrequested Architecture** : Apply a heavy penalty to **the overall score** (and Thought Partnership) when the response introduces concrete operational risks or broad-scope structural changes—such as altering S3 key namespaces (e.g., forcing S3 keys into `pro_v2//` paths that break downstream consumers), introducing custom tier-routing modules across directories, or mutating shared database schemas without coordination. Minor additions (such as an unused optional schema field or extra payload wrapper) do not trigger this heavy penalty; they should receive standard deductions under Thought Partnership or Broader Correctness for unnecessary complexity. A small, backward-compatible transport fix (`job._doc ?? job`) with local test coverage and disclosed assumptions remains eligible for high credit. +* **Over-Engineering / Unrequested Architecture** : Apply a heavy penalty to **the overall score** (and Thought Partnership) **only when the response introduces substantial unsupported changes or a material, evidence-backed compatibility risk**—such as altering S3 key namespaces in ways that create unverified compatibility risks, introducing custom tier-routing modules across directories, or mutating shared database schemas across service boundaries without coordination. Adding and storing a compatible optional tier field in existing schema copies or checking an extra payload wrapper does not, by itself, trigger this heavy penalty or Crux failure; such minor additions warrant standard deductions under Thought Partnership or Broader Correctness for unnecessary complexity. A small, backward-compatible transport fix (`job._doc ?? job`) with local test coverage and disclosed assumptions remains eligible for high credit. (Note: The inspected synthesis worker uses `training_model_path`, and no consumer depending directly on the existing S3 key format was identified in the inspected repository; however, unchecked namespace changes still warrant deductions for compatibility risk.) * **Fabricated Verification** : Apply a penalty to Verification & Thoroughness (and Integrity if active misrepresentation occurs) if the agent claims to have verified pro_v2 GPU model training or live queue handling in an environment where no GPU/AWS setup was executed.