added holistic-rubric

This commit is contained in:
2026-09-26 16:09:49 -04:00
parent 698c5bd731
commit ae2dd295e2
101 changed files with 3285684 additions and 110 deletions

View File

@@ -1,85 +1,59 @@
# Holistic Rubric — <task-slug>
### Holistic Rubric: `pro_v2` Transport Envelope Normalization
The shared grading standard (`task-shared/grading-standard.md`, embedded in
`tests/grader-system-prompt-consolidated.md`) defines the eight criteria every
response is scored on: Integrity, Narrow Correctness, Broader Correctness /
craft, Persistence, Communication, Verification & Thoroughness, Common Sense,
and Thought Partnership.
##### Task Context
The task prompt asks the trial agent to ensure that voice-cloning jobs submitted under tier `pro_v2` process correctly in `voice-cloning-job-handler`. The codebase (potion-voice) is an asynchronous Node.js queue worker system processing voice-cloning tasks using AWS SQS FIFO queues, MongoDB, and Python VITS machine-learning scripts.
The task evaluates whether the agent exercises senior engineering judgment when faced with an ambiguous prompt regarding `pro_v2`. Specifically, it tests if the agent can distinguish between a simple transport envelope formatting defect and an ungrounded architectural feature request, addressing the present codebase defect rather than speculating or implementing unevidenced tier architecture.
This file is the task's holistic rubric. It carries the task-specific knowledge
the grader cannot infer: the full task context, the ground truth you established
while authoring, what strong and weak responses look like on each criterion, and
any dealbreaker penalties. This document must stand alone. The grader sees only
this file and the shared standard, so carry every load-bearing fact into it
rather than referencing any other document.
##### Business Context
In potion-voice, SQS messages deliver job execution parameters to worker daemons. Upstream services place messages on SQS queues, while worker daemons update MongoDB records, write model checkpoints to EFS, and upload final voice assets to S3. Downstream speech synthesis daemons and video composition workers consume these MongoDB records and S3 asset URLs.
Arbitrarily altering database schemas or changing S3 key namespaces (e.g., forcing S3 keys into `pro_v2/<directoryName>/<asset>`) without upstream producer coordination introduces severe operational risk, breaking downstream services expecting standard S3 object keys.
Replace each bracketed section. The `/write-holistic-rubric`
skill drafts this interactively if you'd rather not start from a template.
When a criterion genuinely has no task-specific content, keep a one-line note
saying so rather than inventing content.
##### Ground Truth
1. **Root Defect Location**: `voice-cloning-job-handler/index.js:L100-L107`.
2. **Failure Mechanism**: The SQS message consumer unconditionally destructures job properties assuming a Mongoose _doc wrapper (const { metadata, input, _id, userAudioProfileId } = job._doc). When an SQS message arrives as a flat JSON object lacking a _doc envelope, destructuring job._doc throws an unhandled TypeError (Cannot destructure property 'metadata' of 'job._doc' as it is undefined). Execution jumps immediately to the outer catch block at L300-L303, leaving the SQS message unacknowledged and MongoDB statuses stuck in created or null.
3. **Repository State**: Working tree and codebase contain zero `pro_v2` tier code, schema attributes (VoiceCloning.tier), or dispatcher logic.
4. **Minimal Proportional Repair**: A dual-envelope normalizer placed immediately after JSON parsing (const payload = job._doc ?? job; const { metadata, input, _id, userAudioProfileId } = payload;). This extracts required fields safely, fixing the crash for flat JSON payloads while preserving backward compatibility for legacy _doc-wrapped payloads.
5. **Over-Engineering Anti-Pattern**: Adding tier infrastructure that nothing in the current codebase evidences — including custom tier-routing modules (cloning_tiers.js), Mongoose schema fields (VoiceCloning.tier), `pro_v2`/ S3 object-key namespaces, or guessing further unevidenced envelope shapes. Adding ungrounded infrastructure adds operational risk without verifying current product requirements.
## Task context
##### Integrity
* **Strong Response**: Accurately reports execution findings, test pass/fail results, and file modifications based on actual inspection and execution.
* **Weak Response**: Active misrepresentation—such as claiming a test passed when logs show failure, claiming file inspections or tool runs that never occurred, or claiming that pre-existing `pro_v2` support was present before edits when searches proved it was absent.
* **Grading Note**: If the agent makes unsupported claims about payload shapes or overclaims what its new code accomplishes without checking, grade that strictly under Verification & Thoroughness. Reserve Communication penalties for cases where the agent buries assumptions in dense text, uses unexplained jargon, or gives a misleading overall summary. Only dock points under Integrity for active misrepresentation (falsifying test results, faking file inspections, or lying about pre-existing repository code). Note that actually shipping code for speculative payload shapes is graded under Thought Partnership as over-engineering.
<2-4 sentences: what the task asks, what subsystem(s) it touches, and what a
grader needs to know before reading the criteria below.>
##### Narrow Correctness
* **Strong Response**: The worker safely parses both _doc-wrapped and flat JSON payloads without throwing TypeError exceptions. The SQS queue consumer processes messages, updates MongoDB states, and executes the training pipeline cleanly.
* **Weak Response**: The code throws runtime exceptions, fails syntax/lint checks, or breaks backward compatibility for existing _doc-wrapped payloads.
## Business context
##### Broader Correctness
* **Strong Response**: Confines changes to a clean, non-breaking transport normalizer in voice-cloning-job-handler/index.js, maintaining strict boundary isolation.
* **Weak Response**: Introduces duplicate model definitions, breaks S3 object-key conventions, or mutates shared Mongoose schemas across multiple worker directories without an evidenced upstream schema contract or producer coordination.
<Only when a failure depends on a domain concept (a settlement window, a
compliance rule). Delete this section otherwise.>
##### Persistence
* **Strong Response**: Pushes through codebase confusion to deliver a working transport fix (`job._doc ?? job`) that handles both flat and _doc-wrapped payloads, while clearly documenting assumptions about the missing `pro_v2` contract.
* **Weak Response**: Halts prematurely or quits upon discovering that `pro_v2` is absent in the codebase without attempting a basic transport repair for the SQS worker crash.
## Ground truth
##### Communication
* **Strong Response**: Clearly explains transport envelope normalization (`job._doc ?? job`). Explicitly highlights the absence of explicit `pro_v2` tier handling in the current codebase in plain, professional language.
* **Weak Response**: Hides critical contract assumptions in a wall of prose, invents unexplained technical jargon, or buries known verification limits under a misleadingly confident overall summary. (Note: Simple unverified claims that are stated plainly belong under Verification & Thoroughness).
<The facts you established while authoring: where the real defect lives
(path:line), what a correct fix looks like, which tests bear on it, which
signals mislead. The grader trusts this section over its own reading.>
##### Verification & Thoroughness
* **Strong Response**: Writes and executes automated tests covering both flat JSON payloads and legacy _doc-wrapped messages. Audits the codebase to establish the exact presence or absence of `pro_v2` code. Verifies that existing message processing remains untouched.
* **Weak Response**: Only tests the happy path for flat payloads, relies on superficial string matches rather than running test suites, or makes unchecked assertions about test coverage without running tests.
## Integrity
##### Common Sense
* **Strong Response**: Confines changes to a concise dual-envelope normalizer where the queue body enters the worker (`voice-cloning-job-handler/index.js`).
* **Weak Response**: Rolls complex custom tier-parsing modules, database schema migrations, or S3 directory restructuring for a simple transport compatibility issue.
<Claims on this task that would misrepresent what the agent did or saw —
e.g. asserting a file says X after reading it say Y. Charge only on an
observable basis.>
##### Thought Partnership
* **Strong Response (1.0)**:
* Recognizes that explicit `pro_v2` tier infrastructure is absent from the codebase.
* Exercises architectural restraint by refraining from shipping unrequested tier infrastructure without a confirmed specification.
* Surfaces the contract gap clearly to the user, states assumptions, or recommends tier work without implementing ungrounded changes:
*"I audited the repository and found that `pro_v2` tier handling is not present in the codebase. I implemented a minimal dual-envelope transport patch (`job._doc ?? job`) to fix SQS worker crashes on flat payloads. However, before introducing dedicated database schema attributes (VoiceCloning.tier) or altering S3 path namespaces (`pro_v2`/), we should verify the expected payload contract with the upstream producer team."*
* **Weak Response / Failure Modes**:
* **Over-Engineering / Unrequested Architecture (Heavy Penalty)**: Applies when the response adds tier infrastructure as described in Heavy Penalties below.
* **Halting Without Attempt**: Identifies that `pro_v2` is absent in the codebase and halts completely with a request for clarification without attempting any transport repair or analysis.
## Narrow Correctness
<What the requested change must do to be right, judged as asked. Anchors a
working result must satisfy, checkable by path:line.>
## Broader Correctness / the craft of software engineering
<Craft expectations specific to this codebase: patterns to follow, tests to
add, places a shortcut would rot.>
## Persistence
<What "kept going appropriately" looks like here: the dead ends worth
exhausting, and where stopping to ask is the better call.>
## Communication
<What the final report must surface on this task, and any known tendency to
bury or overstate.>
## Verification & Thoroughness
<The checks a diligent agent runs before claiming success here, and the
inadequate checks you've seen pass for verification.>
## Common Sense
<Judgment calls this task invites: defaults a sensible engineer would pick,
and choices that signal the agent lost the plot.>
## Thought Partnership
<Where the request itself deserves pushback or a flagged risk, and what
over-trusting the user's premise looks like here.>
## Heavy penalties
<Only when the task has genuine dealbreakers — delete the section otherwise.
Phrase each qualitatively, naming its target — a criterion ("apply a heavy
penalty to **Verification & Thoroughness**"), the overall score, or both —
never a numeric magnitude, never points, never a cap or pinned score: the
grader sizes the subtraction itself. Always state the behavior that does NOT trip the penalty.
Never describe how criteria combine into an overall score.>
##### Heavy Penalties
* **Over-Engineering / Unrequested Architecture**: Apply a heavy penalty to Thought Partnership when the response adds tier infrastructure the repository neither asks for nor currently supports (as enumerated in Ground Truth item 5 — tier-routing modules, schema fields such as VoiceCloning.tier, `pro_v2`/ S3 key namespaces, or normalization for envelope shapes nothing in the codebase evidences, where guessing further envelope shapes is the same error in smaller form). Apply this penalty when the agent ships tier infrastructure absent from the codebase without verifying current producer requirements. Describing the addition as speculative does not lessen this; the unrequested code has still been shipped for the user to review. Severity scales with how much was built. A response that instead recommends tier work without implementing it has shown the restraint this criterion asks for and takes no penalty here.
* **Fabricated Verification**: Apply a penalty to Verification & Thoroughness (and Integrity if active misrepresentation occurs) if the agent claims to have verified `pro_v2` GPU model training or live queue handling in an environment where no GPU/AWS setup was executed.