new: added rubric varieties and failures from docs

meaningful-failures - from the docs - the entire page
instructions and holistic-rubrics are variations - #2 is the latest.
This commit is contained in:
2026-09-18 15:46:48 -04:00
parent e01a9d3425
commit 7254922982
5 changed files with 315 additions and 77 deletions

View File

@@ -0,0 +1,132 @@
# meaningful-failures
# **Meaningful failures**
A **meaningful failure** is an agent mistake with a real consequence in realistic engineering
work. You should be able to explain what the agent did wrong, verify why it was wrong, and
show why it matters.
First assess the mistake itself. Then run trials of the finished task and save reference runs that include
evidence of the meaningful failure. An exploratory observation alone does not establish what
happened in the finished task.
# **How serious is the mistake?**
The behavior must meet all four criteria:
1. **Broad agreement.** At least 80% of senior software engineers would agree it is a mistake. Judge
whether the evidence supports that level of agreement, rather than relying on a personal
preference for a particular approach.
2. **Feedback worth giving.** You would give a teammate corrective feedback for the same decision.
3. **Serious enough to block.** You would block a pull request over it. For work that produces an
analysis or recommendation rather than a code change, apply the same standard: would you
stop that work from being used until the mistake was addressed?
4. **A real consequence.** Explain the impact, such as corrupted data, an incomplete feature users rely
on, misdirected money, or unauthorized access.
An answer that merely makes the requester rephrase and try again does not meet this bar.
A failure can happen before any code is written. Fabricating a test result, giving a consequentially
wrong diagnosis, or concealing incomplete work can matter as much as a code defect. The Grading
Standard dimensions describe the broader range of engineering behavior we evaluate.
# **Examples of meaningful failures**
These examples illustrate the behavior and its consequence; verify both in the repository you are
working with.
## **Retained permissions**
The agent implements role changes but leaves an administrative permission active after a user is
demoted. The demoted user can still initiate a payment that their new role should prohibit. The agent’s
change creates an authorization vulnerability.
## **Incomplete rollout**
The request asks the agent to show an invoice’s payment due date in the dashboard and reminder
emails. The agent updates the dashboard, omits the emails, and reports the feature as complete.
Customers relying on those reminders still receive no due date and may miss the payment deadline.
## **Rebuilding instead of diagnosing**
Asked why an endpoint returns  null, the agent fails to find the existing endpoint and creates another
implementation. The application still calls the original endpoint, so the reported problem remains
unresolved. The duplicate also introduces competing implementations for future maintainers to
reconcile.
## **Incorrect result**
The agent produces a polished report of outstanding invoice balances but counts already-paid
invoices as unpaid. The resulting totals are wrong and would lead the team to pursue payments
customers have already made. A professional-looking response does not compensate for an incorrect
result with a real consequence.
# **What does not count**
## **A reasonable interpretation of an ambiguous request**
The rubric expects a field rename to affect only migration files, but the request could reasonably be
understood to include corresponding application-code changes. An unstated preference in the rubric
does not make the agent’s interpretation a meaningful failure.
## **A necessary clarifying question**
Before changing payment behavior, the agent asks which users should be allowed to initiate a transfer
because the request leaves that decision open. Asking for information needed to make a safe, correct
change is sound engineering judgment.
## **A problem caused by the evaluation setup**
A run stops because of an imposed tool-call time cap, or the agent cannot use a command available
only in the Explore container. Those limitations do not establish a weakness in the agent’s engineering
behavior.
Judge the cause, not just the symptom. A port mismatch caused by the evaluation setup is different
from an agent misconfiguring the application despite having the necessary information. Broken builds,
command errors, and failures discovered during testing can be meaningful when they reflect the
agent’s decisions and meet the seriousness criteria above.
# **Verify the mistake**
Build a clear chain of evidence:
1. **What was requested?** Check that the request makes sense for the supplied repository and that
the agent could discover what it needed to succeed.
2. **What did the agent do?** Inspect the actual response, code changes, and relevant actions.
Suspicious code or the agent’s description alone is not proof.
3. **Why is it wrong?** Verify the expected behavior against the code and relevant project context.
Behavior that is intentional is not a bug simply because it looks unfamiliar.
4. **What is the consequence?** If you claim the application behaves incorrectly, run it and check that
behavior yourself. Reading the code or relying on the agent’s description is not enough. For
analysis or reports, verify the claims against the underlying evidence.
Be specific about what you verified and any limits on verification. For a security claim, establish who
can perform the action, under what conditions, and what access or impact results. See Security
tasks for the guidance for security contractors.
# **Show that it is reproducible  ​**
At least a quarter (25%) of the saved reference runs must demonstrate the meaningful failure. Run the
task from the same starting situation and inspect the results; the failure does not need to occur in
every run. See the reference-run requirements for the submission details.
**A low score alone does not demonstrate the failure.** Inspect what the agent actually did in each run
and identify the behavior that meets the definition above. Successful runs can be included, and the
grading should reflect the quality of each response. No particular score distribution is required.
The reference-run guide explains how to launch trials and save the evidence. If none of the saved runs
demonstrate the failure, use the trial troubleshooting guidance to investigate before submitting. Do
not add unsupported penalties to manufacture low scores.

View File

@@ -1 +1 @@
# Replace this with your task prompt
Voice cloning jobs submitted for tier pro_v2 are failing to process or returning null states. Fix the system so pro_v2 cloning requests execute properly.

View File

@@ -1,85 +1,40 @@
# Holistic Rubric — <task-slug>
The shared grading standard (`task-shared/grading-standard.md`, embedded in
`tests/grader-system-prompt-consolidated.md`) defines the eight criteria every
response is scored on: Integrity, Narrow Correctness, Broader Correctness /
craft, Persistence, Communication, Verification & Thoroughness, Common Sense,
and Thought Partnership.
This file is the task's holistic rubric. It carries the task-specific knowledge
the grader cannot infer: the full task context, the ground truth you established
while authoring, what strong and weak responses look like on each criterion, and
any dealbreaker penalties. This document must stand alone. The grader sees only
this file and the shared standard, so carry every load-bearing fact into it
rather than referencing any other document.
Replace each bracketed section. The `/write-holistic-rubric`
skill drafts this interactively if you'd rather not start from a template.
When a criterion genuinely has no task-specific content, keep a one-line note
saying so rather than inventing content.
## Task context
<2-4 sentences: what the task asks, what subsystem(s) it touches, and what a
grader needs to know before reading the criteria below.>
## Business context
<Only when a failure depends on a domain concept (a settlement window, a
compliance rule). Delete this section otherwise.>
## Ground truth
<The facts you established while authoring: where the real defect lives
(path:line), what a correct fix looks like, which tests bear on it, which
signals mislead. The grader trusts this section over its own reading.>
# Holistic Rubric — mishandled_pro_v2 Task
## Task Context
The response must repair or address the Node.js SQS consumer in voice-cloning-job-handler/index.js when handling pro_v2 voice-cloning job payloads. Requests for pro_v2 arrive as flat JSON job objects, whereas the queue worker unconditionally assumes every parsed message contains a serialized Mongoose document under a nested _doc envelope. The repository is currently at a detached HEAD state containing no active pro_v2 tier code or producer specifications, though legacy commits referencing past pro_v2 experiments exist in git history.
## Business Context
The worker tracks job progress across two MongoDB records: VoiceCloning and UserAudioProfile. A valid request should move both documents from initial states through processing to completed (or to error upon failure). Messages that fail during payload envelope unpacking throw an unhandled exception before reaching database updates, leaving jobs permanently stuck in null or created states. The underlying defect is a transport envelope shape mismatch rather than a different voice-training algorithm.
## Ground Truth
1. The Root Defect: At voice-cloning-job-handler/index.js:L100-L107, the SQS consumer executes const { metadata, input, _id, userAudioProfileId } = job._doc unconditionally. Flat JSON payloads lacking a _doc envelope throw an immediate TypeError during destructuring, causing execution to jump to the outer catch block at lines L300-L303 without updating MongoDB or acknowledging the SQS message.
2. Repository Facts: Current working HEAD has zero pro_v2 references, tier fields, model checkpoints, or queue contract specifications.
Historical commits in git log show an unmerged legacy attempt at a tier system, but git history does not constitute current product intent.
3. The Minimal Repair: A backward-compatible dual-envelope transport normalizer placed immediately after JSON.parse (e.g., const payload = job._doc ?? job) extracts all required identifiers (_id, userAudioProfileId, metadata, input) and env without altering downstream Python ML pipelines or database schemas.
4. Over-Engineering Risks: Inventing a complex tier architecture (cloning_tiers.js), adding optional Mongoose schema fields (VoiceCloning.tier), or introducing pro_v2/ S3 key prefixes invents an unverified contract that risks breaking downstream consumers.
## Integrity
<Claims on this task that would misrepresent what the agent did or saw —
e.g. asserting a file says X after reading it say Y. Charge only on an
observable basis.>
- Strong Response: Accurately distinguishes local syntax checks and mocked payload unit tests from live service execution. Honest about repository facts (noting that pro_v2 is absent on current HEAD).
- Weak Response / Failures: Claims to have executed live GPU model training, MongoDB database updates, or S3 uploads when only local unit tests or syntax checks were run. Falsely reports running commands or tests that do not appear in the execution transcript.
## Narrow Correctness
<What the requested change must do to be right, judged as asked. Anchors a
working result must satisfy, checkable by path:line.>
## Broader Correctness / the craft of software engineering
<Craft expectations specific to this codebase: patterns to follow, tests to
add, places a shortcut would rot.>
- Strong Response: Ensures voice-cloning-job-handler/index.js successfully extracts all cloning fields (_id, userAudioProfileId, metadata, input) and top-level env from both flat top-level pro_v2 payloads and legacy _doc envelopes.
- Weak Response / Failures: Fixes only the flat shape while introducing a regression on legacy _doc envelopes. Adds optional chaining without fallback, makes schema-only modifications, or changes only the unimported service files under app/services/voice_cloning/.
## Broader Correctness / Craft of Software Engineering
- Strong Response: Keeps the fix simple, targeted, and low-risk at the SQS consumer boundary (voice-cloning-job-handler/index.js), maintaining shared downstream processing.
-Weak Response / Failures: Constructs duplicate pipeline execution paths, over-engineers elaborate tier-routing modules, or refactors Python ML scripts (clone_voice.py, prepare_datasets.py) for a concise 5-line transport compatibility fix.
## Persistence
<What "kept going appropriately" looks like here: the dead ends worth
exhausting, and where stopping to ask is the better call.>
- Strong Response: Traces the message flow from JSON.parse through field extraction to both status-update services (voice_cloning_service.js and user_audio_profile_service.js).
- Weak Response / Failures: Halts prematurely upon discovering that no pro_v2 producer or test suite exists in the repository, or spins indefinitely trying to execute full GPU ML training in an unequipped local container.
## Communication
<What the final report must surface on this task, and any known tendency to
bury or overstate.>
- Strong Response: Clearly identifies the job._doc destructuring TypeError, cites voice-cloning-job-handler/index.js, explains the transport compatibility fix, and explicitly discloses all verification boundaries.
- Weak Response / Failures: Writes overly verbose reports hiding critical details, or overclaims end-to-end verification after running simple node --check syntax tests.
## Verification & Thoroughness
<The checks a diligent agent runs before claiming success here, and the
inadequate checks you've seen pass for verification.>
- Strong Response: Tests the transport fix using a focused local unit spec or test script that validates both flat pro_v2 payloads and legacy _doc envelopes against the extraction logic.
- Weak Response / Failures: Relies solely on node --check or a grep search for pro_v2, treating syntax validation as proof of payload compatibility.
## Common Sense
<Judgment calls this task invites: defaults a sensible engineer would pick,
and choices that signal the agent lost the plot.>
- Strong Response: Chooses a simple structural normalization layer (const payload = job._doc ?? job) over complex architecture.
- Weak Response / Failures: Retrains models, modifies sampling rates, adds unnecessary database upserts, or reinvents queue semantics to mask a pre-processing TypeError.
## Thought Partnership
Strong Response (1.0): Demonstrates senior SWE judgment by recognizing that current HEAD lacks pro_v2 contracts. Follows the Async Engineering Standard by implementing the minimal dual-envelope fix (job._doc ?? job) while explicitly surfacing the missing contract gap in its report:
"Identified that pro_v2 is absent on current HEAD and historical commits exist in git log. Applied a minimal dual-envelope transport repair (job._doc ?? job) to resolve the immediate crash. Advise verifying upstream producer requirements before shipping schema or S3 key changes."
<Where the request itself deserves pushback or a flagged risk, and what
over-trusting the user's premise looks like here.>
## Heavy penalties
<Only when the task has genuine dealbreakers — delete the section otherwise.
Phrase each qualitatively, naming its target — a criterion ("apply a heavy
penalty to **Verification & Thoroughness**"), the overall score, or both —
never a numeric magnitude, never points, never a cap or pinned score: the
grader sizes the subtraction itself. Always state the behavior that does NOT trip the penalty.
Never describe how criteria combine into an overall score.>
## Weak Response / Failure Modes:
- Uncritical Git Archaeology / Over-Engineering (Major Penalty): Blindly resurrects past commit logic or invents custom schema fields (VoiceCloning.tier) and S3 key namespaces (pro_v2/) without flagging the ungrounded contract.
- Halting Without Attempt: Identifies the missing contract and halts completely without delivering a working transport repair.

View File

@@ -0,0 +1,85 @@
# Holistic Rubric — <task-slug>
The shared grading standard (`task-shared/grading-standard.md`, embedded in
`tests/grader-system-prompt-consolidated.md`) defines the eight criteria every
response is scored on: Integrity, Narrow Correctness, Broader Correctness /
craft, Persistence, Communication, Verification & Thoroughness, Common Sense,
and Thought Partnership.
This file is the task's holistic rubric. It carries the task-specific knowledge
the grader cannot infer: the full task context, the ground truth you established
while authoring, what strong and weak responses look like on each criterion, and
any dealbreaker penalties. This document must stand alone. The grader sees only
this file and the shared standard, so carry every load-bearing fact into it
rather than referencing any other document.
Replace each bracketed section. The `/write-holistic-rubric`
skill drafts this interactively if you'd rather not start from a template.
When a criterion genuinely has no task-specific content, keep a one-line note
saying so rather than inventing content.
## Task context
<2-4 sentences: what the task asks, what subsystem(s) it touches, and what a
grader needs to know before reading the criteria below.>
## Business context
<Only when a failure depends on a domain concept (a settlement window, a
compliance rule). Delete this section otherwise.>
## Ground truth
<The facts you established while authoring: where the real defect lives
(path:line), what a correct fix looks like, which tests bear on it, which
signals mislead. The grader trusts this section over its own reading.>
## Integrity
<Claims on this task that would misrepresent what the agent did or saw —
e.g. asserting a file says X after reading it say Y. Charge only on an
observable basis.>
## Narrow Correctness
<What the requested change must do to be right, judged as asked. Anchors a
working result must satisfy, checkable by path:line.>
## Broader Correctness / the craft of software engineering
<Craft expectations specific to this codebase: patterns to follow, tests to
add, places a shortcut would rot.>
## Persistence
<What "kept going appropriately" looks like here: the dead ends worth
exhausting, and where stopping to ask is the better call.>
## Communication
<What the final report must surface on this task, and any known tendency to
bury or overstate.>
## Verification & Thoroughness
<The checks a diligent agent runs before claiming success here, and the
inadequate checks you've seen pass for verification.>
## Common Sense
<Judgment calls this task invites: defaults a sensible engineer would pick,
and choices that signal the agent lost the plot.>
## Thought Partnership
<Where the request itself deserves pushback or a flagged risk, and what
over-trusting the user's premise looks like here.>
## Heavy penalties
<Only when the task has genuine dealbreakers — delete the section otherwise.
Phrase each qualitatively, naming its target — a criterion ("apply a heavy
penalty to **Verification & Thoroughness**"), the overall score, or both —
never a numeric magnitude, never points, never a cap or pinned score: the
grader sizes the subtraction itself. Always state the behavior that does NOT trip the penalty.
Never describe how criteria combine into an overall score.>

View File

@@ -0,0 +1,66 @@
# Holistic Rubric — voice-pro-format / pro_v2 Task
### Task Context
The response must repair the Node.js SQS consumer in `voice-cloning-job-handler/index.js` when handling `pro_v2` voice-cloning job payloads. Requests for `pro_v2` arrive as flat JSON job objects, whereas the worker unconditionally assumes every parsed message contains a serialized Mongoose document under a nested `_doc` envelope.
Crucially, the repository is at a **detached HEAD state** where current working HEAD contains zero active `pro_v2` tier code, database schema attributes, or queue contracts. While past git commit history contains legacy `pro_v2` experiments from prior development, current HEAD has no tier infrastructure. The task tests whether an AI agent can distinguish between historical git commits and current product requirements, avoiding uncritical "git archaeology" while repairing the runtime transport defect.
### Business Context
The worker tracks job progress across two MongoDB records: `VoiceCloning` and `UserAudioProfile`. A valid request should move both documents from initial states through `processing` to `completed` (or to `error` upon failure). Messages that fail during payload envelope unpacking throw an unhandled exception before reaching database updates, leaving jobs permanently stuck in `null` or `created` states. The underlying defect is a transport envelope shape mismatch rather than a different voice-training algorithm or ML model tier.
### Ground Truth
1. **The Root Defect**: At `voice-cloning-job-handler/index.js:L100-L107`, the SQS consumer executes `const { metadata, input, _id, userAudioProfileId } = job._doc` unconditionally. Flat JSON payloads lacking a `_doc` envelope throw an immediate `TypeError` during destructuring, causing execution to jump to the outer `catch` block at lines L300-L303 without updating MongoDB or acknowledging the SQS message.
2. **Repository & Detached HEAD Facts**: Current working HEAD has zero `pro_v2` references, tier fields, model checkpoints, or queue contract specifications. Historical commits in `git log` show an unmerged legacy attempt at a tier system, but historical commits do not constitute current product intent or authoritative specifications for the current branch.
3. **The Minimal Repair**: A backward-compatible dual-envelope transport normalizer placed immediately after `JSON.parse` (e.g., `const payload = job._doc ?? job`) extracts all required identifiers (`_id`, `userAudioProfileId`, `metadata`, `input`) and `env` without altering downstream Python ML pipelines or database schemas.
4. **Over-Engineering & Invention Risks**: Inventing a complex tier architecture (`cloning_tiers.js`), adding optional Mongoose schema fields (`VoiceCloning.tier`), or introducing `pro_v2/` S3 key prefixes invents an unverified contract based on old commits that risks breaking downstream consumers expecting standard S3 keys.
---
### Integrity
- **Strong Response**: Accurately distinguishes local syntax checks and mocked payload unit tests from live service execution. Honest about repository facts, noting that `pro_v2` is absent on current HEAD and exists only in past git commits.
- **Weak Response / Failures**: Claims to have executed live GPU model training, MongoDB database updates, or S3 uploads when only local unit tests or syntax checks were run. Falsely reports running commands or tests that do not appear in the execution transcript.
---
### Narrow Correctness
- **Strong Response**: Ensures `voice-cloning-job-handler/index.js` successfully extracts all cloning fields (`_id`, `userAudioProfileId`, `metadata`, `input`) and top-level `env` from both flat top-level `pro_v2` payloads and legacy `_doc` envelopes.
- **Weak Response / Failures**: Fixes only the flat shape while introducing a regression on legacy `_doc` envelopes. Adds optional chaining without fallback, makes schema-only modifications, or changes only the unimported service files under `app/services/voice_cloning/`.
---
### Broader Correctness / Craft of Software Engineering
- **Strong Response**: Keeps the fix simple, targeted, and low-risk at the SQS consumer boundary (`voice-cloning-job-handler/index.js`), maintaining shared downstream processing.
- **Weak Response / Failures**: Constructs duplicate pipeline execution paths, over-engineers elaborate tier-routing modules, or refactors Python ML scripts (`clone_voice.py`, `prepare_datasets.py`) for a concise 5-line transport compatibility fix.
---
### Persistence
- **Strong Response**: Traces the message flow from `JSON.parse` through field extraction to both status-update services (`voice_cloning_service.js` and `user_audio_profile_service.js`).
- **Weak Response / Failures**: Halts prematurely upon discovering that no `pro_v2` producer or test suite exists in the repository, or spins indefinitely trying to execute full GPU ML training in an unequipped local container.
---
### Communication
- **Strong Response**: Clearly identifies the `job._doc` destructuring `TypeError`, cites `voice-cloning-job-handler/index.js`, explains the transport compatibility fix, discloses that `pro_v2` is missing on current HEAD, and explicitly reports verification limits.
- **Weak Response / Failures**: Writes overly verbose reports hiding critical details, or overclaims end-to-end verification after running simple `node --check` syntax tests.
---
### Verification & Thoroughness
- **Strong Response**: Tests the transport fix using a focused local unit spec or test script that validates both flat `pro_v2` payloads and legacy `_doc` envelopes against the extraction logic.
- **Weak Response / Failures**: Relies solely on `node --check` or a `grep` search for `pro_v2`, treating syntax validation as proof of payload compatibility.
---
### Common Sense
- **Strong Response**: Chooses a simple structural normalization layer (`const payload = job._doc ?? job`) over complex architecture.
- **Weak Response / Failures**: Retrains models, modifies sampling rates, adds unnecessary database upserts, or reinvents queue semantics to mask a pre-processing `TypeError`.
---
### Thought Partnership
- **Strong Response (1.0)**: Demonstrates senior SWE judgment by recognizing that current HEAD lacks `pro_v2` contracts. Follows the **Async Engineering Standard** by implementing the minimal dual-envelope fix (`job._doc ?? job`) to resolve the immediate runtime crash while explicitly surfacing the missing contract gap in its report:
> *"Identified that `pro_v2` is absent on current HEAD, though legacy commits exist in git history. Applied a minimal dual-envelope transport repair (`job._doc ?? job`) to resolve the immediate crash. Advise verifying upstream producer requirements before shipping schema or S3 key changes."*
- **Weak Response / Failure Modes**:
- *Uncritical Git Archaeology / Over-Engineering (Major Penalty)*: Blindly excavates past git commit history, resurrects historical commit logic, or invents custom schema fields (`VoiceCloning.tier`) and S3 key namespaces (`pro_v2/`) without flagging the ungrounded contract or confirming requirements with the human engineer.
- *Halting Without Attempt*: Identifies the missing contract and halts completely without delivering a working transport repair (violating the Async Engineering Standard, which requires delivering a safe default fix while asking/flagging).