Files
project-work/sources/meaningful-failures.md
Eric Bell 7254922982 new: added rubric varieties and failures from docs
meaningful-failures - from the docs - the entire page
instructions and holistic-rubrics are variations - #2 is the latest.
2026-09-18 15:46:48 -04:00

6.2 KiB
Raw Blame History

meaningful-failures

Meaningful failures

A meaningful failure is an agent mistake with a real consequence in realistic engineering

work. You should be able to explain what the agent did wrong, verify why it was wrong, and

show why it matters.

First assess the mistake itself. Then run trials of the finished task and save reference runs that include evidence of the meaningful failure. An exploratory observation alone does not establish what happened in the finished task.

How serious is the mistake?

The behavior must meet all four criteria:

  1. Broad agreement. At least 80% of senior software engineers would agree it is a mistake. Judge whether the evidence supports that level of agreement, rather than relying on a personal preference for a particular approach.

  2. Feedback worth giving. You would give a teammate corrective feedback for the same decision.

  3. Serious enough to block. You would block a pull request over it. For work that produces an analysis or recommendation rather than a code change, apply the same standard: would you stop that work from being used until the mistake was addressed?

  4. A real consequence. Explain the impact, such as corrupted data, an incomplete feature users rely on, misdirected money, or unauthorized access.

An answer that merely makes the requester rephrase and try again does not meet this bar.

A failure can happen before any code is written. Fabricating a test result, giving a consequentially wrong diagnosis, or concealing incomplete work can matter as much as a code defect. The Grading Standard dimensions describe the broader range of engineering behavior we evaluate.

Examples of meaningful failures

These examples illustrate the behavior and its consequence; verify both in the repository you are working with.

Retained permissions

The agent implements role changes but leaves an administrative permission active after a user is demoted. The demoted user can still initiate a payment that their new role should prohibit. The agent’s change creates an authorization vulnerability.

Incomplete rollout

The request asks the agent to show an invoice’s payment due date in the dashboard and reminder emails. The agent updates the dashboard, omits the emails, and reports the feature as complete. Customers relying on those reminders still receive no due date and may miss the payment deadline.

Rebuilding instead of diagnosing

Asked why an endpoint returns  null, the agent fails to find the existing endpoint and creates another implementation. The application still calls the original endpoint, so the reported problem remains unresolved. The duplicate also introduces competing implementations for future maintainers to reconcile.

Incorrect result

The agent produces a polished report of outstanding invoice balances but counts already-paid invoices as unpaid. The resulting totals are wrong and would lead the team to pursue payments customers have already made. A professional-looking response does not compensate for an incorrect result with a real consequence.

What does not count

A reasonable interpretation of an ambiguous request

The rubric expects a field rename to affect only migration files, but the request could reasonably be understood to include corresponding application-code changes. An unstated preference in the rubric does not make the agent’s interpretation a meaningful failure.

A necessary clarifying question

Before changing payment behavior, the agent asks which users should be allowed to initiate a transfer because the request leaves that decision open. Asking for information needed to make a safe, correct change is sound engineering judgment.

A problem caused by the evaluation setup

A run stops because of an imposed tool-call time cap, or the agent cannot use a command available only in the Explore container. Those limitations do not establish a weakness in the agent’s engineering behavior.

Judge the cause, not just the symptom. A port mismatch caused by the evaluation setup is different from an agent misconfiguring the application despite having the necessary information. Broken builds, command errors, and failures discovered during testing can be meaningful when they reflect the agent’s decisions and meet the seriousness criteria above.

Verify the mistake

Build a clear chain of evidence:

  1. What was requested? Check that the request makes sense for the supplied repository and that the agent could discover what it needed to succeed.

  2. What did the agent do? Inspect the actual response, code changes, and relevant actions. Suspicious code or the agent’s description alone is not proof.

  3. Why is it wrong? Verify the expected behavior against the code and relevant project context. Behavior that is intentional is not a bug simply because it looks unfamiliar.

  4. What is the consequence? If you claim the application behaves incorrectly, run it and check that behavior yourself. Reading the code or relying on the agent’s description is not enough. For analysis or reports, verify the claims against the underlying evidence.

Be specific about what you verified and any limits on verification. For a security claim, establish who can perform the action, under what conditions, and what access or impact results. See Security tasks for the guidance for security contractors.

Show that it is reproducible  ​

At least a quarter (25%) of the saved reference runs must demonstrate the meaningful failure. Run the task from the same starting situation and inspect the results; the failure does not need to occur in every run. See the reference-run requirements for the submission details.

A low score alone does not demonstrate the failure. Inspect what the agent actually did in each run and identify the behavior that meets the definition above. Successful runs can be included, and the grading should reflect the quality of each response. No particular score distribution is required.

The reference-run guide explains how to launch trials and save the evidence. If none of the saved runs demonstrate the failure, use the trial troubleshooting guidance to investigate before submitting. Do not add unsupported penalties to manufacture low scores.