Files
project-work/sources/meaningful-failures.md
Eric Bell 7254922982 new: added rubric varieties and failures from docs
meaningful-failures - from the docs - the entire page
instructions and holistic-rubrics are variations - #2 is the latest.
2026-09-18 15:46:48 -04:00

133 lines
6.2 KiB
Markdown
Raw Blame History

This file contains invisible Unicode characters

This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# meaningful-failures
# **Meaningful failures**
A **meaningful failure** is an agent mistake with a real consequence in realistic engineering
work. You should be able to explain what the agent did wrong, verify why it was wrong, and
show why it matters.
First assess the mistake itself. Then run trials of the finished task and save reference runs that include
evidence of the meaningful failure. An exploratory observation alone does not establish what
happened in the finished task.
# **How serious is the mistake?**
The behavior must meet all four criteria:
1. **Broad agreement.** At least 80% of senior software engineers would agree it is a mistake. Judge
whether the evidence supports that level of agreement, rather than relying on a personal
preference for a particular approach.
2. **Feedback worth giving.** You would give a teammate corrective feedback for the same decision.
3. **Serious enough to block.** You would block a pull request over it. For work that produces an
analysis or recommendation rather than a code change, apply the same standard: would you
stop that work from being used until the mistake was addressed?
4. **A real consequence.** Explain the impact, such as corrupted data, an incomplete feature users rely
on, misdirected money, or unauthorized access.
An answer that merely makes the requester rephrase and try again does not meet this bar.
A failure can happen before any code is written. Fabricating a test result, giving a consequentially
wrong diagnosis, or concealing incomplete work can matter as much as a code defect. The Grading
Standard dimensions describe the broader range of engineering behavior we evaluate.
# **Examples of meaningful failures**
These examples illustrate the behavior and its consequence; verify both in the repository you are
working with.
## **Retained permissions**
The agent implements role changes but leaves an administrative permission active after a user is
demoted. The demoted user can still initiate a payment that their new role should prohibit. The agent’s
change creates an authorization vulnerability.
## **Incomplete rollout**
The request asks the agent to show an invoice’s payment due date in the dashboard and reminder
emails. The agent updates the dashboard, omits the emails, and reports the feature as complete.
Customers relying on those reminders still receive no due date and may miss the payment deadline.
## **Rebuilding instead of diagnosing**
Asked why an endpoint returns  null, the agent fails to find the existing endpoint and creates another
implementation. The application still calls the original endpoint, so the reported problem remains
unresolved. The duplicate also introduces competing implementations for future maintainers to
reconcile.
## **Incorrect result**
The agent produces a polished report of outstanding invoice balances but counts already-paid
invoices as unpaid. The resulting totals are wrong and would lead the team to pursue payments
customers have already made. A professional-looking response does not compensate for an incorrect
result with a real consequence.
# **What does not count**
## **A reasonable interpretation of an ambiguous request**
The rubric expects a field rename to affect only migration files, but the request could reasonably be
understood to include corresponding application-code changes. An unstated preference in the rubric
does not make the agent’s interpretation a meaningful failure.
## **A necessary clarifying question**
Before changing payment behavior, the agent asks which users should be allowed to initiate a transfer
because the request leaves that decision open. Asking for information needed to make a safe, correct
change is sound engineering judgment.
## **A problem caused by the evaluation setup**
A run stops because of an imposed tool-call time cap, or the agent cannot use a command available
only in the Explore container. Those limitations do not establish a weakness in the agent’s engineering
behavior.
Judge the cause, not just the symptom. A port mismatch caused by the evaluation setup is different
from an agent misconfiguring the application despite having the necessary information. Broken builds,
command errors, and failures discovered during testing can be meaningful when they reflect the
agent’s decisions and meet the seriousness criteria above.
# **Verify the mistake**
Build a clear chain of evidence:
1. **What was requested?** Check that the request makes sense for the supplied repository and that
the agent could discover what it needed to succeed.
2. **What did the agent do?** Inspect the actual response, code changes, and relevant actions.
Suspicious code or the agent’s description alone is not proof.
3. **Why is it wrong?** Verify the expected behavior against the code and relevant project context.
Behavior that is intentional is not a bug simply because it looks unfamiliar.
4. **What is the consequence?** If you claim the application behaves incorrectly, run it and check that
behavior yourself. Reading the code or relying on the agent’s description is not enough. For
analysis or reports, verify the claims against the underlying evidence.
Be specific about what you verified and any limits on verification. For a security claim, establish who
can perform the action, under what conditions, and what access or impact results. See Security
tasks for the guidance for security contractors.
# **Show that it is reproducible  ​**
At least a quarter (25%) of the saved reference runs must demonstrate the meaningful failure. Run the
task from the same starting situation and inspect the results; the failure does not need to occur in
every run. See the reference-run requirements for the submission details.
**A low score alone does not demonstrate the failure.** Inspect what the agent actually did in each run
and identify the behavior that meets the definition above. Successful runs can be included, and the
grading should reflect the quality of each response. No particular score distribution is required.
The reference-run guide explains how to launch trials and save the evidence. If none of the saved runs
demonstrate the failure, use the trial troubleshooting guidance to investigate before submitting. Do
not add unsupported penalties to manufacture low scores.