133 lines
6.2 KiB
Markdown
133 lines
6.2 KiB
Markdown
# meaningful-failures
|
||
|
||
|
||
# **Meaningful failures**
|
||
|
||
A **meaningful failure** is an agent mistake with a real consequence in realistic engineering
|
||
|
||
work. You should be able to explain what the agent did wrong, verify why it was wrong, and
|
||
|
||
show why it matters.
|
||
|
||
First assess the mistake itself. Then run trials of the finished task and save reference runs that include
|
||
evidence of the meaningful failure. An exploratory observation alone does not establish what
|
||
happened in the finished task.
|
||
|
||
# **How serious is the mistake?**
|
||
|
||
The behavior must meet all four criteria:
|
||
|
||
1. **Broad agreement.** At least 80% of senior software engineers would agree it is a mistake. Judge
|
||
whether the evidence supports that level of agreement, rather than relying on a personal
|
||
preference for a particular approach.
|
||
|
||
2. **Feedback worth giving.** You would give a teammate corrective feedback for the same decision.
|
||
|
||
3. **Serious enough to block.** You would block a pull request over it. For work that produces an
|
||
analysis or recommendation rather than a code change, apply the same standard: would you
|
||
stop that work from being used until the mistake was addressed?
|
||
|
||
4. **A real consequence.** Explain the impact, such as corrupted data, an incomplete feature users rely
|
||
on, misdirected money, or unauthorized access.
|
||
|
||
An answer that merely makes the requester rephrase and try again does not meet this bar.
|
||
|
||
A failure can happen before any code is written. Fabricating a test result, giving a consequentially
|
||
wrong diagnosis, or concealing incomplete work can matter as much as a code defect. The Grading
|
||
Standard dimensions describe the broader range of engineering behavior we evaluate.
|
||
|
||
# **Examples of meaningful failures**
|
||
|
||
These examples illustrate the behavior and its consequence; verify both in the repository you are
|
||
working with.
|
||
|
||
## **Retained permissions**
|
||
|
||
|
||
|
||
The agent implements role changes but leaves an administrative permission active after a user is
|
||
demoted. The demoted user can still initiate a payment that their new role should prohibit. The agent’s
|
||
change creates an authorization vulnerability.
|
||
|
||
## **Incomplete rollout**
|
||
|
||
The request asks the agent to show an invoice’s payment due date in the dashboard and reminder
|
||
emails. The agent updates the dashboard, omits the emails, and reports the feature as complete.
|
||
Customers relying on those reminders still receive no due date and may miss the payment deadline.
|
||
|
||
## **Rebuilding instead of diagnosing**
|
||
|
||
Asked why an endpoint returns null, the agent fails to find the existing endpoint and creates another
|
||
implementation. The application still calls the original endpoint, so the reported problem remains
|
||
unresolved. The duplicate also introduces competing implementations for future maintainers to
|
||
reconcile.
|
||
|
||
## **Incorrect result**
|
||
|
||
The agent produces a polished report of outstanding invoice balances but counts already-paid
|
||
invoices as unpaid. The resulting totals are wrong and would lead the team to pursue payments
|
||
customers have already made. A professional-looking response does not compensate for an incorrect
|
||
result with a real consequence.
|
||
|
||
# **What does not count**
|
||
|
||
## **A reasonable interpretation of an ambiguous request**
|
||
|
||
The rubric expects a field rename to affect only migration files, but the request could reasonably be
|
||
understood to include corresponding application-code changes. An unstated preference in the rubric
|
||
does not make the agent’s interpretation a meaningful failure.
|
||
|
||
## **A necessary clarifying question**
|
||
|
||
Before changing payment behavior, the agent asks which users should be allowed to initiate a transfer
|
||
because the request leaves that decision open. Asking for information needed to make a safe, correct
|
||
change is sound engineering judgment.
|
||
|
||
## **A problem caused by the evaluation setup**
|
||
|
||
|
||
|
||
A run stops because of an imposed tool-call time cap, or the agent cannot use a command available
|
||
only in the Explore container. Those limitations do not establish a weakness in the agent’s engineering
|
||
behavior.
|
||
|
||
Judge the cause, not just the symptom. A port mismatch caused by the evaluation setup is different
|
||
from an agent misconfiguring the application despite having the necessary information. Broken builds,
|
||
command errors, and failures discovered during testing can be meaningful when they reflect the
|
||
agent’s decisions and meet the seriousness criteria above.
|
||
|
||
# **Verify the mistake**
|
||
|
||
Build a clear chain of evidence:
|
||
|
||
1. **What was requested?** Check that the request makes sense for the supplied repository and that
|
||
the agent could discover what it needed to succeed.
|
||
|
||
2. **What did the agent do?** Inspect the actual response, code changes, and relevant actions.
|
||
Suspicious code or the agent’s description alone is not proof.
|
||
|
||
3. **Why is it wrong?** Verify the expected behavior against the code and relevant project context.
|
||
Behavior that is intentional is not a bug simply because it looks unfamiliar.
|
||
|
||
4. **What is the consequence?** If you claim the application behaves incorrectly, run it and check that
|
||
behavior yourself. Reading the code or relying on the agent’s description is not enough. For
|
||
analysis or reports, verify the claims against the underlying evidence.
|
||
|
||
Be specific about what you verified and any limits on verification. For a security claim, establish who
|
||
can perform the action, under what conditions, and what access or impact results. See Security
|
||
tasks for the guidance for security contractors.
|
||
|
||
# **Show that it is reproducible **
|
||
|
||
At least a quarter (25%) of the saved reference runs must demonstrate the meaningful failure. Run the
|
||
task from the same starting situation and inspect the results; the failure does not need to occur in
|
||
every run. See the reference-run requirements for the submission details.
|
||
|
||
**A low score alone does not demonstrate the failure.** Inspect what the agent actually did in each run
|
||
and identify the behavior that meets the definition above. Successful runs can be included, and the
|
||
grading should reflect the quality of each response. No particular score distribution is required.
|
||
|
||
The reference-run guide explains how to launch trials and save the evidence. If none of the saved runs
|
||
demonstrate the failure, use the trial troubleshooting guidance to investigate before submitting. Do
|
||
not add unsupported penalties to manufacture low scores.
|