remove folder - untrustworthy
This commit is contained in:
@@ -1,65 +0,0 @@
|
||||
---
|
||||
name: detector-meaningful-failure
|
||||
description: |
|
||||
Self-check whether your task tests a real, proportionate, actually-elicited
|
||||
failure. Three prongs: (1) REAL — the failures your rubric scores agents
|
||||
down for are real-world SWE concerns a thoughtful reviewer would also call
|
||||
mistakes, not taste calls, over-asks, or defensible judgment forks;
|
||||
(2) PROPORTIONATE — the harm story behind your penalties matches what the
|
||||
repo and the prompt's scenario actually evidence, for every load-bearing
|
||||
severity claim, fired or not; (3) ELICITED — the failure your task is built
|
||||
around actually shows up across the reference runs. Run this after you have
|
||||
reference runs so the detector can read the grader's per-run reasoning.
|
||||
allowed-tools: Bash, Read, Write
|
||||
---
|
||||
|
||||
# Meaningful-failure detector
|
||||
|
||||
This skill checks one of your tasks against the three-prong quality bar:
|
||||
the rubric points at something *real* (a concrete SWE mistake, not nitty,
|
||||
subjective, or a defensible judgment call), the stakes it claims are
|
||||
*proportionate* (the harm story survives a check against the repo and the
|
||||
prompt's scenario), and the failure is actually *elicited* (it manifests
|
||||
across the reference runs — a task whose runs all score high with the
|
||||
central target never firing documents competent behavior instead of
|
||||
exposing a weakness). Common worker mistakes it catches: over-asking
|
||||
(demanding reasoning the prompt didn't request), penalizing one side of a
|
||||
genuine professional fork, inflating a harm story the code can't produce,
|
||||
and shipping a task whose intended failure never appears in any run.
|
||||
|
||||
**This detector needs reference runs.** Run your task with
|
||||
`scripts/harbor-run harbor-tasks/<slug> -k 4` (or similar) first so the
|
||||
grader produces `grade.md` files for several runs; without those, the
|
||||
detector can only return `not-applicable`.
|
||||
|
||||
Read these before deciding:
|
||||
|
||||
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
|
||||
2. `.claude/skills/detector-meaningful-failure/core.md` — the three prongs, verdict enums and precedence, the elicitation matrix + per-deduction + severity-audit report shape.
|
||||
|
||||
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
|
||||
|
||||
## Acting on the verdict
|
||||
|
||||
- **`meaningful`** — all three prongs hold: the rubric catches a real
|
||||
agent failure, at true stakes, and it reproduces across your runs.
|
||||
Good. Move on to the other detectors.
|
||||
- **`partial`** — one prong is diluted. Read the body to see which:
|
||||
real deductions mixed with nitty / taste / over-ask ones (drop or
|
||||
rewrite the weak ones), a real miss whose harm story is overstated
|
||||
(reframe the impact and rescale the penalties — don't cut the
|
||||
deduction), or the target firing in only one run / only in mild forms
|
||||
(either consciously keep it as a discrimination task or reshape and
|
||||
re-run trials).
|
||||
- **`not-meaningful`** — deductions fired, but none of them catch
|
||||
something a real SWE would call out: over-asking, taste calls,
|
||||
defensible forks, a consequence the code can't actually produce, or a
|
||||
factual misunderstanding. Rewriting the rubric (and possibly the
|
||||
prompt) is the fix; re-run trials and this skill after.
|
||||
- **`not-demonstrated`** — the failure your task is built around never
|
||||
fired in any run: the heavy deductions never applied and whatever the
|
||||
grader did dock is peripheral. The runs document competent behavior.
|
||||
Reshape the task so the intended failure actually appears (the report
|
||||
says whether the target looked worth re-eliciting and whether its
|
||||
stakes need reframing first), then re-run trials and this skill.
|
||||
- **`not-applicable`** — no reference runs yet. Run trials first.
|
||||
Reference in New Issue
Block a user