Files
project-work/worker-toolkit-flaredown/.claude/skills/detector-meaningful-failure/SKILL.md

3.6 KiB

name, description, allowed-tools
name description allowed-tools
detector-meaningful-failure Self-check whether your task tests a real, proportionate, actually-elicited failure. Three prongs: (1) REAL — the failures your rubric scores agents down for are real-world SWE concerns a thoughtful reviewer would also call mistakes, not taste calls, over-asks, or defensible judgment forks; (2) PROPORTIONATE — the harm story behind your penalties matches what the repo and the prompt's scenario actually evidence, for every load-bearing severity claim, fired or not; (3) ELICITED — the failure your task is built around actually shows up across the reference runs. Run this after you have reference runs so the detector can read the grader's per-run reasoning. Bash, Read, Write

Meaningful-failure detector

This skill checks one of your tasks against the three-prong quality bar: the rubric points at something real (a concrete SWE mistake, not nitty, subjective, or a defensible judgment call), the stakes it claims are proportionate (the harm story survives a check against the repo and the prompt's scenario), and the failure is actually elicited (it manifests across the reference runs — a task whose runs all score high with the central target never firing documents competent behavior instead of exposing a weakness). Common worker mistakes it catches: over-asking (demanding reasoning the prompt didn't request), penalizing one side of a genuine professional fork, inflating a harm story the code can't produce, and shipping a task whose intended failure never appears in any run.

This detector needs reference runs. Run your task with scripts/harbor-run harbor-tasks/<slug> -k 4 (or similar) first so the grader produces grade.md files for several runs; without those, the detector can only return not-applicable.

Read these before deciding:

  1. .claude/skills/_detector-worker-shell.md — where to write the report and how to handle re-runs.
  2. .claude/skills/detector-meaningful-failure/core.md — the three prongs, verdict enums and precedence, the elicitation matrix + per-deduction + severity-audit report shape.

Compose the report per the schema in core.md and write it per _detector-worker-shell.md.

Acting on the verdict

  • meaningful — all three prongs hold: the rubric catches a real agent failure, at true stakes, and it reproduces across your runs. Good. Move on to the other detectors.
  • partial — one prong is diluted. Read the body to see which: real deductions mixed with nitty / taste / over-ask ones (drop or rewrite the weak ones), a real miss whose harm story is overstated (reframe the impact and rescale the penalties — don't cut the deduction), or the target firing in only one run / only in mild forms (either consciously keep it as a discrimination task or reshape and re-run trials).
  • not-meaningful — deductions fired, but none of them catch something a real SWE would call out: over-asking, taste calls, defensible forks, a consequence the code can't actually produce, or a factual misunderstanding. Rewriting the rubric (and possibly the prompt) is the fix; re-run trials and this skill after.
  • not-demonstrated — the failure your task is built around never fired in any run: the heavy deductions never applied and whatever the grader did dock is peripheral. The runs document competent behavior. Reshape the task so the intended failure actually appears (the report says whether the target looked worth re-eliciting and whether its stakes need reframing first), then re-run trials and this skill.
  • not-applicable — no reference runs yet. Run trials first.