chore: init commit

in worker.../repo/GITFOLDER.zip is the .git folder.
This commit is contained in:
2026-08-11 14:44:09 -04:00
parent 0012380fd3
commit 2854619bc9
782 changed files with 65944 additions and 0 deletions

View File

@@ -0,0 +1,65 @@
---
name: detector-meaningful-failure
description: |
Self-check whether your task tests a real, proportionate, actually-elicited
failure. Three prongs: (1) REAL — the failures your rubric scores agents
down for are real-world SWE concerns a thoughtful reviewer would also call
mistakes, not taste calls, over-asks, or defensible judgment forks;
(2) PROPORTIONATE — the harm story behind your penalties matches what the
repo and the prompt's scenario actually evidence, for every load-bearing
severity claim, fired or not; (3) ELICITED — the failure your task is built
around actually shows up across the reference runs. Run this after you have
reference runs so the detector can read the grader's per-run reasoning.
allowed-tools: Bash, Read, Write
---
# Meaningful-failure detector
This skill checks one of your tasks against the three-prong quality bar:
the rubric points at something *real* (a concrete SWE mistake, not nitty,
subjective, or a defensible judgment call), the stakes it claims are
*proportionate* (the harm story survives a check against the repo and the
prompt's scenario), and the failure is actually *elicited* (it manifests
across the reference runs — a task whose runs all score high with the
central target never firing documents competent behavior instead of
exposing a weakness). Common worker mistakes it catches: over-asking
(demanding reasoning the prompt didn't request), penalizing one side of a
genuine professional fork, inflating a harm story the code can't produce,
and shipping a task whose intended failure never appears in any run.
**This detector needs reference runs.** Run your task with
`scripts/harbor-run harbor-tasks/<slug> -k 4` (or similar) first so the
grader produces `grade.md` files for several runs; without those, the
detector can only return `not-applicable`.
Read these before deciding:
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
2. `.claude/skills/detector-meaningful-failure/core.md` — the three prongs, verdict enums and precedence, the elicitation matrix + per-deduction + severity-audit report shape.
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
## Acting on the verdict
- **`meaningful`** — all three prongs hold: the rubric catches a real
agent failure, at true stakes, and it reproduces across your runs.
Good. Move on to the other detectors.
- **`partial`** — one prong is diluted. Read the body to see which:
real deductions mixed with nitty / taste / over-ask ones (drop or
rewrite the weak ones), a real miss whose harm story is overstated
(reframe the impact and rescale the penalties — don't cut the
deduction), or the target firing in only one run / only in mild forms
(either consciously keep it as a discrimination task or reshape and
re-run trials).
- **`not-meaningful`** — deductions fired, but none of them catch
something a real SWE would call out: over-asking, taste calls,
defensible forks, a consequence the code can't actually produce, or a
factual misunderstanding. Rewriting the rubric (and possibly the
prompt) is the fix; re-run trials and this skill after.
- **`not-demonstrated`** — the failure your task is built around never
fired in any run: the heavy deductions never applied and whatever the
grader did dock is peripheral. The runs document competent behavior.
Reshape the task so the intended failure actually appears (the report
says whether the target looked worth re-eliciting and whether its
stakes need reframing first), then re-run trials and this skill.
- **`not-applicable`** — no reference runs yet. Run trials first.