lots of change - all to start my 3rd redo

This commit is contained in:
2026-09-26 14:31:52 -04:00
parent 7f4d388e19
commit bceb52e8ee
1046 changed files with 4476 additions and 0 deletions

View File

@@ -0,0 +1,80 @@
---
name: detector-over-hinting
description: |
Self-check whether your task package hints at the answer, on either of two
surfaces you author. (1) **Prompt over-hinting**: your `instruction.md` (or a
standing user turn) gives part of the answer away, or states things any
professional SWE would do unprompted ("be sure to add tests", "cleanly
separate the view logic from the db logic") — converting a judgment the task
could have measured into an instruction the agent merely follows. (2) **Hints
leaking through code comments**: files you add or edit via
`environment/workspace.patch` — often drafted with AI assistance — can carry
over-helpful comments that narrate the obvious, explain intent, or point
straight at the planted defect or the change you expect. Distinguishes
genuine task constraints ("add a retry with exponential backoff capped at
30s" — fine) from giveaways ("hint: the bug is in the retry loop" — not).
Advisory by design: flagged findings are passages to reconsider, not
failures. Reads instruction.md + workspace.patch (+ the resolved holistic
rubric as calibration context); runs before or after reference runs exist.
allowed-tools: Bash, Read, Write
---
# Over-hinting detector
This skill checks one of your tasks for **over-hinting** — places where the
package you author does the test agent's thinking for it, so the agent doesn't
have to exercise the judgment your holistic rubric scores. Two surfaces:
- **Your prompt.** The classic slips are directives any professional follows
unprompted — "be sure to add tests," "cleanly separate the view logic from
the db logic," "remember to handle edge cases" — and outright giveaways:
naming where the bug is, what the fix looks like, or the exact diligence
you're grading. A genuine requirement is different: "add a retry with
exponential backoff capped at 30s" defines *what to build*, like a real
ticket would. The line is whether the sentence pins down the deliverable or
shortcuts the noticing/finding/deciding that is the work.
- **Comments in files you add via `workspace.patch`.** Files drafted with AI
assistance often carry assistant-style comments that are too helpful:
tutorial headers, line-by-line narration, "NOTE: doesn't handle X yet"
sitting exactly on the issue your task plants. The test agent reads the
workspace — a comment that locates the defect or narrates the intended
change is a hint just like one in the prompt, only easier to miss when
packaging.
**This check is advisory.** Over-hinting is a judgment call — real requesters
do sometimes over-specify, and you may keep a hint deliberately. The report
exists so you can *consider hinting less*: each finding quotes the passage,
says what it pre-empts, and offers a concrete de-hinting option, so the
decision stays yours.
Read these before deciding:
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
2. `.claude/skills/detector-over-hinting/core.md` — the two surfaces, the requirement-vs-hint test, what is NOT a finding, verdict enums, and the body schema.
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
## Acting on the verdict
- **`clean`** — your prompt reads like a real request and your workspace
additions speak in-world; nothing pre-empts the graded judgment. Good. Move
on.
- **`partial-hinting`** — mild or borderline hints: SWE-obvious directives, a
directive that restates something your rubric genuinely requires, or
narrating comments away from the graded material. Read each finding and
decide: if the sentence isn't pinning down the deliverable, cut it and let
the rubric measure whether the agent does the professional thing unprompted.
If you keep one deliberately (e.g. your grader hard-requires tests and you
want that unambiguous), that's a legitimate call — the flag is just the
prompt to make it consciously.
- **`clear-hinting`** — something in your prompt or an authored comment points
substantially at the answer your rubric scores — the defect's location, the
expected fix or plan, or the exact diligence being measured. Take the
de-hinting option in each finding: delete the giveaway, move the fact into
your holistic rubric (which the agent never sees), or rewrite it as an
in-world constraint. Comments are usually the easy fix — strip the
over-helpful ones from your patch, keeping what an in-world engineer would
plausibly have written. Then regenerate reference runs if the hint was
live in the ones you have, and re-run this skill.
- **`not-applicable`** — there's no prompt or workspace patch to assess yet.
Draft them first.