4.6 KiB
4.6 KiB
name, description, allowed-tools
| name | description | allowed-tools |
|---|---|---|
| detector-over-hinting | Self-check whether your task package hints at the answer, on either of two surfaces you author. (1) **Prompt over-hinting**: your `instruction.md` (or a standing user turn) gives part of the answer away, or states things any professional SWE would do unprompted ("be sure to add tests", "cleanly separate the view logic from the db logic") — converting a judgment the task could have measured into an instruction the agent merely follows. (2) **Hints leaking through code comments**: files you add or edit via `environment/workspace.patch` — often drafted with AI assistance — can carry over-helpful comments that narrate the obvious, explain intent, or point straight at the planted defect or the change you expect. Distinguishes genuine task constraints ("add a retry with exponential backoff capped at 30s" — fine) from giveaways ("hint: the bug is in the retry loop" — not). Advisory by design: flagged findings are passages to reconsider, not failures. Reads instruction.md + workspace.patch (+ the resolved grader guidance as calibration context); runs before or after reference runs exist. | Bash, Read, Write |
Over-hinting detector
This skill checks one of your tasks for over-hinting — places where the package you author does the test agent's thinking for it, so the agent doesn't have to exercise the judgment your grader guidance scores. Two surfaces:
- Your prompt. The classic slips are directives any professional follows unprompted — "be sure to add tests," "cleanly separate the view logic from the db logic," "remember to handle edge cases" — and outright giveaways: naming where the bug is, what the fix looks like, or the exact diligence you're grading. A genuine requirement is different: "add a retry with exponential backoff capped at 30s" defines what to build, like a real ticket would. The line is whether the sentence pins down the deliverable or shortcuts the noticing/finding/deciding that is the work.
- Comments in files you add via
workspace.patch. Files drafted with AI assistance often carry assistant-style comments that are too helpful: tutorial headers, line-by-line narration, "NOTE: doesn't handle X yet" sitting exactly on the issue your task plants. The test agent reads the workspace — a comment that locates the defect or narrates the intended change is a hint just like one in the prompt, only easier to miss when packaging.
This check is advisory. Over-hinting is a judgment call — real requesters do sometimes over-specify, and you may keep a hint deliberately. The report exists so you can consider hinting less: each finding quotes the passage, says what it pre-empts, and offers a concrete de-hinting option, so the decision stays yours.
Read these before deciding:
.claude/skills/_detector-worker-shell.md— where to write the report and how to handle re-runs..claude/skills/detector-over-hinting/core.md— the two surfaces, the requirement-vs-hint test, what is NOT a finding, verdict enums, and the body schema.
Compose the report per the schema in core.md and write it per _detector-worker-shell.md.
Acting on the verdict
clean— your prompt reads like a real request and your workspace additions speak in-world; nothing pre-empts the graded judgment. Good. Move on.partial-hinting— mild or borderline hints: SWE-obvious directives, a directive that restates something your rubric genuinely requires, or narrating comments away from the graded material. Read each finding and decide: if the sentence isn't pinning down the deliverable, cut it and let the rubric measure whether the agent does the professional thing unprompted. If you keep one deliberately (e.g. your grader hard-requires tests and you want that unambiguous), that's a legitimate call — the flag is just the prompt to make it consciously.clear-hinting— something in your prompt or an authored comment points substantially at the answer your rubric scores — the defect's location, the expected fix or plan, or the exact diligence being measured. Take the de-hinting option in each finding: delete the giveaway, move the fact into your grader guidance (which the agent never sees), or rewrite it as an in-world constraint. Comments are usually the easy fix — strip the over-helpful ones from your patch, keeping what an in-world engineer would plausibly have written. Then regenerate reference runs if the hint was live in the ones you have, and re-run this skill.not-applicable— there's no prompt or workspace patch to assess yet. Draft them first.