lots of change - all to start my 3rd redo

This commit is contained in:
2026-09-26 14:31:52 -04:00
parent 7f4d388e19
commit bceb52e8ee
1046 changed files with 4476 additions and 0 deletions

View File

@@ -0,0 +1,80 @@
---
name: detector-over-hinting
description: |
Self-check whether your task package hints at the answer, on either of two
surfaces you author. (1) **Prompt over-hinting**: your `instruction.md` (or a
standing user turn) gives part of the answer away, or states things any
professional SWE would do unprompted ("be sure to add tests", "cleanly
separate the view logic from the db logic") — converting a judgment the task
could have measured into an instruction the agent merely follows. (2) **Hints
leaking through code comments**: files you add or edit via
`environment/workspace.patch` — often drafted with AI assistance — can carry
over-helpful comments that narrate the obvious, explain intent, or point
straight at the planted defect or the change you expect. Distinguishes
genuine task constraints ("add a retry with exponential backoff capped at
30s" — fine) from giveaways ("hint: the bug is in the retry loop" — not).
Advisory by design: flagged findings are passages to reconsider, not
failures. Reads instruction.md + workspace.patch (+ the resolved holistic
rubric as calibration context); runs before or after reference runs exist.
allowed-tools: Bash, Read, Write
---
# Over-hinting detector
This skill checks one of your tasks for **over-hinting** — places where the
package you author does the test agent's thinking for it, so the agent doesn't
have to exercise the judgment your holistic rubric scores. Two surfaces:
- **Your prompt.** The classic slips are directives any professional follows
unprompted — "be sure to add tests," "cleanly separate the view logic from
the db logic," "remember to handle edge cases" — and outright giveaways:
naming where the bug is, what the fix looks like, or the exact diligence
you're grading. A genuine requirement is different: "add a retry with
exponential backoff capped at 30s" defines *what to build*, like a real
ticket would. The line is whether the sentence pins down the deliverable or
shortcuts the noticing/finding/deciding that is the work.
- **Comments in files you add via `workspace.patch`.** Files drafted with AI
assistance often carry assistant-style comments that are too helpful:
tutorial headers, line-by-line narration, "NOTE: doesn't handle X yet"
sitting exactly on the issue your task plants. The test agent reads the
workspace — a comment that locates the defect or narrates the intended
change is a hint just like one in the prompt, only easier to miss when
packaging.
**This check is advisory.** Over-hinting is a judgment call — real requesters
do sometimes over-specify, and you may keep a hint deliberately. The report
exists so you can *consider hinting less*: each finding quotes the passage,
says what it pre-empts, and offers a concrete de-hinting option, so the
decision stays yours.
Read these before deciding:
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
2. `.claude/skills/detector-over-hinting/core.md` — the two surfaces, the requirement-vs-hint test, what is NOT a finding, verdict enums, and the body schema.
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
## Acting on the verdict
- **`clean`** — your prompt reads like a real request and your workspace
additions speak in-world; nothing pre-empts the graded judgment. Good. Move
on.
- **`partial-hinting`** — mild or borderline hints: SWE-obvious directives, a
directive that restates something your rubric genuinely requires, or
narrating comments away from the graded material. Read each finding and
decide: if the sentence isn't pinning down the deliverable, cut it and let
the rubric measure whether the agent does the professional thing unprompted.
If you keep one deliberately (e.g. your grader hard-requires tests and you
want that unambiguous), that's a legitimate call — the flag is just the
prompt to make it consciously.
- **`clear-hinting`** — something in your prompt or an authored comment points
substantially at the answer your rubric scores — the defect's location, the
expected fix or plan, or the exact diligence being measured. Take the
de-hinting option in each finding: delete the giveaway, move the fact into
your holistic rubric (which the agent never sees), or rewrite it as an
in-world constraint. Comments are usually the easy fix — strip the
over-helpful ones from your patch, keeping what an in-world engineer would
plausibly have written. Then regenerate reference runs if the hint was
live in the ones you have, and re-run this skill.
- **`not-applicable`** — there's no prompt or workspace patch to assess yet.
Draft them first.

View File

@@ -0,0 +1,312 @@
# Over-hinting detector — core
This file is the canonical, context-neutral content for the detector-over-hinting
detector. It defines what the detector looks for, the two surfaces it inspects,
the verdict enums, the patterns to recognize, and the output schema. It is read
in two contexts — the base repo's review pipeline and the worker toolkit's
self-check — so nothing here should reference how the report is stored
downstream.
## What this detector is for
A task is only as good as the judgment it leaves to the agent under test. When
the task package *hints* at the answer — in the prompt, or in comments inside
files the task itself adds to the workspace — the agent doesn't have to
exercise the judgment the rubric scores; it just has to read carefully. The
task gets easier than the author intended, reference runs cluster high, and the
graded behavior stops discriminating.
Two hint surfaces, both authored by the task:
1. **Prompt over-hinting.** `instruction.md` (or the final user turn of a
multi-turn task) gives away part of the answer, or states things any
professional SWE would do unprompted. "Be sure to add tests" and "cleanly
separate the view logic from the db logic" are the canonical examples: a
thoughtful colleague adds tests for a behavior change and keeps view logic
out of the db layer without being told, so saying it in the prompt converts
a judgment the task could have measured into an instruction the agent merely
follows. The stronger form points at the answer itself: "hint: the bug is in
the retry loop," "you'll probably need to touch the serializer," "the fix is
a one-liner."
2. **Hints leaking through code comments.** Files added or modified by
`environment/workspace.patch` are often drafted with AI assistance, and
AI-generated comments tend to be too helpful: they narrate the obvious
line-by-line, explain intent a professional would infer from the code,
carry tutorial-style headers ("Step 1: validate the input"), or — worst —
point straight at the defect or the change the task expects ("NOTE: this
doesn't yet handle the negative-balance case"). The agent under test reads
the workspace; a comment that does its thinking for it is a hint exactly
like one in the prompt, just harder for the author to notice.
**This detector is advisory.** Over-hinting is a judgment call, and a hint is
rarely fatal on its own — the point of flagging is so the author can *consider
hinting less*, not to fail the task. A flagged verdict means "here are the
passages a de-hinting pass should look at," with each finding worded so the
author can decide for themselves whether the hint is doing damage. Realism cuts
both ways: real users do sometimes over-specify, and a task may deliberately
model that. The detector still surfaces the hint; whether to keep it is the
author's call.
## The controlling test: requirement vs. hint
For every candidate passage, ask two questions:
1. **Would a professional SWE have done this anyway, unprompted?** If yes, the
sentence is not conveying information — it's pre-empting a judgment the
task could have measured. "Add tests," "keep the layers separated," "handle
errors gracefully," "make sure it's backwards compatible" are all
SWE-obvious directives when nothing in the task makes them contested.
2. **Is this a genuine task constraint — something the request actually needs
to pin down — or a pointer toward the expected answer?** "Add a retry with
exponential backoff capped at 30s" is a genuine requirement: it defines
*what to build*, the way a real ticket would, and the rubric checks it.
"Hint: the bug is in the retry loop" defines nothing about the deliverable —
it exists only to shortcut the *finding*, which is the work.
A passage is a hint when it fails one of these — when it hands over judgment,
diligence, or discovery that the task is (or could be) measuring. It is a
legitimate constraint when a real, busy requester would plausibly say it to
define the deliverable, and the agent still has to figure out how to satisfy
it.
Use the grader guidance as calibration context, not as a hint surface
(the agent under test never sees it): a directive that restates something the
rubric genuinely requires and checks sits in the borderline zone. "Be sure to
add tests" when the rubric's correctness signal really does hinge on tests is
still worth flagging — the author could let the rubric measure whether the
agent adds them unprompted — but flag it at LOW confidence and say so; the
author may have good reasons to pin it down.
## Inputs
Read from `harbor-tasks/<slug>/`:
- `instruction.md` — the prompt the agent under test receives. The primary
prompt surface. For a snapshot / multi-turn task, also read the load-bearing
user turns in the session history (`environment/session.jsonl` or
`session-full.jsonl`) — a hint in a standing instruction reaches the agent
the same way.
- `environment/workspace.patch` — the diff of files the task adds to or edits
in the workspace. Scan the added/modified lines for hint-bearing comments,
docstrings, TODO/NOTE/FIXME markers, and freshly-authored README/doc prose.
You don't need to build the workspace; the patch text is the surface. The
workspace's pre-existing repo content is out of scope — only the task's own
additions are.
- The grader guidance — calibration context only (see above): what does
the rubric actually score and require? A prompt sentence that merely
restates a graded requirement is borderline, not a clear hint. Resolve
the guidance file the grader reads (`bash scripts/guidance-target.sh
<slug>` prints its path, `tests/grader-guidance-consolidated.md` — the worker shell's
guidance-target resolution) and calibrate against the file it names,
never another document.
- `reference-runs/*/grade.md` — not required, but a useful cross-check when
present: runs that uniformly sail through the intended difficulty, or a
grade that quotes an authored comment as the reason the agent found the
answer, corroborate that a hint is live.
## Patterns to look for
**Prompt surface (`instruction.md` / user turns):**
- **SWE-obvious directives** — instructions any professional would follow
unprompted: "be sure to add tests," "cleanly separate the view logic from
the db logic," "write clean, maintainable code," "remember to handle edge
cases," "don't break existing functionality."
- **Answer giveaways** — the prompt names the defect's location, mechanism, or
fix: "the bug is in X," "check the retry loop," "it's probably a race
condition," "you'll need to update the serializer too."
- **Pre-announced diligence** — the prompt names the exact judgment or
verification the task is meant to measure: "double-check the timezone
handling" on a task whose intended failure is a timezone bug; "make sure the
migration is reversible" when reversibility is the graded catch.
- **Difficulty disclaimers that orient the search** — "this is trickier than
it looks," "the obvious approach won't work here" attached to the specific
place where the intended difficulty lives.
**Workspace-comment surface (`environment/workspace.patch` additions):**
- **Defect pointers** — a comment adjacent to the planted issue that names or
gestures at it: "NOTE: doesn't handle concurrent updates yet," "FIXME:
validation is incomplete," "this assumes the list is sorted" placed exactly
where the assumption breaks.
- **Solution narration** — comments that explain what a change *should* do or
what the next step is, effectively writing the agent's plan: "Step 1:
fetch…, Step 2: validate…," "eventually this should delegate to the
BillingService."
- **Intent narration of the obvious** — line-by-line commentary a professional
would never write ("// increment the counter," "// return the result"), or
a tutorial-style header block that summarizes the file's mechanism in a way
the task expects the agent to work out by reading the code.
- **Assessor's-eye framing** — a comment or doc that describes the file from
outside the scenario ("this is where the interesting part is," "the
important method is below") rather than as something an in-world engineer
would leave.
## What is NOT a finding
- **Genuine requirements and acceptance criteria.** "Add a retry with
exponential backoff capped at 30s," "the endpoint must stay
backwards-compatible with v1 clients," "use the existing PDF pipeline" —
specific asks that define the deliverable, which the rubric checks, are the
task, not hints. Precision about *what to build* is good authoring; the
detector fires on giveaways about *what the agent is supposed to notice,
decide, or find*.
- **Domain context the agent genuinely needs.** Business constraints, in-world
background, a ticket's reproduction steps, what the requester already tried.
A busy user explaining their situation is realism, not hinting — even when
it's detailed.
- **A false or contested premise stated in the prompt.** Tasks legitimately
model a requester who believes something wrong; the prompt asserting that
belief is the scenario, not a hint (the hint would be the prompt *also*
flagging that the belief is wrong).
- **Pre-existing repo comments.** Comments that come from the source repo
unmodified are the codebase the agent must cope with; only the
`workspace.patch` additions/edits are in scope.
- **In-world artifacts that carry the scenario.** A planted TODO or draft doc
can *be* the task's subject (e.g. the task is about an unfinished feature
the TODO marks). The flag condition is a comment that does the agent's
thinking — locates the defect, prescribes the change, or narrates the
judgment being graded — not that an authored comment exists. Ask: would an
in-world engineer plausibly have left this, and does the graded difficulty
survive it?
- **Comments matching the repo's existing style.** Doc headers, license
blocks, docstrings on public APIs — additions that mirror how the codebase
already comments are craft, not hints.
- **Detail level alone.** A long, thorough prompt is not over-hinted; a
two-line prompt can be. The measure is whether graded judgment survives the
text, not how much text there is.
## Verdict definitions
- **`clean`** — neither the prompt nor the authored workspace additions hint
at the answer or pre-empt SWE-obvious judgment. Genuine requirements,
domain context, and in-world artifacts are all clean (see the list above).
- **`partial-hinting`** — mild or borderline hinting worth the author's
attention: SWE-obvious directives ("be sure to add tests," "separate the
view logic from the db logic"), a directive that restates a genuinely graded
requirement, narrating-the-obvious comments away from the graded material,
or a passage you can read either as scenario realism or as a nudge. The task
still works; a de-hinting pass would sharpen it.
- **`clear-hinting`** — the prompt or an authored comment points substantially
at the answer the rubric scores: names the defect or its location,
prescribes the graded fix or plan, pre-announces the exact diligence being
measured, or a workspace comment sits on the planted issue and describes it.
The intended difficulty is materially reduced for any agent that reads
carefully.
- **`not-applicable`** — nothing to assess: `instruction.md` is missing,
empty, or only template/placeholder content, and there is no session
history or `workspace.patch` to inspect. Re-run once the prompt lands.
`clear-hinting` and `partial-hinting` are the flagged outcomes; `clean` and
`not-applicable` are not. All flagged outcomes are advisory: they hand the
author a list of passages to reconsider, and the author may keep any of them
deliberately.
## Confidence
- **HIGH** — the call is unambiguous: a passage plainly gives the answer away
(or plainly nothing does), and the graded behavior is clear from the rubric.
- **MEDIUM** — at least one finding is genuinely two-sided: a reasonable
reviewer might read the passage as legitimate constraint or scenario
realism.
- **LOW** — limited information: the rubric is thin or absent so you can't
tell what's graded, or the directive overlaps a genuine graded requirement
("be sure to add tests" when the rubric requires tests) and the call is the
author's to make.
## Relationship to other detectors
- **vs. detector-answer-obviousness.** Its over-cued shape owns the *fatal*
end of prompt cueing: the prompt names the exact graded behavior, so the
task cannot discriminate at all — a defect verdict about task validity.
This detector owns the *gradient below that*: hint-shaped prose worth a
de-hinting pass whether or not it fully disarms the task (SWE-obvious
directives, partial giveaways, difficulty disclaimers), plus the
workspace-comment surface, which answer-obviousness doesn't read. When a
prompt hint is total, expect both to fire — answer-obviousness on validity,
this detector on the concrete passages to rewrite.
- **vs. detector-snapshot-leakage.** Leakage owns what the agent *inherits*
from a captured session — conversation context that hands over the answer.
This detector owns what the task *authors*: the prompt text and the
comments inside `workspace.patch` additions. A workspace file that leaks
the rubric's answer outright can fire both; each flags its own surface.
- **vs. detector-cross-task-reference.** That detector scans the same
authored surfaces for a different defect (pointers to sibling tasks). A
comment can be both a sibling reference and a hint; the verdicts are
independent.
## Anti-patterns: do not do these
- **Don't flag specificity.** A precise, well-specified ask is good
authoring. The finding is a giveaway about the *graded* judgment,
discovery, or diligence — not detail about the deliverable.
- **Don't flag the scenario's own material.** False premises, planted
in-world TODOs, and requester context are the task. Re-read the "What is
NOT a finding" list before flagging anything in that family.
- **Don't demand a hint be fatal before flagging.** This detector is
advisory by design; a mild SWE-obvious directive is a legitimate
`partial-hinting` finding even though the task still works.
- **Don't treat the flag as a failure either.** Word every finding so the
author can weigh it — quote the passage, say what it pre-empts, and offer
the de-hinted alternative. Never assert the task is broken because a hint
exists.
- **Don't hunt hints in the grader guidance.** The agent under test
never sees it; it's calibration context for you, not a surface.
- **Don't cite evidence you haven't verified in the submitted package.**
Quote passages as they exist in the actual `instruction.md`, session
history, and `workspace.patch` — not as you remember them or as the rubric
paraphrases them.
## Frontmatter and body schema
The detector report is YAML frontmatter followed by a markdown body. Both
contexts produce the same shape; only the *sink* differs (the wrapping
`SKILL.md` tells you where to send the report).
**Frontmatter** — exactly these keys, exactly these enum values:
```yaml
---
detector: detector-over-hinting
verdict: clear-hinting | partial-hinting | clean | not-applicable
confidence: HIGH | MEDIUM | LOW
---
```
**Body sections**, in this order:
```markdown
# Over-hinting check: <slug>
## Findings
One block per finding, strongest first:
### <short label> — <prompt-hint | comment-hint> (<clear | partial>)
- **Where:** the file (and line/section for a patch hunk, or the turn for a
session message).
- **Quote:** the passage verbatim, as a blockquote — never a paraphrase.
- **What it pre-empts:** one or two sentences — the judgment, discovery, or
diligence the passage hands over, tied to what the rubric grades where
possible.
- **De-hinting option:** a concrete alternative — delete the sentence, move
the fact into grader guidance, rewrite the directive as an in-world
constraint — worded so the author can decide whether to take it.
For `clean`, quote the strongest near-miss (a specific requirement, a planted
in-world TODO, a detailed prompt) and say why it was cleared. For
`not-applicable`, name the missing artifacts.
## Overall verdict
2–3 paragraphs reducing the findings to the chosen verdict: which surface(s)
hint and how strongly, whether the graded difficulty survives, and — because
this detector is advisory — what a de-hinting pass would change first. For
`clean`, why the near-misses are constraints or scenario material rather than
hints.
```
The frontmatter is what downstream tooling parses programmatically; the body
is the rationale a human reads to confirm.