added 260907 version of worker toolkit
This commit is contained in:
@@ -0,0 +1,80 @@
|
||||
---
|
||||
name: detector-over-hinting
|
||||
description: |
|
||||
Self-check whether your task package hints at the answer, on either of two
|
||||
surfaces you author. (1) **Prompt over-hinting**: your `instruction.md` (or a
|
||||
standing user turn) gives part of the answer away, or states things any
|
||||
professional SWE would do unprompted ("be sure to add tests", "cleanly
|
||||
separate the view logic from the db logic") — converting a judgment the task
|
||||
could have measured into an instruction the agent merely follows. (2) **Hints
|
||||
leaking through code comments**: files you add or edit via
|
||||
`environment/workspace.patch` — often drafted with AI assistance — can carry
|
||||
over-helpful comments that narrate the obvious, explain intent, or point
|
||||
straight at the planted defect or the change you expect. Distinguishes
|
||||
genuine task constraints ("add a retry with exponential backoff capped at
|
||||
30s" — fine) from giveaways ("hint: the bug is in the retry loop" — not).
|
||||
Advisory by design: flagged findings are passages to reconsider, not
|
||||
failures. Reads instruction.md + workspace.patch (+ the resolved holistic
|
||||
rubric as calibration context); runs before or after reference runs exist.
|
||||
allowed-tools: Bash, Read, Write
|
||||
---
|
||||
|
||||
# Over-hinting detector
|
||||
|
||||
This skill checks one of your tasks for **over-hinting** — places where the
|
||||
package you author does the test agent's thinking for it, so the agent doesn't
|
||||
have to exercise the judgment your holistic rubric scores. Two surfaces:
|
||||
|
||||
- **Your prompt.** The classic slips are directives any professional follows
|
||||
unprompted — "be sure to add tests," "cleanly separate the view logic from
|
||||
the db logic," "remember to handle edge cases" — and outright giveaways:
|
||||
naming where the bug is, what the fix looks like, or the exact diligence
|
||||
you're grading. A genuine requirement is different: "add a retry with
|
||||
exponential backoff capped at 30s" defines *what to build*, like a real
|
||||
ticket would. The line is whether the sentence pins down the deliverable or
|
||||
shortcuts the noticing/finding/deciding that is the work.
|
||||
- **Comments in files you add via `workspace.patch`.** Files drafted with AI
|
||||
assistance often carry assistant-style comments that are too helpful:
|
||||
tutorial headers, line-by-line narration, "NOTE: doesn't handle X yet"
|
||||
sitting exactly on the issue your task plants. The test agent reads the
|
||||
workspace — a comment that locates the defect or narrates the intended
|
||||
change is a hint just like one in the prompt, only easier to miss when
|
||||
packaging.
|
||||
|
||||
**This check is advisory.** Over-hinting is a judgment call — real requesters
|
||||
do sometimes over-specify, and you may keep a hint deliberately. The report
|
||||
exists so you can *consider hinting less*: each finding quotes the passage,
|
||||
says what it pre-empts, and offers a concrete de-hinting option, so the
|
||||
decision stays yours.
|
||||
|
||||
Read these before deciding:
|
||||
|
||||
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
|
||||
2. `.claude/skills/detector-over-hinting/core.md` — the two surfaces, the requirement-vs-hint test, what is NOT a finding, verdict enums, and the body schema.
|
||||
|
||||
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
|
||||
|
||||
## Acting on the verdict
|
||||
|
||||
- **`clean`** — your prompt reads like a real request and your workspace
|
||||
additions speak in-world; nothing pre-empts the graded judgment. Good. Move
|
||||
on.
|
||||
- **`partial-hinting`** — mild or borderline hints: SWE-obvious directives, a
|
||||
directive that restates something your rubric genuinely requires, or
|
||||
narrating comments away from the graded material. Read each finding and
|
||||
decide: if the sentence isn't pinning down the deliverable, cut it and let
|
||||
the rubric measure whether the agent does the professional thing unprompted.
|
||||
If you keep one deliberately (e.g. your grader hard-requires tests and you
|
||||
want that unambiguous), that's a legitimate call — the flag is just the
|
||||
prompt to make it consciously.
|
||||
- **`clear-hinting`** — something in your prompt or an authored comment points
|
||||
substantially at the answer your rubric scores — the defect's location, the
|
||||
expected fix or plan, or the exact diligence being measured. Take the
|
||||
de-hinting option in each finding: delete the giveaway, move the fact into
|
||||
your holistic rubric (which the agent never sees), or rewrite it as an
|
||||
in-world constraint. Comments are usually the easy fix — strip the
|
||||
over-helpful ones from your patch, keeping what an in-world engineer would
|
||||
plausibly have written. Then regenerate reference runs if the hint was
|
||||
live in the ones you have, and re-run this skill.
|
||||
- **`not-applicable`** — there's no prompt or workspace patch to assess yet.
|
||||
Draft them first.
|
||||
@@ -0,0 +1,312 @@
|
||||
# Over-hinting detector — core
|
||||
|
||||
This file is the canonical, context-neutral content for the detector-over-hinting
|
||||
detector. It defines what the detector looks for, the two surfaces it inspects,
|
||||
the verdict enums, the patterns to recognize, and the output schema. It is read
|
||||
in two contexts — the base repo's review pipeline and the worker toolkit's
|
||||
self-check — so nothing here should reference how the report is stored
|
||||
downstream.
|
||||
|
||||
## What this detector is for
|
||||
|
||||
A task is only as good as the judgment it leaves to the agent under test. When
|
||||
the task package *hints* at the answer — in the prompt, or in comments inside
|
||||
files the task itself adds to the workspace — the agent doesn't have to
|
||||
exercise the judgment the rubric scores; it just has to read carefully. The
|
||||
task gets easier than the author intended, reference runs cluster high, and the
|
||||
graded behavior stops discriminating.
|
||||
|
||||
Two hint surfaces, both authored by the task:
|
||||
|
||||
1. **Prompt over-hinting.** `instruction.md` (or the final user turn of a
|
||||
multi-turn task) gives away part of the answer, or states things any
|
||||
professional SWE would do unprompted. "Be sure to add tests" and "cleanly
|
||||
separate the view logic from the db logic" are the canonical examples: a
|
||||
thoughtful colleague adds tests for a behavior change and keeps view logic
|
||||
out of the db layer without being told, so saying it in the prompt converts
|
||||
a judgment the task could have measured into an instruction the agent merely
|
||||
follows. The stronger form points at the answer itself: "hint: the bug is in
|
||||
the retry loop," "you'll probably need to touch the serializer," "the fix is
|
||||
a one-liner."
|
||||
|
||||
2. **Hints leaking through code comments.** Files added or modified by
|
||||
`environment/workspace.patch` are often drafted with AI assistance, and
|
||||
AI-generated comments tend to be too helpful: they narrate the obvious
|
||||
line-by-line, explain intent a professional would infer from the code,
|
||||
carry tutorial-style headers ("Step 1: validate the input"), or — worst —
|
||||
point straight at the defect or the change the task expects ("NOTE: this
|
||||
doesn't yet handle the negative-balance case"). The agent under test reads
|
||||
the workspace; a comment that does its thinking for it is a hint exactly
|
||||
like one in the prompt, just harder for the author to notice.
|
||||
|
||||
**This detector is advisory.** Over-hinting is a judgment call, and a hint is
|
||||
rarely fatal on its own — the point of flagging is so the author can *consider
|
||||
hinting less*, not to fail the task. A flagged verdict means "here are the
|
||||
passages a de-hinting pass should look at," with each finding worded so the
|
||||
author can decide for themselves whether the hint is doing damage. Realism cuts
|
||||
both ways: real users do sometimes over-specify, and a task may deliberately
|
||||
model that. The detector still surfaces the hint; whether to keep it is the
|
||||
author's call.
|
||||
|
||||
## The controlling test: requirement vs. hint
|
||||
|
||||
For every candidate passage, ask two questions:
|
||||
|
||||
1. **Would a professional SWE have done this anyway, unprompted?** If yes, the
|
||||
sentence is not conveying information — it's pre-empting a judgment the
|
||||
task could have measured. "Add tests," "keep the layers separated," "handle
|
||||
errors gracefully," "make sure it's backwards compatible" are all
|
||||
SWE-obvious directives when nothing in the task makes them contested.
|
||||
2. **Is this a genuine task constraint — something the request actually needs
|
||||
to pin down — or a pointer toward the expected answer?** "Add a retry with
|
||||
exponential backoff capped at 30s" is a genuine requirement: it defines
|
||||
*what to build*, the way a real ticket would, and the rubric checks it.
|
||||
"Hint: the bug is in the retry loop" defines nothing about the deliverable —
|
||||
it exists only to shortcut the *finding*, which is the work.
|
||||
|
||||
A passage is a hint when it fails one of these — when it hands over judgment,
|
||||
diligence, or discovery that the task is (or could be) measuring. It is a
|
||||
legitimate constraint when a real, busy requester would plausibly say it to
|
||||
define the deliverable, and the agent still has to figure out how to satisfy
|
||||
it.
|
||||
|
||||
Use the grader guidance as calibration context, not as a hint surface
|
||||
(the agent under test never sees it): a directive that restates something the
|
||||
rubric genuinely requires and checks sits in the borderline zone. "Be sure to
|
||||
add tests" when the rubric's correctness signal really does hinge on tests is
|
||||
still worth flagging — the author could let the rubric measure whether the
|
||||
agent adds them unprompted — but flag it at LOW confidence and say so; the
|
||||
author may have good reasons to pin it down.
|
||||
|
||||
## Inputs
|
||||
|
||||
Read from `harbor-tasks/<slug>/`:
|
||||
|
||||
- `instruction.md` — the prompt the agent under test receives. The primary
|
||||
prompt surface. For a snapshot / multi-turn task, also read the load-bearing
|
||||
user turns in the session history (`environment/session.jsonl` or
|
||||
`session-full.jsonl`) — a hint in a standing instruction reaches the agent
|
||||
the same way.
|
||||
- `environment/workspace.patch` — the diff of files the task adds to or edits
|
||||
in the workspace. Scan the added/modified lines for hint-bearing comments,
|
||||
docstrings, TODO/NOTE/FIXME markers, and freshly-authored README/doc prose.
|
||||
You don't need to build the workspace; the patch text is the surface. The
|
||||
workspace's pre-existing repo content is out of scope — only the task's own
|
||||
additions are.
|
||||
- The grader guidance — calibration context only (see above): what does
|
||||
the rubric actually score and require? A prompt sentence that merely
|
||||
restates a graded requirement is borderline, not a clear hint. Resolve
|
||||
the guidance file the grader reads (`bash scripts/guidance-target.sh
|
||||
<slug>` prints its path, `tests/grader-guidance-consolidated.md` — the worker shell's
|
||||
guidance-target resolution) and calibrate against the file it names,
|
||||
never another document.
|
||||
- `reference-runs/*/grade.md` — not required, but a useful cross-check when
|
||||
present: runs that uniformly sail through the intended difficulty, or a
|
||||
grade that quotes an authored comment as the reason the agent found the
|
||||
answer, corroborate that a hint is live.
|
||||
|
||||
## Patterns to look for
|
||||
|
||||
**Prompt surface (`instruction.md` / user turns):**
|
||||
|
||||
- **SWE-obvious directives** — instructions any professional would follow
|
||||
unprompted: "be sure to add tests," "cleanly separate the view logic from
|
||||
the db logic," "write clean, maintainable code," "remember to handle edge
|
||||
cases," "don't break existing functionality."
|
||||
- **Answer giveaways** — the prompt names the defect's location, mechanism, or
|
||||
fix: "the bug is in X," "check the retry loop," "it's probably a race
|
||||
condition," "you'll need to update the serializer too."
|
||||
- **Pre-announced diligence** — the prompt names the exact judgment or
|
||||
verification the task is meant to measure: "double-check the timezone
|
||||
handling" on a task whose intended failure is a timezone bug; "make sure the
|
||||
migration is reversible" when reversibility is the graded catch.
|
||||
- **Difficulty disclaimers that orient the search** — "this is trickier than
|
||||
it looks," "the obvious approach won't work here" attached to the specific
|
||||
place where the intended difficulty lives.
|
||||
|
||||
**Workspace-comment surface (`environment/workspace.patch` additions):**
|
||||
|
||||
- **Defect pointers** — a comment adjacent to the planted issue that names or
|
||||
gestures at it: "NOTE: doesn't handle concurrent updates yet," "FIXME:
|
||||
validation is incomplete," "this assumes the list is sorted" placed exactly
|
||||
where the assumption breaks.
|
||||
- **Solution narration** — comments that explain what a change *should* do or
|
||||
what the next step is, effectively writing the agent's plan: "Step 1:
|
||||
fetch…, Step 2: validate…," "eventually this should delegate to the
|
||||
BillingService."
|
||||
- **Intent narration of the obvious** — line-by-line commentary a professional
|
||||
would never write ("// increment the counter," "// return the result"), or
|
||||
a tutorial-style header block that summarizes the file's mechanism in a way
|
||||
the task expects the agent to work out by reading the code.
|
||||
- **Assessor's-eye framing** — a comment or doc that describes the file from
|
||||
outside the scenario ("this is where the interesting part is," "the
|
||||
important method is below") rather than as something an in-world engineer
|
||||
would leave.
|
||||
|
||||
## What is NOT a finding
|
||||
|
||||
- **Genuine requirements and acceptance criteria.** "Add a retry with
|
||||
exponential backoff capped at 30s," "the endpoint must stay
|
||||
backwards-compatible with v1 clients," "use the existing PDF pipeline" —
|
||||
specific asks that define the deliverable, which the rubric checks, are the
|
||||
task, not hints. Precision about *what to build* is good authoring; the
|
||||
detector fires on giveaways about *what the agent is supposed to notice,
|
||||
decide, or find*.
|
||||
- **Domain context the agent genuinely needs.** Business constraints, in-world
|
||||
background, a ticket's reproduction steps, what the requester already tried.
|
||||
A busy user explaining their situation is realism, not hinting — even when
|
||||
it's detailed.
|
||||
- **A false or contested premise stated in the prompt.** Tasks legitimately
|
||||
model a requester who believes something wrong; the prompt asserting that
|
||||
belief is the scenario, not a hint (the hint would be the prompt *also*
|
||||
flagging that the belief is wrong).
|
||||
- **Pre-existing repo comments.** Comments that come from the source repo
|
||||
unmodified are the codebase the agent must cope with; only the
|
||||
`workspace.patch` additions/edits are in scope.
|
||||
- **In-world artifacts that carry the scenario.** A planted TODO or draft doc
|
||||
can *be* the task's subject (e.g. the task is about an unfinished feature
|
||||
the TODO marks). The flag condition is a comment that does the agent's
|
||||
thinking — locates the defect, prescribes the change, or narrates the
|
||||
judgment being graded — not that an authored comment exists. Ask: would an
|
||||
in-world engineer plausibly have left this, and does the graded difficulty
|
||||
survive it?
|
||||
- **Comments matching the repo's existing style.** Doc headers, license
|
||||
blocks, docstrings on public APIs — additions that mirror how the codebase
|
||||
already comments are craft, not hints.
|
||||
- **Detail level alone.** A long, thorough prompt is not over-hinted; a
|
||||
two-line prompt can be. The measure is whether graded judgment survives the
|
||||
text, not how much text there is.
|
||||
|
||||
## Verdict definitions
|
||||
|
||||
- **`clean`** — neither the prompt nor the authored workspace additions hint
|
||||
at the answer or pre-empt SWE-obvious judgment. Genuine requirements,
|
||||
domain context, and in-world artifacts are all clean (see the list above).
|
||||
- **`partial-hinting`** — mild or borderline hinting worth the author's
|
||||
attention: SWE-obvious directives ("be sure to add tests," "separate the
|
||||
view logic from the db logic"), a directive that restates a genuinely graded
|
||||
requirement, narrating-the-obvious comments away from the graded material,
|
||||
or a passage you can read either as scenario realism or as a nudge. The task
|
||||
still works; a de-hinting pass would sharpen it.
|
||||
- **`clear-hinting`** — the prompt or an authored comment points substantially
|
||||
at the answer the rubric scores: names the defect or its location,
|
||||
prescribes the graded fix or plan, pre-announces the exact diligence being
|
||||
measured, or a workspace comment sits on the planted issue and describes it.
|
||||
The intended difficulty is materially reduced for any agent that reads
|
||||
carefully.
|
||||
- **`not-applicable`** — nothing to assess: `instruction.md` is missing,
|
||||
empty, or only template/placeholder content, and there is no session
|
||||
history or `workspace.patch` to inspect. Re-run once the prompt lands.
|
||||
|
||||
`clear-hinting` and `partial-hinting` are the flagged outcomes; `clean` and
|
||||
`not-applicable` are not. All flagged outcomes are advisory: they hand the
|
||||
author a list of passages to reconsider, and the author may keep any of them
|
||||
deliberately.
|
||||
|
||||
## Confidence
|
||||
|
||||
- **HIGH** — the call is unambiguous: a passage plainly gives the answer away
|
||||
(or plainly nothing does), and the graded behavior is clear from the rubric.
|
||||
- **MEDIUM** — at least one finding is genuinely two-sided: a reasonable
|
||||
reviewer might read the passage as legitimate constraint or scenario
|
||||
realism.
|
||||
- **LOW** — limited information: the rubric is thin or absent so you can't
|
||||
tell what's graded, or the directive overlaps a genuine graded requirement
|
||||
("be sure to add tests" when the rubric requires tests) and the call is the
|
||||
author's to make.
|
||||
|
||||
## Relationship to other detectors
|
||||
|
||||
- **vs. detector-answer-obviousness.** Its over-cued shape owns the *fatal*
|
||||
end of prompt cueing: the prompt names the exact graded behavior, so the
|
||||
task cannot discriminate at all — a defect verdict about task validity.
|
||||
This detector owns the *gradient below that*: hint-shaped prose worth a
|
||||
de-hinting pass whether or not it fully disarms the task (SWE-obvious
|
||||
directives, partial giveaways, difficulty disclaimers), plus the
|
||||
workspace-comment surface, which answer-obviousness doesn't read. When a
|
||||
prompt hint is total, expect both to fire — answer-obviousness on validity,
|
||||
this detector on the concrete passages to rewrite.
|
||||
- **vs. detector-snapshot-leakage.** Leakage owns what the agent *inherits*
|
||||
from a captured session — conversation context that hands over the answer.
|
||||
This detector owns what the task *authors*: the prompt text and the
|
||||
comments inside `workspace.patch` additions. A workspace file that leaks
|
||||
the rubric's answer outright can fire both; each flags its own surface.
|
||||
- **vs. detector-cross-task-reference.** That detector scans the same
|
||||
authored surfaces for a different defect (pointers to sibling tasks). A
|
||||
comment can be both a sibling reference and a hint; the verdicts are
|
||||
independent.
|
||||
|
||||
## Anti-patterns: do not do these
|
||||
|
||||
- **Don't flag specificity.** A precise, well-specified ask is good
|
||||
authoring. The finding is a giveaway about the *graded* judgment,
|
||||
discovery, or diligence — not detail about the deliverable.
|
||||
- **Don't flag the scenario's own material.** False premises, planted
|
||||
in-world TODOs, and requester context are the task. Re-read the "What is
|
||||
NOT a finding" list before flagging anything in that family.
|
||||
- **Don't demand a hint be fatal before flagging.** This detector is
|
||||
advisory by design; a mild SWE-obvious directive is a legitimate
|
||||
`partial-hinting` finding even though the task still works.
|
||||
- **Don't treat the flag as a failure either.** Word every finding so the
|
||||
author can weigh it — quote the passage, say what it pre-empts, and offer
|
||||
the de-hinted alternative. Never assert the task is broken because a hint
|
||||
exists.
|
||||
- **Don't hunt hints in the grader guidance.** The agent under test
|
||||
never sees it; it's calibration context for you, not a surface.
|
||||
- **Don't cite evidence you haven't verified in the submitted package.**
|
||||
Quote passages as they exist in the actual `instruction.md`, session
|
||||
history, and `workspace.patch` — not as you remember them or as the rubric
|
||||
paraphrases them.
|
||||
|
||||
## Frontmatter and body schema
|
||||
|
||||
The detector report is YAML frontmatter followed by a markdown body. Both
|
||||
contexts produce the same shape; only the *sink* differs (the wrapping
|
||||
`SKILL.md` tells you where to send the report).
|
||||
|
||||
**Frontmatter** — exactly these keys, exactly these enum values:
|
||||
|
||||
```yaml
|
||||
---
|
||||
detector: detector-over-hinting
|
||||
verdict: clear-hinting | partial-hinting | clean | not-applicable
|
||||
confidence: HIGH | MEDIUM | LOW
|
||||
---
|
||||
```
|
||||
|
||||
**Body sections**, in this order:
|
||||
|
||||
```markdown
|
||||
# Over-hinting check: <slug>
|
||||
|
||||
## Findings
|
||||
|
||||
One block per finding, strongest first:
|
||||
|
||||
### <short label> — <prompt-hint | comment-hint> (<clear | partial>)
|
||||
|
||||
- **Where:** the file (and line/section for a patch hunk, or the turn for a
|
||||
session message).
|
||||
- **Quote:** the passage verbatim, as a blockquote — never a paraphrase.
|
||||
- **What it pre-empts:** one or two sentences — the judgment, discovery, or
|
||||
diligence the passage hands over, tied to what the rubric grades where
|
||||
possible.
|
||||
- **De-hinting option:** a concrete alternative — delete the sentence, move
|
||||
the fact into grader guidance, rewrite the directive as an in-world
|
||||
constraint — worded so the author can decide whether to take it.
|
||||
|
||||
For `clean`, quote the strongest near-miss (a specific requirement, a planted
|
||||
in-world TODO, a detailed prompt) and say why it was cleared. For
|
||||
`not-applicable`, name the missing artifacts.
|
||||
|
||||
## Overall verdict
|
||||
|
||||
2–3 paragraphs reducing the findings to the chosen verdict: which surface(s)
|
||||
hint and how strongly, whether the graded difficulty survives, and — because
|
||||
this detector is advisory — what a de-hinting pass would change first. For
|
||||
`clean`, why the near-misses are constraints or scenario material rather than
|
||||
hints.
|
||||
```
|
||||
|
||||
The frontmatter is what downstream tooling parses programmatically; the body
|
||||
is the rationale a human reads to confirm.
|
||||
Reference in New Issue
Block a user