added 260907 version of worker toolkit
This commit is contained in:
@@ -0,0 +1,58 @@
|
||||
---
|
||||
name: detector-cross-task-reference
|
||||
description: |
|
||||
Self-check whether your holistic rubric (or `instruction.md`)
|
||||
references another task — a separate task with its own prompt, workspace, and
|
||||
rubric that this task's grader will never see. The common slip: calibrating a
|
||||
new task against one you wrote earlier ("the failure-mode silhouette is similar
|
||||
to narrowed-too-early", "unlike the webhook-threat task"), which leaves a
|
||||
dangling pointer the grader can't resolve and couples two tasks that must stand
|
||||
alone. Each task has to be fully independent. Reads the holistic rubric file
|
||||
that `bash scripts/guidance-target.sh <slug>` resolves + instruction.md.
|
||||
allowed-tools: Bash, Read, Write
|
||||
---
|
||||
|
||||
# Cross-task-reference detector
|
||||
|
||||
This skill checks whether your task stands on its own — whether your
|
||||
holistic rubric (or `instruction.md`) explains this task's expected
|
||||
behavior by pointing at a **different task**.
|
||||
|
||||
The grader evaluates your task in isolation. It sees only this task's
|
||||
`instruction.md`, its workspace, and your holistic rubric (the file
|
||||
`bash scripts/guidance-target.sh <slug>` resolves) — never any
|
||||
other task. So a sentence like "the failure-mode silhouette is similar to
|
||||
**narrowed-too-early**" or "unlike the webhook-threat task" is a dead end: the
|
||||
grader can't look up what that other task was, and any calibration that hangs off
|
||||
the comparison is lost. It also couples two tasks that are supposed to be
|
||||
independent — if the other task is later changed or dropped, your rubric's
|
||||
meaning silently shifts.
|
||||
|
||||
What is **not** a problem: citing your own source repo (`grep "def as_json"
|
||||
app/models/`), comparing the *product* to real companies ("similar to Earnin /
|
||||
DailyPay"), or naming general concepts, patterns, and libraries. The defect is
|
||||
specifically a pointer to *another task in the set*.
|
||||
|
||||
Read these before deciding:
|
||||
|
||||
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
|
||||
2. `.claude/skills/detector-cross-task-reference/core.md` — what counts as a cross-task reference vs. what doesn't, the verdict enums, and the body schema.
|
||||
|
||||
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
|
||||
|
||||
## Acting on the verdict
|
||||
|
||||
- **`clean`** — your rubric and prompt stand on their own. No references to other
|
||||
tasks. Good. Move on.
|
||||
- **`partial-reference`** — a borderline or low-severity reference (a generic "like
|
||||
other tasks" aside, or a token that might be a sibling-task name). Read the
|
||||
grounding; inline whatever the reference was gesturing at so nothing depends on
|
||||
another task.
|
||||
- **`clear-reference`** — you reference a specific other task (a named sibling, a
|
||||
"similar to / unlike X" comparison, or borrowed calibration). Delete the
|
||||
cross-task comparison and state the point directly in terms of *this* task's
|
||||
own prompt and workspace. Keep any concrete in-this-task guidance (e.g. the
|
||||
exact grep or file to check) — it's only the pointer to the other task that has
|
||||
to go. Re-run after.
|
||||
- **`not-applicable`** — there's no holistic rubric (or prompt) to assess yet.
|
||||
Draft it first.
|
||||
@@ -0,0 +1,220 @@
|
||||
# Cross-task-reference detector — core
|
||||
|
||||
This file is the canonical, context-neutral content for the detector-cross-task-reference
|
||||
detector. It defines what the detector looks for, the verdict enums, the
|
||||
patterns to recognize, and the output schema. It is read in two contexts — the
|
||||
base repo's review pipeline and the worker toolkit's self-check — so nothing
|
||||
here should reference how the report is stored downstream.
|
||||
|
||||
## What this detector is for
|
||||
|
||||
Every task has to stand on its own. The grader evaluates one task in isolation:
|
||||
it sees only that task's `instruction.md`, its workspace, and its grader
|
||||
guidance. It has no access to any other task — not the prompt,
|
||||
not the workspace, not the rubric, not the reference runs of a different task.
|
||||
|
||||
So when a rubric (or a prompt) explains *this* task's expected behavior by
|
||||
pointing at a *different* task — "the failure-mode silhouette is similar to
|
||||
narrowed-too-early", "unlike the webhook-threat task", "score this the way we
|
||||
scored the invoice-OCR task" — two things go wrong:
|
||||
|
||||
1. **The pointer is dangling.** The grader cannot look up what
|
||||
`narrowed-too-early` was, so any calibration that hangs off that comparison is
|
||||
lost. The grader is left guessing what the sentence meant.
|
||||
2. **Two tasks that should be independent are now coupled.** If the referenced
|
||||
task is later revised, renamed, or dropped, this rubric's meaning silently
|
||||
shifts even though nobody touched this file.
|
||||
|
||||
The fix is always the same: inline whatever the cross-reference was trying to
|
||||
convey, so the rubric (or prompt) is self-contained. If the point was "the agent
|
||||
should grep `def as_json` in `app/models/` rather than chase entry points," say
|
||||
*that* directly — don't say "like in narrowed-too-early."
|
||||
|
||||
This detector decides: does *this* submission's authored text — the rubric, the
|
||||
prompt, or any file the submission adds to the workspace — reference another
|
||||
task?
|
||||
|
||||
## Inputs
|
||||
|
||||
Read from `harbor-tasks/<slug>/`:
|
||||
|
||||
- The grader guidance — the rubric. The primary input; this is where
|
||||
cross-task references most often creep in (an author calibrating the new task
|
||||
against one they wrote earlier). Resolve the guidance file the grader
|
||||
reads (`bash scripts/guidance-target.sh <slug>` prints its path,
|
||||
`tests/grader-guidance-consolidated.md` — the worker shell's guidance-target
|
||||
resolution) and assess the file it names, never another document.
|
||||
- `instruction.md` — the prompt the agent under test receives. A cross-task
|
||||
reference here is also a defect (the agent shouldn't learn that other tasks
|
||||
exist, and the reference is just as unresolvable for it). Scan it too.
|
||||
- **Authored workspace additions** — files the submission itself adds to or
|
||||
edits in the workspace, i.e. the `environment/workspace.patch` diff (a
|
||||
task-authored CLAUDE.md, README, design doc, ticket, or similar staged for
|
||||
the test agent to read). Scan the added/modified content in the patch for
|
||||
sibling-task names and set-membership framing — you don't need to build the
|
||||
workspace. A sibling reference here is arguably worse than one in the rubric:
|
||||
it dangles for the grader *and* hands the agent under test context about the
|
||||
task set it should never see (reference-run grades have cited such a file to
|
||||
justify their scores). This has happened in the wild via a workspace
|
||||
CLAUDE.md naming sibling task slugs.
|
||||
|
||||
You do not need the source repo for this call — it's a self-containment check on
|
||||
the authored text, not a fact-check of claims against code. The workspace's
|
||||
pre-existing repo content is out of scope; only the authored additions in the
|
||||
patch are.
|
||||
|
||||
## What counts as a cross-task reference
|
||||
|
||||
A reference to **another task in the set** — a separate task with its own prompt,
|
||||
workspace, and rubric that the grader of this task will never see. Tells:
|
||||
|
||||
- **Naming a sibling task by its slug.** Task slugs are usually behavior-named
|
||||
and hyphenated — `narrowed-too-early`, `missed-blast-radius`,
|
||||
`webhook-threat`, `contact-portal-design`. A hyphenated proper-noun token used
|
||||
to name a task (not a file, branch, or library) is the strongest signal. Watch
|
||||
for the hyphenation + a framing that treats it as a known entity ("similar to
|
||||
narrowed-too-early") rather than a description of behavior ("the agent narrowed
|
||||
its scope too early"). The first is a pointer; the second is prose.
|
||||
- **Comparative framing against another task.** "similar to the X task", "unlike
|
||||
X", "the harder version of X", "as we saw in X", "the same setup as X", "this
|
||||
is the companion to X".
|
||||
- **Borrowing calibration from another task.** "score this the way we scored X",
|
||||
"see X's grader-guidance", "reuse the rubric from X", "apply the same gate as
|
||||
in X".
|
||||
- **Generic-but-load-bearing pointers to siblings.** "the other task", "a
|
||||
sibling task", "another task in this set", "the companion task" — used as if
|
||||
the grader could resolve which one.
|
||||
- **Set-membership framing in an authored workspace file.** A doc the
|
||||
submission adds to the workspace that describes this task from the assessor's
|
||||
point of view — naming the bug pattern under test, the behavior being
|
||||
assessed, or sibling task slugs — instead of speaking as in-world scenario
|
||||
material. The tells are the same as above (sibling slugs, comparative
|
||||
framing); the file is just a different place they leak into.
|
||||
|
||||
## What is NOT a cross-task reference (do not flag these)
|
||||
|
||||
- **This task's own source repo** — file paths, function/class names, modules,
|
||||
grep commands (`grep "def as_json" app/models/`), commit SHAs. That is the
|
||||
task's own material and required context.
|
||||
- **This task's own reference runs / trials.** Over-anchoring the rubric on the
|
||||
observed runs is a real defect, but a *different* one (it's about
|
||||
generalizing to a new agent, not about pointing at a separate task). Don't
|
||||
flag it here.
|
||||
- **Real-world products, companies, or services** used to describe the domain —
|
||||
"an earned-wage-access platform similar to services like Earnin, DailyPay, or
|
||||
Payactiv." Comparing the *product* to real companies is not a reference to
|
||||
another task.
|
||||
- **General named concepts** — design patterns, algorithms, libraries,
|
||||
frameworks, RFCs, CVE IDs, external docs.
|
||||
- **The shared rubric vocabulary** — the eight criteria of the Grading
|
||||
Standard (Integrity, Narrow Correctness, Broader Correctness / craft,
|
||||
Persistence, Communication, Verification & Thoroughness, Common Sense,
|
||||
Thought Partnership) are the project's common language, not other tasks.
|
||||
- **Describing the genre, not a specific sibling** — "in a typical refactoring
|
||||
task", "this kind of audit task". Naming the *category* is fine; it points at
|
||||
nothing the grader needs to look up.
|
||||
- **Authored workspace docs as scenario material.** Many tasks legitimately
|
||||
seed a CLAUDE.md, README, ticket, or design doc into the workspace — that's
|
||||
the scenario, not the defect. The flag condition is that the file names a
|
||||
sibling task or frames membership in a task set, never merely that an
|
||||
authored doc exists.
|
||||
- **The workspace's pre-existing repo content.** Files that come from the
|
||||
source repo unmodified are the task's own material; only the submission's
|
||||
additions/edits (the `environment/workspace.patch` diff) are in scope.
|
||||
|
||||
## Verdict definitions
|
||||
|
||||
- **`clean`** — the rubric, prompt, and authored workspace additions are
|
||||
self-contained. No references to other tasks. (Product-domain comparisons,
|
||||
source-repo citations, and named concepts are all clean — see the list
|
||||
above.)
|
||||
- **`partial-reference`** — a borderline or low-severity cross-task reference.
|
||||
Either: (a) the reference is generic and doesn't name a specific sibling ("a
|
||||
bit like other tasks in this set") so the coupling is vaguer; or (b) a token
|
||||
*might* be a sibling-task slug but could plausibly be a file/branch/concept and
|
||||
you can't tell from context; or (c) the reference sits in a non-load-bearing
|
||||
aside (a parenthetical that doesn't gate any score). Still worth fixing —
|
||||
inline the intent — but not a hard, scoring-relevant dangling pointer.
|
||||
- **`clear-reference`** — an unambiguous reference to a specific other task: a
|
||||
named sibling slug, explicit comparative framing against another task, or
|
||||
borrowed calibration ("score it like X"). Especially when it's load-bearing —
|
||||
a scoring tier or heavy deduction whose meaning depends on knowing the other
|
||||
task.
|
||||
- **`not-applicable`** — there's nothing to assess: the resolved guidance file is
|
||||
missing, empty, or only template/placeholder content (and `instruction.md`
|
||||
likewise has no authored body). Re-run once the rubric lands.
|
||||
|
||||
`clear-reference` and `partial-reference` are the flagged outcomes; `clean` and
|
||||
`not-applicable` are not.
|
||||
|
||||
## Confidence
|
||||
|
||||
- **HIGH** — the call is unambiguous: a clearly-named sibling task, or clearly
|
||||
nothing of the sort.
|
||||
- **MEDIUM** — the token/phrasing is probably a cross-task reference but a
|
||||
reasonable reviewer might read it as a file, concept, or genre.
|
||||
- **LOW** — limited information; verdict is a best guess.
|
||||
|
||||
## Patterns to look for
|
||||
|
||||
- A hyphenated, behavior-shaped proper noun (`narrowed-too-early`,
|
||||
`over-eager-refactor`) that reads as a *name*, not a description — most telling
|
||||
right after a comparative ("similar to", "like", "unlike", "as in").
|
||||
- Sentences that only make sense if the reader already knows a different task:
|
||||
"this is the stricter version", "we calibrated this against the earlier one".
|
||||
- A rubric section headed "Difference From Similar Tasks" (or any heading in
|
||||
that shape). Such a section exists to compare against siblings, and it almost
|
||||
always names them.
|
||||
- A scoring tier or heavy penalty that defers its definition to another task instead
|
||||
of stating the criterion in full.
|
||||
- A workspace file added by the patch that reads like assessment context rather
|
||||
than in-world material — describing what this task tests, its subject or bug
|
||||
pattern, or naming other tasks.
|
||||
|
||||
The clean shape: every calibration the rubric relies on is stated *in this file*,
|
||||
in terms of this task's own prompt, workspace, and expected behavior — and every
|
||||
authored workspace addition speaks only in-world.
|
||||
|
||||
## Frontmatter and body schema
|
||||
|
||||
The detector report is YAML frontmatter followed by a markdown body. Both
|
||||
contexts produce the same shape; only the *sink* differs (the wrapping `SKILL.md`
|
||||
tells you where to send the report).
|
||||
|
||||
**Frontmatter** — exactly these keys, exactly these enum values:
|
||||
|
||||
```yaml
|
||||
---
|
||||
detector: detector-cross-task-reference
|
||||
verdict: clear-reference | partial-reference | clean | not-applicable
|
||||
confidence: HIGH | MEDIUM | LOW
|
||||
---
|
||||
```
|
||||
|
||||
**Body sections**, in this order:
|
||||
|
||||
```markdown
|
||||
# Cross-task-reference check: <slug>
|
||||
|
||||
## Verbatim grounding
|
||||
|
||||
Quote the offending passage(s) from the resolved guidance file,
|
||||
`instruction.md`, or an authored workspace file (name the file the
|
||||
`environment/workspace.patch` diff adds/edits) as blockquotes — don't
|
||||
paraphrase. Name the file and, for each quote, the sibling task it points at.
|
||||
For "clean", quote the strongest near-miss (a product-domain comparison, a
|
||||
source-repo citation, a genre mention, an authored workspace doc that stays
|
||||
in-world) so the reader can confirm it was considered and correctly cleared.
|
||||
For "not-applicable", quote the missing/empty/template artifact.
|
||||
|
||||
## Rationale
|
||||
|
||||
2–4 paragraphs. For a flagged verdict: which passage references which other
|
||||
task, why the grader of this task can't resolve it, and what should be inlined
|
||||
instead so the rubric stands alone. For "clean": why the near-misses are not
|
||||
cross-task references. For "not-applicable": which trigger fired and what needs
|
||||
to land before the detector can run.
|
||||
```
|
||||
|
||||
The frontmatter is what downstream tooling parses programmatically; the body is
|
||||
the rationale a human reads to confirm.
|
||||
Reference in New Issue
Block a user