lots of change - all to start my 3rd redo

This commit is contained in:
2026-09-26 14:31:52 -04:00
parent 7f4d388e19
commit bceb52e8ee
1046 changed files with 4476 additions and 0 deletions

View File

@@ -1,58 +0,0 @@
---
name: detector-cross-task-reference
description: |
Self-check whether your holistic rubric (or `instruction.md`)
references another task — a separate task with its own prompt, workspace, and
rubric that this task's grader will never see. The common slip: calibrating a
new task against one you wrote earlier ("the failure-mode silhouette is similar
to narrowed-too-early", "unlike the webhook-threat task"), which leaves a
dangling pointer the grader can't resolve and couples two tasks that must stand
alone. Each task has to be fully independent. Reads the holistic rubric file
that `bash scripts/guidance-target.sh <slug>` resolves + instruction.md.
allowed-tools: Bash, Read, Write
---
# Cross-task-reference detector
This skill checks whether your task stands on its own — whether your
holistic rubric (or `instruction.md`) explains this task's expected
behavior by pointing at a **different task**.
The grader evaluates your task in isolation. It sees only this task's
`instruction.md`, its workspace, and your holistic rubric (the file
`bash scripts/guidance-target.sh <slug>` resolves) — never any
other task. So a sentence like "the failure-mode silhouette is similar to
**narrowed-too-early**" or "unlike the webhook-threat task" is a dead end: the
grader can't look up what that other task was, and any calibration that hangs off
the comparison is lost. It also couples two tasks that are supposed to be
independent — if the other task is later changed or dropped, your rubric's
meaning silently shifts.
What is **not** a problem: citing your own source repo (`grep "def as_json"
app/models/`), comparing the *product* to real companies ("similar to Earnin /
DailyPay"), or naming general concepts, patterns, and libraries. The defect is
specifically a pointer to *another task in the set*.
Read these before deciding:
1. `.claude/skills/_detector-worker-shell.md` — where to write the report and how to handle re-runs.
2. `.claude/skills/detector-cross-task-reference/core.md` — what counts as a cross-task reference vs. what doesn't, the verdict enums, and the body schema.
Compose the report per the schema in `core.md` and write it per `_detector-worker-shell.md`.
## Acting on the verdict
- **`clean`** — your rubric and prompt stand on their own. No references to other
tasks. Good. Move on.
- **`partial-reference`** — a borderline or low-severity reference (a generic "like
other tasks" aside, or a token that might be a sibling-task name). Read the
grounding; inline whatever the reference was gesturing at so nothing depends on
another task.
- **`clear-reference`** — you reference a specific other task (a named sibling, a
"similar to / unlike X" comparison, or borrowed calibration). Delete the
cross-task comparison and state the point directly in terms of *this* task's
own prompt and workspace. Keep any concrete in-this-task guidance (e.g. the
exact grep or file to check) — it's only the pointer to the other task that has
to go. Re-run after.
- **`not-applicable`** — there's no holistic rubric (or prompt) to assess yet.
Draft it first.

View File

@@ -1,220 +0,0 @@
# Cross-task-reference detector — core
This file is the canonical, context-neutral content for the detector-cross-task-reference
detector. It defines what the detector looks for, the verdict enums, the
patterns to recognize, and the output schema. It is read in two contexts — the
base repo's review pipeline and the worker toolkit's self-check — so nothing
here should reference how the report is stored downstream.
## What this detector is for
Every task has to stand on its own. The grader evaluates one task in isolation:
it sees only that task's `instruction.md`, its workspace, and its grader
guidance. It has no access to any other task — not the prompt,
not the workspace, not the rubric, not the reference runs of a different task.
So when a rubric (or a prompt) explains *this* task's expected behavior by
pointing at a *different* task — "the failure-mode silhouette is similar to
narrowed-too-early", "unlike the webhook-threat task", "score this the way we
scored the invoice-OCR task" — two things go wrong:
1. **The pointer is dangling.** The grader cannot look up what
`narrowed-too-early` was, so any calibration that hangs off that comparison is
lost. The grader is left guessing what the sentence meant.
2. **Two tasks that should be independent are now coupled.** If the referenced
task is later revised, renamed, or dropped, this rubric's meaning silently
shifts even though nobody touched this file.
The fix is always the same: inline whatever the cross-reference was trying to
convey, so the rubric (or prompt) is self-contained. If the point was "the agent
should grep `def as_json` in `app/models/` rather than chase entry points," say
*that* directly — don't say "like in narrowed-too-early."
This detector decides: does *this* submission's authored text — the rubric, the
prompt, or any file the submission adds to the workspace — reference another
task?
## Inputs
Read from `harbor-tasks/<slug>/`:
- The grader guidance — the rubric. The primary input; this is where
cross-task references most often creep in (an author calibrating the new task
against one they wrote earlier). Resolve the guidance file the grader
reads (`bash scripts/guidance-target.sh <slug>` prints its path,
`tests/grader-guidance-consolidated.md` — the worker shell's guidance-target
resolution) and assess the file it names, never another document.
- `instruction.md` — the prompt the agent under test receives. A cross-task
reference here is also a defect (the agent shouldn't learn that other tasks
exist, and the reference is just as unresolvable for it). Scan it too.
- **Authored workspace additions** — files the submission itself adds to or
edits in the workspace, i.e. the `environment/workspace.patch` diff (a
task-authored CLAUDE.md, README, design doc, ticket, or similar staged for
the test agent to read). Scan the added/modified content in the patch for
sibling-task names and set-membership framing — you don't need to build the
workspace. A sibling reference here is arguably worse than one in the rubric:
it dangles for the grader *and* hands the agent under test context about the
task set it should never see (reference-run grades have cited such a file to
justify their scores). This has happened in the wild via a workspace
CLAUDE.md naming sibling task slugs.
You do not need the source repo for this call — it's a self-containment check on
the authored text, not a fact-check of claims against code. The workspace's
pre-existing repo content is out of scope; only the authored additions in the
patch are.
## What counts as a cross-task reference
A reference to **another task in the set** — a separate task with its own prompt,
workspace, and rubric that the grader of this task will never see. Tells:
- **Naming a sibling task by its slug.** Task slugs are usually behavior-named
and hyphenated — `narrowed-too-early`, `missed-blast-radius`,
`webhook-threat`, `contact-portal-design`. A hyphenated proper-noun token used
to name a task (not a file, branch, or library) is the strongest signal. Watch
for the hyphenation + a framing that treats it as a known entity ("similar to
narrowed-too-early") rather than a description of behavior ("the agent narrowed
its scope too early"). The first is a pointer; the second is prose.
- **Comparative framing against another task.** "similar to the X task", "unlike
X", "the harder version of X", "as we saw in X", "the same setup as X", "this
is the companion to X".
- **Borrowing calibration from another task.** "score this the way we scored X",
"see X's grader-guidance", "reuse the rubric from X", "apply the same gate as
in X".
- **Generic-but-load-bearing pointers to siblings.** "the other task", "a
sibling task", "another task in this set", "the companion task" — used as if
the grader could resolve which one.
- **Set-membership framing in an authored workspace file.** A doc the
submission adds to the workspace that describes this task from the assessor's
point of view — naming the bug pattern under test, the behavior being
assessed, or sibling task slugs — instead of speaking as in-world scenario
material. The tells are the same as above (sibling slugs, comparative
framing); the file is just a different place they leak into.
## What is NOT a cross-task reference (do not flag these)
- **This task's own source repo** — file paths, function/class names, modules,
grep commands (`grep "def as_json" app/models/`), commit SHAs. That is the
task's own material and required context.
- **This task's own reference runs / trials.** Over-anchoring the rubric on the
observed runs is a real defect, but a *different* one (it's about
generalizing to a new agent, not about pointing at a separate task). Don't
flag it here.
- **Real-world products, companies, or services** used to describe the domain —
"an earned-wage-access platform similar to services like Earnin, DailyPay, or
Payactiv." Comparing the *product* to real companies is not a reference to
another task.
- **General named concepts** — design patterns, algorithms, libraries,
frameworks, RFCs, CVE IDs, external docs.
- **The shared rubric vocabulary** — the eight criteria of the Grading
Standard (Integrity, Narrow Correctness, Broader Correctness / craft,
Persistence, Communication, Verification & Thoroughness, Common Sense,
Thought Partnership) are the project's common language, not other tasks.
- **Describing the genre, not a specific sibling** — "in a typical refactoring
task", "this kind of audit task". Naming the *category* is fine; it points at
nothing the grader needs to look up.
- **Authored workspace docs as scenario material.** Many tasks legitimately
seed a CLAUDE.md, README, ticket, or design doc into the workspace — that's
the scenario, not the defect. The flag condition is that the file names a
sibling task or frames membership in a task set, never merely that an
authored doc exists.
- **The workspace's pre-existing repo content.** Files that come from the
source repo unmodified are the task's own material; only the submission's
additions/edits (the `environment/workspace.patch` diff) are in scope.
## Verdict definitions
- **`clean`** — the rubric, prompt, and authored workspace additions are
self-contained. No references to other tasks. (Product-domain comparisons,
source-repo citations, and named concepts are all clean — see the list
above.)
- **`partial-reference`** — a borderline or low-severity cross-task reference.
Either: (a) the reference is generic and doesn't name a specific sibling ("a
bit like other tasks in this set") so the coupling is vaguer; or (b) a token
*might* be a sibling-task slug but could plausibly be a file/branch/concept and
you can't tell from context; or (c) the reference sits in a non-load-bearing
aside (a parenthetical that doesn't gate any score). Still worth fixing —
inline the intent — but not a hard, scoring-relevant dangling pointer.
- **`clear-reference`** — an unambiguous reference to a specific other task: a
named sibling slug, explicit comparative framing against another task, or
borrowed calibration ("score it like X"). Especially when it's load-bearing —
a scoring tier or heavy deduction whose meaning depends on knowing the other
task.
- **`not-applicable`** — there's nothing to assess: the resolved guidance file is
missing, empty, or only template/placeholder content (and `instruction.md`
likewise has no authored body). Re-run once the rubric lands.
`clear-reference` and `partial-reference` are the flagged outcomes; `clean` and
`not-applicable` are not.
## Confidence
- **HIGH** — the call is unambiguous: a clearly-named sibling task, or clearly
nothing of the sort.
- **MEDIUM** — the token/phrasing is probably a cross-task reference but a
reasonable reviewer might read it as a file, concept, or genre.
- **LOW** — limited information; verdict is a best guess.
## Patterns to look for
- A hyphenated, behavior-shaped proper noun (`narrowed-too-early`,
`over-eager-refactor`) that reads as a *name*, not a description — most telling
right after a comparative ("similar to", "like", "unlike", "as in").
- Sentences that only make sense if the reader already knows a different task:
"this is the stricter version", "we calibrated this against the earlier one".
- A rubric section headed "Difference From Similar Tasks" (or any heading in
that shape). Such a section exists to compare against siblings, and it almost
always names them.
- A scoring tier or heavy penalty that defers its definition to another task instead
of stating the criterion in full.
- A workspace file added by the patch that reads like assessment context rather
than in-world material — describing what this task tests, its subject or bug
pattern, or naming other tasks.
The clean shape: every calibration the rubric relies on is stated *in this file*,
in terms of this task's own prompt, workspace, and expected behavior — and every
authored workspace addition speaks only in-world.
## Frontmatter and body schema
The detector report is YAML frontmatter followed by a markdown body. Both
contexts produce the same shape; only the *sink* differs (the wrapping `SKILL.md`
tells you where to send the report).
**Frontmatter** — exactly these keys, exactly these enum values:
```yaml
---
detector: detector-cross-task-reference
verdict: clear-reference | partial-reference | clean | not-applicable
confidence: HIGH | MEDIUM | LOW
---
```
**Body sections**, in this order:
```markdown
# Cross-task-reference check: <slug>
## Verbatim grounding
Quote the offending passage(s) from the resolved guidance file,
`instruction.md`, or an authored workspace file (name the file the
`environment/workspace.patch` diff adds/edits) as blockquotes — don't
paraphrase. Name the file and, for each quote, the sibling task it points at.
For "clean", quote the strongest near-miss (a product-domain comparison, a
source-repo citation, a genre mention, an authored workspace doc that stays
in-world) so the reader can confirm it was considered and correctly cleared.
For "not-applicable", quote the missing/empty/template artifact.
## Rationale
2–4 paragraphs. For a flagged verdict: which passage references which other
task, why the grader of this task can't resolve it, and what should be inlined
instead so the rubric stands alone. For "clean": why the near-misses are not
cross-task references. For "not-applicable": which trigger fired and what needs
to land before the detector can run.
```
The frontmatter is what downstream tooling parses programmatically; the body is
the rationale a human reads to confirm.