lots of change - all to start my 3rd redo
This commit is contained in:
@@ -1,139 +0,0 @@
|
||||
---
|
||||
name: brainstorm-product-arcs
|
||||
description: Invent big, plausible product directions ("arcs") for a source repo and decompose each into a backlog of concrete tasks that can actually be built and verified with no network access. Use when you want task ideas that ladder into a coherent product story instead of one-off commits.
|
||||
allowed-tools: Read, Glob, Grep, Bash, Write, Edit, WebSearch, WebFetch, Task
|
||||
---
|
||||
|
||||
# Brainstorm Product Arcs
|
||||
|
||||
## What this is
|
||||
|
||||
A method for going from "what could this company build next?" to a backlog of concrete, buildable raccoon tasks. Instead of mining the git history for a single commit to recreate, you invent a **product arc** — a big, plausible direction the company would pursue — and decompose it into many tasks that share one story.
|
||||
|
||||
Use this when:
|
||||
|
||||
- you want a set of tasks that ladder into a coherent theme, not scattered one-offs;
|
||||
- you're starting from the product ("what's the next feature?") rather than from a commit;
|
||||
- you want forward-looking features (things the codebase doesn't have yet), not historical changes.
|
||||
|
||||
An **arc** is a product-scale initiative (e.g. "let workers build credit", "self-serve employer onboarding"), not a single feature. One arc spawns 4–8 tasks.
|
||||
|
||||
## The three lenses
|
||||
|
||||
Every arc has to pass all three. Most ideas die on lens 3.
|
||||
|
||||
1. **Plausible** — obviously something _this_ company would do. The test: is it an expansion of what they already do, or a pivot "into making printers"? Ground it in the real product, not the brand.
|
||||
2. **Differentiated** — a sharp, concrete delta against both (a) what the product does _today_ and (b) the _workaround_ a user reaches for now (a named competitor or a manual process).
|
||||
3. **Buildable with no network** — the substance has to be exercisable by a test suite in a sandbox with no internet. This is the gate, and it's the heart of this skill (Step 4).
|
||||
|
||||
## Step 1 — Map the product surface first (go deep; don't guess)
|
||||
|
||||
This step is the foundation: a shallow or guessed map produces wrong "today" baselines and implausible arcs, and every later step inherits the error. **Take the time to actually read the code, and verify every claim against a file you've opened** — there's no token or time budget to protect here, and depth pays for itself.
|
||||
|
||||
Read the code first and write down:
|
||||
|
||||
- the main data models and what they represent in product terms;
|
||||
- the feature areas (from directory / route names) and what each does for the user;
|
||||
- who the end user is;
|
||||
- the external integrations and what each one powers;
|
||||
- the repo's existing **mocking patterns** — how it already fakes those integrations in tests (provider / adapter interfaces with fakes, recorded HTTP cassettes, a swappable HTTP client + JSON fixtures, or service-object stubs in the consuming spec). You'll reuse these in Step 4, so note where they live and which canonical file to copy;
|
||||
- how the company makes money.
|
||||
|
||||
Dispatch several Explore subagents in parallel for breadth, then read the load-bearing files yourself. Everything you propose later must cite real files / models — never a guess, and never a memory of "how apps like this usually work." That's what keeps lens 1 honest and the deltas accurate. This mapping is your own groundwork — it feeds each arc's Delta; it does **not** become a shared "current state" section in the output. Every arc must stand alone (see Step 5).
|
||||
|
||||
## Step 2 — Generate arcs (lens 1: plausible)
|
||||
|
||||
Heuristics that produce arcs that read as "obviously them":
|
||||
|
||||
- **Widen a proven mechanic.** The strongest arcs generalize something the product already does _narrowly_ (a rent-smoothing engine pointed at any bill; a one-step approval grown into multi-step policies). The company has already proven the mechanic; you're just broadening it.
|
||||
- **Follow the asset.** What does this company uniquely have — a data set, a relationship, a captured flow? Build on that.
|
||||
- **Keep arcs independent.** Each arc — and each task it spawns — should stand on its own, so the set can be fanned out to different people and built in parallel. Avoid arcs (or tasks) that only make sense once another one ships.
|
||||
- **New markets count as arcs** (a new segment, vertical, or user type) — as long as they reuse infrastructure the company already has.
|
||||
|
||||
Apply the printer test ruthlessly. Write down, for calibration, 2–3 ideas that would _not_ scan, so the boundary is explicit.
|
||||
|
||||
## Step 3 — Sharpen the delta (lens 2: differentiated)
|
||||
|
||||
For each arc, write these four things. A vague "better X" is not a delta.
|
||||
|
||||
- **In-product today:** what exists now, citing code — the baseline _this_ arc changes (keep it inside the arc; see Step 5).
|
||||
- **Workaround today:** the named competitor or the manual process a user uses to get the same outcome right now. Name it; link it.
|
||||
- **Without it / With it:** a concrete scenario each way, written as a **numbered list** — the steps the user actually goes through, in order. Numbered steps read far better here than a dense paragraph.
|
||||
- **The delta:** one sentence — "what's actually new."
|
||||
|
||||
Name the real external services and link them. They're load-bearing twice over: they make the delta concrete, _and_ they're where lens 3 gets decided.
|
||||
|
||||
**Write for a non-expert reader.** Define business-domain terms (what a credit bureau is, what "KYB" means) on first use — but don't explain general SWE concepts (mock, fixture, state machine); the reader already knows those.
|
||||
|
||||
## Step 4 — The buildability filter ("simulate the protocol, not the product")
|
||||
|
||||
The sandbox that runs a finished task has **no outbound network access** — you can confirm this yourself by running the task under `harbor-run`. So any external service the feature depends on must be faked locally; there's no calling the real API at grade time. The question is never "does it touch the network" — it's whether a _faithful_ local mock is possible.
|
||||
|
||||
Grade every feature into one of three buckets:
|
||||
|
||||
- **Build directly (internal logic).** The substance is logic a test suite exercises with static inputs: state machines, money math, eligibility / validation rules, routing / waterfalls, parsing a fixtured payload and mutating state.
|
||||
- **Build with a mock (a documented protocol).** The external dependency is a _contract_: forms, file formats, return / webhook codes, ledger APIs, list lookups. The hard work is on _our_ side (build the request, parse the response, reconcile state, handle the documented failures). Stand up a faithful local mock — a fixture service, a small local server, a seeded table. It **must be adversarial**: a mock that only ever returns success is fake even for a great protocol; the difficulty lives in the realistic _failures_ it throws (rejects, returns, conflicts, async-then-callback, partial failures).
|
||||
- **Don't build it (a product / experience / black-box model).** The substance is a client SDK + device + UX (a mobile wallet), proprietary model behavior (OCR accuracy, fraud scoring), or market mechanics (FX pricing). A local mock collapses to a cartoon and deletes the only hard part. Skip it, or scope down to the protocol slice.
|
||||
|
||||
**The author's test:** _"Could I write this mock's spec straight from public documentation, AND would a correct integration against my mock also be correct against the real service?"_ Two yeses → build the mock. If the honest answer is "my mock would be a cartoon of the real thing" → don't.
|
||||
|
||||
**The split move:** most "integration" features decompose into a buildable protocol slice plus a non-buildable product slice. "Add card payments" = [skip: the wallet / SDK] + [build: verify the signed webhook, apply the fee, transition state]. "Pay overseas" = [skip: FX execution] + [build: multi-currency modeling]. Scope the task to the buildable slice and host a faithful mock for the boundary.
|
||||
|
||||
> **Follow the repo's existing mocking patterns.** Most of these source repos already have a way to fake their external dependencies — a provider / adapter interface with a fake implementation, test doubles, recorded fixtures, or a local stub server (you noted it in Step 1). Build any new mock the _same_ way, wired through the same seam, rather than inventing a new style — and point your coding agent at the existing example to copy. Matching the repo's convention matters more than the technique you'd pick from scratch. (For anything bank-related, an existing fake banking-as-a-service provider is usually the template — extend it.)
|
||||
|
||||
## Step 5 — Decompose each arc into a task backlog
|
||||
|
||||
This is the deliverable. For each arc, produce:
|
||||
|
||||
- **Pitch** — one line on what it is.
|
||||
- **Why it's them** — the plausibility argument.
|
||||
- **External services** — named and linked.
|
||||
- **Delta** — _in-product today_ (the baseline this arc changes — folded in here, not in a shared section) / _workaround today_ / numbered _without_-vs-_with_.
|
||||
- **Buildability — what to mock, and how** — name which parts are plain internal logic (built directly), then each external system that must be mocked: what it is, a link to learn its contract, the **repo's existing mock pattern to follow** (point to it), and a concrete pointer for standing up the mock with a coding agent (what to read, what to generate, which failure cases to seed). When the answer is "nothing external to mock," say so — it's a strength.
|
||||
- **Tasks it spawns** — 4–8 concrete tasks, each a candidate to author. Mark which need a mock and what it models.
|
||||
|
||||
Each line in "tasks it spawns" should be a real task you could hand to someone.
|
||||
|
||||
**Keep every arc self-contained.** A reader should get the whole idea from its one section, top to bottom — so the "today" baseline lives in that arc's Delta, never in a shared "current state" section. And don't frame buildability as a yes/no question: by the time an arc is in the backlog it has already passed the Step 4 filter, so describe _how_ it's built, not _whether_.
|
||||
|
||||
## Step 6 — Hand off to authoring
|
||||
|
||||
A task idea isn't a task until it has a verifier. The buildability filter is exactly what makes a verifier possible offline: a build-directly task is verified by tests over internal logic; a build-with-a-mock task is verified by tests over the local mock's behavior. When you write the holistic rubric for one of these, see [[write-holistic-rubric]].
|
||||
|
||||
## A worked example (full template)
|
||||
|
||||
This is the Step 5 output shape for a single arc — copy this structure, including the numbered _Without it_ / _With it_ lists.
|
||||
|
||||
**Arc — "Credit Builder"** (for a worker-banking app)
|
||||
|
||||
- **Pitch:** let workers build a credit history through the on-time payments they already make, reported automatically from their paycheck.
|
||||
- **Why it's them:** the app already issues cards and runs repayment ledgers — this reuses both — and building credit is a natural next step for the paycheck-to-paycheck users it serves.
|
||||
- **External services:** [Experian](https://www.experian.com) / [Equifax](https://www.equifax.com) / [TransUnion](https://www.transunion.com) (the credit bureaus); [Metro 2](https://www.cdiaonline.org/metro-2/) (the file format used to report to them); [e-OSCAR](https://www.e-oscar.org) (the dispute system). Products a user would otherwise use: [Self](https://www.self.inc), [Kikoff](https://kikoff.com), [Chime Credit Builder](https://www.chime.com/credit/credit-builder/).
|
||||
- **Delta — in-product today:** the app issues debit cards and tracks repayments, but reports nothing to the bureaus, so none of that activity builds the user's credit.
|
||||
- **Delta — workaround today:** the user signs up for a separate credit-builder app (Self, Kikoff, Chime) that isn't connected to their paycheck.
|
||||
- **Delta — without it:**
|
||||
1. The worker gets paid and takes the occasional advance in the app.
|
||||
2. None of it is reported to the bureaus, so their credit score doesn't move.
|
||||
3. To build credit they open a second app (e.g. Self) and commit to a fixed monthly payment.
|
||||
4. They manage it as a separate account, with a separate payment to remember.
|
||||
- **Delta — with it:**
|
||||
1. The worker turns on "Credit Builder" in the app — no second account.
|
||||
2. Each pay cycle, the app sets aside the scheduled payment from the incoming paycheck.
|
||||
3. The app records it as on-time and reports it to the three bureaus that month.
|
||||
4. The worker builds credit inside the app their paycheck already lands in.
|
||||
- **Buildability — what to mock, and how:** the ledger, on-time / late logic, utilization, and the deposit hold build directly. The only external touchpoint — the credit bureaus — is a documented protocol: **Metro 2** is a published file format ([CDIA](https://www.cdiaonline.org/metro-2/)), so have your coding agent generate a valid file from the ledger plus a "bureau" that returns realistic field-level rejects, seeded with good and bad records; **e-OSCAR** disputes ([e-oscar.org](https://www.e-oscar.org)) are a verify-request → response-code exchange. Build both the way this repo already fakes its bank provider — find that fake and follow its pattern (extend it if it covers the rails) rather than starting fresh.
|
||||
- **Tasks it spawns:** Metro 2 file builder (+ handle the mock's rejects); on-time / late / charge-off determination with grace periods; secured-deposit hold and release; dispute (ACDV) state machine; utilization and credit-limit-increase rules; payment-allocation order (fees → interest → principal).
|
||||
|
||||
> Another domain, same shape: an accounts-payable app's **"1099 e-file"** arc — aggregating each vendor's annual payments and producing the tax form is internal logic, and the IRS e-file boundary is a documented protocol, so a local mock validates the filing and returns accept / reject-with-error-code.
|
||||
|
||||
## Common mistakes
|
||||
|
||||
- **A shallow or guessed map** (brainstorming from the brand, or from how "apps like this usually work") → implausible arcs and wrong "today" baselines. Go deep in Step 1 and verify every claim against a file you've opened.
|
||||
- **A mushy delta** ("make X better") → a delta is a named workaround plus a concrete numbered without / with.
|
||||
- **A shared "current state" section** → keep each arc self-contained; its _in-product today_ line carries the baseline, so a reader never has to look elsewhere.
|
||||
- **Arcs (or tasks) that depend on each other** → they can't be fanned out to workers in parallel. Make each one stand alone.
|
||||
- **A mock built in a new style** → if the repo already fakes its integrations a certain way, follow that pattern and wiring; don't invent a parallel one.
|
||||
- **Treating "touches an external service" as disqualifying** → it isn't; the question is protocol-vs-product fidelity.
|
||||
- **A mock that only returns success** → fake even for a good protocol. Simulate the failures.
|
||||
- **Explaining SWE basics** (what a mock or a fixture is) → the reader knows them; spend the words on business-domain terms instead.
|
||||
- **A feature whose verifier needs live external state** (a real balance, a real model's output, a live rate) → unbuildable; either it's a don't-build, or you haven't found the buildable slice yet.
|
||||
Reference in New Issue
Block a user