Compare commits

..

7 Commits

Author SHA1 Message Date
c8eb9114ba new Schwab mcp servers 2026-09-16 12:41:48 -04:00
b5bff5bd87 added ICM file and new AGENT.md 2026-09-16 12:39:59 -04:00
c34ec8af08 Moved files around, updated README
Various files from live trading sessions were dumped in the root.
Moved these files sources/misc/ (new folder).
Updated README
2026-09-16 10:59:54 -04:00
0e55edf917 added rules from failures 2026-09-01 15:33:36 -04:00
c0c770bcd2 added runbooks 2026-09-01 14:19:33 -04:00
7c7546f1da added build-watcher; updated README 2026-08-31 16:17:08 -04:00
81ae138a00 var - backtest related and spy runbook 2026-08-28 15:10:09 -04:00
40 changed files with 31796 additions and 15 deletions

6
.gitmodules vendored Normal file
View File

@@ -0,0 +1,6 @@
[submodule "resources/mcps/jkoelker-schwab-mcp"]
path = resources/mcps/jkoelker-schwab-mcp
url = https://github.com/jkoelker/schwab-mcp.git
[submodule "resources/mcps/vsoneji-schwab-mcp"]
path = resources/mcps/vsoneji-schwab-mcp
url = https://github.com/vsoneji/schwab-mcp.git

38
AGENT.md Normal file
View File

@@ -0,0 +1,38 @@
# ICM (Interpretable Context Methodology) – Process Overview
ICM is a workflow for orchestrating AI agent tasks using the filesystem as the coordination mechanism, eliminating the need for complex agent frameworks.
## Core Idea
- Numbered folders represent sequential stages of a workflow.
- Each stage contains a `CONTEXT.md` file that defines:
- **Inputs** (what files to read from reference and working layers)
- **Process** (what the agent should do)
- **Outputs** (what files to write)
- Plain markdown and JSON files carry prompts and context; local scripts handle non‑AI tasks.
- The agent reads the appropriate files at each stage, producing intermediate outputs that humans can inspect, edit, and approve before proceeding.
## Five‑Layer Context Hierarchy
1. **Layer 0 – CLAUDE.md**: Workspace identity and routing.
2. **Layer 1 – CONTEXT.md**: Workspace‑level task routing (which stage to run).
3. **Layer 2 – Stage CONTEXT.md**: Stage‑specific contract (inputs, process, outputs).
4. **Layer 3 – Reference material** (`references/`, `_config/`): Stable rules, voice guides, design systems (the “factory”).
5. **Layer 4 – Working artifacts** (`output/`): Per‑run intermediate results (the “product”).
## Workflow Characteristics
- **Sequential**: Stage n+1 reads the output of Stage n.
- **Human‑in‑the‑loop**: After each stage, a human can review and edit the output file before the next stage runs.
- **Editable & observable**: All prompts, context, and intermediate results are plain text files in folders—easy to inspect, version, and modify.
- **Portable**: A workspace is just a folder; it can be copied, versioned with Git, or shared without extra configuration.
- **Focused context loading**: Each stage loads only the context it needs (Layers 0‑4 relevant to that stage), keeping the model’s context window small and relevant.
## Benefits
- Replaces multi‑agent orchestration frameworks with simple folder conventions.
- Enables non‑technical users to modify prompts and workflows by editing markdown files.
- Provides inherent audit trails and review gates, supporting human oversight and debugging.
- Scales token usage efficiently by avoiding irrelevant context.
## Typical Use Cases
Content production pipelines (research → script → animation), slide‑deck generation, research analysis, policy workflows—any repeatable, sequential process where human review at each step adds value.
## License
ICM is open source under the MIT license; a workspace‑builder tool is included to scaffold new workspaces.

202
README.md
View File

@@ -2,15 +2,31 @@
This README catalogs the tools and resources in this directory. It is intended for humans and LLM agents to quickly identify what is available and when to use it.
## Jurisdiction (set by member, 2026-09-01)
- `workflows/` = **ICM** (Interpretable Context Methodology, per `workflows/ICM.md`). Member invokes these by name; human review gates between stages. MARI never applies ICM structure to live trade execution.
- `runbooks/` = **MARI-native execution procedures.** This is how MARI runs when the member says "go" — no stage gates except the platform's own (Arm/Enter trade-action cards). Per MEMORY.md RUN RULES, the relevant runbook is read as the FIRST tool call at T+0 of any execution.
## Catalog
| Tool / Folder | Purpose | Key Files | When to Use |
| --- | --- | --- | --- |
| `runbooks/spy-0dte-scalp.md` | MARI-native e2e scalp procedure — T+0 gate, recon burst, schema-exact record, arm handoff, manage, stale-kill, post-mortem discipline. | `runbooks/spy-0dte-scalp.md` | FIRST tool call of every live scalp (per MEMORY.md RUN RULES). Not ICM — no stage gates. |
| `workflows/tool-catalog-maintainer/` | Creates and maintains catalog READMEs for directories of tools/resources. | `SKILL.md`, `CONTEXT.md`, `references/catalog-format.md` | Use when asked to catalog, index, or update documentation for a folder of tools. |
| `workflows/research-dfns/` | DFNS research folder — alerts logs, chat log, and TODO tracking. | `TODO.md`, `AlertsLog_*.txt`, `ChatLog_*.txt` | Use when reviewing or continuing DFNS research. |
| `sources/260731-1330-credit-spread.md` | SPY Bull Put Credit Spread trade plan (0DTE, July 31). | `260731-1330-credit-spread.md` | Use for today's SPY options trade — bullish slow grind thesis. |
| `workflows/research-dfns/` | DFNS research folder — alerts logs, chat log, and TODO tracking from July 27, 2026. | `TODO.md`, `AlertsLog_*.txt`, `ChatLog_*.txt` | Use when reviewing or continuing DFNS research. |
| `workflows/backtest-strategy/` | Orchestrates rules-based strategy backtesting for core styles using TTG data and a pure-Python engine. | `CLAUDE.md`, `CONTEXT.md`, `_config/`, `stages/`, `shared/`, `pyproject.toml` | Use when backtesting trading strategies needing review-gated workflow, data caching, and honest reporting. |
| `workflows/scan-watchlist-for-equities/` | Scans equities for trade setups with confluence using a mechanical data-fetch script. | `CLAUDE.md`, `CONTEXT.md`, `_config/`, `shared/scripts/fetch_scan_data.py`, `stages/` | Use when scanning equities for confluence, generating mechanical data briefs, then interpreting for trade plans. |
| `workflows/scan-watchlist-for-options/` | Scans options for trade setups with confluence using a mechanical data-fetch script. | `CLAUDE.md`, `CONTEXT.md`, `_config/`, `shared/` (CSV watchlists, `scan-watchlist.md`), `stages/` | Use when scanning options for confluence, generating mechanical data briefs, then interpreting for trade plans. |
| `sources/260731-1330-credit-spread.md` | SPY Bull Put Credit Spread trade plan (0DTE, July 31, 2026). Sell $740 put / Buy $738 put spread. | `260731-1330-credit-spread.md` | Use for today's SPY options trade — bullish slow grind thesis. |
| `sources/iron-condor-45dte.md` | SPX Iron Condor trade plan — 45DTE entry, 21DTE time exit, Schwab broker. | `iron-condor-45dte.md` | Use when setting up or managing SPX iron condor spreads. |
| `sources/cha-martin-watchlist.ms.csv` | Cha Martin watchlist export (CSV). | `cha-martin-watchlist.ms.csv` | Use for ticker watchlist reference or scan filtering. |
| `sources/cha-martin-watchlist.ms.csv` | Cha Martin watchlist export (CSV) for ticker reference or scan filtering. | `cha-martin-watchlist.ms.csv` | Use for ticker watchlist reference or scan filtering. |
| `sources/2026-08-11-SFY-IC-Debrief.md` | Debrief of a SPY Iron Condor trade on 2026-08-11 with observations and key takeaways. | `2026-08-11-SFY-IC-Debrief.md` | Use for reviewing past iron condor trades to learn and improve. |
| `sources/HANDOFF_scan_workflow_scripts.md` | Documentation of the mechanical data-fetch script for scan workflows (API, MCP, output format). | `HANDOFF_scan_workflow_scripts.md` | Use when understanding or implementing the fetch_scan_data.py script. |
| `sources/servers.md` | List of MCP servers available in the MARI environment (ttg-stocks, ttg-benzinga, etc.). | `servers.md` | Use when referencing which data servers are available for MCP tool calls. |
| `sources/misc/` | Miscellaneous files of interest that may be used as sources (e.g., trade plans, analysis). | Various files (e.g., `xpon-squeeze-2026-08-24.md`) | For referencing miscellaneous source-like files. |
| `tos-short-squeeze/` | Contains a TSV file for short squeeze scanning. | `ShortSqueezeScanner.tsv` | Use for scanning short squeeze opportunities using the provided TSV data. |
| `backtest-workflow-PLAN.md` | Draft plan for a backtesting workflow (approved and built) describing architecture and stages. | `backtest-workflow-PLAN.md` | Use for understanding the design of the backtest-strategy workspace. |
| `build-watcher.md` | Build spec/prompt for a scheduled 9EMA×VWAP cross alert watcher on MESU6 futures. | `build-watcher.md` | Use when building, reviewing, or resuming the MESU6 9EMA/VWAP cross scheduled-alert watcher (or as a template for a similar cross-based futures alert). |
## Tools
@@ -50,13 +66,81 @@ Reviewing or continuing DFNS research. Note: per ticker tracking rule, DFNS shou
**Setup / dependencies:**
None noted.
### `workflows/backtest-strategy/`
**Purpose:**
Orchestrates rules-based strategy backtesting for core styles (SPY 0DTE credit spreads, single-leg equity/options day trades) using TTG data servers as the only market-data source and a local pure-Python engine.
**Contents:**
- `CLAUDE.md` — workspace identity and entry point
- `CONTEXT.md` — workspace-level routing and stage description
- `_config/` — stable references: `risk-params.md`, `strategy-spec-TEMPLATE.md`
- `stages/` — five stage folders (00_clarify_strategy through 04_report) each with review gates
- `shared/` — data cache (`shared/data/`) and engine scripts (`shared/scripts/`)
- `pyproject.toml` — zero dependencies (stdlib-only)
**Use when:**
Backtesting trading strategies using TTG data, needing a review-gated workflow with data caching, honest reporting, and no look-ahead.
**Setup / dependencies:**
None noted (stdlib-only, zero installs; the operator manages packaging with `uv`). Market data comes ONLY from TTG data servers.
**Notes:**
Follows ICM.md — every stage ends at a review gate. Output is read and optionally edited before the next stage runs. Pure Python 3.12 stdlib (json, csv, sqlite3, statistics, math, datetime, urllib). No historical greeks/IV series; backtests are price/levels-driven.
### `workflows/scan-watchlist-for-equities/`
**Purpose:**
Scans tickers for trade setups with confluence, using a mechanical data-fetch script to gather price/volume/bars/S-R data from TTG, producing a data brief for interpretation.
**Contents:**
- `CLAUDE.md` — workspace identity
- `CONTEXT.md` — workspace routing and stage description
- `_config/` — references such as ticker list, default parameters
- `shared/scripts/fetch_scan_data.py` — stdlib-only Python script for fetching and computing mechanical data
- `shared/` — watchlist CSVs (e.g., `cha-martin-watchlist.ms.csv`, date-named CSVs)
- `stages/` — five stage folders (00_clarify through 04_summary_card)
**Use when:**
Scanning equities for trade setups with confluence, generating mechanical data briefs (no interpretation), then interpreting the brief for trade plans.
**Setup / dependencies:**
Python 3.11+, access to TTG data API via `MARI_MCP_CONFIG` environment variable. The fetch script is stdlib-only (urllib, json, etc.).
**Notes:**
The workflow includes a mechanical script (`fetch_scan_data.py`) that fetches and computes data without interpretation. Stages: 00_clarify (ask clarifying questions), 01_broad_filter (runs the script), 02_confluence_analysis, 03_trade_plan, 04_summary_card. The script writes `data_brief.md` and `raw_data.json` to `stages/01_broad_filter/output/`.
### `workflows/scan-watchlist-for-options/`
**Purpose:**
Scans options tickers for trade setups with confluence using a mechanical data-fetch script similar to the equities version.
**Contents:**
- `CLAUDE.md` — workspace identity
- `CONTEXT.md` — workspace routing and stage description (identical to equities version)
- `_config/` — `trade_plan_template.md`
- `shared/` — CSV watchlists (e.g., `cha-martin-watchlist.ms.csv`, date-named CSVs) and `scan-watchlist.md`
- `stages/` — five stage folders (00_clarify through 04_summary_card)
**Use when:**
Scanning options for trade setups with confluence, generating mechanical data briefs, then interpreting the brief for trade plans.
**Setup / dependencies:**
Python 3.11+, access to TTG data API via `MARI_MCP_CONFIG` environment variable. The fetch script is stdlib-only.
**Notes:**
Very similar to the equities workflow but focused on options. The shared folder contains CSV watchlists and a `scan-watchlist.md` file. The mechanical script (if present) would fetch options data; verify the exact script name and location.
### `sources/260731-1330-credit-spread.md`
**Purpose:**
Active trade plan for a SPY Bull Put Credit Spread (0DTE, July 31, 2026). Sell $740 put / Buy $738 put spread, targeting 20%+ profit with bullish slow grind thesis.
**Contents:**
- `260731-1330-credit-spread.md` — trade plan with entry, exit, profit target, stop loss, and monitoring notes.
**Use when:**
Executing or monitoring today's SPY options trade.
Executing or monitoring today's SPY options trade (July 31, 2026).
**Setup / dependencies:**
None noted. Trade-specific, not reusable.
@@ -66,8 +150,11 @@ None noted. Trade-specific, not reusable.
**Purpose:**
Trade plan for SPX Iron Condor through Schwab broker. 45DTE entry, 30-point wings, forced close at 21DTE. Includes profit target (25% of net credit), stop loss (50% of max loss), and monitoring cadence.
**Contents:**
- `iron-condor-45dte.md` — trade plan with entry, exit, risk management, and monitoring details.
**Use when:**
Setting up or managing SPX iron condor spreads with Schwab.
Setting up or managing SPX iron condor spreads with Schwab broker.
**Setup / dependencies:**
Schwab broker session required.
@@ -77,12 +164,117 @@ Schwab broker session required.
**Purpose:**
CSV export of a Cha Martin watchlist — likely contains tickers for scanning or monitoring.
**Contents:**
- `cha-martin-watchlist.ms.csv` — plain CSV with one ticker per line (or columns).
**Use when:**
Referencing tickers for scans, short squeeze candidates, or watchlist filtering.
**Setup / dependencies:**
None noted.
### `sources/2026-08-11-SFY-IC-Debrief.md`
**Purpose:**
Debrief of a SPY Iron Condor trade on 2026-08-11, summarizing trade details, observations, key takeaways, and next times.
**Contents:**
- `2026-08-11-SFY-IC-Debrief.md` — trade summary, plan, observations, asymmetric structure notes, NL3 feed issue, key takeaways.
**Use when:**
Reviewing past iron condor trades to learn from observations and improve future trades.
**Setup / dependencies:**
None noted.
### `sources/HANDOFF_scan_workflow_scripts.md`
**Purpose:**
Documentation of a mechanical data-fetch script for the scan-watchlist-for-equities workflow, detailing API facts, script design, environment, MCP protocol, verified responses, output format, and integration notes.
**Contents:**
- `HANDOFF_scan_workflow_scripts.md` — detailed handoff covering goal, design, environment facts, MCP-over-HTTP protocol, verified response shapes, data brief format, stitching into CONTEXT.md, verification targets, pitfalls, resume checklist, and current state.
**Use when:**
Understanding or implementing the `fetch_scan_data.py` script, or integrating mechanical data fetching into scan workflows.
**Setup / dependencies:**
None noted.
### `sources/servers.md`
**Purpose:**
List of MCP servers available in the MARI environment for tool calls (e.g., market data, news, chat, etc.).
**Contents:**
- `servers.md` — plain list of server names: `chrome-devtools`, `mari-cell`, `ttg-benzinga`, `ttg-chat`, `ttg-crypto`, `ttg-economy`, `ttg-finviz-elite`, `ttg-forex`, `ttg-futures`, `ttg-holygrail`, `ttg-indices`, `ttg-options`, `ttg-platform`, `ttg-stocks`, `ttg-uw`.
**Use when:**
Referencing which data servers are available for MCP tool calls in workflows or scripts.
**Setup / dependencies:**
None noted.
### `sources/misc/`
**Purpose:**
Miscellaneous files of interest that may be used as sources (e.g., trade plans, analysis, data dumps). These are not primary source files but could be referenced for insights or as starting points.
**Contents:**
- `xpon-squeeze-2026-08-24.md` — XPON Short Squeeze — Premarket signal reconstruction (2026-08-24) documenting a micro-cap short squeeze setup with volume spike, range break, premarket gap, and squeeze mechanics.
- Various JSON files containing live trading cockpit data, error fixes, and trade plans from September 2026.
**Use when:**
Referring to miscellaneous source-like files for reference, analysis, or as a starting point for creating new sources.
**Setup / dependencies:**
None noted.
### `tos-short-squeeze/`
**Purpose:**
Contains a TSV file for short squeeze scanning, likely a precomputed list of candidates.
**Contents:**
- `ShortSqueezeScanner.tsv` — tab-separated values with columns likely including ticker, metrics, etc.
**Use when:**
Scanning for short squeeze opportunities using the provided TSV data as input or reference.
**Setup / dependencies:**
None noted.
### `backtest-workflow-PLAN.md`
**Purpose:**
Draft plan for a backtesting workflow (approved and built) describing the architecture, stages, and rules for backtesting core styles using TTG data and a pure-Python engine.
**Contents:**
- `backtest-workflow-PLAN.md` — goal, architecture (data, engine, reporting layers), proposed workspace layout, stage contracts, decision points, known limitations, credit/cost profile, and out-of-scope items.
**Use when:**
Understanding the design of the `backtest-strategy` workspace or as a reference for building similar review-gated backtesting workflows.
**Setup / dependencies:**
None noted.
### `build-watcher.md`
**Purpose:**
Prompt/spec for building a scheduled-alert watcher that fires on a 9EMA×VWAP cross for MESU6 (MES Sep 2026 front month). Signal: 9EMA on 3-minute bar closes (built from 1-minute bars) crossing session VWAP anchored at the 9:30 AM ET open. Cross up = long trigger, cross down = short trigger. Explicitly scoped as a scheduled-alert build, not a Holy Grail analysis.
**Contents:**
- `build-watcher.md` — signal definition, data-source rules, alert discipline, schedule parameters, and credit-cost note.
**Use when:**
Building, reviewing, or resuming work on the MESU6 9EMA/VWAP cross scheduled alert, or as a template for a similar cross-based futures alert watcher.
**Setup / dependencies:**
Requires the TTG futures MCP server (`futures_aggs` for bars, `futures_snapshot` for last price) and a Scheduled job runner. The SCHEDULE section (active window, hard stop, fire cap) contains bracketed placeholders in the current draft — these must be filled in before a job is created, and the confirm card must be reviewed and approved before creation.
**Notes:**
Fires only on an actual state change (no re-alerts); permanently silent after the watch ends or the fire cap is hit. Job pause/delete is user-owned via the Scheduled sidebar — the agent must never pause/delete it itself. Each fired alert costs ~50 Active Trader credits; polls are free.
## Maintenance Notes
When adding or updating a tool folder, update this README with:

94
README.md.backup Normal file
View File

@@ -0,0 +1,94 @@
# Tool Catalog
This README catalogs the tools and resources in this directory. It is intended for humans and LLM agents to quickly identify what is available and when to use it.
## Catalog
| Tool / Folder | Purpose | Key Files | When to Use |
| --- | --- | --- | --- |
| `workflows/tool-catalog-maintainer/` | Creates and maintains catalog READMEs for directories of tools/resources. | `SKILL.md`, `CONTEXT.md`, `references/catalog-format.md` | Use when asked to catalog, index, or update documentation for a folder of tools. |
| `workflows/research-dfns/` | DFNS research folder — alerts logs, chat log, and TODO tracking. | `TODO.md`, `AlertsLog_*.txt`, `ChatLog_*.txt` | Use when reviewing or continuing DFNS research. |
| `sources/260731-1330-credit-spread.md` | SPY Bull Put Credit Spread trade plan (0DTE, July 31). | `260731-1330-credit-spread.md` | Use for today's SPY options trade — bullish slow grind thesis. |
| `sources/iron-condor-45dte.md` | SPX Iron Condor trade plan — 45DTE entry, 21DTE time exit, Schwab broker. | `iron-condor-45dte.md` | Use when setting up or managing SPX iron condor spreads. |
| `sources/cha-martin-watchlist.ms.csv` | Cha Martin watchlist export (CSV). | `cha-martin-watchlist.ms.csv` | Use for ticker watchlist reference or scan filtering. |
## Tools
### `workflows/tool-catalog-maintainer/`
**Purpose:**
Creates and maintains a `README.md` catalog for a directory containing subfolders of tools, prompts, scripts, docs, workflows, skills, or other reusable resources. The README is a living catalog that improves over time.
**Contents:**
- `SKILL.md` — skill definition and workflow
- `CONTEXT.md` — full context: inputs, process, outputs, verification
- `references/catalog-format.md` — catalog README structure and cataloging rules
- `output/` — generated artifacts
**Use when:**
You need to catalog, index, summarize, or update documentation for a folder of tools so humans or LLM agents can quickly choose the right resource.
**Setup / dependencies:**
None noted. Works with file read/write tools.
**Notes:**
Always use this tool when asked to catalog a directory — inspect subfolders, read their key files, and write the README using `references/catalog-format.md`.
### `workflows/research-dfns/`
**Purpose:**
Research folder for DFNS (Digital Frontier Acquisition Corp.) — contains alert logs, chat transcripts, and tracking notes from July 27, 2026.
**Contents:**
- `TODO.md` — research tracking tasks
- `AlertsLog_Mon Jul 27 2026*.txt` — alert logs (3 files)
- `ChatLog_Mon Jul 27 2026.txt` — chat transcript
**Use when:**
Reviewing or continuing DFNS research. Note: per ticker tracking rule, DFNS should not be actively tracked unless a fresh positive reason arises.
**Setup / dependencies:**
None noted.
### `sources/260731-1330-credit-spread.md`
**Purpose:**
Active trade plan for a SPY Bull Put Credit Spread (0DTE, July 31, 2026). Sell $740 put / Buy $738 put spread, targeting 20%+ profit with bullish slow grind thesis.
**Use when:**
Executing or monitoring today's SPY options trade.
**Setup / dependencies:**
None noted. Trade-specific, not reusable.
### `sources/iron-condor-45dte.md`
**Purpose:**
Trade plan for SPX Iron Condor through Schwab broker. 45DTE entry, 30-point wings, forced close at 21DTE. Includes profit target (25% of net credit), stop loss (50% of max loss), and monitoring cadence.
**Use when:**
Setting up or managing SPX iron condor spreads with Schwab.
**Setup / dependencies:**
Schwab broker session required.
### `sources/cha-martin-watchlist.ms.csv`
**Purpose:**
CSV export of a Cha Martin watchlist — likely contains tickers for scanning or monitoring.
**Use when:**
Referencing tickers for scans, short squeeze candidates, or watchlist filtering.
**Setup / dependencies:**
None noted.
## Maintenance Notes
When adding or updating a tool folder, update this README with:
- purpose
- key files and entry points
- usage guidance
- setup requirements
- notable changes or cautions

107
backtest-workflow-PLAN.md Normal file
View File

@@ -0,0 +1,107 @@
# Backtest Workflow — DRAFT PLAN (for review — nothing built yet)
**Status:** APPROVED (2026-08-27 ~11 PM ET) and BUILT — workspace live at
`workflows/backtest-strategy/`. No packages installed (zero-dep pyproject.toml
included; operator manages packaging with uv). No data pulled yet. Engine is
authored at stage-03 exec time, not pre-built.
**Date drafted:** Thu 2026-08-27, ~11 PM ET
**Proposed workspace:** `workflows/backtest-strategy/` (follows ICM.md — folder structure as orchestrator)
---
## Goal
Rules-based backtesting for Gump's core styles, using TTG data servers as the only market-data source and a local pure-Python engine:
1. **SPY 0DTE credit put spreads** (bread and butter — premium-selling, % profit targets)
2. **Single-leg equity/options day trades** (VWAP-pullback longs, 50MA-fade shorts)
3. Multi-target management exactly as traded live: T1/T2/T3 partials, stop tightening, hard time exit (15:45 ET), optional breakeven-after-T1
## Architecture — three layers
### 1. Data layer (TTG servers only — never scraped)
| Need | Source (verified 08/27) | Notes |
|---|---|---|
| Equity bars | custom OHLC bars, daily summaries, grouped-daily (all tickers) | minute-level intraday |
| Options bars | per-contract OHLC bars | minute→month timespans, **back to 2014-06-02**, ≤50,000 bars/pull |
| Contract enumeration | contract list + specs | which strikes/expiries existed on date X |
| Historical quotes | per-contract quote history | **depth unverified — stage 01 verifies before any promise** |
| Historical ticks | per-contract trade history | fallback/verification for quote data |
| Snapshots (live) | chain w/ greeks+IV, per-contract, unified | current-tape only — not for backtests |
**Cache design:** all pulls land in `shared/data/` as SQLite (stdlib `sqlite3`), keyed by (contract, timespan, window). Re-runs read cache first — pulls happen once per window ever. Cache caps enforced so the folder doesn't balloon.
### 2. Engine layer — pure Python 3.12 stdlib (v1: zero installs)
- Bar-by-bar event loop over cached data (`json`, `csv`, `sqlite3`, `statistics`, `math`, `datetime` only)
- Strategies are **declarative spec files** (`_config/strategy-*.md` + params block) — human-readable, reviewable, diffable
- Position model:
- Single leg (long/short equity or option)
- **Two-leg credit spread** = short leg + long leg bars stitched; credit = leg diff at entry; P&L tracked per bar
- Exit model: stop, T1/T2/T3 partial scale-outs, hard time exit, optional breve-after-T1, EOD flat (15:45 ET per house rule)
- Costs: configurable slippage (default **conservative**: mid ± half-spread per fill), optional commissions; RTH default / extended optional flag per session-language rule
- Validation: in-sample/out-of-sample split by date; no look-ahead (signals computed on bars ≤ current bar only)
### 3. Reporting layer
- `trades.csv` — every simulated trade: entry/exit, legs, MAE/MFE, time-in-trade, exit reason
- `metrics.json` — win rate, expectancy, avg win/loss, profit factor, max drawdown, per-hour-of-day and per-weekday breakdowns
- Final stage renders a readable report + summary card
## Proposed ICM workspace layout
```
workflows/backtest-strategy/
├── CLAUDE.md # entry point + exec protocol (read ICM.md fresh, clear output/ on re-run)
├── CONTEXT.md # workspace context: routing + data rules (TTG-only, no fabrication)
├── _config/
│ ├── strategy-spec-TEMPLATE.md
│ └── risk-params.md # default slippage, session, sizing
├── shared/
│ └── data/ # SQLite cache (built during runs)
└── stages/
├── 00_clarify_strategy/ # capture rules in plain English → strategy_spec.md
├── 01_verify_data/ # enumerate contracts, verify quote-history depth → data_manifest.md
├── 02_fetch_cache/ # pull bars/quotes → shared/data/ → fetch_log.md
├── 03_run_backtest/ # engine executes spec → trades.csv + metrics.json
└── 04_report/ # human-readable report + summary card → report.md
```
Every stage ends at a **review gate** — output/ is read and (if needed) edited before the next stage runs, per ICM.
## Stage contracts (summary)
| Stage | Reads | Does | Writes |
|---|---|---|---|
| 00_clarify | _config templates, member Q&A | freeze strategy rules + params | strategy_spec.md |
| 01_verify_data | strategy_spec.md | enumerate contracts; **test-pull quote history depth**; flag gaps | data_manifest.md |
| 02_fetch_cache | data_manifest.md | pull bars/quotes → cache (single pass, no repeats) | fetch_log.md |
| 03_run_backtest | cache + spec | run engine, no look-ahead, costs applied | trades.csv, metrics.json |
| 04_report | trades + metrics | render findings honestly (incl. small-sample warnings) | report.md |
## Decision points — need Gump's call
1. **v1 compute:** pure stdlib, zero installs (recommended) — or approve `pip install pandas numpy` now?
2. **First strategy to backtest:** (a) SPY 0DTE credit put spread *(recommended — core style)*, (b) VWAP-pullback equity long, (c) 50MA fade short?
3. **Data window:** propose **last 6 months** of SPY 0DTE contracts for v1 (cache stays lean; extendable later)?
4. **Slippage default:** conservative mid ± half-spread, tunable in `_config/risk-params.md` — OK?
5. **Sizing model:** fixed 1-contract (clean signal measurement) vs fixed-dollar risk? Propose fixed 1-contract for v1.
## Known limitations (stated up front)
- **No historical greeks/IV series** — backtests are price/levels-driven; IV-rank conditions are not testable
- **No order-book replay** — fills modeled at bar close with slippage knob; inherently approximate, slightly optimistic
- **Quote-history depth TBD** — stage 01 measures it before we rely on it
- **0DTE dailies** only exist for the era they traded; earlier "0DTE" = nearest weekly
- Small-sample honesty: 6 months of 0DTE ≈ ~125 trading days — stage 04 will flag low-N results as such
## Credit/cost profile
- All pulls are read-only TTG market-data calls (conservation mode is OFF per member).
- Heavy stage is 02; the cache means any re-run/re-parameterization costs **zero** additional pulls.
- Rough v1 pull count: ~125 contracts × 2 legs × 1 window, each well under the 50k-bar cap.
## Explicitly out of scope for v1
- Portfolio-level / multi-strategy simulation, options greeks modeling, intrabar stop sequencing (stops checked bar-by-bar, close-based), live-paper forwarding (backtest ≠ trade plan — any live trade still goes through the normal cockpit path)

31
build-watcher.md Normal file
View File

@@ -0,0 +1,31 @@
Build and implement a 9EMA×VWAP cross watcher for MESU6 (MES Sep 2026 front month).
Do not run a Holy Grail analysis — this is a scheduled-alert build. Plan it, show me
the confirm card, and I approve before anything is created.
SIGNAL
- Compute 9EMA (my TL) on 3-minute bar closes, built from 1-minute MESU6 bars.
- Compute session VWAP anchored at the 9:30 AM ET day-session open.
- Cross UP (9EMA above VWAP, was below) = long trigger. Cross DOWN = short trigger.
- Use the TL×VWAP crossover as the signal, exactly as drawn on my 3-min MES charts.
DATA RULES
- All market data from TTG futures MCP only: futures_aggs for bars, futures_snapshot
for the exact last price ({ticker: "MESU6"}). Never estimate or approximate prices.
ALERT DISCIPLINE
- Alert only on an actual cross (state change). No re-alerts while state is unchanged.
- First-match-only, no redundant snapshot calls when nothing changed.
- Alert text: under 3 sentences, exact price, e.g. "MESU6 9EMA crossed UP over VWAP
@ 7691.25 (long trigger) — 3:42 PM ET."
- Once the watch ends, it goes permanently silent — every later run is a no-op.
- Job pause/delete is owned by me via the Scheduled sidebar; you never pause it yourself.
SCHEDULE
- Poll every 1 minute.
- Active window: [START — e.g., "tomorrow 9:30 AM ET" or "now, evening session"]
- Hard stop: [DEADLINE — e.g., "8:45 PM ET" / "4:00 PM ET"], after which permanent silence.
- Fire cap: [MAX ALERTS — e.g., "4 per session", then silent until next session].
CREDITS
- Each fired alert costs ~50 Active Trader credits; polls are free. Confirm no
per-run billing on the confirm card before creating.

View File

@@ -0,0 +1,64 @@
# SPY 0DTE Scalp Runbook — E2E Fast Path (MARI-native)
*v2 — 2026-09-01 13:45 ET. v1 built 2026-08-28 from spy-1787942969793 (closed un-filled at 3:00 wall). v2 folds in the 2026-09-01 double-failure post-mortem (spy-1788280948132, spy-1788282204586 — both discarded un-filled, zero fills, $0 cost). Lives in `runbooks/` — MARI-native execution jurisdiction, NOT ICM: no stage-review gates; the platform's Arm/Enter cards are the only human gates. Per MEMORY.md RUN RULES, reading this file is the FIRST tool call of every scalp, before recon, before record — no exceptions.*
## The speed rule
**Never discover inside a trade window.** Design + record the plan BEFORE the intended entry window; arm = one pasted line. Measured 2026-09-01: prompt→ready 5m41s (two full ~700KB chain pulls + one schema retry) then ~21 min dead in the arm handoff = setup near-invalidation by ready time. Targets under this runbook: **prompt→plan-ready ≤90s, prompt→armed ≈2–3 min.** The Arm card click is the irreducible floor (app-owned pack preflight, safety rail) — everything else compresses.
## 0. T+0 gate + handshake (v2 — this failed twice before it was written down)
- **First tool call of the run = read this runbook.** It must land visibly in the transcript before any recon or record call.
- **Handshake:** member parks the Live Trading window OPEN + CONNECTED *before* saying go. The go prompt carries symbol + account, nothing else.
- MARI never asks "which window are you in" mid-flight. If the cockpit is not connected at arm time, the window is already burning → stale-kill per §4, re-stage, don't negotiate.
## 1. Recon burst (≤60s, ONE parallel tool block)
- Same block: `snapshot(SPY)` + `ttg_trade_workflow_prepare_context` + `options_snapshot_chain` for **ONE side only** — the side the snapshot's direction call already picked. Never both chains. **Never a mid-flow re-pull, including at stale checks** (the 9/1 re-validation pull at 12:53 was pure waste; snapshot + last known quotes answer staleness).
- Direction call: price vs VWAP + day-low/high structure. Trend-day tape → trade the break (continuation), not the bounce. Mid-range chop → stand down. Bounce in progress toward invalidation → do NOT chase; record a CONDITIONAL plan whose entry zone only fills on a re-press (9/1: zone 0.42–0.58 on 761P after a 4-minute bounce), or stand down entirely.
- Liquidity check (both expiries, Friday rule):
- `ttg-options` → `options_snapshot_chain`, params: `underlyingAsset: "SPY"`, `contract_type: "put"|"call"`, `expiration_date` EXACT string "2026-09-01"-style (range objects = HTTP 500), `limit: 250`, `sort: "strike_price"`.
- 0DTE near-money: 1–2¢ spreads, 300–600K vol = elite. Next-week: premium ~$4 → +20% in 15 min needs ~$1.60 SPY move vs ~$0.30 on 0DTE. 0DTE is the 15-min vehicle; next-week is for plays/swings.
- Vehicle pick: ATM-ish strike, delta −0.3 to −0.5 at expected trigger price, spread ≤ 2¢.
## 2. Record the plan (schema-exact v2, FIRST TRY)
`ttg_trade_workflow_record_plan` — template (validated shape, both 8/28 and 9/1):
```
brokerId: "tos-paper", connectedBrokerPackId: "tos-paper",
accountCode: "D-67336185" // 185; "D-67336186" = 186
assetClass: "option", symbol: "SPY", side: "buy",
tradeStyle: "scalp", // NOT "style"
instrumentRef: {symbol: "SPY", strike: 761, type: "put",
expiration: "2026-09-01"}, // MUST be object, never string
levels: {
entry: {zone: {low: 0.42, high: 0.58}, type: "limit"}, // premium zone
stop: {offset: 0.10}, // POSITIVE number — it's a distance below fill.
// 9/1 cost 45s on "-0.1": rejected, full retry cycle.
targets: [{offset: 0.12, portionPct: 50}, {offset: 0.18, portionPct: 50}], // max 2
timeStopSec: 900, // REQUIRED for scalp/day; in-trade from fill
breakevenAfterT1: true
},
risk: {maxDollars: 20}, // engine computes qty — NEVER send qty
timeLimitSec: 600, // ENTRY WINDOW seconds (NOT the time stop)
rationale: "≥40 chars — SPY-level triggers, invalidations, walls, order-safety directives here"
```
Validation gotchas (each has cost an iteration): `tradeStyle` not `style`; no top-level entry/stop/target/qty; `instrumentRef` object not string; **stop offset positive**; engine enforces **R:R > 1 — every target offset must exceed the stop offset**.
Response: planId `spy-<epoch-ms>`, status draft in PLANS/. SPY-level triggers go in rationale.
## 3. Arm — CONFIRMED 2026-09-01 (twice), no longer a hypothesis
- **Main-chat "arm" CANNOT fire the card.** The Arm card exists ONLY in the Live Trading window's own MARI chat. Member types `arm <planId>` THERE → card renders → Arm click (preflight auto-runs) → Enter.
- MARI waits via `await_event` (revision change), never sleep-polling; plan-ready message carries planId + the literal arm line so the paste is 5 seconds.
- **NEVER fake arming from main chat** (`update_plan status:"open"` skips preflight/stream verify — dangerous). Main-window MARI cannot arm; the cockpit orchestrator owns execution.
- **Stale-kill:** if not armed by the entry-window deadline → `ttg_trade_workflow_discard_plan` (verified live 3× on 9/1), stand down, never manage a dead window, never chase the bounce afterward.
- **Post-flight rule (added 9/1 14:31, after third un-armed discard):** stale-kill ends the flight. Debrief question AFTER the kill, never mid-flight: was the member in the cockpit at plan-ready? (a) Yes + arm line didn't render → platform delta, fold into §3. (b) No → handshake breach: "go" means member is IN the cockpit chat at that moment. MARI never asks mid-flight, never re-stages on a spent "go" — re-stage only on the member's renewed go.
- **ARM RENDER CONFIRMED 9/1 14:38 (debrief):** cockpit paste → card DID render, with lag ("eventually"). The 3rd kill was a RACE: card in flight when the deadline hit. Fix = paste-confirm handshake: plan-ready message asks for one-word reply "pasted" in MAIN chat the moment the arm line is in the cockpit chat. "Pasted" ⇒ kill deadline extends +2 min (render-lag budget: card render + Arm click + preflight). No "pasted" by deadline ⇒ kill stands, no renegotiation.
- **Card-after-kill:** if the Arm card renders after a stale-kill, the plan is already discarded (discard is authoritative; verified status=draft at kill). DO NOT click Arm — dismiss/ignore the card. No orders can exist; flat is guaranteed.
## 4. Manage (standing protocol — member directives, encode in rationale every time)
1. On fill: verify BOTH working and filled orders. Stop must read **SELL TO CLOSE <qty>** — never sell-to-open, never a short entry.
2. T1 fills → scale stop to remaining qty (+ breakevenAfterT1 moves it to BE). T2/stop-out → **cancel ALL working orders the instant flat or profit taken.** Working + filled checked both after every event. No orphans, ever.
3. Walls: `timeStopSec` from fill + explicit wall-clock hard flat in rationale. Bail on tape deterioration before the stop (slow grind back = failed setup).
4. Management loop = `await_event` cycles at 1s belt/watcher cadence; terse timestamped updates only while the window is live.
## 5. Post-mortem discipline (v2)
After every e2e run — fill or not: procedure deltas fold into THIS file; MEMORY.md Lessons gets a one-line pointer only. **Lessons = history + pointers. Runbook = procedure.** A procedure that lives only in Lessons will lose to run momentum — verified 9/1 when the loaded lesson was ignored mid-flight.
## 15-min target math (10–25% premium band)
- 0DTE ATM-ish premium P, delta Δ: +20% ≈ 0.20·P / Δ SPY move. At P≈$0.85, Δ≈−0.50 → ~$0.34 SPY. Routine in 15 min on a break.
- Engine R:R rule forces target offset > stop offset: size stop −10 to −16% and T1 +19 to +22% to stay inside the band AND clear R:R>1 (verified combos: stop 0.18 / T1 0.20 / T2 0.40 on ~$0.84 fill; stop 0.10 / T1 0.12 / T2 0.18 on ~$0.54 mid).

1305
sources/ICM.md Normal file

File diff suppressed because it is too large Load Diff

View File

@@ -0,0 +1,29 @@
New watchlist. Interview me, then build and run it.
INTERVIEW — exactly 6 questions, quick-answer buttons (or numbered choices in chat), one at a time:
1. Size bucket: Large ($10B+) | Mid ($2B–10B) | Small ($300M–2B) | Micro (<$300M) | Mixed
2. Setup (max 2): Short squeeze (high SI + gap) | Gap & momentum | Near 50MA confluence |
VWAP pullback | Base/measured-move breakout | Volume breakout
3. Universe: Full US market | Cha Martin watchlist | My scratchpad file | Sector/theme (free text)
4. Liquidity floor: $1M | $5M | $20M | $100M avg daily dollar volume | No floor
5. List size: 5 | 10 | 15 | 20 tickers
6. Cadence: On-demand only | Daily pre-market 9:15 ET (scheduled job, needs my approval) |
Twice daily (9:15 + 11:30 ET) | Intraday every 2h RTH
Derive the profile name from the answers (e.g. "smallcap-squeeze"). Do not ask more than 6
questions; fold anything else into sensible defaults and state them.
ARTIFACTS — after the answers:
- Save the profile as .mari/scratchpad/watchlists/configs/<name>.json
(fields: name, created, updated, size_bucket, horizon, screens, min_dollar_volume,
market_cap bounds, universe, list_size, sector_rule "avoid_concentration", cadence, notes).
- Never edit _template/config.json; copy its schema.
- Run the first scan immediately against live TTG data (snapshot day.vw = session VWAP,
day.dv = today dollar volume; market cap verified per finalist via stocks_reference_ticker;
ETFs excluded unless I asked for them).
- Write results to .mari/scratchpad/watchlists/results/<name>-<YYYY-MM-DD>.md plus -latest.md
mirror, and show me the table in chat: Ticker | Price | VWAP or ref level | % from level |
$Vol | Mkt Cap | one-line note, plus one alternate and any rejections with reasons.
RERUN SEMANTICS — going forward, "run my <name> watchlist" = same screens, fresh data, new
dated file. "Patch <name>: <change>" = edit one field of the JSON, keep history.
If the button-card renderer fails, fall back to numbered in-chat choices — same flow.

View File

@@ -0,0 +1,36 @@
Run my largecap-vwap-pullback watchlist.
PROFILE (the 6 answers, saved at .mari/scratchpad/watchlists/configs/largecap-vwap-pullback.json):
1. Size bucket: Large cap, market cap ≥ $10B (verify per candidate via ticker details — never assume)
2. Setup: VWAP pullback — price within ±0.35% of live session VWAP, session high at least
+0.5% above VWAP (ran up, pulled back to the line), price still above/near VWAP
3. Universe: full US market (exclude ETFs — require type CS or ADRC in ticker details)
4. Liquidity: today's dollar volume ≥ $5M (use $25M to pre-narrow, verify $5M floor on finalists)
5. List size: top 5, ranked by push-above-VWAP depth then dollar volume; max 2 per sector
6. Cadence: on-demand only ("run my <name> watchlist")
DATA PATH (TTG MCP servers only):
- One call to ttg-stocks stocks_snapshot_all {include_otc:false} → per-ticker day.vw IS the
live session VWAP, day.dv is today's dollar volume, day.h is the session high. No per-ticker
bar calls needed. Parse: find '{"mcp_ui"' in the dump → json.loads → ["json"]["tickers"].
- Market cap only comes from stocks_reference_ticker (market_cap field) — call it per finalist,
after narrowing, never before.
- Screen passes → market-cap + type verification → reject anything under $10B or any ETF
(watch for lookalikes: leveraged ETFs like QID/LABD/AAPD, commodity/crypto ETFs like BITO/ETHA).
ARTIFACTS:
- Save/refresh results at .mari/scratchpad/watchlists/results/<name>-<YYYY-MM-DD>.md,
mirror to <name>-latest.md, keep dated history.
- Table columns: Ticker | Price | VWAP | % vs VWAP | % High-over-VWAP | $Vol | Mkt Cap | Note.
- Include an alternate (6th name) and list any near-miss rejections with the reason (e.g.
"AAL cut: $8.6B mcap under the $10B bar").
OUTPUT: the table + one-line trigger note (VWAP reclaim/hold = long trigger, stop under session
low). Terse. No re-asking the 6 questions — they live in the config file.
Variants — swap only line 1 and the market-cap bar:
Small-cap list: size bucket: small ($300M–2B) → bar becomes ≥$300M, <$2B
Mid-cap: ≥$2B, <$10B
Different setup: replace the screen definition in item 2 (e.g., "gap & momentum: day open ≥2% above prev close, dollar volume ≥ 2× 20-day average") — everything else stays
The profile JSON on disk is the source of truth; this prompt is just its human-readable twin. Say "run my largecap-vwap-pullback watchlist" any time and I execute exactly this.

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

View File

@@ -0,0 +1,86 @@
ShortSqueeze SqueezeScan # TOS Short Squeeze Scanner
# Run in: TOS -> Tools -> Run Scan (or save to Scan Library)
# Data period: Pre-Market (select in scan data settings)
# Columns: Symbol, Ticker, SI%(manual), Squeeze, Gap%, PremkVol, PremkVolPct, PremkRange%, PremkAboveVWAP, PremkBid, PremkAsk, PremkMid
#
# HOW TO USE:
# 1. TOS -> Tools -> Run Scan. Paste this entire file into the "ThinkScript" box.
# 2. Set data period: Pre-Market. Exclusions: None.
# 3. Universe: US stocks, price > $1, volume > 0.
# (For a tighter list, add a static pre-filter for float < 30M shares.)
# 4. Run. Sort by SqueezeScore.
# 5. Manually enter each candidate's short float % (SI%) using TOS stock
# details page, S3 Partners, or your usual source. This column feeds
# the SqueezeScore but the technical flags work without it.
# 6. Best candidates: Squeeze=1 AND PremkAboveVWAP=0 AND PremkVolPct > 1.5
# AND SI% > 25.
input gapUpMin = 1.0; # min % gap up from prev close
input gapUpMax = 12.0; # max % gap up (above this = news/pop already done)
input volPctMin = 1.5; # min premarket volume as % of 1-day avg
input vwAPBand = 0.75; # % band around VWAP where "hug" is valid
input squeezeScoreMin = 1; # min score to show (1-5 scale)
# ---- Price levels (premarket context) ----
def prevClose = Close[-1];
def open = Open;
def high = High;
def low = Low;
def close = Close;
def vol = Volume;
def gapPct = if prevClose > 0 then (open - prevClose) / prevClose * 100.0 else Double.NaN;
# ---- Pre-market VWAP approximation ----
# Built from the day's bars so far, so it includes premarket prints.
def vwapCalc = if Sum(vol, 1) > 0 then Sum(close*vol, 1) / Sum(vol, 1) else Double.NaN;
def vwapPctFromClose = if vwapCalc > 0 then (close - vwapCalc) / vwapCalc * 100.0 else Double.NaN;
# ---- Pre-market volume as % of 1-day average ----
def avgVol1D = Average(vol, 22); # ~1 trading day of 1-min bars
def volPct = if avgVol1D > 0 then vol / avgVol1D * 100.0 else Double.NaN;
# ---- Pre-market range as % of prev close (narrow range = coiling) ----
def rangePct = if prevClose > 0 then (high - low) / prevClose * 100.0 else Double.NaN;
# ---- Is price hugging/just under VWAP? ----
def huggingVWAP = vwapCalc > 0 and vwapPctFromClose >= -vwAPBand and vwapPctFromClose <= vwAPBand;
# ---- Squeeze flags ----
def flagGap = gapPct >= gapUpMin and gapPct <= gapUpMax;
def flagVol = volPct >= volPctMin;
def flagHug = huggingVWAP;
def flagNarrow = rangePct <= 2.5; # narrow range = coiling, ready to spring
# ---- Squeeze score: 0-5 (add SI% manually in the column) ----
def score = (if flagGap then 1 else 0)
+ (if flagVol then 1 else 0)
+ (if flagHug then 1 else 0)
+ (if flagNarrow then 1 else 0)
+ 0; # SI% is manual - add +1 in your head if SI > 25
# ---- Composite squeeze signal ----
def squeezeSignal = flagGap and flagVol and flagHug and squeezeScore >= squeezeScoreMin;
# ---- Output columns ----
plot Ticker = GetSymbol();
plot Symbol = GetSymbol();
plot SI_Pct = Double.NaN; # MANUAL: paste short float % per ticker
plot Squeeze = squeezeSignal ? 1 : 0;
plot GapPctOut = gapPct;
plot PremkVol = vol;
plot PremkVolPct = volPct;
plot PremkRangePct = rangePct;
plot PremkAboveVWAP = if vwapPctFromClose > 0 then 1 else 0;
plot PremkVWAPPct = vwapPctFromClose;
plot PremkBid = BidPrice;
plot PremkAsk = AskPrice;
plot PremkMid = (BidPrice + AskPrice) / 2;
plot SqueezeScore = score;
# ---- Scanner conditions (filter rows) ----
# Adjust to taste. Starting: at least a gap-up + elevated volume + VWAP hug.
condition SqueezeSetup = squeezeSignal;
condition GapUpOnly = flagGap and flagVol;
condition VolSpikeOnly = flagVol;
Can't render this file because it has a wrong number of fields in line 2.

1
tri_scan_cands.json Normal file

File diff suppressed because one or more lines are too long

19
watchlists/README.md Normal file
View File

@@ -0,0 +1,19 @@
# Watchlists — reusable scan configs
Each watchlist is a small JSON config. Say "run my <name> watchlist" (or "update my <name> watchlist")
and MARI re-runs the scan against the saved config, then refreshes the results file.
Say "new watchlist" to answer the 7-question card again and create a new config.
## Layout
- `_template/config.json` — schema/template (copy, never edit)
- `configs/<name>.json` — saved watchlist definitions
- `results/<name>-<YYYY-MM-DD>.md` — dated scan output; latest is also mirrored to `results/<name>-latest.md`
## Re-run semantics
- Same config → same screens → dated results file. Old results are kept for history.
- "Update my watchlist" = re-run with today's data. Config can be patched via chat ("make it midcaps now").
- MARI pulls market data only from TTG MCP servers; screens use ADV, price, gap, short interest,
distance-to-50MA, VWAP structure, and news-catalyst checks as the config directs.
## Config fields (see _template/config.json)
size_bucket, horizon, screens, min_dollar_volume, price_range, catalyst_rule, list_size, sector_rule, notes.

View File

@@ -0,0 +1,14 @@
{
"name": "",
"created": "",
"updated": "",
"size_bucket": "large | mid | small | micro | mixed",
"horizon": "intraday | swing | position",
"screens": ["short_squeeze", "gap_momentum", "near_50ma", "vwap_pullback", "base_breakout", "volume_breakout"],
"min_dollar_volume": 0,
"price_range": { "min": 0, "max": null },
"catalyst_rule": "require_news | exclude_earnings_5d | no_filter",
"list_size": 10,
"sector_rule": "no_constraint | avoid_concentration | prefer:<sector>",
"notes": ""
}

View File

@@ -0,0 +1,23 @@
{
"name": "largecap-vwap-pullback",
"created": "2026-09-02",
"updated": "2026-09-02",
"size_bucket": "large",
"horizon": "intraday",
"screens": [
"vwap_pullback"
],
"min_dollar_volume": 5000000,
"price_range": {
"min": null,
"max": null
},
"market_cap_min_usd": 10000000000,
"catalyst_rule": "no_filter",
"list_size": 5,
"universe": "full_us_market",
"sector_rule": "avoid_concentration",
"cadence": "on_demand",
"notes": "Member intake 2026-09-02 via 6-question flow. Large caps ($10B+), full US market, >=$5M ADV dollar volume, top 5, VWAP pullback setup (price above VWAP earlier in session, pulling back to/through VWAP with hold). Rerun phrase: 'run my largecap-vwap-pullback watchlist'.",
"lastRun": "2026-09-03T14:28-04:00"
}

View File

@@ -0,0 +1,20 @@
# Watchlist: largecap-vwap-pullback — 2026-09-02 (3:05 PM ET re-run)
Profile: large caps ($10B+), full US market, ≥$5M dollar volume, VWAP-pullback setup, top 5, on-demand refresh.
Screen: price within ±0.35% of session VWAP, session high ≥ +0.5% above VWAP, dollar volume ≥ $5M, market cap ≥ $10B verified per finalist, ETFs excluded, max 2 per sector, ranked by push-above-VWAP depth.
## Top 5
| # | Ticker | Price | VWAP | vs VWAP | High vs VWAP | $ Vol | Mkt Cap | Note |
|---|--------|-------|------|---------|--------------|-------|---------|------|
| 1 | SMCI | 36.60 | 36.55 | +0.15% | +3.64% | $25M | $24B | Deepest push on the board, back at the line |
| 2 | NIO | 3.88 | 3.88 | -0.24% | +2.20% | $52M | $10.2B | Right at VWAP, big intraday push earlier |
| 3 | TSLA | 352.71 | 353.54 | -0.23% | +2.00% | $26M | mega | Slightly under VWAP — needs reclaim |
| 4 | SPCX | 140.11 | 140.48 | -0.26% | +1.86% | $38M | $1.87T | Repeat visitor; testing VWAP from above all day |
| 5 | NVDA | 224.40 | 224.38 | +0.01% | +1.59% | $120M | mega | Pinched on VWAP, highest $volume by 3x |
Alternate: NOK 9.83 (+0.08%, push +1.55%, $38M, $56B) — held out to keep tech count at 2.
Rejected: SNXX (Tradr 2x Long SNDK ETF — not a stock), AAL ($8.6B mcap, under the $10B bar), F (auto-sector concentration — would be the 3rd auto with NIO+TSLA), CDE (market-cap screen).
Note vs 2:35 PM run: AAPL, PFE, VALE have drifted >0.35% off VWAP or lost the push; NVDA entered as it settled onto VWAP.
Trigger template: reclaim/hold of VWAP, stop under session low or VWAP-minus.

View File

@@ -0,0 +1,20 @@
# largecap-vwap-pullback — 2026-09-03 (run ~2:28 PM ET, afternoon re-run)
Universe: full US market, non-OTC. Screen: price within ±0.35% of live session VWAP, session high ≥+0.5% above VWAP, price ≥ VWAP−0.35%, $dv ≥ $5M (pre-narrow $25M → relaxed to $5M to fill list), mcap ≥ $10B verified per finalist, ETFs excluded (CS/ADRC only), max 2/sector. Ranked by push-above-VWAP depth, then $vol.
| # | Ticker | Price | VWAP | % vs VWAP | % High-over-VWAP | $Vol | Mkt Cap | Note |
|---|--------|-------|------|-----------|------------------|------|---------|------|
| 1 | BSX | 47.22 | 47.36 | -0.30% | +4.01% | $11M | $70.1B | Med devices, day -2.4%, low 46.97 |
| 2 | MRNA | 145.22 | 145.10 | +0.08% | +3.51% | $9M | $60.2B | Biotech, day -4.0%, low 142.78 |
| 3 | U | 42.295 | 42.19 | +0.24% | +3.19% | $7M | $17.9B | Unity/software, day +3.8%, low 41.00 |
| 4 | KHC | 25.475 | 25.47 | +0.03% | +2.48% | $13M | $31.1B | Consumer staples, day -3.0%, low 25.30 |
| 5 | CCL | 23.415 | 23.42 | -0.02% | +2.27% | $13M | $32.5B | Cruises, day -1.4%, low 23.26 |
| 6 (alt) | BBD | 3.46 | 3.467 | -0.19% | +1.83% | $30M | $36.5B | Brazil bank ADR (ADRC); also F $56.4B push +0.8% $28M |
**Rejected (near-miss):** PLTD/METU/DRAM-type ETFs; KEEL $1.9B, PLUG $2.9B, AAL $8.7B, PATH $9.3B all under $10B bar; NTSK/FLO/GRAB/OPEN/RXRX/KLAR/SOUN under bar or unverified; GLXY unverified ticker (likely ETF-lookalike).
**Afternoon vs 10:23 run:** MU/IREN/NOK ran well above VWAP band (no longer pullbacks); CNH faded to -0.54% vs VWAP (off screen).
**Trigger:** long only on VWAP reclaim/hold (tick back above live VWAP with follow-through); stop under session low. No reclaim = no trade.
Data: stocks_snapshot_all @ ~14:28 ET; mcap/type via stocks_reference_ticker per finalist.

View File

@@ -0,0 +1,18 @@
# largecap-vwap-pullback — 2026-09-03 (run ~10:23 ET)
Universe: full US market, non-OTC. Screen: price within ±0.35% of live session VWAP, session high ≥+0.5% above VWAP, price ≥ VWAP−0.35%, $dv ≥ $5M (pre-narrow $25M → relaxed to $5M to fill list), mcap ≥ $10B verified per finalist, ETFs excluded, max 2/sector. Ranked by push-above-VWAP depth, then $vol.
| # | Ticker | Price | VWAP | % vs VWAP | % High-over-VWAP | $Vol | Mkt Cap | Note |
|---|--------|-------|------|-----------|------------------|------|---------|------|
| 1 | MU | 936.84 | 938.22 | -0.15% | +2.30% | $8M | $1.08T | Semis, day -2.0%, low 918.8801 |
| 2 | CNH | 13.6284 | 13.59 | +0.30% | +2.23% | $5M | $22.0B | Ag/Industrials, day -0.2%, low 13.38 |
| 3 | IREN | 39.57 | 39.49 | +0.21% | +1.72% | $9M | $15.6B | Crypto-miner/AI, day -0.1%, low 39.03 |
| 4 | NOK | 9.675 | 9.66 | +0.13% | +1.27% | $21M | $55.1B | Telecom equip, day -1.6%, low 9.54 |
| 5 | F | 14.225 | 14.21 | +0.07% | +0.95% | $6M | $56.4B | Autos, day +0.6%, low 14.06 |
| 6 (alt) | INTC | 89.34 | 89.07 | +0.30% | +0.86% | $16M | $476B | 2nd semis slot; NVDA cut: sector cap (3rd semis) |
**Rejected (near-miss):** SPWR $0.34 microcap; SOXL/GDX/SCHD/QID ETFs; NVD $3.68 / ONDS / PLUG / AUR under $10B; PATH cut: $9.3B mcap under bar; BTG cut: $7.1B mcap under bar; DRAM unverified ticker (likely ETF); AUR also member-designated exclusion; NVDA sector cap.
**Trigger:** long only on VWAP reclaim/hold (tick back above live VWAP with follow-through); stop under session low. No reclaim = no trade.
Data: stocks_snapshot_all @ ~10:23 ET; mcap/type via stocks_reference_ticker per finalist.

View File

@@ -0,0 +1,18 @@
# largecap-vwap-pullback — 2026-09-03 (run ~14:32 ET)
Universe: full US market, non-OTC. Screen: price within ±0.35% of live session VWAP, session high ≥+0.5% above VWAP, $dv ≥ $5M, mcap ≥ $10B verified per finalist, ETFs excluded, max 2/sector. Ranked by push-above-VWAP depth, then $vol. 53 raw hits → verified finalists.
| # | Ticker | Price | VWAP | % vs VWAP | % High-over-VWAP | $Vol | Mkt Cap | Note |
|---|--------|-------|------|-----------|------------------|------|---------|------|
| 1 | BSX | 47.22 | 47.36 | -0.30% | +4.02% | $11M | $70.1B | Med-device, day -2.4%, deep push reclaim off low 46.97 |
| 2 | MRNA | 144.82 | 145.10 | -0.20% | +3.51% | $9M | $60.2B | Biotech, day -3.8%, big intraday push off low 142.78 |
| 3 | U | 42.295 | 42.19 | +0.24% | +3.19% | $7M | $17.9B | Software, day +3.9%, trending name holding above VWAP |
| 4 | KHC | 25.475 | 25.47 | +0.03% | +2.48% | $13M | $31.1B | Packaged food, day -3.0%, tight coil at VWAP |
| 5 | CCL | 23.415 | 23.42 | -0.02% | +2.27% | $13M | $32.5B | Cruises, day -1.4%, dead-on VWAP |
| 6 (alt) | RKLB | 63.43 | 63.53 | -0.16% | +2.16% | $15M | $40.3B | Space; kept as alt only — 2-day fresh momentum already extended |
**Rejected (near-miss):** MARA $4.0B mcap; CAG $7.8B mcap; NCLH $7.15B mcap; NTSK/PLTD/FLO/GRAB/KLAR/GLXY/HL/SOUN/CDE/KEEL/FRMI/EQX/GGB/BBD/SBS/ABEV/JBLU/HTZ under $10B or microcap; XLB/METU ETFs; PFE cut on sector cap (3rd healthcare w/ BSX+MRNA); NFLX/ITUB/CMCSA/NVO lower push depth.
**Trigger:** long only on VWAP reclaim/hold (tick back above live VWAP with follow-through); stop under session low. No reclaim = no trade.
Data: stocks_snapshot_all @ ~14:31 ET; mcap/type via stocks_reference_ticker per finalist.

View File

@@ -0,0 +1,32 @@
# Backtest Strategy ICM Workspace
This workspace orchestrates rules-based strategy backtesting for the operator's core
styles (SPY 0DTE credit spreads, single-leg equity/options day trades) using TTG data
servers as the only market-data source and a local pure-Python engine.
The agent follows the numbered stages to freeze strategy rules, verify data depth,
build the data cache, run the backtest, and produce an honest report.
Folder structure:
- CLAUDE.md (Layer 0): workspace identity
- CONTEXT.md (Layer 1): workspace-level routing
- stages/: numbered stage folders
- 00_clarify_strategy/: freeze rules into a spec (with operator)
- 01_verify_data/: enumerate contracts + verify historical depth
- 02_fetch_cache/: pull bars/quotes into the shared cache
- 03_run_backtest/: run the engine over cached data
- 04_report/: render findings with small-sample honesty
- _config/: Layer 3 reference material (stable across runs)
- shared/: data cache (SQLite) + engine scripts
- Each stage's output/ holds Layer 4 working artifacts for handoff to next stage.
## Hard rules (apply to every stage)
- Market data comes ONLY from TTG data servers. Never scrape or substitute public sites.
- Python is stdlib-only. NO package installs. The operator manages packaging with uv
(pyproject.toml in the workspace root has zero dependencies by design).
- No look-ahead: a signal on bar N may only use bars <= N.
- Every stage ends at a review gate. output/ is read (and optionally edited by the
operator) before the next stage runs.
- Clear a stage's output/ before re-running it.
- Backtest results are research, not trade plans. Any live trade still goes through
the normal cockpit path.

View File

@@ -0,0 +1,51 @@
# Workspace Context: Strategy Backtest
## Routing
Given a strategy to backtest, the workflow proceeds through stages:
1. **00_clarify_strategy** - Freeze the strategy rules with the operator into a
declarative spec (structure, entries, exits, session, sizing, costs).
2. **01_verify_data** - Enumerate the contract universe, then verify historical
depth (minute-bar range per contract leg, quote-history lookback) with small
test pulls. Produce a data manifest with honest date ranges and gaps.
3. **02_fetch_cache** - Single-pass bulk pull of bars/quotes into the shared
SQLite cache, keyed by (contract, timespan, window). Cache-first: existing
rows are never re-pulled.
4. **03_run_backtest** - Author/run the stdlib-only engine over cached data.
No look-ahead. Costs from risk-params. Output trades.csv + metrics.json.
5. **04_report** - Human-readable report: win rate, expectancy, drawdown,
per-hour breakdowns. Small-sample results flagged as small-sample.
## Shared Resources
- _config/: strategy spec template + risk parameters (stable reference).
- shared/data/: SQLite cache of fetched bars/quotes (built during runs).
- shared/scripts/: engine + helper scripts (stdlib only, authored at stage exec).
## Data rules (verified 2026-08-27)
- Options per-contract OHLC bars: minute→month timespans, history back to
2014-06-02, max 50,000 bars per pull. Daily summaries + previous-day also exist.
- Options quote history + tick trade history exist per contract; quote-history
lookback depth is UNVERIFIED until stage 01 measures it.
- Equity bars: custom OHLC (minute-level), daily summaries, grouped-daily
(all tickers), previous-day.
- No historical greeks/IV series anywhere in the catalog: backtests are
price/levels-driven. IV-rank conditions are not testable.
- Live chain snapshots (greeks/IV) are current-tape only — never a backtest source.
## Engine rules
- Pure Python 3.12 stdlib (json, csv, sqlite3, statistics, math, datetime, urllib).
NO pip installs — the operator manages packaging with uv; pyproject.toml
declares zero dependencies.
- A credit spread is two leg pulls stitched: credit = short-leg price − long-leg
price at entry; P&L tracked per bar on the leg diff.
- Costs: conservative slippage default (mid ± half-spread per fill), from
_config/risk-params.md.
- Exits modeled per house style: stop, T1/T2/T3 partials, hard 15:45 ET time exit,
optional breakeven-after-T1.
- RTH is the default session; extended only if the spec says so explicitly.
## Honesty rules
- Never present a backtest result without N (trade count) and the tested window.
- Low-N results (<30 trades) must be labeled low-N in the report.
- If data has gaps (no bars for an era/strike), say so in the manifest and report —
never silently interpolate through them.

View File

@@ -0,0 +1,40 @@
# Risk Parameters (stable defaults — edit sparingly, diffs matter)
# Stage 03 reads this file; the strategy spec may override individual values.
# Data window (v1)
window:
start: 2026-02-27 # ~6 months back; stage 01 narrows to contract reality
end: 2026-08-27
# Costs
costs:
slippage_model: mid_half_spread # each fill at mid ± half observed spread
min_tick: 0.01 # options tick floor
commissions_per_contract: 0.0 # 0 by default; operator may set
# Sessions
session:
default: RTH # 09:30–16:00 ET
hard_exit_et: "15:45" # house rule: flat before the close
extended_allowed: false # only if strategy spec explicitly enables
# Sizing
sizing:
mode: fixed_contract
contracts: 1
# Exits (defaults; spec overrides)
exits:
breakeven_after_t1: false
# Cache
cache:
path: shared/data/backtest_cache.sqlite3
key: (contract, timespan, window) # cache-first; never re-pull existing rows
max_rows_per_table: 5000000
# Validation
validation:
no_lookahead: true
low_n_threshold: 30 # reports must flag results below this trade count
split: none # optional in-sample/out-of-sample date split

View File

@@ -0,0 +1,49 @@
# Strategy Spec Template
# Fill one copy per strategy: _config/strategy-spec-<name>.md
# The frozen spec (stage 00 output) is the single source of truth for the engine.
# Name
name: <short-slug>
# Underlying & structure
underlying: <SPY>
structure: <single_leg | credit_spread | debit_spread>
# For spreads, list legs explicitly:
legs:
- role: short
type: put
selection: <e.g., ~30-delta proxy: strike nearest 0.5% OTM of spot>
- role: long
type: put
selection: <e.g., $5 wide below short strike>
expiration: <e.g., same-day (0DTE) — nearest daily expiry>
# Session
session: <RTH (default) | extended>
entry_window: <e.g., 09:45–11:00 ET>
hard_exit: <e.g., 15:45 ET>
# Entry rules (plain English, bar-level. NO look-ahead allowed.)
entry:
- <rule 1 — e.g., price pulls back to rising VWAP>
- <rule 2 — optional confirmation>
all_required: true # every rule must hold on the entry bar
# Exit rules
exits:
stop: <underlying level or premium % — define precisely>
targets:
- t1: <profit % of credit or premium>
scale_out: <fraction, e.g., 50%>
- t2: <...>
scale_out: <...>
breakeven_after_t1: false
time_exit: 15:45 ET
# Sizing & costs
sizing:
mode: fixed_contract # v1 default: 1 contract
contracts: 1
costs:
slippage: mid_half_spread # conservative default; see risk-params.md
commissions: 0 # set if the operator wants them modeled

View File

@@ -0,0 +1,9 @@
[project]
name = "backtest-strategy"
version = "0.1.0"
description = "ICM backtest workspace - pure stdlib engine. Operator manages packaging with uv; zero dependencies by design."
requires-python = ">=3.12"
dependencies = []
# NOTE: intentionally dependency-free. If a future need for pandas/numpy arises,
# that is an explicit operator decision - not an agent action.

View File

@@ -0,0 +1,8 @@
# shared/data/
Holds the backtest cache (SQLite) built by stage 02_fetch_cache.
- `backtest_cache.sqlite3` — keyed by (contract, timespan, window). Cache-first:
existing rows are never re-pulled from TTG data servers.
- This folder may grow large. Cap discipline lives in `_config/risk-params.md`.
- Delete the .sqlite3 file to force a full re-pull (stage 02 will rebuild it).

View File

@@ -0,0 +1,28 @@
# Stage 00 Clarify Strategy: Freeze the Rules
Purpose: turn the operator's strategy idea into a single declarative spec that the
engine can execute verbatim. Nothing downstream runs until this file exists and the
operator approves it. This is a review gate.
## Inputs
- Layer 3 (reference): ../../_config/strategy-spec-TEMPLATE.md
- Layer 3 (reference): ../../_config/risk-params.md
- Layer 4 (working): operator's description of the strategy (from conversation)
## Process
1. Read the template and risk params.
2. Interview the operator until every template field is answerable: structure
(single leg / credit spread / debit spread), leg selection, expiration rule,
entry window, entry rules (bar-level, no look-ahead), stop, targets with
scale-out fractions, time exit, sizing.
3. Restate the rules back in plain English and get explicit operator confirmation
before freezing. Ambiguity is resolved by the operator, never guessed.
4. Default candidate if the operator asks for a starting point: SPY 0DTE credit
put spread (short ~0.5% OTM, long $5 wider, 09:45–11:00 ET entries, 15:45 hard
exit, T1/T2 partials). This is a proposal, not a decision.
5. Write the frozen spec as a filled copy of the template.
6. Stop and hand off to the operator for review. Stage 01 does not start until
the operator approves strategy_spec.md.
## Outputs
- strategy_spec.md -> output/

View File

@@ -0,0 +1,34 @@
# Stage 01 Verify Data: Contract Enumeration + Depth Check
Purpose: establish, with small test pulls only, exactly what data exists for the
frozen spec's universe — before any bulk fetching. Honesty about gaps is the whole
point of this stage. Review gate.
## Inputs
- Layer 4 (working): ../00_clarify_strategy/output/strategy_spec.md
- Layer 3 (reference): ../../CONTEXT.md (Data rules section)
- Layer 3 (reference): ../../_config/risk-params.md (window)
## Process
1. Read the strategy spec to determine the universe: underlying, leg selection
rule, expiration cadence, session.
2. Enumerate the contract universe for the window using the options/stocks
reference endpoints (browse with retrieve_all, confirm shapes with params
before any execute — house rule for every new endpoint).
3. For 2–3 sample contracts (one recent, one mid-window, one oldest needed):
- pull a small bar window (e.g., 1 day of minute bars) per leg to confirm
minute-bar availability within the spec's entry window;
- pull a small quote-history window and record the ACTUAL earliest timestamp
returned. Quote-history lookback depth is unverified — this measures it.
4. Record per-contract-leg findings in the manifest: earliest/latest verified bar
dates, quote-history earliest date, gaps, holidays in window, and any contract
the spec's selection rule would pick that has no data.
5. If the spec's window is not fully coverable (e.g., quote history shallower than
the window), state the impact plainly and propose the largest fully-coverable
window. Do not silently shrink the test.
6. No bulk pulls in this stage. Keep total pulls small (roughly a dozen).
7. Stop for operator review of the manifest before stage 02 fetches anything.
## Outputs
- data_manifest.md -> output/ (universe table, verified depth per leg, gaps,
proposed final window)

View File

@@ -0,0 +1,29 @@
# Stage 02 Fetch Cache: Single-Pass Bulk Pull
Purpose: materialize every bar/quote series the manifest calls for into the shared
SQLite cache — once. Re-runs of later stages must never re-pull data. Review gate.
## Inputs
- Layer 4 (working): ../01_verify_data/output/data_manifest.md
- Layer 3 (reference): ../../_config/risk-params.md (window, cache path, caps)
- Layer 3 (reference): ../../CONTEXT.md (Data rules)
## Process
1. Read the manifest's final (operator-approved) universe + window.
2. Author a stdlib-only fetch script into ../../shared/scripts/ (urllib for HTTP,
sqlite3 for the cache, json/csv for any side exports). No third-party packages.
3. Cache-first: for each (contract, timespan, window), check the cache and skip
rows already present. Only missing ranges are fetched.
4. Pull legs bar-by-bar: for each contract, each leg, minute bars for the spec's
session window across the manifest's date list. Respect the 50,000-bar per-pull
cap by splitting multi-month pulls into monthly sub-windows.
5. Optional per spec: pull quote history only for the eras stage 01 verified.
6. Log every pull (contract, timespan, date range, rows returned, gaps found) to
the fetch log. A pull returning zero bars is a logged fact, not an error to hide.
7. Sanity-check the cache: row counts per contract vs expected session days; flag
any contract with <50% expected coverage.
8. Stop for operator review before the engine runs.
## Outputs
- fetch_log.md -> output/ (pull table, coverage stats, anomalies)
- shared/data/backtest_cache.sqlite3 (the cache itself)

View File

@@ -0,0 +1,38 @@
# Stage 03 Run Backtest: Execute the Spec Over the Cache
Purpose: author and run the stdlib-only engine against the cached data, producing
a complete trade list and metrics. The spec is law; the engine never improvises.
Review gate.
## Inputs
- Layer 4 (working): ../00_clarify_strategy/output/strategy_spec.md
- Layer 4 (working): ../02_fetch_cache/output/fetch_log.md (coverage caveats)
- Layer 3 (reference): ../../_config/risk-params.md (costs, exits defaults)
- Layer 4 (working): ../../shared/data/backtest_cache.sqlite3
## Process
1. Author the engine script into ../../shared/scripts/ (pure Python 3.12 stdlib:
sqlite3, csv, json, statistics, math, datetime). No third-party packages —
the operator manages packaging with uv; pyproject.toml has zero dependencies.
2. Engine mechanics:
- Iterate bars chronologically per session; signals on bar N may only use
bars <= N (no look-ahead, hard rule).
- Entries only inside the spec's entry window and session (RTH default).
- Credit spread = short leg + long leg stitched; credit at entry = short
price − long price; per-bar P&L tracked on the leg diff.
- Exits in precedence order: hard time exit (15:45 ET) > stop > targets
(T1/T2/T3 scale-outs per spec) > optional breakeven-after-T1.
- Fills at bar close ± slippage from risk-params (mid ± half-spread default).
3. Emit one row per simulated trade: dates, entry/exit, leg prices, credit,
P&L, MAE/MFE, exit reason, session day.
4. Compute metrics: N, win rate, expectancy, avg win/loss, profit factor, max
drawdown, per-hour-of-day and per-weekday breakdowns, equity curve series.
5. Also run a costs-off variant (same trades, zero slippage) and store both —
the gap between them is the cost drag, and it must be visible.
6. If cached coverage has gaps inside the window, exclude affected days from
stats and list them — never interpolate through a gap.
7. Stop for operator review before the report stage.
## Outputs
- trades.csv -> output/
- metrics.json -> output/ (with-costs and costs-off variants)

View File

@@ -0,0 +1,29 @@
# Stage 04 Report: Honest Findings, Operator-Readable
Purpose: turn trades.csv + metrics.json into a report a trader can act on — or
consciously discard. No cherry-picking, no curve-fit praise. Final review gate.
## Inputs
- Layer 4 (working): ../03_run_backtest/output/trades.csv
- Layer 4 (working): ../03_run_backtest/output/metrics.json
- Layer 4 (working): ../01_verify_data/output/data_manifest.md (coverage caveats)
## Process
1. Lead with the headline numbers: N trades, win rate, expectancy per trade,
profit factor, max drawdown — for the with-costs run. Costs-off appears only
as the visible cost-drag comparison, never as the headline.
2. State the tested window, contract universe size, and data coverage honestly,
including any excluded gap days.
3. Flag low-N explicitly: below 30 trades the report must say the sample is too
small to trust and must not recommend going live on it.
4. Breakdowns: per-hour-of-day, per-weekday, exit-reason mix (stop vs target vs
time exit), and the stop-vs-target balance. These drive spec tuning.
5. MAE/MFE distribution: how deep winners usually dip before working — this is
what sets realistic stop placement and profit-target spacing.
6. End with a plain recommendation set: keep / tune (which parameter, which
direction) / discard. A losing strategy is a valid, useful result — say so.
7. Remind the reader: backtest output is research. Live execution still goes
through the normal cockpit path with its own preflight.
## Outputs
- report.md -> output/