# Backtest Workflow — DRAFT PLAN (for review — nothing built yet) **Status:** APPROVED (2026-08-27 ~11 PM ET) and BUILT — workspace live at `workflows/backtest-strategy/`. No packages installed (zero-dep pyproject.toml included; operator manages packaging with uv). No data pulled yet. Engine is authored at stage-03 exec time, not pre-built. **Date drafted:** Thu 2026-08-27, ~11 PM ET **Proposed workspace:** `workflows/backtest-strategy/` (follows ICM.md — folder structure as orchestrator) --- ## Goal Rules-based backtesting for Gump's core styles, using TTG data servers as the only market-data source and a local pure-Python engine: 1. **SPY 0DTE credit put spreads** (bread and butter — premium-selling, % profit targets) 2. **Single-leg equity/options day trades** (VWAP-pullback longs, 50MA-fade shorts) 3. Multi-target management exactly as traded live: T1/T2/T3 partials, stop tightening, hard time exit (15:45 ET), optional breakeven-after-T1 ## Architecture — three layers ### 1. Data layer (TTG servers only — never scraped) | Need | Source (verified 08/27) | Notes | |---|---|---| | Equity bars | custom OHLC bars, daily summaries, grouped-daily (all tickers) | minute-level intraday | | Options bars | per-contract OHLC bars | minute→month timespans, **back to 2014-06-02**, ≤50,000 bars/pull | | Contract enumeration | contract list + specs | which strikes/expiries existed on date X | | Historical quotes | per-contract quote history | **depth unverified — stage 01 verifies before any promise** | | Historical ticks | per-contract trade history | fallback/verification for quote data | | Snapshots (live) | chain w/ greeks+IV, per-contract, unified | current-tape only — not for backtests | **Cache design:** all pulls land in `shared/data/` as SQLite (stdlib `sqlite3`), keyed by (contract, timespan, window). Re-runs read cache first — pulls happen once per window ever. Cache caps enforced so the folder doesn't balloon. ### 2. Engine layer — pure Python 3.12 stdlib (v1: zero installs) - Bar-by-bar event loop over cached data (`json`, `csv`, `sqlite3`, `statistics`, `math`, `datetime` only) - Strategies are **declarative spec files** (`_config/strategy-*.md` + params block) — human-readable, reviewable, diffable - Position model: - Single leg (long/short equity or option) - **Two-leg credit spread** = short leg + long leg bars stitched; credit = leg diff at entry; P&L tracked per bar - Exit model: stop, T1/T2/T3 partial scale-outs, hard time exit, optional breve-after-T1, EOD flat (15:45 ET per house rule) - Costs: configurable slippage (default **conservative**: mid ± half-spread per fill), optional commissions; RTH default / extended optional flag per session-language rule - Validation: in-sample/out-of-sample split by date; no look-ahead (signals computed on bars ≤ current bar only) ### 3. Reporting layer - `trades.csv` — every simulated trade: entry/exit, legs, MAE/MFE, time-in-trade, exit reason - `metrics.json` — win rate, expectancy, avg win/loss, profit factor, max drawdown, per-hour-of-day and per-weekday breakdowns - Final stage renders a readable report + summary card ## Proposed ICM workspace layout ``` workflows/backtest-strategy/ ├── CLAUDE.md # entry point + exec protocol (read ICM.md fresh, clear output/ on re-run) ├── CONTEXT.md # workspace context: routing + data rules (TTG-only, no fabrication) ├── _config/ │ ├── strategy-spec-TEMPLATE.md │ └── risk-params.md # default slippage, session, sizing ├── shared/ │ └── data/ # SQLite cache (built during runs) └── stages/ ├── 00_clarify_strategy/ # capture rules in plain English → strategy_spec.md ├── 01_verify_data/ # enumerate contracts, verify quote-history depth → data_manifest.md ├── 02_fetch_cache/ # pull bars/quotes → shared/data/ → fetch_log.md ├── 03_run_backtest/ # engine executes spec → trades.csv + metrics.json └── 04_report/ # human-readable report + summary card → report.md ``` Every stage ends at a **review gate** — output/ is read and (if needed) edited before the next stage runs, per ICM. ## Stage contracts (summary) | Stage | Reads | Does | Writes | |---|---|---|---| | 00_clarify | _config templates, member Q&A | freeze strategy rules + params | strategy_spec.md | | 01_verify_data | strategy_spec.md | enumerate contracts; **test-pull quote history depth**; flag gaps | data_manifest.md | | 02_fetch_cache | data_manifest.md | pull bars/quotes → cache (single pass, no repeats) | fetch_log.md | | 03_run_backtest | cache + spec | run engine, no look-ahead, costs applied | trades.csv, metrics.json | | 04_report | trades + metrics | render findings honestly (incl. small-sample warnings) | report.md | ## Decision points — need Gump's call 1. **v1 compute:** pure stdlib, zero installs (recommended) — or approve `pip install pandas numpy` now? 2. **First strategy to backtest:** (a) SPY 0DTE credit put spread *(recommended — core style)*, (b) VWAP-pullback equity long, (c) 50MA fade short? 3. **Data window:** propose **last 6 months** of SPY 0DTE contracts for v1 (cache stays lean; extendable later)? 4. **Slippage default:** conservative mid ± half-spread, tunable in `_config/risk-params.md` — OK? 5. **Sizing model:** fixed 1-contract (clean signal measurement) vs fixed-dollar risk? Propose fixed 1-contract for v1. ## Known limitations (stated up front) - **No historical greeks/IV series** — backtests are price/levels-driven; IV-rank conditions are not testable - **No order-book replay** — fills modeled at bar close with slippage knob; inherently approximate, slightly optimistic - **Quote-history depth TBD** — stage 01 measures it before we rely on it - **0DTE dailies** only exist for the era they traded; earlier "0DTE" = nearest weekly - Small-sample honesty: 6 months of 0DTE ≈ ~125 trading days — stage 04 will flag low-N results as such ## Credit/cost profile - All pulls are read-only TTG market-data calls (conservation mode is OFF per member). - Heavy stage is 02; the cache means any re-run/re-parameterization costs **zero** additional pulls. - Rough v1 pull count: ~125 contracts × 2 legs × 1 window, each well under the 50k-bar cap. ## Explicitly out of scope for v1 - Portfolio-level / multi-strategy simulation, options greeks modeling, intrabar stop sequencing (stops checked bar-by-bar, close-based), live-paper forwarding (backtest ≠ trade plan — any live trade still goes through the normal cockpit path)