Files
ite-workflows/backtest-workflow-PLAN.md

6.6 KiB
Raw Blame History

Backtest Workflow — DRAFT PLAN (for review — nothing built yet)

Status: APPROVED (2026-08-27 ~11 PM ET) and BUILT — workspace live at workflows/backtest-strategy/. No packages installed (zero-dep pyproject.toml included; operator manages packaging with uv). No data pulled yet. Engine is authored at stage-03 exec time, not pre-built. Date drafted: Thu 2026-08-27, ~11 PM ET Proposed workspace: workflows/backtest-strategy/ (follows ICM.md — folder structure as orchestrator)


Goal

Rules-based backtesting for Gump's core styles, using TTG data servers as the only market-data source and a local pure-Python engine:

  1. SPY 0DTE credit put spreads (bread and butter — premium-selling, % profit targets)
  2. Single-leg equity/options day trades (VWAP-pullback longs, 50MA-fade shorts)
  3. Multi-target management exactly as traded live: T1/T2/T3 partials, stop tightening, hard time exit (15:45 ET), optional breakeven-after-T1

Architecture — three layers

1. Data layer (TTG servers only — never scraped)

Need Source (verified 08/27) Notes
Equity bars custom OHLC bars, daily summaries, grouped-daily (all tickers) minute-level intraday
Options bars per-contract OHLC bars minute→month timespans, back to 2014-06-02, ≤50,000 bars/pull
Contract enumeration contract list + specs which strikes/expiries existed on date X
Historical quotes per-contract quote history depth unverified — stage 01 verifies before any promise
Historical ticks per-contract trade history fallback/verification for quote data
Snapshots (live) chain w/ greeks+IV, per-contract, unified current-tape only — not for backtests

Cache design: all pulls land in shared/data/ as SQLite (stdlib sqlite3), keyed by (contract, timespan, window). Re-runs read cache first — pulls happen once per window ever. Cache caps enforced so the folder doesn't balloon.

2. Engine layer — pure Python 3.12 stdlib (v1: zero installs)

  • Bar-by-bar event loop over cached data (json, csv, sqlite3, statistics, math, datetime only)
  • Strategies are declarative spec files (_config/strategy-*.md + params block) — human-readable, reviewable, diffable
  • Position model:
    • Single leg (long/short equity or option)
    • Two-leg credit spread = short leg + long leg bars stitched; credit = leg diff at entry; P&L tracked per bar
  • Exit model: stop, T1/T2/T3 partial scale-outs, hard time exit, optional breve-after-T1, EOD flat (15:45 ET per house rule)
  • Costs: configurable slippage (default conservative: mid ± half-spread per fill), optional commissions; RTH default / extended optional flag per session-language rule
  • Validation: in-sample/out-of-sample split by date; no look-ahead (signals computed on bars ≤ current bar only)

3. Reporting layer

  • trades.csv — every simulated trade: entry/exit, legs, MAE/MFE, time-in-trade, exit reason
  • metrics.json — win rate, expectancy, avg win/loss, profit factor, max drawdown, per-hour-of-day and per-weekday breakdowns
  • Final stage renders a readable report + summary card

Proposed ICM workspace layout

workflows/backtest-strategy/
├── CLAUDE.md                 # entry point + exec protocol (read ICM.md fresh, clear output/ on re-run)
├── CONTEXT.md                # workspace context: routing + data rules (TTG-only, no fabrication)
├── _config/
│   ├── strategy-spec-TEMPLATE.md
│   └── risk-params.md        # default slippage, session, sizing
├── shared/
│   └── data/                 # SQLite cache (built during runs)
└── stages/
    ├── 00_clarify_strategy/  # capture rules in plain English → strategy_spec.md
    ├── 01_verify_data/       # enumerate contracts, verify quote-history depth → data_manifest.md
    ├── 02_fetch_cache/       # pull bars/quotes → shared/data/ → fetch_log.md
    ├── 03_run_backtest/      # engine executes spec → trades.csv + metrics.json
    └── 04_report/            # human-readable report + summary card → report.md

Every stage ends at a review gate — output/ is read and (if needed) edited before the next stage runs, per ICM.

Stage contracts (summary)

Stage Reads Does Writes
00_clarify _config templates, member Q&A freeze strategy rules + params strategy_spec.md
01_verify_data strategy_spec.md enumerate contracts; test-pull quote history depth; flag gaps data_manifest.md
02_fetch_cache data_manifest.md pull bars/quotes → cache (single pass, no repeats) fetch_log.md
03_run_backtest cache + spec run engine, no look-ahead, costs applied trades.csv, metrics.json
04_report trades + metrics render findings honestly (incl. small-sample warnings) report.md

Decision points — need Gump's call

  1. v1 compute: pure stdlib, zero installs (recommended) — or approve pip install pandas numpy now?
  2. First strategy to backtest: (a) SPY 0DTE credit put spread (recommended — core style), (b) VWAP-pullback equity long, (c) 50MA fade short?
  3. Data window: propose last 6 months of SPY 0DTE contracts for v1 (cache stays lean; extendable later)?
  4. Slippage default: conservative mid ± half-spread, tunable in _config/risk-params.md — OK?
  5. Sizing model: fixed 1-contract (clean signal measurement) vs fixed-dollar risk? Propose fixed 1-contract for v1.

Known limitations (stated up front)

  • No historical greeks/IV series — backtests are price/levels-driven; IV-rank conditions are not testable
  • No order-book replay — fills modeled at bar close with slippage knob; inherently approximate, slightly optimistic
  • Quote-history depth TBD — stage 01 measures it before we rely on it
  • 0DTE dailies only exist for the era they traded; earlier "0DTE" = nearest weekly
  • Small-sample honesty: 6 months of 0DTE ≈ ~125 trading days — stage 04 will flag low-N results as such

Credit/cost profile

  • All pulls are read-only TTG market-data calls (conservation mode is OFF per member).
  • Heavy stage is 02; the cache means any re-run/re-parameterization costs zero additional pulls.
  • Rough v1 pull count: ~125 contracts × 2 legs × 1 window, each well under the 50k-bar cap.

Explicitly out of scope for v1

  • Portfolio-level / multi-strategy simulation, options greeks modeling, intrabar stop sequencing (stops checked bar-by-bar, close-based), live-paper forwarding (backtest ≠ trade plan — any live trade still goes through the normal cockpit path)