Files
ite-workflows/backtest-workflow-PLAN.md

108 lines
6.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Backtest Workflow — DRAFT PLAN (for review — nothing built yet)
**Status:** APPROVED (2026-08-27 ~11 PM ET) and BUILT — workspace live at
`workflows/backtest-strategy/`. No packages installed (zero-dep pyproject.toml
included; operator manages packaging with uv). No data pulled yet. Engine is
authored at stage-03 exec time, not pre-built.
**Date drafted:** Thu 2026-08-27, ~11 PM ET
**Proposed workspace:** `workflows/backtest-strategy/` (follows ICM.md — folder structure as orchestrator)
---
## Goal
Rules-based backtesting for Gump's core styles, using TTG data servers as the only market-data source and a local pure-Python engine:
1. **SPY 0DTE credit put spreads** (bread and butter — premium-selling, % profit targets)
2. **Single-leg equity/options day trades** (VWAP-pullback longs, 50MA-fade shorts)
3. Multi-target management exactly as traded live: T1/T2/T3 partials, stop tightening, hard time exit (15:45 ET), optional breakeven-after-T1
## Architecture — three layers
### 1. Data layer (TTG servers only — never scraped)
| Need | Source (verified 08/27) | Notes |
|---|---|---|
| Equity bars | custom OHLC bars, daily summaries, grouped-daily (all tickers) | minute-level intraday |
| Options bars | per-contract OHLC bars | minute→month timespans, **back to 2014-06-02**, ≤50,000 bars/pull |
| Contract enumeration | contract list + specs | which strikes/expiries existed on date X |
| Historical quotes | per-contract quote history | **depth unverified — stage 01 verifies before any promise** |
| Historical ticks | per-contract trade history | fallback/verification for quote data |
| Snapshots (live) | chain w/ greeks+IV, per-contract, unified | current-tape only — not for backtests |
**Cache design:** all pulls land in `shared/data/` as SQLite (stdlib `sqlite3`), keyed by (contract, timespan, window). Re-runs read cache first — pulls happen once per window ever. Cache caps enforced so the folder doesn't balloon.
### 2. Engine layer — pure Python 3.12 stdlib (v1: zero installs)
- Bar-by-bar event loop over cached data (`json`, `csv`, `sqlite3`, `statistics`, `math`, `datetime` only)
- Strategies are **declarative spec files** (`_config/strategy-*.md` + params block) — human-readable, reviewable, diffable
- Position model:
- Single leg (long/short equity or option)
- **Two-leg credit spread** = short leg + long leg bars stitched; credit = leg diff at entry; P&L tracked per bar
- Exit model: stop, T1/T2/T3 partial scale-outs, hard time exit, optional breve-after-T1, EOD flat (15:45 ET per house rule)
- Costs: configurable slippage (default **conservative**: mid ± half-spread per fill), optional commissions; RTH default / extended optional flag per session-language rule
- Validation: in-sample/out-of-sample split by date; no look-ahead (signals computed on bars ≤ current bar only)
### 3. Reporting layer
- `trades.csv` — every simulated trade: entry/exit, legs, MAE/MFE, time-in-trade, exit reason
- `metrics.json` — win rate, expectancy, avg win/loss, profit factor, max drawdown, per-hour-of-day and per-weekday breakdowns
- Final stage renders a readable report + summary card
## Proposed ICM workspace layout
```
workflows/backtest-strategy/
├── CLAUDE.md # entry point + exec protocol (read ICM.md fresh, clear output/ on re-run)
├── CONTEXT.md # workspace context: routing + data rules (TTG-only, no fabrication)
├── _config/
│ ├── strategy-spec-TEMPLATE.md
│ └── risk-params.md # default slippage, session, sizing
├── shared/
│ └── data/ # SQLite cache (built during runs)
└── stages/
├── 00_clarify_strategy/ # capture rules in plain English → strategy_spec.md
├── 01_verify_data/ # enumerate contracts, verify quote-history depth → data_manifest.md
├── 02_fetch_cache/ # pull bars/quotes → shared/data/ → fetch_log.md
├── 03_run_backtest/ # engine executes spec → trades.csv + metrics.json
└── 04_report/ # human-readable report + summary card → report.md
```
Every stage ends at a **review gate** — output/ is read and (if needed) edited before the next stage runs, per ICM.
## Stage contracts (summary)
| Stage | Reads | Does | Writes |
|---|---|---|---|
| 00_clarify | _config templates, member Q&A | freeze strategy rules + params | strategy_spec.md |
| 01_verify_data | strategy_spec.md | enumerate contracts; **test-pull quote history depth**; flag gaps | data_manifest.md |
| 02_fetch_cache | data_manifest.md | pull bars/quotes → cache (single pass, no repeats) | fetch_log.md |
| 03_run_backtest | cache + spec | run engine, no look-ahead, costs applied | trades.csv, metrics.json |
| 04_report | trades + metrics | render findings honestly (incl. small-sample warnings) | report.md |
## Decision points — need Gump's call
1. **v1 compute:** pure stdlib, zero installs (recommended) — or approve `pip install pandas numpy` now?
2. **First strategy to backtest:** (a) SPY 0DTE credit put spread *(recommended — core style)*, (b) VWAP-pullback equity long, (c) 50MA fade short?
3. **Data window:** propose **last 6 months** of SPY 0DTE contracts for v1 (cache stays lean; extendable later)?
4. **Slippage default:** conservative mid ± half-spread, tunable in `_config/risk-params.md` — OK?
5. **Sizing model:** fixed 1-contract (clean signal measurement) vs fixed-dollar risk? Propose fixed 1-contract for v1.
## Known limitations (stated up front)
- **No historical greeks/IV series** — backtests are price/levels-driven; IV-rank conditions are not testable
- **No order-book replay** — fills modeled at bar close with slippage knob; inherently approximate, slightly optimistic
- **Quote-history depth TBD** — stage 01 measures it before we rely on it
- **0DTE dailies** only exist for the era they traded; earlier "0DTE" = nearest weekly
- Small-sample honesty: 6 months of 0DTE ≈ ~125 trading days — stage 04 will flag low-N results as such
## Credit/cost profile
- All pulls are read-only TTG market-data calls (conservation mode is OFF per member).
- Heavy stage is 02; the cache means any re-run/re-parameterization costs **zero** additional pulls.
- Rough v1 pull count: ~125 contracts × 2 legs × 1 window, each well under the 50k-bar cap.
## Explicitly out of scope for v1
- Portfolio-level / multi-strategy simulation, options greeks modeling, intrabar stop sequencing (stops checked bar-by-bar, close-based), live-paper forwarding (backtest ≠ trade plan — any live trade still goes through the normal cockpit path)