var - backtest related and spy runbook

This commit is contained in:
2026-08-28 15:10:09 -04:00
parent dd8e4cca0c
commit 81ae138a00
16 changed files with 948 additions and 94 deletions

107
backtest-workflow-PLAN.md Normal file
View File

@@ -0,0 +1,107 @@
# Backtest Workflow — DRAFT PLAN (for review — nothing built yet)
**Status:** APPROVED (2026-08-27 ~11 PM ET) and BUILT — workspace live at
`workflows/backtest-strategy/`. No packages installed (zero-dep pyproject.toml
included; operator manages packaging with uv). No data pulled yet. Engine is
authored at stage-03 exec time, not pre-built.
**Date drafted:** Thu 2026-08-27, ~11 PM ET
**Proposed workspace:** `workflows/backtest-strategy/` (follows ICM.md — folder structure as orchestrator)
---
## Goal
Rules-based backtesting for Gump's core styles, using TTG data servers as the only market-data source and a local pure-Python engine:
1. **SPY 0DTE credit put spreads** (bread and butter — premium-selling, % profit targets)
2. **Single-leg equity/options day trades** (VWAP-pullback longs, 50MA-fade shorts)
3. Multi-target management exactly as traded live: T1/T2/T3 partials, stop tightening, hard time exit (15:45 ET), optional breakeven-after-T1
## Architecture — three layers
### 1. Data layer (TTG servers only — never scraped)
| Need | Source (verified 08/27) | Notes |
|---|---|---|
| Equity bars | custom OHLC bars, daily summaries, grouped-daily (all tickers) | minute-level intraday |
| Options bars | per-contract OHLC bars | minute→month timespans, **back to 2014-06-02**, ≤50,000 bars/pull |
| Contract enumeration | contract list + specs | which strikes/expiries existed on date X |
| Historical quotes | per-contract quote history | **depth unverified — stage 01 verifies before any promise** |
| Historical ticks | per-contract trade history | fallback/verification for quote data |
| Snapshots (live) | chain w/ greeks+IV, per-contract, unified | current-tape only — not for backtests |
**Cache design:** all pulls land in `shared/data/` as SQLite (stdlib `sqlite3`), keyed by (contract, timespan, window). Re-runs read cache first — pulls happen once per window ever. Cache caps enforced so the folder doesn't balloon.
### 2. Engine layer — pure Python 3.12 stdlib (v1: zero installs)
- Bar-by-bar event loop over cached data (`json`, `csv`, `sqlite3`, `statistics`, `math`, `datetime` only)
- Strategies are **declarative spec files** (`_config/strategy-*.md` + params block) — human-readable, reviewable, diffable
- Position model:
- Single leg (long/short equity or option)
- **Two-leg credit spread** = short leg + long leg bars stitched; credit = leg diff at entry; P&L tracked per bar
- Exit model: stop, T1/T2/T3 partial scale-outs, hard time exit, optional breve-after-T1, EOD flat (15:45 ET per house rule)
- Costs: configurable slippage (default **conservative**: mid ± half-spread per fill), optional commissions; RTH default / extended optional flag per session-language rule
- Validation: in-sample/out-of-sample split by date; no look-ahead (signals computed on bars ≤ current bar only)
### 3. Reporting layer
- `trades.csv` — every simulated trade: entry/exit, legs, MAE/MFE, time-in-trade, exit reason
- `metrics.json` — win rate, expectancy, avg win/loss, profit factor, max drawdown, per-hour-of-day and per-weekday breakdowns
- Final stage renders a readable report + summary card
## Proposed ICM workspace layout
```
workflows/backtest-strategy/
├── CLAUDE.md # entry point + exec protocol (read ICM.md fresh, clear output/ on re-run)
├── CONTEXT.md # workspace context: routing + data rules (TTG-only, no fabrication)
├── _config/
│ ├── strategy-spec-TEMPLATE.md
│ └── risk-params.md # default slippage, session, sizing
├── shared/
│ └── data/ # SQLite cache (built during runs)
└── stages/
├── 00_clarify_strategy/ # capture rules in plain English → strategy_spec.md
├── 01_verify_data/ # enumerate contracts, verify quote-history depth → data_manifest.md
├── 02_fetch_cache/ # pull bars/quotes → shared/data/ → fetch_log.md
├── 03_run_backtest/ # engine executes spec → trades.csv + metrics.json
└── 04_report/ # human-readable report + summary card → report.md
```
Every stage ends at a **review gate** — output/ is read and (if needed) edited before the next stage runs, per ICM.
## Stage contracts (summary)
| Stage | Reads | Does | Writes |
|---|---|---|---|
| 00_clarify | _config templates, member Q&A | freeze strategy rules + params | strategy_spec.md |
| 01_verify_data | strategy_spec.md | enumerate contracts; **test-pull quote history depth**; flag gaps | data_manifest.md |
| 02_fetch_cache | data_manifest.md | pull bars/quotes → cache (single pass, no repeats) | fetch_log.md |
| 03_run_backtest | cache + spec | run engine, no look-ahead, costs applied | trades.csv, metrics.json |
| 04_report | trades + metrics | render findings honestly (incl. small-sample warnings) | report.md |
## Decision points — need Gump's call
1. **v1 compute:** pure stdlib, zero installs (recommended) — or approve `pip install pandas numpy` now?
2. **First strategy to backtest:** (a) SPY 0DTE credit put spread *(recommended — core style)*, (b) VWAP-pullback equity long, (c) 50MA fade short?
3. **Data window:** propose **last 6 months** of SPY 0DTE contracts for v1 (cache stays lean; extendable later)?
4. **Slippage default:** conservative mid ± half-spread, tunable in `_config/risk-params.md` — OK?
5. **Sizing model:** fixed 1-contract (clean signal measurement) vs fixed-dollar risk? Propose fixed 1-contract for v1.
## Known limitations (stated up front)
- **No historical greeks/IV series** — backtests are price/levels-driven; IV-rank conditions are not testable
- **No order-book replay** — fills modeled at bar close with slippage knob; inherently approximate, slightly optimistic
- **Quote-history depth TBD** — stage 01 measures it before we rely on it
- **0DTE dailies** only exist for the era they traded; earlier "0DTE" = nearest weekly
- Small-sample honesty: 6 months of 0DTE ≈ ~125 trading days — stage 04 will flag low-N results as such
## Credit/cost profile
- All pulls are read-only TTG market-data calls (conservation mode is OFF per member).
- Heavy stage is 02; the cache means any re-run/re-parameterization costs **zero** additional pulls.
- Rough v1 pull count: ~125 contracts × 2 legs × 1 window, each well under the 50k-bar cap.
## Explicitly out of scope for v1
- Portfolio-level / multi-strategy simulation, options greeks modeling, intrabar stop sequencing (stops checked bar-by-bar, close-based), live-paper forwarding (backtest ≠ trade plan — any live trade still goes through the normal cockpit path)