var - backtest related and spy runbook

This commit is contained in:
2026-08-28 15:10:09 -04:00
parent dd8e4cca0c
commit 81ae138a00
16 changed files with 948 additions and 94 deletions

View File

@@ -0,0 +1,32 @@
# Backtest Strategy ICM Workspace
This workspace orchestrates rules-based strategy backtesting for the operator's core
styles (SPY 0DTE credit spreads, single-leg equity/options day trades) using TTG data
servers as the only market-data source and a local pure-Python engine.
The agent follows the numbered stages to freeze strategy rules, verify data depth,
build the data cache, run the backtest, and produce an honest report.
Folder structure:
- CLAUDE.md (Layer 0): workspace identity
- CONTEXT.md (Layer 1): workspace-level routing
- stages/: numbered stage folders
- 00_clarify_strategy/: freeze rules into a spec (with operator)
- 01_verify_data/: enumerate contracts + verify historical depth
- 02_fetch_cache/: pull bars/quotes into the shared cache
- 03_run_backtest/: run the engine over cached data
- 04_report/: render findings with small-sample honesty
- _config/: Layer 3 reference material (stable across runs)
- shared/: data cache (SQLite) + engine scripts
- Each stage's output/ holds Layer 4 working artifacts for handoff to next stage.
## Hard rules (apply to every stage)
- Market data comes ONLY from TTG data servers. Never scrape or substitute public sites.
- Python is stdlib-only. NO package installs. The operator manages packaging with uv
(pyproject.toml in the workspace root has zero dependencies by design).
- No look-ahead: a signal on bar N may only use bars <= N.
- Every stage ends at a review gate. output/ is read (and optionally edited by the
operator) before the next stage runs.
- Clear a stage's output/ before re-running it.
- Backtest results are research, not trade plans. Any live trade still goes through
the normal cockpit path.

View File

@@ -0,0 +1,51 @@
# Workspace Context: Strategy Backtest
## Routing
Given a strategy to backtest, the workflow proceeds through stages:
1. **00_clarify_strategy** - Freeze the strategy rules with the operator into a
declarative spec (structure, entries, exits, session, sizing, costs).
2. **01_verify_data** - Enumerate the contract universe, then verify historical
depth (minute-bar range per contract leg, quote-history lookback) with small
test pulls. Produce a data manifest with honest date ranges and gaps.
3. **02_fetch_cache** - Single-pass bulk pull of bars/quotes into the shared
SQLite cache, keyed by (contract, timespan, window). Cache-first: existing
rows are never re-pulled.
4. **03_run_backtest** - Author/run the stdlib-only engine over cached data.
No look-ahead. Costs from risk-params. Output trades.csv + metrics.json.
5. **04_report** - Human-readable report: win rate, expectancy, drawdown,
per-hour breakdowns. Small-sample results flagged as small-sample.
## Shared Resources
- _config/: strategy spec template + risk parameters (stable reference).
- shared/data/: SQLite cache of fetched bars/quotes (built during runs).
- shared/scripts/: engine + helper scripts (stdlib only, authored at stage exec).
## Data rules (verified 2026-08-27)
- Options per-contract OHLC bars: minute→month timespans, history back to
2014-06-02, max 50,000 bars per pull. Daily summaries + previous-day also exist.
- Options quote history + tick trade history exist per contract; quote-history
lookback depth is UNVERIFIED until stage 01 measures it.
- Equity bars: custom OHLC (minute-level), daily summaries, grouped-daily
(all tickers), previous-day.
- No historical greeks/IV series anywhere in the catalog: backtests are
price/levels-driven. IV-rank conditions are not testable.
- Live chain snapshots (greeks/IV) are current-tape only — never a backtest source.
## Engine rules
- Pure Python 3.12 stdlib (json, csv, sqlite3, statistics, math, datetime, urllib).
NO pip installs — the operator manages packaging with uv; pyproject.toml
declares zero dependencies.
- A credit spread is two leg pulls stitched: credit = short-leg price − long-leg
price at entry; P&L tracked per bar on the leg diff.
- Costs: conservative slippage default (mid ± half-spread per fill), from
_config/risk-params.md.
- Exits modeled per house style: stop, T1/T2/T3 partials, hard 15:45 ET time exit,
optional breakeven-after-T1.
- RTH is the default session; extended only if the spec says so explicitly.
## Honesty rules
- Never present a backtest result without N (trade count) and the tested window.
- Low-N results (<30 trades) must be labeled low-N in the report.
- If data has gaps (no bars for an era/strike), say so in the manifest and report —
never silently interpolate through them.

View File

@@ -0,0 +1,40 @@
# Risk Parameters (stable defaults — edit sparingly, diffs matter)
# Stage 03 reads this file; the strategy spec may override individual values.
# Data window (v1)
window:
start: 2026-02-27 # ~6 months back; stage 01 narrows to contract reality
end: 2026-08-27
# Costs
costs:
slippage_model: mid_half_spread # each fill at mid ± half observed spread
min_tick: 0.01 # options tick floor
commissions_per_contract: 0.0 # 0 by default; operator may set
# Sessions
session:
default: RTH # 09:30–16:00 ET
hard_exit_et: "15:45" # house rule: flat before the close
extended_allowed: false # only if strategy spec explicitly enables
# Sizing
sizing:
mode: fixed_contract
contracts: 1
# Exits (defaults; spec overrides)
exits:
breakeven_after_t1: false
# Cache
cache:
path: shared/data/backtest_cache.sqlite3
key: (contract, timespan, window) # cache-first; never re-pull existing rows
max_rows_per_table: 5000000
# Validation
validation:
no_lookahead: true
low_n_threshold: 30 # reports must flag results below this trade count
split: none # optional in-sample/out-of-sample date split

View File

@@ -0,0 +1,49 @@
# Strategy Spec Template
# Fill one copy per strategy: _config/strategy-spec-<name>.md
# The frozen spec (stage 00 output) is the single source of truth for the engine.
# Name
name: <short-slug>
# Underlying & structure
underlying: <SPY>
structure: <single_leg | credit_spread | debit_spread>
# For spreads, list legs explicitly:
legs:
- role: short
type: put
selection: <e.g., ~30-delta proxy: strike nearest 0.5% OTM of spot>
- role: long
type: put
selection: <e.g., $5 wide below short strike>
expiration: <e.g., same-day (0DTE) — nearest daily expiry>
# Session
session: <RTH (default) | extended>
entry_window: <e.g., 09:45–11:00 ET>
hard_exit: <e.g., 15:45 ET>
# Entry rules (plain English, bar-level. NO look-ahead allowed.)
entry:
- <rule 1 — e.g., price pulls back to rising VWAP>
- <rule 2 — optional confirmation>
all_required: true # every rule must hold on the entry bar
# Exit rules
exits:
stop: <underlying level or premium % — define precisely>
targets:
- t1: <profit % of credit or premium>
scale_out: <fraction, e.g., 50%>
- t2: <...>
scale_out: <...>
breakeven_after_t1: false
time_exit: 15:45 ET
# Sizing & costs
sizing:
mode: fixed_contract # v1 default: 1 contract
contracts: 1
costs:
slippage: mid_half_spread # conservative default; see risk-params.md
commissions: 0 # set if the operator wants them modeled

View File

@@ -0,0 +1,9 @@
[project]
name = "backtest-strategy"
version = "0.1.0"
description = "ICM backtest workspace - pure stdlib engine. Operator manages packaging with uv; zero dependencies by design."
requires-python = ">=3.12"
dependencies = []
# NOTE: intentionally dependency-free. If a future need for pandas/numpy arises,
# that is an explicit operator decision - not an agent action.

View File

@@ -0,0 +1,8 @@
# shared/data/
Holds the backtest cache (SQLite) built by stage 02_fetch_cache.
- `backtest_cache.sqlite3` — keyed by (contract, timespan, window). Cache-first:
existing rows are never re-pulled from TTG data servers.
- This folder may grow large. Cap discipline lives in `_config/risk-params.md`.
- Delete the .sqlite3 file to force a full re-pull (stage 02 will rebuild it).

View File

@@ -0,0 +1,28 @@
# Stage 00 Clarify Strategy: Freeze the Rules
Purpose: turn the operator's strategy idea into a single declarative spec that the
engine can execute verbatim. Nothing downstream runs until this file exists and the
operator approves it. This is a review gate.
## Inputs
- Layer 3 (reference): ../../_config/strategy-spec-TEMPLATE.md
- Layer 3 (reference): ../../_config/risk-params.md
- Layer 4 (working): operator's description of the strategy (from conversation)
## Process
1. Read the template and risk params.
2. Interview the operator until every template field is answerable: structure
(single leg / credit spread / debit spread), leg selection, expiration rule,
entry window, entry rules (bar-level, no look-ahead), stop, targets with
scale-out fractions, time exit, sizing.
3. Restate the rules back in plain English and get explicit operator confirmation
before freezing. Ambiguity is resolved by the operator, never guessed.
4. Default candidate if the operator asks for a starting point: SPY 0DTE credit
put spread (short ~0.5% OTM, long $5 wider, 09:45–11:00 ET entries, 15:45 hard
exit, T1/T2 partials). This is a proposal, not a decision.
5. Write the frozen spec as a filled copy of the template.
6. Stop and hand off to the operator for review. Stage 01 does not start until
the operator approves strategy_spec.md.
## Outputs
- strategy_spec.md -> output/

View File

@@ -0,0 +1,34 @@
# Stage 01 Verify Data: Contract Enumeration + Depth Check
Purpose: establish, with small test pulls only, exactly what data exists for the
frozen spec's universe — before any bulk fetching. Honesty about gaps is the whole
point of this stage. Review gate.
## Inputs
- Layer 4 (working): ../00_clarify_strategy/output/strategy_spec.md
- Layer 3 (reference): ../../CONTEXT.md (Data rules section)
- Layer 3 (reference): ../../_config/risk-params.md (window)
## Process
1. Read the strategy spec to determine the universe: underlying, leg selection
rule, expiration cadence, session.
2. Enumerate the contract universe for the window using the options/stocks
reference endpoints (browse with retrieve_all, confirm shapes with params
before any execute — house rule for every new endpoint).
3. For 2–3 sample contracts (one recent, one mid-window, one oldest needed):
- pull a small bar window (e.g., 1 day of minute bars) per leg to confirm
minute-bar availability within the spec's entry window;
- pull a small quote-history window and record the ACTUAL earliest timestamp
returned. Quote-history lookback depth is unverified — this measures it.
4. Record per-contract-leg findings in the manifest: earliest/latest verified bar
dates, quote-history earliest date, gaps, holidays in window, and any contract
the spec's selection rule would pick that has no data.
5. If the spec's window is not fully coverable (e.g., quote history shallower than
the window), state the impact plainly and propose the largest fully-coverable
window. Do not silently shrink the test.
6. No bulk pulls in this stage. Keep total pulls small (roughly a dozen).
7. Stop for operator review of the manifest before stage 02 fetches anything.
## Outputs
- data_manifest.md -> output/ (universe table, verified depth per leg, gaps,
proposed final window)

View File

@@ -0,0 +1,29 @@
# Stage 02 Fetch Cache: Single-Pass Bulk Pull
Purpose: materialize every bar/quote series the manifest calls for into the shared
SQLite cache — once. Re-runs of later stages must never re-pull data. Review gate.
## Inputs
- Layer 4 (working): ../01_verify_data/output/data_manifest.md
- Layer 3 (reference): ../../_config/risk-params.md (window, cache path, caps)
- Layer 3 (reference): ../../CONTEXT.md (Data rules)
## Process
1. Read the manifest's final (operator-approved) universe + window.
2. Author a stdlib-only fetch script into ../../shared/scripts/ (urllib for HTTP,
sqlite3 for the cache, json/csv for any side exports). No third-party packages.
3. Cache-first: for each (contract, timespan, window), check the cache and skip
rows already present. Only missing ranges are fetched.
4. Pull legs bar-by-bar: for each contract, each leg, minute bars for the spec's
session window across the manifest's date list. Respect the 50,000-bar per-pull
cap by splitting multi-month pulls into monthly sub-windows.
5. Optional per spec: pull quote history only for the eras stage 01 verified.
6. Log every pull (contract, timespan, date range, rows returned, gaps found) to
the fetch log. A pull returning zero bars is a logged fact, not an error to hide.
7. Sanity-check the cache: row counts per contract vs expected session days; flag
any contract with <50% expected coverage.
8. Stop for operator review before the engine runs.
## Outputs
- fetch_log.md -> output/ (pull table, coverage stats, anomalies)
- shared/data/backtest_cache.sqlite3 (the cache itself)

View File

@@ -0,0 +1,38 @@
# Stage 03 Run Backtest: Execute the Spec Over the Cache
Purpose: author and run the stdlib-only engine against the cached data, producing
a complete trade list and metrics. The spec is law; the engine never improvises.
Review gate.
## Inputs
- Layer 4 (working): ../00_clarify_strategy/output/strategy_spec.md
- Layer 4 (working): ../02_fetch_cache/output/fetch_log.md (coverage caveats)
- Layer 3 (reference): ../../_config/risk-params.md (costs, exits defaults)
- Layer 4 (working): ../../shared/data/backtest_cache.sqlite3
## Process
1. Author the engine script into ../../shared/scripts/ (pure Python 3.12 stdlib:
sqlite3, csv, json, statistics, math, datetime). No third-party packages —
the operator manages packaging with uv; pyproject.toml has zero dependencies.
2. Engine mechanics:
- Iterate bars chronologically per session; signals on bar N may only use
bars <= N (no look-ahead, hard rule).
- Entries only inside the spec's entry window and session (RTH default).
- Credit spread = short leg + long leg stitched; credit at entry = short
price − long price; per-bar P&L tracked on the leg diff.
- Exits in precedence order: hard time exit (15:45 ET) > stop > targets
(T1/T2/T3 scale-outs per spec) > optional breakeven-after-T1.
- Fills at bar close ± slippage from risk-params (mid ± half-spread default).
3. Emit one row per simulated trade: dates, entry/exit, leg prices, credit,
P&L, MAE/MFE, exit reason, session day.
4. Compute metrics: N, win rate, expectancy, avg win/loss, profit factor, max
drawdown, per-hour-of-day and per-weekday breakdowns, equity curve series.
5. Also run a costs-off variant (same trades, zero slippage) and store both —
the gap between them is the cost drag, and it must be visible.
6. If cached coverage has gaps inside the window, exclude affected days from
stats and list them — never interpolate through a gap.
7. Stop for operator review before the report stage.
## Outputs
- trades.csv -> output/
- metrics.json -> output/ (with-costs and costs-off variants)

View File

@@ -0,0 +1,29 @@
# Stage 04 Report: Honest Findings, Operator-Readable
Purpose: turn trades.csv + metrics.json into a report a trader can act on — or
consciously discard. No cherry-picking, no curve-fit praise. Final review gate.
## Inputs
- Layer 4 (working): ../03_run_backtest/output/trades.csv
- Layer 4 (working): ../03_run_backtest/output/metrics.json
- Layer 4 (working): ../01_verify_data/output/data_manifest.md (coverage caveats)
## Process
1. Lead with the headline numbers: N trades, win rate, expectancy per trade,
profit factor, max drawdown — for the with-costs run. Costs-off appears only
as the visible cost-drag comparison, never as the headline.
2. State the tested window, contract universe size, and data coverage honestly,
including any excluded gap days.
3. Flag low-N explicitly: below 30 trades the report must say the sample is too
small to trust and must not recommend going live on it.
4. Breakdowns: per-hour-of-day, per-weekday, exit-reason mix (stop vs target vs
time exit), and the stop-vs-target balance. These drive spec tuning.
5. MAE/MFE distribution: how deep winners usually dip before working — this is
what sets realistic stop placement and profit-target spacing.
6. End with a plain recommendation set: keep / tune (which parameter, which
direction) / discard. A losing strategy is a valid, useful result — say so.
7. Remind the reader: backtest output is research. Live execution still goes
through the normal cockpit path with its own preflight.
## Outputs
- report.md -> output/