Research

Systematic Trading with Claude

Strategy and code assisted by Claude. Every trade is deterministic Python. No LLM in the trade loop.


Live (paper account)

Dashboard

Since inception -2.92%
from 2026-06-17
Indexed equity 97.08
peak 100
Max drawdown -4.95%
2026-07-29
Trading days 57
79 calendar days
Open positions 10/10
AAPL, AMD, AMZN, GLD, GOOGL, IWM, MSFT, SPY, TSLA, USO
Trades 16
13 enter / 3 exit
Indexed equity (start = 100)
Recent trades

Last 16 actions

Date Symbol Action Qty Price Reason
2026-09-04 IWM ENTER 109 296.01 pullback
2026-09-02 GLD ENTER 69 402.78 pullback
2026-08-06 AMD ENTER 13 489.28 pullback
2026-07-31 GOOGL ENTER 21 356.13 pullback
2026-07-31 AMZN ENTER 17 271.58 pullback
2026-07-30 MSFT ENTER 11 451.1 pullback
2026-07-14 USO ENTER 56 120.17 pullback
2026-07-14 NVDA EXIT 49 104.35 trail close<EMA50
2026-07-09 AMZN ENTER 58 247.04 pullback
2026-07-08 NVDA ENTER 49 204.12 pullback
2026-07-01 GOOGL ENTER 22 361.21 pullback
2026-07-01 AAPL ENTER 33 294.38 pullback
2026-06-29 TSLA ENTER 15 411.84 pullback
2026-06-29 SPY ENTER 27 741 pullback
2026-06-23 NVDA EXIT 128 200.04 trail close<EMA50
2026-06-22 GOOGL EXIT 84 349.68 trail close<EMA50

Updated 2026-09-05T08:30:48+00:00 · equity values indexed to protect account size · paper trading account (no real money).


Rules

Strategy

Long-only, daily. The engine runs the same seven checks every market close — same inputs, same outputs, no model in the loop.

01

Trend filter. — Price above the 200-EMA and the 50-EMA is rising.

02

Pullback. — Price touched the 50-EMA within the last 3 bars.

03

Trigger candle. — A bullish reversal — hammer, engulfing, or close back above the 50-EMA.

04

Exit. — Trail — sell on a daily close below the 50-EMA. Hard stop at swing_low−1×ATR(20).

05

Position sizing. — 1% of equity risked per trade. Max 10 positions, no leverage (gross ≤ 100%). 3% daily-loss stop halts new entries.

06

Event risk. — Never hold a single stock through earnings. ETFs are exempt.

07

Concentration cap. — Skip a new buy whose 60-day trailing returns correlate > 0.75 with an existing holding.


Evidence

Backtest results

Daily mark-to-market portfolio simulation, 1993–2026. Event-driven: signal on completed close, fill at next open. 5 bps/side, no leverage, same live ruleset applied throughout.

CAGR Max drawdown Sharpe Calmar
Strategy — full (1993–2026) 10.5% −27.3% 0.91 0.38
Strategy — OOS (last 20%) 15.2% −21.7% 1.04 0.70
Buy & hold SPY (full) 10.9% −55.2% 0.65 0.20
200-EMA filter only (full) 13.3% −47.3% 0.82 0.28

Statistical significance (block bootstrap, 3,000 resamples, 21-day blocks): Sharpe 0.98 [0.68–1.27] · P(beats SPY Sharpe) = 97.5% · P(smaller drawdown than SPY) = 92%.

Regime robustness (26 of 34 calendar years positive): 2008 −1% (SPY −36%) · 2020 +38% (SPY −34% intra-year) · 2022 −12% (SPY −19%).

Honest limits: costs modelled at 5 bps/side only (no market impact, borrow, or slippage); early years use a thinner universe; equity-heavy basket (correlation cap mitigates, does not eliminate). Live track record is just starting.


How we tested it

Validation

Each round exposed a flaw in the last. What started as a basic IS/OOS split became a 12-step adversarial battery — including permutation tests, block bootstrap, factor regression, random-entry ablation, and an independent reimplementation. Several things we were attached to washed out and were removed.

# Test Key finding
1 Trade-level IS/OOS expectancy +0.38R OOS gross — but no costs, wrong exit, equity-curve drawdown
2 Overfitting controls (naive baseline) Naive gate: +0.736R IS → +0.048R OOS. OOS split catches overfit
3 Risk-adjusted metrics CAGR alone misleads — Sharpe/Sortino/Calmar added throughout
4 Quant-review tearsheet Attack surface enumerated; set the agenda for rounds 5–12
5 Permutation test (Markov gate) Gate: p ≈ 0.49 — not significant. Removed from live strategy
6 Daily MTM portfolio, 5 bps/side OOS Sharpe 1.04. Beat buy&hold on every risk metric
7 Block bootstrap (3,000 resamples) P(beats SPY Sharpe) = 97.5%. Gate still not justified
8 Walk-forward across 34 years 26/34 years positive. 2008: −1% while SPY −36%
9 Correlation cap sensitivity 0.75 cap: max drawdown −32% → −27%. Adopted
10 Factor regression (FF5, HAC t-stats) β ≈ 0.28; alpha +4–5%/yr, t = 2.5–3.1, survives FF5
11 Random-entry ablation Entry signal has no edge. Random matches at Sharpe 0.94 vs 0.91
12 Independent vectorbt reimplementation 4,079 trades, both engines. +1.50 vs +1.51%/trade. Exact replication

Three things the data removed: the Markov regime gate (rounds 5/7/11), confidence in the entry signal as the source of edge (round 11), and any doubt about look-ahead bugs (round 12).

Methodology record: knowledge/validation-methodology.md · Scripts: research/


In production

Live verification

The cron log is a second test environment. Within the first three weeks of live trading it surfaced three failure modes the backtests never saw.

01

Stops never placed. — The original code submitted the protective stop as a separate order immediately after the market buy. Alpaca rejected it: "no position yet". GOOGL and NVDA entered with no hard stop for two sessions. Fix: OTO (One-Triggers-Other) bracket — the stop attaches to the buy and activates when the fill happens.

02

Weekend re-entry. — The runner checked current positions but not pending orders. Friday’s buy left a pending bracket; over the weekend the cron saw an empty position and tried to re-enter the same symbol each night. Fix: open_order_symbols() — skip any symbol with an unfilled order already in flight.

03

Broker snapshot anomaly. — On 2026-07-07 the paper account reported equity dropping −45.66% ($98k → $53k) in one session, then recovering overnight with no position changes. Cause: Alpaca paper-account snapshot bug, not market movement. The cron log caught it immediately; the publish pipeline now filters any snapshot with |dayPnL| > 20%.

Bugs 1–2 fixed in commit 7d2ac39 (2026-06-23). Anomaly filter: src/trading_mf/publish.py.


Approach

How this works

01

Claude helps with research and code. — Strategy design, backtesting, tests, docs, and bug fixes are all AI-assisted.

02

The trade loop is deterministic. — Every entry, exit, and position size is decided by plain Python from public market data. Same inputs, same outputs. No LLM in production.

03

Adversarial review by a second model. — Quant reports were written up and delivered to GPT-5.1, prompted as an adversarial lead quant reviewing a junior's work. The peer review found real holes, and a random-forest test confirmed the finding: the entry signal has no significant edge.

04

The edge is risk management, not signal. — OOS Sharpe 1.04 vs buy-and-hold 0.65. Max drawdown −22% vs SPY −55%. In 2008 the strategy finished −1% while SPY closed −36%. The strategy earns its place by not losing.

05

Paper account only. — This is research. No real money.