Subject: a published mean-reversion strategy (long-only, RSI-based dip buying), replicated and stress-tested for small-cap deployment. This sample audits my own strategy candidate, with the same treatment a client engagement receives, demonstrated on a system I wanted to deploy and killed instead. Every number traces to a versioned report in my repository; reproduction instructions available on request.
An internally correct backtest of a non-existent edge is the most common expensive object in retail quant. Separating these verdicts is the audit.
Entry and exit computed on completed bars only; fills next-open. No same-bar signal-to-fill path exists in the engine. PASS.
Survivorship: corpses included (§1). Membership determined by ranking-date data only. PASS.
Fills at next session's open with slippage; dead-while-held positions force-exit at final traded price (no slippage; conservative direction disclosed). 10 forced exits across 9,887 trades: the graveyard's damage came through loss asymmetry, not frozen holdings, a finding that reversed my own pre-audit narrative.
Identical replay, one variable: costs on vs. off.
| arm | $1M becomes | CAGR | max drawdown | win rate | edge/trade |
|---|---|---|---|---|---|
| realistic costs | $14,685 | −18.5%/yr | 98.7% | 53.8% | −0.38% |
| zero costs (diagnostic) | $5,614,441 | +8.7%/yr | 49.0% | 63.3% | +0.22% |
The raw edge (~22 bps/trade) is smaller than the round-trip toll (~60 bps). Expectancy does not shrink under costs: it changes sign. Breakeven requires costs implausible for the universe. The zero-cost arm is a diagnostic, not an investable result.
9,887 trades: the sign of the result is not noise. The magnitude is window-dependent (one 19-year period).
Published defaults only; zero parameters tuned by me; evaluation gates and tripwires pre-registered before the run. The result cannot be a product of my search because there was no search.
Cross-universe: SPY 78.9% win rate (inside the published band), S&P basket 70.0%, small-caps-with-corpses 53.8%. The edge is a monotone function of the size ladder: a structure, not a data accident.
This engagement was in-sample by design (a deployment decision, not an edge claim). House policy for edge claims: a sealed holdout, opened once, judged by pre-written rules. The same protocol caught my own momentum tunings inverting out-of-sample (in-sample Sharpe 0.81, sealed-window 0.38).
What we know: the implementation is correct; the small-cap deployment loses money under any realistic cost assumption; the large-cap edge is real but thin (~0.7%/trade family-wide, per published and replicated results).
What we don't know: whether the large-cap variant survives forward costs and regime change; whether the 19-year window flatters or punishes the family.
The one experiment that most efficiently reduces uncertainty: a 6-month forward paper test of the large-cap variant at realistic cost assumptions, with evaluation thresholds written before it starts. Cost: $0 and patience. This is what I did instead of deploying.