Why market regimes belong in your sample
4 min read
Sample completeness should include both venue coverage and regime coverage. Cross‑venue research shows market behavior differs across venues; the Backtest Overfitting Audit detects fit‑to‑sample artifacts; and the Deflated Sharpe Ratio provides a stricter lens on risk‑adjusted performance after multiple testing.
A backtest can be broad in one dimension and still be fragile in another. Covering multiple venues helps reduce venue-specific bias, but coverage in time matters just as much. A sample that never includes stressed conditions has not tested how a strategy behaves when the market environment changes sharply.
Sonar’s cross-venue data research makes the venue side of that argument explicit: market data quality and behavior vary across trading venues, and cross-venue analysis is necessary to avoid drawing conclusions from one incomplete slice of the market. The same logic applies across time. If execution conditions, liquidity, spreads, and price formation differ by venue, they also differ by regime. A sample restricted to calm conditions is therefore incomplete in a way that is structurally similar to a sample restricted to one venue.
The methodological reason this matters is not just realism but inference. Sonar’s Backtest Overfitting Audit is designed to evaluate how likely it is that an apparent backtest edge is the result of overfitting rather than a durable relationship. In that framework, the quality of the sample matters because the audit depends on the evidence contained in the tested history. If the history omits adverse periods, then the strategy has not been confronted with a full range of conditions, and any in‑sample strength may reflect tuning to a narrower regime set. A regime‑limited sample can therefore increase overfitting risk by making the tested environment easier to fit.
This is where regime coverage and venue coverage meet. Cross‑venue data broadens the market microstructure context. Regime coverage broadens the temporal context. Both aim to prevent a researcher from mistaking conditional success for general validity.
A useful way to express that distinction is through risk‑adjusted evaluation rather than raw headline outcomes. Sonar’s glossary entry on the Deflated Sharpe Ratio explains why ordinary Sharpe‑based evaluation can be misleading when many trials, parameter choices, or model variants are involved. The deflated Sharpe ratio adjusts for multiple testing and non‑normality, asking whether an observed Sharpe is still statistically credible after accounting for the research process. That matters for regime selection as well: if a researcher studies only a favorable window, the observed Sharpe may look stronger than it would in a sample that includes difficult periods. Deflation is one way to pressure‑test whether apparent quality survives a more skeptical statistical lens.
The practical implication is straightforward. A backtest that spans many venues but only benign periods is not comprehensive. It has wider spatial coverage but narrow temporal coverage. For strategies whose behavior depends on liquidity, volatility, or dislocations, omitting stress periods leaves an important failure mode unobserved.
If you are evaluating a backtest methodology, the key standard is not whether the sample is large in the abstract, but whether it is representative across the dimensions that can change the result. Venue is one such dimension. Regime is another. Leaving out stressed periods means leaving out part of the market reality the strategy may eventually face.
Drafted with AI assistance from cited sources. Reviewed and approved by Sonar Sciences Quant & Research Team.