Stationarity is the property that a time series keeps the same statistical behaviour over time.
Stationarity is the property that a time series keeps the same statistical behaviour over time. In a stationary series, quantities such as the mean, variance, and dependence structure do not drift materially across the sample. Many quantitative modelling techniques are built around this assumption because they estimate patterns from history and project them forward. If the underlying process changes, the estimated relationships can become unstable and the model can appear valid in sample while failing out of sample.
This matters in markets because raw asset prices often evolve in ways that violate stationarity. A series can trend, change volatility regime, or shift in autocorrelation structure. When that happens, a model may fit a mixture of different regimes rather than one stable process. Apparent predictive power can then come from transient structure, data mining, or sample-specific noise rather than a durable relation.
Sonar’s strategy validation research frames this problem through the lens of overfitting and multiple testing. The core issue is that repeated strategy search on historical data increases the chance of selecting a model that looks strong only because it matched noise in the sample. When the data-generating process is unstable, this risk is harder to detect because non-stationary inputs can create spuriously attractive backtest characteristics that do not generalize. Sonar’s validation framework therefore emphasizes statistical controls that adjust for selection effects rather than accepting naïve backtest outputs at face value.
The backtest overfitting audit tool operationalizes this idea by analysing a set of trials and estimating whether the selected result is likely to be overstated by the search process. The tool is designed to help identify when apparent edge may be explained by repeated testing and unstable historical structure. In practical terms, non-stationary inputs can inflate naïve performance metrics because a model can lock onto regime-specific behaviour that is not persistent. An audit that accounts for the breadth of search and the robustness of the result is a way to reduce that error.
One of the clearest examples is the difference between a raw Sharpe ratio and a deflated Sharpe ratio. Sonar’s glossary defines the deflated Sharpe ratio as a Sharpe-ratio adjustment intended to account for non-normal returns and multiple testing. The adjustment asks a stricter question than the raw Sharpe ratio asks. Instead of treating the observed Sharpe ratio as if it came from a single clean test under stable assumptions, it discounts the statistic for the fact that many variants may have been tried and that return distributions may depart from idealized conditions. This is directly relevant when data are non-stationary, because unstable environments can increase the odds that a high raw Sharpe ratio reflects sample-specific conditions rather than a persistent effect.
The mechanism is straightforward. A model is estimated on historical observations. If those observations come from a process whose properties drift over time, the model parameters are partly estimates of past regimes rather than stable features. During research, testing many parameter sets on such data can produce a subset with strong in-sample metrics by chance. Raw summary measures, including the ordinary Sharpe ratio, do not by themselves correct for that search process. Sonar’s validation materials focus on correcting for this gap by using procedures that assess whether a result remains statistically credible after accounting for overfitting risk.
For quantitative traders and researchers, the practical implication is that stationarity is not just a textbook assumption. It is a condition that determines whether historical inference is likely to transfer forward. When market data are non-stationary, model validation needs to be more demanding, especially when many ideas, parameters, and filters have been explored. In that setting, naïve metrics can overstate robustness, while adjusted validation measures such as the deflated Sharpe ratio provide a more conservative assessment of whether an apparent signal is distinguishable from chance.
The supplied Sonar sources support the general relationship between unstable data, backtest overfitting risk, and the need to adjust naïve performance statistics. They do not provide representative asset-price examples or reported Augmented Dickey-Fuller or KPSS test results, so that specific empirical demonstration should not be asserted from these materials alone.
Covered in depth in the Strategy validation & overfitting pillar hub.