Start freeSign in
Research/Glossary/Backtest

Backtest

Reference

A backtest is a simulation of a strategy’s rules on historical data.

A backtest is a simulation of a strategy’s rules on historical data. Its usefulness does not come from the fact that it uses the past; it comes from whether the past is represented honestly enough to stress the strategy under conditions that could actually have been faced at the time. Research quality is based on careful handling of data, assumptions, and evaluation discipline, and two practical failure modes in strategy evaluation are overstated statistical significance and backtest overfitting. A backtest’s reliability depends on the integrity of the historical inputs and on whether implementation frictions and statistical corrections are modeled rather than ignored.

The first pillar is historical data integrity. A backtest consumes a historical record of prices, signals, and any conditioning information used by the strategy. If those records are incomplete, cleaned in a way that uses future knowledge, or restricted to instruments that survived until the end of the sample, the test stops being a realistic historical simulation and becomes a distorted reconstruction. Robust quantitative research requires disciplined treatment of data and assumptions. So the defensible claim is methodological rather than numeric: if the data pipeline excludes adverse historical outcomes or injects information unavailable at the decision time, the backtest becomes less reliable.

The second pillar is trading cost realism. Gross performance is not the same thing as implementable performance. A historical simulation that omits commissions, fees, financing, spreads, market impact, or slippage will generally describe an easier world than the one a live strategy must face. Assumptions matter, and research should distinguish between an abstract rule set and the frictions involved in implementation. If cost assumptions are understated or missing, apparent performance is inflated relative to a more realistic implementation model.

This connects directly to overfitting risk. Sonar’s backtest-overfitting audit tool is relevant because a strategy can look strong in-sample even when its apparent edge is largely a byproduct of repeated testing, selective retention of variants, or insufficient correction for multiple trials. In that setting, the problem is not only the rule set itself but the full research process: the data segmentation, the number of ideas tried, the choice of final specification, and the assumptions used to convert paper trades into net results. The audit tool is therefore best understood as a control against false confidence. It is not a guarantee that a strategy is valid; it is a way to examine whether the reported backtest may owe too much to the search process.

The statistical side of that control is clarified by the deflated Sharpe ratio. A plain Sharpe ratio can look persuasive even when many variants were tested or when the sample is not long enough to justify confidence. The deflated Sharpe ratio adjusts interpretation by accounting for effects such as non-normality and multiple testing, making it a more conservative lens for judging whether a reported Sharpe is likely to reflect genuine skill rather than luck. In practical research terms, this matters because a backtest with optimistic data handling and optimistic cost assumptions can already be overstated before one even asks whether the measured Sharpe survives statistical deflation. A high raw Sharpe produced by favorable assumptions is not the same thing as a robust signal. The deflated Sharpe ratio is useful precisely because it pushes evaluation away from headline metrics and toward evidence that can survive more realistic scrutiny.

The trustworthiness of a backtest is governed by the honesty of its construction. That includes whether the historical dataset is assembled without look-ahead contamination and survivorship distortions, whether execution frictions are modeled rather than idealized away, and whether the reported performance is evaluated with tools that account for multiple testing and overfitting.

For a quantitative trader or research analyst, the implication is straightforward. A backtest is not evidence because it is precise; it is evidence only to the extent that the data, costs, and statistical interpretation are honestly specified. Clean but biased data can mislead. Detailed but unrealistic cost settings can mislead. Strong-looking risk-adjusted metrics that ignore the breadth of the research search can mislead. Sonar’s materials point toward a more defensible standard: test on historically faithful data, include realistic implementation assumptions, and audit the resulting statistics for overfitting and multiple-testing effects before treating the output as informative.

Covered in depth in the Strategy research fundamentals pillar hub.

Apply this and the related checks to your own results with the Backtest Overfitting Audit.Open the audit
ShareXLinkedInFacebookEmail