Start freeSign in
Research/Glossary/Multiple testing problem

Multiple testing problem

Reference

A multiple testing problem appears when many hypotheses are tested at once.

A multiple testing problem appears when many hypotheses are tested at once. In strategy research, each backtest is effectively a hypothesis test about whether a trading rule contains real signal rather than noise. If enough backtests are run, some will produce strong looking metrics by chance alone, even when no true edge exists. This is why the evidentiary bar for accepting a strategy must rise with the number of trials.

Sonar’s strategy validation research states the core mechanism directly. As the number of backtests tried increases, the probability of seeing an apparently significant result from luck also increases. A strategy that looks compelling after searching across many variants can therefore be a false discovery unless the evaluation accounts for the size of the search process. The practical implication is that a raw performance statistic is not enough. It must be interpreted in the context of how many tests were run and how much selection occurred during research.

This is the same statistical issue addressed by multiple testing corrections such as Bonferroni style adjustments and false discovery rate controls. The general idea is to tighten the acceptance threshold as the number of trials grows, because a fixed threshold no longer preserves the same false positive risk once many tests are conducted. Sonar’s research frames this in backtesting terms: when many strategy variants are explored, the observed best result is biased upward by selection, so stronger evidence is required before calling that result robust.

The effect can be described using performance metrics such as the Sharpe ratio. If a researcher evaluates many independent candidate strategies, the maximum observed Sharpe ratio across the set tends to increase even if the candidates are only random. That means a high in sample Sharpe ratio is not automatically persuasive when it emerged from a broad search. Sonar’s strategy validation material highlights this point by emphasizing that the more trials a researcher performs, the more likely it is that one trial will look unusually good by chance. The false positive rate therefore grows with the number of backtests unless the threshold is adjusted.

Sonar’s glossary defines the deflated Sharpe ratio as a statistic designed to account for this problem. It adjusts the interpretation of an observed Sharpe ratio for factors including multiple testing and non normality, with the goal of estimating whether the observed result is genuinely distinguishable from luck. In plain terms, the deflated Sharpe ratio asks a harder question than the ordinary Sharpe ratio. It does not simply ask whether returns look good relative to volatility. It asks whether the observed Sharpe ratio remains convincing after considering how many opportunities there were to find an attractive result by searching.

That adjustment is exactly why the required evidence threshold rises with the number of backtests. As the count of tested variants increases, the benchmark for statistical credibility becomes more demanding. A result that might look notable when it is the only strategy tested may look ordinary once it is understood as the best outcome from a large research sweep. Sonar’s glossary usage of the deflated Sharpe ratio reflects this logic by treating multiple trials as part of the evidence calculation rather than as irrelevant background.

Sonar’s Backtest Overfitting Audit tool applies this concept operationally. The tool evaluates a backtest in light of the number of trials and related overfitting inputs, and it reports an adjusted standard for judging robustness rather than relying on the raw headline metric alone. The tool’s purpose is to audit whether a result may be explained by overfitting from repeated testing. In that framework, increasing the number of backtest trials increases the adjustment burden, so the strategy needs stronger statistical support to clear the robustness bar.

This creates a simple rule for research interpretation. The best backtest result in a large batch should not be judged by the same standard as a result that emerged from a small and disciplined set of tests. The first has a much higher chance of being a random winner. The second has less selection bias. Sonar’s materials consistently present this as a validation problem, not just a performance reporting problem. Robustness depends on the path taken to arrive at the result, including how many ideas were tried and discarded.

For quantitative traders and strategy developers, the practical lesson is conceptual. Search intensity changes the meaning of evidence. As the number of backtests rises, so does the chance of a spurious standout, and the threshold for statistical credibility must rise accordingly. Metrics that explicitly adjust for multiple testing, such as the deflated Sharpe ratio, are useful because they formalize that intuition instead of ignoring it.

Covered in depth in the Strategy validation & overfitting pillar hub.

Apply this and the related checks to your own results with the Backtest Overfitting Audit.Open the audit
ShareXLinkedInFacebookEmail