Start freeSign in
Research/Glossary/False discovery rate

False discovery rate

Reference

When a research process evaluates many backtests, the main statistical problem is not whether one strategy looks good in isolation.

When a research process evaluates many backtests, the main statistical problem is not whether one strategy looks good in isolation. It is how many of the selected strategies are false positives created by repeated testing. In that setting, the false discovery rate, or FDR, is the quantity that directly answers the practical question: among the strategies you are calling discoveries, what fraction should be expected to be false.

Sonar’s strategy-validation research frames this as a multiple-testing problem created by large-scale strategy search. As the number of trials rises, the chance of finding apparently attractive results by luck also rises, even when no true edge exists. The article emphasizes that validation must therefore adjust for the breadth of the search rather than rely on naive single-test significance logic. In that framing, FDR control is the operational metric for mass research because it targets the expected proportion of claimed discoveries that are false, instead of only controlling the probability of at least one false positive.

This distinction matters for systematic trading research. In a broad search, some false positives are nearly unavoidable; the question is whether their share among selected candidates remains acceptable. Family-wise error control can be too strict for exploratory research because it aims to avoid any false positive at all. FDR control is designed for the more realistic objective of limiting the expected false share among selected results, which makes it more aligned with high-throughput strategy development workflows described in Sonar’s research materials.

Sonar’s backtest overfitting audit tool is presented as a way to quantify overfitting risk produced by strategy selection. Its purpose is to audit the research process, not just the final chosen backtest, by examining how search and selection affect the credibility of apparent discoveries. That aligns naturally with FDR-based thinking: once many variants, filters, or parameter combinations have been explored, the relevant output is not merely a headline performance statistic but an estimate of how selection has inflated the number of apparent winners.

The supplied Sonar sources also support comparison with the deflated Sharpe ratio, but only at the level of role and interpretation. Sonar’s glossary defines the deflated Sharpe ratio as a Sharpe-ratio adjustment intended to account for non-normal returns, short samples, and multiple testing. That makes it a useful single-strategy diagnostic for asking whether an observed Sharpe remains statistically credible after accounting for data‑mining effects. But this is not identical to FDR control. The deflated Sharpe ratio evaluates whether an individual strategy’s Sharpe is likely to be genuine after deflation for testing and distributional issues, whereas FDR addresses the expected false share across the set of strategies selected from a large search. They are complementary rather than interchangeable.

For that reason, a mass-search workflow can use the two ideas at different layers. FDR control governs the list of discoveries produced by the search process. The deflated Sharpe ratio provides a more conservative interpretation of each candidate’s Sharpe after acknowledging non‑normality, sample length, and multiple trials. Sonar’s materials support that conceptual division.

In mass strategy search, the key validation target is not simply whether selected backtests look profitable, but the expected proportion of those selections that are false discoveries. That is the role of FDR. Sonar’s validation framework and overfitting audit are consistent with this perspective because they focus on the distortion introduced by large‑scale search and selection. The deflated Sharpe ratio adds a complementary check on individual strategies by adjusting apparent Sharpe for multiple testing and distributional effects.

Covered in depth in the Strategy validation & overfitting pillar hub.

Apply this and the related checks to your own results with the Backtest Overfitting Audit.Open the audit
ShareXLinkedInFacebookEmail