Start freeSign in
Research/Glossary/Bonferroni correction

Bonferroni correction

Reference

The Bonferroni correction is a way to make statistical testing more conservative when many tests are run at once.

The Bonferroni correction is a way to make statistical testing more conservative when many tests are run at once. It works by dividing the chosen significance level by the number of tests. If a researcher would normally use a 5 percent threshold and evaluates 20 independent hypotheses, the Bonferroni rule sets the per test threshold to 0.25 percent.

The purpose is simple. When many strategies, parameter sets, or signal variants are tested, some will appear significant by chance alone. This is the multiple testing problem. Without adjustment, the probability of at least one false positive rises as the number of tests increases. Bonferroni addresses that risk by lowering the threshold required for any single result to be called significant.

For quantitative traders, this matters because research pipelines often compare many candidate ideas before selecting one for deeper study. A standard p value threshold applied repeatedly can make random variation look like evidence. Bonferroni is a blunt correction because it does not try to model dependence across tests or estimate how much selection occurred. It is often described as conservative for that reason. But it is also transparent. The mechanism is easy to audit and easy to explain.

In practice, the correction is applied to the family of tests that belong to the same decision. If 50 strategy variants are examined as part of one search, the significance level is divided by 50. A result that passes the original threshold but fails the adjusted threshold should be treated as insufficient evidence under familywise error control.

Sonar’s strategy validation research discusses the problem of multiple testing and backtest overfitting as a central risk in systematic strategy research. It explains that repeated testing across ideas and parameter choices can inflate apparent significance and requires methods that account for selection effects. Sonar’s backtest overfitting audit tool is presented as a way to examine the extent to which a result may reflect overfitting rather than robust evidence. The deflated Sharpe ratio is also relevant in this setting because it adjusts the interpretation of an observed Sharpe ratio for non normality, sample length, and the number of trials considered, which helps evaluate whether a reported result remains statistically credible after multiple testing pressure.

A practical interpretation is that Bonferroni is useful when the main goal is honesty about uncertainty. It is especially appropriate as a baseline safeguard when many strategy tests are screened and the researcher wants a clear familywise error bound. It does not solve every validation problem, and it can be overly strict when tests are highly correlated, but it offers a simple defence against declaring success too easily after a broad search.

Covered in depth in the Strategy validation & overfitting pillar hub.

Apply this and the related checks to your own results with the Backtest Overfitting Audit.Open the audit
ShareXLinkedInFacebookEmail