Start freeSign in

Why adding rules usually adds overfitting

ST
Sonar Sciences Quant & Research Team · Quant & Research Team The research desk of Sonar Sciences · Publications and reviewed work
Published 7 Aug 2026
3 min read

Adding rule‑based filters increases strategy flexibility, which raises the risk of fitting historical noise rather than generalizable structure. Sonar’s sources support this qualitatively through their emphasis on out‑of‑sample validation, overfitting audits, and the deflated Sharpe ratio as a correction for multiple testing and false discoveries.

Why adding rules usually adds overfitting: a wordless annotated mechanism illustration
Why adding rules usually adds overfitting: a wordless annotated mechanism illustration

Systematic traders often add filters to improve a backtest: another regime check, another volatility gate, another timing condition. The statistical problem is that each added rule increases model flexibility. More flexibility gives a strategy more ways to fit historical noise, not just historical structure. In Sonar’s framing, this is a validation problem: a strategy can look strong in-sample while failing to generalize out-of-sample, and that gap is a core symptom of backtest overfitting.[1][2]

Research on strategy validation emphasizes that evaluating a systematic idea requires more than a single in-sample backtest. The key question is whether observed performance survives procedures designed to test robustness, including out-of-sample evaluation and explicit overfitting checks.[1] That directly supports the central intuition behind rule proliferation: when the researcher keeps adding conditions, the backtest has more chances to align with quirks of the sample. Even if each rule appears sensible in isolation, the combined rule set can become a highly tuned description of one historical path rather than a durable process.[1][2]

The Backtest Overfitting Audit tool is specifically described as a way to examine whether a strategy’s apparent edge is likely to be the result of overfitting.[2] In practical terms, this matters because rule‑heavy systems create more opportunities to search over variants. If many combinations of filters are tried and the best‑looking one is selected, the reported in‑sample result can be inflated by selection itself. The validation framework treats that inflation as something to be measured and discounted rather than accepted at face value.[1][2]

This is where the deflated Sharpe ratio becomes relevant. The glossary defines the deflated Sharpe ratio as a statistic intended to adjust a reported Sharpe ratio for multiple testing and non‑normal returns, so that the probability of a false discovery is reduced.[3] That concept is directly aligned with the claim that more rules raise overfitting risk. Adding independent rule conditions typically increases the effective number of strategy variations that could have been explored. As the number of tested or testable variants rises, a raw Sharpe ratio becomes less informative on its own, because some apparently attractive result can emerge by chance. The deflated Sharpe ratio is meant to correct for that reality.[3]

Under this lens, each additional condition effectively consumes degrees of freedom. The strategy is no longer just expressing one hypothesis; it is expressing a larger family of possible mappings from historical inputs to historical trades. The validation materials support the principle that stronger claims require stronger robustness checks precisely because flexible strategies can fit noise.[1][2] A filter that improves in‑sample metrics is therefore not strong evidence by itself. The improvement may reflect genuine structure, but it may also reflect the strategy using another degree of freedom to memorize the sample.

The practical implication is methodological. A new rule should be treated as an increase in model complexity that raises the burden of proof. The strategy‑validation framework points toward out‑of‑sample testing and overfitting diagnostics as the mechanism for meeting that burden.[1][2] Likewise, the description of the deflated Sharpe ratio implies that performance evaluation should be adjusted for the breadth of specification search, not just for the headline metric produced by the final chosen variant.[3]

Claim register 3 claims · all sourced
Why adding rules usually adds overfitting https://sonar-sci.com/research/strategy-validation/
Why adding rules usually adds overfitting https://sonar-sci.com/tools/backtest-overfitting-audit
Why adding rules usually adds overfitting https://sonar-sci.com/research/glossary/deflated-sharpe-ratio
Run the Backtest Overfitting Audit on your own results Eight questions about your sample, your process, and your cost model. No signup, and you get a written verdict at the end.
Open the audit

Drafted with AI assistance from cited sources. Reviewed and approved by Sonar Sciences Quant & Research Team.