Start freeSign in
Research/Glossary/Deflated Sharpe ratio

Deflated Sharpe ratio

Reference

A Sharpe ratio adjusted for the number of strategy variants tested before one was chosen, so that a strong result found across many trials is not mistaken for skill.

In practice

Sweep two parameters across ten values each and you have a hundred trials. The best of those hundred will show a strong Sharpe ratio even on random data. Deflation asks how strong the best result would have to be, across that many trials, before it stopped being ordinary, and reports the difference as a probability rather than a score.

inputs: trial count · Sharpe variance across trials · series length · skew · kurtosis

The expression, with each input labelled. Derivation and a worked example in the pillar article.

The problem the correction solves is selection. A Sharpe ratio is computed on one return series, but a research process produces many: every parameter setting, every filter toggled, every restart after a disappointing run. Reporting the best of those as if it were the only one overstates the evidence, and the more you searched, the worse the overstatement. The deflated Sharpe ratio, introduced by Bailey and Lopez de Prado in 2014, makes the search itself an input.

Mechanically, it estimates the Sharpe ratio the best of that many unskilled trials would be expected to show, given the variance of Sharpe ratios across the trials, then asks whether the chosen result clears that bar once the length, skew and kurtosis of its returns are accounted for. The output is a probability that the observed result reflects skill rather than search. A raw Sharpe ratio of two can deflate to something indistinguishable from noise once a large enough search is priced in; the arithmetic is unforgiving about trial counts.

The common failure is under-counting those trials. The count is every configuration evaluated, not the number kept, and correlated variants still contribute: a sweep across neighbouring lookbacks is closer to one trial than to five, so an honest count takes judgment as well as bookkeeping. If the count was never recorded, the deflated figure cannot be computed at all, and that absence is itself a finding about the research process.

Covered in depth in the Strategy research fundamentals pillar hub.

Related terms
This term is one of the eight checks in the Backtest Overfitting Audit.Open the audit
ShareXLinkedInFacebookEmail