Start freeSign in

What is a deflated Sharpe ratio?

ST
Sonar Sciences Quant & Research Team · Quant & Research Team The research desk of Sonar Sciences · Publications and reviewed work
Published 7 Aug 2026
9 min read

A deflated Sharpe ratio is a Sharpe ratio corrected for the number of strategy variants you tested before choosing this one. It asks a different question than the raw figure: not how good the result looks, but how likely a result this good is to appear by chance, given how many times you looked. Try enough variants and a strong raw figure becomes ordinary, which is why the raw number cannot tell you whether you found an edge or a coincidence.

What is a deflated Sharpe ratio?: a wordless concept illustration
Best-of-N Sharpe ratios on returns with no edge. The dashed line is where a single trial would land.

The correction comes from Bailey and Lopez de Prado, who published it in 2014 in The Journal of Portfolio Management. It takes the number of trials you ran, the variance of the Sharpe ratios across those trials, and the length, skew and kurtosis of the return series you kept, then reports the probability that the best result you found reflects real skill rather than search. In practice the input people get wrong is the trial count: it is every parameter combination you evaluated, not the number of strategies you kept.

Nothing about the idea is exotic. Multiple testing is the same problem a medical trial faces when it checks many endpoints and reports the one that moved. The finance literature has been explicit about the scale of it: Harvey, Liu and Zhu argue that a newly reported factor should clear a t-statistic near three rather than the classical two, precisely because the field as a whole has run so many searches. The deflated Sharpe ratio is the same instinct turned into a procedure you can apply to your own search history, on your own results, before anyone else has to.

Why does a raw Sharpe ratio overstate an edge?

Because the ratio is computed on one series, and you chose that series. Sweep two parameters across ten values each and you have a hundred candidate series; the best of a hundred noise processes still looks impressive. The Sharpe ratio has no term for how many alternatives you rejected, so nothing in the arithmetic can distinguish a lucky survivor from a considered result.

Bailey, Borwein, Lopez de Prado and Zhu showed the effect formally: the expected maximum Sharpe ratio across trials keeps rising as the trial count grows, even when every strategy tested has exactly zero skill. The lead figure plots that distribution for one, ten, a hundred and five hundred trials on random returns. Selection does the work an edge is supposed to do, which is why the raw number flatters every search, and flatters the largest searches most.

How many trials counts as too many?

There is no clean threshold, and any number you are given is a heuristic. What matters more is whether the trials were independent. Sweeping a lookback from eighteen to twenty-two gives you five highly correlated results, not five trials; the honest unit is a roughly independent region of the parameter space. Count those regions, be honest about restarts after a bad result, and include the variants you abandoned without finishing, because abandoning is also selection.

Under-reporting the count is the most common way a deflated Sharpe ratio ends up as optimistic as the raw one. The count also cannot be reconstructed afterwards, which is the practical lesson of the whole literature: record trials as you run them, in the experiment log, at the time. An unrecorded search is indistinguishable from an unrun one, and it will later be priced at whichever number suits the person doing the reporting.

How do you compute it on your own results?

You need four ingredients: the Sharpe ratios of every variant you evaluated, the length of the return series at the frequency the ratio was computed, and the skew and kurtosis of the returns of the variant you kept. From the trial ratios you estimate what the best of that many unskilled attempts would be expected to show. The deflation then asks whether your chosen result clears that expectation once the shape of its return distribution is taken into account, and reports the answer as a probability rather than a score.

Use it as a gate, not a garnish. Below whatever threshold you set in advance, the result goes back into research rather than forward into anything that matters. The failure modes of the calculation itself are mundane: annualising the trial ratios inconsistently, mixing frequencies between the series length and the ratio, and averaging away the variance across trials are the usual three, and each one biases the deflation toward mercy.

What does the deflated figure not tell you?

It does not tell you the strategy will work, and it does not replace a holdout. Deflation prices the search you ran on the data you had. Out-of-sample testing checks behaviour on data the search never saw, and walk-forward validation checks whether the fitting procedure keeps working as the window moves. The three answer different questions, and a strategy can clear all of them and still fail live, because live trading adds costs, latency and regime change that no historical procedure fully rehearses.

Treat them as a sequence, with deflation first. It is the cheapest of the three and the most often skipped, because it requires an honest trial count, and the trial count is the most embarrassing number in quantitative research. That is exactly why it belongs in the report: a result presented together with its search history is evidence, and the same result presented without one is an anecdote.

Claim register 3 claims · all sourced
The deflated Sharpe ratio corrects for selection bias under multiple testing and for non-normal returns, taking the trial count, the variance of Sharpe ratios across trials, and the length, skew and kurtosis of the chosen return series as inputs Bailey and Lopez de Prado (2014), The Journal of Portfolio Management 40(5), 94-107
The expected maximum Sharpe ratio across trials rises with the number of trials even when every strategy tested has zero skill Bailey, Borwein, Lopez de Prado and Zhu (2014), Notices of the AMS 61(5), 458-471
A newly reported factor should clear a multiple-testing hurdle near a t-statistic of three rather than the classical two Harvey, Liu and Zhu (2016), The Review of Financial Studies 29(1), 5-68
Run the Backtest Overfitting Audit on your own results Eight questions about your sample, your process, and your cost model. No signup, and you get a written verdict at the end.
Open the audit

Drafted with AI assistance from cited sources. Reviewed and approved by Sonar Sciences Quant & Research Team.