Start freeSign in
Research/Glossary/White's reality check

White's reality check

Reference

White's reality check is a statistical test for a specific research problem.

White's reality check is a statistical test for a specific research problem. It asks whether the strategy chosen from a large set of candidates performed well because it contains genuine predictive information, or whether it looks strong only because the researcher searched across many alternatives and selected the best result.

This matters because strategy development usually involves repeated testing. A researcher may vary rules, filters, parameters, universes, holding periods, and execution assumptions. Even if none of the candidates has real edge, the best observed backtest can still look impressive by chance alone.

White's reality check addresses that selection effect directly. The null hypothesis is not merely that one named strategy has zero skill in isolation. The null is that the best‑performing strategy in the tested family is no better than what would be expected from searching across many candidates with no genuine skill. In plain terms, it compares the chosen strategy to a best‑of‑many benchmark produced by chance.

The mechanism is a statistical audit over the whole research set, not just the winner. The audit begins with the full set of candidate strategies that were evaluated during the research process. This requirement is important because the size and diversity of the search are the source of the bias. If only the final strategy is examined, the test cannot measure the inflation created by the discarded alternatives.

Next, the researcher records the original out‑of‑sample performance metric for the selected strategy. What matters is that the same metric be computed consistently across the full candidate set and in the bootstrap procedure.

The key step is to generate the null distribution. Under the null hypothesis of no genuine skill, the observed relation between strategy signals and realized returns must be broken while preserving enough of the data structure to make the comparison realistic. This can be done by reshuffling returns or by using block‑bootstrap methods. A block bootstrap is especially relevant when returns are dependent over time, because it preserves some serial structure instead of treating every observation as fully independent.

For each bootstrap sample, the researcher recomputes the performance of every candidate strategy and then takes the best value among them. Repeating this many times produces a distribution for the best performance that would be expected when many strategies are tried under the null. This is the core idea of White's reality check. The winner in the real data is judged against a null distribution built from the winner‑takes‑all selection process itself.

The result is reported as a p‑value or confidence assessment. The p‑value measures how often the bootstrapped best‑of‑many benchmark matches or exceeds the selected strategy's original out‑of‑sample performance. A small p‑value indicates that the chosen strategy did better than would typically be expected from data mining alone. A large p‑value indicates that the apparent superiority of the chosen strategy is plausibly explained by selection bias.

This is why White's reality check is stricter than a naive significance test on one strategy. A naive test ignores the research path that produced the final choice. White's reality check conditions on the fact that many candidates were tried and only the best one was kept. In that sense, it is a correction for the hidden multiple‑testing burden embedded in systematic strategy research.

Covered in depth in the Strategy validation & overfitting pillar hub.

Apply this and the related checks to your own results with the Backtest Overfitting Audit.Open the audit
ShareXLinkedInFacebookEmail