Start freeSign in
Tool

Backtest Overfitting Audit

Eight questions about the sample you tested on, the process you followed, and the costs you modelled. Answer them against your own results and you get a written verdict, with the specific weakness named. Nothing is sent anywhere and no account is needed.

0 of 8 checks satisfied
Not evidence yet The gaps below are the ones that most often turn a good backtest into a bad first month.
The sample
The process
Costs and fills
Answers stay in your browser. Nothing is uploaded.

How this audit works

Where each group of checks catches a failure: sample before process, process before costs.

The eight checks are the failure modes that appear most often when a backtest is re-run out of sample: too short a window, a window with only one regime in it, a trial count that was never recorded, parameters re-tuned after the out-of-sample peek, and costs applied as a flat percentage of notional. Each check is binary on purpose: a partial answer here is how a weak sample gets talked into passing.

The score is a count, not a rating, and it carries no claim about what your strategy will earn. Eight out of eight means the obvious ways of fooling yourself have been ruled out. It does not mean the edge is real. The three bands exist to tell you whether you have evidence yet, and if not, which check to go and fix first.

Sources: the deflation and trial-count logic follows Bailey and López de Prado (2014); the cost-modelling checks follow the execution-realism rules Sonar Sciences applies to submitted strategies, which are published in full in the execution realism guide.