Start freeSign in

How to report a backtest honestly

ST
Sonar Sciences Quant & Research Team · Quant & Research Team The research desk of Sonar Sciences · Publications and reviewed work
Published 7 Aug 2026
4 min read

Sonar’s cited materials support honest backtest reporting as a disclosure problem: publish the number of trials, sample period, and modeling assumptions such as costs, and do not hide the selection process behind a single favorable result. Sonar’s Backtest Overfitting Audit and glossary on the deflated Sharpe ratio specifically support disclosing trial count because multiple testing can inflate a raw Sharpe ratio. Sonar’s research-to-publishing guidance supports documenting sample windows and assumptions for transparency and reproducibility.

How to report a backtest honestly: a wordless annotated mechanism illustration
How to report a backtest honestly: a wordless annotated mechanism illustration

Honest backtest reporting means publishing the context that can change how a result is interpreted, not just the headline metric. That context includes at least the number of trials behind the reported specification, the sample period used, realistic transaction costs, and the full set of outcomes rather than only the favorable ones.

Sonar’s Backtest Overfitting Audit is built around the idea that a reported result can look stronger than it really is when many variants were tried before selecting the final one. The tool evaluates backtest overfitting by incorporating the number of trials and uses that information to estimate a deflated Sharpe ratio rather than relying only on the raw Sharpe ratio. In Sonar’s glossary, the deflated Sharpe ratio is defined as a Sharpe ratio adjusted for multiple testing and non-normal returns, specifically to account for selection effects that arise when many strategies or parameterizations are tested. That directly supports disclosure of trial count: without it, readers cannot judge how much selection bias may be embedded in the final result, and the raw performance metric can be overstated relative to a multiple-testing-adjusted one.

The same logic applies to the sample window. Sonar’s research-to-publishing resource emphasizes documenting methodology and assumptions so research can be evaluated and reproduced. A backtest result without the sample period omits the conditions under which it was produced. Because regime, market structure, and data availability vary across time, the sample window is not a cosmetic detail; it is part of the experiment definition. If the period is undisclosed, readers cannot assess whether the result depends on a particular environment or verify the test independently.

Transaction costs are another mandatory part of honest reporting because they determine whether simulated results reflect implementable trading conditions. Sonar’s publishing guidance stresses clear documentation of assumptions and model inputs. Costs, slippage, and related execution assumptions are among the inputs that materially affect backtest outputs. Omitting them can make reported metrics appear better than a cost-adjusted simulation would show. The cited Sonar pages support the principle that research should include the assumptions necessary for interpretation and replication.

Failures and unfavorable outcomes also belong in the report. Sonar’s research-to-publishing materials focus on making research transparent and reproducible rather than presenting only a polished result. That implies reporting the path taken to the final specification, including discarded trials and negative findings where they are relevant to understanding model selection. This matters for the same reason the deflated Sharpe ratio matters: when only the surviving result is shown, the audience sees the selected winner but not the selection process. The omission itself can be misleading because it hides how many attempts did not work.

The clearest Sonar-specific support for this article’s claim comes from the relationship between trial disclosure and the deflated Sharpe ratio. Sonar’s Backtest Overfitting Audit requires information about testing breadth to evaluate whether the observed Sharpe ratio remains persuasive after accounting for multiple testing. Sonar’s glossary explains why: the deflated Sharpe ratio is designed to reduce the inflation that can occur when researchers search across many variants. In practice, that means a backtest report that omits trial count may present a raw metric that is not comparable to one adjusted for data-mining risk.

The reporting standard that follows is straightforward:

  • State how many strategy, parameter, or rule variations were tried before selecting the reported specification.
  • State the exact sample period used for the test.
  • State the transaction cost and execution assumptions applied.
  • Report the full research context, including unsuccessful or discarded variants where needed to understand the degree of selection.
  • When presenting Sharpe-based results, consider whether a multiple-testing-aware measure such as the deflated Sharpe ratio is more appropriate than a raw Sharpe ratio alone.

Caution: without numeric examples, the claims remain conceptual rather than demonstrated.

Claim register 3 claims · all sourced
Run the Backtest Overfitting Audit on your own results Eight questions about your sample, your process, and your cost model. No signup, and you get a written verdict at the end.
Open the audit

Drafted with AI assistance from cited sources. Reviewed and approved by Sonar Sciences Quant & Research Team.