Start freeSign in

How to run an independent strategy review

ST
Sonar Sciences Quant & Research Team · Quant & Research Team The research desk of Sonar Sciences · Publications and reviewed work
Published 7 Aug 2026
5 min read

An independent review starts from complete reproducibility inputs: code, data specification, parameters and documentation. The reviewer replicates the backtest, checks data integrity, and runs statistical sanity tests before forming any opinion.

How to run an independent strategy review: a wordless annotated mechanism illustration
How to run an independent strategy review: a wordless annotated mechanism illustration

An independent strategy review is designed to answer a narrow but important question: if someone other than the author reruns the strategy, using the same materials and a standard validation process, do the reported results still appear? That logic is supported by three elements: the reviewer must receive complete reproducibility inputs; the rerun should begin with basic replication and audit checks; and the evaluation should include a metric that adjusts for multiple testing and overfitting risk rather than relying on a raw Sharpe ratio alone.

The reviewer’s required inputs start with reproducibility. Sonar’s research-to-publishing guidance says a submission must include the materials needed for another party to reproduce the research, including strategy code, data specification, parameter settings, and documentation of how the results were generated. That matters because an independent review cannot test independence if the reviewer is forced to infer missing assumptions or reconstruct undocumented steps. A reviewer needs the exact implementation, the exact dataset definition or source description, the parameter values used in the reported run, and enough procedural detail to rerun the backtest in the same way the author did. Without those inputs, differences between the original and rerun may reflect missing information rather than a substantive problem in the strategy.

What the reviewer checks first is not originality or economic intuition, but whether the reported result can be reproduced and whether the inputs look internally consistent. Sonar’s backtest overfitting audit tool emphasizes early checks around replication of the backtest, validation of data integrity, and basic statistical sanity tests. In practice, that means confirming that the same code and settings generate the same headline backtest outputs; checking that the data used in the test is the intended data, with no obvious integrity problems; and applying simple statistical diagnostics before making broader claims about robustness. These first steps are a gate, not a full verdict. If the backtest cannot be replicated or the data foundation is unclear, there is little value in moving immediately to more elaborate analysis.

The author’s absence from the rerun is the point because it removes an important source of hidden dependency. If a strategy only works when the original author is present to explain undocumented preprocessing, repair data issues on the fly, reinterpret ambiguous settings, or make ad hoc adjustments during execution, then the reported performance is not fully contained in the research package. An independent rerun tests exactly that boundary. Sonar’s framework supports comparing the original backtest outcome with the independently reproduced outcome to see whether the result survives outside the author’s direct involvement. If the two align, the strategy is less likely to depend on hidden interventions. If they diverge, the gap can reveal omitted assumptions, live fixes, or discretionary steps that were not captured in the submitted materials.

This is also why reviewer checklists should be standardized. A fixed review sequence reduces the chance that the reviewer unconsciously compensates for missing details, and it creates a cleaner comparison between the author’s reported run and the independent rerun. The checklist begins with complete inputs, replication of the stated backtest, data verification, and basic statistical checks for plausibility and overfitting risk. Only after those pass does it make sense to interpret the strategy more deeply.

A further safeguard is to evaluate the strategy with the deflated Sharpe ratio rather than treating the ordinary Sharpe ratio as sufficient. Sonar’s glossary describes the deflated Sharpe ratio as a Sharpe-based measure adjusted for the effects of non-normal returns, sample length, and multiple testing. That adjustment is relevant in strategy review because a published backtest may be the survivor of many trials, specifications, or parameter searches. In that setting, a strong raw Sharpe ratio can overstate the evidence. The deflated Sharpe ratio asks a harder question: after accounting for the number of opportunities to find a seemingly attractive result, does the observed performance still clear a meaningful statistical threshold?

Used in an independent review, the deflated Sharpe ratio helps separate reproducibility from credibility. A reviewer may be able to reproduce the exact backtest and still conclude that the result is fragile if its apparent quality is largely explained by multiple testing bias. Conversely, if the independent rerun matches the original and the deflated Sharpe ratio remains supportive after those adjustments, the evidence that the strategy’s performance is not just an artifact of repeated searching is stronger. The deflated Sharpe ratio is used for that purpose, though no universal cutoff applies to every strategy review.

Putting the pieces together, an independent strategy review on Sonar’s terms is not merely a courtesy second look. It is a structured attempt to determine whether the strategy exists as a reproducible research object rather than as tacit knowledge held by its author. The reviewer needs the full set of reproducibility inputs described in Sonar’s research-to-publishing guidance. The reviewer begins with straightforward rerun and audit steps from the backtest overfitting tool: replicate the reported backtest, verify the integrity of the data and setup, and run basic statistical sanity checks. Then the reviewer compares the independent outcome with the original and evaluates the result with the deflated Sharpe ratio to account for overfitting and multiple testing bias.

That is the clearest justification for excluding the author from the rerun: independence is how the review detects hidden assumptions, undocumented interventions, and discretionary adjustments that would otherwise remain invisible. If a strategy can be reproduced and still looks statistically credible under those conditions, the case for its research integrity is materially stronger. If it cannot, the review has still succeeded by identifying exactly where reproducibility breaks down.

Claim register 3 claims · all sourced
How to run an independent strategy review https://sonar-sci.com/research/research-to-publishing/
How to run an independent strategy review https://sonar-sci.com/tools/backtest-overfitting-audit
How to run an independent strategy review https://sonar-sci.com/research/glossary/deflated-sharpe-ratio
Run the Backtest Overfitting Audit on your own results Eight questions about your sample, your process, and your cost model. No signup, and you get a written verdict at the end.
Open the audit

Drafted with AI assistance from cited sources. Reviewed and approved by Sonar Sciences Quant & Research Team.