Start freeSign in

Why a track record needs an audit trail

ST
Sonar Sciences Quant & Research Team · Quant & Research Team The research desk of Sonar Sciences · Publications and reviewed work
Published 7 Aug 2026
4 min read

A track record is verifiable only when it includes an audit trail. The cited Sonar sources support three parts of that standard: traceable research-to-publication workflow, independent backtest overfitting audit outputs, and metrics like the deflated Sharpe ratio that depend on transparency about multiple testing and selection. The sources support the need for timestamps, version lineage, and third-party records at the principle level, though they do not provide detailed public specifications for logging schemas or version-control implementation.

A chain of sealed evidence boxes connected by cords, each seal intact, leading back to an original source box
A chain of sealed evidence boxes, each seal intact, leading back to the original source.

A research track record is only verifiable when it can be reconstructed from independent evidence. In practice, that means more than a summary equity curve or a selected set of results. It requires an audit trail: timestamps showing when runs occurred, version records showing what code and data produced each result, and third-party records that document how the research was evaluated.

Sonar’s research-to-publishing workflow explicitly frames publication as a process built around traceability rather than presentation alone. The workflow describes a path from research to publication in which strategy research, backtests, and related outputs are organized so they can be reviewed and published with supporting records rather than as isolated claims. That matters because a track record without provenance can be curated after the fact, while a track record with preserved research artifacts can be checked against the underlying process that created it.[1]

Timestamps are the first part of that provenance. If each backtest run is logged with a time of execution, a reviewer can distinguish between contemporaneous research and retrospective selection. A timestamped sequence helps answer basic verification questions: when was a model tested, in what order were variants explored, and which result existed at a given point in time?[1]

Version control is the second part. A backtest result is only meaningful if a reviewer can tie it to the exact research state that generated it. That includes the code version, the data version, and any parameterization used at the time. Without version history, a published result can be detached from the actual research process and turned into a curated snapshot. With version history, changes over time become part of the record. Publication should be connected to the underlying research record, not merely to a final chart or metric.[1]

Third-party records are the final layer that separates a verifiable history from a self-curated one. Sonar’s Backtest Overfitting Audit is relevant here because it produces an external audit output focused on whether a backtest appears robust or likely overfit. The tool describes an audit process centered on overfitting diagnostics, including outputs intended to help evaluate whether apparent strategy quality may be inflated by selection effects. A third-party audit does not replace timestamps or version records, but it creates an additional record outside the researcher’s own narrative. That outside record is what makes it harder to rewrite history around a preferred result.[2]

This is especially important for metrics that depend on understanding the research path, not just the final score. Sonar’s glossary entry on the deflated Sharpe ratio explains that the metric is designed to adjust Sharpe-ratio interpretation in the presence of multiple testing and selection effects. In other words, if many variants were tried and only the best result was highlighted, the naive Sharpe ratio can overstate evidence. The deflated Sharpe ratio addresses that problem by accounting for the statistical inflation introduced by searching across many trials.[3]

That connection is the key reason audit trails matter. A metric such as the deflated Sharpe ratio is most meaningful when the underlying research process is transparent enough to assess how much searching occurred. If the history of experiments is incomplete, curated, or reconstructed after publication, then even a sophisticated metric can lose interpretive value because the extent of model selection is obscured. Transparent records of timestamps, version changes, and third-party audits help establish the context those metrics need.[2][3]

For quantitative traders and research analysts, the practical distinction is straightforward. A curated history shows selected outputs. A verifiable history shows the chain of evidence behind those outputs.

Claim register 3 claims · all sourced
Why a track record needs an audit trail https://sonar-sci.com/research/research-to-publishing/
Why a track record needs an audit trail https://sonar-sci.com/tools/backtest-overfitting-audit
Why a track record needs an audit trail https://sonar-sci.com/research/glossary/deflated-sharpe-ratio
Run the Backtest Overfitting Audit on your own results Eight questions about your sample, your process, and your cost model. No signup, and you get a written verdict at the end.
Open the audit

Drafted with AI assistance from cited sources. Reviewed and approved by Sonar Sciences Quant & Research Team.