Start freeSign in
Research/Strategy research fundamentals

What do the terms in a backtest report actually mean?

Exactly what they are defined to mean, which is narrower than how they are used. A Sharpe ratio is a ratio of mean excess return to volatility, not a quality score. Drawdown is a property of a path, not a feeling about a bad month. Out-of-sample means the data played no part in choosing the rules, and the moment it influences a decision it stops being out-of-sample. Most bad arguments about strategies are two people using the same word for different quantities.

This pillar is the reference layer for everything else on the site: short definitional entries, each with the definition in the first sentence, a worked example second, and a link to the fuller article that uses the term properly. The checklists and utilities live here too, because a checklist is just a set of definitions with a verdict attached. Nothing in this pillar needs to be read in order. Come here when another article uses a term you want pinned down, or before you read a vendor's performance page anywhere on the internet, ours included. A reader who holds these definitions precisely is harder to mislead, which is the entire purpose of the section.

The vocabulary of a backtest report, grouped by what each term measures: the sample, the process, the costs.

Suggested reading order

1Why arithmetic and geometric returns differArithmetic and geometric averages differ because returns compound multiplicatively. A sequence like a five percent gain then a negative five percent loss has a zero arithmetic average but still turns 100 into 99.75, so realized growth is below the simple average. For multi period portfolio growth, the geometric mean is the relevant measure because volatility creates a gap between average period returns and compounded wealth growth.3 min2How trading costs compound over timeSmall per-trade costs compound because they are paid on every trade just as the strategy’s edge is earned on every trade. Net expectancy equals gross edge minus average cost, so the break-even cost threshold is simply the gross edge per trade. When costs are modeled realistically in backtests, they reduce average returns and can materially depress the deflated Sharpe ratio, making thin-edge, high-turnover strategies especially vulnerable.3 min3How does backtesting workBacktesting works by replaying a strategy’s rules over historical data under a specific simulation design. The result is highly sensitive to the chosen dataset, transaction cost model, and event sequence because each setting changes the simulated trade path and portfolio statistics. Sonar’s materials present these choices as core parts of research design and use the backtest overfitting audit tool and deflated Sharpe ratio to help interpret results after multiple testing and specification search.5 min4How is drawdown calculatedDrawdown is calculated as the percentage decline from the running equity peak to the current value, and maximum drawdown is the largest such peak to trough drop in the sample. It is path dependent because the ordering of returns determines peaks and troughs. A backtest measures drawdown on only one realized historical path, so the deepest observed drawdown is not a bound on future risk. A broad numerical study quantifying the average gap between backtest and future maximum drawdown is not provided.4 min5How leverage changes risk of ruinLeverage linearly scales exposure, but risk of ruin rises non-linearly because losses shrink the capital base that future gains must rebuild. The recovery requirement after a loss L is G=L/(1-L), which grows disproportionately as losses deepen. Sonar’s fundamentals support the definitions of leverage and risk of ruin, while Sonar’s backtest-overfitting audit and deflated Sharpe ratio materials show why empirical ruin estimates should be treated skeptically unless the underlying backtest survives overfitting and statistical credibility checks.5 min6How long should you paper trade a strategySonar Sciences materials support the methodological case for ending paper trading based on a predefined minimum number of executed trades rather than a fixed calendar period. Their common thread is that validation quality depends on evidential sufficiency and protection against overfitting, not arbitrary elapsed time. The Deflated Sharpe Ratio especially supports the idea that confidence in performance statistics depends on the amount of underlying data.5 min7How to annualize returns and volatilityAnnualized volatility uses √T scaling and annualized mean return uses T scaling under IID or Brownian increment assumptions.3 min8How to calculate position sizeCalculate position size by dividing a predefined account risk budget by the price distance to invalidation, adjusted for instrument value and trading costs. This method makes quantity responsive to the actual risk of the setup, unlike fixed lot sizing, which leaves risk uneven when stop distances vary. Evaluation of sizing rules should use risk‑adjusted metrics and overfitting controls; direct empirical evidence comparing fixed‑lot and risk‑based sizing is not provided.4 min9How to compare two equity curvesComparing two equity curves takes three controls: the same date range, a common risk scale, and risk-adjusted statistics that account for multiple testing. Raw return differences without those controls mostly measure different risk budgets.6 min10How to read a backtest reportReading a backtest report in layers: prioritize out-of-sample and overfitting-aware evidence because in-sample quality can be misleading. Give early attention to the deflated Sharpe ratio, as it adjusts Sharpe interpretation for multiple testing and selection effects. Use core risk-adjusted measures such as Sharpe and Calmar to evaluate return efficiency under volatility and drawdown lenses. Treat descriptive statistics as context rather than primary proof of robustness.6 min11How to read an equity curveA steep, unusually smooth equity curve with shallow drawdowns should not be read as automatic evidence of a robust strategy. Overfitting and repeated optimization can create visually impressive in‑sample curves that do not generalize, and measures such as the deflated Sharpe ratio exist to adjust for this selection bias.6 min12Can you separate skill from luck in trading resultsA trading record is more consistent with skill than luck when it holds up under several tests at once: adequate sample size, stability across multiple periods, out‑of‑sample validation, and an explicit overfitting audit. The framework explains why metrics such as the deflated Sharpe ratio are important when many strategy variants have been tested. It does not support any universal minimum trade count or a single metric threshold that proves skill with certainty.6 min13What counts as a trading strategySonar’s materials support a strict definition of a trading strategy as a complete, testable rule set specifying entry, exit, and position sizing. The fundamentals research states these as core components of strategy design, the backtest-overfitting audit implies that robust evaluation requires a fully specified backtest object, and the deflated Sharpe ratio glossary entry shows that meaningful performance measurement depends on a rule-defined return series.Under that standard, anything missing one of the three components is not yet a complete strategy.3 min14What is a good Sharpe ratioThere is no universal cutoff for a good Sharpe ratio. Its interpretation depends on how many variants were tested, how long the backtest sample is, and whether commissions, slippage, and fees are included. Sonar Sciences’ Deflated Sharpe Ratio material and Backtest Overfitting Audit both emphasize that multiple testing and overfitting can make a raw Sharpe ratio look stronger than it really is. A more credible assessment uses net results, longer and more informative samples, and bias adjusted evaluation rather than a fixed rule of thumb.4 min15What is risk per tradeRisk per trade is the planned loss if a position reaches its stop. It is determined by the distance to the stop and the position size, so the effective dollar risk is stop distance multiplied by units held. Margin describes collateral, not loss at the stop. In Sonar’s framework, position sizing rules govern this risk by converting a chosen loss limit into an allowable trade size.3 min16Why losses hurt more than gains helpA loss of fraction L reduces capital to 1−L, so the gain needed to recover solves (1−L)(1+G)=1, giving G=L/(1−L). This makes recovery arithmetic asymmetric: 20% down requires 25% up, 30% down requires about 42.9% up, and 50% down requires 100% up.3 min
Tool · free, no signupBacktest Overfitting AuditEight questions about your sample, your process, and your cost model. Answer them and you get a written verdict you can keep.

Terms used in this pillar