Start freeSign in

How to compare two equity curves

ST
Sonar Sciences Quant & Research Team · Quant & Research Team The research desk of Sonar Sciences · Publications and reviewed work
Published 7 Aug 2026
6 min read

Comparing two equity curves takes three controls: the same date range, a common risk scale, and risk-adjusted statistics that account for multiple testing. Raw return differences without those controls mostly measure different risk budgets.

How to compare two equity curves: a wordless annotated mechanism illustration
How to compare two equity curves: a wordless annotated mechanism illustration

Two equity curves can look similar while representing very different exposures, sample periods, and levels of statistical credibility. A visual comparison is useful as a first pass, but it is not a reliable basis for inference. A more defensible comparison starts by aligning the observation window, then matching risk, and then evaluating risk adjusted statistics that account for multiple testing and non normality, such as the deflated Sharpe ratio.

An equity curve is the cumulative value path of a strategy over time. Its shape reflects both the return sequence and the amount of risk taken to produce that path. Because of that, two curves are only comparable if they are observed over a common interval and interpreted under comparable risk assumptions. The Sonar fundamentals material frames this directly: comparisons based on raw backtest outputs can be distorted by differences in regime exposure, leverage, drawdowns, and the number of trials used to obtain a strategy variant. The practical implication is that eyeballing one line against another can mislead when the lines are not conditioned on the same data and the same risk scale.

The first step is period alignment. If curve A starts earlier than curve B, then part of A reflects market conditions that B never experienced. That alone can change the apparent slope, drawdown profile, and ending value. To compare them, restrict both series to their overlapping dates. Once both curves start on the same date and end on the same date, rebasing each to a common initial value lets the analyst inspect path differences that arise within the same market window rather than across different windows.

A simple step-by-step example shows the mechanism of alignment. Suppose curve A covers January 2020 through December 2023 and curve B covers January 2022 through December 2023. If A trends strongly upward in 2020 to 2021 and then flattens, while B only exists during the flatter 2022 to 2023 period, the full sample chart will tend to make A appear superior because it includes an earlier favorable regime. After aligning both to January 2022 through December 2023 and rebasing them to the same starting value, the visual ranking can change. What looked like a stronger strategy may simply have benefited from being measured over an easier interval. This is the core reason alignment matters: an equity curve is regime dependent, and unmatched regimes contaminate visual inference.

The second step is risk normalization. Even on the same dates, two curves may be generated under very different volatility or drawdown budgets. A curve with a steeper upward slope may simply be taking more risk. Sonar fundamentals emphasizes that the level of returns is not interpretable in isolation from the path risks required to achieve them. In practice, one common normalization is volatility scaling. If strategy A has realized volatility twice as high as strategy B, then A can be rescaled so that its return stream has the same target volatility as B before cumulative performance is compared. This asks a cleaner question: what would the curves look like if both had been run at the same risk level.

The same logic applies to drawdown based normalization. Maximum drawdown is the largest peak to trough decline on the equity curve. If one strategy tolerates materially deeper drawdowns, then raw cumulative performance embeds a larger loss budget. Scaling or filtering comparisons so that both strategies are assessed under a similar drawdown tolerance helps separate skill from aggression. Sonar fundamentals treats these path dependent risk measures as necessary context for interpreting backtest output rather than optional annotations.

The effect on metrics is straightforward. If returns are scaled down to a lower volatility target, the cumulative path becomes less steep but also less variable. Metrics based on raw return levels change because the strategy is now being compared on a common risk basis. A plain return comparison may favor the higher risk curve, while a risk normalized comparison may show that the lower risk curve delivers similar reward per unit of risk.

The third step is statistical comparison. A plain Sharpe ratio summarizes average excess return per unit of volatility, but on its own it can overstate evidence when a researcher has tried many variants or when return distributions depart from ideal assumptions. Sonar's glossary on the deflated Sharpe ratio explains that the deflated Sharpe ratio adjusts the interpretation of an observed Sharpe ratio for non normality and for the selection bias that arises from multiple trials. The important mechanism is not just that a strategy has a high Sharpe ratio, but whether that Sharpe ratio remains statistically credible after accounting for the fact that many candidate models may have been tested.

If curve A has a higher plain Sharpe ratio than curve B, that does not automatically make A the more reliable result. If A emerged from a much larger search over parameters or signal definitions, then its observed Sharpe may be more exposed to backtest overfitting. The deflated Sharpe ratio is designed to answer a more demanding question: how likely is it that the observed Sharpe exceeds what could have arisen by chance under the research process that produced it. In that sense it is a stronger comparator than the plain Sharpe ratio when the goal is to assess whether two curves are genuinely comparable rather than merely visually attractive.

The Backtest Overfitting Audit tool is relevant here because it is built to examine whether a backtest result is robust or likely inflated by the research process. Sonar describes the tool as an audit framework for evaluating overfitting risk in strategy development. That makes it useful when two equity curves look compelling but were generated under different levels of model search, parameter tuning, or in sample optimization. An empirical use of the tool can show how an attractive unadjusted curve may lose credibility once the audit quantifies overfitting risk. In that case, visual smoothness or a strong terminal value is not enough. The audit supplies evidence about whether the curve is the product of a stable process or a favorable selection from many attempts.

The practical workflow is therefore sequential. First, intersect the date ranges and compare only the common sample. Second, rebase both curves to the same starting value so the path comparison is not dominated by absolute capital differences. Third, normalize returns to a common risk target, such as matched realized volatility or a comparable drawdown budget. Fourth, compute risk adjusted metrics. A plain Sharpe ratio can be included as a baseline summary, but the deflated Sharpe ratio is more appropriate when the strategies come from a broader research pipeline with multiple variants. Fifth, use an audit process aimed at backtest overfitting to assess whether one curve's apparent superiority survives scrutiny of the model selection process.

The advantage of this framework is that each step removes a specific source of distortion. Period alignment removes regime mismatch. Risk normalization removes leverage and risk budget mismatch. The deflated Sharpe ratio reduces false confidence from selection effects and non normal return behavior. Overfitting audit methods test whether the backtest is robust rather than merely attractive in hindsight.

The sources support the principles behind this claim, but they do not provide enough numerical detail in the cited materials to reproduce a full worked calculation for two specific equity curves here. What they do establish is the correct comparative standard: compare like with like in time, compare like with like in risk, and judge the resulting statistics with methods that account for overfitting and selection bias.

Claim register 3 claims · all sourced
Run the Backtest Overfitting Audit on your own results Eight questions about your sample, your process, and your cost model. No signup, and you get a written verdict at the end.
Open the audit

Drafted with AI assistance from cited sources. Reviewed and approved by Sonar Sciences Quant & Research Team.