The in-sample period is the segment of historical data used to develop a strategy.
The in-sample period is the segment of historical data used to develop a strategy. It is where rules are chosen, parameters are tuned, and model variants are compared. Because those choices are made on that same data, strong in-sample results are expected. They show that the strategy has been fit to the sample, not that it will generalize beyond it.
Sonar’s strategy-validation research defines the in-sample period as the portion of data used for model selection and calibration, in contrast to the out-of-sample period used for validation. This distinction matters because the development process searches across rules and parameters until something looks good on the training segment. The more alternatives tested, the easier it becomes to find a specification that matches noise as well as signal. For that reason, in-sample performance is a tool for fitting. It is not evidence of future or external performance by itself.
The mechanism is straightforward. A researcher starts with a universe of possible entry rules, exit rules, filters, parameter values, and portfolio settings. Each trial is evaluated on the in-sample segment. Poor trials are discarded and better-looking trials are retained. After enough search, the surviving specification will often look strong on the same data that guided the search. This can happen even when the apparent edge is partly or mostly the result of over-fitting. The in-sample period therefore rewards adaptation to the sample, including accidental adaptation to its noise.
Sonar’s backtest-overfitting audit is designed to test that risk directly. The tool evaluates whether reported in-sample results are likely to be inflated by the research process rather than supported by persistent structure. Its purpose is not to celebrate a high in-sample metric, but to ask how believable that metric remains after accounting for multiple trials and selection effects. In practice, that means the audit examines the extent to which the final backtest may owe its strength to repeated searching over configurations. A favorable in-sample result can therefore be downgraded when the audit indicates that many opportunities existed to discover a good-looking but fragile specification.
The same logic appears in the deflated Sharpe ratio. Sonar’s glossary describes the deflated Sharpe ratio as a Sharpe-ratio adjustment that accounts for non-normal returns, sample length, and the number of trials considered during strategy development. The purpose of the adjustment is to correct the raw in-sample Sharpe for the fact that the best result among many tested variants is biased upward. A raw Sharpe observed after extensive searching is not on equal footing with the same Sharpe produced after little or no search. The deflated Sharpe ratio lowers the evidential weight of the in-sample figure when the development process creates a meaningful over-fitting risk.
This is why a good in-sample result proves very little on its own. It may be necessary for strategy development, because rules and parameters must be chosen somewhere. But necessity is not validation. The in-sample period is where a strategy is fit. Evidence begins only when the strategy faces data that did not participate in that fitting process, and when the apparent strength of the in-sample result survives tools that adjust for over-fitting risk such as Sonar’s audit framework and the deflated Sharpe ratio.
Covered in depth in the Strategy validation & overfitting pillar hub.