Working note No. 5 — on holding-period sweeps and the gains that evaporate out of sample.
Abstract. A recurring temptation in short-horizon strategy research is the holding-period extension: sweep the exit clock, observe that longer holds earn more in the backtest, and conclude that the strategy has been leaving performance on the table. We ran the sweep. In sample, patience was profitable at every extension tested; out of sample, every increment of the apparent gain disappeared. We discuss the mechanism — longer holds load the position onto risk the entry signal knows nothing about — and propose a general suspicion: in backtests, patience is among the easiest virtues to counterfeit.
1. Introduction
The experiment is irresistible because it requires no new idea. The strategy already exists; the only question is whether its exit clock is too hasty. One sweeps the holding period across a grid, the equity curve rises with the horizon, and the conclusion writes itself: we have been amputating our own trades.
We report, in the qualitative terms this series permits, why we no longer believe that conclusion — and why our exit clock stayed short.
2. The sweep and its undoing
In sample, the pattern was textbook: monotone improvement with horizon, comfortable significance at the longer holds, an annualized uplift large enough to justify the meeting. Under the full protocol — combinatorial folds, purged boundaries, criteria fixed ex ante — the ordering collapsed. Fig. 1 shows the two curves in stylized form. The longer-hold variants were not merely weaker out of sample; their advantage was distributed exactly as one would expect if it had been assembled from a handful of in-sample episodes, present in the folds that contained those episodes and absent from every fold that did not.1
Fig. 1. This figure shows strategy performance as a function of holding period under two evaluation regimes. In sample (dashed gray), performance rises monotonically with the horizon — patience appears free. Out of sample under the full protocol (solid red, with dispersion band), the ordering reverses: performance is flat to declining in the horizon, and the chosen clock (black marker) remains at the short end. Curves are synthetic and illustrative; tick labels are intentionally omitted.
The diagnosis, once seen, is unremarkable. A short-horizon entry signal is a claim about the next hour, not the next day. Extending the hold does not extend the claim; it simply appends, to a brief period of informed exposure, a long period of uninformed exposure — beta with extra steps. In a rising sample, uninformed exposure is paid. The sweep was not measuring the signal at longer horizons. It was measuring the sample.
3. A general suspicion
We suspect the result generalizes, for a structural reason: any parameter whose extension increases time in market will inherit the sample's drift. Holding periods are the purest case, but the family is large — looser stops, wider re-entry windows, slower rebalancing clocks. Each of these, swept in a favorable sample, will produce the same seductive monotone curve, and for the same reason a longer beach holiday produces a better tan. The multiple-testing arithmetic compounds the problem — a swept grid is a family of trials, and the best point of a family is an order statistic, not an estimate2 — which is why the sweep's winner faces the deflated-Sharpe accounting of Bailey and López de Prado (2014) before it faces the folds.
The overnight-intraday decomposition of Lou, Polk, and Skouras (2019) offers a complementary caution from the academic side: expected returns are not homogeneous across the clock — their decomposition finds entire anomaly premia concentrated in one component of the close-to-close return, with the other component earning nothing or the opposite sign. A hold extended across periods with different return-generating clienteles is not "more of the same trade"; it is a different trade wearing the first one's entry ticket.3
4. Conclusion
Our exit clock remains short — how short stays home, as usual. The general finding travels well enough without the number: when a backtest tells you that patience is free, the first hypothesis should be that the sample, not the signal, is the patient one.
Notes
- Fold-level attribution is, in our experience, the single most informative diagnostic the protocol produces: an effect whose profitability concentrates in the folds containing a particular episode is not an effect with an episode problem, it is an episode with a marketing department.
- Harvey, Liu, and Zhu (2016) make the same point at the scale of the published literature; the program-level version differs only in that the trials are at least countable, if one has the constitution to count them.
- Carver (2015) arrives at the operational corollary from the systematic-management side: the parameters of a rule-based system should be set by design-time evidence and then frozen, precisely because the temptation to "let winners run a little longer" is a post-hoc edit whose evidential support is always, on inspection, the recent sample. Our configuration-freeze discipline follows his reasoning.
References
Bailey, D.H., López de Prado, M., 2014. The deflated Sharpe ratio: correcting for selection bias, backtest overfitting, and non-normality. Journal of Portfolio Management 40 (5), 94–107.
Carver, R., 2015. Systematic Trading: A Unique New Method for Designing Trading and Investing Systems. Harriman House, Petersfield.
Harvey, C.R., Liu, Y., Zhu, H., 2016. …and the cross-section of expected returns. Review of Financial Studies 29, 5–68.
Lou, D., Polk, C., Skouras, S., 2019. A tug of war: overnight versus intraday expected returns. Journal of Financial Economics 134, 192–213.
Keywords: holding period, parameter sweeps, in-sample drift, deflated Sharpe ratio, overnight returns.