← All articles
Research note

negative results · backtesting · research process

Your Failed Backtests Are Worth More Than Your Winners

Working note No. 1 — on the economics of negative results in systematic strategy research.

Abstract. Over an extended period we subjected a large number of candidate trading signals to a fixed validation protocol. The overwhelming majority did not survive. We argue that this mortality rate is not evidence of failure but the central product of the research process: a strategy graveyard, honestly maintained, is the closest thing a systematic researcher has to a moat. We relate the observed mortality to the multiple-testing arithmetic documented in the asset-pricing literature, discuss why the economics of public discourse push practitioners to write only about survivors, and argue that the discipline of burying ideas — quickly, cheaply, and without sentiment — matters more than the ideas themselves.

1. Introduction

A textbook account of quantitative research describes a pipeline: hypotheses enter, evidence accumulates, and a small number of robust effects emerge. What the textbook omits is the ratio. In our experience the ratio is brutal, and it is supposed to be.

Over the past several years we tested, by our own count, several dozen distinct families of ideas against a fixed out-of-sample protocol. These included conditioning schemes drawn from market microstructure, alternative ranking constructions, regime filters, exit overlays of every conventional flavor, and signals derived from the options market. Nearly all of them died. Some died loudly — inverting their sign out of sample, a familiar humiliation. Most died quietly, indistinguishable from noise at any confidence level a referee would accept.1

We document no individual result here, deliberately. Our interest is in the aggregate phenomenon: what it means for a research program that the modal outcome of a well-run experiment is a funeral.

2. The base rate has a literature

The mortality rate that surprises practitioners has been quantified by academics. Harvey, Liu, and Zhu (2016) catalog the factors published in the cross-sectional asset-pricing literature and conclude that, once the accumulated volume of testing is accounted for, the conventional t-statistic hurdle of two is far too permissive — a newly proposed factor, they argue, should clear a hurdle of three or more, and by that standard a large fraction of the published record is likely false. Bailey, Borwein, López de Prado, and Zhu (2014) make the practitioner's version of the argument with a sharper instrument: given enough configurations, a strategy with no skill whatsoever can be tuned to an arbitrarily attractive backtest, and the probability of exactly this manufacture rises mechanically with the number of trials.2

The implication runs opposite to instinct. A research program whose candidates mostly survive is not a skilled program; it is one that is either testing too little or rejecting too gently. The observed mortality in our own pipeline is not a cost of doing business. It is the business.

3. Why the graveyard is the product

It is tempting to regard a dead strategy as wasted effort. We take the opposite view, for three reasons.

First, a validated no is information with a long shelf life. Markets recycle ideas; the same intuition returns every few quarters wearing different clothes. A graveyard with good record-keeping converts each burial into a standing rejection that prices future temptation at close to zero. The second time an idea shows up, the cost of dismissing it is a lookup, not a study.

Second, negative results discipline the prior. Every failed conditioning scheme is a small measurement of how efficient the studied margin actually is. After enough funerals, one develops a calibrated sense — not a slogan, a posterior — of where edge is unlikely to live. This calibration is invisible in any single backtest and decisive in allocation of research time.

Third, and least appreciated: the graveyard is what makes the survivors credible. A strategy that outlived dozens of siblings under an unchanging protocol carries a different epistemic weight than a strategy that was born validated. Survivorship means little if nothing was allowed to die.

4. The publication bias of practitioners

Academic finance has a documented file-drawer problem; practitioner writing has a worse one. The blogs and threads of our industry are a museum of survivors — entries, exits, and annotated equity curves, each implying a research process with a hit rate no honest research program has ever achieved. The incentives are obvious: survivors are legible, funerals are not.

We suspect this bias does real pedagogical damage, and not only statistically. Duke (2018) gives the behavioral half of the diagnosis a useful name: resulting — the habit of grading a decision by its outcome rather than by the process that produced it. A reader calibrated on survivor-only writing is being trained to result: to treat a green equity curve as evidence of a good process and a funeral as evidence of a bad one, when the mapping in a noisy domain runs through the protocol or it does not run at all.3 When his thirtieth idea fails he concludes that he is bad at this, when in fact he is — for the first time — experiencing the true base rate.

5. Conclusion

We keep the graveyard for the same reason a serious lab keeps its failed experiments: it is the record of contact with reality. The ideas were plausible. Many were elegant. The evidence said no, and the evidence was allowed to say no — that last clause being, we would argue, the entire game.

In subsequent notes we examine particular residents of the graveyard: the protective exits that quietly sold our best trades, the signals that never cleared the noise floor, and the seductive gains that existed only in sample. The survivors, the reader will understand, stay home.


Notes

  1. The protocol under which these verdicts were reached — combinatorially purged cross-validation with ex-ante kill criteria — is described in note No. 4 of this series. Counts are approximate by design; we see no reader who is served by knowing whether the graveyard holds forty residents or sixty.
  2. The title of Bailey et al. (2014) — "Pseudo-mathematics and financial charlatanism" — is not rhetorical excess; the paper's point is that an unreported number of trials converts legitimate-looking mathematics into an instrument of (possibly self-) deception. Our pipeline's insistence on counting every trial, including the embarrassing ones, is downstream of exactly this argument.
  3. Duke's framing also supplies the correct reading of a funeral: a strategy that died out of sample under a pre-registered protocol was a good decision to test and a good decision to kill. The outcome was negative; the process was not.

References

Bailey, D.H., Borwein, J.M., López de Prado, M., Zhu, Q.J., 2014. Pseudo-mathematics and financial charlatanism: the effects of backtest overfitting on out-of-sample performance. Notices of the American Mathematical Society 61, 458–471.

Duke, A., 2018. Thinking in Bets: Making Smarter Decisions When You Don't Have All the Facts. Portfolio/Penguin, New York.

Harvey, C.R., Liu, Y., Zhu, H., 2016. …and the cross-section of expected returns. Review of Financial Studies 29, 5–68.

López de Prado, M., 2018. Advances in Financial Machine Learning. Wiley, Hoboken.

Keywords: negative results, backtesting, research process, publication bias, multiple testing, resulting.