← All articles
Research note

noise floor · random baselines · null model

Can Your Signal Beat a Coin Flip? Ours Couldn't

Working note No. 3 — on random baselines and the signals that cannot beat them.

Abstract. Before any candidate signal earns a place in our validation queue, it must clear a threshold most public backtests never compute: it must outperform a randomized version of itself. We describe this "noise floor" practice, report — in qualitative terms — how a family of plausible microstructure signals fared against it, and argue that the coin-flip benchmark, properly constructed, is the cheapest and most honest referee a systematic researcher can hire. We relate the practice to the multiple-testing corrections now standard in the academic factor literature, of which it is the poor man's — and, we will argue, the honest man's — implementation.

1. Introduction

Every practitioner knows the p-value near 0.05 that arrives after months of work, carrying its faint perfume of hope. Less discussed is its uglier sibling: the p-value near 0.5. We have come to regard the second as the more instructive number, because producing it reliably requires an apparatus most research pipelines lack — an explicit model of what nothing looks like.

2. The random-override benchmark

The construction is simple enough to describe in a paragraph. Take the baseline strategy. Identify the decision the candidate signal proposes to improve — an entry filter, a ranking, a veto. Replace the candidate's output with a random draw of matching frequency: if the signal would have intervened in some fraction of cases, let a coin intervene in the same fraction. Run the full out-of-sample protocol on both. Formally, for candidate c and its frequency-matched randomization , the object of interest is not the raw improvement Δ(c) but the exceedance

Δ(c) − quantile_q[ Δ(c̃) ],

the candidate's margin over the upper tail of its own shadow's distribution.1 The candidate's claim to existence is that margin, and nothing else.

The benchmark is embarrassing by design. It embodies the null hypothesis not as an abstraction but as a running competitor, and it prices in something conventional significance tests miss: the possibility that any intervention of that shape and frequency — however arbitrary — would have helped in the studied period, because the period happened to reward interference.

3. What the floor caught

We report one episode in aggregate terms. A family of signals derived from intraday price geometry — the kind of constructions that decorate every technical forum: relative weakness measures, gap-behavior conditionals, open-close asymmetries — produced, in naive backtests, improvements that a less suspicious pipeline would have shipped. Against the random override, every member of the family landed inside the noise band. Not near it: inside it. Fig. 1 gives the picture. The coin, given the same number of interventions, did as well.

Candidate signals inside the random-override noise band

Fig. 1. This figure shows the out-of-sample improvement of five candidate microstructure signals over the baseline (red markers), plotted against the distribution of improvements achieved by frequency-matched random overrides — the "coin" (gray density). Every candidate lies within the body of the null distribution: no member of the family demonstrates an exceedance over its own shadow. All quantities are synthetic renderings of the qualitative result; units are intentionally omitted.

We emphasize what this does and does not establish. It does not establish that intraday geometry carries no information; the literature on intraday return patterns — Lou, Polk, and Skouras (2019) on the overnight-intraday decomposition, for instance — documents structure at horizons adjacent to ours. It establishes that our implementations, at our horizon, on our instruments were indistinguishable from ritual. That narrower statement is the only one a validation protocol can honestly make, and it was sufficient: the family is buried, and the lookup costs nothing when its members reappear in new costumes.

4. Relation to the multiple-testing literature

The floor is our local answer to a global problem quantified by Harvey, Liu, and Zhu (2016): after decades of collective search over the same data, the effective number of trials behind any "new" signal is enormous, and the conventional t ≈ 2 hurdle has been spent many times over. Their prescription — raise the hurdle to t ≥ 3 and treat marginal discoveries as presumptively false — is aimed at the published literature, but the arithmetic applies with more force to a research program's private search, where nothing obliges the researcher to count his own trials.2 The random override enforces the counting mechanically: it does not ask how many things have been tried — a number nobody reports honestly, ourselves included — it asks the only question that number was a proxy for: does this candidate beat the luck available to its own silhouette?

5. The economics of the floor

The noise floor changes research economics in a way we did not anticipate. Ideas below the floor die in days rather than months, before they accumulate advocates. And ideas that clear the floor arrive at the expensive stage of validation pre-shrunk: the population reaching combinatorial cross-validation is small enough that the multiple-testing penalty — the quiet killer of practitioner research — stays governable.3

A referee who works for the price of a random seed, never tires, and cannot be argued with. We know of no better hire in this line of work.


Notes

  1. The quantile q is fixed ex ante with the rest of the acceptance region and is not disclosed. The distribution of Δ(c̃) is estimated from a number of independent randomizations large enough that its upper tail is stable across re-runs; we have verified the family's verdicts are unchanged under reasonable variation of this count.
  2. Bailey, Borwein, López de Prado, and Zhu (2014) derive the same conclusion from the opposite direction: the expected maximum Sharpe ratio among N skill-free trials grows with N, so an unreported trial count converts an ordinary backtest into an upward-biased order statistic. The random override can be read as an estimate of exactly that order statistic, purchased at runtime rather than assumed away.
  3. The floor also disciplines a subtler failure: the researcher who, denied one signal, resubmits it with cosmetic changes. Against a conventional pipeline this is a second free draw; against the floor it is a second race against a fresh shadow, and the shadow's price never inflates.

References

Bailey, D.H., Borwein, J.M., López de Prado, M., Zhu, Q.J., 2014. Pseudo-mathematics and financial charlatanism: the effects of backtest overfitting on out-of-sample performance. Notices of the American Mathematical Society 61, 458–471.

Harvey, C.R., Liu, Y., Zhu, H., 2016. …and the cross-section of expected returns. Review of Financial Studies 29, 5–68.

Lou, D., Polk, C., Skouras, S., 2019. A tug of war: overnight versus intraday expected returns. Journal of Financial Economics 134, 192–213.

Keywords: noise floor, random baselines, null model, multiple testing, t-statistic hurdle, microstructure signals.