A small text model reads headlines beside a live trading lab. It may write stamps after the close. It may never touch a number. New paper, free PDF below; posted on SSRN.
A text model can sit next to a trading system the way a thermometer sits next to a patient: it reads, it never prescribes. That is the whole design. The paper follows one such model, a small commercial judge pinned to one version, that answers only typed questions about headlines: a probability for a yes-or-no question, or one word from a fixed list. Never free text.
Before it was allowed to answer anything that mattered, it had to pass six gates. They were assembled over three days, most of them after a failure. We present them as the checklist we would now apply up front.

Gate A. Read a plain fact against a real label
A text signal must first show it can read a fact from headlines. Ours was plain: a scheduled company event is reported or imminent for this name this morning. The first test failed its pre-registered bar by 0.013 against a label that turned out to be wrong: a price-derived proxy that caught about half of the true event mornings and raised as many false flags. We parsed a real earnings calendar from the news feed's own daily schedule article and held a second, pre-registered hearing on rows nobody had opened. The judge scored 0.981 on 39 rows. A sanity check, not a measurement; the paper says so.
Gate B. Beat the dumbest baseline
For "is there an event", the dumbest baseline is to count the headlines. On the same 39 rows the count alone scores 0.889. The judge's margin is 0.092, and on a bootstrap it touches zero. The paper calls it borderline rather than rounding it up. The house example of losing to this baseline came later: a "news regime" reading that passed its bar, until we built the baseline from the judge's own input and the headline count reproduced the whole effect. The reading became a counter.
Gate C. No memory of famous mornings
Hand the judge six widely reported event mornings from 2022 to 2024 with the headlines removed and the real date kept. Every one came back at 0.03, identical to quiet controls. A refresh on four more public mornings gave the same. The consequence is narrow: this judge, at this version, does not recall these events without their text. It does not prove old-headline backtests free of look-ahead.
Gate D. Give the same answer twice
Repeat noise is about plus or minus 0.05 on confident answers. A fixed panel of 50 stamped mornings is re-asked weekly, after the close. Tone agreement below 0.90 is declared drift and moves the since-date of every row that depends on the judge.
Gate E. The date and the window travel with every headline
The judge reads dates: the same scheduled-event headline scored 0.97 with the real decision date and 0.4 with a dummy one. So the decision-time date is passed on every call, and the window is part of the question. The same rows read under two window clocks gave two different answers.
Gate F. Test the forward judge before it ships
Every forward-tracked item carries a verdict rule written before its data. One such rule, run against random flag sets, said "integrate" 49.8% of the time. The fixed rule requires the flagged mean at or below the fifth percentile of random same-size sets, a minimum flagged count, and fixed looks instead of nightly re-evaluation. The five text rows have not yet been run against random flags; until they are, their verdicts carry the flaw this gate found.
Where it may act, and where it never acts
Allowed: stamps written after the close, outside the live process, counted forward on a board where every row has a pre-written judge and a fixed review count. Never: thresholds, sizing, edge validation, risk measurement, parameter search. A judge inside a parameter search launders overfitting: it supplies a plausible story for whichever cell the search prefers. No live code reads any of its fields.
The plumbing
None of the failures that cost a verdict was about the model. An unpaged feed turned every six-hour window into the last hour, two thirds of it analyst notes; a study's "nothing" was withdrawn and re-run. A window ending at the wrong clock changed the answer on identical rows. A stamp test that searched the source for a function name tested nothing; a stray comment would have failed every held name silently. Three "tracked" items had no code behind them. The rule now: any forward claim names, in the same entry, the file that writes it, the health check that proves it alive, and the row that counts it.
What the stamps have shown
Small, and honestly measured. One fact read well. One small effect below its pre-registered bar, with a consistent sign. Several that vanished. Table 1 in the paper lists all 22 uses of the judge and their verdicts, failures included, because a reader needs the full count to judge any single pass.
Where to read it
The full paper, eleven pages with the question texts and the gate protocols in the appendix, is on its abstract page as a PDF, with BibTeX and the Zenodo archive. It is also on SSRN as abstract 7589261, DOI 10.2139/ssrn.7589261. It is the companion to the parity paper: that one asked whether the backtest runs the same system as live; this one asks what a text instrument is allowed to do next to it.
One sentence from the abstract that we mean exactly as written: this is a paper about checking data and the instrument, not about profit. We make no claim about either strategy's returns.
Keywords: large language models, text classification, look-ahead bias, pre-registration, research infrastructure.