Five years of classified financial headlines, tested against forward returns at one to three days. A hundred and twenty tests. Nothing survives.
“Aren’t most of these headlines just commentary on a move that already happened?”
Yes. Measurably so — and it changes nothing. Headlines arriving within four hours of another land on a stock that has already moved 1.30× its normal amount; those arriving after three days of silence land on one that has been unusually quiet, at 0.79×. Isolating that clean 1.4% of the corpus produces an edge of +0.2 basis points.
See the test →“What if I only need the occasional big winner — cap every loss at 2% and let the rare 20% run?”
News does fatten the tails, and the effect grows the further out you look: 1.35× the normal rate of +20% moves. But −20% moves fatten to 1.33× at the same time. Both ends widen together. Run as a strategy over ten days, the 2% stop is hit 71.1% of the time and the 20% target 3.6%.
See the test →“Is there anything here I could actually act on?”
One effect survived every test: news filed before the open produces a gap that keeps going during the session. +20.6 bps per event across 7,781 events, positive in all six years. It is real, it is documented elsewhere in the literature, and it is almost exactly the size of the cost of harvesting it.
See the test →Each test asks one question: after subtracting the stock’s own drift and the market’s move, does this category of headline predict which way the price goes next? Every answer came back no.
p < 0.05 bar.
That sounds like a finding until you count: 120 tests of pure noise produce six such
results on average, purely by chance. Eight is not a discovery. It is what an empty
room looks like when you take 120 photographs of it. After correcting for the number
of looks, the strongest result in the entire study lands at
q = 0.454.Analyst upgrades, earnings beats, guidance changes, dilutive offerings — these are the headlines with genuine information in them, and they are unambiguous about direction. The five-year sample is large enough to have caught a real effect in most of them. The detectable column is the smallest edge each test had an 80% chance of finding.
| category | events | observed | detectable | q |
|---|---|---|---|---|
| legal_trouble | 1,191 | −6.5b | 18.4b | 0.928 |
| analyst_upgrade | 967 | +3.0b | 25.9b | 0.936 |
| earnings_beat | 313 | −10.1b | 58.0b | 0.936 |
| analyst_downgrade | 308 | +10.8b | 45.3b | 0.936 |
| dilution | 205 | −11.8b | 49.1b | 0.936 |
| guidance_up | 158 | +9.5b | 98.6b | 0.936 |
| buyback | 131 | −11.7b | 51.3b | 0.936 |
| guidance_down | 117 | +17.5b | 58.7b | 0.918 |
| pt_cut | 107 | −6.6b | 58.8b | 0.936 |
| acquisition | 76 | −58.7b | 64.6b | 0.454 |
| pt_raise | 66 | −31.3b | 70.4b | 0.789 |
| earnings_miss | 56 | +27.6b | 119.9b | 0.936 |
Analyst upgrades are the clearest case. Across 967 of them, a 26 basis point edge would have shown up. The measurement came back at three. That is not a sample too small to see anything — it is a good look at nothing.
Four events. Four coin flips landing the same way, which happens one time in sixteen.
Corrected for the 120 tests it sits inside, it becomes q = 0.498
— a coin flip about a coin flip.
Reported on its own, stripped of the other 119 tests it was found among, this is a strategy with a perfect record. That framing is not a lie about the arithmetic. It is just a lie about where the number came from.
| category | events | edge | vs. cost | overlap |
|---|---|---|---|---|
| move_up | 14,925 | −6.3b | under 9.0b | 92% |
| move_down | 7,954 | +16.6b | marginal | 86% |
| lexicon_positive | 8,110 | −7.3b | under 9.0b | 86% |
| lexicon_negative | 5,023 | +3.5b | under 9.0b | 79% |
A round trip costs about nine basis points in spread and slippage on liquid names. Every measured edge here sits at or below that line, so even taking the numbers at face value there is nothing left after paying to trade them.
The overlap column is the share of events whose measurement windows share price bars with a neighbour. At 86–92%, these are not independent observations, which means the p-values are already flattering. The null is stronger than it looks.
A null on direction is not a null on everything. Three patterns in this data are real, and none of them is an edge.
Whatever a headline does to a price, it has finished doing it well before the close. At two and three days out the figures are 0.995× and 1.000× — indistinguishable from an ordinary Tuesday.
| category | events | vs. normal | q |
|---|---|---|---|
| guidance_up | 158 | 1.36× | 0.036 |
| earnings_beat | 313 | 1.28× | 0.007 |
| move_down | 7,954 | 1.11× | <0.001 |
| move_up | 14,925 | 1.09× | <0.001 |
| lexicon_negative | 5,023 | 0.89× | <0.001 |
| legal_trouble | 1,191 | 0.82× | <0.001 |
| approval | 198 | 0.81× | 0.007 |
| acquisition | 76 | 0.79× | 0.045 |
Earnings and guidance arrive during genuinely turbulent stretches. Lawsuits, FDA approvals, buybacks and acquisitions are followed by days quieter than average — a fifth below baseline, at vanishing q-values.
The tempting reading is that boring news makes for boring days. It is more likely backwards. A newswire covers large caps continuously, so administrative filler gets written because nothing else is happening. The category is describing the day it was published into, not forecasting the one that follows. With publisher timestamps that confound cannot be separated out here.
| after “stock fell” news | median | mean | share up |
|---|---|---|---|
| 1 day | −8.6b | −2.5b | 48.0% |
| 3 days | −13.6b | +2.3b | 48.2% |
At three days the typical outcome drifts down and the average outcome drifts up. Occasional sharp rebounds outweigh the many small continued declines.
A strategy shorting after down-news would be right slightly more often than not and still lose money, because the losses when it is wrong are far larger than the gains when it is right. Any backtest reporting hit rate rather than mean return would show this as a winner. It is the second way these numbers can lie to you, and it is quieter than the first.
If a wire writes because a price moved, then most headlines are stale by construction and the null above is an artifact of measuring commentary. That is the strongest argument against everything on this page, so it gets tested directly: sort every headline by how long the stock had been quiet before it arrived.
| gap since last headline | events | move before it | vs. normal |
|---|---|---|---|
| first mention (>72h quiet) | 549 | 146b | 0.79× |
| 24–72h | 2,047 | 192b | 1.05× |
| 4–24h | 9,501 | 226b | 1.23× |
| follow-up (<4h) | 28,451 | 238b | 1.30× |
The gradient is clean and monotonic. The less time since the last story, the more the price has already done. Reverse causality, measured rather than asserted — and it is most of the corpus: only 1.4% of headlines are first mentions, while seventy percent arrive within four hours of another.
| bucket | events | forward mean | median | share up | p vs. follow-ups |
|---|---|---|---|---|---|
| first mention | 549 | +7.9b | −5.1b | 48.5% | 0.512 |
| 24–72h | 2,047 | −2.5b | +0.1b | 50.1% | 0.463 |
| 4–24h | 9,501 | +3.5b | +1.5b | 50.4% | 0.585 |
| follow-up | 28,451 | +1.7b | −1.3b | 49.6% | — |
549 events. Edge +0.2 basis points, 95% interval
[−18.7, +20.4], p = 0.986. Not small — absent.
This matters more than another null, because it removes the best remaining explanation for all of them. The clean 1.4% behaves exactly like the polluted 98.6%.
The honest limit: 549 events can detect a 26 basis point edge at 80% power, so something between 10 and 25 basis points could still be hiding — tradable after costs, invisible at this sample size. Sharpening it means a wider universe, because the first-mention filter throws away almost everything.
An average of zero can still hide a payoff worth having. If news made enormous gains likelier without making enormous losses likelier, a portfolio of small positions would profit from the asymmetry no matter what the mean said. So: how often does each tail actually arrive?
| 3-day move | after news | baseline | ratio |
|---|---|---|---|
| > +5% | 8.92% | 7.79% | 1.15× |
| < −5% | 8.27% | 7.03% | 1.18× |
| > +10% | 2.41% | 1.92% | 1.26× |
| < −10% | 1.59% | 1.23% | 1.29× |
| > +20% | 0.43% | 0.32% | 1.35× |
| < −20% | 0.10% | 0.08% | 1.33× |
News genuinely does fatten the tails, and the effect strengthens the further out you go — from 1.15× at five percent to 1.35× at twenty. That is a real finding, and it is the first thing in this study that grows rather than vanishes under scrutiny.
It is also useless, because the downside fattens by the same amount. At every threshold the negative ratio is equal to or larger than the positive one. Measured skew after news is +0.65 against a baseline of +0.74: news events are marginally less right-leaning than ordinary days. You are not buying a lottery ticket. You are buying a wider distribution, in both directions at once.
| target / stop | entered on news | entered at random | difference | hit target |
|---|---|---|---|---|
| +5% / −2%, 3d | +16.6b | +14.3b | +2.3b | 18.7% |
| +10% / −2%, 5d | +38.6b | +32.1b | +6.5b | 8.7% |
| +20% / −2%, 10d | +67.5b | +59.1b | +8.4b | 3.6% |
| +5% / −2%, 10d | +23.5b | +24.0b | −0.6b | 30.1% |
Every one of these configurations makes money. All of them make almost exactly as much money entered on random dates. The profit is five years of rising large caps, collected by being long — not by reading anything. The news increment is 2 to 8 basis points, and in one configuration it is negative.
Without that control column, the +20% strategy reads as +67.5 bps per trade and a working system. It is drift wearing a strategy’s clothes.
And the structure itself deserves a plain look. Over ten days, a 2% stop paired with a 20% target is hit — the stop, not the target — 71.1% of the time. The target arrives 3.6% of the time. A 2% stop on a stock that routinely moves 2% in a day is not risk control; it is a near-guarantee of being removed from the position before anything can happen. Ties were resolved in the strategy’s favour nowhere: when both barriers were breached in one session, the stop was counted first.
Everything above is a null. This is not. News filed before the opening bell produces a gap, and then the price continues in the same direction through the session — the opening auction underreacts.
The effect is symmetric, which rules out the obvious artifact: positive news gaps up then drifts up a further 23.6 bps, negative news gaps down then drifts down 21.6 bps. A constant intraday bias cannot produce both.
| subset | events | edge | p |
|---|---|---|---|
| all pre-market news | 7,781 | +20.6b | <0.0001 |
| descriptive only | 3,982 | +30.2b | <0.0001 |
| excluding descriptive | 3,799 | +10.5b | 0.0006 |
| templated only | 1,053 | +13.2b | 0.0116 |
| lexicon only | 2,746 | +9.5b | 0.0093 |
| year | events | edge | p |
|---|---|---|---|
| 2021 | 531 | +12.3b | 0.157 |
| 2022 | 1,175 | +27.3b | 0.0002 |
| 2023 | 1,338 | +16.2b | 0.0030 |
| 2024 | 1,588 | +17.7b | 0.0005 |
| 2025 | 1,875 | +16.9b | 0.0002 |
| 2026 | 1,274 | +31.6b | <0.0001 |
This pipeline just independently rediscovered a documented anomaly — underreaction at the open, the same family as post-earnings-announcement drift. That is a positive control. A measuring instrument that returns nothing everywhere might simply be broken; one that returns nothing in 120 places and then finds a known effect in the 121st is working.
| assumed cost | total | annualised | max drawdown | days positive |
|---|---|---|---|---|
| 9 bps per trade | +432.3% | +37.8% | −21.8% | 56.0% |
| 20 bps per trade | +42.0% | +7.0% | −31.1% | 50.3% |
7,781 trades over 1,202 trading days. Per trade: mean +22.8 bps, median +13.6 bps, win rate 54.0%, best +25.1%, worst −19.8%.
Read the two rows against each other. Eleven basis points of assumed cost is the difference between 37.8% a year and 7.0% a year — and at 7.0% with a 31% drawdown you have taken considerably more risk than an index fund to earn less. The entire result lives inside the execution assumption, and it must be executed at the opening auction, the widest-spread moment of the trading day.
Entry at the open and exit at the close is a day trade. Below $25,000 in equity that is capped at three per five rolling days, against a median of six signals a day. Two thousand simulations of an account taking three a week at 25% position size:
| outcome over five years | result |
|---|---|
| median simulation | +7.4% |
| 5th percentile | −14.8% |
| 95th percentile | +34.9% |
| simulations that lost money | 30.8% |
| unconstrained, all signals, same costs | +42.0% |
The edge is 22.8 basis points on a 54% win rate. Nothing about a single trade is reliable; only the average over thousands is. Throttle it to three trades a week and the same positive-expectancy system loses money in 31% of simulated histories.
Which inverts the intuition this whole study started from. The value is not in reading the news well. It is in reading all of it, mechanically, and never once deciding which story matters. The moment judgment enters, the sample size that made the average real goes with it.
A null is only interesting if you can say what produced it. These four steps are ordered, because each one removes part of whatever edge the previous step left standing. By the fourth there is nothing to remove.
Newswires sell machine-readable feeds that deliver structured headlines in single-digit milliseconds to systems sitting in the same buildings as the exchange matching engines. A headline is not a secret you found; it is a broadcast, and some recipients are thousands of times closer to the market than a person reading a website.
removes: the speed advantage
Measured above: a story arriving within four hours of the last one lands on a stock that has already moved 1.30× its normal amount. Only 1.4% of headlines follow genuine quiet. The rest are reporting on a price that has already responded to whatever caused it.
removes: 98.6% of the corpus as novel information
Earnings beats and guidance raises are followed by 1.28× and 1.36× normal movement, so they plainly carry information. They are also the single most anticipated, most instrumented events in the market, with dedicated systems parsing them the instant they cross. Carrying information and offering an opportunity are different properties.
removes: the categories worth trading
A round trip costs roughly nine basis points on liquid names. Every measured edge in this study sits at or below that line, so even taking the numbers at face value there is nothing left after paying to trade them. An edge has to be not merely real but larger than the cost of collecting it.
removes: the remainder
None of this makes markets mysterious or the study a failure. It is what an efficiently priced market is supposed to look like from the outside: information that is genuinely valuable, distributed to everyone simultaneously, absorbed faster than a person can read it, leaving a residue too small to pay for its own collection.
The finding is not that news does not matter. News moves prices — that is visible throughout these results. The finding is that by the time you can act on a public headline, the moving is done.
Every event is stamped when the wire filed it, not when a reader could act. The study was handed a head start of seconds to minutes on every headline and still found nothing. That makes these results an upper bound, and the true numbers worse.
The classifier is pattern-matching plus a finance lexicon. Most copy is macro commentary that fits no template. A better classifier tests more of the corpus — though the categories it does catch are the ones with real information.
Mega-cap news is read by everyone with a faster feed. A null here says little about small caps, where coverage is thinner and attention scarcer.
Nothing here rules out an effect that lives and dies inside the first minutes after a headline. It rules out one that persists into the following days.
Event returns are measured from the first bar at or after the headline — never before — to one, two, and three days out, minus SPY over the same window. They are compared against the same stock’s unconditional return distribution over the full five years, not against zero: in a rising market, almost anything “predicts” gains. Significance is a Welch t-test with a bootstrap interval, then Benjamini–Hochberg across all 120 tests.
This is a measurement of historical data. It is not investment advice, a forecast, or a recommendation about any security.