Research findings
Do earnings-day gaps predict forward returns?
Tested whether earnings-day price gaps predict forward returns across 10,135 events in the S&P 500. After controlling for general market movement, no tradeable edge survives out-of-sample testing.
Underlying data last refreshed August 27, 2026.
What this rules out — and what it doesn't
Ruled out
- −A gap-specific momentum or reversal effect independent of general market movement
- −Any of the in-sample patterns replicating out-of-sample
- −A tradeable edge that survives realistic transaction costs
Not ruled out
- ?Sector- or size-factor-driven patterns (real estate, small-caps) — not independently tested out-of-sample
- ?A genuine effect too small or noisy to detect with 5 years of data and a SPY-only benchmark
- ?A cleaner result from a size/sector-matched benchmark instead of SPY alone
This distinction is the difference between rigorous research and a failed experiment: a null result reached by controlling for market beta and testing out-of-sample is a real finding, even though (especially because) it doesn't confirm a tradeable edge.
Methodology, in brief
Universe & data
Current S&P 500 constituents (survivorship-biased), daily OHLCV and historical earnings dates via yfinance, 2021-07-29 to 2026-08-05. 544 companies, 10,135 earnings events, 713,281 price rows.
Reaction-day detection
Earnings timestamps don't reliably say before/after market close, so the reaction day is inferred from a volume-spike heuristic, aggregated per company with a confidence score rather than trusted event-by-event.
Gap buckets & forward returns
Five configurable buckets by overnight gap size; forward returns from the post-earnings open at 1, 5, 20, and 60 trading days.
Market adjustment & validation
Excess return = raw return minus SPY's own return over the identical window. In-sample (2021-07–2023-12) patterns re-tested unmodified against a held-out out-of-sample period (2024-01–2026-07).
1.Raw returns are mostly market beta
Same gap buckets, split by whether the S&P 500 itself was up or down over the same forward window. Every bucket swings by a similar magnitude in lockstep with the market — the bucket itself explains very little.
Excess return nets out general market movement. Effect sizes shrink by roughly an order of magnitude once beta is removed — this is the metric that actually isolates a gap-specific effect, if one exists. Error bars are 95% confidence intervals.
3.Out-of-sample validation — the most important check
In-sample (2021-07 to 2023-12) vs. out-of-sample (2024-01 to 2026-07), on excess returns. Zero of the 20 bucket x horizon combinations were significant, same-signed, and significant in both periods. Error bars are 95% confidence intervals.
4.Transaction cost check
A flat 0.1% slippage-and-spread assumption per side (0.2% round trip), subtracted from every individual excess return before re-testing significance.
| Bucket | Gross return | Net of cost | p-value (net) | Survives? |
|---|---|---|---|---|
| Big gap down | -0.29% | -0.49% | 0.113 | No |
| Big gap up | 0.23% | 0.03% | 0.910 | No |
| Flat | -0.39% | -0.59% | 0.000 | No |
| Small gap down | -0.40% | -0.60% | 0.000 | No |
| Small gap up | -0.12% | -0.32% | 0.032 | No |
No bucket at this horizon survives a 0.2% round-trip cost assumption net of costs.
5.Sector & market-cap segmentation
More significant hits than pure chance would predict against SPY alone — but SPY doesn't net out sector or size exposure. Re-tested below against a matched benchmark (sector ETF, or a size-tercile ETF) with the same out-of-sample bar used everywhere else on this page.
Against SPY alone (pooled, not yet out-of-sample tested)
| Bucket | Sector | n | Mean return | p-value |
|---|---|---|---|---|
| Flat | Real Estate | 236 | -1.73% | 0.0000 |
| Small gap down | Health Care | 311 | -1.51% | 0.0011 |
| Flat | Consumer Discretionary | 184 | -1.76% | 0.0022 |
| Small gap up | Real Estate | 170 | -1.34% | 0.0026 |
| Flat | Health Care | 226 | -1.44% | 0.0040 |
| Big gap down | Consumer Discretionary | 189 | -2.02% | 0.0098 |
| Big gap up | Information Technology | 355 | 1.82% | 0.0140 |
| Small gap down | Real Estate | 124 | -1.29% | 0.0238 |
| Big gap up | Materials | 45 | -2.18% | 0.0365 |
| Small gap up | Consumer Discretionary | 232 | -1.15% | 0.0458 |
| Flat | Energy | 140 | 1.81% | 0.0462 |
Re-tested: matched benchmark + out-of-sample
Of 220 sector × bucket × horizon combinations, 1held up out-of-sample against its own sector's benchmark (same-signed and significant in both periods). Of 60 market-cap combinations, 0held up. Both are close to the ~5% false-positive rate chance alone would produce at this sample size — the SPY-flagged effects mostly don't survive a fairer test. The table below shows exactly what happened to each SPY-flagged effect from above.
| Bucket | Sector | n | vs. SPY | vs. matched benchmark | Held up OOS? |
|---|---|---|---|---|---|
| Flat | Real Estate | 236 | -1.73% (p=<0.001) | -0.58% (p=0.115) | No |
| Small gap down | Health Care | 311 | -1.51% (p=0.001) | -1.19% (p=0.004) | No |
| Flat | Consumer Discretionary | 184 | -1.76% (p=0.002) | -1.44% (p=0.025) | No |
| Small gap up | Real Estate | 170 | -1.34% (p=0.003) | -0.52% (p=0.176) | No |
| Flat | Health Care | 226 | -1.44% (p=0.004) | -0.59% (p=0.206) | No |
| Big gap down | Consumer Discretionary | 189 | -2.02% (p=0.010) | -1.37% (p=0.072) | No |
| Big gap up | Information Technology | 355 | 1.82% (p=0.014) | 1.03% (p=0.141) | No |
| Small gap down | Real Estate | 124 | -1.29% (p=0.024) | -0.75% (p=0.119) | No |
| Big gap up | Materials | 45 | -2.18% (p=0.036) | -1.83% (p=0.055) | No |
| Small gap up | Consumer Discretionary | 232 | -1.15% (p=0.046) | -0.51% (p=0.403) | No |
| Flat | Energy | 140 | 1.81% (p=0.046) | 1.00% (p=0.104) | No |
6.Does the EPS surprise itself predict anything?
Every event already carries an EPS surprise % (reported vs. estimated EPS). Tested separately from gap size: does the magnitude of the surprise correlate with the gap, or with forward excess returns?
Surprise magnitude vs. gap size
A handful of extreme outliers (surprise_pct ranges from about −10,000% to +15,900% in this dataset — almost certainly near-zero EPS estimates blowing up the percentage) can dominate a raw correlation, so three views are reported side by side rather than picking one. This one holds up: Spearman correlation is strong both in-sample (ρ=0.35) and out-of-sample (ρ=0.29) — bigger surprises produce bigger gaps, consistently.
| Period | n | Pearson r (raw) | Pearson r (winsorized) | Spearman ρ | Spearman p |
|---|---|---|---|---|---|
| Overall | 10125 | 0.023 (p=0.021) | 0.184 (p=<0.001) | 0.319 | <0.001 |
| In-sample | 5018 | 0.032 (p=0.023) | 0.220 (p=<0.001) | 0.347 | <0.001 |
| Out-of-sample | 5107 | 0.018 (p=0.202) | 0.143 (p=<0.001) | 0.297 | <0.001 |
Trend line fit on winsorized surprise % (clipping extreme outliers) across all 10,125 events: R² = 0.034, slope = 0.00034 (p < 0.001). Chart display clips |surprise| to 100% and samples up to 2,000 points for readability — the fit statistics use every event.
Surprise magnitude vs. 20-day excess return
Correlation with forward returns is much weaker than with gap size, and mostly doesn't hold up out-of-sample — consistent with the project's main finding.
| Period | n | Pearson r (raw) | Pearson r (winsorized) | Spearman ρ | Spearman p |
|---|---|---|---|---|---|
| Overall | 10124 | 0.001 (p=0.956) | 0.014 (p=0.151) | 0.027 | 0.007 |
| In-sample | 5018 | 0.000 (p=0.988) | -0.001 (p=0.952) | 0.015 | 0.274 |
| Out-of-sample | 5106 | 0.001 (p=0.944) | 0.036 (p=0.011) | 0.035 | 0.011 |
Excess return by surprise quintile, 20d
Faded bars are not statistically significant at this horizon. Zero of the 20 quintile × horizon combinations held up out-of-sample(same sign and significant in both periods) — including the most-positive-surprise quintile at 60 days, which looked like a survivor in an earlier pass but stopped holding up once survivorship bias was reduced by adding back companies removed from the S&P 500 during the study window (see Methodology). Consistent with the project's main finding either way.
7.Does a proper multifactor benchmark change anything?
SPY-only adjustment nets out broad market beta but not sector, size, value, or profitability exposure. Re-testing against a Fama-French 5-factor model — each stock's own factor exposure estimated from a trailing year of returns ending before each event, never using future data — surfaces slightly more than the SPY-only comparison.
2 of 20bucket × horizon combinations held up out-of-sample under the multifactor adjustment (vs. 0 of 20under SPY-only) — still close to the ~1-in-20 rate expected by chance at a 5% threshold, so this isn't strong evidence of a real effect, but worth naming exactly:
- Big gap up, 1 day: a small negativefactor-adjusted return, significant in both periods (−0.53% in-sample, −0.40% out-of-sample) — consistent with a short-term reversal after a large positive factor-adjusted gap, not continuation.
- Small gap down, 60 days: a larger negative factor-adjusted return, significant in both periods (−1.37% in-sample, −2.17% out-of-sample) — further underperformance, not a rebound.
Both are negative (predicting further under-performance, not a buyable edge) and both are exploratory given 2-of-20 is barely above the chance rate — flagged here for completeness, not promoted to a confirmed finding.
Caveats
- Survivorship bias — the universe is the current S&P 500 list applied retroactively; companies removed from the index during the window are absent.
- Short window, one market regime mix — 5 years, dominated by a bull run with one sharp bear year.
- SPY-only market adjustment nets out broad beta but not sector or size-factor exposure — re-tested in Finding 5 against a matched benchmark, and most segmented effects vanish, consistent with them being factor exposure rather than a gap-specific effect. The size-tercile benchmark is itself an approximation (see Finding 5), and headline Findings 1–4 still use SPY only.
- Reaction-day detection is a heuristic, not ground truth — confidence-scored, and a robustness check excluding the lowest-confidence tickers didn't change conclusions.
- Multiple comparisons — 220 sector-level and 60 market-cap-level tests were run against SPY; after a Benjamini-Hochberg FDR correction, 18 of 220 sector combinations and 12of 60 market-cap combinations remain significant (down from the raw 5% threshold count) — and per Finding 5, even those don't hold up out-of-sample against a matched benchmark.
- Daily bars only, no intraday data.