Study of 3,963 US Economic Releases Finds Forecast Misses Do Not Predict Market Reaction

Sixteen years of data show the largest payroll surprises moved currency prices no further than the smallest ones — and that several releases flagged “high impact” produce quieter-than-average hours

A new study of 3,963 high-impact US economic releases has found that the size of a data surprise — the gap between the figure economists forecast and the figure actually published — has no measurable relationship to how far markets move when it lands.

The research, conducted by the FxBacktest team, joined a 95,799-row economic calendar covering 2007 to 2026, carrying the actual, forecast and previous value of each release, to hourly currency, metals and equity index price data over the same period. Each release hour was measured against the average range of that same clock hour on weekdays containing no high-impact US release at all, giving every event a like-for-like baseline rather than a comparison against the trading day as a whole.

Across 180 non-farm payroll releases, sorting the results into three groups by surprise size produced average euro-dollar price ranges of 60.7, 64.6 and 65.5 pips in the release hour — a pip being the fourth decimal place in most currency quotes. Median figures were flatter still: 56.8 pips for the smallest third of surprises and 56.9 pips for the largest. The group with a median miss of 114,000 jobs moved the market one tenth of a pip further than the group with a median miss of 17,000.

“The forecast miss is the most visible number in the room on release day, so it gets credited with the move,” said Vasil K., CEO of FxBacktest. “What the data shows is that the market is repricing a scheduled moment of uncertainty, not the number itself. Positioning is cleared around a known event at roughly the same scale whether the print lands close to consensus or a long way from it.”

The study identifies a structural reason for the result. The payrolls report is not a single figure: the unemployment rate and average hourly earnings are published in the same instant, so a headline figure above forecast can arrive alongside a weak internal reading. The market’s response is a reading of the entire release rather than of the one line the consensus forecast was written against.

Direction proved symmetric as well. The 100 releases that came in above forecast averaged 62.0 pips of range; the 78 that came in below averaged 66.5.

Which releases actually move markets

The study does not conclude that scheduled data fails to move markets. It finds instead that the moving is concentrated in a small number of events, and that the calendar’s own impact ratings are a poor guide to which.

The Federal Reserve rate decision hour averaged 73.2 pips of euro-dollar range across 91 decisions — 5.21 times an ordinary hour at the same time of day, the largest multiple of any scheduled event in the sample. Non-farm payrolls averaged 63.6 pips, or 2.69 times normal, across 180 releases. Minutes of the Federal Open Market Committee came in at 2.67 times across 99 publications, and the consumer price index at 2.08 times across 82.

Below that, the multiples fall away quickly. Retail sales measured 1.59 times an ordinary hour, manufacturing survey data 1.46 times, gross domestic product 1.35 times, durable goods orders 1.33 times, producer prices 1.31 times and consumer confidence 1.27 times. Weekly unemployment claims — the most frequently published release the calendar flags as high impact — managed 1.21 times, barely a blip. Pooled across all 2,279 high-impact releases in the hourly sample, the average was 1.82 times.

“One event on the list runs above three times a normal hour, three run above two, and the median release runs 1.82,” said Vasil K. “A calendar that prints all of them in the same red typeface is describing the release, not the reaction to it.”

The same release, six different markets

The second finding has broader reach for anyone tracking more than one asset class, because the same scheduled event was found to produce very different responses depending on the instrument.

Weekly unemployment claims moved the euro-dollar rate 1.21 times a normal hour, gold 0.97 times, and the Nasdaq 100 index 0.49 times. That last figure is below one — meaning the hour containing a release flagged as high impact was, for that index, calmer than an ordinary hour at the same time of day. Manufacturing survey data showed the same pattern, running at 0.98 times on the same index.

The reverse also appears. The consumer price index moved the Nasdaq 100 considerably more than it moved the currency market — 2.68 times against 2.08 — a result consistent with an inflation reading being priced primarily as an interest-rate event and an equity index behaving as a long-duration asset. The dollar-yen exchange rate proved the most payroll-sensitive instrument in the set at 3.33 times normal, ahead of the euro at 2.69 and sterling at 2.28.

Across the pooled set of high-impact releases, the ranking by sensitivity ran dollar-yen at 1.84 times, euro-dollar at 1.82, sterling-dollar at 1.65, gold at 1.31, the S&P 500 at 1.23 and the Nasdaq 100 at 1.12.

Timing explains part of the headline figure

The research team cautions that the Federal Reserve’s 5.21 multiple is partly a statement about when the decision is published rather than about its importance relative to other events.

A 2:00 p.m. Eastern decision lands during an hour when the euro-dollar baseline range is roughly 13 to 14 pips, among the quietest of the trading day. Payrolls, published at 8:30 a.m. Eastern, arrive in an hour whose baseline runs 23 to 27 pips — already one of the busiest, and therefore with far less room to multiply. In absolute terms the two events are much closer than the ratios suggest, at 73.2 pips against 63.6.

The study publishes both readings side by side rather than choosing between them, on the grounds that they answer different questions: the absolute range describes how far price travelled, while the multiple describes how unusual the hour was relative to its own norm.

What happens after the release

A smaller section of the study, using 15-minute price data, examined whether the initial reaction persisted. On the pooled set of 327 releases, the direction established in the first 15 minutes was still intact four hours later 63.5% of the time.

The team stresses the limits of that figure. It says nothing about how far price travelled in the opposite direction in the interim, and the average displacement 60 minutes after a release — 23.2 pips — was smaller than the release bar’s own range of 29.5 pips, which the study describes as the signature of a spike that partially retraces. The finer-grained price files reach back only to approximately 2022, leaving individual event samples small: the consumer price index reading of 50.0% persistence rests on 18 observations, and is presented in the study as a sample-size caveat rather than a finding.

A data-quality defect in the source calendar

The research also documents a problem in the underlying calendar data that affects any study of this type, and which the team says is rarely disclosed by publishers of event statistics.

Nearly a quarter of the calendar rows — 24,699, or 25.8% — carried a midnight placeholder rather than an actual publication time. For releases issued at fixed US Eastern times, the study reconstructed the moment from the date under US Eastern daylight-saving rules, then validated that reconstruction against the rows that did carry a timestamp, accepting it only where it landed in the correct hour at least 95% of the time.

Non-farm payrolls validated at 100% of its 59 timed rows and producer prices at 100% of 48. Others failed and were used from timed rows only: the Philadelphia Fed index validated at just 19.1%, because its publication time moved from 10:00 to 8:30 a.m. Eastern during the sample period, and the federal funds rate at 88.1%, because the Committee published at 2:15 p.m. Eastern prior to 2013.

A second defect involved incorrect daylight-saving offsets on a minority of rows, which places a release in the neighbouring hour. Releases whose recorded hour fell outside the hours holding at least 15% of that event’s own history were discarded. Of the final 3,963 resolved releases, 3,367 came from the source clock and 596 from validated reconstruction. Releases published simultaneously, such as the several lines of an inflation report, were collapsed into a single event so they counted once rather than three times.

Stated limits

The team emphasises that price range is not a measure of profitability. The figures are drawn from one side of the market and exclude the cost of transacting, which widens sharply at precisely these moments, and they make no allowance for the difference between a quoted and an executed price during a fast move. Every figure in the study, the team notes, should be read as a ceiling on price movement rather than an estimate of what was capturable.

Historical statistics also describe the period measured and do not forecast future behaviour. The full dataset — release-hour ranges across six instruments, the payroll surprise analysis and the post-release persistence tables — is published free for reuse with attribution.