Part one · The math nobody shows you · Chapter 4

WHY WINNERS LIE

A statistically significant winner can still be false, and a real winner is usually smaller than it looked. Three habits make both worse.

“Significant at 95%” sounds like “right 95% of the time.” It isn’t. It means that if the change did nothing, a difference at least this big, in either direction, would show up by chance less than one time in twenty. How often a winner is real depends on something the p-value doesn’t know: how many of your ideas work in the first place.

False winners

Kohavi, Alex Deng and Lukas Vermeer worked it out in a 2022 paper written to correct common misreadings of A/B tests. Run 100 tests at the usual settings, 80% power and a two-sided 5% significance threshold, where one idea in ten really works. About eight of the ten real improvements show up as winners. And about two (2.25, on average) of the ninety ideas that did nothing also show up as winners, by chance. So about 22% of the winners are false Published.

Share of ideas that really workShare of winners that are false, at 80% power
33% (Microsoft)5.9%
20%11.1%
15% (Bing)15.0%
10% (Booking.com, Google, Netflix)22.0%
8% (Airbnb search)26.4%

PublishedKohavi, Deng and Vermeer, “A/B Testing Intuition Busters,” 2022. Counts a winner as a significant result in the right direction.

That’s at 80% power. Cut the power and it gets much worse, because fewer real improvements show up while the chance winners keep coming. Goodson’s 2014 illustration: in 100 tests with 10 real effects, tests stopped at two weeks with under 30% power produce about three real winners and five false ones. “63% of your winning tests are completely imaginary” Reported. (Goodson counts 5% of the ideas that did nothing as winners; on the stricter count above it would be about two.) The 2022 paper takes apart a widely shared test that claimed a 337% lift from 82 and 75 visitors. Its power to detect a 10% change was 3%. Even if one idea in three worked, a win from a test like that would be false 63% of the time Published.

Run your numbers

How many of your winners are real?

Example numbers. Replace with yours. Take power from the planner in chapter 3; a test stopped early has less than planned.
real winners, per 100 tests
false winners, per 100 tests
of your winners are false
real improvements missed, per 100 tests
A false winner here is a test of a change that did nothing, significant in the winning direction by chance: half the threshold. This assumes you don’t stop early; peeking adds more false winners on top.

Real winners shrink

Even when a win is real, the lift you measured is usually too big. A small test only reaches significance when chance happens to push the result up, so the winners it produces are the lucky draws. The 2022 paper calls it the winner’s curse: “the ‘lucky’ experimenter who finds an effect in a low power setting” is “cursed by finding an inflated effect” Published. The lower the power, the worse it gets: a significant win from a test with 50% power overstates the true lift by about 40% on average, and at 25% power it roughly doubles it Derived. Plan on real lifts being smaller than the test said, and treat a surprising one with Twyman’s law, which Kohavi quotes often: “Any figure that looks interesting or different is usually wrong” Published.

The bigger the lift a small test reports, the less you should believe it.

Three habits that make it worse

  1. PeekingChecking results as they come in and stopping when one side looks ahead. The statistician Evan Miller showed in 2010 that checking after every visitor in a small test of a change that did nothing pushed the false-positive rate from 5% to 26.1% Reported. A team at Optimizely, which sells a testing tool, simulated millions of tests comparing identical pages and found more than 57% declared a winner or loser at some point during the run Reported. Johari and colleagues found the error rate can easily rise fivefold even at 10,000 visitors Published.
  2. Many metrics, many segmentsLook at 20 metrics or 20 customer segments and, on average, one will show a “significant” difference by chance. Pick the primary metric before launch. Treat segment findings as ideas for the next test, not results.
  3. Believing the dashboard’s method fixes itSome tools use sequential statistics built for checking results as they arrive. Optimizely has since 2015, and VWO added a correction for peeking to its Bayesian engine in December 2024 Reported. That works, so find out which method your tool uses. But a Bayesian readout isn’t immune to peeking on its own: in David Robinson’s 2015 simulation of a Bayesian stopping rule, checking daily raised false acceptances from 2.5% to 11.8% Reported. And no method creates power the traffic doesn’t have.

Do this

This is one chapter of The Honest Test, which is free and readable in full on a single page with no form in front of it.