A statistically significant winner can still be false, and a real winner is usually smaller than it looked. Three habits make both worse.
“Significant at 95%” sounds like “right 95% of the time.” It isn’t. It means that if the change did nothing, a difference at least this big, in either direction, would show up by chance less than one time in twenty. How often a winner is real depends on something the p-value doesn’t know: how many of your ideas work in the first place.
Kohavi, Alex Deng and Lukas Vermeer worked it out in a 2022 paper written to correct common misreadings of A/B tests. Run 100 tests at the usual settings, 80% power and a two-sided 5% significance threshold, where one idea in ten really works. About eight of the ten real improvements show up as winners. And about two (2.25, on average) of the ninety ideas that did nothing also show up as winners, by chance. So about 22% of the winners are false Published.
| Share of ideas that really work | Share of winners that are false, at 80% power |
|---|---|
| 33% (Microsoft) | 5.9% |
| 20% | 11.1% |
| 15% (Bing) | 15.0% |
| 10% (Booking.com, Google, Netflix) | 22.0% |
| 8% (Airbnb search) | 26.4% |
PublishedKohavi, Deng and Vermeer, “A/B Testing Intuition Busters,” 2022. Counts a winner as a significant result in the right direction.
That’s at 80% power. Cut the power and it gets much worse, because fewer real improvements show up while the chance winners keep coming. Goodson’s 2014 illustration: in 100 tests with 10 real effects, tests stopped at two weeks with under 30% power produce about three real winners and five false ones. “63% of your winning tests are completely imaginary” Reported. (Goodson counts 5% of the ideas that did nothing as winners; on the stricter count above it would be about two.) The 2022 paper takes apart a widely shared test that claimed a 337% lift from 82 and 75 visitors. Its power to detect a 10% change was 3%. Even if one idea in three worked, a win from a test like that would be false 63% of the time Published.
Even when a win is real, the lift you measured is usually too big. A small test only reaches significance when chance happens to push the result up, so the winners it produces are the lucky draws. The 2022 paper calls it the winner’s curse: “the ‘lucky’ experimenter who finds an effect in a low power setting” is “cursed by finding an inflated effect” Published. The lower the power, the worse it gets: a significant win from a test with 50% power overstates the true lift by about 40% on average, and at 25% power it roughly doubles it Derived. Plan on real lifts being smaller than the test said, and treat a surprising one with Twyman’s law, which Kohavi quotes often: “Any figure that looks interesting or different is usually wrong” Published.
The bigger the lift a small test reports, the less you should believe it.
This is one chapter of The Honest Test, which is free and readable in full on a single page with no form in front of it.