At the best-run testing programs in the world, somewhere between one idea in three and one in twelve improves the number it was meant to improve. Plan for that.
Every test starts with someone who believes in the idea. Ron Kohavi, who ran experimentation at Microsoft and later at Airbnb, spent two decades counting how often that belief is right. The answer is the most useful number in conversion work.
| Where | Share of tested ideas that improved their target metric |
|---|---|
| Microsoft | About one in three |
| Bing | About 10% to 20% |
| Booking.com | About 10% |
| Google (2009) | About 10% of roughly 12,000 experiments led to a change being made |
| Airbnb search | 8%: 20 of 250 ideas |
PublishedKohavi, Crook and Longbotham, 2009; Kohavi and colleagues, 2012, 2014 and 2022, the last citing Stefan Thomke for Booking.com and Jim Manzi for Google. Full references in Appendix C.
These are products tuned by thousands of engineers, so the easy wins are long gone, and a small DTC store can reasonably expect a higher hit rate on its first ideas. But the pattern holds: most ideas won’t work. At Microsoft, the two-thirds of ideas that didn’t improve their metric either made no measurable difference or made things worse. Nobody could reliably tell which ideas would work before testing them.
The clearest example comes from Bing. An engineer suggested showing more of an ad’s text in its headline. It was a small change, rated as low priority, and it sat in the backlog for more than six months. When someone finally tested it, revenue rose 12%, worth more than $100 million a year in the US Published. Kohavi and Stefan Thomke later called it the best revenue-generating idea in Bing’s history. Nobody had thought it would matter.
Your confidence in an idea tells you almost nothing about whether it will work. That’s why you test, and why the test has to be one you can believe.
This is one chapter of The Honest Test, which is free and readable in full on a single page with no form in front of it.