Every ad platform reports the sales it touched. Only an experiment tells you the sales it caused. Three studies show how far apart those can be.
An ad platform’s dashboard answers the question “which sales did our ads come near?” You want the answer to a different question: “which sales would not have happened without the ads?” The gap between the two is where most wasted ad money lives.
In 2012 eBay’s economists ran a field experiment on their own search advertising. First they stopped buying ads on searches for the word “eBay” on MSN and Bing. Almost nothing happened. The people who would have clicked the ad clicked the free listing right underneath it instead, and organic search recovered 99.5% of the traffic Published. eBay had been paying for visits it would have received anyway.
Then they switched off non-brand search ads, the ads on searches like “used guitar,” in 68 of the 210 US TV markets for about two months, and compared those markets with the rest. The ads did bring in new and infrequent buyers. But most of the spend went to frequent eBay users who would have bought regardless. The measured return on that spending was −63%. An ordinary, non-experimental analysis of the same data had estimated +4,173% Published. That’s Blake, Nosko and Tadelis, published in Econometrica in 2015.
An observational estimate said +4,173%. The experiment said −63%. Same ads, same company.
A few years later, researchers at Northwestern and Facebook took 15 large US advertising experiments run on Facebook, where the true effect was known because a random control group never saw the ads, and asked how close standard non-experimental methods would have come. In half the studies, the methods were off by a factor of three Published. In one checkout study the experiment measured a 73% lift; simply comparing people who saw the ads with people who didn’t gave 316%. The methods mostly overestimated, but not always (Gordon, Zettelmeyer, Bhargava and Chapsky, Marketing Science, 2019).
The reason is simple. The people an ad system chooses to show ads to aren’t a random sample. They’re the people most likely to buy. Compare them with everyone else and you measure the targeting, not the ad.
Randall Lewis and Justin Rao looked at 25 large advertising experiments with US retailers and brokerages and found a sobering problem: individual sales are so noisy that even big experiments often can’t tell a good campaign from a break-even one Published. The median confidence interval on return was more than 100 percentage points wide. Only 3 of the 25 had enough data to show that a campaign earning a strong +50% return beat break-even. Their conclusion was that informative tests can need more than ten million person-weeks.
You don’t need ten million people to use this. You need to stop treating small differences as findings, and run tests long enough, on big enough groups, to see the effect you care about.
In April 2021 Apple released iOS 14.5, which asked iPhone users whether apps could track them across other companies’ apps and websites. Many said no. The ad platforms lost much of the data they used to connect an ad to a purchase. In February 2022 Meta’s chief financial officer estimated that Apple’s iOS changes would cost the company on the order of $10 billion in revenue that year Published. Since then, platform-reported results have leaned more heavily on modeled conversions: estimates of sales the platform believes it caused but can’t observe. The dashboard became more confident and less checkable at the same moment.
This is one chapter of The Whole Machine, which is free and readable in full on a single page with no form in front of it.