Part two · Winning the stranger · Chapter 8

WHAT THE PLATFORM SAYS YOU CAUSED

Every ad platform reports the sales it touched. Only an experiment tells you the sales it caused. Three studies show how far apart those can be.

An ad platform’s dashboard answers the question “which sales did our ads come near?” You want the answer to a different question: “which sales would not have happened without the ads?” The gap between the two is where most wasted ad money lives.

eBay turns off its search ads

In 2012 eBay’s economists ran a field experiment on their own search advertising. First they stopped buying ads on searches for the word “eBay” on MSN and Bing. Almost nothing happened. The people who would have clicked the ad clicked the free listing right underneath it instead, and organic search recovered 99.5% of the traffic Published. eBay had been paying for visits it would have received anyway.

Then they switched off non-brand search ads, the ads on searches like “used guitar,” in 68 of the 210 US TV markets for about two months, and compared those markets with the rest. The ads did bring in new and infrequent buyers. But most of the spend went to frequent eBay users who would have bought regardless. The measured return on that spending was −63%. An ordinary, non-experimental analysis of the same data had estimated +4,173% Published. That’s Blake, Nosko and Tadelis, published in Econometrica in 2015.

An observational estimate said +4,173%. The experiment said −63%. Same ads, same company.

Facebook’s own experiments

A few years later, researchers at Northwestern and Facebook took 15 large US advertising experiments run on Facebook, where the true effect was known because a random control group never saw the ads, and asked how close standard non-experimental methods would have come. In half the studies, the methods were off by a factor of three Published. In one checkout study the experiment measured a 73% lift; simply comparing people who saw the ads with people who didn’t gave 316%. The methods mostly overestimated, but not always (Gordon, Zettelmeyer, Bhargava and Chapsky, Marketing Science, 2019).

The reason is simple. The people an ad system chooses to show ads to aren’t a random sample. They’re the people most likely to buy. Compare them with everyone else and you measure the targeting, not the ad.

Why experiments are hard, too

Randall Lewis and Justin Rao looked at 25 large advertising experiments with US retailers and brokerages and found a sobering problem: individual sales are so noisy that even big experiments often can’t tell a good campaign from a break-even one Published. The median confidence interval on return was more than 100 percentage points wide. Only 3 of the 25 had enough data to show that a campaign earning a strong +50% return beat break-even. Their conclusion was that informative tests can need more than ten million person-weeks.

You don’t need ten million people to use this. You need to stop treating small differences as findings, and run tests long enough, on big enough groups, to see the effect you care about.

2021 made it worse

In April 2021 Apple released iOS 14.5, which asked iPhone users whether apps could track them across other companies’ apps and websites. Many said no. The ad platforms lost much of the data they used to connect an ad to a purchase. In February 2022 Meta’s chief financial officer estimated that Apple’s iOS changes would cost the company on the order of $10 billion in revenue that year Published. Since then, platform-reported results have leaned more heavily on modeled conversions: estimates of sales the platform believes it caused but can’t observe. The dashboard became more confident and less checkable at the same moment.

A measurement ladder you can actually climb

  1. Run the business on blended numbersMER, new-customer MER and new-customer cost from chapter 3. They can’t be double-counted, and they move when anything real changes.
  2. Hold out the cheap wins firstBrand search, retargeting and existing-customer audiences are where platforms most often take credit for sales that would have happened anyway. Pause brand search in a test region, or hold out a random share of the retargeting audience, and see what really changes.
  3. Test your biggest channel once a yearA platform conversion-lift study or a geo test, where some regions go dark and the rest continue. Write down beforehand what result would change the budget.
  4. Ask the customerA one-question post-purchase survey, “How did you first hear about us?”, is imprecise but independent of every platform. It’s the best early warning when channels you can’t track, like creators or podcasts, are doing more than the dashboard shows.
  5. Model it when spend is largeMedia mix models estimate each channel’s effect from spend and sales over time. They need a lot of history and should be checked against your experiments, not the other way round.
Run your numbers

Read a holdout test

Example test: 90% of an audience saw the ads, a random 10% didn’t. Replace with yours.
lift in purchase rate
purchases the ads caused
cost per purchase the ads caused
platform claim versus what the test found
Uses a normal approximation for the difference between two rates, with a 95% interval. It assumes people were split at random before the test started. If the holdout had fewer than about 30 purchases, treat any result as a hint, not a finding.

Do this

This is one chapter of The Whole Machine, which is free and readable in full on a single page with no form in front of it.