Part five · Knowing what works · Chapter 21

TESTS THAT TEACH

A test that wins without a written reason teaches you to repeat a result you don’t understand. A test that loses with one teaches you something you’ll use for years.

Most teams test. Few learn. The difference isn’t statistics; it’s paperwork. A test that started with a written hypothesis and a success metric chosen in advance produces a lesson whether it wins or loses. A test that started with “let’s try a new subject line” produces a winner at best, and nobody can say why.

One card per test

Before anything launches, fill in five lines.

  1. Because we sawThe observation that prompted the test. “Cart recovery converts at half the rate of welcome.”
  2. We believe thatThe change and who it’s for. “Sending the first cart email after one hour instead of ten minutes, for all cart abandoners.”
  3. Will moveOne metric, chosen now. “Placed order rate per recipient of the cart flow.”
  4. By at leastThe smallest change that would make you act. If you can’t name one, you don’t know why you’re testing.
  5. Measured overHow long, and on how many people, before anyone looks. Decide this before launch and don’t peek.

After the test, add two more lines: what happened, and what changed because of it. The card goes in a test log that anyone on the team can search. A template is in Appendix B.

The test log is worth more than any single winner in it.

What to test first

Test where the money is and where the effect is likely to be large. From my work, the order usually looks like this: offers and flow timing first, because they move orders directly; flow structure and segmentation next; creative angles in paid social alongside; subject lines and button colors last, because the effects are small and hard to measure since opens became unreliable.

How big, how long

Lewis and Rao’s 25 advertising experiments in chapter 8 are the warning. Individual behavior is so noisy that small tests can’t see small effects. Three practical rules follow.

Even famous results need checking

In a study published in 2000, Sheena Iyengar and Mark Lepper described a jam-tasting table in a grocery store. When it offered 24 jams, more people stopped, but only 3% of those who stopped bought. When it offered six, 30% bought Published. “Choice overload” became one of the most quoted findings in marketing. Ten years later Benjamin Scheibehenne and colleagues combined 50 experiments on the same question and found the average effect was essentially zero, with big differences between studies Published. Too much choice can hurt, under some conditions. It isn’t a law.

The lesson for your own tests is the same. One win is a hint. A result that holds when you run it again, on a different cohort, in a different month, is a finding.

A pace to hold

From my work: one new test launched every two weeks across the program, each with a card, is a pace most teams can sustain and it compounds. Twenty-five tests a year with written lessons is a playbook no competitor can copy, because it’s about your customers.

Do this

This is one chapter of The Whole Machine, which is free and readable in full on a single page with no form in front of it.