For DTC founders and operators · A field guide

THEHONESTTEST

Conversion work for stores without Amazon’s traffic, and without room to fool themselves.

Andrew LauchnerAuthor of The Second Order and The Whole MachineSeptember 2026 · 15 chapters · About an hour

A note before you start

Most DTC stores run A/B tests they can’t read. The tool shows a winner, the team ships it, and six months later the conversion rate is where it started. Nobody lied. The math was never there.

8%
of ideas tested in Airbnb’s search experiments improved the metric they targeted: 20 of 250 (Ron Kohavi)
~1,600
orders in each version of a test, at the usual settings, to detect a 10% lift in conversion rate reliably

Those two numbers are the guide in miniature. The first says that at a company with some of the best product people and data in the world, most ideas didn’t work. Most of yours won’t either, and that’s normal. The second says what it costs to find out: about 1,600 orders in each version to see a 10% lift, at any typical store conversion rate. A store taking 400 orders a month through its product pages would need eight months to run that test once.

So this is a guide for the rest of us. It shows the math the testing tools don’t lead with, and then what to do when the math says you can’t test: which research finds the leaks, which fixes don’t need a test at all, and how to change a product page with confidence when a clean experiment isn’t available.

A result you can’t read is not a result. It’s a decision you made anyway, with a chart attached.

It builds on The Whole Machine, whose chapters on the product page and on tests are a shorter version of parts one and three here.

How to read it

Start with The Testing Audit, or with the traffic table just below, which tells you in one look what your traffic can support. Or follow a path:

Four tools run in the page. Nothing you type leaves your browser.

What’s proven and what isn’t

Examples that open with Say or Picture use made-up round numbers. Every source is listed in Appendix C.

Andrew LauchnerScottsdale, Arizona
Front

TEN POSITIONS

What this guide argues, and what would prove each claim wrong.

A position says what would prove it wrong. Test each on your own store.

  1. Most ideas don’t move the number they target, even at the best companies.Wrong if more than half of your fully powered tests produced a win that held up when you tested it again.
    Most Ideas Don’t Work
  2. Traffic decides what you can test, before any idea does.Wrong if your store can detect a 5% lift on its product pages in a month.
    How Much Traffic a Test Needs
  3. A winner from an underpowered test is more likely to be noise than a winner from a well-powered one, and its lift is exaggerated either way.Wrong if your winners’ lifts hold at full size when you run them again.
    Why Winners Lie
  4. Stopping a test when it looks good multiplies false winners.Wrong if, across many A/A tests checked daily and stopped at the first significant result, no more than one in twenty finds a winner.
    Why Winners Lie
  5. Check the split before the result.Wrong if none of your past tests had a traffic split off by more than chance allows.
    Before You Believe a Result
  6. When you can’t test, research tells you what to change and a careful before-and-after tells you whether it helped.Wrong if five customer sessions on your product page find nothing you didn’t already know.
    The Ladder of Evidence
  7. Bugs, slowness and missing information get fixed, not tested.Wrong if fixing known checkout bugs has, on balance, lowered your checkout completion rate.
    Fix It, Don’t Test It
  8. Most product pages lose sales by hiding answers: the total cost, the delivery date, the returns, the reviews.Wrong if your post-purchase survey’s “what almost stopped you” answers are about something else.
    The Product Page
  9. Judge conversion changes on contribution, not conversion rate.Wrong if your free-shipping threshold or discount test raised conversion and margin every time.
    Shipping, Offers and Price Tests
  10. A test that wasn’t written down before it ran teaches nothing, even when it wins.Wrong if your team can say what it learned from its last ten tests without looking them up.
    The Brief and the Log
Front

YOUR TRAFFIC ON ONE PAGE

Find your row. It tells you the smallest improvement a four-week test on that page could reliably see, and so which kind of evidence your store can afford.

Count the visitors who reach the page you want to change, not the site’s total. Then read across.

Every chapter, on its own page

The whole book is above and always will be. These are the same chapters addressed individually, for linking to one idea rather than to ninety.

  1. TEN POSITIONSWhat this guide argues, and what would prove each claim wrong.
  2. YOUR TRAFFIC ON ONE PAGEFind your row. It tells you the smallest improvement a four-week test on that page could reliably see, and so which kind of evidence your store can afford.
  3. THE TESTING AUDITTwelve checks on whether your store can learn from changes, and whether it has been. About forty minutes with your analytics, your testing tool and a phone.
  4. MOST IDEAS DON’T WORKAt the best-run testing programs in the world, somewhere between one idea in three and one in twelve improves the number it was meant to improve. Plan for that.
  5. HOW MUCH TRAFFIC A TEST NEEDSOne formula, sixteen times a number over a number squared, tells you before you start whether a test can find what you’re looking for.
  6. WHY WINNERS LIEA statistically significant winner can still be false, and a real winner is usually smaller than it looked. Three habits make both worse.
  7. BEFORE YOU BELIEVE A RESULTCheck the split, check the size, read the range, then ask whether the story holds. In that order, every time.
  8. THE LADDER OF EVIDENCEAn A/B test is the top rung, not the only one. Pick the highest rung your traffic can reach, and use it with care.
  9. RESEARCH THAT FINDS THE LEAKFive customer sessions, one survey question, your funnel by device and a hard look at three competitors. Most stores find more in a week of this than in a year of tests.
  10. FIX IT, DON’T TEST ITBugs, slowness and hidden information don’t need an experiment. Testing them spends your scarcest resource proving what you already know.
  11. THE PRODUCT PAGEMost product pages lose sales by hiding answers. The fixes are rarely clever. They’re usually a line of text moved closer to the button.
  12. THE MOBILE GAPPhones bring most of the traffic and convert at well under the desktop rate. Some of that gap is how people shop. Some of it is your site.
  13. SHIPPING, OFFERS AND PRICE TESTSThe changes that move conversion most are the ones that cost margin. Judge them on what they leave, and test prices in a way you’d be comfortable explaining.
  14. QUIZZES, POP-UPS AND SEARCHThree features sold with big conversion numbers. Each can help. Most of the headline numbers don’t measure what the feature caused.
  15. WHAT ONE GOOD TEST LOOKS LIKEThe most famous website test in politics measured a 40.6% lift. It also produced a $60 million headline that nobody measured. Both halves are the lesson.
  16. THE BRIEF AND THE LOGA page of paperwork before each test and a line after it. That’s the difference between a team that tests and a team that learns.
  17. THE FIRST THIRTY DAYSNumbers first, then customers, then fixes, then one test worth running. Four weeks, in that order.
  18. DAY ONESix things the person who owns conversion needs on the first day.
  19. THE SHELFThe books and papers this guide leans on, and what to take from each.
  20. ABOUT THE AUTHOR
  21. FOR YOUR ANALYSTThe formulas behind the three calculators, and three queries every store should be able to run.
  22. TEMPLATESFive one-page forms. Copy them into whatever your team already uses.
  23. SOURCESEvery external source, by chapter. Web sources were read in September 2026.
Visitors to the page in 4 weeksSmallest lift a 4-week test can reliably detect, at 2% conversionAt 3% conversionWeeks to see a 10% lift, at 2%
10,00040%32%63
25,00025%20%26
50,00018%14%13
100,00013%10%7
250,0008%6%3
1,000,0004%3%1

DerivedTwo versions with traffic split evenly, a two-sided 5% significance level and 80% power, using the standard rule of thumb for sample size in chapter 3. Lifts are relative: a 20% lift takes 2.0% to 2.4%. Weeks are rounded up to whole weeks.

Now hold that against what real changes do. At Bing, one of the most tested products in the world, a change to how ad headlines were displayed raised revenue by 12%, worth over $100 million a year in the US; the book that tells the story says simple changes that big happen there “only once every few years” Published. If your row says the smallest lift you can see is 18% or more, a clean A/B test can’t reliably see most of the changes you’re likely to make. It will still show you winners. They’ll mostly be noise.

Your traffic decides what you can learn from a test before your ideas do.

What each row can do

Do this

Start here · Chapter 1

THE TESTING AUDIT

Twelve checks on whether your store can learn from changes, and whether it has been. About forty minutes with your analytics, your testing tool and a phone.

The audit isn’t about how many tests you run. It’s about whether the results you act on are true, and whether the changes you make without a test are chosen by evidence or by whoever spoke last in the meeting.

Open your analytics, your testing tool if you have one, your store’s checkout settings, and your last ten tests or site changes. Score each check 0 to 2: 0 if it failed or nobody can answer it, 1 if partly true, 2 if clean. A tool’s “95% chance to win” doesn’t count as an answer to any of these.

A dashboard’s “chance to win” is a claim, not a check.

The twelve checks

  1. You know each key page’s traffic and conversion rate · 4 minLook at: Four weeks of visitors and conversion rate for your home page, top product page and cart, by device.
    Good: Someone can produce the six numbers today, and knows which row of the traffic table each page sits in.
    Cost if wrong: You start tests that can’t finish, and believe the ones that stop early.
    Read next: How Much Traffic a Test Needs
  2. Every test has a planned size before it starts · 3 minLook at: Your last five tests.
    Good: Each had a sample size and end date written down before launch, based on the smallest lift worth acting on.
    Cost if wrong: Tests end when someone gets bored or excited, which means they stop when the noise looks best.
    Read next: How Much Traffic a Test Needs
  3. One primary metric, chosen in advance · 3 minLook at: The same five tests.
    Good: Each named one metric that would decide it, before it ran. Usually orders or contribution per visitor.
    Cost if wrong: With enough metrics, every test finds a winner somewhere.
    Read next: Why Winners Lie
  4. Nobody stops a test early because it looks good · 3 minLook at: Whether any recent test ended before its planned size.
    Good: None did, or the tool uses a method built for checking results as they arrive, and the team knows which.
    Cost if wrong: Peeking can turn a 5% false-winner rate into 25% or more.
    Read next: Why Winners Lie
  5. Winners are confirmed · 3 minLook at: What happened after your last three winners shipped.
    Good: Each was re-run, held back from a slice of traffic, or checked against the forecast after launch.
    Cost if wrong: You bank lifts that were never there.
    Read next: Why Winners Lie
  6. The split is checked on every test · 3 minLook at: Visitors in each version of your last three tests.
    Good: Someone checked the split was as planned before reading the result.
    Cost if wrong: A broken split invalidates the result, and it happens more often than teams expect.
    Read next: Before You Believe a Result
  7. An A/A test in the last year · 2 minLook at: Whether you’ve ever run two identical versions against each other on your testing tool.
    Good: Yes, and the split was even and no winner held up on a rerun.
    Cost if wrong: You’re trusting a setup nobody has checked.
    Read next: Before You Believe a Result
  8. Research in the last quarter · 4 minLook at: Recordings or notes from customer sessions, and your post-purchase survey.
    Good: You watched at least five people use your site on a phone, and a live post-purchase question asks buyers what almost stopped them.
    Cost if wrong: Your test ideas come from opinions instead of from customers.
    Read next: Research That Finds the Leak
  9. Known problems are fixed, not tested · 4 minLook at: Your list of known bugs, slow pages and errors.
    Good: There’s a list, it’s short, and nothing on it is waiting for a test.
    Cost if wrong: You spend traffic proving a broken thing is broken.
    Read next: Fix It, Don’t Test It
  10. The product page answers the four questions · 5 minLook at: Your top product page, on a phone.
    Good: Total cost, delivery date, returns and reviews visible near the add-to-cart button, without scrolling far or opening anything.
    Cost if wrong: Shoppers find the answers at checkout, and about 70% of carts are abandoned before an order.
    Read next: The Product Page
  11. Mobile has its own number and an owner · 3 minLook at: Your weekly report.
    Good: Mobile conversion rate is reported separately from desktop, and one person owns closing the gap.
    Cost if wrong: The device that brings most of your traffic gets the least attention.
    Read next: The Mobile Gap
  12. A test log · 3 minLook at: Where past tests and site changes are recorded.
    Good: One searchable log, with the hypothesis, the planned size, the result and what changed because of it, for every test and every major change.
    Cost if wrong: You’ll run the same losing test again next year.
    Read next: The Brief and the Log

Score as you go; your band appears when all twelve are in.

Run your numbers

Score the twelve checks

0: failed, or nobody can answer it. 1: partly true. 2: clean. Scores stay in this browser.
0
of 24 points
0 of 12
checks scored

Read your score

ScoreWhat it meansRead next
20–24You can trust what you learn. Your job now is learning faster: bolder tests on the pages that can carry them.What One Good Test Looks Like, then The Brief and the Log
14–19Some of what you believe is true and some isn’t, and you can’t yet tell which. Fix the zeros first.The chapter linked from your lowest check, then Why Winners Lie
8–13You’re changing the site on opinion, with a testing tool to make it look like evidence.Part one, starting at How Much Traffic a Test Needs
0–7Stop testing for a month. Fix what’s broken and talk to five customers.Fix It, Don’t Test It, then The First Thirty Days

If you’ve never run a test, checks 2 to 7 score 0 and that’s fine: they’re about running tests well. Ignore the band and look at checks 1 and 8 to 12, which apply to every store. A clean 12 on those six is a strong start.

Part one · The math nobody shows you · Chapter 2

MOST IDEAS DON’T WORK

At the best-run testing programs in the world, somewhere between one idea in three and one in twelve improves the number it was meant to improve. Plan for that.

Every test starts with someone who believes in the idea. Ron Kohavi, who ran experimentation at Microsoft and later at Airbnb, spent two decades counting how often that belief is right. The answer is the most useful number in conversion work.

The rates

WhereShare of tested ideas that improved their target metric
MicrosoftAbout one in three
BingAbout 10% to 20%
Booking.comAbout 10%
Google (2009)About 10% of roughly 12,000 experiments led to a change being made
Airbnb search8%: 20 of 250 ideas

PublishedKohavi, Crook and Longbotham, 2009; Kohavi and colleagues, 2012, 2014 and 2022, the last citing Stefan Thomke for Booking.com and Jim Manzi for Google. Full references in Appendix C.

These are products tuned by thousands of engineers, so the easy wins are long gone, and a small DTC store can reasonably expect a higher hit rate on its first ideas. But the pattern holds: most ideas won’t work. At Microsoft, the two-thirds of ideas that didn’t improve their metric either made no measurable difference or made things worse. Nobody could reliably tell which ideas would work before testing them.

Nobody can pick the winners

The clearest example comes from Bing. An engineer suggested showing more of an ad’s text in its headline. It was a small change, rated as low priority, and it sat in the backlog for more than six months. When someone finally tested it, revenue rose 12%, worth more than $100 million a year in the US Published. Kohavi and Stefan Thomke later called it the best revenue-generating idea in Bing’s history. Nobody had thought it would matter.

Your confidence in an idea tells you almost nothing about whether it will work. That’s why you test, and why the test has to be one you can believe.

What this means for a small store

Do this

Part one · The math nobody shows you · Chapter 3

HOW MUCH TRAFFIC A TEST NEEDS

One formula, sixteen times a number over a number squared, tells you before you start whether a test can find what you’re looking for.

Testing tools will happily start any test. None of them leads with the question that decides whether it’s worth starting: how many visitors does it take to see the lift you care about? The answer comes from a rule of thumb that statisticians have used for decades.

The rule

For a test with two versions and traffic split evenly, the number of visitors each version needs is about:

Visitors per version

16 × p × (1 − p) ÷ d²

p is the conversion rate now. d is the change you want to detect, in absolute terms: to see a 10% lift on a 2% conversion rate, d is 0.2 percentage points, or 0.002.

This is Lehr’s rule, and Kohavi and his co-authors give it for exactly this use: it provides a two-sided 5% significance level and 80% power, which means that if the lift is real and that size, the test will detect it four times in five. Replace the 16 with 21 for 90% power Published.

Work it through and a shortcut appears. To see a 10% lift you need about 1,600 orders in each version, and to see a 5% lift about 6,300, almost regardless of your conversion rate Derived. The statistician Martin Goodson, then at the testing company Qubit, gave the same rule of thumb in 2014: 1,600 conversions per group for a 10% lift, 6,000 for 5% Reported. So the fastest way to know if you can test a page is to count its orders. A product page that produces 400 orders a month needs about eight months to see a 10% lift.

Count the orders, not the visitors. About 1,600 per version to see a 10% lift.

Kohavi’s own summary is blunt: A/B tests are useful for effects of reasonable size when you have “at least, thousands of active users, preferably tens of thousands” Published. Compare that with what some tools will accept. VWO, a popular testing tool, says its engine needs at least 25 conversions per version and 1,500 visitors in total before it will call a result Reported. That’s enough to show a winner. It’s nowhere near enough to know if the winner is real.

Revenue needs more than orders

Conversion rate is the cheapest metric to test, because every visitor either orders or doesn’t. Revenue per visitor is far noisier, because a few large orders swing it. In Kohavi’s worked example, detecting a 5% change in revenue per user took 3.3 times as many users as detecting a 5% change in conversion rate Published. In studies of retail advertising, a customer’s sales commonly varied by ten times their average Published. So test on orders per visitor, and check that order value didn’t fall. Changes that aim straight at order value, like bundles, price points and free-shipping thresholds, need revenue or contribution as the metric, and that needs a big store. Chapter 11 has the alternatives.

Test where the traffic is

The visitors that count are the ones who see the change. A new product page layout is seen by product page visitors, not by the whole site. A checkout change is seen only by people who reach checkout, a small fraction of your traffic but with a high conversion rate, which helps. Run the numbers for the page, and prefer changes that apply across many pages at once, such as a template, where the traffic adds up.

Run your numbers

Can this page carry a test?

Example numbers. Replace with yours. Visitors means people who reach the page you’d change, both versions together.
visitors needed in each version
to run, at your traffic
smallest lift you can reliably see in 4 weeks
chance of detecting your lift if you stop at 4 weeks
extra revenue a year, if the lift is real (before costs)
Two versions, traffic split evenly, a two-sided 5% significance level, 80% power. Round the result up to whole weeks, and run whole weeks: weekday and weekend shoppers behave differently.

Look at the fourth number in the example. A test stopped at four weeks, on a page that needed twenty-one, has about a one-in-four chance of detecting a real 10% lift. Put that power into the tool in the next chapter and see what it does to the winners.

Do this

Part one · The math nobody shows you · Chapter 4

WHY WINNERS LIE

A statistically significant winner can still be false, and a real winner is usually smaller than it looked. Three habits make both worse.

“Significant at 95%” sounds like “right 95% of the time.” It isn’t. It means that if the change did nothing, a difference at least this big, in either direction, would show up by chance less than one time in twenty. How often a winner is real depends on something the p-value doesn’t know: how many of your ideas work in the first place.

False winners

Kohavi, Alex Deng and Lukas Vermeer worked it out in a 2022 paper written to correct common misreadings of A/B tests. Run 100 tests at the usual settings, 80% power and a two-sided 5% significance threshold, where one idea in ten really works. About eight of the ten real improvements show up as winners. And about two (2.25, on average) of the ninety ideas that did nothing also show up as winners, by chance. So about 22% of the winners are false Published.

Share of ideas that really workShare of winners that are false, at 80% power
33% (Microsoft)5.9%
20%11.1%
15% (Bing)15.0%
10% (Booking.com, Google, Netflix)22.0%
8% (Airbnb search)26.4%

PublishedKohavi, Deng and Vermeer, “A/B Testing Intuition Busters,” 2022. Counts a winner as a significant result in the right direction.

That’s at 80% power. Cut the power and it gets much worse, because fewer real improvements show up while the chance winners keep coming. Goodson’s 2014 illustration: in 100 tests with 10 real effects, tests stopped at two weeks with under 30% power produce about three real winners and five false ones. “63% of your winning tests are completely imaginary” Reported. (Goodson counts 5% of the ideas that did nothing as winners; on the stricter count above it would be about two.) The 2022 paper takes apart a widely shared test that claimed a 337% lift from 82 and 75 visitors. Its power to detect a 10% change was 3%. Even if one idea in three worked, a win from a test like that would be false 63% of the time Published.

Run your numbers

How many of your winners are real?

Example numbers. Replace with yours. Take power from the planner in chapter 3; a test stopped early has less than planned.
real winners, per 100 tests
false winners, per 100 tests
of your winners are false
real improvements missed, per 100 tests
A false winner here is a test of a change that did nothing, significant in the winning direction by chance: half the threshold. This assumes you don’t stop early; peeking adds more false winners on top.

Real winners shrink

Even when a win is real, the lift you measured is usually too big. A small test only reaches significance when chance happens to push the result up, so the winners it produces are the lucky draws. The 2022 paper calls it the winner’s curse: “the ‘lucky’ experimenter who finds an effect in a low power setting” is “cursed by finding an inflated effect” Published. The lower the power, the worse it gets: a significant win from a test with 50% power overstates the true lift by about 40% on average, and at 25% power it roughly doubles it Derived. Plan on real lifts being smaller than the test said, and treat a surprising one with Twyman’s law, which Kohavi quotes often: “Any figure that looks interesting or different is usually wrong” Published.

The bigger the lift a small test reports, the less you should believe it.

Three habits that make it worse

  1. PeekingChecking results as they come in and stopping when one side looks ahead. The statistician Evan Miller showed in 2010 that checking after every visitor in a small test of a change that did nothing pushed the false-positive rate from 5% to 26.1% Reported. A team at Optimizely, which sells a testing tool, simulated millions of tests comparing identical pages and found more than 57% declared a winner or loser at some point during the run Reported. Johari and colleagues found the error rate can easily rise fivefold even at 10,000 visitors Published.
  2. Many metrics, many segmentsLook at 20 metrics or 20 customer segments and, on average, one will show a “significant” difference by chance. Pick the primary metric before launch. Treat segment findings as ideas for the next test, not results.
  3. Believing the dashboard’s method fixes itSome tools use sequential statistics built for checking results as they arrive. Optimizely has since 2015, and VWO added a correction for peeking to its Bayesian engine in December 2024 Reported. That works, so find out which method your tool uses. But a Bayesian readout isn’t immune to peeking on its own: in David Robinson’s 2015 simulation of a Bayesian stopping rule, checking daily raised false acceptances from 2.5% to 11.8% Reported. And no method creates power the traffic doesn’t have.

Do this

Part one · The math nobody shows you · Chapter 5

BEFORE YOU BELIEVE A RESULT

Check the split, check the size, read the range, then ask whether the story holds. In that order, every time.

A test result is only as good as the machinery that produced it. Two checks decide whether there’s a result at all. Two more decide how much to believe it.

1. The split

If you asked for 50/50 and got 52/48 on a big test, something is broken. Aleksander Fabijan and colleagues found this problem, called a sample ratio mismatch, in about 6% of experiments at Microsoft and about 10% of analyses at LinkedIn. It “in most cases completely invalidates experiment results” Published, because whatever pushed visitors out of one version usually pushed out a particular kind of visitor. Microsoft flags a mismatch when a standard chi-square test on the counts gives a p-value below 0.0005 Published. On a store, the usual causes are a version that redirects to a new page and loses people during the extra load, bot filtering that treats the versions differently, and a script that fails to fire in one version.

2. The size

Did the test reach the size planned before launch, and run whole weeks? If not, it isn’t finished, whatever the dashboard says. Never stop early because it looks like a win, unless your tool uses a sequential method designed for it and you know that it does. Stopping early because the split is broken or a guardrail is in trouble is fine.

3. The range

Read the confidence interval, not the single number. A lift of 11% with a range from 1% to 22% says the change probably helps, by an amount you don’t know well. Plan on the low end. A range that includes zero says you can’t tell, which isn’t the same as saying there’s no difference.

4. The story

Does the result make sense, and does it hold within the test? Look at it week by week. A lift that’s big in week one and gone by week three may be novelty: Kohavi and colleagues describe returning users who “investigate the new feature, click everywhere, and thus introduce a ‘novelty’ bias that dies quickly” Published. Most store visitors are new, so novelty matters less on a store than on a product people use daily, but check anyway. Check guardrails: order value, returns, contribution. And if the result is surprising, run it again. The 2022 paper recommends a stricter threshold, 0.01 or 0.005, before believing surprising results Published.

Check the split before the result. A broken split means there is no result.

From my workOne more check belongs before the test starts: is the baseline clean? On one program, a broken product block was dragging down the numbers of the very thing we planned to redesign. Fixing the bug and testing the redesign at the same time would have credited the redesign with the fix. The rule I hold to: fix first, in both versions, then give the page two to three weeks so you know its new conversion rate before you size the test. Never let a repair ride along in only one version, and never plan a test on a baseline you haven’t measured.

Run your numbers

Read a finished test

Example numbers. Replace with yours. A is the current version, B the change.
the split
conversion rate, A then B
measured lift
95% confidence interval for the lift
p-value, two-sided
Uses a two-proportion test for the p-value, a log-ratio interval for the lift, and a chi-square test on the split; the verdict follows the interval. It assumes you didn’t stop early. For tests with more than two versions, compare each with A and use a stricter threshold.

Run an A/A test

Before you trust a testing setup, test it against itself. Split traffic between two identical versions for two weeks. You should see an even split, and usually no winner. A single A/A winner happens one time in twenty by chance, so rerun it. If it wins again, or if the split fails the check, fix the setup before running anything that matters.

Do this

Part two · When you can’t test · Chapter 6

THE LADDER OF EVIDENCE

An A/B test is the top rung, not the only one. Pick the highest rung your traffic can reach, and use it with care.

If the traffic table told you your store can’t test most changes, you still have to change things. The choice isn’t between a perfect experiment and guessing. There’s a ladder of methods between them, each cheaper and weaker than the one above.

  1. An A/B test to its planned sizeRandom split, planned size, one primary metric, read with the checks in chapter 5. The strongest evidence, and the most expensive in traffic. Use it for bold changes on your highest-traffic pages.
  2. A holdout rolloutShip the change to most visitors and keep a random slice, say 10%, on the old version for a longer period. If the change helps, only the held-back slice misses out while you learn. If it hurts, most visitors bear it, so watch the guardrails weekly. A 90/10 split needs almost three times the visitors of a 50/50 test to reach the same power, so it runs for months, not weeks. Assign the holdout by customer ID where you can, because browser cookies expire over months.
  3. On-off weeksSwitch the change on and off a whole week at a time, for six to eight weeks, deciding which week of each pair is on by a coin flip rather than strict alternation. It’s weaker than randomizing visitors: a sale or a press mention can land in one kind of week, returning shoppers see both versions, and you have only a handful of weeks to compare. Compare week by week rather than pooling visitors, and log everything else that happened. Never use it for prices.
  4. Before and after, with guardsCompare four whole weeks before the change with four after, with the guards below. The weakest rung that still deserves the name evidence.
  5. Research, then a fixFor problems that research shows plainly, like a broken button or a missing delivery date, fix them and move on. Chapter 8.

A weaker method used carefully beats a stronger one used badly.

The guards for a before-and-after

A plain before-and-after comparison fools you whenever anything else changed at the same time: the season, the traffic mix, a promotion, a competitor’s sale. The guards don’t make it an experiment. They make it harder to fool yourself.

Say you rebuild the mobile product page. Over four weeks before and four after, mobile conversion goes from 1.6% to 1.8%, a 12.5% lift. Desktop, untouched, went from 2.8% to 2.9% over the same weeks, about 3.6%. Mobile rose about 9% more than desktop (1.125 ÷ 1.036), and that’s your best estimate of what the rebuild did. It’s weaker than a test, and much better than the 12.5% you’d have reported without the comparison. But check it against noise: at 50,000 mobile visitors a month, a result like this could plausibly be anywhere from a small loss to a 25% gain. That’s what the last guard is for.

Test only bold changes

If the page gets about 50,000 visitors or more in four weeks, you can still A/B test, as long as the change is big enough to clear the smallest lift you can see. That rules out button colors and word tweaks and rules in whole new pages, a different offer, a different way of presenting price or delivery. A bold change that loses teaches you more than a timid one you can’t read.

Do this

Part two · When you can’t test · Chapter 7

RESEARCH THAT FINDS THE LEAK

Five customer sessions, one survey question, your funnel by device and a hard look at three competitors. Most stores find more in a week of this than in a year of tests.

Testing tells you whether a change worked. Research tells you what to change. Small stores need research more than big ones, because when you can only test a few things a year, the quality of your ideas decides everything.

The funnel, by device

Start with the numbers you already have. For the last four weeks, split by phone and desktop: sessions, product page views, add to cart, reached checkout, entered shipping, entered payment, ordered. Find the biggest single drop on each device. That step is where you watch people in the next section.

Watch five people

In 1993 Jakob Nielsen and Thomas Landauer published a model of how many usability problems a test finds as you add participants. Nielsen’s popular summary is that five users uncover about 85% of the problems, and that three rounds of five, fixing between rounds, beat one round of fifteen Published. It’s a model, not a guarantee: when Laura Faulkner drew random groups of five from 60 participants, some groups found 99% of the problems and others 55% Published. But the point stands. You don’t need many people to find the big problems.

Recruit five people who fit your customer and have never used your site. Give them a phone and a task, like “buy something for your sister’s birthday under $50,” and ask them to think out loud. Don’t help. Watch where they hesitate, what they tap that isn’t a link, what they look for and can’t find. Record it if they agree. The first session usually finds something the whole team missed.

You don’t need many people to find the big problems. You need to watch, and not help.

Ask the people who just bought

A one-question survey on the order confirmation page, or in the first email after purchase, finds the objections that almost cost you the sale. The conversion research firm CXL suggests surveying recent first-time buyers within days of purchase, aiming for at least 100 responses, and asking questions like “What made you almost not buy from us?” Reported. Read every answer and group them. The biggest groups are your product page’s missing answers.

Four more sources

A heuristic review, where people who know usability judge your pages against a checklist, is useful too, and Nielsen Norman Group recommends three to five reviewers working separately. It also warns that such reviews “cannot replace user research” Reported. Use it to prepare for the five sessions, not instead of them.

What comes out

One ranked list of problems, each with its evidence: how many of the five sessions hit it, how many survey answers mention it, where it sits in the funnel. Problems that are plainly broken go to chapter 8. Problems with an uncertain fix go to the ladder in chapter 6.

Do this

Part two · When you can’t test · Chapter 8

FIX IT, DON’T TEST IT

Bugs, slowness and hidden information don’t need an experiment. Testing them spends your scarcest resource proving what you already know.

Some changes are uncertain: nobody knows whether a new layout will help. Others aren’t: a button that doesn’t work on some phones is costing you orders, and nobody needs a test to know it. Treating the second kind like the first is one of the most expensive habits in conversion work.

What goes straight on the fix list

Speed, with the evidence sorted

Speed is where the evidence is most often overstated, so sort it. The strongest example is Vodafone’s: a true A/B test in Italy, with half the traffic sent to an optimized landing page. Its main content loaded 31% sooner, by the measure Google calls Largest Contentful Paint, and sales rose 8% Reported. The widely quoted Deloitte study for Google, which found that a 0.1-second improvement went with 8.4% more retail conversions, is observational: it watched speed vary naturally across 37 brands Reported. And the old line that 53% of mobile visits are abandoned after three seconds comes from 2016 Google data with thin methods; it’s too dated and loosely sourced to plan around. Speed matters. Measure your own before and after.

If five of five people trip on it, fix it. Save your traffic for questions you can’t answer by watching.

When a fix should still be tested

A few “obvious” fixes have real downsides and deserve a test or a holdout: anything that changes price or what the customer pays for shipping, anything that removes information some shoppers rely on, and anything that trades conversion for margin. Everything else, fix.

And fix before you test anything else on the same page. A test run on top of a broken baseline credits the new version with the repair. Log each fix with its date, so the before-and-after in chapter 6 has something to read.

Do this

Part three · The page · Chapter 9

THE PRODUCT PAGE

Most product pages lose sales by hiding answers. The fixes are rarely clever. They’re usually a line of text moved closer to the button.

A stranger on a product page asks the same questions in the same order: what is it, is it for me, can I trust it, what will it cost me in total, when will it arrive, and what happens if I’m wrong. Most pages answer the first at length and bury the rest in tabs, footers and the checkout.

What the benchmark finds

Baymard Institute reviews the product pages of leading ecommerce sites against its usability guidelines. In its benchmark of more than 155 sites, updated in March 2026, 52% of desktop and 62% of mobile product pages rated “mediocre or worse” Reported. Some of the gaps it counts:

GapShare of sites
Don’t show the total cost near the buy section67%
Don’t respond to negative reviews89%
Size choices not shown as buttons57%
No link to the returns policy from the product page44%
No images that show the product’s scale37%

ReportedBaymard Institute, “The current state of ecommerce product page UX,” updated March 2026.

Read the list as a set of questions left unanswered. The shopper who can’t see the total cost finds it at checkout, and extra costs are the most common reason US shoppers give for abandoning one: 40% in Baymard’s survey Reported. The shopper who can’t find the returns policy wonders what happens if the size is wrong, and leaves to think about it.

Reviews

Northwestern University’s Spiegel Research Center analyzed review data supplied by PowerReviews, a review vendor. At one high-end gift retailer, a product with five reviews was 270% more likely to be bought than one with none, and displaying reviews raised conversion 190% for lower-priced products and 380% for higher-priced ones. Across several product categories, purchase likelihood peaked for average ratings between 4.0 and 4.7 and fell as ratings approached 5.0 Reported. The data is observational and much of it comes from one store, so don’t bank the percentages. Take the pattern: reviews matter more as the price rises, and perfection looks suspicious.

Show the imperfect reviews and answer them. Perfection reads as a filter.

The first screen on a phone

On a phone, many visitors never scroll past the first screen, and Contentsquare finds nearly two-thirds of visitors who land on a product page leave without going further Reported. What belongs on that first screen, in this order:

  1. The product, in useA photo that shows scale and context, not only the product on white.
  2. Who it’s for, in one lineThe problem it solves, in the customer’s words from your reviews and survey.
  3. The review count and averageLinked to the reviews, including the critical ones.
  4. The price, and the price per use for consumables“$38, about $1.27 a serving.” If there’s a subscription option, its price and terms beside it.
  5. The delivery date and the return promise“Arrives Thursday. Free returns for 30 days.” Next to the button, not in a tab.
  6. The add-to-cart button, and an express walletBig enough for a thumb.

Everything else, the story, the ingredients, the founder’s note, goes below for the people who scroll. They’re the people who were going to read it anyway.

Fewer choices isn’t always better

Many teams cut variants because of the famous jam study, in which 30% of shoppers who stopped at a display of six jams bought, against 3% at a display of 24 Published. A 2010 analysis of 50 experiments on the same question found the average effect was “virtually zero,” with big differences between studies Published, and a later review found that too much choice hurts under particular conditions: complex options, unclear preferences, hard decisions Published. So don’t cut variants on principle. Watch your five sessions: if people struggle to choose, add a recommendation (“most people start with the 30-day size”) before you remove anything.

Do this

Part three · The page · Chapter 10

THE MOBILE GAP

Phones bring most of the traffic and convert at well under the desktop rate. Some of that gap is how people shop. Some of it is your site.

For most DTC brands, the phone is where the customer meets the store, and it’s where the store performs worst. Closing part of that gap is often the largest conversion opportunity I find in an audit.

The size of the gap

Contentsquare’s 2026 benchmark, drawn from 99 billion sessions on more than 6,500 sites in the last quarter of 2025, found phones brought 69.9% of traffic. Desktop converted at 3.4% and mobile at 2.0%; for retail sites, 3.7% against 2%. Mobile sessions lasted 2 minutes 20 seconds, desktop 4 minutes 46 Reported. Contentsquare’s customers skew toward large sites. For Shopify stores, Littledata’s 2023 benchmark of about 2,800 sites put the average conversion rate at 1.4%, with mobile at 1.2% and desktop at 1.9% Reported. Both are vendors reporting on their own customers.

Don’t aim for parity. Part of the gap is behavior: people browse on their phones in spare minutes and finish on a laptop, and some of those orders get counted as desktop. But part of it is friction, and the benchmarks say there’s a lot. Baymard’s review of more than 150 mobile sites, updated July 2026, rated 75% of them “mediocre.” 93% didn’t use adaptive error messages on forms, and 54% had no address lookup or validation Reported.

From my workA finding I see often in store audits: phones bring most of the visits and convert at a fraction of the desktop rate, while the team reviews the site on a laptop. The plan that follows puts the mobile product page and mobile checkout first, and gives mobile conversion its own line on the weekly report and its own owner.

If the team reviews the site on laptops and customers shop on phones, the team is reviewing the wrong site.

Where the thumb gives up

Measure it properly

Because some phone shoppers finish on a laptop, mobile orders understate what mobile does. Watch the mobile funnel’s steps too: add to cart, reached checkout, completed checkout. A fix that raises mobile checkout completion is working even if some of the extra orders show up under desktop.

Do this

Part three · The page · Chapter 11

SHIPPING, OFFERS AND PRICE TESTS

The changes that move conversion most are the ones that cost margin. Judge them on what they leave, and test prices in a way you’d be comfortable explaining.

Free shipping, a discount, a lower price: each will usually raise conversion. That’s what makes them dangerous in a conversion program. A test that reports conversion rate will declare each of them a winner, and the business can be worse off.

Free shipping

Extra costs are the most common reason US shoppers give for abandoning a checkout: 40% in Baymard’s survey named shipping, tax or fees, and 12% said they couldn’t see or calculate the total up front Reported. So showing the shipping cost early is a fix, not a test.

Making shipping free is a different decision. Michael Lewis, Vishal Singh and Scott Fay studied an online retailer that tried many shipping-fee schedules and found shoppers very sensitive to shipping charges, in whether they ordered and how much. Free shipping and free-shipping thresholds generated extra sales. But the lost shipping revenue, and the segments that didn’t respond, made those promotions unprofitable for that retailer Published. A 2020 study by Edlira Shehu, Dominik Papies and Scott Neslin found that free-shipping promotions led people to buy riskier products and return more of them, which also made the promotions unprofitable Published.

A conversion test on a free-shipping offer will almost always say yes. Ask the contribution question instead.

If you set a free-shipping threshold, put it a little above your typical order, show shoppers in the cart how far they are from it, and judge it on contribution per visitor, returns included. The Whole Machine has a contribution calculator, and The First Offer covers discounts and offers in depth.

Price tests

Random price tests aren’t generally illegal in the US, but the rules vary by state and country and are changing. They’re also easy to get wrong in ways customers remember. In September 2000, Amazon was found to be showing different customers randomly different discounts, 20% to 40% off, on 68 DVD titles over five days. After the backlash it refunded 6,896 customers an average of $3.10 each, and Jeff Bezos said, “We’ve never tested and we never will test prices based on customer demographics.” Amazon also said any future random price test would give buyers the lowest test price automatically Reported.

The law has moved since. New York’s Algorithmic Pricing Disclosure Act, in force since November 2025, requires a disclosure when a price is set by an algorithm using the customer’s personal data Published. Whether a random price test that uses no personal data falls under it hasn’t been settled; ask counsel if you sell into New York. At the federal level, the FTC published initial findings in January 2025 from a study of how companies use personal data to set individual prices, without declaring the practice illegal Published. One Shopify price-testing app, Intelligems, sets the store’s list price to the highest price in the test and applies test prices at checkout, so ads and shopping feeds never show a lower price than the store Reported.

Four rules keep a price test defensible:

  1. Randomize; never targetAssign prices at random, not by device, location or anything you know about the person.
  2. One price per personA visitor who comes back sees the same price.
  3. Honor the lowestBe ready to refund the difference to anyone who paid more than the lowest tested price, and consider doing it automatically when the test ends, as Amazon promised in 2000.
  4. Judge on contribution per visitorNot conversion rate. A higher price can lose orders and still win.

And be realistic about size. A price test is a revenue test, and chapter 3 showed that revenue needs several times the traffic of conversion. Most stores should test prices in bold steps, across a product line at once so the test has enough traffic, and only with visitor-level randomization that keeps each person on one price. Don’t use on-off weeks for price: returning shoppers would see the price change from week to week.

Do this

Part three · The page · Chapter 12

QUIZZES, POP-UPS AND SEARCH

Three features sold with big conversion numbers. Each can help. Most of the headline numbers don’t measure what the feature caused.

Quizzes, pop-ups and site search share a statistical trap. The people who use them were already different from the people who don’t. Compare users with non-users and the feature gets credit for the difference.

Quizzes

Quiz vendors report that quiz takers convert several times better than other visitors. The quiz vendor Octane AI’s site claims “4x higher conversion,” with no sample, period or comparison group disclosed Reported. Shoppers who take a quiz are engaged by definition. The only way to learn what a quiz adds is to show its entry points to a random half of visitors and compare the halves.

A quiz’s lasting value is often the answers, not the conversion. Store each shopper’s current answers on their profile in your email platform, so segments and emails can use them, and store each completed quiz as a dated event, so you can see answers change and trigger flows from them. Version the quiz, so a changed question doesn’t silently break every segment built on it. And drop any question you can’t name a use for. Every extra question costs completions. I’ve written the full data model up as a note: Most brands that have quiz data cannot build one segment from it.

Pop-ups

Omnisend, an email platform that sells pop-ups, reports that across 1.24 billion pop-up views in 2025, the average sign-up rate was 2.1%: 2.0% for a standard form, 2.3% for a multi-step one and 3.5% for a spin-to-win wheel Reported. A pop-up grows the list, which is worth a lot. It can also cost orders by covering the product on arrival, and Google discourages full-screen pop-ups on mobile (chapter 10).

So on phones, prefer a small banner, or a pop-up that waits until a second page view or real scrolling. Judge the pop-up on both sides of the trade: sign-ups gained, and conversion on the pages where it shows, against a random slice of visitors who never see it.

Users of a feature were different before they used it. Compare random halves, not users with non-users.

Search

“Shoppers who search convert two to three times as often” traces to blog posts from 2013 and 2015 that don’t say how it was measured Reported, and it measures intent: people who search already know what they want. The better reason to care about search is that it’s often broken. Baymard’s 2026 benchmark of more than 170 sites found mediocre-or-worse search on 46% of desktop and 58% of mobile sites. 66% failed to handle searches for things other than products, like “returns” or “shipping,” and 54% failed on abbreviations or symbols Reported.

The fix is a weekly habit. Pull the searches that returned no results, add synonyms and redirects for the common ones, and make sure the search box is visible on mobile, not hidden behind a menu.

Do this

Part four · Running a program · Chapter 13

WHAT ONE GOOD TEST LOOKS LIKE

The most famous website test in politics measured a 40.6% lift. It also produced a $60 million headline that nobody measured. Both halves are the lesson.

In December 2007, Dan Siroker, then working on Barack Obama’s presidential campaign and later a co-founder of the testing company Optimizely, ran a test on the campaign’s splash page, the page that asked visitors to sign up with their email address before entering the site. He wrote it up in 2010 on Optimizely’s blog, and the story has been told in marketing talks ever since.

What was tested

Two parts of the page changed at once. Four versions of the button text (“Sign Up,” “Learn More,” “Join Us Now” and “Sign Up Now”) were crossed with six pieces of media: three photos and three videos. That’s 24 combinations, shown to 310,382 visitors. The winner, a “Learn More” button with a family photo, raised the sign-up rate from 8.26% to 11.6%, a lift of 40.6%. And every video did worse than every photo Reported.

Why it was a good test

What it didn’t measure

The post’s title says the test raised $60 million. It didn’t measure that. Siroker estimated that about 10 million people signed up over the campaign, assumed the lift held the whole time, and calculated 2.88 million extra email addresses. He then assumed 10% of sign-ups volunteered, for 288,000 extra volunteers, and that each address was worth an average of $21 in donations, for $60 million Reported. Each step is reasonable. None was measured by the test, and the post reports no significance test or range for the lift itself. With 24 combinations, the best one also benefits from the winner’s curse in chapter 4: picking the top of 24 flatters it.

Keep the measurement and the extrapolation in separate sentences. The first is evidence. The second is a forecast.

That’s not a criticism of the test, which was well designed and clearly worth running. It’s a lesson in how results travel. A measured 40.6% lift on sign-ups became “$60 million” in the headline, and the headline is what people remember. In your own reports, write the measured number first, with its range, and put the projection after it with each assumption named.

Do this

Part four · Running a program · Chapter 14

THE BRIEF AND THE LOG

A page of paperwork before each test and a line after it. That’s the difference between a team that tests and a team that learns.

A test that won without a written reason teaches you to repeat a result you don’t understand. A test that lost with one teaches you something about your customers you’ll use for years. The brief is where the reason goes.

The brief

Fill it in before anything launches. The Whole Machine has a shorter version; this one adds the decisions that part one of this guide showed are easy to get wrong.

  1. Because we sawThe evidence from research: sessions, survey answers, the funnel. “Three of five people couldn’t find the delivery date.”
  2. We believeThe change, and who sees it.
  3. Primary metricOne. Usually orders per visitor, or contribution per visitor for anything touching price or shipping.
  4. GuardrailsWhat mustn’t get worse: order value, returns, contribution, page speed.
  5. Smallest lift worth acting onAnd the planner’s answer: visitors per version, weeks, and the power you’ll have.
  6. The rungA/B test, holdout, on-off weeks or before-and-after, from chapter 6.
  7. Decided in advanceWhat you’ll do if it wins, if it loses, and if it can’t tell. Writing the third answer down is what stops a team from rerunning the same test until it wins.

Decide what you’ll do with every possible result before you see any of them.

The log

One row per test and per major site change, in a shared sheet anyone can search: the date, the page, the brief, the rung, the planned and actual size, the result with its range, and what changed because of it. Before anyone proposes a test, they search the log. A year in, the log is the most valuable document in your conversion program, because it’s the only record of what your changes did, as opposed to what someone remembers they did.

How many tests

A page that needs four weeks per test can carry at most thirteen tests a year, one after another, and fewer once you allow for launches and sales. Tests on different pages can run at the same time, as long as each assigns visitors at random on its own and the two changes can’t clash (two tests both changing the delivery message, say). Counted against the traffic table, most DTC stores I see have room for somewhere between two and twenty real tests a year. Spend them on bold changes drawn from research, not on small ideas drawn from opinion.

From my workFor experiences a testing tool doesn’t reach, like email flows and post-purchase pages, I stamp every customer with a random digit from 0 to 9 when they’re first recorded, and never change it. Holding out digit 0 gives a standing 10% control group, with no setup per test. It measures everything you’ve shipped to the other 90% together, so read it as the combined effect of your program. To read one change on its own, give it its own random split.

Do this

Part four · Running a program · Chapter 15

THE FIRST THIRTY DAYS

Numbers first, then customers, then fixes, then one test worth running. Four weeks, in that order.

Whether you’re starting conversion work or restarting it, the order is the same: learn what your traffic can support, learn what’s wrong from the people who shop, fix what’s broken, and only then spend traffic on a test.

  1. Week one: the numbersFour weeks of visitors, conversion and orders for your top three pages, by device, placed in the traffic table (front). The funnel by device (chapter 7). Buy from your own store on a phone and time it. Add the post-purchase question. If you have a testing tool, start a two-week A/A test on your busiest page (chapter 5).
  2. Week two: the customersFive sessions on a phone, starting at your biggest funnel drop. A competitor teardown. Read the pre-purchase questions in support tickets and your 3-star reviews. End the week with one ranked problem list.
  3. Week three: the fixesShip the top fixes from the list, each logged with its date (chapter 8). The pages you fixed now need two to three weeks to settle into a new baseline.
  4. Week four: one test and the logPick one bold change for a page that can carry it and that you didn’t fix this month, or wait until a fixed page has two weeks of new baseline. Run the planner, write the brief (chapter 14) and launch once the A/A test is clean. Start the log. Re-score the audit.

At day thirty you’ll have fewer tests running than most teams, and more reason to believe the ones you have. The fixes will already be live and logged. And the ranked problem list will keep you busy for a quarter.

Research first, fixes second, tests third. Most teams run it backwards.

Do this

Close

DAY ONE

Six things the person who owns conversion needs on the first day.

Whoever owns conversion, a new hire, an agency, or you on the Monday you decide to stop guessing, needs six things on day one. Without them, the first month goes on hunting for things they should have been handed.

  1. Analytics accessWith page-level traffic, conversion and a funnel by device, and someone who can explain how it’s configured.
  2. The testing tool, if there is oneWith every past test, its settings and its results.
  3. A list of site changesEvery redesign, app install, theme change, price change and promotion of the last year, with dates. Gaps are findings.
  4. Customer wordsSupport tickets, chat logs, reviews and any survey answers, unfiltered.
  5. Contribution per orderFrom finance, so offer and shipping changes can be judged on what they leave.
  6. Authority to fixA developer’s time and the right to ship a fix without a test when research shows it’s broken.

Do this

Close

THE SHELF

The books and papers this guide leans on, and what to take from each.

And the research: Kohavi, Deng and Vermeer (2022) on false winners; Fabijan and colleagues (2019) on broken splits; Johari and colleagues (2022) on peeking; Deng, Xu, Kohavi and Walker (2013) on variance reduction; Nielsen and Landauer (1993) and Faulkner (2003) on how many users to test; Lewis, Singh and Fay (2006) and Shehu, Papies and Neslin (2020) on free shipping; Iyengar and Lepper (2000), Scheibehenne and colleagues (2010) and Chernev and colleagues (2015) on choice. Full references are in Appendix C.

Close

ABOUT THE AUTHOR

Andrew Lauchner runs Growth Legend, embedding inside consumer brands to own lifecycle, email and SMS, and revenue operations. He wrote The Second Order, on turning first-time buyers into second-time buyers; Close the Loop, on getting customers to bring the next customer; The First Offer, on the offer that wins the first order; The Whole Machine, on the fundamentals of DTC growth; and The Standing Order, on subscription programs.

As Senior Director of Growth and Retention Marketing at Gallery Furniture, he rebuilt the customer journey and the sales playbooks together. He has worked on growth and retention at Binance and 3Commas, and has been Head of Growth and Retention at Greatness Wins and at Nexus Agriscience.

The methods marked “from my work” come from auditing client stores and running their programs. Clients aren’t named and their numbers aren’t here.

What colleagues say

“Andrew led retention, lifecycle, and email/SMS, but what separates him from most in this space is how deeply he understands the role retention plays in the overall growth engine.”

Akram Khan, Head of Marketing at Gallery Furniture, senior to Andrew but didn’t manage Andrew directly

Andrew answers every note from people trying to make their stores convert, including those looking for someone to own it. Write to andrew@growthlegend.com or message him on LinkedIn.

Appendix A

FOR YOUR ANALYST

The formulas behind the three calculators, and three queries every store should be able to run.

The formulas

ForFormulaNotes
Visitors per versionn = 16 × p(1 − p) / d²p: current conversion rate. d: absolute change to detect. 95% confidence, 80% power; use 21 for 90% power.
Smallest detectable liftL = √(16(1 − p) / (p × n))Relative lift, for n visitors per version.
Power at nΦ(d × √(n / (2p(1 − p))) − 1.96)n is visitors per version. Φ is the standard normal cumulative distribution function.
False winners(1 − s)(α/2) / ((1 − s)(α/2) + s × power)s: share of ideas that work. α: two-sided threshold. Kohavi, Deng and Vermeer, 2022.
Split checkχ² = Σ (observed − expected)² / expectedOne degree of freedom for two versions. Flag p < 0.0005.
Two-version resultz = (pB − pA) / √(p̄(1 − p̄)(1/nA + 1/nB))p̄: pooled conversion rate, for the p-value. The range is for the ratio pB/pA: exp(ln(pB/pA) ± 1.96√((1 − pA)/xA + (1 − pB)/xB)) − 1, where x is orders. The verdict follows the range.

Randomize and count at the same level. If the tool assigns visitors, analyze visitors, not sessions or page views. A visitor who comes back five times is one visitor, and counting their sessions separately makes the result look more certain than it is.

The funnel by device

-- sessions reaching each step, last 28 days, by device
-- events: one row per event, with session_id, device, name, ts
SELECT device,
       COUNT(DISTINCT session_id)                                         AS sessions,
       COUNT(DISTINCT session_id) FILTER (WHERE name = 'view_item')       AS viewed_product,
       COUNT(DISTINCT session_id) FILTER (WHERE name = 'add_to_cart')     AS added_to_cart,
       COUNT(DISTINCT session_id) FILTER (WHERE name = 'begin_checkout')  AS began_checkout,
       COUNT(DISTINCT session_id) FILTER (WHERE name = 'add_shipping_info') AS entered_shipping,
       COUNT(DISTINCT session_id) FILTER (WHERE name = 'add_payment_info') AS entered_payment,
       COUNT(DISTINCT session_id) FILTER (WHERE name = 'purchase')        AS purchased
FROM events
WHERE ts >= CURRENT_DATE - INTERVAL '28 days'
GROUP BY device;

The event names follow Google Analytics 4’s recommended ecommerce events; rename them to match your setup. The syntax is Postgres; in BigQuery, write COUNT(DISTINCT IF(name = 'add_to_cart', session_id, NULL)). Divide each column by the one before to get each step’s pass rate, and compare phones with desktop step by step. Express wallets can skip the shipping and payment events, so a step above 100% means an event is missing, not a miracle.

The split check

-- visitors assigned to each version, counted once each
SELECT variant, COUNT(DISTINCT visitor_id) AS visitors
FROM exposures
WHERE experiment_id = 'pdp-delivery-date'
GROUP BY variant;

Put the two counts into the reader in chapter 5. Run the check daily in the first days of any test, not only at the end: a broken split caught on day two costs two days.

Page traffic for the traffic table

-- visitors, orders and conversion for visitors who saw a page type
SELECT page_type,
       COUNT(DISTINCT s.visitor_id) AS visitors,
       COUNT(DISTINCT o.order_id)   AS orders,
       COUNT(DISTINCT o.order_id)::numeric
         / COUNT(DISTINCT s.visitor_id) AS orders_per_visitor
FROM page_views s
LEFT JOIN orders o
  ON o.visitor_id = s.visitor_id
 AND o.ordered_at BETWEEN s.viewed_at AND s.viewed_at + INTERVAL '1 day'
WHERE s.viewed_at >= CURRENT_DATE - INTERVAL '28 days'
GROUP BY page_type
ORDER BY visitors DESC;

This counts an order toward a page if it came within a day of the visit. That’s an approximation, and a generous one, but it’s the right shape for deciding which pages can carry a test.

Appendix B

TEMPLATES

Five one-page forms. Copy them into whatever your team already uses.

The test brief

BECAUSE WE SAW     evidence from research (sessions, survey, funnel):
WE BELIEVE         the change, and who sees it:
PRIMARY METRIC     one:
GUARDRAILS         order value / returns / contribution / speed:
SMALLEST LIFT      worth acting on:        %
PLANNED SIZE       visitors per version:        weeks:        power:
RUNG               A/B / holdout / on-off weeks / before-and-after
IF IT WINS         we will:
IF IT LOSES        we will:
IF IT CAN'T TELL   we will:
OWNER              name, and the date it will be read:

The session script

BEFORE    "We're testing the site, not you. Please think out loud.
           There are no wrong answers, and I won't be able to help."
TASK 1    "You're shopping for [need]. Find something you'd buy."
TASK 2    "Buy it. Stop when you'd have to pay."
TASK 3    "Before buying, find out when it would arrive and how
           returns work."
AFTER     "What was the most confusing moment?"
          "What almost made you leave?"
NOTE      Every hesitation over five seconds, every tap on
          something that isn't a link, every question asked aloud.

The post-purchase question

ON THE CONFIRMATION PAGE, OR THE FIRST EMAIL AFTER PURCHASE
One question, free text:
  "What almost stopped you from buying today?"
Optional second question, multiple choice:
  "Where did you first hear about us?"
Read every answer weekly. Group them. Count the groups.

The test log

DATE | PAGE | CHANGE | RUNG | BRIEF LINK | PLANNED SIZE
     | ACTUAL SIZE | SPLIT CHECK | RESULT (WITH RANGE)
     | DECISION | WHAT WE LEARNED | OWNER

The before-and-after record

CHANGE              what, where, live date:
METRIC              one, decided before the change:
BEFORE WEEKS        whole weeks, no promotions:
AFTER WEEKS         same length, no promotions:
COMPARISON          what the change didn't touch (other device, similar
                    products):
TRAFFIC MIX CHECK   share by channel, before vs after:
LAST YEAR           same weeks' change last year:
RESULT              change relative to the comparison (ratio), and the
                    same figure for past periods with no change:
Appendix C

SOURCES

Every external source, by chapter. Web sources were read in September 2026.

Experimentation and statistics (chapters 2 to 5, 13, 14)

Research methods (chapters 6, 7)

The page, mobile and speed (chapters 8 to 10, 12)

Shipping and price (chapter 11)