Conversion work for stores without Amazon’s traffic, and without room to fool themselves.
Most DTC stores run A/B tests they can’t read. The tool shows a winner, the team ships it, and six months later the conversion rate is where it started. Nobody lied. The math was never there.
Those two numbers are the guide in miniature. The first says that at a company with some of the best product people and data in the world, most ideas didn’t work. Most of yours won’t either, and that’s normal. The second says what it costs to find out: about 1,600 orders in each version to see a 10% lift, at any typical store conversion rate. A store taking 400 orders a month through its product pages would need eight months to run that test once.
So this is a guide for the rest of us. It shows the math the testing tools don’t lead with, and then what to do when the math says you can’t test: which research finds the leaks, which fixes don’t need a test at all, and how to change a product page with confidence when a clean experiment isn’t available.
A result you can’t read is not a result. It’s a decision you made anyway, with a chart attached.
It builds on The Whole Machine, whose chapters on the product page and on tests are a shorter version of parts one and three here.
Start with The Testing Audit, or with the traffic table just below, which tells you in one look what your traffic can support. Or follow a path:
Four tools run in the page. Nothing you type leaves your browser.
Examples that open with Say or Picture use made-up round numbers. Every source is listed in Appendix C.
What this guide argues, and what would prove each claim wrong.
A position says what would prove it wrong. Test each on your own store.
Find your row. It tells you the smallest improvement a four-week test on that page could reliably see, and so which kind of evidence your store can afford.
Count the visitors who reach the page you want to change, not the site’s total. Then read across.
The whole book is above and always will be. These are the same chapters addressed individually, for linking to one idea rather than to ninety.
| Visitors to the page in 4 weeks | Smallest lift a 4-week test can reliably detect, at 2% conversion | At 3% conversion | Weeks to see a 10% lift, at 2% |
|---|---|---|---|
| 10,000 | 40% | 32% | 63 |
| 25,000 | 25% | 20% | 26 |
| 50,000 | 18% | 14% | 13 |
| 100,000 | 13% | 10% | 7 |
| 250,000 | 8% | 6% | 3 |
| 1,000,000 | 4% | 3% | 1 |
DerivedTwo versions with traffic split evenly, a two-sided 5% significance level and 80% power, using the standard rule of thumb for sample size in chapter 3. Lifts are relative: a 20% lift takes 2.0% to 2.4%. Weeks are rounded up to whole weeks.
Now hold that against what real changes do. At Bing, one of the most tested products in the world, a change to how ad headlines were displayed raised revenue by 12%, worth over $100 million a year in the US; the book that tells the story says simple changes that big happen there “only once every few years” Published. If your row says the smallest lift you can see is 18% or more, a clean A/B test can’t reliably see most of the changes you’re likely to make. It will still show you winners. They’ll mostly be noise.
Your traffic decides what you can learn from a test before your ideas do.
Twelve checks on whether your store can learn from changes, and whether it has been. About forty minutes with your analytics, your testing tool and a phone.
The audit isn’t about how many tests you run. It’s about whether the results you act on are true, and whether the changes you make without a test are chosen by evidence or by whoever spoke last in the meeting.
Open your analytics, your testing tool if you have one, your store’s checkout settings, and your last ten tests or site changes. Score each check 0 to 2: 0 if it failed or nobody can answer it, 1 if partly true, 2 if clean. A tool’s “95% chance to win” doesn’t count as an answer to any of these.
A dashboard’s “chance to win” is a claim, not a check.
Score as you go; your band appears when all twelve are in.
| Score | What it means | Read next |
|---|---|---|
| 20–24 | You can trust what you learn. Your job now is learning faster: bolder tests on the pages that can carry them. | What One Good Test Looks Like, then The Brief and the Log |
| 14–19 | Some of what you believe is true and some isn’t, and you can’t yet tell which. Fix the zeros first. | The chapter linked from your lowest check, then Why Winners Lie |
| 8–13 | You’re changing the site on opinion, with a testing tool to make it look like evidence. | Part one, starting at How Much Traffic a Test Needs |
| 0–7 | Stop testing for a month. Fix what’s broken and talk to five customers. | Fix It, Don’t Test It, then The First Thirty Days |
If you’ve never run a test, checks 2 to 7 score 0 and that’s fine: they’re about running tests well. Ignore the band and look at checks 1 and 8 to 12, which apply to every store. A clean 12 on those six is a strong start.
At the best-run testing programs in the world, somewhere between one idea in three and one in twelve improves the number it was meant to improve. Plan for that.
Every test starts with someone who believes in the idea. Ron Kohavi, who ran experimentation at Microsoft and later at Airbnb, spent two decades counting how often that belief is right. The answer is the most useful number in conversion work.
| Where | Share of tested ideas that improved their target metric |
|---|---|
| Microsoft | About one in three |
| Bing | About 10% to 20% |
| Booking.com | About 10% |
| Google (2009) | About 10% of roughly 12,000 experiments led to a change being made |
| Airbnb search | 8%: 20 of 250 ideas |
PublishedKohavi, Crook and Longbotham, 2009; Kohavi and colleagues, 2012, 2014 and 2022, the last citing Stefan Thomke for Booking.com and Jim Manzi for Google. Full references in Appendix C.
These are products tuned by thousands of engineers, so the easy wins are long gone, and a small DTC store can reasonably expect a higher hit rate on its first ideas. But the pattern holds: most ideas won’t work. At Microsoft, the two-thirds of ideas that didn’t improve their metric either made no measurable difference or made things worse. Nobody could reliably tell which ideas would work before testing them.
The clearest example comes from Bing. An engineer suggested showing more of an ad’s text in its headline. It was a small change, rated as low priority, and it sat in the backlog for more than six months. When someone finally tested it, revenue rose 12%, worth more than $100 million a year in the US Published. Kohavi and Stefan Thomke later called it the best revenue-generating idea in Bing’s history. Nobody had thought it would matter.
Your confidence in an idea tells you almost nothing about whether it will work. That’s why you test, and why the test has to be one you can believe.
One formula, sixteen times a number over a number squared, tells you before you start whether a test can find what you’re looking for.
Testing tools will happily start any test. None of them leads with the question that decides whether it’s worth starting: how many visitors does it take to see the lift you care about? The answer comes from a rule of thumb that statisticians have used for decades.
For a test with two versions and traffic split evenly, the number of visitors each version needs is about:
16 × p × (1 − p) ÷ d²
p is the conversion rate now. d is the change you want to detect, in absolute terms: to see a 10% lift on a 2% conversion rate, d is 0.2 percentage points, or 0.002.
This is Lehr’s rule, and Kohavi and his co-authors give it for exactly this use: it provides a two-sided 5% significance level and 80% power, which means that if the lift is real and that size, the test will detect it four times in five. Replace the 16 with 21 for 90% power Published.
Work it through and a shortcut appears. To see a 10% lift you need about 1,600 orders in each version, and to see a 5% lift about 6,300, almost regardless of your conversion rate Derived. The statistician Martin Goodson, then at the testing company Qubit, gave the same rule of thumb in 2014: 1,600 conversions per group for a 10% lift, 6,000 for 5% Reported. So the fastest way to know if you can test a page is to count its orders. A product page that produces 400 orders a month needs about eight months to see a 10% lift.
Count the orders, not the visitors. About 1,600 per version to see a 10% lift.
Kohavi’s own summary is blunt: A/B tests are useful for effects of reasonable size when you have “at least, thousands of active users, preferably tens of thousands” Published. Compare that with what some tools will accept. VWO, a popular testing tool, says its engine needs at least 25 conversions per version and 1,500 visitors in total before it will call a result Reported. That’s enough to show a winner. It’s nowhere near enough to know if the winner is real.
Conversion rate is the cheapest metric to test, because every visitor either orders or doesn’t. Revenue per visitor is far noisier, because a few large orders swing it. In Kohavi’s worked example, detecting a 5% change in revenue per user took 3.3 times as many users as detecting a 5% change in conversion rate Published. In studies of retail advertising, a customer’s sales commonly varied by ten times their average Published. So test on orders per visitor, and check that order value didn’t fall. Changes that aim straight at order value, like bundles, price points and free-shipping thresholds, need revenue or contribution as the metric, and that needs a big store. Chapter 11 has the alternatives.
The visitors that count are the ones who see the change. A new product page layout is seen by product page visitors, not by the whole site. A checkout change is seen only by people who reach checkout, a small fraction of your traffic but with a high conversion rate, which helps. Run the numbers for the page, and prefer changes that apply across many pages at once, such as a template, where the traffic adds up.
Look at the fourth number in the example. A test stopped at four weeks, on a page that needed twenty-one, has about a one-in-four chance of detecting a real 10% lift. Put that power into the tool in the next chapter and see what it does to the winners.
A statistically significant winner can still be false, and a real winner is usually smaller than it looked. Three habits make both worse.
“Significant at 95%” sounds like “right 95% of the time.” It isn’t. It means that if the change did nothing, a difference at least this big, in either direction, would show up by chance less than one time in twenty. How often a winner is real depends on something the p-value doesn’t know: how many of your ideas work in the first place.
Kohavi, Alex Deng and Lukas Vermeer worked it out in a 2022 paper written to correct common misreadings of A/B tests. Run 100 tests at the usual settings, 80% power and a two-sided 5% significance threshold, where one idea in ten really works. About eight of the ten real improvements show up as winners. And about two (2.25, on average) of the ninety ideas that did nothing also show up as winners, by chance. So about 22% of the winners are false Published.
| Share of ideas that really work | Share of winners that are false, at 80% power |
|---|---|
| 33% (Microsoft) | 5.9% |
| 20% | 11.1% |
| 15% (Bing) | 15.0% |
| 10% (Booking.com, Google, Netflix) | 22.0% |
| 8% (Airbnb search) | 26.4% |
PublishedKohavi, Deng and Vermeer, “A/B Testing Intuition Busters,” 2022. Counts a winner as a significant result in the right direction.
That’s at 80% power. Cut the power and it gets much worse, because fewer real improvements show up while the chance winners keep coming. Goodson’s 2014 illustration: in 100 tests with 10 real effects, tests stopped at two weeks with under 30% power produce about three real winners and five false ones. “63% of your winning tests are completely imaginary” Reported. (Goodson counts 5% of the ideas that did nothing as winners; on the stricter count above it would be about two.) The 2022 paper takes apart a widely shared test that claimed a 337% lift from 82 and 75 visitors. Its power to detect a 10% change was 3%. Even if one idea in three worked, a win from a test like that would be false 63% of the time Published.
Even when a win is real, the lift you measured is usually too big. A small test only reaches significance when chance happens to push the result up, so the winners it produces are the lucky draws. The 2022 paper calls it the winner’s curse: “the ‘lucky’ experimenter who finds an effect in a low power setting” is “cursed by finding an inflated effect” Published. The lower the power, the worse it gets: a significant win from a test with 50% power overstates the true lift by about 40% on average, and at 25% power it roughly doubles it Derived. Plan on real lifts being smaller than the test said, and treat a surprising one with Twyman’s law, which Kohavi quotes often: “Any figure that looks interesting or different is usually wrong” Published.
The bigger the lift a small test reports, the less you should believe it.
Check the split, check the size, read the range, then ask whether the story holds. In that order, every time.
A test result is only as good as the machinery that produced it. Two checks decide whether there’s a result at all. Two more decide how much to believe it.
If you asked for 50/50 and got 52/48 on a big test, something is broken. Aleksander Fabijan and colleagues found this problem, called a sample ratio mismatch, in about 6% of experiments at Microsoft and about 10% of analyses at LinkedIn. It “in most cases completely invalidates experiment results” Published, because whatever pushed visitors out of one version usually pushed out a particular kind of visitor. Microsoft flags a mismatch when a standard chi-square test on the counts gives a p-value below 0.0005 Published. On a store, the usual causes are a version that redirects to a new page and loses people during the extra load, bot filtering that treats the versions differently, and a script that fails to fire in one version.
Did the test reach the size planned before launch, and run whole weeks? If not, it isn’t finished, whatever the dashboard says. Never stop early because it looks like a win, unless your tool uses a sequential method designed for it and you know that it does. Stopping early because the split is broken or a guardrail is in trouble is fine.
Read the confidence interval, not the single number. A lift of 11% with a range from 1% to 22% says the change probably helps, by an amount you don’t know well. Plan on the low end. A range that includes zero says you can’t tell, which isn’t the same as saying there’s no difference.
Does the result make sense, and does it hold within the test? Look at it week by week. A lift that’s big in week one and gone by week three may be novelty: Kohavi and colleagues describe returning users who “investigate the new feature, click everywhere, and thus introduce a ‘novelty’ bias that dies quickly” Published. Most store visitors are new, so novelty matters less on a store than on a product people use daily, but check anyway. Check guardrails: order value, returns, contribution. And if the result is surprising, run it again. The 2022 paper recommends a stricter threshold, 0.01 or 0.005, before believing surprising results Published.
Check the split before the result. A broken split means there is no result.
From my workOne more check belongs before the test starts: is the baseline clean? On one program, a broken product block was dragging down the numbers of the very thing we planned to redesign. Fixing the bug and testing the redesign at the same time would have credited the redesign with the fix. The rule I hold to: fix first, in both versions, then give the page two to three weeks so you know its new conversion rate before you size the test. Never let a repair ride along in only one version, and never plan a test on a baseline you haven’t measured.
Before you trust a testing setup, test it against itself. Split traffic between two identical versions for two weeks. You should see an even split, and usually no winner. A single A/A winner happens one time in twenty by chance, so rerun it. If it wins again, or if the split fails the check, fix the setup before running anything that matters.
An A/B test is the top rung, not the only one. Pick the highest rung your traffic can reach, and use it with care.
If the traffic table told you your store can’t test most changes, you still have to change things. The choice isn’t between a perfect experiment and guessing. There’s a ladder of methods between them, each cheaper and weaker than the one above.
A weaker method used carefully beats a stronger one used badly.
A plain before-and-after comparison fools you whenever anything else changed at the same time: the season, the traffic mix, a promotion, a competitor’s sale. The guards don’t make it an experiment. They make it harder to fool yourself.
Say you rebuild the mobile product page. Over four weeks before and four after, mobile conversion goes from 1.6% to 1.8%, a 12.5% lift. Desktop, untouched, went from 2.8% to 2.9% over the same weeks, about 3.6%. Mobile rose about 9% more than desktop (1.125 ÷ 1.036), and that’s your best estimate of what the rebuild did. It’s weaker than a test, and much better than the 12.5% you’d have reported without the comparison. But check it against noise: at 50,000 mobile visitors a month, a result like this could plausibly be anywhere from a small loss to a 25% gain. That’s what the last guard is for.
If the page gets about 50,000 visitors or more in four weeks, you can still A/B test, as long as the change is big enough to clear the smallest lift you can see. That rules out button colors and word tweaks and rules in whole new pages, a different offer, a different way of presenting price or delivery. A bold change that loses teaches you more than a timid one you can’t read.
Five customer sessions, one survey question, your funnel by device and a hard look at three competitors. Most stores find more in a week of this than in a year of tests.
Testing tells you whether a change worked. Research tells you what to change. Small stores need research more than big ones, because when you can only test a few things a year, the quality of your ideas decides everything.
Start with the numbers you already have. For the last four weeks, split by phone and desktop: sessions, product page views, add to cart, reached checkout, entered shipping, entered payment, ordered. Find the biggest single drop on each device. That step is where you watch people in the next section.
In 1993 Jakob Nielsen and Thomas Landauer published a model of how many usability problems a test finds as you add participants. Nielsen’s popular summary is that five users uncover about 85% of the problems, and that three rounds of five, fixing between rounds, beat one round of fifteen Published. It’s a model, not a guarantee: when Laura Faulkner drew random groups of five from 60 participants, some groups found 99% of the problems and others 55% Published. But the point stands. You don’t need many people to find the big problems.
Recruit five people who fit your customer and have never used your site. Give them a phone and a task, like “buy something for your sister’s birthday under $50,” and ask them to think out loud. Don’t help. Watch where they hesitate, what they tap that isn’t a link, what they look for and can’t find. Record it if they agree. The first session usually finds something the whole team missed.
You don’t need many people to find the big problems. You need to watch, and not help.
A one-question survey on the order confirmation page, or in the first email after purchase, finds the objections that almost cost you the sale. The conversion research firm CXL suggests surveying recent first-time buyers within days of purchase, aiming for at least 100 responses, and asking questions like “What made you almost not buy from us?” Reported. Read every answer and group them. The biggest groups are your product page’s missing answers.
A heuristic review, where people who know usability judge your pages against a checklist, is useful too, and Nielsen Norman Group recommends three to five reviewers working separately. It also warns that such reviews “cannot replace user research” Reported. Use it to prepare for the five sessions, not instead of them.
One ranked list of problems, each with its evidence: how many of the five sessions hit it, how many survey answers mention it, where it sits in the funnel. Problems that are plainly broken go to chapter 8. Problems with an uncertain fix go to the ladder in chapter 6.
Bugs, slowness and hidden information don’t need an experiment. Testing them spends your scarcest resource proving what you already know.
Some changes are uncertain: nobody knows whether a new layout will help. Others aren’t: a button that doesn’t work on some phones is costing you orders, and nobody needs a test to know it. Treating the second kind like the first is one of the most expensive habits in conversion work.
Speed is where the evidence is most often overstated, so sort it. The strongest example is Vodafone’s: a true A/B test in Italy, with half the traffic sent to an optimized landing page. Its main content loaded 31% sooner, by the measure Google calls Largest Contentful Paint, and sales rose 8% Reported. The widely quoted Deloitte study for Google, which found that a 0.1-second improvement went with 8.4% more retail conversions, is observational: it watched speed vary naturally across 37 brands Reported. And the old line that 53% of mobile visits are abandoned after three seconds comes from 2016 Google data with thin methods; it’s too dated and loosely sourced to plan around. Speed matters. Measure your own before and after.
If five of five people trip on it, fix it. Save your traffic for questions you can’t answer by watching.
A few “obvious” fixes have real downsides and deserve a test or a holdout: anything that changes price or what the customer pays for shipping, anything that removes information some shoppers rely on, and anything that trades conversion for margin. Everything else, fix.
And fix before you test anything else on the same page. A test run on top of a broken baseline credits the new version with the repair. Log each fix with its date, so the before-and-after in chapter 6 has something to read.
Most product pages lose sales by hiding answers. The fixes are rarely clever. They’re usually a line of text moved closer to the button.
A stranger on a product page asks the same questions in the same order: what is it, is it for me, can I trust it, what will it cost me in total, when will it arrive, and what happens if I’m wrong. Most pages answer the first at length and bury the rest in tabs, footers and the checkout.
Baymard Institute reviews the product pages of leading ecommerce sites against its usability guidelines. In its benchmark of more than 155 sites, updated in March 2026, 52% of desktop and 62% of mobile product pages rated “mediocre or worse” Reported. Some of the gaps it counts:
| Gap | Share of sites |
|---|---|
| Don’t show the total cost near the buy section | 67% |
| Don’t respond to negative reviews | 89% |
| Size choices not shown as buttons | 57% |
| No link to the returns policy from the product page | 44% |
| No images that show the product’s scale | 37% |
ReportedBaymard Institute, “The current state of ecommerce product page UX,” updated March 2026.
Read the list as a set of questions left unanswered. The shopper who can’t see the total cost finds it at checkout, and extra costs are the most common reason US shoppers give for abandoning one: 40% in Baymard’s survey Reported. The shopper who can’t find the returns policy wonders what happens if the size is wrong, and leaves to think about it.
Northwestern University’s Spiegel Research Center analyzed review data supplied by PowerReviews, a review vendor. At one high-end gift retailer, a product with five reviews was 270% more likely to be bought than one with none, and displaying reviews raised conversion 190% for lower-priced products and 380% for higher-priced ones. Across several product categories, purchase likelihood peaked for average ratings between 4.0 and 4.7 and fell as ratings approached 5.0 Reported. The data is observational and much of it comes from one store, so don’t bank the percentages. Take the pattern: reviews matter more as the price rises, and perfection looks suspicious.
Show the imperfect reviews and answer them. Perfection reads as a filter.
On a phone, many visitors never scroll past the first screen, and Contentsquare finds nearly two-thirds of visitors who land on a product page leave without going further Reported. What belongs on that first screen, in this order:
Everything else, the story, the ingredients, the founder’s note, goes below for the people who scroll. They’re the people who were going to read it anyway.
Many teams cut variants because of the famous jam study, in which 30% of shoppers who stopped at a display of six jams bought, against 3% at a display of 24 Published. A 2010 analysis of 50 experiments on the same question found the average effect was “virtually zero,” with big differences between studies Published, and a later review found that too much choice hurts under particular conditions: complex options, unclear preferences, hard decisions Published. So don’t cut variants on principle. Watch your five sessions: if people struggle to choose, add a recommendation (“most people start with the 30-day size”) before you remove anything.
Phones bring most of the traffic and convert at well under the desktop rate. Some of that gap is how people shop. Some of it is your site.
For most DTC brands, the phone is where the customer meets the store, and it’s where the store performs worst. Closing part of that gap is often the largest conversion opportunity I find in an audit.
Contentsquare’s 2026 benchmark, drawn from 99 billion sessions on more than 6,500 sites in the last quarter of 2025, found phones brought 69.9% of traffic. Desktop converted at 3.4% and mobile at 2.0%; for retail sites, 3.7% against 2%. Mobile sessions lasted 2 minutes 20 seconds, desktop 4 minutes 46 Reported. Contentsquare’s customers skew toward large sites. For Shopify stores, Littledata’s 2023 benchmark of about 2,800 sites put the average conversion rate at 1.4%, with mobile at 1.2% and desktop at 1.9% Reported. Both are vendors reporting on their own customers.
Don’t aim for parity. Part of the gap is behavior: people browse on their phones in spare minutes and finish on a laptop, and some of those orders get counted as desktop. But part of it is friction, and the benchmarks say there’s a lot. Baymard’s review of more than 150 mobile sites, updated July 2026, rated 75% of them “mediocre.” 93% didn’t use adaptive error messages on forms, and 54% had no address lookup or validation Reported.
From my workA finding I see often in store audits: phones bring most of the visits and convert at a fraction of the desktop rate, while the team reviews the site on a laptop. The plan that follows puts the mobile product page and mobile checkout first, and gives mobile conversion its own line on the weekly report and its own owner.
If the team reviews the site on laptops and customers shop on phones, the team is reviewing the wrong site.
Because some phone shoppers finish on a laptop, mobile orders understate what mobile does. Watch the mobile funnel’s steps too: add to cart, reached checkout, completed checkout. A fix that raises mobile checkout completion is working even if some of the extra orders show up under desktop.
The changes that move conversion most are the ones that cost margin. Judge them on what they leave, and test prices in a way you’d be comfortable explaining.
Free shipping, a discount, a lower price: each will usually raise conversion. That’s what makes them dangerous in a conversion program. A test that reports conversion rate will declare each of them a winner, and the business can be worse off.
Extra costs are the most common reason US shoppers give for abandoning a checkout: 40% in Baymard’s survey named shipping, tax or fees, and 12% said they couldn’t see or calculate the total up front Reported. So showing the shipping cost early is a fix, not a test.
Making shipping free is a different decision. Michael Lewis, Vishal Singh and Scott Fay studied an online retailer that tried many shipping-fee schedules and found shoppers very sensitive to shipping charges, in whether they ordered and how much. Free shipping and free-shipping thresholds generated extra sales. But the lost shipping revenue, and the segments that didn’t respond, made those promotions unprofitable for that retailer Published. A 2020 study by Edlira Shehu, Dominik Papies and Scott Neslin found that free-shipping promotions led people to buy riskier products and return more of them, which also made the promotions unprofitable Published.
A conversion test on a free-shipping offer will almost always say yes. Ask the contribution question instead.
If you set a free-shipping threshold, put it a little above your typical order, show shoppers in the cart how far they are from it, and judge it on contribution per visitor, returns included. The Whole Machine has a contribution calculator, and The First Offer covers discounts and offers in depth.
Random price tests aren’t generally illegal in the US, but the rules vary by state and country and are changing. They’re also easy to get wrong in ways customers remember. In September 2000, Amazon was found to be showing different customers randomly different discounts, 20% to 40% off, on 68 DVD titles over five days. After the backlash it refunded 6,896 customers an average of $3.10 each, and Jeff Bezos said, “We’ve never tested and we never will test prices based on customer demographics.” Amazon also said any future random price test would give buyers the lowest test price automatically Reported.
The law has moved since. New York’s Algorithmic Pricing Disclosure Act, in force since November 2025, requires a disclosure when a price is set by an algorithm using the customer’s personal data Published. Whether a random price test that uses no personal data falls under it hasn’t been settled; ask counsel if you sell into New York. At the federal level, the FTC published initial findings in January 2025 from a study of how companies use personal data to set individual prices, without declaring the practice illegal Published. One Shopify price-testing app, Intelligems, sets the store’s list price to the highest price in the test and applies test prices at checkout, so ads and shopping feeds never show a lower price than the store Reported.
Four rules keep a price test defensible:
And be realistic about size. A price test is a revenue test, and chapter 3 showed that revenue needs several times the traffic of conversion. Most stores should test prices in bold steps, across a product line at once so the test has enough traffic, and only with visitor-level randomization that keeps each person on one price. Don’t use on-off weeks for price: returning shoppers would see the price change from week to week.
Three features sold with big conversion numbers. Each can help. Most of the headline numbers don’t measure what the feature caused.
Quizzes, pop-ups and site search share a statistical trap. The people who use them were already different from the people who don’t. Compare users with non-users and the feature gets credit for the difference.
Quiz vendors report that quiz takers convert several times better than other visitors. The quiz vendor Octane AI’s site claims “4x higher conversion,” with no sample, period or comparison group disclosed Reported. Shoppers who take a quiz are engaged by definition. The only way to learn what a quiz adds is to show its entry points to a random half of visitors and compare the halves.
A quiz’s lasting value is often the answers, not the conversion. Store each shopper’s current answers on their profile in your email platform, so segments and emails can use them, and store each completed quiz as a dated event, so you can see answers change and trigger flows from them. Version the quiz, so a changed question doesn’t silently break every segment built on it. And drop any question you can’t name a use for. Every extra question costs completions. I’ve written the full data model up as a note: Most brands that have quiz data cannot build one segment from it.
Omnisend, an email platform that sells pop-ups, reports that across 1.24 billion pop-up views in 2025, the average sign-up rate was 2.1%: 2.0% for a standard form, 2.3% for a multi-step one and 3.5% for a spin-to-win wheel Reported. A pop-up grows the list, which is worth a lot. It can also cost orders by covering the product on arrival, and Google discourages full-screen pop-ups on mobile (chapter 10).
So on phones, prefer a small banner, or a pop-up that waits until a second page view or real scrolling. Judge the pop-up on both sides of the trade: sign-ups gained, and conversion on the pages where it shows, against a random slice of visitors who never see it.
Users of a feature were different before they used it. Compare random halves, not users with non-users.
“Shoppers who search convert two to three times as often” traces to blog posts from 2013 and 2015 that don’t say how it was measured Reported, and it measures intent: people who search already know what they want. The better reason to care about search is that it’s often broken. Baymard’s 2026 benchmark of more than 170 sites found mediocre-or-worse search on 46% of desktop and 58% of mobile sites. 66% failed to handle searches for things other than products, like “returns” or “shipping,” and 54% failed on abbreviations or symbols Reported.
The fix is a weekly habit. Pull the searches that returned no results, add synonyms and redirects for the common ones, and make sure the search box is visible on mobile, not hidden behind a menu.
The most famous website test in politics measured a 40.6% lift. It also produced a $60 million headline that nobody measured. Both halves are the lesson.
In December 2007, Dan Siroker, then working on Barack Obama’s presidential campaign and later a co-founder of the testing company Optimizely, ran a test on the campaign’s splash page, the page that asked visitors to sign up with their email address before entering the site. He wrote it up in 2010 on Optimizely’s blog, and the story has been told in marketing talks ever since.
Two parts of the page changed at once. Four versions of the button text (“Sign Up,” “Learn More,” “Join Us Now” and “Sign Up Now”) were crossed with six pieces of media: three photos and three videos. That’s 24 combinations, shown to 310,382 visitors. The winner, a “Learn More” button with a family photo, raised the sign-up rate from 8.26% to 11.6%, a lift of 40.6%. And every video did worse than every photo Reported.
The post’s title says the test raised $60 million. It didn’t measure that. Siroker estimated that about 10 million people signed up over the campaign, assumed the lift held the whole time, and calculated 2.88 million extra email addresses. He then assumed 10% of sign-ups volunteered, for 288,000 extra volunteers, and that each address was worth an average of $21 in donations, for $60 million Reported. Each step is reasonable. None was measured by the test, and the post reports no significance test or range for the lift itself. With 24 combinations, the best one also benefits from the winner’s curse in chapter 4: picking the top of 24 flatters it.
Keep the measurement and the extrapolation in separate sentences. The first is evidence. The second is a forecast.
That’s not a criticism of the test, which was well designed and clearly worth running. It’s a lesson in how results travel. A measured 40.6% lift on sign-ups became “$60 million” in the headline, and the headline is what people remember. In your own reports, write the measured number first, with its range, and put the projection after it with each assumption named.
A page of paperwork before each test and a line after it. That’s the difference between a team that tests and a team that learns.
A test that won without a written reason teaches you to repeat a result you don’t understand. A test that lost with one teaches you something about your customers you’ll use for years. The brief is where the reason goes.
Fill it in before anything launches. The Whole Machine has a shorter version; this one adds the decisions that part one of this guide showed are easy to get wrong.
Decide what you’ll do with every possible result before you see any of them.
One row per test and per major site change, in a shared sheet anyone can search: the date, the page, the brief, the rung, the planned and actual size, the result with its range, and what changed because of it. Before anyone proposes a test, they search the log. A year in, the log is the most valuable document in your conversion program, because it’s the only record of what your changes did, as opposed to what someone remembers they did.
A page that needs four weeks per test can carry at most thirteen tests a year, one after another, and fewer once you allow for launches and sales. Tests on different pages can run at the same time, as long as each assigns visitors at random on its own and the two changes can’t clash (two tests both changing the delivery message, say). Counted against the traffic table, most DTC stores I see have room for somewhere between two and twenty real tests a year. Spend them on bold changes drawn from research, not on small ideas drawn from opinion.
From my workFor experiences a testing tool doesn’t reach, like email flows and post-purchase pages, I stamp every customer with a random digit from 0 to 9 when they’re first recorded, and never change it. Holding out digit 0 gives a standing 10% control group, with no setup per test. It measures everything you’ve shipped to the other 90% together, so read it as the combined effect of your program. To read one change on its own, give it its own random split.
Numbers first, then customers, then fixes, then one test worth running. Four weeks, in that order.
Whether you’re starting conversion work or restarting it, the order is the same: learn what your traffic can support, learn what’s wrong from the people who shop, fix what’s broken, and only then spend traffic on a test.
At day thirty you’ll have fewer tests running than most teams, and more reason to believe the ones you have. The fixes will already be live and logged. And the ranked problem list will keep you busy for a quarter.
Research first, fixes second, tests third. Most teams run it backwards.
Six things the person who owns conversion needs on the first day.
Whoever owns conversion, a new hire, an agency, or you on the Monday you decide to stop guessing, needs six things on day one. Without them, the first month goes on hunting for things they should have been handed.
The books and papers this guide leans on, and what to take from each.
And the research: Kohavi, Deng and Vermeer (2022) on false winners; Fabijan and colleagues (2019) on broken splits; Johari and colleagues (2022) on peeking; Deng, Xu, Kohavi and Walker (2013) on variance reduction; Nielsen and Landauer (1993) and Faulkner (2003) on how many users to test; Lewis, Singh and Fay (2006) and Shehu, Papies and Neslin (2020) on free shipping; Iyengar and Lepper (2000), Scheibehenne and colleagues (2010) and Chernev and colleagues (2015) on choice. Full references are in Appendix C.
Andrew Lauchner runs Growth Legend, embedding inside consumer brands to own lifecycle, email and SMS, and revenue operations. He wrote The Second Order, on turning first-time buyers into second-time buyers; Close the Loop, on getting customers to bring the next customer; The First Offer, on the offer that wins the first order; The Whole Machine, on the fundamentals of DTC growth; and The Standing Order, on subscription programs.
As Senior Director of Growth and Retention Marketing at Gallery Furniture, he rebuilt the customer journey and the sales playbooks together. He has worked on growth and retention at Binance and 3Commas, and has been Head of Growth and Retention at Greatness Wins and at Nexus Agriscience.
The methods marked “from my work” come from auditing client stores and running their programs. Clients aren’t named and their numbers aren’t here.
“Andrew led retention, lifecycle, and email/SMS, but what separates him from most in this space is how deeply he understands the role retention plays in the overall growth engine.”
Akram Khan, Head of Marketing at Gallery Furniture, senior to Andrew but didn’t manage Andrew directly
Andrew answers every note from people trying to make their stores convert, including those looking for someone to own it. Write to andrew@growthlegend.com or message him on LinkedIn.
The formulas behind the three calculators, and three queries every store should be able to run.
| For | Formula | Notes |
|---|---|---|
| Visitors per version | n = 16 × p(1 − p) / d² | p: current conversion rate. d: absolute change to detect. 95% confidence, 80% power; use 21 for 90% power. |
| Smallest detectable lift | L = √(16(1 − p) / (p × n)) | Relative lift, for n visitors per version. |
| Power at n | Φ(d × √(n / (2p(1 − p))) − 1.96) | n is visitors per version. Φ is the standard normal cumulative distribution function. |
| False winners | (1 − s)(α/2) / ((1 − s)(α/2) + s × power) | s: share of ideas that work. α: two-sided threshold. Kohavi, Deng and Vermeer, 2022. |
| Split check | χ² = Σ (observed − expected)² / expected | One degree of freedom for two versions. Flag p < 0.0005. |
| Two-version result | z = (pB − pA) / √(p̄(1 − p̄)(1/nA + 1/nB)) | p̄: pooled conversion rate, for the p-value. The range is for the ratio pB/pA: exp(ln(pB/pA) ± 1.96√((1 − pA)/xA + (1 − pB)/xB)) − 1, where x is orders. The verdict follows the range. |
Randomize and count at the same level. If the tool assigns visitors, analyze visitors, not sessions or page views. A visitor who comes back five times is one visitor, and counting their sessions separately makes the result look more certain than it is.
-- sessions reaching each step, last 28 days, by device
-- events: one row per event, with session_id, device, name, ts
SELECT device,
COUNT(DISTINCT session_id) AS sessions,
COUNT(DISTINCT session_id) FILTER (WHERE name = 'view_item') AS viewed_product,
COUNT(DISTINCT session_id) FILTER (WHERE name = 'add_to_cart') AS added_to_cart,
COUNT(DISTINCT session_id) FILTER (WHERE name = 'begin_checkout') AS began_checkout,
COUNT(DISTINCT session_id) FILTER (WHERE name = 'add_shipping_info') AS entered_shipping,
COUNT(DISTINCT session_id) FILTER (WHERE name = 'add_payment_info') AS entered_payment,
COUNT(DISTINCT session_id) FILTER (WHERE name = 'purchase') AS purchased
FROM events
WHERE ts >= CURRENT_DATE - INTERVAL '28 days'
GROUP BY device;
The event names follow Google Analytics 4’s recommended ecommerce events; rename them to match your setup. The syntax is Postgres; in BigQuery, write COUNT(DISTINCT IF(name = 'add_to_cart', session_id, NULL)). Divide each column by the one before to get each step’s pass rate, and compare phones with desktop step by step. Express wallets can skip the shipping and payment events, so a step above 100% means an event is missing, not a miracle.
-- visitors assigned to each version, counted once each SELECT variant, COUNT(DISTINCT visitor_id) AS visitors FROM exposures WHERE experiment_id = 'pdp-delivery-date' GROUP BY variant;
Put the two counts into the reader in chapter 5. Run the check daily in the first days of any test, not only at the end: a broken split caught on day two costs two days.
-- visitors, orders and conversion for visitors who saw a page type
SELECT page_type,
COUNT(DISTINCT s.visitor_id) AS visitors,
COUNT(DISTINCT o.order_id) AS orders,
COUNT(DISTINCT o.order_id)::numeric
/ COUNT(DISTINCT s.visitor_id) AS orders_per_visitor
FROM page_views s
LEFT JOIN orders o
ON o.visitor_id = s.visitor_id
AND o.ordered_at BETWEEN s.viewed_at AND s.viewed_at + INTERVAL '1 day'
WHERE s.viewed_at >= CURRENT_DATE - INTERVAL '28 days'
GROUP BY page_type
ORDER BY visitors DESC;
This counts an order toward a page if it came within a day of the visit. That’s an approximation, and a generous one, but it’s the right shape for deciding which pages can carry a test.
Five one-page forms. Copy them into whatever your team already uses.
BECAUSE WE SAW evidence from research (sessions, survey, funnel): WE BELIEVE the change, and who sees it: PRIMARY METRIC one: GUARDRAILS order value / returns / contribution / speed: SMALLEST LIFT worth acting on: % PLANNED SIZE visitors per version: weeks: power: RUNG A/B / holdout / on-off weeks / before-and-after IF IT WINS we will: IF IT LOSES we will: IF IT CAN'T TELL we will: OWNER name, and the date it will be read:
BEFORE "We're testing the site, not you. Please think out loud.
There are no wrong answers, and I won't be able to help."
TASK 1 "You're shopping for [need]. Find something you'd buy."
TASK 2 "Buy it. Stop when you'd have to pay."
TASK 3 "Before buying, find out when it would arrive and how
returns work."
AFTER "What was the most confusing moment?"
"What almost made you leave?"
NOTE Every hesitation over five seconds, every tap on
something that isn't a link, every question asked aloud.
ON THE CONFIRMATION PAGE, OR THE FIRST EMAIL AFTER PURCHASE One question, free text: "What almost stopped you from buying today?" Optional second question, multiple choice: "Where did you first hear about us?" Read every answer weekly. Group them. Count the groups.
DATE | PAGE | CHANGE | RUNG | BRIEF LINK | PLANNED SIZE
| ACTUAL SIZE | SPLIT CHECK | RESULT (WITH RANGE)
| DECISION | WHAT WE LEARNED | OWNER
CHANGE what, where, live date:
METRIC one, decided before the change:
BEFORE WEEKS whole weeks, no promotions:
AFTER WEEKS same length, no promotions:
COMPARISON what the change didn't touch (other device, similar
products):
TRAFFIC MIX CHECK share by channel, before vs after:
LAST YEAR same weeks' change last year:
RESULT change relative to the comparison (ratio), and the
same figure for past periods with no change:
Every external source, by chapter. Web sources were read in September 2026.