Part four · Chapter 18

THE SINGLE-DIGIT STOP

If a test can't expect ten conversions per arm, don't run it. Decide and say so.

Before your next test ships, multiply its baseline rate by each arm's size. The result is the number of conversions each arm can expect. Under ten, don't run the test, because no method turns eight conversions against eleven into a finding. Decide another way and write down how you decided.

The Single-Digit Stop

Call the rule the Single-Digit Stop. Baseline rate × arm size = expected conversions per arm; under 10, don't run the test. For example, a 2% conversion rate on 400 people per arm expects eight orders in each. That test doesn't ship.

Two things set what a test can see. Noise shrinks with the square root of the count, so quadrupling the sample only halves the smallest effect you can detect. That is why one more week rarely rescues a test that was too small on day one. Rarity matters as much, since a rate is only as solid as the number of events behind it.

Small samples can still see large effects. What they lose is the small ones. A new subject line is a small effect; a different entry product can be a large one.

What one file can read

One skincare brand's six-year file shows the limit. Each figure below assumes 80% power and a two-sided 95% test.

+5.4 pts
smallest readable lift at the lipstick row's 12.2% repeat baseline, 1,379 customers across both arms
930
customers across both arms to see +2 points on the 0.20% rate at which travel-size buyers graduated to full size
2,180
customers across both arms to see +1 point on the same 0.20% base

At the lipstick row, repeat purchase has to climb from 12.2% to about 17.6% before a result is readable. A two-point gain, which would be a good result for a flow rewrite, stays invisible at that size. The interval table in Cost per Returner shows the same limit row by row.

The event count decides what you can test, whatever the headcount. When I ran sales at mostdope, a CRM for roofing and solar contractors, one quarter produced twelve new customers. Twelve is a good quarter and a hopeless sample. Subscription software sits at the other end, because every subscriber makes a renewal decision every month. Pick an outcome that happens more often.

Run the check

  1. Write the baselineTake it from your own file, for the outcome the test will be graded on, at the definition you'll use. No industry benchmark.
  2. Count each armUse last month's real entry volume. Subtract anyone suppressed, bounced or already in another test.
  3. Multiply and stopBaseline × arm size. In single digits, stop here.
  4. Write your expected effectWrite it before you look at the detectable effect, so the calculator can't talk you into its number.
  5. Get the detectable effectAny sample-size calculator at 80% power and a two-sided 95% test. If it beats the effect you expect, the test can't answer.

When the answer is no

"Inconclusive, deciding on judgment" is a legitimate verdict. It records what was measured and leaves the reasoning where the next person can find it, so the call can be reversed when better evidence arrives. A judgment call dressed as statistics does more damage. "Variant B won" on a few hundred people per arm gets written into a playbook and used a year later against a finding that holds.

Wrong for you if

Multiply your baseline rate by your arm size. If you expect well over ten conversions per arm, run the test; this chapter is for smaller files and rarer outcomes.

Do this

This is one chapter of The Second Order, which is free and readable in full on a single page with no form in front of it.