If a test can't expect ten conversions per arm, don't run it. Decide and say so.
Before your next test ships, multiply its baseline rate by each arm's size. The result is the number of conversions each arm can expect. Under ten, don't run the test, because no method turns eight conversions against eleven into a finding. Decide another way and write down how you decided.
Call the rule the Single-Digit Stop. Baseline rate × arm size = expected conversions per arm; under 10, don't run the test. For example, a 2% conversion rate on 400 people per arm expects eight orders in each. That test doesn't ship.
Two things set what a test can see. Noise shrinks with the square root of the count, so quadrupling the sample only halves the smallest effect you can detect. That is why one more week rarely rescues a test that was too small on day one. Rarity matters as much, since a rate is only as solid as the number of events behind it.
Small samples can still see large effects. What they lose is the small ones. A new subject line is a small effect; a different entry product can be a large one.
One skincare brand's six-year file shows the limit. Each figure below assumes 80% power and a two-sided 95% test.
At the lipstick row, repeat purchase has to climb from 12.2% to about 17.6% before a result is readable. A two-point gain, which would be a good result for a flow rewrite, stays invisible at that size. The interval table in Cost per Returner shows the same limit row by row.
The event count decides what you can test, whatever the headcount. When I ran sales at mostdope, a CRM for roofing and solar contractors, one quarter produced twelve new customers. Twelve is a good quarter and a hopeless sample. Subscription software sits at the other end, because every subscriber makes a renewal decision every month. Pick an outcome that happens more often.
"Inconclusive, deciding on judgment" is a legitimate verdict. It records what was measured and leaves the reasoning where the next person can find it, so the call can be reversed when better evidence arrives. A judgment call dressed as statistics does more damage. "Variant B won" on a few hundred people per arm gets written into a playbook and used a year later against a finding that holds.
Multiply your baseline rate by your arm size. If you expect well over ten conversions per arm, run the test; this chapter is for smaller files and rarer outcomes.
This is one chapter of The Second Order, which is free and readable in full on a single page with no form in front of it.