Part one · How defaults work · Chapter 3

HOW BIG, REALLY

Defaults hold up when the evidence is checked for publication bias. Most other nudges shrink, and the ones that look like your emails shrink most.

Behavioral science had a hard decade. Famous findings failed to replicate, and journals turned out to have published the lucky results. Before you plan around a nudge, know which effects survived the checking.

The meta-analyses

Jachimowicz and colleagues’ 2019 meta-analysis of defaults found a large average effect, d = 0.68, with wide variation: most studies found positive effects, several found none, two found negative ones. They tested for publication bias, estimated about eight studies were likely missing, and found the effect held Published. Defaults in consumer settings were stronger than average.

In 2022 Stephanie Mertens and colleagues pooled over 200 studies of all kinds of nudges, with more than two million participants, and found d = 0.43 overall, with “decision structure” nudges, which include defaults, strongest at d = 0.54 Published. Months later Maximilian Maier and colleagues re-analyzed the same data with a method that corrects for publication bias. The corrected average for all nudges was d = 0.04: no evidence of an effect. For the structure category, which holds the defaults, the evidence was “undecided,” and the authors noted that the spread of results means “some nudges might be effective, even when there is evidence against the mean effect” Published.

PublishedJachimowicz, Duncan, Weber and Johnson, Behavioural Public Policy, 2019; Mertens, Herberz, Hahnel and Brosch, PNAS, 2022; Maier, Bartoš, Stanley, Shanks, Harris and Wagenmakers, PNAS, 2022. The bias-corrected estimate for structure nudges alone was reported as undecided, not as a number.

The nudge units

The most useful study for a DTC operator is the one that compared journals with the real world. Stefano DellaVigna and Elizabeth Linos collected every trial run by two large US government nudge units, 126 trials covering 23 million people, and compared them with nudges published in academic journals. In the journals, the average nudge raised take-up by 8.7 percentage points, a 33.4% lift. In the nudge units’ full set of trials, it was 1.4 points, an 8.0% lift Published. About 70% of the gap came from selective publication combined with small samples.

Look at what the nudge units sent. About 90% of their nudges were emails, letters and postcards. Only one trial used defaults, and it was left out of the main analysis Published. So the realistic benchmark for a reminder email is 8%, not 33%, and it says nothing about defaults, a different and stronger tool.

Plan a reminder email at single-digit lifts. Save the big expectations for the choices you preselect.

What this means on Monday

Say a brand sends a replenishment reminder to 20,000 customers a month and 5% would reorder that week anyway. An 8% lift, strong by the nudge-unit standard, adds 80 orders. Someone expecting 33% was counting on 330, and will spend months rewriting copy to chase a number that was never there.

Defaults are your biggest lever, for better or worse, so they need the most care. Messages are small levers: test them with a holdout (The Honest Test covers sizing) and distrust any double-digit claim for a phrase. Chapter 11 has an example.

Do this

This is one chapter of The Free Choice, which is free and readable in full on a single page with no form in front of it.