Luck, variance and other people’s playbooks, and how to tell a signal from a bad week.
Every Monday, somebody explains last week. Revenue fell 9% because the new creative is tired. The welcome flow’s conversion rate jumped because of the subject line. Repeat rate dipped because of the shipping delay. Most weeks, most of those stories are about noise.
Those two numbers are the guide in miniature. The first is what happens when you learn only from the companies still standing: the lesson you draw can be the opposite of the truth. The second is what happens when you act on every movement: the most active traders did worst. Operators make both mistakes every week. We copy the brand that won and never meet the ten that did the same thing and vanished. We rewrite the flow, the ad and the offer because a number moved, when the number would have moved anyway.
A number that moves isn’t news. A number that moves more than it usually does might be.
The noise floor is a term from audio engineering: the background hiss of a system, below which you can’t hear a real signal. Every metric on your scorecard has one. This guide shows you how to measure it, how to hear what rises above it, and how to make decisions before you know how they’ll turn out: from base rates, with kill criteria set in advance, and with a written record you can learn from later.
It doesn’t teach A/B testing; The Honest Test does. It assumes you know your contribution margin and cohorts, which The Whole Machine covers.
Start with The Judgment Audit. Your lowest checks name the chapters to read first. Or follow a path:
Three tools run in the page. Nothing you type leaves your browser.
Examples that open with Say or Picture use made-up round numbers. Every source is listed in Appendix C.
What this guide argues, and what would prove each claim wrong.
A position says what would prove it wrong. Test each on your own numbers.
Eight questions operators ask every month. Each has a trap that makes the obvious answer wrong, and a tool that fixes it.
Find the question you asked last week. Read across.
The whole book is above and always will be. These are the same chapters addressed individually, for linking to one idea rather than to ninety.
| The question | The trap | The fix | Chapter |
|---|---|---|---|
| Did revenue really drop this week? | Explaining noise | A process behavior chart with limits from your last 12 weeks | 3 |
| Should we rewrite the flow? | Tampering with a steady process | Change only on a signal or a plan | 4 |
| Is this campaign a winner? | Regression to the mean | Shrink the result toward your average | 5 |
| Is this channel, agency or hire good? | Calling luck skill | Check whether rank persists from one period to the next | 6 |
| Should we copy what they did? | The halo effect and survivors | Ask the evidence questions first | 7 |
| Will this launch hit plan? | The inside view | Forecast from a reference class, then adjust | 9 |
| When do we stop? | Deciding under sunk cost | Kill criteria written before launch | 10 |
| Was that a good call? | Judging by the outcome | A decision journal written at the time | 12 |
The first four rows are about reading what already happened. The last four are about deciding before you know. Both halves rest on the same idea: an outcome is a draw from a range of possible outcomes, and you only ever see one draw. The job is to estimate the range, act on the range, and learn from the draws without treating each one as a verdict.
You only see one draw. Decide on the range.
Twelve checks on whether your team reacts to signals or to noise, and whether it decides before it knows or after. About forty minutes with your scorecard, your flow history and your last three launch plans.
This audit isn’t about whether your numbers are good. It’s about whether the way you read them, and the way you decide, can tell a real change from a normal week, and a good decision from a lucky one.
Open your weekly scorecard, the change log for your email and SMS flows, your last three launch or campaign plans, and your notes from the last few weekly meetings. Score each check 0 to 2: 0 if it failed or nobody can answer it, 1 if partly true, 2 if clean. “We usually talk about it” scores 0. The check is whether it’s written down.
If it isn’t written down before the result comes in, it didn’t happen.
Score as you go; your band appears when all twelve are in.
| Score | What it means | Read next |
|---|---|---|
| 20–24 | Your team reads signals, not noise, and decides before it knows. Keep the records going and review them each quarter. | The Signal Scorecard, then The Decision Journal |
| 14–19 | You read the numbers well in places and react to noise in others. Fix the zeros first. | The chapter linked from your lowest check, then The Process Behavior Chart |
| 8–13 | Most of what gets explained on Monday is noise, and most big bets are judged by how they turned out. | Part one, starting at Most Movement Is Noise |
| 0–7 | Stop explaining weekly moves for a month. Put limits on five numbers and write kill criteria for every live bet. | The Process Behavior Chart, then The First Thirty Days |
If checks 9 to 12 averaged lower than checks 1 to 8, start with part three. Reading numbers is something your team already practices every Monday. Writing things down before the result arrives is a new habit, and it’s the cheaper one to build.
Before anyone explains a number, ask how much it moves when nothing has changed. The answer is usually more than the move you’re explaining.
Every weekly number is the sum of thousands of small causes: who happened to be paid that Friday, the weather in Texas, what the ad auction did on Tuesday, which day the email went out, a creator’s post you never saw. Those causes don’t stop. They make every number wobble, every week, with no story behind the wobble.
Start with the smallest source of wobble: pure chance. Say a brand ships 10,000 orders a month, about 2,300 a week. Even if nothing about the store, the traffic or the customers changed, the weekly count would bounce around 2,300 the way coin flips bounce around half heads. The 95% range from chance alone is about 94 orders either way, or 4%. Two weeks compared with each other can differ by almost 6% by chance alone Derived.
That’s the floor, and it’s the best case: a big count and a simple metric. Smaller denominators wobble far more.
| Number | Typical week | Chance alone, either way | As a share of the number |
|---|---|---|---|
| Weekly orders | 2,300 orders | ±94 orders | ±4% |
| Site conversion rate | 2.5% of 90,000 sessions | ±0.10 pts | ±4% |
| Email click rate | 1.5% of 20,000 recipients | ±0.17 pts | ±11% |
| Repeat rate, one month’s cohort | 25% of 400 customers | ±4.2 pts | ±17% |
| One SKU’s sales | 40 units | ±12 units | ±31% |
| A campaign’s order rate | 0.15% of 12,000 recipients | ±0.07 pts | ±46% |
Derived95% ranges: 1.96 × √(p(1 − p) / n) for rates and 1.96 × √n for counts. This is sampling chance only. Real weeks vary more.
Read the last two rows again. A campaign to 12,000 people that normally produces 18 orders can produce 26 or 10 with nothing different about it. A new SKU selling 40 units a week can sell 28 or 52. Those are the numbers most teams make their fastest decisions on.
The smaller the denominator, the louder the noise. The loudest numbers get the fastest decisions.
The table understates it, because the thousands of small causes add their own variation on top of sampling chance. Walter Shewhart of Bell Laboratories, who invented the control chart, called them chance causes, and the rare big ones assignable causes; W. Edwards Deming later called the second kind special causes Reported. A good week and a bad week from the same steady system are both made of chance causes. Nobody can find “the reason” for one of them, because there isn’t one.
You can’t compute that total wobble from a formula. You measure it from your own history, which is what the chart in the next chapter does. How wide it is depends on the business. In the example in the next chapter, the limits sit about 20% either side of average weekly revenue, and that store has nothing wrong with it.
Open rate is the noisiest number in the building, and not only because of chance. Apple’s Mail Privacy Protection downloads remote content “in the background by default,” whether or not the reader engages with the email, and loading that content is what makes an open register. Apple says it stops senders from learning when or how many times a message was opened Reported. So a change in opens can come from how many of your recipients use Apple Mail on a given send. Judge subject lines and sends on clicks, orders or revenue per recipient. For deliverability, which does use opens as one input, see The Whole Machine.
Campaign results carry the small-denominator problem too. Most campaigns produce a few dozen orders, and a few dozen orders move by a third on chance alone. One campaign doesn’t tell you whether an offer, a subject line or a day of the week works. Twenty campaigns might.
A monthly cohort of 400 first-time buyers can show a 30-day repeat rate of 21% or 29% from the same underlying behavior. Read repeat rate by cohort and over at least three cohorts before you credit a new post-purchase flow or blame a shipping delay. The Second Order covers how to build the cohort view.
There are two mistakes, and fixing one makes the other more likely. You can treat noise as a signal: explain it, act on it, change something that wasn’t broken. Or you can miss a real signal because every week already has a story and the real one looks like the rest. Shewhart’s chart was designed to balance the two, so that you react rarely enough to avoid chasing noise and often enough to catch real change. Gut feel guards reliably against neither.
Two lines drawn from your own last twelve weeks tell you which moves deserve a meeting. The math fits on an index card.
On May 16, 1924, Walter Shewhart sent his superiors at Bell Laboratories a memo proposing a simple chart Reported. A century later, the version Donald Wheeler calls the process behavior chart, or XmR chart, is still the best tool for a weekly scorecard. It tells you how much a number moves when nothing has changed, from the number’s own history.
PublishedWheeler, “What Makes the XmR Chart Work?” Quality Digest, 2012; Understanding Variation, 1993. The 2.66 is 3 divided by 1.128, the constant that converts an average moving range into an estimate of the standard deviation, so the limits sit about three standard deviations from the average.
Why the moving range and not the standard deviation of all twelve weeks? Because if the number shifted partway through, the overall standard deviation includes the shift and the limits come out too wide to show it. The moving range looks only at week-to-week changes, which a shift barely touches. It’s also why the constant is 2.66 and not 2 or 3: people who swap in a rounder number get limits that are too tight and flag noise, or too loose and miss signals.
Wheeler’s book gives three rules for spotting a signal Reported:
For the moving range, use only the first rule: a single week-to-week jump above its upper limit Published. Everything else is noise. It gets noted, not explained.
That’s the pattern the chart exists to catch. The big drop felt like news and wasn’t. The real change was small, steady and invisible on a week-over-week report, because every one of those weeks, compared with the week before, looked normal.
With the defaults, this week’s $53,900 is 13% below the average and 15% below last week, and it’s noise: the limits run from $49,255 to $74,645 and no rule fires. Change this week to 48,000 and the verdict flips to a signal. The rule-of-eight shift in the figure above doesn’t show up here, because the tool judges one new week; for runs, keep charting week after week.
Changing a steady process because of a normal week doesn’t just waste effort. It adds variation. Investors show the cost in hard numbers.
Deming had a name for adjusting a steady process in response to noise: tampering. He demonstrated it with his funnel experiment, in which people try to hit a target by moving the funnel after each drop, and scatter the drops wider than if they’d left it alone Reported.
The arithmetic is simple. Suppose each week lands on the true average plus some random error. If you leave the process alone, the weeks scatter with that error’s spread. If you “correct” each week by the amount it missed last time, every result now carries two errors: this week’s, and the reversal of last week’s. The spread of results doubles in variance, about 41% wider in the units you see Derived. The team that responds to every dip isn’t steering. It’s shaking the wheel.
Responding to every dip isn’t steering. It’s shaking the wheel.
Picture a welcome flow that converts about 6% of new subscribers, week in, week out, with ordinary wobble. One week it shows 5.1%. Someone rewrites the second email. The next week it’s back to 6.2%, and the rewrite gets the credit. A month later another dip, another rewrite. After a year the flow has had eight versions, nobody knows which one was best, and the flow is exactly as good as it was, with more variation and a longer change log.
Each rewrite has costs that don’t show on the chart: the hours, the review cycle, the risk of a broken link or a wrong discount code, and the loss of any clean comparison. Worse, the regression you’ll meet in the next chapter makes the habit self-reinforcing. A bad week is usually followed by a better one, so the fix always seems to work.
Brad Barber and Terrance Odean studied 66,465 households with accounts at a large discount broker from 1991 to 1996. The fifth of households that traded most earned 11.4% a year after costs; the market returned 17.9%. The average household turned over 75% of its portfolio a year Published. In a follow-up on more than 35,000 households, men traded 45% more than women, and trading cut their net returns by 2.65 percentage points a year, against 1.72 for women Published. Their explanation was overconfidence: people who trade most are surest that they know what the next move will be.
A store isn’t a portfolio, and a flow rewrite isn’t a stock trade. But the mechanism carries over. Each action feels informed, has a small cost, and responds to movement that’s mostly noise. Added up, the activity costs more than it earns. The operator who rewrites the flow every month is trading on noise and paying the commission in hours and in lost learning.
The best campaign, SKU or channel of the quarter was probably good and lucky. The luck doesn’t come with it into next quarter.
In 1886 Francis Galton published the heights of 928 adult children of 205 pairs of parents. Tall parents had tall children, but less tall than themselves: on average, a child’s height stood about two-thirds as far from the average as the parents’ did Published. He called it regression towards mediocrity. We call it regression to the mean, and it applies to every number that mixes something real with something random.
Daniel Kahneman called it “the most satisfying Eureka experience” of his career. He was teaching Israeli Air Force flight instructors that praise works better than punishment for learning a skill, when a senior instructor objected. In his experience, cadets he praised for a clean maneuver usually did worse on the next try, and cadets he screamed at usually did better Published.
The instructor’s observation was right and his explanation was wrong. He praised after unusually good flights and screamed after unusually bad ones, and unusual flights are followed by more ordinary ones whatever the instructor says. Kahneman’s conclusion: because we reward others when they do well and punish them when they do badly, “we are statistically punished for rewarding others and rewarded for punishing them.” He then had the instructors toss two coins each at a target behind their backs, without feedback, to show the same pattern appearing with no instruction at all Published.
Praise follows a lucky week. So does disappointment.
Kahneman’s rule: “Whenever the correlation between two scores is imperfect, there will be regression to the mean” Published. How far a result falls back depends on how much of the spread between results is real. If your campaigns differ a lot in true quality and each is sent to a big audience, a winner keeps most of its lead. If they’re similar in quality and the audiences are small, most of the lead is noise and goes away.
You can estimate the split from your own history. The spread of past campaigns’ results is real differences plus chance. Chance you can compute from the audience size. Take it away, and what’s left is the real spread. The share of a winner’s lead you should expect to keep is the real spread divided by the real spread plus the chance in that one result Derived. The tool does it for you.
With the defaults, a campaign to 12,000 people that produced 30 orders, a 0.25% order rate against a usual 0.15%, should be expected to do about 0.20% next time: half its lead is likely real and half was luck. That’s 24 orders, not 30. Double the audience of the winner to 24,000 with 60 orders and about two-thirds of the lead holds, because a bigger audience carries less chance.
DerivedExpected = 0.15% + 0.5 × (observed − 0.15%), where 0.5 is the share of the lead that holds with the default inputs. Bars scaled to 0.30%.
The winners are still winners. They’re just closer to the pack than the dashboard says, and the gap between them mostly disappears.
Skill repeats and luck doesn’t. So the test for skill is whether a ranking holds from one period to the next.
Michael Mauboussin’s The Success Equation (2012) starts from a simple model: every result is part skill and part luck. You can’t see the parts in one result. You can see them across results, because skill carries forward and luck doesn’t.
In a 2012 Harvard Business Review article, Mauboussin set out what makes a statistic useful. One test is that it’s persistent: the same action produces a similar outcome from one time to the next Published. Apply that to anything you’re judging on results: channels, creatives, agencies, creators, people.
Say you ran ten creators last quarter and ten again this quarter, the same ten. Rank them by cost per order in each quarter and compare the two rankings. If the top three last quarter are near the top again, results are persistent and the ranking reflects something real. If last quarter’s top three are scattered through this quarter’s list, the ranking was mostly luck. Spreadsheets do this in one line: the correlation between the two columns of ranks.
That correlation does double duty. It’s also roughly the share of a winner’s lead you should expect next period, which is the same shrinkage as chapter 5‘s tool, measured a different way Derived.
A ranking that doesn’t repeat isn’t a ranking of skill.
Mauboussin’s paradox of skill: as people get better at an activity, the gaps between the best, the average and the worst become “much narrower” Published. When everyone is skilled, luck decides more of the ranking. I’d expect paid social to work this way. When every brand in a category uses the same ad platform, the same creative formats and similar agencies, the difference in skill between them shrinks, and the week-to-week ranking of who’s winning becomes noisier. The edge moves to the places competitors can’t copy quickly: the product, the offer, the retention program.
Mauboussin put the sports version plainly: a player with a great statistical year was almost always “skillful and lucky” Reported. The same holds for a founder who had one great year, or an agency whose case study features its best client. Both parts were there. Only one of them is for hire.
Results over one quarter are a poor measure of a person whose results depend on campaigns, markets and budgets they don’t control. Judge on two things instead: results over several periods, and the process measures they do control, like whether tests were set up properly, whether launches shipped on time, whether the scorecard was kept. Process measures tend to persist, which is why they’re fairer.
Case studies and founder threads are written after the outcome is known, about the companies that survived. Both facts bend the lesson.
A founder posts the seven things that took the brand to $50 million. Every item is plausible. Some may even be true. But the post is built from two biased samples at once: it describes a winner, and it describes the winner as seen by someone who already knows it won.
Phil Rosenzweig’s The Halo Effect (2007) takes apart the most popular business research. The halo effect is “the tendency to make specific inferences on the basis of a general impression” Published. When a company is doing well, observers describe its strategy as clear, its culture as strong and its leaders as decisive. When the same company does badly, the same observers find confusion, arrogance and drift. His example is Cisco, praised for its strategy in the late 1990s and criticized by many of the same observers when the bubble burst, though the company hadn’t fundamentally changed. The data behind much of the popular research comes from “retrospective interviews, articles from the business press, and business school case studies,” all colored by the outcome Published.
Rosenzweig’s test is one question: if I didn’t know how the company was performing, what would I think of its culture, its execution and its customer focus? Published
Jerker Denrell showed why learning from successful firms goes wrong even without the halo. Failed firms disappear from view, so the sample you learn from is survivors. His 2003 paper shows how a risky practice with no link to performance across all firms can seem positively related to performance among the survivors Published. A bold strategy produces big winners and big losers. Look only at the survivors and it looks like a strategy for winning.
DTC is full of these. Raise big, spend big on brand, open stores, launch a second category. Some brands did those things and thrived. More did them and closed, and nobody writes the thread about them.
Peter Golder and Gerard Tellis tested the belief that pioneers win. Earlier studies had relied on databases of surviving firms. Using historical records for about 500 brands in 50 categories, they found that 47% of market pioneers failed, their mean market share was 10%, and only 11% were still category leaders. Earlier studies had put pioneers’ share near 30% and found almost half of them leading. The leaders they found entered on average 13 years after the pioneer Published. The pioneers who survived long enough to be studied made first-mover advantage look real.
The advice comes from the survivors. The evidence is with the ones who didn’t make it.
Jeffrey Pfeffer and Robert Sutton made the case for evidence-based management in Harvard Business Review and in Hard Facts, Dangerous Half-Truths and Total Nonsense (both 2006): as medicine learned to do, managers should decide from the best evidence about what actually works Published. The questions below are mine, not theirs. They’re what I’d ask before copying anything from another brand:
Three DTC brands that were case studies in both directions, and one celebrated turnaround whose numbers weren’t real. What to take from them, and what not to.
These aren’t stories about foolish founders. Each brand built something customers liked, and each was written up as a model before it was written up as a warning. The point is how much of both stories was hindsight.
Casper went public in February 2020 at $12 a share Filed. Its revenue grew from $358 million in 2018 to $439 million in 2019 and $497 million in 2020, while it lost about $90 million in each of those years Filed. In January 2022, Durational Capital Management completed its acquisition of the company for $6.90 a share in cash, and Casper left the New York Stock Exchange Filed. The losses were in the filings the whole time. The story around them changed.
Allbirds priced its IPO in November 2021 at $15 a share Filed. Net revenue rose from $194 million in 2019 to $298 million in 2022, then fell each year to $152 million in 2025 Filed. In 2026 the company disclosed substantial doubt about its ability to continue as a going concern, sold its footwear business’s assets, including the brand, to American Exchange Group in a sale that closed on June 9, and renamed itself Smartbird, with a new business acquiring and monetizing the GPU chips used in AI computing. Its 10-Q for the second quarter says the sale proceeds and new financing alleviated the doubt Filed. Sustainability, the wool runner and the founders’ story were once cited as the reasons it grew. They were the same things when it shrank.
Glossier was valued at $1.8 billion in July 2021. In January 2022 it cut more than 80 corporate jobs, about a third of its corporate staff, many in technology, and in May 2022 its founder Emily Weiss stepped down as chief executive to become executive chair. In her staff email about the layoffs, she said the company had “got ahead of ourselves on hiring” and wrote, “these missteps are on me” Reported. That’s rarer than it should be: a founder’s own account, given at the time, of which decisions went wrong. It’s more useful than any case study written afterward.
The cautionary case is older. Al Dunlap’s Mean Business (1997) was the playbook of a celebrated turnaround artist. At Sunbeam, the SEC charged, at least $60 million of the company’s reported $189 million in 1997 earnings from continuing operations before taxes came from accounting fraud. In 2002, without admitting or denying the charges, Dunlap agreed to pay a $500,000 civil penalty and to a permanent bar from serving as an officer or director of a public company. He also paid $15 million of his own money to settle a related shareholder class action Filed. Readers who copied his methods were copying a result that, at Sunbeam, the SEC said was partly manufactured.
A case study is only as good as the numbers it was built on, and it was written after the ending was known.
Before you forecast from your plan, look at what happened to plans like yours. Then move from there, not from the plan.
Daniel Kahneman once helped write a new curriculum. Partway through, he asked the team how long it would take to finish. The estimates ranged from 18 to 30 months. Then he asked the curriculum expert on the team how long similar projects had taken. About 40% had given up, he said, and of the rest, he couldn’t think of one finished in less than seven years or more than ten. The project was finished about eight years later Published.
Kahneman and Dan Lovallo told that story in a 1993 paper that named the problem. The inside view forecasts “by focusing on the case at hand”: the plan, its obstacles, scenarios of how it will unfold. The outside view “essentially ignores the details of the case at hand” and looks instead at the statistics of a class of similar cases Published. The expert knew the base rate. He still gave an inside-view estimate until someone asked the other question.
Bent Flyvbjerg turned the outside view into a method, reference-class forecasting, and built a database of more than 16,000 big projects. In it, only 8.5% came in on budget and on time, and 0.5% on budget, on time and with the benefits promised Published. Your launches aren’t bridges. But the pattern of plans beating results is the same, and so is the fix.
Your plan is one scenario. The reference class is what usually happens.
Your own history is the best reference class because it shares your team, product and customers. The ratio of actual to plan is often the most useful single number in it: if your launches have delivered 60% of plan on average, multiply the next plan by 0.6 before you buy inventory.
The outside view is where you start, not where you stop. Your specifics matter, to the degree your specifics have predicted results before. The adjustment is the same regression as chapter 5: begin at the reference-class median and move toward your own estimate by the share that your past forecasts have tracked results. If they’ve tracked well, move most of the way. If they haven’t, barely move Derived.
With the defaults, a team forecasting $400,000 for a launch whose reference class usually does $120,000, and whose forecasts have tracked results only loosely, should plan on about $172,000. Eight in ten launches like it land between about $54,000 and $554,000, and the $400,000 plan has about an 18% chance of happening. That doesn’t mean don’t launch. It means don’t buy inventory for $400,000.
Write down what would make you stop, and when you’ll check, before the money is spent and the team is attached.
Every project is easiest to judge before it starts. Nobody’s reputation is tied to it yet, no money is sunk, and the team can still say what a failure would look like without it sounding like a verdict on anyone. After launch, all of that changes, and the question “should we keep going?” gets answered by hope.
Annie Duke’s Quit (2022) is a book-length argument for this. Among its tools are quitting contracts, set up in advance, and a warning about escalation of commitment: the pull to keep investing in something because you already have Reported. Kill criteria are the operator’s version: a written condition, a date and a decision-maker, agreed before launch.
The best time to decide when to stop is before you start. The second best is now.
Say a brand with a $60 contribution per first order launches a new paid channel with $30,000 over eight weeks. Its last three channel tests reached a cost per first order about 40% above the plan in their first two months. The plan says $50. The kill criteria, written before launch:
Nothing here stops the team from learning. It stops the team from arguing about the threshold in week 6, when everyone already knows which answer they want.
Before launch, assume it failed and write down why. Twenty minutes that surface the risks everyone saw and nobody said.
Gary Klein described the method in Harvard Business Review in 2007. “A premortem is the hypothetical opposite of a postmortem.” Instead of asking what might go wrong, the leader tells the team the project “has failed spectacularly” and asks why Published.
PublishedKlein, “Performing a Project Premortem,” Harvard Business Review, September 2007.
Klein’s article cites a 1989 study by Deborah Mitchell, Jay Russo and Nancy Pennington, and says it found that imagining an event has already happened “increases the ability to correctly identify reasons for future outcomes by 30%” Published. That’s not quite what the study measured. It counted reasons, not correct ones. People told an outcome was certain gave longer explanations, more tied to specific events, than people told it was only possible; a 2010 paper co-written by Klein summarized the effect as about 30% more reasons Published. The 1989 paper’s own abstract adds that certainty, not setting the event in the future or the past, drove the difference. The quoted figure is about quantity, not accuracy.
Later tests are small but encouraging. In that 2010 study, by Beth Veinott, Klein and Sterling Wiggins, with 178 students judging a flu-response plan, groups that ran a premortem lowered their confidence in it by 25 points on average, against 12 to 14 points for groups that listed pros and cons or only cons Published. More reasons and less overconfidence is what a launch plan needs. Treat it as a cheap habit with a sound mechanism, not a proven 30% gain.
Certainty loosens tongues. “It failed” gets more honest answers than “could it fail?”
Write down what you expected, and why, at the time. It’s the only way to learn from decisions instead of from luck.
A good decision can turn out badly and a bad one can turn out well. If you judge decisions only by outcomes, you’ll learn the wrong lessons from both, and you’ll learn them confidently.
Annie Duke, a former professional poker player, built Thinking in Bets (2018) on this idea. Her summary: “The quality of our lives is the sum of decision quality plus luck” Reported. You control the first term. The second arrives on its own schedule.
| Good outcome | Bad outcome | |
|---|---|---|
| Good decision | Deserved success. Repeat the process. | Bad luck. Keep the process; check the risk was sized right. |
| Bad decision | Dumb luck. The most dangerous box: it teaches the wrong lesson. | Deserved. Fix the process. |
Without a record, everything gets filed by outcome, and the top-right and bottom-left boxes disappear. The dumb-luck box is where overconfident habits come from.
Judge the decision on what you knew and how you decided, not on how it turned out.
One entry per decision that matters: a launch, a big buy, a hire, a channel, a price change. Written before the result, dated, short.
Once a quarter, read the entries whose review dates have passed. For each, put it in a box of the table above. Then look across them. Are your ranges too narrow? If fewer than about eight in ten results fell inside ranges you called 80% likely, they are. Do you overrate one kind of reason? Which calls turned out right for reasons you didn’t write down? That’s where the learning is, and it’s invisible without the page.
One page, every week. Each number with its limits and its denominator, the live bets against their kill criteria, and the decisions due for review.
Most weekly scorecards show this week, last week and the change between them. That’s the format most likely to produce a story about noise. The signal scorecard shows each number against its own limits, so the page tells you where to look before anyone starts explaining.
| Number | Shown as | Read |
|---|---|---|
| Revenue | Process behavior chart, limits from 12 normal weeks; promotion weeks charted separately | Weekly |
| Orders and average order value | A chart each, so you can see which one moved | Weekly |
| Site conversion rate | Chart, with sessions beside it | Weekly |
| Cost per first order | Chart, by channel for the two biggest | Weekly |
| Revenue per recipient, campaigns and top flows | Chart each; any series under about 100 orders a week marked “read monthly” | Weekly, acted on monthly |
| Repeat rate, latest full cohort | Chart by cohort month, with the cohort size | Monthly |
| Live bets | Each one’s metric against its kill line, with the check date | At each check date |
| Decisions due for review | Journal entries past their review date | Quarterly |
No explanations for numbers inside their limits. Not even good ones.
First, no rate appears without its denominator. A conversion rate that rose because sessions fell is a different story from one that rose with sessions flat, and the page should show which. Second, every chart carries the dated notes from chapter 3: launches, price changes, flow rewrites, site changes. When a signal appears, the first place to look is the nearest note.
Limits, then winners, then forecasts and kill criteria, then the premortem and the journal. Four weeks, in that order.
Start with what the team looks at every week, because that’s where the most time goes to noise. Then the habits that need the new charts to work. By day thirty, the Monday meeting runs on signals and every live bet has a written exit.
At day thirty, you won’t know yet whether your decisions are better. That takes a quarter of journal entries and a few kill dates. What you’ll have is a meeting that spends its time on signals, forecasts that start from what usually happens, and a written record that will show you, in ninety days, where your judgment is good and where it’s lucky.
Limits before stories. Base rates before plans. Records before hindsight.
What whoever runs the weekly numbers needs on the first day.
Whoever owns the scorecard, a new analyst, an operator, or you on the Monday you decide to stop explaining noise, needs six things on day one. Without them, the first month goes on hunting for things they should have been handed.
The books worth reading next, and what to take from each.
And the research: Galton (1886) on regression; Kahneman and Lovallo (1993) on the inside and outside view; Mitchell, Russo and Pennington (1989) and Veinott, Klein and Wiggins (2010) on premortems; Denrell (2003) on survivors; Golder and Tellis (1993) on pioneers; Barber and Odean (2000, 2001) on overtrading. Full references are in Appendix C.
Andrew Lauchner runs Growth Legend, embedding inside consumer brands to own lifecycle, email and SMS, and revenue operations. He is the author of The Second Order, on turning first-time buyers into second-time buyers, and The Whole Machine, on the fundamentals of DTC growth, along with a series of field guides for DTC operators at andrewlauchner.com.
As Senior Director of Growth and Retention Marketing at Gallery Furniture, he rebuilt the customer journey and the sales playbooks together. He has worked on growth and retention at Binance and 3Commas, and has been Head of Growth and Retention at Greatness Wins and at Nexus Agriscience.
“Andrew led retention, lifecycle, and email/SMS, but what separates him from most in this space is how deeply he understands the role retention plays in the overall growth engine.”
Akram Khan, Head of Marketing at Gallery Furniture, senior to Andrew but didn’t manage Andrew directly
Andrew answers every note from operators working on this, including those looking for someone to own it. Write to andrew@growthlegend.com or message him on LinkedIn.
The formulas behind the three tools, and the queries that feed them.
| For | Formula | Notes |
|---|---|---|
| Natural process limits | X̄ ± 2.66 × mR̄ | X̄: average of the baseline weeks. mR̄: average of |Xt − Xt−1|. 2.66 = 3 / 1.128. |
| Moving range limit | 3.268 × mR̄ | Only rule for the moving ranges: a point above it. |
| Signal rules | 1 point outside; 3 of 4 beyond X̄ ± 1.33 mR̄; 8 in a row on one side | 1.33 mR̄ is halfway from the central line to a limit. |
| Chance-alone range | 1.96 × √(p(1 − p) / n); counts: 1.96 × √n | 95% range from sampling alone. |
| Share of a lead that holds | w = (s² − p̄(1 − p̄)/m) / (s² − p̄(1 − p̄)/m + p̄(1 − p̄)/n) | s: spread of past rates; m: typical past audience; n: the winner’s audience. Floor the numerator at 0. Expected = p̄ + w(observed − p̄). |
| Reference-class adjustment | μ = ln(median) + r(ln(forecast) − ln(median)) | σ = (ln P80 − ln P20) / 1.683; residual σ√(1 − r²). Range: exp(μ ± 1.2816 × residual). Plan figure: exp(μ − 0.6745 × residual). |
-- weekly revenue and orders, excluding promotion weeks from the baseline
-- orders: order_id, created_at, total_price, cancelled_at
-- promo_weeks: week_start (one row per promotion week)
WITH weekly AS (
SELECT DATE_TRUNC('week', created_at)::date AS week_start,
COUNT(*) AS orders,
SUM(total_price) AS revenue
FROM orders
WHERE cancelled_at IS NULL
GROUP BY 1
),
base AS (
SELECT w.*, ABS(revenue - LAG(revenue) OVER (ORDER BY week_start)) AS mr
FROM weekly w
WHERE week_start NOT IN (SELECT week_start FROM promo_weeks)
AND week_start < DATE_TRUNC('week', CURRENT_DATE)
ORDER BY week_start DESC
LIMIT 12
)
SELECT AVG(revenue) AS central_line,
AVG(revenue) - 2.66 * AVG(mr) AS lower_limit,
AVG(revenue) + 2.66 * AVG(mr) AS upper_limit,
3.268 * AVG(mr) AS moving_range_limit
FROM base;
The moving ranges here are computed before the twelve weeks are chosen, so the oldest week’s range reaches back one normal week; that’s fine. If a promotion week sits between two normal weeks, the range spans it, which is also fine. Swap revenue for orders or a rate to chart the others. Refunds belong in a separate series from the refunds table, not netted into revenue, so a returns problem shows on its own chart.
-- last 30 email campaigns: order rate per recipient, its spread, typical audience
-- campaign_sends: campaign_id, channel, sent_at, recipients
-- campaign_orders: campaign_id, order_id (orders attributed to the campaign)
WITH c AS (
SELECT s.campaign_id, s.recipients,
COUNT(o.order_id)::numeric / NULLIF(s.recipients, 0) AS rate
FROM campaign_sends s
LEFT JOIN campaign_orders o USING (campaign_id)
WHERE s.channel = 'email'
GROUP BY s.campaign_id, s.recipients, s.sent_at
ORDER BY s.sent_at DESC
LIMIT 30
)
SELECT AVG(rate) * 100 AS avg_rate_pct,
STDDEV_SAMP(rate) * 100 AS spread_pts,
AVG(recipients) AS typical_audience
FROM c;
Use the same attribution window for every campaign. Mixing sends to the whole list with sends to small segments inflates the spread; run it separately for each.
-- first-90-day revenue against plan, for every launch
-- launches: launch_id, product_id, launched_on, plan_revenue_90d
-- order_lines: order_id, product_id, price, quantity; orders: order_id, created_at
WITH actual AS (
SELECT l.launch_id, l.plan_revenue_90d,
SUM(ol.price * ol.quantity) AS revenue_90d
FROM launches l
JOIN order_lines ol ON ol.product_id = l.product_id
JOIN orders o ON o.order_id = ol.order_id
AND o.created_at >= l.launched_on
AND o.created_at < l.launched_on + INTERVAL '90 days'
WHERE l.launched_on < CURRENT_DATE - INTERVAL '90 days'
GROUP BY 1, 2
)
SELECT PERCENTILE_CONT(0.2) WITHIN GROUP (ORDER BY revenue_90d) AS p20,
PERCENTILE_CONT(0.5) WITHIN GROUP (ORDER BY revenue_90d) AS median,
PERCENTILE_CONT(0.8) WITHIN GROUP (ORDER BY revenue_90d) AS p80,
PERCENTILE_CONT(0.5) WITHIN GROUP
(ORDER BY revenue_90d / NULLIF(plan_revenue_90d, 0)) AS median_actual_to_plan,
CORR(LN(plan_revenue_90d), LN(revenue_90d)) AS forecast_tracking
FROM actual
WHERE revenue_90d > 0 AND plan_revenue_90d > 0;
forecast_tracking is the tracking figure for the tool in chapter 9. With fewer than about ten launches it’s rough; round it down.
-- does last quarter's ranking by cost per first order hold this quarter?
-- channel_quarters: channel, quarter, spend, first_orders
WITH r AS (
SELECT channel, quarter,
RANK() OVER (PARTITION BY quarter
ORDER BY spend / NULLIF(first_orders, 0)) AS rnk
FROM channel_quarters
WHERE quarter IN ('2026-Q1', '2026-Q2') AND first_orders >= 30
)
SELECT CORR(a.rnk, b.rnk) AS rank_persistence
FROM r a
JOIN r b ON a.channel = b.channel
WHERE a.quarter = '2026-Q1' AND b.quarter = '2026-Q2';
The same query works for creators, ad sets or campaign types; change the table. The first-orders floor keeps tiny channels, whose rank is almost pure chance, out of the comparison.
Five one-page forms. Copy them into whatever your team already uses.
BET what we're launching, and its budget: OWNER who runs it: DECIDES who applies the criteria (and who can overrule, in writing): METRIC one number the bet controls: BASE RATE what similar bets did (reference class, chapter 9): MINIMUM SAMPLE orders or weeks before the criteria count: CHECK DATE on the calendar: KILL IF metric worse than: SCALE IF metric better than: CONTINUE IF in between, for how long, and then decide with no extension:
20 MINUTES, BEFORE LAUNCH
2 min Brief the final plan and its kill criteria.
1 min "It's [date, six months out]. This failed badly."
6 min Everyone writes every reason, alone. Include the ones
you wouldn't normally say.
8 min Round the room, one reason per turn, until the lists
are empty. The founder goes last.
3 min Owner picks the reasons that change the plan today,
and the ones that become signals to watch.
AFTER List goes in the decision journal entry for this launch.
DATE today: DECISION what we chose: OPTIONS what else we considered, including doing nothing: BASE RATE what usually happens in cases like this: EXPECT a number and a range (low / most likely / high): CONFIDENCE % that the result lands in the range: WHY the two or three reasons that decided it: CHANGE MY MIND the signal or kill criterion that would reopen it: REVIEW DATE when we'll read this again: -- at review -- RESULT what happened: BOX good or bad decision / good or bad outcome: LESSON what we'd do differently, about the process:
1. SIGNALS numbers that broke a rule. Owner and one question:
"what happened that week?"
2. NOISE everything else, read out as "inside the limits."
No discussion.
3. BETS live bets at their check date. Apply the criteria.
4. DECISIONS anything decided this week goes in the journal
before the meeting ends.
TACTIC what, and whose:
1. CAUSE did it cause their success, or come with it?
2. FAILURES who did the same and failed? (name them, or say
you didn't look)
3. HALO would we rate it the same if they were struggling?
4. CONDITIONS what made it work there that we don't share?
5. SMALL TEST cheapest way to find out here:
6. DOWNSIDE what it costs if we're wrong:
Every external source, by chapter. Web sources were read in September 2026.