Delivery you can count on, contacts you never get, and what to do when it breaks anyway.
Every order is a promise: this thing, in this condition, by this date. Keep it and the customer barely notices. Break it and they notice twice, once when it happens and again when they decide whether to order from you next time.
Those two numbers are the guide in miniature. The first says a broken promise costs repeat revenue even when the product itself was fine: the car came, the rider got there, just later than the app had said. The second says the contacts a broken promise creates are not a fact of life. Most of them come from things other teams did: a date the site shouldn’t have shown, a package that left the warehouse late, a tracking page that said nothing. Fix the cause and the contact never happens.
So this guide treats delivery, support contacts and recovery as one system: count every failure in one weighted number, promise dates you can keep, tell customers bad news first, send each contact reason back to the team that causes it, and make things right without believing that a good recovery beats no failure at all.
The promise kept is the retention metric. Everything else is a report on it.
It builds on The Second Order, which covers the second purchase and cohorts, and The Whole Machine, which covers contribution margin and the core flows. Returns have their own guide in this series, at /returns; they appear here only where a wrong or damaged order creates one.
Start with The Promise Audit. Your lowest checks name the chapters to read first. Or follow a path:
Four tools run in the page. Nothing you type leaves your browser.
Examples that open with Say or Picture use made-up round numbers. Every source is listed in Appendix C. The section on the FTC’s shipping rule is an operator’s summary as of September 2026, not legal advice.
What this guide argues, and what would prove each claim wrong.
A position says what would prove it wrong. Test each on your own store.
Six moments between the checkout and the next order. Each has a promise, a number that says whether it was kept, and a team that owns it.
Most brands measure the first moment and the fifth. The four in between are where promises break.
The whole book is above and always will be. These are the same chapters addressed individually, for linking to one idea rather than to ninety.
| Moment | The promise | The number | Owner |
|---|---|---|---|
| 1. The checkout | A delivery date, stated or implied | Share of orders delivered by the date shown | Ecommerce, with operations |
| 2. The warehouse | The right items, packed to survive, out on time | Ship-on-time rate; wrong-item and short-ship rate | Fulfillment or the 3PL |
| 3. The carrier | Transit in the days the date assumed | Transit time at the median and the 90th percentile, by lane | Operations |
| 4. The doorstep | Arrives intact, all of it, once | Damaged, lost and split-shipment rates | Operations and packaging |
| 5. The contact | An answer without chasing | Contacts per order, by reason | The team that caused the reason |
| 6. The recovery | Made right, fast, once | Repurchase after each failure type | Support, with a budget |
The moments feed each other. An overconfident date at moment 1 becomes a late delivery at moment 3, a contact at moment 5 and a credit at moment 6.
The contact queue is where the business’s broken promises arrive. It’s rarely where they’re made.
Twelve checks on whether your store keeps its delivery promises, knows when it doesn’t, and makes it right. About forty minutes with your order data, your help desk and a phone.
The audit isn’t about how fast you ship. It’s about whether the checkout promise matches what arrives, whether you hear about the gap before the customer does, and whether anyone outside support feels it when it breaks.
Open your order export, shipping platform, help desk and your checkout on a phone. Score each check 0 to 2: 0 if it failed or nobody can answer it, 1 if partly true, 2 if clean.
A promise nobody measures is a guess you made on the customer’s behalf.
Score as you go; your band appears when all twelve are in.
| Score | What it means | Read next |
|---|---|---|
| 20–24 | You keep your promises and know when you don’t. Your job now is pricing each failure and paying the whole team on the index. | What a Broken Promise Costs, then Paying for the Promise |
| 14–19 | The basics are in place, but some failures reach customers before they reach you. Fix the zeros first. | The chapter linked from your lowest check, then One Number for Every Failure |
| 8–13 | Support is absorbing failures made elsewhere, and nobody can say what they cost. Start counting. | Part one, starting at What a Broken Promise Costs |
| 0–7 | You don’t yet know which promises you’re breaking. Measure on-time against the date you showed, and send delay notices, this month. | An Accurate Date Beats a Fast One, then The First Thirty Days |
If you use a 3PL, checks 1, 2 and 4 depend on data they hold. Put it in the contract: order-level ship scans, delivery scans and error codes, daily.
The refund and the reship are the small part. The large part is the next order that never comes, and you can measure it on your own store.
Ask a team what a late order costs and they’ll add up the credit, the agent’s time and maybe a reship. The cost they can’t see is the customer who quietly orders less often afterward. It shows up months later, in a cohort report nobody connects to the day the package was late.
Two details matter later. Harter’s team, and a study of ratings at a South American online retailer by Serkan Akturk and colleagues, both found that the damage from lateness grows more slowly as the delay gets longer Published. The first day late does a large share of the harm, which makes the promised date in chapter 4 the cheapest lever in this guide. And Norvell’s result previews chapter 10: recovery softens a failure but doesn’t erase it.
A late order doesn’t cost you a refund. It costs you part of the next order, and the one after that.
Studies give the direction. Your own data gives the size, and the size wins budget. The method is a failure cohort:
Matching narrows the bias without removing it. The cleanest evidence comes from failures you didn’t choose, like a carrier’s regional outage: customers caught in it were failed more or less at random. Keep a dated list of such events. The query is in Appendix A.
Say a brand ships 10,000 orders a month and 6% of them fail in some way: 600 orders. Customers with no failure come back within a year at 30%. The failure cohort shows that a failure cuts that rate by a fifth, to 24%. So of the 600 failed customers, 36 who would have come back don’t. Each returning customer places 2.5 more orders in the year, at $25 of contribution each.
That’s 36 × 2.5 × $25 = $2,250 of contribution lost each month, from one month’s failures, before any refund or reship. At $12 of direct cost per failure, the direct cost is $7,200 a month. The hidden cost is almost a quarter of the total, and it’s the quarter nobody reports Derived.
With the defaults, each failure costs $15.75, 24% of it in lost repeat contribution, and one point off the failure rate is worth about $18,900 a year. The 20% drop is a placeholder; replace it from your failure cohort, since it decides the answer. For scale, the Uber study’s 5% to 10% fall in spending followed a car ride that arrived late; a damaged or lost order is a bigger failure than that.
FedEx counted every kind of failure, weighted by how much it hurt, in one daily number. The idea fits a DTC brand almost unchanged.
An on-time rate of 96% sounds good and tells you almost nothing. It hides the 4% that were late and treats a day-late package the same as one that never arrived. It says nothing about damage, wrong items or the customer who had to write in twice. A team told to raise it will raise it, sometimes by shipping faster and packing worse.
According to a Federal Highway Administration review of its methods, until 1989 FedEx assumed that on-time delivery was what its customers valued most, and customer research showed they expected much more Reported. So it built the Service Quality Indicator: twelve kinds of failure, each counted every day and multiplied by a weight that reflected how much it hurt customer satisfaction. When FedEx won the Malcolm Baldrige National Quality Award in 1990, the award profile noted that SQI reports went daily to workers at every site, management met daily to discuss the previous day’s performance, and executives were evaluated on the SQI Published.
| Failure | Weight |
|---|---|
| Right day, late | 1 |
| Wrong day, late | 5 |
| Traces not answered | 1 |
| Complaints reopened by customers | 5 |
| Missing proofs of delivery | 1 |
| Invoice adjustments requested | 1 |
| Missed pickups | 10 |
| Lost packages | 10 |
| Damaged packages | 10 |
| Aircraft delay, in minutes | 5 |
| Overgoods (packages that lost their labels) | 5 |
| Abandoned calls | 1 |
ReportedThe weights as given in case materials on Federal Express. A version cited by the Federal Highway Administration, from a 2003 article in California CPA, lists an international indicator in place of aircraft delay. The structure, twelve items weighted by their effect on satisfaction, is from the 1990 Baldrige award profile.
Two things are worth copying. The weights are blunt, 1, 5 and 10, not decimals from a regression. And two items are about the service around the delivery: a complaint the customer had to reopen, and an abandoned call. FedEx counted the customer having to chase as a failure in its own right.
A raw count treats a lost package like a late one. A weighted count tells the team which failure to fix first.
Here’s the index I’d start on, with FedEx’s 1, 5, 10 scale until your failure cohort from chapter 2 lets you set each weight in proportion to what that failure costs.
| Failure | Counted when | Starting weight |
|---|---|---|
| Late, 1 to 2 days | Up to two days past the date shown | 1 |
| Late, 3 days or more | Three or more days past it, or past a date that mattered | 5 |
| Split without warning | Arrived in unannounced pieces | 1 |
| Wrong or missing item | Any line wrong or short | 5 |
| Damaged | Arrived unusable | 10 |
| Lost | Never received | 10 |
| Contact reopened | The customer had to come back about it | 5 |
Report it weekly as points per 1,000 orders, beside the perfect-order rate: the share of orders with no failure at all. Show the board the perfect-order rate. Run the operation on the index, because it says where the points come from.
With the defaults, the month scores 1,885 points, or 188.5 per 1,000 orders, with a perfect-order rate of at least 94.0%. The biggest source is orders late by three days or more: 24% of the index from 14% of the failures. The chart shows why the weights matter.
Derived650 failures and 1,885 points from the tool’s example counts and starting weights. Bars are scaled to 60%.
By raw count, the team would spend the quarter on short delays. By points, it looks at long delays and damage first, and the damage fix is often a packaging change costing cents per order.
Speed sells. A missed date costs more than speed earns. Set the checkout date from your own delivery data, at the share of orders you’re willing to be wrong about.
You can improve a delivery promise by making delivery faster, which costs warehouses and carrier upgrades, or by making the promise truer, which costs a little conversion. Most brands spend on the first. The research says the second is where the retention is.
When a US apparel retailer opened a new distribution center that cut delivery times to its western customers, Marshall Fisher, Santiago Gallino and Joseph Xu found online sales rose about 1.45% for each business day saved, from a starting point of seven business days, with a spillover to the retailer’s stores Published. On Alibaba’s Tmall, Vinayak Deshpande and Pradeep Pendem estimated that cutting three-day deliveries to two days would lift average daily sales for third-party sellers by 13.3% Published. In a business-to-business online store, Shin Oblander and Kinshuk Jerath found each day off the promised time lifted demand 1.82%, like a 2.21% discount, though buyers barely reacted to promises under a week Published.
One study separates the promise from the delivery. Ruomeng Cui, Zhikun Lu, Tianshu Sun and Joseph Golden worked with Collage.com, which sells custom photo products, to change the delivery estimates shown on the site while the actual delivery speed stayed the same. A faster promise raised sales and profits. It also raised returns and reduced customer retention Published. The faster date bought orders today with broken promises that cost orders later.
Recall from chapter 2 that lateness hurts repurchase more than earliness helps. And when Nooshin Salari, Sheng Liu and Zuo-Jun Max Shen built a model on JD.com data that forecasts the whole spread of possible delivery times and sets each promise with the cost of being late in mind, their simulations put the sales gain over JD.com’s existing policy at 6.1% Published. A promise set from data, not a template, is worth money in both directions.
A fast promise wins the order. A kept promise wins the next one.
Most checkout promises are set from the typical delivery: “usually 3 to 5 business days.” But the slow tail, not the typical order, generates the contacts. If half your orders arrive in four days and one in ten takes seven or more, a five-day promise breaks far more often than the team thinks. Pick the share you’re willing to deliver late, then promise the date that share implies. I’d start at 90% to 95% on time against the date shown, by region.
With the defaults, a five-day promise on a lane with a four-day median and a seven-day 90th percentile is kept about 70% of the time: roughly 3,000 broken promises a month on 10,000 orders. Hitting 95% would take a nine-day promise, or a tail short enough that one order in ten takes no more than about 4.8 days. Neither is comfortable. The template “3 to 5 business days” hides that choice; the tool puts it on the table.
Shipping price and free-shipping thresholds are covered in The First Offer. This chapter is only about the date.
The worst way for a customer to learn an order is late is by waiting for it. Tell them the day you know, with a new date and a choice. The law already requires most of this.
Every delay gets discovered. The question is who discovers it first. If it’s you, you explain it and offer a choice. If it’s the customer, they find it on a tracking page that hasn’t changed in three days, and the first thing you hear is a complaint.
David Maister’s 1985 essay on waiting set out propositions that have held up: uncertain waits feel longer than known ones, unexplained waits longer than explained ones, and anxiety makes waits seem longer Published. Shirley Taylor’s 1994 study of delayed airline passengers found delays hurt evaluations mainly through anger and uncertainty Published. A delay notice attacks exactly those: it makes the wait known, explains it, and shows someone is in control.
Order matters too. Angela Legg and Kate Sweeny found that people receiving mixed news mostly wanted the bad news first, while the people delivering it had to be nudged to give it that way; hearing it first left recipients less worried Published. The delay goes in the subject line and the first sentence.
Put the bad news in the subject line. Everything after it is the part they’ll actually read.
A caution: I haven’t found a published field experiment isolating the effect of a proactive delay email on ecommerce repurchase. The case rests on the waiting research and the studies in chapter 2. So hold back a random group, as below, and measure it yourself.
Send the notice the day the delay becomes likely, not the day it’s certain. Five triggers catch most delays:
Templates are in Appendix B. This is a service message, not marketing, so it goes to every buyer. Keep promotions out of it, and send by SMS only to customers who gave you a number for order updates.
In the US, the FTC’s Mail, Internet, or Telephone Order Merchandise Rule sets a floor under all of this. It was last amended in 2014 Published. In summary:
Violations can carry civil penalties per violation; the inflation-adjusted maximum is $53,088, as set in January 2025 Published. A delivery date shown at checkout is also a claim under the FTC Act’s ban on deception. States and other countries add their own rules. This is an operator’s summary as of September 2026, not legal advice; have counsel review your preorder, backorder and delay flows.
Published16 CFR Part 435, as amended at 79 FR 55619 (September 17, 2014); FTC, Business Guide to the FTC’s Mail, Internet, or Telephone Order Merchandise Rule. Full references in Appendix C.
Hold back 10% of delayed orders, at random, from the early notice for one month. They still get everything the law requires. Compare contacts per delayed order, cancellations and 90-day repeat rate. If the notice doesn’t cut contacts, rewrite it before you blame the idea.
Customers value work they can see. A tracking page that shows real progress buys patience; a page that says “label created” for three days spends it.
Between the order confirmation and the doorstep, most brands go dark. The customer’s only window is a carrier page written in scan codes. Ryan Buell’s research at Harvard Business School explains why that silence is expensive.
In experiments on travel and dating websites, Buell and Michael Norton found people valued a service more when it showed its work, even if that meant waiting longer for the same result. They called it the labor illusion Published.
In food-service field experiments, Buell, Tami Kim and Chia-Jung Tsay let customers and cooks see each other. Customer-reported quality rose 22.2% and throughput times fell 19.2% Published. And in Boston, with Ethan Porter, Buell found that residents who reported a problem through the city’s app and got a photograph of the fix submitted 60% more requests afterward; people shown the government’s work were 14% more trusting and 12% more supportive of it Published. It worked because the city was actually responding. Transparency shows the work; it can’t replace it.
Transparency multiplies whatever it shows. Show real work and it earns patience. Show a stalled order and it earns a contact.
Transparency is not a stream of notifications. Send a message only when it changes what the customer expects or answers a question they’d ask. Four emails for an on-time order train customers to ignore the one about the delay. How the flows are built is in The Whole Machine.
Support teams are usually measured on how fast they answer. The number that matters is how often customers need to ask, and why.
Response times and handle times measure how well support copes with contacts. They say nothing about why the contacts exist. The best service, as Bill Price and David Jaffe titled their book, is no service: the contact that never has to happen.
Price was Amazon’s first global vice president of customer service. The metric his team put in place was contacts per order: every call, email and chat, divided by orders. By his own account, it fell 70% during his almost three years at the company Reported. Amazon kept watching it after he left. Jeff Bezos’s letter to shareholders for 2002 called it “our most sensitive measure of customer satisfaction” and reported a 13% improvement that year Reported.
It’s sensitive because almost nobody contacts a store for fun. A contact means something didn’t work. Count contacts per order, split them by reason, and you have a map of every broken promise in the business, drawn by the customers who hit them.
Every contact is a customer telling you, at your expense, where the business broke a promise.
Price and Jaffe’s seven principles, in their order: eliminate dumb contacts; create engaging self-service; be proactive; make it easy to contact the company; own the actions across the company; listen and act; deliver great service experiences Reported. The order matters. Most brands start at the second, with a chatbot, which moves contacts to software without asking why they exist. Start at the first. Chapter 5 applies the third to delivery, and chapter 8 the fifth.
Most help desks have forty tags, applied inconsistently, several per ticket. What works:
A starting list is in Appendix B.
Say a brand ships 10,000 orders a month and receives 1,800 contacts, four in ten about where an order is. With those defaults, contacts per order falls from 0.180 to 0.128 and about 70 agent hours a month are freed, worth about $25,000 a year in agent time alone Derived. That’s modest next to the cost of the failures behind the contacts, which is the point: the support saving is the smaller half. Replace the made-up shares with your own reason counts before quoting it.
A help-center article that answers “where is my order?” deflects the contact. The customer still had the worry and still found your promise broken. Report contacts per order and self-service sessions per order side by side. If the second rises while the first falls, you’ve moved the problem, not solved it.
Support answers the contact. The team that caused it has to remove it. Put a name, a target and a date next to every reason code.
Most contact reasons are decisions made somewhere else: the date by ecommerce, the packaging by operations, the discount rule by marketing. Support meets the consequences and has no power to change them. So each reason gets an owner in the team that causes it, measured on that reason’s contacts per order.
| Reason | Usual cause | Owner | First fix |
|---|---|---|---|
| Where is my order | Silence after checkout; a vague date | Ecommerce, with operations | A dated promise, an owned tracking page |
| Late against the date shown | A promise the lane can’t keep | Operations | Promise by region; delay notice |
| Damaged | Packaging and handling | Operations and packaging | Drop-test the top five products |
| Wrong or missing item | Picking errors | Fulfillment or the 3PL | Scan-to-verify at pack |
| Change or cancel an order | No self-serve edit window | Ecommerce | An edit window on the order page |
| Discount or code didn’t work | Unclear rules | Marketing | One rule per offer, printed with the code |
| Product question before buying | The page doesn’t answer it | Merchandising | Add the top three questions to the page |
| How to use the product | Missing instructions | Product | An insert and a first-use email |
| Unexpected charge | Descriptor, split charges, renewals | Finance or retention | A clear descriptor; renewal reminders |
| Return or exchange | Fit, expectations | Merchandising and ops | See the returns guide |
Subscription contacts belong to whoever runs the program (The Standing Order); product-page questions overlap with The Honest Test; first-use questions with the first-use guide.
Support should own the answer. The team that caused the question should own the number.
Once a month, for forty-five minutes, each owner brings one slide: their reasons’ contacts per order for three months, one fix shipped with its date and expected drop, and one planned. Support reads ten verbatim contacts from the top reason aloud. The next review checks whether each fix moved its number. A fix that didn’t isn’t a fix.
Owners move faster when a contact has a cost in their own budget. Put a notional price on each contact reason: the agent time from the tool in chapter 7, plus, for failures, the repeat contribution from chapter 2. Say the damaged reason runs 150 contacts a month; at $4 of agent time and $3.75 of lost repeat contribution each, that’s about $1,160 a month charged to packaging. A sturdier box at $0.15 more on 10,000 orders costs $1,500. Now the owner has a real decision to make, with both sides priced.
Making things easy for customers is the right goal. The famous single-question survey built to measure it predicts retention poorly. Measure effort by what customers had to do.
In 2010 Matthew Dixon, Karen Freeman and Nick Toman published “Stop Trying to Delight Your Customers” in Harvard Business Review. Its argument has held up. Its metric hasn’t.
From a study of more than 75,000 people who had used contact centers or self-service, the authors concluded that going above and beyond made little difference to loyalty; customers wanted their problem solved simply. Among customers who reported low effort, 94% said they intended to buy again, 88% that they’d spend more, and only 1% that they’d speak badly of the company Reported. Their prescriptions: head off the next likely problem, stop making customers switch channels, attend to how the customer feels, learn from those who struggled, and put solving ahead of speed. A customer who writes twice or repeats their order number to a second agent has done work that was yours to do.
The article proposed a Customer Effort Score, one survey question on how much effort the customer spent, reported to predict loyalty better than satisfaction or the Net Promoter Score. That claim failed an independent test. Evert de Haan, Peter Verhoef and Thorsten Wiesel compared the three metrics on customers of 93 firms in 18 industries and tracked who stayed. The effort score “in itself has little to no predictive power and performs the worst of all” the metrics studied, even among customers who had actually contacted the company. Top-two-box satisfaction, the share who gave one of the two highest scores, predicted retention best Published.
Reduce effort as a design rule. Don’t adopt a one-question effort score as your loyalty metric.
The authors’ caution is about scope: the effort question asks about one past interaction, while retention depends on the whole relationship, including the delivery that caused the contact. The contact is a symptom; its effort is a symptom of the symptom.
Your help desk already records effort:
| Measure | Defined as | What it catches |
|---|---|---|
| Repeat contact rate | Contacts about the same order within 7 days of a previous one, over all contacts | First answers that didn’t resolve it |
| Contacts per resolution | Contacts, over issues marked solved | Back-and-forth to get one thing done |
| Channel switches | Issues touched in more than one channel, over all issues | Chat that sends people to email, email that sends them to phone |
| Transfers | Contacts handed to a second person, over all contacts | Agents without the authority to finish |
Repeat contacts are FedEx’s “complaints reopened,” weighted 5 in its index and in the one in chapter 3. If you survey after contacts, report the top-two-box share, which predicted retention best in de Haan’s study.
A great recovery can leave a customer more satisfied than no failure at all. It doesn’t leave them more likely to buy again. Plan on recovery softening the loss, never reversing it.
Every support team has a story about the customer whose order went wrong, got an extraordinary fix, and became the brand’s biggest fan. The research calls it the service recovery paradox, and it’s used to justify treating failures as opportunities. The evidence says treat them as costs.
In 2007, Celso de Matos, Jorge Henrique and Carlos Rossi pooled the studies that had tested the paradox. The effect was real for satisfaction: on average, customers who had a failure and a good recovery rated their satisfaction higher than customers who had no failure. For repurchase intentions, word of mouth and the company’s image, the pooled effect was not significant: no evidence of a paradox at all Published. Even the satisfaction effect varied with study design, such as whether subjects were students.
The field studies since then point the same way. Stefan Michel and Matthew Meuter tested the paradox with more than 11,000 interviews with bank customers about real service encounters. It was a rare event, and where it appeared the differences were small Published. And the restaurant study in chapter 2, which followed behavior over years, found recovered customers drifting down toward those whose problems were never fixed Published.
Recovery raises how customers feel about the fix. It doesn’t raise how often they come back above what they’d have done if nothing broke.
If you believed the paradox, you’d spend freely on recovery and worry less about prevention. Since it doesn’t hold for repurchase, the math runs the other way: recovery wins back part of the repeat contribution a failure puts at risk, and prevention protects all of it. Set recovery spending against the lost repeat contribution per failure from the cost tool, and prevention spending against the whole cost. Recovery isn’t optional; an unrecovered failure is worse in every study above. It’s something to do well, at a known cost, not something to celebrate.
What each failure gets, written down, so recovery doesn’t depend on who picks up or how loudly the customer complains. Plus a budget each agent can spend without asking.
Without a grid, recovery is set by things that shouldn’t matter: which agent answers and how hard the customer pushes. The polite customer with a damaged order gets less than the furious one with a late order. The grid fixes that; the budget handles what it can’t foresee.
| Failure | Fix | Remedy, first failure | Remedy, repeat within 12 months |
|---|---|---|---|
| Late 1 to 2 days, notified first | New date | None beyond the notice | Shipping refunded |
| Late 1 to 2 days, not notified | New date, apology | Shipping refunded | Credit of about 10% of the order |
| Late 3+ days, or past a date that mattered | Upgrade or reship if faster | Credit of about 10% of the order | Credit of about 20%, from a person |
| Split without warning | Say what’s coming, when | None unless late | Shipping refunded |
| Wrong or missing item | Ship the right item now | Credit of about 10% | Credit of about 20% |
| Damaged | Replace now | Credit of about 10% | Refund the item too |
| Lost | Replace now, before the carrier claim | Credit of about 15% | Refund the order too |
These are starting points, not findings. Set them so the expected remedy per failure sits below the lost repeat contribution per failure from chapter 2. A credit toward the next order gives a reason to come back, and the Uber study found apologies worked best with a future-trip credit attached. Keep credits tied to failures, or customers learn to complain for them; The First Offer covers how discounts train customers.
A grid can’t foresee the gift that arrived after the birthday or the third failure in a month. For those, give each agent a monthly budget to spend without asking anyone.
Size it from the value at stake. Say an agent handles 150 failure contacts a month, each putting $3.75 of repeat contribution at risk, as in the example in chapter 2: about $560 a month. A budget of a fifth to a third of that, $110 to $190, covers the cases that matter most Derived.
A written grid makes recovery fair. A budget makes it human. You need both.
A field experiment with 1.5 million Uber riders tested apologies for late trips. Words alone did little. A credit helped. Apologizing again and again made things worse.
Most apology templates are written by instinct: say sorry, say it warmly, promise it won’t happen again. A very large field experiment suggests the instinct is wrong about the last part.
Basil Halperin, Benjamin Ho, John List and Ian Muir worked with Uber on riders whose trips arrived later than the app had estimated. Riders were randomly assigned to messages like these, some carrying a $5 credit toward a future ride, or to no message at all Published:
| Type | The message |
|---|---|
| Basic apology | “Your trip took longer than we estimated, and we know that’s not ok.” |
| Status apology | “We underestimated how long your trip would take, and that’s our fault.” |
| Commitment apology | “We’re working hard to give you arrival times that you can count on.” |
| No apology | “You have places to go and people to see. Enjoy $5 off your next ride.” |
PublishedExcerpts from each message, as reported in the paper. Outcome: riders’ net spending over the next twelve weeks.
Three findings matter. First, an apology in words alone had little effect, and was sometimes counterproductive. Second, “money speaks louder than words”: the messages that came with a credit did best, and the credit alone raised riders’ net spending by about 1.5% over the first week and about 0.8% over twelve weeks. Third, apologizing repeatedly to the same rider after repeated bad trips reduced their future spending. By the third late trip, an apology with a credit had a significantly negative effect Published. The authors’ advice: use apologies sparingly, and ideally only after outcomes that were unexpectedly bad.
The commitment apology deserves its own warning. The paper names apologies that promise to do better, with repeated ones, as cases where apologizing can be worse than sending nothing Published. A promise to improve is another promise, and a rider who’s late again has now seen two broken.
An apology is a promise about the future. Don’t make one you haven’t already kept.
Rule 6 needs the failure history that the grid in chapter 11 records. Templates are in Appendix B.
In 1995 Continental Airlines promised a monthly bonus to every hourly employee, all 35,000 of them, if the company hit one shared goal. Theory said a bonus shared that widely shouldn’t work. It did.
When Gordon Bethune became chief executive of Continental in October 1994, the airline ranked last in most of the industry’s performance measures Reported. He told the story of the turnaround in his book From Worst to First. The best evidence on why one part of it worked comes from two economists who studied it afterward.
In February 1995, Continental introduced an incentive scheme that promised monthly bonuses to all 35,000 of its hourly employees in any month the company achieved a firm-wide performance goal Published. It was one goal for the whole company, not a goal for the gate agents or the mechanics: one number for everybody, paid to everybody.
On paper, that shouldn’t work. With 35,000 people sharing a goal, one employee’s effort barely affects whether the bonus is paid, so theory predicts people will coast: the free-rider problem.
Marc Knez and Duncan Simester studied the scheme and found it did raise employee performance. Their explanation: employees worked in autonomous groups, like the crew at one airport, where people could see each other’s work and pressed each other to do their part. They call it mutual monitoring. The firm-wide bonus gave every group the same reason to care, and the small group gave each person someone watching Published. Continental went on from last in most performance categories to winning more J.D. Power customer satisfaction awards than any other airline Reported.
Pay everyone on the same kept promise, in teams small enough to see each other work.
Say 25 people each get $75 in any month the perfect-order rate beats target: $22,500 a year if it’s hit every month. At the cost tool’s default, where each point of failure rate costs about $18,900 a year, the bonus pays for itself if it takes about 1.2 points off the failure rate, say from 6% to 4.8% Derived. Publish the number daily where the team can see it, the way FedEx sent its SQI to every site.
One page, every week. Each number with its denominator, split by the lane where it’s worst, and one quarterly line that says what failures cost in repeat orders.
Delivery usually shows up in the weekly meeting as shipping cost per order, and support as response time. Neither says whether the business kept its promises. This page does.
| Number | Defined as | What it catches |
|---|---|---|
| Perfect-order rate | Orders with no failure, over orders delivered | The headline |
| Failure index | Points per 1,000 orders, top two sources, worst lane | Which failure to fix first |
| On time against the date shown | Delivered by the checkout date, by region | Promises the lane can’t keep |
| Ship on time | With the carrier by the promised ship date | Warehouse delays, early |
| Order-to-door days | Median and 90th percentile, by region | A lengthening tail |
| Told first | Delayed orders notified before any contact | Delay triggers that stopped firing |
| Contacts per order | Total, and for each of the top five reasons | A new reason appearing, a fix that didn’t hold |
| Repeat contact rate | Same-order contacts within 7 days | Effort: answers that didn’t resolve |
| Delivery check “no” rate | “No” answers to the delivery check | Damage and errors customers didn’t report |
| Recovery spend | Remedies and budget use per failure | Drift in what failures cost to fix |
| Repurchase after failure (quarterly) | 12-month repeat rate by failure type, against no failure | What each failure costs in the orders that follow |
A number without its denominator is a mood, and an average without its worst lane is an alibi.
First, every rate shows its worst lane beside it: the carrier, region or product where it’s worst. Customers on that lane don’t experience the average. Second, watch the “told first” line. A delay trigger that fires zero times in a week isn’t a good week; it’s usually a broken integration.
If the perfect-order rate improves while the index worsens, fewer orders fail but those that do fail worse. Contacts lag failures by about a week. When you change something, write the date on the page.
Measure, then promise, then route the contacts, then write down how to make it right. Four weeks, in that order.
The order of work is the same whoever you are. Find out which promises you’re breaking. Stop making the ones you can’t keep. Send each contact back to its cause. Then make recovery consistent.
At day thirty you won’t yet know what the changes did to repeat orders; that takes a quarter. You’ll have a store that knows which promises it breaks, tells customers first, and knows who owns each reason they write in.
Measure before you promise. Promise before you apologize.
Six things whoever owns the promise needs on the first day.
Whoever owns delivery reliability and support, a new operations lead or you on the Monday you decide the ticket queue is telling you something, needs six things on day one.
The books and papers worth reading next, and what to take from each.
The research behind each chapter is listed in Appendix C.
Andrew Lauchner runs Growth Legend, embedding inside consumer brands to own lifecycle, email and SMS, and revenue operations. He is the author of The Second Order, on turning first-time buyers into second-time buyers, and The Whole Machine, on the fundamentals of DTC growth, along with a series of field guides for DTC operators at andrewlauchner.com.
As Senior Director of Growth and Retention Marketing at Gallery Furniture, he rebuilt the customer journey and the sales playbooks together. He has worked on growth and retention at Binance and 3Commas, and has been Head of Growth and Retention at Greatness Wins and at Nexus Agriscience.
“Andrew led retention, lifecycle, and email/SMS, but what separates him from most in this space is how deeply he understands the role retention plays in the overall growth engine.”
Akram Khan, Head of Marketing at Gallery Furniture, senior to Andrew but didn’t manage Andrew directly
Andrew answers every note from operators working on this, including those looking for someone to own it. Write to andrew@growthlegend.com or message him on LinkedIn.
The formulas behind the four calculators, and four queries every store should be able to run.
| For | Formula | Notes |
|---|---|---|
| Cost per failure | r × d × n × c + x | r: repeat rate, no failure. d: relative drop. n: orders per returner. c: contribution. x: direct cost. |
| Failure index | Σ (counti × weighti) / orders × 1,000 | Perfect-order rate = 1 − failed orders / orders. |
| On time at a promise of D days | Φ((ln D − ln m) / σ), σ = (ln p90 − ln m) / 1.2816 | m: median days. p90: 90th percentile. Lognormal fit; Φ is the standard normal CDF. |
| Days to promise for target t | ⌈ m × ezt σ ⌉ | zt: normal quantile for t (1.645 for 95%). |
| Contacts avoided | C × w × r + C × (1 − w) × x | C: contacts. w: status share. r: share removed. x: other share removed. |
Check the lognormal fit against last month’s actual on-time share.
-- share of delivered orders that arrived by the date shown at checkout
-- orders.promised_delivery_date must be stored when the order is placed
SELECT o.shipping_region,
f.carrier_service,
COUNT(*) AS delivered,
AVG(CASE WHEN f.delivered_at::date <= o.promised_delivery_date
THEN 1.0 ELSE 0 END) AS on_time_share,
PERCENTILE_CONT(0.5) WITHIN GROUP
(ORDER BY EXTRACT(EPOCH FROM f.delivered_at - o.created_at) / 86400) AS median_days,
PERCENTILE_CONT(0.9) WITHIN GROUP
(ORDER BY EXTRACT(EPOCH FROM f.delivered_at - o.created_at) / 86400) AS p90_days
FROM orders o
JOIN fulfillments f ON f.order_id = o.id
WHERE f.delivered_at >= CURRENT_DATE - INTERVAL '30 days'
GROUP BY o.shipping_region, f.carrier_service
ORDER BY on_time_share;
The syntax is Postgres. If you don’t store the promised date, start today; until then, reconstruct it from the shipping method and order time, and call the result an estimate. The median and 90th-percentile columns feed the date tool in chapter 4.
-- one row per order with its worst failure, for the index and perfect-order rate
WITH flags AS (
SELECT o.id AS order_id,
CASE WHEN f.delivered_at IS NULL AND f.lost_at IS NOT NULL THEN 'lost'
WHEN EXISTS (SELECT 1 FROM tickets t WHERE t.order_id = o.id
AND t.reason = 'damaged') THEN 'damaged'
WHEN EXISTS (SELECT 1 FROM tickets t WHERE t.order_id = o.id
AND t.reason = 'wrong_or_missing') THEN 'wrong_or_missing'
WHEN f.delivered_at::date >= o.promised_delivery_date + 3 THEN 'late_3plus'
WHEN f.delivered_at::date > o.promised_delivery_date THEN 'late_1_2'
WHEN o.split_unannounced THEN 'split'
ELSE 'none' END AS failure
FROM orders o JOIN fulfillments f ON f.order_id = o.id
WHERE o.created_at >= DATE_TRUNC('month', CURRENT_DATE) - INTERVAL '1 month'
AND o.created_at < DATE_TRUNC('month', CURRENT_DATE)
)
SELECT failure, COUNT(*) AS orders
FROM flags GROUP BY failure;
Count reopened contacts separately and add them in the tool. Each order counts once, at its worst failure, which keeps the perfect-order rate honest.
-- 12-month repeat rate of first-time customers, by what happened to the first order
WITH firsts AS (
SELECT DISTINCT ON (customer_id) customer_id, id AS order_id, created_at,
shipping_region, first_product_id
FROM orders ORDER BY customer_id, created_at
)
SELECT DATE_TRUNC('month', fo.created_at) AS cohort_month,
fo.shipping_region,
fl.failure,
COUNT(*) AS customers,
AVG(CASE WHEN EXISTS (
SELECT 1 FROM orders o2
WHERE o2.customer_id = fo.customer_id
AND o2.created_at > fo.created_at
AND o2.created_at <= fo.created_at + INTERVAL '12 months')
THEN 1.0 ELSE 0 END) AS repeat_12m
FROM firsts fo
JOIN flags fl ON fl.order_id = fo.order_id -- the flags logic above, run for these orders
WHERE fo.created_at < CURRENT_DATE - INTERVAL '12 months'
GROUP BY 1, 2, 3
ORDER BY 1, 2, 3;
Compare each failure group with the ‘none’ group within the same month and region, then average the gaps, weighted by group size. Add whether a remedy was given to get the comparison in chapter 10. Pool quarters until groups reach a few hundred.
-- weekly contacts per order, by primary reason
SELECT DATE_TRUNC('week', t.created_at) AS week,
t.reason,
COUNT(*)::numeric / NULLIF(w.orders, 0) AS contacts_per_order
FROM tickets t
JOIN (SELECT DATE_TRUNC('week', created_at) AS week, COUNT(*) AS orders
FROM orders GROUP BY 1) w
ON w.week = DATE_TRUNC('week', t.created_at)
WHERE t.created_at >= CURRENT_DATE - INTERVAL '12 weeks'
AND t.direction = 'inbound'
GROUP BY 1, 2, w.orders
ORDER BY 1, 3 DESC;
Table and column names are generic; rename them to match your exports. Count conversations, not messages. Repeat contact rate follows the same pattern: contacts with an earlier contact on the same order in the previous 7 days, over all contacts.
Delay notices, apologies, reason codes and the delivery check. Copy them into whatever your team already uses.
These messages are transactional: about an order the customer placed, sent to every buyer, with no promotion. Keep your standard footer with your address and a preferences link. Have counsel review the option notice.
SUBJECT Your Order Is Running Late
PREVIEW new date inside, and your options
Hi [first name],
Your order [#1234] is running late. It should now arrive
by [Thursday, October 8], instead of [Monday, October 5].
What happened: [the carrier missed its pickup from our
warehouse on Tuesday].
What you can do:
Wait for it Nothing to do. We'll write again if
the date moves.
Cancel for a refund [Cancel order] (one tap, full refund
within 7 working days)
Swap it [Choose something in stock]
[Only if the recovery grid calls for it:]
We've added a [$X] credit to your account for your next order.
[Name], [Brand] customer care
Reply to this email to reach a person.
[Brand, postal address] | Email preferences
[Brand]: Your order #1234 is running late. New date: Thu Oct 8. Wait, cancel for a refund, or swap: [short link]. We're sorry for the delay. Reply STOP to opt out.
Send by SMS only to customers who opted in to order updates by text.
SUBJECT We Can't Ship Your Order On Time PREVIEW please choose: wait or cancel Hi [first name], We can't ship [item] from order [#1234] by [the date we promised]. We now expect to ship it by [revised date]. [Or: We don't know yet when we can ship it.] You can: Wait [I'll wait] Cancel [Cancel for a full refund] If you cancel, we'll refund [$X] within [7 working days]. [If the delay is 30 days or less and this is the first delay:] If we don't hear from you, we'll ship it when it's ready. [If longer, indefinite, or a second delay:] If we don't hear from you by [date], we'll cancel this item and refund you in full. [Brand, postal address] | Email preferences
SUBJECT Your Order Arrived Damaged. A Replacement Is On Its Way. PREVIEW shipped this morning, arriving by [date] Hi [first name], Your [item] arrived damaged. That's on us. A replacement shipped this morning and should arrive by [date]. You don't need to send anything back. We've also added a [$X] credit to your account for your next order. [What we've changed, only if it's true: We've moved this item to a sturdier box.] [Name], [Brand] customer care [Brand, postal address] | Email preferences
WHERE IS MY ORDER no update, customer asking LATE AGAINST DATE SHOWN past the checkout date DAMAGED arrived unusable WRONG OR MISSING ITEM any line wrong, short or absent LOST never arrived, or marked delivered CHANGE OR CANCEL ORDER after checkout, before shipping DISCOUNT OR CODE didn't apply, or unclear PRODUCT QUESTION, BEFORE pre-purchase question PRODUCT USE, AFTER how to use, results, fit BILLING unexpected or unclear charge SUBSCRIPTION skip, change, cancel RETURN OR EXCHANGE start, status, refund ACCOUNT login, details OTHER review monthly; split anything over 5%
IN THE DELIVERY CONFIRMATION EMAIL "Did everything arrive as expected?" [Yes] [No, something's wrong] "No" opens a short form: damaged / wrong or missing / not received / other, with an optional photo. Every "no" gets a same-day reply and a grid remedy.
Every external source, by chapter. Web sources were read in September 2026.