FedEx counted every kind of failure, weighted by how much it hurt, in one daily number. The idea fits a DTC brand almost unchanged.
An on-time rate of 96% sounds good and tells you almost nothing. It hides the 4% that were late and treats a day-late package the same as one that never arrived. It says nothing about damage, wrong items or the customer who had to write in twice. A team told to raise it will raise it, sometimes by shipping faster and packing worse.
According to a Federal Highway Administration review of its methods, until 1989 FedEx assumed that on-time delivery was what its customers valued most, and customer research showed they expected much more Reported. So it built the Service Quality Indicator: twelve kinds of failure, each counted every day and multiplied by a weight that reflected how much it hurt customer satisfaction. When FedEx won the Malcolm Baldrige National Quality Award in 1990, the award profile noted that SQI reports went daily to workers at every site, management met daily to discuss the previous day’s performance, and executives were evaluated on the SQI Published.
| Failure | Weight |
|---|---|
| Right day, late | 1 |
| Wrong day, late | 5 |
| Traces not answered | 1 |
| Complaints reopened by customers | 5 |
| Missing proofs of delivery | 1 |
| Invoice adjustments requested | 1 |
| Missed pickups | 10 |
| Lost packages | 10 |
| Damaged packages | 10 |
| Aircraft delay, in minutes | 5 |
| Overgoods (packages that lost their labels) | 5 |
| Abandoned calls | 1 |
ReportedThe weights as given in case materials on Federal Express. A version cited by the Federal Highway Administration, from a 2003 article in California CPA, lists an international indicator in place of aircraft delay. The structure, twelve items weighted by their effect on satisfaction, is from the 1990 Baldrige award profile.
Two things are worth copying. The weights are blunt, 1, 5 and 10, not decimals from a regression. And two items are about the service around the delivery: a complaint the customer had to reopen, and an abandoned call. FedEx counted the customer having to chase as a failure in its own right.
A raw count treats a lost package like a late one. A weighted count tells the team which failure to fix first.
Here’s the index I’d start on, with FedEx’s 1, 5, 10 scale until your failure cohort from chapter 2 lets you set each weight in proportion to what that failure costs.
| Failure | Counted when | Starting weight |
|---|---|---|
| Late, 1 to 2 days | Up to two days past the date shown | 1 |
| Late, 3 days or more | Three or more days past it, or past a date that mattered | 5 |
| Split without warning | Arrived in unannounced pieces | 1 |
| Wrong or missing item | Any line wrong or short | 5 |
| Damaged | Arrived unusable | 10 |
| Lost | Never received | 10 |
| Contact reopened | The customer had to come back about it | 5 |
Report it weekly as points per 1,000 orders, beside the perfect-order rate: the share of orders with no failure at all. Show the board the perfect-order rate. Run the operation on the index, because it says where the points come from.
With the defaults, the month scores 1,885 points, or 188.5 per 1,000 orders, with a perfect-order rate of at least 94.0%. The biggest source is orders late by three days or more: 24% of the index from 14% of the failures. The chart shows why the weights matter.
Derived650 failures and 1,885 points from the tool’s example counts and starting weights. Bars are scaled to 60%.
By raw count, the team would spend the quarter on short delays. By points, it looks at long delays and damage first, and the damage fix is often a packaging change costing cents per order.
This is one chapter of The Kept Promise, which is free and readable in full on a single page with no form in front of it.