Reviews, claims and what customers believe, and the evidence behind every word you publish.
Almost everything a DTC brand says about itself is something the customer can’t check before paying. Made in Vermont. Clinically tested. Handmade. Loved by 40,000 customers. Shoppers know this, so they discount the words and look for proof, and the proof they find first is usually your reviews.
Those two numbers are the guide in miniature. The first says review scores aren’t a clean measure of what customers think: one early vote, assigned by a coin flip, moved where ratings ended up. The second says the old shortcuts, buying a few reviews, holding back the bad ones, rounding a claim up, now carry a price per instance.
So this guide does two jobs. It shows, from the research, how reviews and claims shape belief. And it shows how to build the record that lets you say true things with confidence: a proof file with one row for every claim on your site and in your email.
A claim without a file behind it is a bet that nobody asks.
It builds on The Honest Test, on the product page, and The Whole Machine, on the email flows the review ask lives in.
Start with The Proof Audit, or with the map just below, which shows every kind of claim and what proof it needs. Or follow a path:
Three tools run in the page. Nothing you type leaves your browser.
Examples that open with Say or Picture use made-up round numbers. Every source is listed in Appendix C. Where the law comes up, it’s dated and it isn’t legal advice: have counsel review anything you’ll publish.
What this guide argues, and what would prove each claim wrong.
A position says what would prove it wrong. Test each on your own store.
Find the kind of claim you’re making. The row tells you whether the customer can check it, what belongs in your file, and which rules apply.
The further down a claim sits, the less the customer can check it, and the more your reviews, your file and regulators have to say about it.
The whole book is above and always will be. These are the same chapters addressed individually, for linking to one idea rather than to ninety.
| Kind of claim | Can the buyer check it? | What goes in the file | Rules that bite (US, EU, UK) |
|---|---|---|---|
| Product facts: size, material, weight, count | Before buying, mostly | Spec sheet, supplier certificate, your own measurement | General deception law |
| Performance: lasts 12 hours, fits true to size | Only after use | Your test method and results, dated; return and review data | FTC substantiation policy (1984) |
| Social proof: 4.8 stars, 40,000 customers | Partly, by reading reviews | Review export, how the average is computed, order counts | FTC review rule (2024); EU review rules (2022); UK DMCC Act (2025) |
| Testimonials, influencers, expert quotes | No | Signed consent, any payment or free product, typicality evidence | FTC Endorsement Guides (2023); FTC review rule |
| Origin: Made in USA, made in Vermont | No | Bill of materials with country of each input, where each step happens | FTC Made in USA rule (2021) |
| Process: handmade, small batch, family owned | No | Who makes it, where, how, and the share done by hand | General deception law |
| Health and efficacy: clinically shown, supports sleep | No | Studies on your formula or its active doses, ideally randomized trials | FTC health products guidance (2022) |
| Environmental: recyclable, eco, carbon neutral | No | Certifications, recycling access data, emissions accounts | FTC Green Guides (2012); EU Directive 2024/825 from Sept 27, 2026 |
PublishedDates are when each rule or guide took effect or was last revised; chapter 11 has the details and sources. Economists have sorted claims this way since Phillip Nelson (1970) and Michael Darby and Edi Karni (1973): search qualities you can check before buying, experience qualities you learn by using, and credence qualities you may never be able to verify.
Most DTC brands put their strongest words in the bottom four rows, because that’s where the differentiation lives. Those are also the rows where customers depend most on your word, and where a regulator is most likely to ask for your file.
Your best claims are the ones customers can’t check. That’s why they need the best proof.
Twelve checks on whether customers have reason to believe what you say, and whether you could show a regulator why. About forty-five minutes with your site, your email platform and your review app.
The audit isn’t about how many reviews you have or how high the stars are. It’s about whether the things customers read about you are true, whether you can show it, and whether your review system reports what customers think or what you’d like them to think.
Open your top product pages, last five campaign emails, post-purchase flow and review app settings. Score each check 0 to 2: 0 if it failed or nobody can answer it, 1 if partly true, 2 if clean.
If nobody can find the proof in a day, you don’t have proof. You have a memory.
Score as you go; your band appears when all twelve are in.
| Score | What it means | Read next |
|---|---|---|
| 20–24 | Customers have reason to believe you, and you could show why. Your job now is speed: more reviews sooner on new products, and a quarterly check of the file. | The Ask, then The Proof Scorecard |
| 14–19 | Most of what you say is probably true, but you couldn’t prove all of it on a deadline. Fix the zeros first. | The chapter linked from your lowest check, then The Proof File |
| 8–13 | Your claims run on memory and your reviews on defaults. Build the file before the next campaign goes out. | Part one, starting at Claims Nobody Can Check |
| 0–7 | Stop adding claims this month. List what you say, pull what you can’t support, and turn off any review filter. | The Rules, Dated, then The First Thirty Days |
Checks 2, 3, 9, 10 and 11 carry legal exposure, not just lost sales. If any of those five scored 0, treat it as this week’s work whatever your band says.
Shoppers sort what you say into what they can see, what they’ll learn by using it, and what they’ll never know. They discount the last two, and they’re right to.
A shopper can see a hoodie’s color. They can’t see whether it will pill after ten washes, and they’ll never see where the cotton was grown. Every claim falls into one of those three bins.
Economists named the bins fifty years ago. Phillip Nelson (1970) split product qualities into search qualities, which you can check before buying, and experience qualities, which you only learn by using. Michael Darby and Edi Karni (1973) added credence qualities, which a buyer may never be able to verify even after use: whether a supplement contains what it says, whether a candle was poured by hand, whether a mailer is recyclable in your town Published.
The theory predicts that buyers trust search claims most, because lying about them is pointless. In 1990 Gary Ford, Darlene Smith and John Swasy tested that directly. Consumers were more skeptical of experience claims than search claims, and more skeptical of subjective claims (“luxurious”) than objective ones (“100% cotton”). The researchers didn’t find the extra step they expected for credence claims over experience claims Published. Read that as: once a shopper can’t check a claim before paying, the discount is already applied. It doesn’t matter much to them whether they’ll find out next week or never.
The discount starts the moment a claim can’t be checked on the page.
They look for someone else who has checked. For most DTC products that means reviews: other buyers with no reason to flatter you. PowerReviews, which sells review software, reports that 95% of shoppers consult reviews before buying Reported (vendor data). The other checkers are certifiers, publications, and your own guarantee, which is a claim you pay for if it’s false. A generous return policy says “we’re confident” in a way copy can’t; the returns guide covers what it costs.
That’s why reviews and claims belong in one book. A claim your reviews contradict (“runs small,” says every third review of your “true to size” jeans) costs you twice: the sale, and belief in your next claim.
The Federal Trade Commission’s 1984 policy statement on substantiation says advertisers need a reasonable basis for a claim before it runs. If an ad says or implies a level of support, like “tests show” or “doctors recommend,” the advertiser must have at least that level Published. For health claims, the FTC’s December 2022 guidance sets the bar at “competent and reliable scientific evidence” and says randomized controlled trials are generally what experts would require. Testimonials, however sincere, are not a substitute Published.
Notice the overlap. The claims customers discount most are the ones regulators scrutinize most. Both are reacting to the same fact: only you know whether it’s true.
You can’t make a credence claim checkable, but you can make it more specific, name who checked it, and show the document. Each step moves it closer to a search claim.
| Instead of | Say | Why it works |
|---|---|---|
| Premium cotton | 100% cotton, 280 gsm, knit in Portugal | Objective, and partly checkable on arrival |
| Eco-friendly packaging | Mailer is 80% recycled paper, certified by [scheme]; recyclable where paper is collected | Specific, names the checker, states the limit |
| Clinically proven | In a 12-week randomized trial of 60 adults, [result]. Study summary linked. | Matches the claim to the study, with the size shown |
| Loved by thousands | 4.6 from 2,314 reviews, including 131 one-star. Read them. | Checkable on the page, and the negatives add credibility |
The figures are illustrations and the brackets are placeholders. Never fill them with a number, certifier or study you don’t have on file.
One row per claim, on the site and in email: what it says, where it runs, what supports it, and who checks it. The cheapest insurance a brand can buy.
Brands rarely lie on purpose. They drift. A founder says “all natural” in an interview, a copywriter repeats it, the formula changes, and three years later nobody knows where the claim came from. The proof file stops the drift.
A spreadsheet is enough. Appendix B has the columns ready to paste.
The column that matters most is the one that says how far the evidence goes.
Say a candle brand’s footer reads “Hand-poured in the USA,” it sends 12 campaigns a month to 150,000 subscribers, and its product pages get 200,000 views a month. That footer claim is delivered about 1.8 million times a month in email alone, before counting the site. Now say the wax blend is imported and the pouring moved to a contract manufacturer last spring, which uses a filling machine for the 16-ounce size. The claim was true when written. For at least one product it no longer is, in two ways: origin and process.
At the review date, the “what the evidence supports” column would read “poured in the USA from imported wax; 8-ounce size poured by hand.” Still a good claim, just not the one in the footer.
The scorer weighs four things: how little the buyer can check, how thin your evidence is, how far the wording outruns it, and whether a specific rule covers the term. It’s my weighting, not a legal test; evidence counts most because you control it.
With the example’s numbers, the “Made in USA” footer scores 73, in the band to pull or qualify now, and the first fix is the evidence: a supplier’s email is not a bill of materials. At two million views a month it would be seen about 24 million times in a year. Score your own top ten claims and fix them in score order.
Three different research designs, on books, restaurants and Amazon products, reach the same answer. Most of the effect comes from the first handful of reviews.
Products with more reviews sell more, but that could be sales causing reviews. Researchers have found three ways around the problem, and all three say reviews cause sales.
In 2006 Judith Chevalier and Dina Mayzlin compared the same books on Amazon and Barnes & Noble’s website. Anything about a book itself, its author, its marketing, its cover, is the same on both sites, so a change in reviews on one site but not the other isolates the effect of the reviews. An improvement in a book’s reviews raised its relative sales on that site. And the effect was lopsided: a one-star review hurt more than a five-star review helped Published.
Reviews were also overwhelmingly positive: 67% of reviews on Barnes & Noble were five stars, and 53% on Amazon Published. Hold that thought for chapter 5.
A one-star review costs more than a five-star review earns. Plan your review system around the bad ones.
Northwestern’s Spiegel Research Center analyzed data from PowerReviews, a review platform, across many retailers. A product with five reviews had a purchase likelihood 270% greater than the same kind of product with none. The benefit of each extra review fell off quickly after the first five, and showing reviews lifted conversion more for higher-priced products Reported (vendor data, analyzed by a university research center).
Treat the exact numbers with care. This is observational data, and products that collect reviews may differ from products that don’t. But it fits the peer-reviewed work: Ewa Maslowska, Edward Malthouse and Vijay Viswanathan found, using data from two retailers, that the rating’s effect on purchase was strongest when a product had many reviews, when the shopper read them, and when the product was expensive Published. For a DTC brand that means the review gap that costs most is the one on new and expensive products.
So coverage beats average: the reviews worth most are the first five on products that have none. Show stars, count and a link near the price and the button; The Honest Test covers the product page.
Not your customers. A small, self-selected slice of them: the delighted, the furious, and a few people who review everything.
Your average rating looks like a survey result. It isn’t. A survey asks everyone; reviews come from whoever chose to write.
Jakob Nielsen called it participation inequality in 2006: in most online communities, about 90% of users only read, 9% contribute occasionally, and 1% contribute most of the content. His example from Amazon: one reviewer alone had written 12,423 book reviews Reported. Clay Shirky’s Here Comes Everybody (2008) makes the same point at book length: contributions follow a steep power law, and the average contributor doesn’t exist.
Plot most products’ reviews by star and you get a J: a tall bar at five stars, a smaller bump at one, and little in between. Nan Hu, Paul Pavlou and Jennifer Zhang found that 78% of book ratings, 73% of DVD ratings and 72% of video ratings on Amazon were four stars or higher Published.
Then they ran the test that settles it. In a controlled experiment, everyone rated the same music CD. Their ratings formed an ordinary hump in the middle. The same CD’s ratings on Amazon formed a J Published. Same product, different shape, because on Amazon nobody was required to rate it.
Two filters explain it. Purchasing bias: people who expect to dislike a product don’t buy it, so never review it. Under-reporting bias: the delighted and the furious “brag or moan”; the mildly satisfied stay quiet. So the average is a poor proxy for quality; look at the spread and the two peaks too Published.
The middle of your customer base is the part that doesn’t write.
Leif Brandes, David Godes and Dina Mayzlin (2022) added a third filter: attrition. Reviewers with moderate experiences are more likely to stop reviewing over time than those with extreme ones. In a large field experiment with an online travel platform, plain review-request emails, with no incentive, made the reviews that came in less extreme Published. Asking everyone brings back the middle. That’s the strongest research case for a systematic ask, and chapter 8 is about how.
Eric Anderson and Duncan Simester found that about 5% of reviews at a large private-label retailer came from customers with no record of buying the product, and those reviews were significantly more negative. Many came from the firm’s most loyal customers Published. Some are gift recipients, not fakes. But it’s a reason to label verified buyers. Spiegel’s analysis of PowerReviews data found a verified-buyer badge raised the odds of purchase by 15% Reported (vendor data).
One random early vote changed where ratings ended up, in a randomized experiment on a hundred thousand comments. Launch like the first reviews matter, because they do.
If people rated only what they saw, the order of reviews wouldn’t matter. It does. People read what came before and adjust, and the first rating sets the direction.
Lev Muchnik, Sinan Aral and Sean Taylor ran a randomized experiment on a social news site like Reddit, published in Science in 2013. Over five months, 101,281 comments were randomly assigned at birth to get one up-vote, one down-vote, or nothing. The votes had nothing to do with the content. The researchers then watched 308,515 ratings from real users Published.
Positive herding compounds; negative herding gets corrected, at least on that site. In Matthew Salganik, Peter Dodds and Duncan Watts’s MusicLab experiment with 14,341 participants, showing what others downloaded made hits bigger and success less predictable. The best songs rarely did badly and the worst rarely did well, “but any other result was possible” Published.
Your first reviews don’t just describe the product. They steer every review that follows.
The first reviews come from whoever buys first, and they aren’t typical. Xinxin Li and Lorin Hitt showed that early buyers’ particular tastes shape later buyers’ decisions and can mislead them Published. Your loyalists produce fans; cold paid traffic produces mismatches. Neither is your month-six customer.
Fake early reviews work briefly, which is why people buy them. He, Hollenbeck and Proserpio found the effect was short-lived: after products stopped buying fake reviews, their ratings fell and their share of one-star reviews rose significantly, especially for young products Published. In 2019 the FTC brought its first case challenging fake paid reviews on an independent retail website: Cure Encapsulations had paid a website to post Amazon reviews of a weight-loss supplement and asked it to keep the product at five stars. The judgment was $12.8 million, suspended on payment of $50,000 Filed. Since October 2024, fake reviews carry civil penalties per violation (chapter 11).
A 4.9 average from 12 reviews looks better than a 4.6 from 840. It isn’t, yet. One more one-star review would pull the first to 4.6, while the second would barely move. The fix is to rank by a score that shrinks small counts toward your store’s typical rating and then asks how low the true average could plausibly be. It’s a simplified version of the Bayesian approach the statistician Evan Miller describes for star ratings Reported.
With the defaults, product A’s 4.9 adjusts to 4.67 with a lowest plausible score of 4.29, while B’s 4.6 adjusts to 4.60 with a floor of 4.54. B is the one to feature. A needs about 27 more reviews at 4.9 to pass it, and one more one-star review would show it at 4.6 Derived. Use the same score to sort your “top rated” collection.
As reviews pile up, each new one tends to come in a little lower. Usually that isn’t the product getting worse. It’s the review system working as designed.
A product launches at 4.8 and a year later sits at 4.5. Often nothing went wrong with the product. The later buyers are different people, reading a longer and messier set of reviews, and they rate differently.
David Godes and José Silva (2012) separated two clocks: how long a product has been on sale, and how many reviews came before a given one. Once they controlled for calendar time, the time pattern actually rose. But the sequence pattern fell: each successive rating tended to be lower than the one before, controlling for the product and the reviewer Published.
Their explanation: as reviews pile up, it gets harder to tell which apply to you, so more people buy something that doesn’t suit them, and rate it lower Published. Two other findings point the same way. Early buyers have unusual tastes, so the first reviews can mislead later buyers (Li and Hitt). And frequent reviewers tend to differentiate themselves from the crowd, pulling ratings down over time, while occasional reviewers follow it (Wendy Moe and David Schweidel) Published.
Drift is mostly mismatch. The fix is helping the wrong buyer notice before they buy.
Most reviews exist because someone asked. When you ask, who you ask and what you offer decide what the reviews say.
PowerReviews, a review platform, reports that 70% of its clients’ reviews come from emails sent after purchase Reported (vendor data). So the review request is not a small flow. It’s the machine that writes most of your social proof, and its settings shape the result.
So: ask everyone, once, with one reminder; use a true social-norm line (“2,314 customers have reviewed this”); and expect a reward to buy more short reviews, not better ones.
A review is only useful if the writer has used the product. Asks timed to fulfillment arrive before the product and produce reviews about the courier, or nothing. The right trigger is delivery plus the time to form a view: days for a T-shirt, weeks for skincare or supplements.
Say a brand ships 3,000 orders a month, delivery takes four days, and most customers have used the product enough to judge it about two weeks after it arrives. Say 8% of customers who have used it answer an ask, 3% of those who haven’t still reply with something about delivery, and interest halves every 30 days after the right moment. The tool shows what the send day does.
With those defaults, an ask on day 7 produces about 51 product reviews a month plus 71 about delivery, so 58% of what comes in says nothing about the product. Moving it to day 18 produces about 240 product reviews a month, and a new product selling 150 a month reaches five reviews in about two weeks instead of two months Derived. Your own rates will differ; the shape usually won’t.
Trigger the ask from delivery plus use, never from the shipping label.
If results take a month to show, as with skincare, send a first-impressions ask at two weeks and a results ask at five, labeled on the review. Text only customers who agreed to texts. Appendix B has both versions.
A perfect score reads as a filter. A few visible negatives, answered well, make the positives believable. And hiding them is now clearly illegal.
Shoppers go looking for the bad reviews. PowerReviews data cited by the Spiegel Research Center says 82% of shoppers specifically seek out negative reviews Reported (vendor data). If they can’t find any, they don’t conclude the product is perfect. They conclude something is being hidden.
In the Spiegel Research Center’s analysis of PowerReviews data, purchase likelihood peaked for products rated between 4.0 and 4.7 stars, then fell as ratings approached 5.0 Reported (vendor data, analyzed by a university research center). The researchers’ reading: negative reviews “establish credibility and authenticity.” It’s observational, so treat the range as a guide.
There’s experimental support. Danit Ein-Gar, Baba Shiv and Zakary Tormala (2012) found in four studies that adding a small dose of negative information to an otherwise positive description made people like a product more, which they called the blemishing effect. The conditions matter: it worked when the negative came after the positive and when people weren’t reading closely Published. A product page fits both. Lead with the summary and the praise; put a real critical review one tap away.
A 4.6 with visible one-stars is more believable than a 5.0 with none.
In January 2022, the FTC announced that the fast-fashion retailer Fashion Nova would pay $4.2 million to settle charges that it blocked negative reviews. From late 2015 through November 2019, according to the FTC, the company used a third-party review interface that automatically posted four- and five-star reviews and held lower-rated ones for approval that never came. Hundreds of thousands of lower-rated reviews were never published. It was the FTC’s first case about a company concealing negative reviews. The same day, the FTC sent letters to ten companies that provide review-management services and published guidance on collecting and showing reviews Filed.
The lesson is in the mechanism. Nobody had to delete a review; a moderation setting did the work. Most review apps have one. Find yours.
You can hold back some reviews, but the rules are narrow and must apply to positive and negative reviews alike:
Write a moderation policy, publish it next to the reviews, apply it to every star level, and log every removal with the reason. Appendix B has a starting draft.
A reply to a one-star review is written for the next hundred shoppers. Thank them, state the fact without arguing, say what you changed, and offer a person to talk to, within two business days. Never offer anything for changing or removing the review.
Where and how a thing was made changes what people will pay for it. That’s why these claims are worth making, and why regulators check them.
Customers can’t check where a product was made or whose hands made it, yet they pay for those facts. Three lines of research explain why; one FTC rule sets the cost of getting them wrong.
George Newman and Ravi Dhar (2014) found that people see products made in a brand’s original factory as carrying the “essence” of the brand, a belief in contagion, and so as more authentic and more valuable than identical products made elsewhere. People more sensitive to contagion showed the effect more strongly, and priming the idea of contagion strengthened it Published. Their studies used everyday branded goods, including Levi’s jeans Reported (Yale School of Management).
The founding workshop and the town on the label carry that essence. That’s real value, and it’s why moving production without changing the story is a problem.
Christoph Fuchs, Martin Schreier and Stijn van Osselaer (2015) found across four studies that “handmade” makes products more attractive, largely because people feel handmade products symbolically contain the maker’s love. In one study, people asked what they would pay for a bar of soap for their mother offered $6.56 when it was described as handmade and $5.63 when machine-made Published, about 17% more Derived. The effect held for gifts to loved ones but not to distant recipients, and was stronger when the goal was to show love rather than to get the best-performing product Published. So “handmade” earns most on gifts, and on the occasions people buy them.
Ryan Buell and Michael Norton (2011) called it the labor illusion. In five experiments with travel and dating websites, people who were shown the work being done on their behalf valued the service more, and could even prefer a longer wait to an instant result, with identical results Published. In later field experiments in food service, Buell, Tami Kim and Chia-Jung Tsay found that when cooks and customers could see each other, customer-reported quality rose 22.2% and service got 19.2% faster Published.
Show the work you really do: sourcing, batch testing, hand-finishing, fit sessions. Shown effort is valued; it just has to be real.
The claims customers value most are about things they can’t see. Keep the records they’d want to see.
The FTC’s Made in USA Labeling Rule took effect on August 13, 2021. It covers labels and mail-order promotional material, including material sent electronically. An unqualified “Made in USA” claim requires that final assembly or processing happen in the US, that all significant processing happen in the US, and that all or virtually all ingredients or components be made and sourced in the US Published. Violations carry civil penalties. A qualified claim, like “Made in USA from imported leather,” is the honest alternative when the inputs come from abroad.
Enforcement is active. From 2021 to 2024 the FTC brought 11 Made in USA actions with $15.75 million in judgments Filed. In July 2025 it sent warning letters to four companies, and to Amazon and Walmart about third-party sellers’ US-origin claims Filed.
In 2020, Williams-Sonoma settled FTC charges that it had falsely advertised several product lines as all or virtually all made in the USA, and agreed to an order requiring truthful origin claims. In April 2024 it agreed to pay a $3.175 million civil penalty for violating that order, which the FTC called the largest ever in a Made in USA case. According to the FTC, mattress pads sold under its Pottery Barn Teen brand were marketed as “Crafted in America from domestic and imported materials” but were made in China, and six other products advertised as made in the USA were also deceptive Filed.
A company with a legal department, under an order naming the exact issue, still let the claim run on products it didn’t fit. Copy outlives the facts. That’s what the proof file is for.
Between 2021 and 2026 the US, the EU and the UK each wrote specific rules for reviews, origin and green claims. Here is what each says, and from when.
Review and claim rules used to say only: don’t deceive. Now they name specific practices and set penalties per violation. This is a dated map as of September 2026, not legal advice; have counsel review your practices against it.
| Rule | Date | What it means for a brand |
|---|---|---|
| FTC Rule on the Use of Consumer Reviews and Testimonials (16 CFR Part 465) | Announced Aug 14, 2024; effective Oct 21, 2024 | Bans fake reviews and testimonials, including AI-generated ones; buying reviews conditioned on sentiment; undisclosed reviews by officers and employees; company-controlled review sites posing as independent; suppressing reviews by unfounded legal threats or intimidation, or implying you show most or all reviews when you’ve suppressed negative ones; and buying fake followers or views. Civil penalties for knowing violations. |
| FTC Endorsement Guides, revised | Approved June 29, 2023 | Don’t procure, suppress, boost, organize or edit reviews in a way that distorts what customers think. Disclose material connections with reviewers and endorsers. |
| Consumer Review Fairness Act | 2016 | Form contracts can’t bar honest reviews, fine customers for them, or claim rights in the review’s content. |
| Made in USA Labeling Rule (16 CFR Part 323) | Effective Aug 13, 2021 | Unqualified US-origin claims on labels and mail-order material must meet the “all or virtually all” standard. Civil penalties. |
| Green Guides (16 CFR Part 260) | Last revised 2012; review opened Dec 2022 | Guidance on “recyclable,” “compostable,” “eco” and similar terms. Still the 2012 version as of September 2026. |
| Health Products Compliance Guidance | Dec 2022 | Health claims need competent and reliable scientific evidence, generally randomized controlled trials. Testimonials don’t substitute. |
PublishedFTC, Federal Register and eCFR texts listed in Appendix C. Green Guides status: no revision in the Federal Register, and none listed by the BWD Strategic tracker, checked September 2026.
The review rule has started to bite. On December 22, 2025, the FTC sent warning letters to ten companies about possible violations, noting penalties of up to $53,088 per violation Filed. That is the 2025 inflation-adjusted cap; the Federal Register showed no 2026 adjustment as of September 2026 Published. Before the rule, in October 2021, it had already sent a Notice of Penalty Offenses on endorsements to more than 700 companies, which makes it easier to seek civil penalties from recipients who later use those practices Filed.
If you sell into the EU, act on the green claims rules now: “carbon neutral shipping” backed by offsets is exactly what the new list bans. A separate EU proposal on substantiating green claims is a different law; ask counsel where it stands.
The Digital Markets, Competition and Consumers Act 2024 lists practices that are unfair in all circumstances, including submitting or commissioning fake reviews, concealing that a review was incentivized, and publishing reviews or review information in a misleading way. Those provisions came into force on April 6, 2025 Published. The Competition and Markets Authority’s guidance, published two days earlier, says businesses that publish reviews must take “reasonable and proportionate steps” to prevent and remove banned content, may use incentives only if the review says so and still reflects a real experience, and must not suppress genuine negative or positive reviews Published. Penalties under the Act can reach £300,000 or 10% of turnover, whichever is higher, and the CMA can now impose fines itself instead of going to court Published.
The question used to be whether a practice was deceptive. Now it’s often whether it’s on a list.
Green claims get enforced too. In January 2022, Keurig Canada agreed to pay a C$3 million penalty to settle the Canadian Competition Bureau’s concerns about its K-Cup pods. The Bureau found that recyclability claims were false or misleading, because municipal recycling programs outside British Columbia and Quebec didn’t accept the pods, and that the instructions for preparing them for recycling didn’t match what many municipalities required. Keurig also donated C$800,000 to an environmental charity, paid C$85,000 in costs, changed its packaging claims and published corrective notices, including to its subscribers by email Filed.
The claim was one word on the box. The evidence it needed was a map of where the pod could be recycled. That’s the “what the evidence supports” column, filled in by a regulator.
A guide about proof should grade its own. Here’s which findings in this field have held up, which are mixed, and which rest on a single paper.
Marketing runs on famous findings, and some didn’t survive a second look. Before one goes into your playbook, ask: was it tested in the field, and has anyone besides the original team found it again?
| Finding | Original | Since then | Verdict |
|---|---|---|---|
| Better reviews raise sales | Chevalier and Mayzlin, 2006 | Same answer from a rounding design on Yelp (Luca) and a policy shift on Amazon (He and colleagues) | Held up |
| Early ratings steer later ones | Muchnik, Aral and Taylor, 2013 | Randomized; matches Salganik’s randomized MusicLab study (2006). Negative herding was corrected on Muchnik’s site | Held up, for positive herding |
| Rating distributions are J-shaped because of who writes | Hu, Pavlou and Zhang, 2009 | Field experiment: plain requests made reviews less extreme (Brandes, Godes and Mayzlin, 2022) | Held up |
| Ratings fall as reviews accumulate | Godes and Silva, 2012 | Consistent with Li and Hitt (2008) and Moe and Schweidel (2012); reasons debated | Held up in observational data |
| Five reviews lift purchase likelihood 270%; 4.0 to 4.7 is the sweet spot | Spiegel Research Center, 2017 | Observational vendor data; direction fits peer-reviewed work, the exact numbers aren’t causal | Direction yes, numbers no |
| “Most guests reuse their towels” beats a green appeal | Goldstein, Cialdini and Griskevicius, 2008 | A German replication found the norm message did no better than the standard one (Bohner and Schlüter, 2014), though both beat no message | Mixed |
| Too many choices stop people buying (the jam study) | Iyengar and Lepper, 2000 | Meta-analysis of 50 studies: average effect virtually zero, with large variation (Scheibehenne, Greifeneder and Todd, 2010) | Didn’t hold as a rule |
| A little negative information raises liking | Ein-Gar, Shiv and Tormala, 2012 | Four studies in one paper, with narrow conditions. I found no independent replication | Promising, one paper |
| Handmade raises value because it “contains love” | Fuchs, Schreier and van Osselaer, 2015 | Four studies in one paper | Promising, one paper |
| Original-factory products carry brand essence | Newman and Dhar, 2014 | Studies in one paper | Promising, one paper |
| Shown effort raises perceived value | Buell and Norton, 2011 | Field experiments in food service (Buell, Kim and Tsay, 2017), same lead author | Consistent, one research group |
| Rewards buy review count, norms buy length | Burtch and colleagues, 2018 | One field experiment plus an online one | Promising, one field test |
PublishedAll references in Appendix C. The verdicts are my reading of the evidence I could verify for this guide, in September 2026.
Trust findings that were randomized, run in the field, and found twice.
Build on what held up: coverage, the first five reviews, the systematic ask. Don’t cite mixed findings to justify a decision; test the tactic if your traffic allows (The Honest Test). Use one-paper findings only where they’re cheap and true anyway, like showing a critical review or the real work behind the product, and don’t promise a lift.
Twelve numbers, once a month, on one page. Half tell you whether customers have reason to believe you; half tell you whether you could prove it.
Review programs get checked when something goes wrong. The scorecard makes the check routine, so problems show up as a number turning the wrong way instead of a crisis.
| Number | How to count it | Good sign |
|---|---|---|
| Claims with a current file row | Live claims with evidence and an unexpired review date, over all live claims | 100% |
| Claims scoring 45 or more | From the scorer in chapter 3 | Zero, or each with a fix date |
| Product reviews per 100 orders | Reviews that discuss the product, over orders delivered in the ask window | Rising, then steady |
| Reviews about delivery | Share of new reviews about shipping or packaging only | Falling |
| Revenue under five reviews | Share of revenue from products with fewer than five reviews | Falling |
| Days to fifth review | For products launched in the last six months | Under 30 |
| New-review average, top products | Monthly average of that month’s new reviews | No step of 0.3 stars or more |
| Three-star share | Share of new reviews at three stars | Not near zero; the middle is writing |
| Reviews held or removed | Count by star, with the logged reason | Every one logged; no pattern by star |
| Reply time to one- and two-star reviews | Median business days to a public reply | Two or fewer |
| Verified-buyer share | Reviews matched to an order | Steady or rising |
| Incentivized reviews disclosed | Reviews with any reward that carry the disclosure | 100% |
The first two and the last three are compliance numbers: a slip there goes straight to the owner named in chapter 11. Read the middle seven together: if reviews per 100 orders jump while the three-star share falls, check that nobody switched on a filter or a reward.
A review program you only check in a crisis is a crisis you scheduled.
Appendix A has the queries. The one field you’ll usually add is a flag for reviews about delivery.
Claims first, then the review settings, then the ask, then the record. Four weeks, in that order.
The order matters. The claims carry the legal exposure, so they come first. The review settings come next, because a filter running today is a problem today. Then the ask, which takes weeks to show results. Then the routine that keeps it all true.
At day thirty you’ll have fewer claims, each backed by a document someone can find, and reviews that are slightly less perfect and more believable.
Prove what you say, publish what they say, and ask when they know.
Six things the person who owns reviews and claims needs on the first day.
Whoever owns proof, a new hire, an agency or you, needs six things on day one. Without them, the first month goes on hunting for documents.
The books and papers this guide leans on, and what to take from each.
Andrew Lauchner runs Growth Legend, embedding inside consumer brands to own lifecycle, email and SMS, and revenue operations. He is the author of The Second Order, on turning first-time buyers into second-time buyers, and The Whole Machine, on the fundamentals of DTC growth, along with a series of field guides for DTC operators at andrewlauchner.com.
As Senior Director of Growth and Retention Marketing at Gallery Furniture, he rebuilt the customer journey and the sales playbooks together. He has worked on growth and retention at Binance and 3Commas, and has been Head of Growth and Retention at Greatness Wins and at Nexus Agriscience.
“Andrew led retention, lifecycle, and email/SMS, but what separates him from most in this space is how deeply he understands the role retention plays in the overall growth engine.”
Akram Khan, Head of Marketing at Gallery Furniture, senior to Andrew but didn’t manage Andrew directly
Andrew answers every note from operators working on this, including those looking for someone to own it. Write to andrew@growthlegend.com or message him on LinkedIn.
The formulas behind the three tools, and the queries behind the scorecard.
| For | Formula | Notes |
|---|---|---|
| Adjusted rating | (k × m + n × r) / (k + n) | r: product average. n: review count. m: store-wide average. k: weight on the store average, in reviews (10 is a sensible start). |
| Lowest plausible rating | adjusted − 1.645 × s / √(k + n) | s: standard deviation of single ratings, from your export. A one-sided 95% bound. Sort “top rated” by this. |
| Product reviews from an ask on day t | orders × p × F(t) × 0.5^(max(0, t − d − u) / h) | F(t): share able to judge by day t, rising evenly from 0 at delivery d to 1 at d + u. p: response rate once they can judge. h: halving time. |
| Delivery-only reviews | orders × q × (1 − F(t)) | q: response rate before they can judge. |
| Claim risk score | 20·C/3 + 35·E/3 + 25·W/3 + 20·R/3 | C checkability, E evidence, W wording, R regulated term, each 0 to 3. |
-- revenue in the last 90 days by product, with published review count
-- reviews: review_id, product_id, order_id, rating, created_at, status
SELECT p.product_id, p.title,
SUM(ol.price * ol.quantity) AS revenue_90d,
COALESCE(r.reviews, 0) AS reviews,
CASE WHEN COALESCE(r.reviews, 0) < 5 THEN 1 ELSE 0 END AS under_five
FROM order_lines ol
JOIN orders o ON o.order_id = ol.order_id
JOIN products p ON p.product_id = ol.product_id
LEFT JOIN (SELECT product_id, COUNT(*) AS reviews
FROM reviews WHERE status = 'published'
GROUP BY product_id) r ON r.product_id = p.product_id
WHERE o.created_at >= CURRENT_DATE - INTERVAL '90 days'
GROUP BY p.product_id, p.title, r.reviews
ORDER BY revenue_90d DESC;
Sum revenue where under_five = 1 and divide by the total for the scorecard’s “revenue under five reviews.”
-- reviews per 100 delivered orders, by delivery month, split by topic
-- reviews.about_delivery: true when the review only discusses shipping or packaging
SELECT DATE_TRUNC('month', f.delivered_at) AS month,
COUNT(DISTINCT f.order_id) AS delivered,
100.0 * COUNT(r.review_id) FILTER (WHERE NOT r.about_delivery)
/ COUNT(DISTINCT f.order_id) AS product_reviews_per_100,
100.0 * COUNT(r.review_id) FILTER (WHERE r.about_delivery)
/ COUNT(DISTINCT f.order_id) AS delivery_reviews_per_100
FROM fulfillments f
LEFT JOIN reviews r ON r.order_id = f.order_id
WHERE f.delivered_at >= CURRENT_DATE - INTERVAL '12 months'
GROUP BY 1 ORDER BY 1;
Count reviews against the month the order was delivered, not the month the review arrived, or a change in ask timing will look like a change in response. Leave the latest two months out of any comparison until their reviews have had time to come in.
-- monthly average of new reviews, and average by position in the sequence
WITH seq AS (
SELECT r.*, ROW_NUMBER() OVER (PARTITION BY product_id ORDER BY created_at) AS n
FROM reviews r WHERE status = 'published'
)
SELECT product_id,
DATE_TRUNC('month', created_at) AS month,
COUNT(*) AS new_reviews,
ROUND(AVG(rating), 2) AS avg_new,
ROUND(AVG(rating) FILTER (WHERE n <= 10), 2) AS avg_first_10,
ROUND(AVG(rating) FILTER (WHERE n BETWEEN 11 AND 100), 2) AS avg_11_to_100
FROM seq
GROUP BY 1, 2 ORDER BY 1, 2;
Chart avg_new by month with supplier, formula and packaging changes marked. A step of 0.3 stars or more at a change date is a product question; a slow slide is drift.
-- for products published in the last 6 months
WITH ranked AS (
SELECT product_id, created_at,
ROW_NUMBER() OVER (PARTITION BY product_id ORDER BY created_at) AS n
FROM reviews WHERE status = 'published'
)
SELECT p.product_id, p.title, p.published_at::date AS launched,
(r.created_at::date - p.published_at::date) AS days_to_fifth
FROM products p
LEFT JOIN ranked r ON r.product_id = p.product_id AND r.n = 5
WHERE p.published_at >= CURRENT_DATE - INTERVAL '6 months'
ORDER BY days_to_fifth NULLS FIRST;
The syntax is Postgres. In BigQuery, replace FILTER (WHERE ...) with COUNTIF or AVG(IF(..., rating, NULL)), and date subtraction with DATE_DIFF. Column names follow a generic Shopify-style export; rename to match yours.
The file, the ask, the reply and the policy. Copy them into whatever your team already uses, and have counsel review the policy before you publish it.
CLAIM (EXACT WORDS) | WHERE IT RUNS (URLs, templates, feeds, packaging) KIND (map row) | EVIDENCE (link to the document itself) SOURCE AND DATE | INDEPENDENT OF US? (yes / no) WHAT IT SUPPORTS | the widest wording the evidence allows OWNER | NEXT REVIEW DATE | RISK SCORE (chapter 3) STATUS | live / qualified / pulled, with date
SUBJECT How Is Your [Product] Working So Far? PREVIEW two minutes, and every review gets published, good or bad Hi [First name], You've had your [product] for about [two weeks] now. Other customers decide whether to buy it by reading what people like you say, so we'd be grateful for a few honest lines. What worked, what didn't, and who you'd recommend it to. [ Write a review ] We publish every review that follows our review policy [link], whatever the rating. If something isn't right, reply to this email and a person will help, and you're still welcome to review it. [Brand] · [Address] You're receiving this because you ordered from us. [Unsubscribe]
No reward, no suggested rating, same email to everyone. For the reminder a week later, add the true review count as a social-norm line.
[Brand]: Hi [First name], it's been two weeks with your [product]. How's it going? Tell other customers here: [short link] Reply STOP to opt out
Send only to customers who have agreed to marketing texts, within your sending hours.
ON THE REVIEW "This reviewer received [a free product / a 10% code]
for writing a review. The reward did not depend on
the rating or what they wrote."
Thank you, [name], for telling us and other customers. [Fact, without arguing: e.g. "The medium does run small in this style."] We've [what you did or changed: e.g. "added a sizing note to the page"]. If you'd like a replacement or a refund, write to [email] and [first name] on our team will sort it out. [First name], [Brand]
Never offer anything in exchange for changing or removing the review.
HOW WE COLLECT REVIEWS We email every customer [N] days after delivery to ask for a review. We don't choose who is asked based on how we expect them to rate us. [If applicable: Reviewers who received a free product or reward are labeled, and the reward never depended on their rating.] HOW WE CHECK THEM Reviews marked "Verified buyer" match an order. [Say what you do with reviews that don't match an order.] WHAT WE PUBLISH Every review, positive or negative, except those that contain: personal information, abuse or harassment, content unrelated to the product, or content we can show is false. The same rules apply at every rating. We log every review we don't publish and why. HOW WE CALCULATE THE RATING The average of all published reviews for this product[, including those imported from (source)].
For [product / component], please send by [date]: 1. Where it is made, and where each significant step happens. 2. Country of origin of each input over [5]% of cost. 3. Certificates and test reports for [claim], with dates. 4. Notice of any change to the above, before it ships.
Every external source, by chapter. Web sources were read in September 2026.