A guide about proof should grade its own. Here’s which findings in this field have held up, which are mixed, and which rest on a single paper.
Marketing runs on famous findings, and some didn’t survive a second look. Before one goes into your playbook, ask: was it tested in the field, and has anyone besides the original team found it again?
| Finding | Original | Since then | Verdict |
|---|---|---|---|
| Better reviews raise sales | Chevalier and Mayzlin, 2006 | Same answer from a rounding design on Yelp (Luca) and a policy shift on Amazon (He and colleagues) | Held up |
| Early ratings steer later ones | Muchnik, Aral and Taylor, 2013 | Randomized; matches Salganik’s randomized MusicLab study (2006). Negative herding was corrected on Muchnik’s site | Held up, for positive herding |
| Rating distributions are J-shaped because of who writes | Hu, Pavlou and Zhang, 2009 | Field experiment: plain requests made reviews less extreme (Brandes, Godes and Mayzlin, 2022) | Held up |
| Ratings fall as reviews accumulate | Godes and Silva, 2012 | Consistent with Li and Hitt (2008) and Moe and Schweidel (2012); reasons debated | Held up in observational data |
| Five reviews lift purchase likelihood 270%; 4.0 to 4.7 is the sweet spot | Spiegel Research Center, 2017 | Observational vendor data; direction fits peer-reviewed work, the exact numbers aren’t causal | Direction yes, numbers no |
| “Most guests reuse their towels” beats a green appeal | Goldstein, Cialdini and Griskevicius, 2008 | A German replication found the norm message did no better than the standard one (Bohner and Schlüter, 2014), though both beat no message | Mixed |
| Too many choices stop people buying (the jam study) | Iyengar and Lepper, 2000 | Meta-analysis of 50 studies: average effect virtually zero, with large variation (Scheibehenne, Greifeneder and Todd, 2010) | Didn’t hold as a rule |
| A little negative information raises liking | Ein-Gar, Shiv and Tormala, 2012 | Four studies in one paper, with narrow conditions. I found no independent replication | Promising, one paper |
| Handmade raises value because it “contains love” | Fuchs, Schreier and van Osselaer, 2015 | Four studies in one paper | Promising, one paper |
| Original-factory products carry brand essence | Newman and Dhar, 2014 | Studies in one paper | Promising, one paper |
| Shown effort raises perceived value | Buell and Norton, 2011 | Field experiments in food service (Buell, Kim and Tsay, 2017), same lead author | Consistent, one research group |
| Rewards buy review count, norms buy length | Burtch and colleagues, 2018 | One field experiment plus an online one | Promising, one field test |
PublishedAll references in Appendix C. The verdicts are my reading of the evidence I could verify for this guide, in September 2026.
Trust findings that were randomized, run in the field, and found twice.
Build on what held up: coverage, the first five reviews, the systematic ask. Don’t cite mixed findings to justify a decision; test the tactic if your traffic allows (The Honest Test). Use one-paper findings only where they’re cheap and true anyway, like showing a critical review or the real work behind the product, and don’t promise a lift.
This is one chapter of The Proof File, which is free and readable in full on a single page with no form in front of it.