One random early vote changed where ratings ended up, in a randomized experiment on a hundred thousand comments. Launch like the first reviews matter, because they do.
If people rated only what they saw, the order of reviews wouldn’t matter. It does. People read what came before and adjust, and the first rating sets the direction.
Lev Muchnik, Sinan Aral and Sean Taylor ran a randomized experiment on a social news site like Reddit, published in Science in 2013. Over five months, 101,281 comments were randomly assigned at birth to get one up-vote, one down-vote, or nothing. The votes had nothing to do with the content. The researchers then watched 308,515 ratings from real users Published.
Positive herding compounds; negative herding gets corrected, at least on that site. In Matthew Salganik, Peter Dodds and Duncan Watts’s MusicLab experiment with 14,341 participants, showing what others downloaded made hits bigger and success less predictable. The best songs rarely did badly and the worst rarely did well, “but any other result was possible” Published.
Your first reviews don’t just describe the product. They steer every review that follows.
The first reviews come from whoever buys first, and they aren’t typical. Xinxin Li and Lorin Hitt showed that early buyers’ particular tastes shape later buyers’ decisions and can mislead them Published. Your loyalists produce fans; cold paid traffic produces mismatches. Neither is your month-six customer.
Fake early reviews work briefly, which is why people buy them. He, Hollenbeck and Proserpio found the effect was short-lived: after products stopped buying fake reviews, their ratings fell and their share of one-star reviews rose significantly, especially for young products Published. In 2019 the FTC brought its first case challenging fake paid reviews on an independent retail website: Cure Encapsulations had paid a website to post Amazon reviews of a weight-loss supplement and asked it to keep the product at five stars. The judgment was $12.8 million, suspended on payment of $50,000 Filed. Since October 2024, fake reviews carry civil penalties per violation (chapter 11).
A 4.9 average from 12 reviews looks better than a 4.6 from 840. It isn’t, yet. One more one-star review would pull the first to 4.6, while the second would barely move. The fix is to rank by a score that shrinks small counts toward your store’s typical rating and then asks how low the true average could plausibly be. It’s a simplified version of the Bayesian approach the statistician Evan Miller describes for star ratings Reported.
With the defaults, product A’s 4.9 adjusts to 4.67 with a lowest plausible score of 4.29, while B’s 4.6 adjusts to 4.60 with a floor of 4.54. B is the one to feature. A needs about 27 more reviews at 4.9 to pass it, and one more one-star review would show it at 4.6 Derived. Use the same score to sort your “top rated” collection.
This is one chapter of The Proof File, which is free and readable in full on a single page with no form in front of it.