Replication Crisis Casualties

Facts and insights about replication crisis casualties.

Replication Crisis: Beginning around 2011, scientists discovered that a large share of published findings in psychology, medicine and biology could not be reproduced when other labs repeated the experiments. The scandal exposed weak statistics, publication bias, flexible analysis and outright fraud, and it triggered reforms such as preregistration, open data and registered reports.

Open Science Collaboration: In 2015, 270 researchers tried to repeat 100 published psychology studies. Roughly 36 percent of the repeats produced a statistically significant result, compared with 97 percent of the originals, and the effects that survived were about half the original size.

Stanford Prison Experiment: Philip Zimbardo's 1971 study claimed ordinary students spontaneously turned cruel as guards. Later audio recordings and archives showed that experimenters coached the guards to be tough, and participant Douglas Korpi admitted he exaggerated his breakdown.

Power Posing: A 2010 study by Dana Carney, Amy Cuddy and Andy Yap claimed two minutes of expansive posture raised testosterone and risk taking. A larger 2015 replication found no hormonal effect, and in 2016 Carney publicly stated she no longer believed the effect was real.

Ego Depletion: Roy Baumeister's 1998 radish and cookie experiment suggested willpower works like a muscle that tires. A 2016 multi lab replication across 23 labs and more than 2,000 participants found an effect statistically indistinguishable from zero.

Facial Feedback Hypothesis: Fritz Strack's 1988 pen in the mouth study claimed that forcing a smile makes cartoons seem funnier. A 2016 Registered Replication Report spanning 17 labs failed to reproduce the result.

Daryl Bem: His 2011 paper Feeling the Future reported nine experiments suggesting people could sense events before they happened. It was published in a top psychology journal, and when others failed to repeat it, the field was forced to question how such a result passed peer review.

John Bargh: His 1996 Yale study claimed students who unscrambled words about old age walked more slowly down a hallway. In 2012, Stephane Doyen's team used infrared timers and found the effect only appeared when experimenters expected it.

Marshmallow Test: Walter Mischel's famous experiment linked a child's ability to wait for a second treat with later success. A 2018 replication of about 900 children found an effect roughly half as large, and most of it vanished once family background was controlled.

Diederik Stapel: This Dutch social psychologist was exposed in 2011 for inventing data. An investigation found fraud in at least 55 of his publications, and he admitted fabricating entire datasets at home and handing them to students.

Marc Hauser: The Harvard psychologist studied morality and cognition in monkeys. In 2010 Harvard found him responsible for eight instances of scientific misconduct, and he resigned in 2011.

Kathleen Vohs: Her 2006 Science paper claimed that subtle reminders of money made people more self reliant and less helpful. Large replication attempts later failed to find the effect, making money priming a textbook casualty.

Reproducibility Project: Cancer Biology: This eight year effort tried to repeat 50 experiments from 23 high profile cancer papers. The final 2021 report found that effect sizes were on average about 85 percent smaller than originally reported.

Glenn Begley: In 2012 the former Amgen scientist reported his team could confirm the findings of only 6 of 53 landmark cancer studies. Some original authors told him the result was real only in their best run.

Bayer HealthCare: In 2011 company scientists reported that they could reproduce only roughly a quarter of the published academic findings they tried to build drug projects on. The internal audit became an early warning that medicine had a reproducibility problem.

Hurricane Names Study: A 2014 PNAS paper claimed hurricanes with female names killed more people because they seemed less threatening. Statisticians challenged the analysis, and the result did not hold up under scrutiny.

Israeli Parole Judges: A 2011 study reported that favorable rulings dropped from about 65 percent at the start of a session to nearly zero before a break, blamed on hungry judges. Critics showed that the order of cases was not random, which could explain much of the pattern.

Oxytocin Trust Study: Michael Kosfeld's 2005 Nature paper reported that sniffing oxytocin made people more trusting with money. Later reviews and replications found the effect small or absent, and the nickname love hormone faded.

STAP Cells: In January 2014, Haruko Obokata's team in Japan claimed ordinary cells could become stem cells with an acid bath. Labs worldwide could not repeat it, and the Nature papers were retracted by July 2014.

GFAJ-1: In 2010 a NASA funded team announced a Mono Lake bacterium that swapped phosphorus for arsenic in its DNA. Independent labs in 2012 found no arsenic in the DNA, undermining a claim that was billed as redefining life.

Cold Fusion: In March 1989, Martin Fleischmann and Stanley Pons announced room temperature nuclear fusion at the University of Utah. Laboratories around the world rushed to repeat it and almost none could, in one of the most public replication failures in history.

Mozart Effect: A 1993 Nature note by Frances Rauscher and colleagues reported brief spatial reasoning gains after 36 students listened to Mozart. Georgia's governor later funded classical music CDs for newborns, though follow up research found no lasting boost in intelligence.

Dan Ariely: A famous 2012 PNAS study said signing an honesty pledge at the top of a form reduced cheating. A 2021 analysis found the insurance dataset in the paper showed signs of fabrication, and Ariely denied knowingly falsifying data.

Bottomless Soup Bowl: Brian Wansink's Cornell lab claimed people given a self refilling bowl ate about 73 percent more soup without noticing. After a 2017 and 2018 investigation, the paper was retracted and Cornell found misconduct, and he resigned in 2019.

Implicit Association Test: Introduced in 1998, the IAT promised to measure hidden bias in seconds. Its test retest reliability is only moderate, and studies find it poorly predicts real world behavior, which sparked an ongoing debate about what it truly measures.