The Replication Crisis in Psychology: Why Famous Findings Keep Falling Apart

Every few years, an example in a psychology textbook quietly disappears from the next edition without much explanation. That is usually a sign of the replication crisis in psychology at work, a decade-long reckoning over how many published findings actually hold up when other researchers try to reproduce them. Some of the field’s most cited studies, the ones that show up in TED talks and intro textbooks alike, turned out to rest on shakier ground than anyone expected.

Key Takeaways

  • A landmark 2015 project attempted to redo 100 published psychology studies and found that only about a third produced the same statistically significant result the second time.
  • Several widely taught findings, including power posing and ego depletion, failed large-scale replication attempts run by independent labs.
  • Small sample sizes, flexible statistical choices, and a publishing culture that rewards novelty over confirmation all contributed to the problem.
  • Psychology has responded with preregistration, open data sharing, and Registered Reports, changes that are gradually rebuilding trust in the field’s findings.

What Replication Actually Means

Before getting into specific cases, it helps to be precise about the vocabulary, since these terms get used loosely in casual conversation. A direct replication repeats the original study as closely as possible, using the same procedures and materials, to see whether the same result appears again. A conceptual replication tests the same underlying idea using a different method, which is useful for confirming that an effect is not just an artifact of one particular setup.

Neither of these is quite the same as reproducibility, which refers to whether someone else can take the original raw data and analysis code and arrive at the same reported numbers. A study can be perfectly reproducible in that narrow sense, meaning the math checks out, while still failing to replicate when a new sample of participants is tested. Generalizability is a separate question again: even a finding that replicates reliably in one population, say American undergraduates, may not hold in a different age group, culture, or context. Confusing these terms is part of why public discussion of the replication crisis in psychology gets muddled so quickly.

Replication vs Reproducibility vs Generalizability

How the Numbers Added Up

The moment that turned a simmering concern into a full-blown crisis was the 2015 Reproducibility Project: Psychology, a coordinated effort by the Open Science Collaboration involving more than 270 researchers. The team selected 100 studies published in three respected psychology journals and attempted to redo each one using the original materials and appropriately powered sample sizes.

The results were sobering. Ninety-seven percent of the original studies had reported statistically significant effects, which is exactly what you would expect given how journals favor positive findings. But only 36 percent of the replication attempts reached statistical significance, and the effects that did replicate were, on average, about half the size of what the original papers had claimed (Open Science Collaboration, 2015). Social psychology studies fared worse than cognitive psychology studies in this particular sample, though both showed meaningful declines.

That gap between 97 percent and 36 percent is the number most people remember, and it is a fair one to remember, but it also does not mean two-thirds of psychology is fabricated or worthless. It means the field had been running on optimistic assumptions about statistical power and effect sizes for a long time, and 2015 was when those assumptions got tested directly.

Replication Crisis Examples That Reshaped the Field

Numbers on a page do not stick with students the way specific stories do, so it is worth walking through a few of the most consequential replication crisis examples individually.

Precognition and the Bem studies. In 2011, respected social psychologist Daryl Bem published nine experiments in a mainstream, peer-reviewed journal claiming evidence that people could sense future events before they happened (Bem, 2011). The paper followed conventional statistical methods, which was exactly the point critics raised: if standard practices in the field could produce apparent support for precognition, something was wrong with the standard practices, not just this one paper. Independent teams that tried to replicate the effect using preregistered designs consistently came up empty (Ritchie et al., 2012). Bem’s paper became a turning point precisely because it was so methodologically ordinary.

Power posing. Amy Cuddy’s 2010 study, coauthored with Dana Carney and Andy Yap, claimed that standing in an expansive posture for two minutes could shift hormone levels and increase confidence. It became one of the most viewed TED talks of all time. A 2015 replication with a much larger sample, led by Eva Ranehill and colleagues, reproduced the self-reported feelings of power but found no evidence for the hormonal or behavioral effects the original study had emphasized. Carney, the study’s original lead author, later published a public statement saying she no longer believed the effect was real.

Ego depletion. The idea that willpower draws from a limited, exhaustible resource had shaped decades of research on self-control. In 2016, a Registered Replication Report led by Martin Hagger coordinated 23 laboratories and more than 2,000 participants to retest the effect using a shared, preregistered protocol. The combined result was close to zero. Notably, 22 of the 23 participating labs had predicted beforehand that they would successfully replicate the effect, which says something about how confident the field had become in a finding that did not hold up under close inspection.

These examples span different subfields and different decades of research, but the pattern across them is consistent: an intriguing original finding, wide public attention, and then a large, carefully designed replication attempt that failed to reproduce the original result.

Social priming effects. A related cluster of findings came from social priming research, which claimed that subtle exposure to certain words or concepts could measurably change behavior. One widely cited study reported that participants primed with words related to old age walked more slowly down a hallway afterward, without being aware of the connection. Several attempts by other labs to reproduce this specific effect did not find it, and the broader unease around social priming became significant enough that Nobel laureate Daniel Kahneman, who had praised some of this work in his own writing, publicly raised concerns about whether the underlying effects could be trusted. Priming research did not disappear as a result, but it became far more cautious about which specific effects it treated as established.

What ties these cases together is not that the researchers involved were acting in bad faith. Most were working within norms that the entire field considered acceptable at the time: modest sample sizes, some flexibility in how data got analyzed, and a strong pull toward publishing whatever came out statistically significant. The replication crisis in psychology exposed how much those norms mattered, not because any single researcher broke the rules, but because the rules themselves left too much room for chance findings to look like real discoveries.

Why This Kept Happening

None of this happened because researchers were lying. Most of what drove the replication crisis in psychology came from ordinary incentives and small, defensible-seeming decisions that added up to a systemic problem.

Small sample sizes were common, especially in social psychology, because running fewer participants is cheaper and faster. Smaller samples produce noisier estimates, which makes it easier for a study to stumble into a significant result that will not hold up later. Researchers also had a lot of flexibility in how they analyzed their data, deciding after the fact which variables to include, which outliers to drop, or when to stop collecting data. Simmons and colleagues demonstrated that this kind of flexibility, often called p-hacking, can push the false positive rate for a study well above the nominal 5 percent threshold researchers assume they are working with.

Publication bias compounded the problem. Journals have historically preferred novel, positive findings over null results or straightforward replications, so studies that failed to find an effect often never got published at all. That created a distorted literature where the successful, surprising, headline-friendly studies were overrepresented, and the quieter failures that would have offered a more balanced picture were sitting in file drawers. Anyone digging into measurement issues in psychological research will recognize a related theme: a test or a study can look clean on paper while still producing conclusions that do not hold up under scrutiny.

Some of the studies exposed by this era were also connected to broader ethical concerns in psychological research more generally. Reexaminations of older, celebrated work, including scrutiny of how Zimbardo ran the Stanford Prison Experiment and later disclosures about experimenter influence in Milgram’s obedience research, fed into the same broader conversation about how much confidence the field should place in dramatic, singular findings that were never designed to be replicated at scale.

How Psychology Is Fixing Itself

The response to all this has been substantial, and it is one of the more encouraging parts of the story. Preregistration, where researchers publicly commit to their hypotheses and analysis plan before collecting data, has become far more common and removes much of the after-the-fact flexibility that fueled p-hacking. Registered Reports take this further by having journals evaluate and accept a study based on its methods before the results are even known, which removes the incentive to chase a positive finding.

The Center for Open Science and its Open Science Framework have made it standard practice for many researchers to share raw data, materials, and analysis code alongside their published papers, so other labs can check the work directly rather than taking it on faith. Multi-lab collaborations, like the ones that tested ego depletion and the original Reproducibility Project itself, have become a normal part of how the field settles disputed effects, replacing the older model where a single lab’s finding could stand largely unchallenged for years.

Journals have adjusted their own incentives too. Some now accept papers that report null results, which was rare before the crisis, and several prominent outlets have added badges or formal recognition for studies that share open data and preregistered materials. None of these reforms make replication automatic, and plenty of debate continues over how strict preregistration should be or how many labs a Registered Replication Report really needs to be convincing. But the direction of change has been consistent: less trust placed in any single study, and more weight given to findings that have been tested repeatedly across different labs and samples.

None of this makes psychology unreliable as a science. If anything, a field willing to publicly test and revise its own most famous findings is behaving exactly the way science is supposed to. Students encountering the replication crisis in psychology for the first time should walk away less with distrust and more with a sharper sense of how to read a study: check the sample size, ask whether it has been replicated, and treat single, dramatic findings with healthy skepticism until other labs have had a chance to weigh in.

References

How to cite this article:

The Psychology Notes Headquarters. (2026). The Replication Crisis in Psychology: Why Famous Findings Keep Falling Apart. Retrieved from https://www.psychologynoteshq.com/replication-crisis-in-psychology/

Leave a Reply

Your email address will not be published. Required fields are marked *

Post comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.