Collection: How Psychology Knows

Deep DiveESTABLISHED

What the Replication Crisis Changed

Psychology discovered that many famous findings looked weaker when scientists tried to replay them. The real story is what happened after the shock.

Parallel psychology experiments produce different patterns in repeated trials.

In 2015, hundreds of researchers tried something uncomfortable.

They went back.

Instead of searching for the next surprising psychological effect, they attempted to reproduce 100 findings that had already made it into prominent journals.

The original papers looked impressive.

Ninety-seven percent of the selected original studies reported statistically significant results.

When the experiments were repeated, only 36 percent of the replication studies reached statistical significance.

The average replicated effect was roughly half the size of the original average effect.

Those numbers spread around the world.

And almost instantly, a complicated scientific project was compressed into a brutal headline:

Psychology doesn't replicate.

That headline was too simple.

The underlying problem was not.

What was the replication crisis?

Replication is one of science's basic error-correction mechanisms.

A finding should not become credible only because a famous researcher reported it once.

Other researchers should be able to collect new data under comparable conditions and observe evidence consistent with the claim.

That does not mean every replication must produce the same p-value or identical effect size.

Samples vary.

Contexts differ.

Measurements contain noise.

But if a claimed effect repeatedly disappears, shrinks dramatically or changes direction under rigorous replication, confidence should change.

By the early 2010s, psychology had accumulated enough suspicious patterns that researchers began organizing large, transparent replication projects.

The 2015 Reproducibility Project: Psychology became the symbol.

The 100-study project

The Open Science Collaboration selected 100 experimental and correlational studies published in 2008 across three psychology journals.

Replication teams used original materials when available, worked with original authors, preregistered protocols and used relatively high-powered designs.

The result was not one number.

It was several.

Thirty-six percent of replications produced statistically significant results.

Forty-seven percent of original effect sizes fell inside the replication study's 95 percent confidence interval.

Thirty-nine percent of effects were subjectively rated by replication teams as having replicated the original result.

The mean replication effect size was about half the original mean.

Different definitions of “replicated” produced different answers.

That matters.

Why the 36% number is not “36% of psychology is true”

The project did not randomly sample every topic, journal, method or decade in psychology.

It examined 100 findings from three journals.

It also intentionally targeted published findings, which already passed through the publication system.

And replication success is not binary in the way a light switch is binary.

An effect can be smaller but still present.

A confidence interval can include both a meaningful effect and zero.

A replication can differ because of population or context.

The original study can be wrong.

The replication can be noisy.

Or both studies can be estimating a context-sensitive effect.

The correct conclusion is not that exactly 64 percent of psychology was false.

The correct conclusion was more disturbing and more useful:

published confidence was often stronger than the evidence deserved.

The hidden incentive problem

Why would that happen?

Imagine two researchers run equally careful studies.

One finds a dramatic statistically significant effect.

The other finds nothing.

Which paper is easier to publish?

For decades, journals, careers and media attention tended to reward positive, surprising results more strongly than null findings.

That creates a selection system.

Even if every scientist is honest, the literature can become distorted because exciting results are more likely to appear in print.

This is publication bias.

The file drawer fills with boring results.

The published shelf fills with winners.

If you only see the shelf, reality looks more dramatic than it is.

Researcher degrees of freedom

There was another problem.

Many studies allowed a large number of defensible analytic choices.

Which participants should be excluded?

Which outcome should be primary?

Should the analysis control for age?

Should data collection stop at 40 participants or continue to 60?

Which subgroup should be examined?

Each decision may seem reasonable.

But if researchers see the data before finalizing the decisions, flexibility can slowly bend toward significance.

This can happen intentionally.

It can also happen without conscious deception.

P-hacking describes practices that increase the chance of obtaining statistically significant results by trying multiple analyses, stopping rules or selections until something crosses the threshold.

HARKing means hypothesizing after the results are known, then presenting the discovered pattern as if it had been predicted in advance.

The problem is not exploration.

Exploration is essential.

The problem is disguising exploration as confirmation.

Small studies create dramatic winners

Small samples are noisy.

They can miss real effects.

They can also exaggerate the size of effects that happen to cross the significance threshold.

If only the significant small studies are published, the literature can become filled with inflated effect estimates.

Later, larger replications may find the effect is much smaller.

This is one reason the replication crisis was also a statistical-power crisis.

The original finding may not vanish completely.

It may simply shrink from headline-sized to human-sized.

The crisis changed replication itself

Before the crisis, direct replication often had low prestige.

Scientists were rewarded for novelty.

Repeating someone else's experiment could be treated as uncreative.

That culture began to change.

Large multi-lab projects made replication visible.

Some journals explicitly began treating rigorous replications as valuable scientific contributions.

And a deeper principle returned to the center:

a claim becomes stronger when it survives people who are not invested in discovering it.

Preregistration: write the plan before seeing the ending

Preregistration is one reform.

Researchers specify hypotheses, outcomes, exclusions, sample plans and analyses before the final data are examined.

This does not prevent exploratory analysis.

It creates a boundary between what was predicted and what was discovered afterward.

That distinction is powerful because the statistical meaning of a test changes when the test was selected after seeing the data.

Preregistration is not a truth machine.

A bad plan can be preregistered.

Researchers can deviate from it.

Complex studies may require flexibility.

But the plan makes the researcher's decision path more visible.

Visibility is the point.

Registered Reports: publish the question before the result

Registered Reports go further.

In this format, journals evaluate the research question and methods before the results are known.

A strong protocol can receive in-principle acceptance before data collection.

If researchers follow the approved plan and interpret the results responsibly, publication does not depend on whether the result is positive, negative or statistically significant.

That flips an old incentive.

Instead of asking, “Did you find something exciting?” the system first asks, “Did you design a study capable of answering an important question?”

Nature Neuroscience described Registered Reports as a way to reduce publication bias and questionable practices such as selective reporting, p-hacking and HARKing.

A 2021 study comparing 29 Registered Reports with conventional papers found that reviewers rated the Registered Reports higher on methodological and analytic rigor while finding no clear loss of novelty or creativity.

Promising.

Not perfect.

Open data, open materials, open code

A finding becomes easier to inspect when researchers share how it was produced.

Open materials allow others to repeat the experiment.

Open data allow others to verify analyses, where privacy and ethics permit.

Open code exposes analytic decisions.

None of this guarantees a correct conclusion.

Transparent mistakes are still mistakes.

But transparency improves the chance that errors can be found.

Science does not become reliable by eliminating mistakes.

It becomes reliable by making mistakes easier to detect and correct.

The crisis was also a theory crisis

Better statistics cannot rescue a theory that predicts almost anything.

Several psychologists argued that weakly specified theories contributed to the problem.

If a theory can explain a positive result, a null result and the opposite result after the fact, replication cannot meaningfully test it.

That is why the post-crisis conversation expanded from methods to theory.

What exactly does the theory predict?

Under which conditions?

What result would count against it?

Can competing theories make different predictions?

A field can become more reproducible and still remain theoretically vague.

Replication is necessary error correction.

It is not the entire scientific project.

Did the reforms solve the crisis?

No.

Science changes slower than a checklist.

Career incentives remain.

Open practices vary by field.

Preregistration can become ceremonial.

Registered Reports fit some research questions better than others.

Data sharing can be limited by privacy, ownership and consent.

Replication itself can become biased if only spectacular failures attract attention.

A 2026 Nature Reviews Psychology commentary emphasized systemic forces in the crisis, not just individual researcher behavior.

The problem was never only “bad scientists.”

It was a system that often rewarded a certain kind of result.

Systems require structural fixes.

The most important thing the crisis changed

The replication crisis did something paradoxical.

It damaged trust in psychology.

And it also gave psychology better reasons to deserve trust.

A mature science should be able to discover that some of its celebrated findings were overstated.

It should be able to update publication practices.

It should be able to make replication prestigious.

It should expose analysis plans earlier.

It should publish informative failures.

A field that never finds problems in its own literature would be more suspicious than a field that does.

The crisis was not proof that psychology cannot know anything.

It was proof that knowing requires a stronger machine for being wrong.

Key Takeaways

  1. The 2015 Reproducibility Project repeated 100 findings from three psychology journals, not all of psychology.
  2. 97% of the selected originals were statistically significant versus 36% of replications; replication effect sizes averaged roughly half the originals.
  3. Replication success has multiple definitions, so the crisis cannot be reduced to one percentage.
  4. Publication bias can make the visible literature look more positive than the full body of conducted research.
  5. Researcher degrees of freedom can inflate false-positive risk when analysis choices are made after seeing data.
  6. Small studies can produce exaggerated published effect sizes.
  7. Preregistration separates planned confirmatory tests from later exploration.
  8. Registered Reports move peer review and publication commitment before the result is known.
  9. Open data, materials and code improve inspectability but do not guarantee correctness.
  10. The replication crisis also exposed problems in weakly specified psychological theory.
  11. The most constructive legacy of the crisis is stronger scientific error correction, not blanket cynicism.