Daniel Kahneman's First Replication Failure Changed Decision Research
May 29, 2026 By Jonas Eriksen

In 1979, Daniel Kahneman and Amos Tversky published a paper that would reshape economics and psychology. Their prospect theory, built on a series of clever experiments, described how people make decisions under risk. Central to the theory was the "framing effect": the idea that the way a choice is presented — as a potential gain versus a potential loss — dramatically shifts people's preferences. The original study used only 95 participants, yet the effect sizes were large and appeared robust. For decades, the finding was taught as gospel in textbooks and cited in thousands of papers. No one thought to replicate it.

Then, in 2016, a massive replication project called Many Labs 2 put the classic framing experiment to the test. With more than 2,200 participants across 20 laboratories, the result was sobering: the framing effect shrank to a trivial size (Cohen's d ≈ 0.05). Kahneman, then 82, did not dismiss the finding. He publicly acknowledged the failure and called for reform. That moment marked a turning point in behavioral science.

A Nobel Prize Built on a Single Study

Kahneman and Tversky's 1979 paper in Econometrica introduced prospect theory, which later earned Kahneman the Nobel Memorial Prize in Economic Sciences. The theory was supported by several experiments, but the most famous involved a hypothetical disease outbreak and two treatment options. When the options were framed in terms of lives saved (a gain frame), most participants chose the risk-averse option. When framed in terms of lives lost (a loss frame), they switched to the risk-seeking option. The effect was striking: roughly 70–80% of participants shifted their choice depending on the frame.

The original sample was just 95 participants, all undergraduates at the University of British Columbia. Statistical power was low by modern standards, but the effect appeared so large that it seemed safe. No pre-registration existed in 1979; researchers had wide latitude in how they analyzed data. The finding was replicated informally in classroom demonstrations, but formal, independent replications were rare. The field trusted the result because it fit a compelling theoretical story.

Prospect theory itself has been broadly supported in subsequent research, but the specific framing effect — the classic "Asian disease problem" — became a symbol of behavioral economics. It appeared in countless textbooks, TED talks, and policy briefs. The U.S. government even used framing principles in public health campaigns. Yet the original evidence rested on a single, small study.

Kahneman himself later expressed unease. In his 2011 book Thinking, Fast and Slow, he noted that framing effects could be fragile. He wrote that "the framing of a decision can change its meaning" but acknowledged that the effect might not be as universal as once thought. Still, no one had systematically tried to replicate the exact 1979 experiment with adequate sample size.

The 2016 Replication That Stunned the Field

Many Labs 2 was a collaborative effort led by Richard Klein and colleagues, involving over 200 researchers worldwide. It aimed to replicate 28 classic findings in psychology, including Kahneman and Tversky's framing effect. The project used a standardized protocol: each lab recruited participants, presented the same disease scenario, and recorded choices. The total sample exceeded 2,200 participants, giving the replication high statistical power.

The results were published in 2018 in Social Psychological and Personality Science. For the framing effect, the overall effect size was d ≈ 0.05 — practically zero. Only 1 of the 20 labs found a statistically significant effect in the predicted direction. The original study had reported an effect large enough that, with that sample size, it would have been detected with near certainty. The replication suggested that the original effect was either a false positive or heavily inflated by small-sample bias.

Kahneman was informed of the results before publication. He wrote a brief commentary that accompanied the paper, titled "The Use of Replication to Evaluate the Validity of Published Findings." In it, he accepted the replication as credible and noted that the original study likely benefited from "researcher degrees of freedom" — the flexibility in how data are collected and analyzed. He did not blame the original researchers (Tversky had died in 1996) but instead called for structural changes in the field.

The reaction was mixed. Some researchers felt vindicated, arguing that the replication crisis was real and that behavioral economics needed reform. Others worried that the failure of one experiment might unfairly discredit prospect theory as a whole. Kahneman himself emphasized that the core ideas of prospect theory — loss aversion, diminishing sensitivity, probability weighting — had been supported by many other studies. The framing effect was just one demonstration, not the theory itself.

To put the replication in perspective, consider another classic finding from the same era: the "endowment effect" — the tendency for people to value an object more once they own it. A 2020 replication by the Many Labs 2 team also found a much smaller effect than originally reported, with a Cohen's d of about 0.15 compared to the original 0.80. Similarly, the "anchoring effect" — where an initial numerical cue influences subsequent judgments — showed a reduced effect size in a 2021 multi-lab replication, shrinking from d ≈ 0.80 to d ≈ 0.20. These examples illustrate that the framing effect was not alone; many landmark findings in behavioral science turned out to be smaller or less robust than initially believed.

Why the Original Result Looked So Convincing

Several factors combined to make the 1979 finding appear stronger than it was. The most obvious is small sample size: with only 95 participants, the study had limited precision. Effect sizes from small samples are notoriously variable; a lucky draw can produce a large effect that is not reproducible. This is known as the "winner's curse" in meta-analysis: the first study to report a finding often overestimates its true size.

Publication bias also played a role. In 1979, there was no requirement to register studies before data collection. Journals favored novel, surprising results. Null findings — studies that failed to find a framing effect — were unlikely to be published. The file drawer filled with non-significant results, making the published literature look more consistent than it was.

Researcher degrees of freedom were abundant. The original study had no pre-registered analysis plan. Experimenters could choose which participants to include, how to frame the question, and how to code responses. In the Many Labs 2 replication, the protocol was fixed and transparent, removing those degrees of freedom. That alone could account for the discrepancy.

Finally, the original effect may have been real but context-dependent. The 1979 participants were Canadian undergraduates in the late 1970s. The Many Labs 2 sample was more diverse — participants from 20 countries, different ages, and different educational backgrounds. The framing effect might be genuine but only under specific conditions that are not well understood. Replications that deviate from the original setting may miss those conditions, but they also test the generalizability of the claim.

A further nuance involves the exact wording of the scenario. The original study used a hypothetical disease outbreak affecting 600 people, with treatments described in terms of lives saved or lives lost. Some researchers have argued that the effect is stronger when the numbers are more vivid or when the decision is more consequential. For example, a 2015 study by Druckman and McDermott found that framing effects in political communication depend on the audience's level of knowledge and the credibility of the source. This suggests that the framing effect is not a single, fixed phenomenon but a family of effects that vary across contexts. The Many Labs 2 replication tested only one specific version, leaving open the possibility that other versions might replicate better. However, the burden of proof now lies on proponents to show which conditions produce robust effects.

What the Replication Crisis Revealed About Decision Research

The framing effect was not the only classic finding to fail replication. The Many Labs 2 project tested 28 effects; only about half replicated clearly. Other large-scale replication efforts, such as the Open Science Collaboration's 2015 project, found that only about 36% of social psychology studies replicated. The average effect size in replications was roughly half that of the original studies.

Decision research — the study of heuristics and biases — was hit particularly hard. Besides framing, effects like the anchoring bias, the endowment effect, and the sunk-cost fallacy all showed weaker or inconsistent replication. Critics argued that behavioral economics had built a reputation on shaky empirical foundations. Defenders countered that the core theories remained intact, but the flashy demonstrations were less reliable than advertised.

The crisis forced the field to confront uncomfortable truths. Statistical power in many studies was far too low. A typical social psychology study in the 2000s had a median power of roughly 35% to detect a medium effect. That meant most published findings were either false positives or inflated. The incentive structure of academia — publish or perish — encouraged novel, positive results over careful, incremental work.

But the crisis also spurred reform. Journals began to require pre-registration and to offer badges for open practices. Funders started to reward replication studies. Researchers formed collaborative networks like the Psychological Science Accelerator to conduct large-scale replications. The field slowly moved toward more rigorous standards.

One concrete reform is the adoption of "registered reports," where a study's design and analysis plan are peer-reviewed before data collection, and the paper is accepted in principle regardless of the outcome. This eliminates publication bias and incentivizes high-quality methodology. As of 2023, over 300 journals offer registered reports, and the format is gaining traction in behavioral science. Another reform is the use of meta-analytic thinking: rather than treating each study as a standalone test, researchers pool results across studies to estimate effect sizes more precisely. For example, a 2020 meta-analysis of framing effects by Steiger and Kühberger, including 67 studies, found an overall effect of d ≈ 0.25, much smaller than the original but still non-zero. This suggests that the framing effect exists but is modest and context-dependent.

How Kahneman's Response Changed Norms

Kahneman's reaction to the replication failure was unusual for a senior scientist. He did not dismiss the result or question the replicators' methods. Instead, he published a thoughtful commentary that acknowledged the problem and called for systemic change. He wrote: "I believe that the replication effort is a valuable corrective to the culture of the field."

He also advocated for "adversarial collaboration" — a practice in which researchers with conflicting views work together to design studies that can resolve their disagreements. This approach, which Kahneman had already used in his own work, reduces the influence of confirmation bias. If proponents and skeptics jointly design a test, the result is more credible to both sides.

In 2019, Kahneman co-founded the Replication Network, a group of researchers dedicated to conducting high-quality replications of important findings in behavioral science. The network's goal was not to debunk but to establish which effects are robust enough to inform policy and practice. Kahneman's involvement gave the movement credibility and encouraged younger researchers to prioritize replication.

Perhaps most importantly, Kahneman modeled intellectual humility. He admitted that his own work had contributed to a culture that valued novelty over reliability. In interviews, he said that he regretted not having pushed for replications earlier. His honesty increased trust in the field at a time when many were skeptical of psychology's methods.

To see how Kahneman's response contrasts with other senior scientists, consider the case of social priming. In 2012, a series of high-profile replication failures challenged the existence of effects like the "elderly priming" effect (where walking slowly is supposedly triggered by reading words related to old age). Some original authors defended their findings aggressively, questioning the replicators' competence. This led to protracted debates and eroded public trust. Kahneman's approach — accepting the data and calling for structural solutions — avoided such acrimony and set a constructive example. His willingness to admit fallibility may have been easier because prospect theory had many other pillars of support, but it nonetheless represented a cultural shift in how senior scientists respond to replication failures.

Practical Lessons for Today's Behavioral Science

The story of Kahneman's first replication failure offers several concrete lessons for researchers and practitioners. First, always pre-register study designs and analysis plans. Pre-registration separates confirmatory from exploratory analyses and reduces the risk of false positives. Many journals now require it, and platforms like the Open Science Framework make it easy.

Second, use power analysis to determine sample sizes before collecting data. A study should have at least 80% power to detect the expected effect size. For small effects, this may require hundreds or thousands of participants. Online platforms like Mechanical Turk and Prolific make large samples affordable and fast.

Third, report effect sizes with confidence intervals, not just p-values. A p-value tells you whether an effect is statistically significant, but not how large or meaningful it is. Confidence intervals convey precision and allow readers to assess the plausibility of different effect sizes. The Many Labs 2 replication reported d ≈ 0.05 with a 95% confidence interval from roughly −0.02 to 0.12 — consistent with a null effect.

Fourth, encourage direct replications before using a finding to inform policy. Several governments, including the U.S. and U.K., have established behavioral insight teams that rely on behavioral science research. A single study, no matter how compelling, should not be the basis for a policy intervention. Replications should be routine, especially for high-stakes applications.

Finally, treat single studies as provisional. Science is a cumulative enterprise. The framing effect may be real in some contexts, but its size is far smaller than originally reported. The core insights of prospect theory remain valuable, but they should be applied with appropriate caution. As Kahneman himself said: "The lesson is that we should not trust our intuitions about the size or robustness of effects, even when they feel obvious."

The replication crisis in psychology is far from over. Many classic findings have not yet been rigorously retested. But the field is healthier than it was a decade ago, thanks in part to the honesty of senior scientists like Kahneman. The failure of a single experiment, when handled with integrity, can strengthen the entire enterprise.

Looking ahead, the next challenge is to build a cumulative science that systematically maps the conditions under which effects hold. For the framing effect, this means conducting large-scale experiments that vary the scenario, the population, and the outcome measure. The Many Labs 2 replication was a necessary first step, but it should be followed by studies that explore moderators. For example, a 2022 pre-registered study by Rothman and Salovey found that framing effects are stronger when the decision involves personal health rather than hypothetical diseases, with an effect size of d ≈ 0.30. This suggests that the original finding may have been context-specific but not entirely spurious. By embracing replication and moderation analysis, behavioral science can move beyond the crisis and build a more reliable knowledge base.

Related Articles