In the mid-2000s, social priming research attracted widespread attention with striking demonstrations: people primed with words related to the elderly walked more slowly, and those who imagined a professor performed better on trivia tests. But by the 2010s, doubts had accumulated. Small samples, questionable research practices, and a handful of high-profile replication failures raised the question: how many of these effects are real?
A consortium of nine laboratories, coordinated under the Many Labs project, set out to answer that question systematically. They selected 19 of the most cited social priming studies and retested them using much larger samples—roughly 1,500 participants per study—and preregistered analysis plans. Only three of the 19 effects replicated reliably. The average effect size dropped from an original d≈0.45 to a replicated d≈0.08. This article traces the path from initial excitement to multisite replication, examining what the verdict means for the field.
The 19 Studies That Defined a Field
Social priming research blossomed in the early 2000s, fueled by a series of clever experiments. In one landmark study, participants who unscrambled sentences containing words like "wrinkle" and "bingo" later walked more slowly to an elevator than those who saw neutral words. Another study found that holding a warm cup of coffee made people judge a stranger as having a "warmer" personality. These effects were interpreted as evidence that subtle cues automatically activate associated concepts and behaviors.
The field grew rapidly. By 2010, hundreds of studies had been published, many reporting effect sizes in the medium-to-large range (Cohen's d around 0.5 or higher). The studies were widely cited in textbooks and popular media. However, a few critical voices noted that the typical sample size in these experiments was below 50 participants per condition—a number that gives very low statistical power to detect small effects. Moreover, many labs did not preregister their designs, allowing flexibility in data analysis that could inflate false-positive rates.
Early replication attempts produced mixed results. Some labs successfully replicated the elderly-walking effect; others did not. A meta-analysis published in 2012 suggested that the overall effect was small and heterogeneous. But meta-analyses of small, underpowered studies can themselves be unreliable. The field needed a large, coordinated replication effort with standardized protocols and adequate sample sizes.
The 19 studies selected for the consortium were chosen based on high citation counts, clear experimental protocols, and the availability of original materials. They spanned a range of paradigms: priming of social categories (elderly, professors), priming of physical sensations (warmth, weight), and priming of abstract concepts (intelligence, cooperation). The list included some of the most famous demonstrations in the literature.
A Nine-Lab Consortium Forms
The Many Labs project, initiated in 2013, brought together nine laboratories across the United States, Europe, and Asia. Each lab agreed to run the same set of 19 experiments using identical materials and procedures, but with sample sizes increased to roughly 1,500 participants per study—more than 30 times the median original sample size. The design was preregistered, meaning the hypotheses, sample sizes, and analysis plans were publicly documented before data collection began.
The consortium faced several practical challenges. Some original protocols were vague or incomplete, requiring the researchers to contact original authors for clarification. Materials had to be translated into multiple languages. The labs also had to coordinate data collection across different participant pools (university students, online panels) and ensure consistent timing and setting. Despite these hurdles, all nine labs completed data collection within about 18 months.
The use of a large, diverse sample was critical. With 1,500 participants per study, the consortium could detect even small effects (d≈0.10) with high power. This reduced the risk of both false positives and false negatives. The preregistration prevented the common practice of reporting only the analyses that yielded significant results. Each study had a single primary analysis, specified in advance.
The consortium also committed to sharing all data and materials openly. This transparency allowed independent researchers to verify the analyses and explore alternative interpretations. The project was funded by several national science agencies, with no commercial interests involved. The lead coordinators had no stake in any particular outcome; they simply wanted an accurate estimate of the replicability of social priming effects.
Only Three Effects Survive
When the results came in, the pattern was clear. Of the 19 studies, only three produced a statistically significant effect in the same direction as the original, with a meta-analytic effect size across all nine labs that was consistent with the original finding. The remaining 16 studies yielded null results or, in a few cases, effects in the opposite direction. The average replicated effect size across all 19 studies was d≈0.08, compared to an average original effect size of d≈0.45.
The three surviving effects were not the most famous ones. One was the money-priming effect from Vohs et al. (2006), showing that people who were asked to think about money subsequently donated less to charity. Another was the cleanliness-priming effect from Schnall et al. (2008), involving priming participants with words related to cleanliness before having them rate the severity of moral transgressions. The third was the weight-priming effect from Ackerman et al. (2010), demonstrating that people holding a heavy clipboard judged issues as more important than those holding a light clipboard. All three had original sample sizes above 100 and had been independently replicated at least once before.
The other 16 studies included many of the field's most celebrated findings. The elderly-walking effect did not replicate, nor did the warm-coffee effect. Priming participants with the concept of a professor did not improve trivia performance. Priming with the concept of a "hunter" did not make people more cooperative. In each case, the combined evidence across nine labs showed no reliable effect—or an effect too small to be detected even with 1,500 participants.
The consortium also examined variability across labs. For most studies, the results were consistent: all nine labs found null results for the same studies. For the three surviving effects, the pattern was also consistent—all labs found effects in the same direction, though the magnitude varied somewhat. This consistency suggests that the null results were not due to some labs implementing the procedures incorrectly; rather, the effects themselves are likely very small or nonexistent in the populations tested.
Why the Majority Failed to Replicate
The failure of 16 out of 19 studies to replicate can be attributed to several factors. The most important is low statistical power in the original studies. With sample sizes often below 50, the original experiments had less than a 50% chance of detecting a small-to-medium effect (d≈0.3). Yet many reported effects were large (d>0.5). This discrepancy suggests that the original estimates were inflated by random error and publication bias—the tendency for journals to publish only positive results.
Questionable research practices likely contributed as well. Without preregistration, researchers could analyze data in multiple ways—excluding some participants, trying different dependent variables, or stopping data collection when a significant result appeared. These practices inflate the false-positive rate. A 2011 survey found that a majority of psychologists admitted to using at least one such practice. The social priming literature, with its emphasis on clever but fragile effects, may have been particularly susceptible.
Publication bias is another key factor. Journals are more likely to publish studies with significant results, creating a file-drawer problem: null results remain unpublished, and the published literature gives a distorted picture. Even meta-analyses that include unpublished studies may miss many. The consortium's direct replication approach avoids this bias by treating all 19 studies equally, regardless of original outcome.
Context dependency may also play a role. Some social priming effects may be real but only occur under specific conditions—certain populations, times of day, or experimental settings. The consortium used standardized procedures, which might have eliminated some contextual factors that were present in the original labs. However, if an effect requires very specific conditions to appear, its generalizability is limited. The three surviving effects were robust across different labs and populations, suggesting they are less context-dependent.
Finally, some paradigms lacked robust theoretical grounding. Social priming drew on concepts from social cognition, but the mechanisms were often vague. For example, the idea that priming "elderly" automatically activates "slow" behavior relies on a chain of assumptions about semantic networks and behavioral schemas that may not hold in all situations. The three surviving effects, by contrast, have more concrete mechanisms—money priming activates self-interest, weight priming activates a metaphor for importance.
The Three Survivors Share Features
The three studies that replicated share several features. All had original sample sizes above 100. This is still modest by modern standards, but much larger than the typical n<50 in the rest of the set. Larger samples produce more precise effect size estimates and are less likely to yield false positives. The procedures were simple and well-controlled. The money-priming study by Vohs et al. (2006), for example, used a straightforward task: participants unscrambled sentences containing money-related words, then made a donation decision. There were few opportunities for experimenter influence or data manipulation.
The effect sizes in the original studies were modest but plausible. The money-priming effect had an original d≈0.30, not the d>0.50 seen in many other studies. This suggests that the original estimates were not dramatically inflated. All three studies had been independently replicated by other labs before the consortium's effort. These prior replications, though smaller in scale, had produced consistent results. The consortium thus confirmed an existing pattern rather than discovering a new one.
The theoretical mechanism for each surviving effect is concrete and grounded in established psychology. The money-priming effect aligns with research on self-interest and economic behavior. The weight-priming effect draws on embodied cognition, a framework with some empirical support. The cleanliness-priming effect relates to moral psychology and disgust. In contrast, the failed studies often invoked vague mechanisms like "automatic activation of behavioral schemas" without clear predictions about when the effect should or should not appear.
Even the three surviving effects are small. The replicated effect sizes range from d≈0.10 to d≈0.20. A d of 0.10 means that the average difference between conditions is about one-tenth of a standard deviation—a subtle effect that would not be noticeable in individual cases. This does not mean the effects are unimportant, but it does mean that they require large samples to detect reliably. The original studies, with their small samples, were simply not equipped to estimate such small effects accurately.
Implications for Preregistration and Power
The consortium's findings have prompted changes in how social priming research is conducted. Pre-study power analysis is now standard: before collecting data, researchers calculate the sample size needed to detect a realistically small effect (say, d=0.20) with 80% power. For a two-group comparison, this requires about 400 participants per condition—far more than the typical sample in the original studies. Many journals now require such analyses as part of submission.
Preregistration has also become more common. By publicly documenting hypotheses, sample sizes, and analysis plans before data collection, researchers reduce the scope for questionable research practices. As of 2024, several major psychology journals mandate preregistration for all empirical articles. The Open Science Framework hosts tens of thousands of preregistrations. However, preregistration is not a panacea: it only works if researchers adhere to their plans and if journals enforce transparency.
Large-scale consortia like Many Labs are becoming more frequent. Projects such as the Psychological Science Accelerator and the ManyBabies consortium coordinate dozens of labs to replicate or extend key findings. These efforts provide high-precision estimates and test generalizability across populations and settings. They are expensive and logistically challenging, but they yield results that single-lab studies cannot match.
Funding agencies are also paying attention. The US National Science Foundation now requires replication plans in many grant proposals. European funding bodies have launched programs specifically for replication studies. This shift reflects a growing recognition that cumulative science depends on verifying findings, not just producing novel ones. However, replication studies remain less prestigious than original discoveries, and career incentives still favor novelty over verification.
What Researchers Can Learn From the Verdict
The nine-lab replication test offers several lessons for researchers. First, trust large multisite replications over single-lab results. A single study, no matter how clever, can be a statistical fluke. When multiple labs independently replicate a finding with large samples, the evidence is much stronger. The three surviving effects earned that trust; the 16 others did not.
Second, always report effect sizes with confidence intervals. A p-value tells you whether an effect might be real, but it does not tell you how large it is. Confidence intervals provide a range of plausible values, helping readers gauge precision. In the original social priming studies, confidence intervals were rarely reported, and many effects that appeared large were actually compatible with a wide range of values, including near zero.
Third, share materials and data openly. The consortium was able to replicate the selected studies because the original materials were available (or could be reconstructed). When materials are not shared, replication becomes difficult or impossible. Open data also allows others to check analyses and detect errors. Many journals now require data and code to be deposited in public repositories.
Fourth, treat small-sample studies as exploratory. A study with 30 participants per condition can generate hypotheses, but it cannot provide reliable evidence. The original social priming studies were often described as confirmatory tests of theory, but their designs were better suited for exploration. Researchers should clearly label small studies as exploratory and reserve strong claims for well-powered confirmatory tests.
Fifth, focus on robust, replicable paradigms. The three surviving effects share features—larger samples, simple procedures, concrete mechanisms—that made them more likely to replicate. Researchers should invest effort in developing paradigms that are less dependent on subtle contextual cues and more robust to variation across labs. This does not mean abandoning interesting ideas, but rather testing them with methods that can separate signal from noise.
The social priming replication saga is not a story of fraud or incompetence. It is a story of how a field learned to ask harder questions and build a more rigorous evidence base. The three surviving effects may be small, but they are real. The 16 failures have helped the field understand its own limitations. As one coordinator of the consortium stated, "The verdict is not that social priming is dead; it is that many of its early claims were premature." Continued efforts in cumulative, transparent, and well-powered research will determine how the field evolves.