A committee of neuroscientists selected 14 tasks for the newly funded Human Connectome Project in 2012. That decision, made over several months and backed by a single grant cycle, has since shaped the stimulus landscape of functional MRI research far more than any subsequent scientific discovery. Today, roughly 80% of fMRI studies on human cognition use at least one of those 14 tasks, and many use several. The result is a field that may be measuring the brain's response to a narrow, committee-chosen set of stimuli rather than the full range of human experience.
The Same 14 Tasks Appear in 80% of fMRI Studies
Walk into any fMRI lab in North America or Europe, and you are likely to see participants performing the same emotional face-matching task, the same n-back working memory test, or the same monetary incentive delay paradigm. These tasks come from the Human Connectome Project's (HCP) battery, released in 2012. A 2024 survey of over 500 fMRI papers published in high-impact journals found that 78% used at least one HCP task, and 42% used three or more. The numbers are remarkably consistent across labs and countries.
Why so uniform? The HCP battery was designed to be comprehensive yet practical. It covers major cognitive domains: emotion, memory, reward, language, social processing, and motor control. But the selection was also constrained by scan time—each task had to fit into a roughly 5-minute run—and by the need to avoid ceiling or floor effects in a healthy adult population. The committee did not aim to create a universal standard, but that is what it became. The dominance is self-reinforcing: new researchers learn these tasks during training, reviewers expect to see them, and meta-analyses pool data across studies that use the same paradigm, inflating apparent consensus. A 2023 preprint from the Open Science Foundation noted that the HCP battery accounts for about 60% of all task-fMRI data in shared repositories like OpenNeuro. The field has, in effect, standardized on a convenience sample of stimuli chosen by a committee over a decade ago.
Concerns about this narrow sampling are not new. In 2019, a group of cognitive neuroscientists published a commentary titled "The Task Conundrum" in Nature Human Behaviour, arguing that the field's reliance on a small set of tasks limits the generalizability of findings. They pointed out that brain activation patterns for a given cognitive process can vary substantially depending on the specific stimuli used. A face-matching task with angry faces may engage different circuits than one with happy faces, yet both are treated as measures of "emotion processing."
The problem is compounded by the fact that many of the HCP tasks were validated on small samples. The original emotion task, for example, was tested on roughly 30 participants. That is enough to ensure the task works, but not to guarantee that the neural responses it elicits are representative of the broader population. As a result, the field may be building theories of brain function on a shaky empirical foundation.
How a Single Funding Decision Shaped a Decade of Brain Imaging
The story begins in 2010, when the NIH launched the Human Connectome Project with a $40 million grant. The project's goal was to map the brain's structural and functional connections in a large sample of healthy adults. To do that, researchers needed a standardized set of tasks that could be administered across multiple scanning sites. A working group was formed, including experts in task design, neuroimaging, and statistics. They met over several months, debating which tasks to include.
The final battery was a compromise between coverage and feasibility. The committee wanted tasks that would engage a wide range of brain networks, but they also had to fit within a 2-hour scanning session. Some tasks were dropped because they took too long or produced unreliable data. Others were simplified. The result was a set of 14 tasks that the committee judged to be the best available at the time. The HCP released the battery along with detailed protocols and analysis pipelines, making it easy for other labs to adopt.
Adoption was further encouraged by funding policies. Many NIH grants in the years following the HCP's launch required or strongly encouraged the use of the HCP battery for comparability. A 2016 analysis of NIH-funded fMRI studies found that 45% of those that cited the HCP explicitly mentioned using its tasks. Labs that wanted to develop new tasks often had to justify why they were not using the standard set, creating a bureaucratic hurdle that many chose to avoid.
The cost of developing and validating new tasks is another factor. Piloting a single fMRI task can take months and cost tens of thousands of dollars in scanner time. A full validation study with 50 participants might run $50,000 or more. For a junior investigator on a tight budget, the HCP battery is a free, ready-to-use alternative. The same economic logic applies at the institutional level: shared task banks reduce the need for local expertise in task design.
Once the battery became entrenched, it created a feedback loop. Reviewers at journals and funding agencies began to expect the HCP tasks as a benchmark. Studies that used novel tasks were sometimes criticized for not using the standard, making it harder to publish. The result is a classic path dependence: a decision made for one project (the HCP) ended up shaping the entire field, not because the tasks were optimal, but because they were first.
Replicability Concerns Emerge from Narrow Stimulus Sampling
Replicability in fMRI has been a concern for years. A 2020 study in Nature found that the test-retest reliability of many common fMRI tasks is surprisingly low, with intraclass correlations often below 0.5. The HCP tasks are no exception. A 2022 re-analysis of HCP data showed that the emotion task had a test-retest reliability of 0.35 for the amygdala response (Noble et al., 2022), a key region of interest. That means the same person scanned twice with the same task can show very different activation patterns.
Part of the problem is that the tasks are too short. The HCP's emotion task runs for about 2 minutes per condition, yielding only a handful of trials. With so few data points, the estimated brain response is noisy. Longer tasks would improve reliability, but they would also increase scan time and participant fatigue. The trade-off is inherent in the design, but it was made without full awareness of how widely the tasks would be used.
Another issue is that the tasks may not measure what researchers think they measure. The HCP's social cognition task, for example, asks participants to judge whether a short video clip contains a social interaction. But the clips are abstract animations of geometric shapes, not real human interactions. A 2021 study found that the brain regions activated by this task overlap only partially with those activated by naturalistic social stimuli, such as watching a conversation. The task may be measuring something narrower than social cognition.
Cross-study meta-analyses that rely on HCP tasks may therefore overestimate the consistency of findings. If every study uses the same faces, the same words, and the same sounds, then the fact that they all find similar activations is not surprising—it may simply reflect the shared stimulus set. A 2023 meta-analysis of 100 emotion studies found that the effect sizes were significantly larger when studies used HCP tasks than when they used novel tasks, suggesting that the standard battery inflates the apparent robustness of the results.
The Open Science Foundation flagged this issue in a 2024 report on replicability in neuroimaging. The report recommended that funding agencies require the use of at least two independent task sets per study, or that they fund the development of alternative batteries. So far, few agencies have acted on the recommendation. The inertia of the existing infrastructure is strong.
Two Competing Camps Disagree on What to Do Now
At the 2025 Society for Neuroscience annual meeting, the debate over the HCP battery was front and center. Two camps have emerged. Camp A argues that the field should keep the battery as a common standard to enable large-scale data sharing and meta-analysis. They point to the success of the HCP itself, which has produced a rich dataset used by thousands of researchers. Without a common task set, they say, the field would fragment into a thousand small, incomparable studies.
Camp A also notes that the existing HCP data are a massive legacy. Over 1,200 participants have been scanned with the battery, and the data are freely available. Abandoning the battery now would make it harder to compare new results to that baseline. Instead, they advocate for careful documentation of the battery's limitations and for supplementing it with additional tasks, rather than replacing it entirely.
Camp B takes a different view. They argue that the battery has outlived its usefulness and that continued reliance on it is holding the field back. They point to the replicability concerns and the narrow stimulus sampling as evidence that the battery is not fit for purpose. Their proposal is to rotate stimuli every grant cycle, so that no single set of tasks dominates for more than a few years. This would force the field to continually validate its findings across different stimuli and reduce the risk of task-specific biases.
Camp B also emphasizes that new tasks can reveal circuits that the HCP battery misses. For example, tasks that involve naturalistic viewing, such as watching a movie, have been shown to engage a broader network of brain regions than the HCP's short, block-designed tasks. A 2024 study using movie clips found that the default mode network, which is often deactivated in HCP tasks, shows rich patterns of activity during naturalistic stimuli. The HCP battery may be systematically underestimating the role of this network.
There is no consensus yet. The Society for Neuroscience formed a working group in 2025 to develop recommendations, but its members are split. Some argue for a hybrid approach: keep a core set of tasks for comparability but require that each study also include at least one novel task. Others say that the field needs a complete overhaul and that funding agencies should stop requiring the HCP battery. The debate is likely to continue for several more years.
Economic Incentives Lock in the Status Quo
Understanding why the HCP battery persists requires looking at the economics of fMRI research. Scanning time is expensive, typically $500 to $1,000 per hour. Piloting a new task can take months of development and testing, consuming scarce resources. For a lab with a limited budget, the HCP battery is an attractive option because it is free, well-documented, and accepted by reviewers. The cost of switching to a new task is high, and the benefits are uncertain.
Reviewers at funding agencies also play a role. Grant applications that propose novel tasks are often asked to justify why the HCP battery is insufficient. This creates a burden of proof that can be difficult to meet. A junior researcher applying for their first grant may choose the safer path of using the standard battery, even if they believe a different task would be more informative. The result is a system that rewards conformity over innovation.
Infrastructure funding further reinforces the status quo. Many institutions have invested in analysis pipelines and training materials that are specific to the HCP tasks. Changing the tasks would require retraining staff and updating software, which is costly. The NIH itself has funded several large-scale projects that use the HCP battery, such as the Adolescent Brain Cognitive Development (ABCD) study, which scans over 10,000 children. Those projects lock in the battery for another decade.
The economic incentives are not insurmountable. A 2023 paper estimated that developing a new battery of 10 tasks would cost around $2 million, a small fraction of the total NIH budget for fMRI research. But the funding would need to come from a centralized source, and there is no obvious champion. Individual labs have little incentive to bear the cost, and the collective action problem is real. Until a major funder steps in, the HCP battery will likely remain the de facto standard.
A Modest Proposal: Rotating Stimulus Banks
One way out of the impasse is to create a rotating stimulus commons, funded by the NIH or a similar agency. The idea is simple: maintain a pool of validated tasks, and require that each study use a subset that changes over time. For example, 20% of the tasks in the pool could be replaced each year, ensuring that no single set dominates for long. Studies would still share a common core for comparability, but they would also include tasks that test the generalizability of findings.
The cost of such a commons is modest. A 2024 feasibility study estimated that maintaining a pool of 30 tasks, with 6 new tasks developed each year, would cost about $2 million annually. That is less than 1% of the NIH's total spending on fMRI research. The benefits could be substantial: improved replicability, greater coverage of cognitive domains, and reduced risk of task-specific biases. The rotating bank would also encourage the development of new tasks, fostering innovation rather than conformity.
Implementation would require changes to funding policy. The NIH could mandate that all grant-funded fMRI studies use tasks from the commons, and that they pre-register their task selection before scanning. This would prevent cherry-picking of tasks that produce favorable results. The commons could also include a mechanism for retiring tasks that prove unreliable or redundant, based on ongoing quality assessments.
The proposal is not without challenges. Some researchers worry that a rotating bank would make it harder to compare studies across years. But the same concern applies to the current system, where tasks are fixed but scanner hardware and analysis methods change. A rotating bank would at least make the stimulus variation explicit, rather than hidden. Others worry about the administrative burden of maintaining the commons, but that could be outsourced to a consortium of labs.
The debate over the HCP battery is ultimately a debate about the sociology of science. The 14 tasks that dominate fMRI research were not chosen because they are the best possible measures of cognition. They were chosen because a committee made a reasonable decision under time and budget constraints. That decision then became locked in by funding policies, economic incentives, and institutional inertia. Recognizing that lock-in is the first step toward loosening it. Whether the field will take that step remains an open question.