One Sieve Mesh Size Alters Five Foraminifera-Based Temperature Reconstructions
May 29, 2026 By Renu Shah

Foraminifera—tiny, shell-building protists—are among the most widely used thermometers in paleoceanography. By counting the species in a sediment sample and applying a transfer function, researchers can estimate past sea-surface temperatures (SSTs) with a precision that has underpinned our understanding of glacial-interglacial cycles, ocean circulation changes, and millennial-scale climate variability. But a 2025 study from the University of Cambridge, led by Babette Hoogakker, has exposed a subtle methodological choice that can shift those estimates by as much as 2.5 °C: the mesh size of the sieve used to pick the shells.

One Sieve Size, Five Divergent Records

Hoogakker and colleagues re-analyzed five independent foraminifera-based SST records from North Atlantic sediment cores. All five had been published in peer-reviewed journals over the past decade, and each used a slightly different sieve mesh size for picking the fossils. The team applied a consistent correction based on the mesh size used in each original study. Three of the five records lost their millennial-scale variability—the wiggles that had been interpreted as rapid climate shifts—and two showed a reduced amplitude of glacial-interglacial cycles. The offset between the original and corrected SST estimates ranged from 1.2 °C to 2.5 °C, a range that, as Hoogakker noted in a presentation, is large enough to alter the interpretation of past ocean conditions.

The finding echoes a broader theme in paleoclimate science: that seemingly minor procedural details can propagate into large differences in reconstructed climate. A similar study on sediment sampling earlier this year showed that the depth interval used for sampling could shift SST estimates by up to 1 °C. The mesh-size effect is larger, and it affects a wider range of records because the sieve is used in nearly every foraminifera-based study.

Why Foraminifera Shells Record Ocean Heat

Planktic foraminifera build their calcium carbonate tests (shells) in the surface waters of the ocean. Different species thrive in different temperature ranges; for instance, Globigerina bulloides is common in cool, productive waters, while Globigerinoides ruber prefers warm, stratified conditions. When the organisms die, their shells sink and accumulate in seafloor sediments. By counting the relative abundance of each species in a sediment sample, paleoceanographers can estimate the mean annual SST at the time the shells were deposited.

The conversion from species counts to temperature is done via a transfer function, typically a statistical model calibrated using core-top samples (the uppermost sediment layer) and modern instrumental SST data. The most widely used calibration dataset, the MARGO (Multiproxy Approach for the Reconstruction of the Glacial Ocean) database, relies on a 150 μm sieve as the standard for picking foraminifera. That standard, however, has not been uniformly adopted. Many labs use a 125 μm or even a 63 μm sieve, especially when sediment samples are small or when they want to increase the number of specimens available for counting. The rationale is pragmatic: finer sieves capture more shells, reducing counting uncertainty. But as Hoogakker's study shows, that pragmatism comes at a cost.

The Mesh-Size Experiment That Revealed the Bias

Hoogakker et al. (2025) conducted a controlled experiment using the same sediment samples from three North Atlantic cores. They split each sample into three fractions: one sieved at 150 μm (the standard), one at 125 μm, and one at 63 μm. For each fraction, they picked and identified all foraminifera shells larger than the mesh size. The results were striking. The 63 μm sieve retrieved roughly 30% more juvenile shells (small, thin-walled specimens that are often excluded at coarser meshes). Those juveniles belonged predominantly to cold-water species such as Neogloboquadrina pachyderma and Turborotalita quinqueloba, which are small in adult size and produce even smaller juveniles. In contrast, warm-water species like G. ruber are larger and were already well represented in the 150 μm fraction.

The consequence is a systematic bias: finer sieves overrepresent cold-water species, leading to an artificially cool SST estimate. The magnitude of the bias depends on the temperature gradient at the site. In the subpolar North Atlantic, where cold and warm waters mix, the bias reached 2.5 °C. In tropical cores, where the species pool is dominated by large warm-water taxa, the bias was smaller—around 0.5 °C—but still non-negligible.

How the Bias Propagates Into Paleoclimate Curves

The bias is not constant through time. If a lab uses a 150 μm sieve for the top of a core and a 63 μm sieve for deeper sections (perhaps because the deeper sediment is harder to disaggregate), the resulting down-core SST curve will show an artificial cooling trend as the mesh size changes. Such shifts in mesh size are common in practice. Hoogakker's team surveyed 30 published foraminifera-based SST records from the North Atlantic and found that only 12 reported the sieve mesh size used; among those, six used a consistent mesh throughout the core, while the other six changed mesh size at some depth. The records that changed mesh size showed, on average, a larger difference between the original and corrected SST estimates than those that stayed consistent.

When the team re-analyzed the five published records, they applied a correction factor derived from their experimental data. The correction reduced the amplitude of glacial-interglacial cycles in two records by roughly 20–30%, and it eliminated millennial-scale variability in three others. For one core from the Norwegian Sea, the original record had shown a series of abrupt warming events during the last deglaciation, which had been linked to changes in ocean circulation. After correction, those events were barely distinguishable from noise. The implication is that some of the rapid climate shifts inferred from foraminifera-based SST records may be artifacts of changing mesh size, not real climate signals.

A similar issue was identified in a previous analysis of electrophysiology studies, where a seemingly minor change in equipment purchase rules altered outcomes. Here, the procedural detail is the sieve, and its effect is equally profound.

Peer Review Missed This Methodological Glitch

How did such a systematic bias escape detection for decades? Part of the answer lies in the way paleoceanographic studies are reviewed. Peer reviewers typically focus on the interpretation of the data—the correlation to other records, the alignment with ice-core or speleothem archives—rather than on the laboratory methods used to generate the data. Many reviews do not request the sieve mesh size because it is not considered a critical parameter. The original studies that Hoogakker re-analyzed passed peer review without any mention of mesh size in their methods sections. In some cases, the mesh size was reported in a supplementary table, but it was not discussed as a potential source of bias.

Another factor is that the calibration datasets themselves are heterogeneous. The MARGO database includes core-top samples processed with a range of mesh sizes, from 63 μm to 150 μm. When a transfer function is built from such mixed data, it implicitly averages over the mesh-size bias, but it does not remove it. A study that uses a 63 μm sieve and applies a transfer function calibrated on 150 μm data will inherit a systematic offset. Hoogakker's team showed that this offset can be as large as 1.5 °C in the calibration itself, meaning that even the reference SST estimates for core-top samples are not fully reliable.

The community is now grappling with the implications. At a recent workshop of the PAGES (Past Global Changes) Foraminifera Working Group, held in April 2025, a session was devoted to methodological standards. Several participants argued that all future studies should report the sieve mesh size, the number of specimens counted, and the size distribution of the picked shells. Others went further, calling for a re-analysis of existing records using a consistent 150 μm mesh or a correction factor. The debate is ongoing, and no consensus has yet emerged.

Trade-Offs and Counter-Arguments

While the bias is clear, some researchers question whether the correction factor is always appropriate. For instance, in regions where the foraminifera assemblage is dominated by a single species, the effect of mesh size may be smaller than in diverse assemblages. A study from the South Atlantic, published in 2023, used a 125 μm sieve and found that the species composition did not change significantly when re-sieved at 150 μm, possibly because the dominant species Globorotalia inflata has a large adult size that is captured efficiently at both meshes. This suggests that the bias may be site-specific, and a one-size-fits-all correction could introduce errors in some locations.

Another counter-argument is that the correction factor relies on size distributions measured in the North Atlantic, which may not apply globally. For example, in the equatorial Pacific, foraminifera shells tend to be smaller due to lower nutrient availability, so a 63 μm sieve might capture a higher proportion of adults rather than juveniles. Preliminary data from a 2024 cruise in the western Pacific warm pool indicate that the size distribution of G. ruber is shifted toward smaller sizes compared to the North Atlantic, which could alter the bias magnitude. Hoogakker acknowledges this limitation and has called for a global calibration effort, but until that is completed, the correction factor should be used with caution.

There is also a practical trade-off: using a standard 150 μm sieve reduces the number of specimens available, which increases counting uncertainty. For a typical sample, a 150 μm sieve may yield 200–300 specimens, whereas a 63 μm sieve can yield over 500. With fewer specimens, the statistical error in the species percentages increases, potentially masking real climate signals. This trade-off is particularly acute in sediment cores with low foraminifera abundance, such as those from the deep ocean or from periods of high dissolution. In such cases, the benefit of a finer sieve (more specimens, lower counting error) must be weighed against the risk of bias. Some researchers argue that the bias can be corrected after the fact, whereas low specimen counts cannot be remedied, so a finer sieve is preferable as long as a correction is applied.

A Practical Fix for Future Reconstructions

Hoogakker's study also offers a practical solution. The team has published an open-access correction code on GitHub that takes as input the sieve mesh size and the species counts, and outputs a corrected species assemblage that approximates what would have been obtained with a 150 μm sieve. The correction is based on a simple statistical model that accounts for the size distribution of each species, derived from the experimental data. The code is written in R and includes a vignette that walks users through the correction process. As of May 2025, the repository had been forked roughly 40 times, indicating interest from the community.

The PAGES Foraminifera Working Group has endorsed the correction method as a temporary standard while a more comprehensive calibration is developed. However, some researchers are cautious. They point out that the correction assumes that the size distribution of foraminifera shells is constant across the ocean, which is not true. In regions with high productivity, shells tend to be larger; in oligotrophic regions, they are smaller. The correction may work well for the North Atlantic, where it was calibrated, but its performance in other basins is unknown. A global calibration, using core-top samples from all major ocean basins, would be needed to fully resolve the bias.

In the meantime, the simplest fix is to use a consistent 150 μm sieve for all picking, as originally recommended by the MARGO community. That recommendation, however, has not been enforced, and many labs continue to use finer sieves out of habit or necessity. The cost of switching is not trivial: a 150 μm sieve captures fewer specimens, so longer cores or more samples are needed to achieve the same statistical confidence. For small-volume samples, such as those from deep-sea drilling expeditions, a finer sieve may be the only way to get enough shells for a reliable count. In those cases, the correction factor provides a workaround, but it is not a perfect substitute for standardization.

The broader lesson is that methodological transparency matters. A recent replication effort in social psychology showed that subtle differences in procedure can change outcomes; the same applies in paleoclimate. The mesh-size effect is one of many hidden variables that can distort reconstructions. Others include the number of specimens counted (typically 300–400, but sometimes fewer), the taxonomic resolution (species versus morphotypes), and the choice of transfer function (modern analog technique versus weighted averaging). Each of these choices introduces uncertainty, and the cumulative effect may be larger than the 2.5 °C bias seen here.

Named Examples and Specific Data Points

To illustrate the impact, consider the core MD99-2284 from the Norwegian Sea, which was included in the re-analysis. The original record, published in 2018, showed a rapid warming of about 4 °C between 14,500 and 13,000 years ago, interpreted as a collapse of the Scandinavian ice sheet. After correction for a mesh size change from 150 μm to 63 μm at a depth of 300 cm, the warming reduced to about 2 °C, and the timing became less abrupt. This suggests that the original interpretation may have been exaggerated by the mesh-size artifact.

Another example is core ODP-983 from the Iceland Basin, where the original SST reconstruction showed a pronounced cooling at the onset of the Younger Dryas (around 12,900 years ago) of about 6 °C. The correction, based on a consistent 125 μm sieve used throughout, reduced the cooling to about 4.5 °C, bringing it closer to estimates from alkenone-based proxies. This convergence across proxies lends support to the corrected foraminifera record.

Hoogakker's team also examined the effect on the last glacial maximum (LGM) SST estimates. For the North Atlantic, the original LGM cooling relative to modern was about 8–10 °C in the five records. After correction, the cooling was 6–8 °C, which is more consistent with climate model simulations. This is a critical finding because LGM SST estimates are used to evaluate the sensitivity of climate models to greenhouse gas forcing.

Future Directions and Community Response

The study has spurred several follow-up initiatives. The PAGES working group is planning a community-wide effort to re-analyze all North Atlantic foraminifera-based SST records using the correction factor, with a target completion date of 2027. A similar effort is being considered for the Southern Ocean, where foraminifera are less abundant but the bias could be significant due to the prevalence of small cold-water species. In addition, a group at the University of Bremen has begun collecting size-distribution data from core tops in the Atlantic, Pacific, and Indian Oceans to build a global calibration.

There is also discussion about incorporating the mesh-size correction into the next version of the MARGO database. If adopted, this would allow future studies to automatically account for the bias when using the database. However, this would require reprocessing thousands of core-top samples, a task that would take years and substantial funding. Some researchers have suggested that a simpler approach is to include the sieve mesh size as a variable in the transfer function, allowing the model to implicitly account for the bias. This idea is still in the conceptual stage.

Despite the challenges, the study has already changed practice in some labs. At the University of Cambridge, all new foraminifera picking is now done with a 150 μm sieve, and the correction factor is applied to any existing data that used a different mesh. Other labs, such as those at the University of Bordeaux and the University of Tokyo, have announced similar changes. The community is moving toward greater standardization, but the transition will take time.

Conclusion: A Call for Transparency

The Hoogakker study does not invalidate foraminifera-based paleothermometry. It does, however, underscore the need for rigorous, well-documented methods. As the climate community works to refine past temperature estimates, the sieve mesh size will not be overlooked again. The challenge now is to apply the correction to the hundreds of existing SST records, a task that will require collaboration among labs and a commitment to open data. Whether the community will adopt that commitment remains to be seen, but the path forward is clear: report the mesh, and where possible, use the standard.

Related Articles