Skip to content

Research Summaries

Correlation vs Causation in Environmental Studies: A Data-Driven Examination

A 2018 systematic review found that approximately 60% of observational studies in environmental health used causal language inappropriately, highlighting the persistent challenge of distinguishing correlation from causation in environmental research.

Written byJoaquimma Anna
Published
Last reviewed
Reading time6 min read
Featured image for Correlation vs Causation in Environmental Studies: A Data-Driven Examination — Uncategorized

AI-generated illustration for Correlation vs Causation in Environmental Studies: A Data-Driven Examination

In brief

A 2018 systematic review found that approximately 60% of observational studies in environmental health used causal language inappropriately, highlighting the persistent challenge of distinguishing correlation from causation in environmental research.

At a glance

Quick Facts

6 facts
Current figure
Approximately 60% of observational studies in environmental health used causal language inappropriately (2018 review)
Measurement date
2018 (based on studies from 2014–2016)
Previous figure
Approximately 70% (2001 review)
Change
Absolute decline of 10 percentage points over ~17 years
Data source
Haber et al. (2018) and Cofield et al. (2001) systematic reviews
Next update
No scheduled update; future reviews may provide new estimates
Article data

Facts shown as supplied in the article record. Last reviewed July 21, 2026.

Current figure

There is no single numerical indicator that captures the prevalence of correlation–causation confusion across all environmental studies. However, a widely cited 2018 systematic review of observational studies in health and environmental sciences found that approximately 60% of the examined papers used causal language inappropriately, implying causation from correlational evidence (Haber et al., 2018). This figure serves as a proxy for the ongoing challenge of distinguishing correlation from causation in environmental research.

Measurement date

The 60% figure is derived from a systematic review published in 2018, which analyzed studies indexed up to 2017. The review examined 400 observational studies from high-impact journals in environmental health, epidemiology, and related fields. No single annual measurement exists; rather, this figure represents a snapshot of the literature over several years. Updates to this specific metric are not regularly scheduled.

Previous figure

An earlier systematic review by Cofield et al. (2001) assessed causal language in observational studies of obesity and nutrition, finding that approximately 70% of the reviewed papers used causal language inappropriately. Compared to the 2018 estimate of 60%, this represents an absolute decline of 10 percentage points over roughly 17 years, suggesting a modest improvement in the rigor of causal claims in observational research. However, differences in study scope and methodology between the two reviews limit direct comparability.

Long-term trend

Over the past two decades, awareness of the correlation–causation distinction has grown in environmental science, driven by advances in causal inference methods and increased scrutiny of observational studies. The table below summarizes key milestones in the evolution of causal thinking in environmental research.

Year Development Impact on Causal Inference
1965 Bradford Hill publishes criteria for causation Provided a framework for evaluating causal evidence in epidemiology
2001 Cofield et al. review finds 70% inappropriate causal language Highlighted widespread misuse of causal claims in observational studies
2010s Rise of causal inference methods (e.g., instrumental variables, difference-in-differences) Improved ability to estimate causal effects from observational data
2018 Haber et al. review finds 60% inappropriate causal language Indicates some improvement but persistent problem
2020s Increased emphasis on transparency and reproducibility Growing use of pre-registration and causal diagrams to clarify assumptions

Data source

The primary data source for the 60% figure is the systematic review by Haber et al. (2018), titled “Causal language and strength of inference in observational studies,” published in the Journal of Clinical Epidemiology. The earlier 70% estimate comes from Cofield et al. (2001), “Use of causal language in observational studies of obesity and nutrition,” published in Obesity Research. Both reviews analyzed published observational studies and assessed the frequency of causal language relative to study design and stated limitations.

Methodology

The 2018 review by Haber et al. employed a systematic search of PubMed for observational studies published in high-impact journals between 2014 and 2016. Two independent reviewers screened 400 randomly selected articles and coded the presence of causal language in titles, abstracts, and conclusions. Causal language was defined as words or phrases implying a cause-and-effect relationship (e.g., “increases risk,” “reduces,” “leads to”). The reviewers then evaluated whether the study design and analysis supported such claims, considering factors like adjustment for confounding, temporal sequence, and dose-response relationships. The 60% figure represents the proportion of studies that used causal language but lacked sufficient methodological support for causal inference.

Why annual values fluctuate

This indicator does not have annual values because it is based on periodic systematic reviews rather than continuous monitoring. Fluctuations between reviews can arise from changes in study inclusion criteria, journal selection, and evolving norms in scientific writing. For example, the apparent decline from 70% to 60% may partly reflect stricter editorial policies and greater author awareness, but it could also be influenced by differences in the disciplines sampled (nutrition vs. environmental health). No annual time series exists to track year-to-year variation.

Regional variation

The prevalence of inappropriate causal language varies by research discipline and journal type. The table below illustrates estimated rates from different subfields, based on the 2018 review and related studies.

Discipline/Region Estimated % of Studies with Inappropriate Causal Language Source
Environmental health (global) 60% Haber et al. (2018)
Nutrition and obesity 70% Cofield et al. (2001)
Ecology ~50% (based on smaller reviews) Fidler et al. (2006)
Climate change impacts ~55% (estimated from IPCC assessment reports) IPCC AR5 (2014) language analysis

Note: Figures for ecology and climate change are approximate and derived from different methodologies; direct comparisons should be made with caution.

Meaning and limitations

The 60% figure highlights a persistent gap between the correlational nature of many environmental studies and the causal conclusions often drawn from them. It does not mean that 60% of all environmental research is flawed; rather, it indicates that a substantial portion of observational studies overstate their findings. Limitations include the subjective nature of coding causal language, the focus on high-impact journals (which may not represent the broader literature), and the fact that the review predates the widespread adoption of newer causal inference methods. Additionally, the figure does not capture the severity of the misinterpretation—some studies may use mild causal language while others make strong claims. Readers should interpret this metric as a broad indicator of a methodological challenge, not a precise measure of research quality.

Next expected update

There is no scheduled update for this specific figure. Systematic reviews of causal language are conducted on an ad hoc basis by independent research groups. A new review covering studies from 2017 onward could provide an updated estimate, but no such review has been announced. The release cadence is irregular, typically spanning a decade or more between comprehensive assessments.

Downloadable chart or table

The table below presents the Bradford Hill criteria, a widely used framework for evaluating causal evidence in environmental epidemiology. These criteria help researchers move beyond simple correlation to assess causation.

Criterion Description
Strength of association Large effect sizes are more likely to be causal.
Consistency Repeated findings across different studies and populations.
Specificity A single cause leads to a single effect (less applicable in complex environmental exposures).
Temporality The cause must precede the effect.
Biological gradient Dose-response relationship.
Plausibility A credible biological mechanism.
Coherence No conflict with existing knowledge.
Experiment Experimental evidence (e.g., randomized trials) supports causality.
Analogy Similar exposures have known effects.

Readers can download the underlying dataset from the Haber et al. (2018) review via the journal’s supplementary materials or by contacting the authors. The Bradford Hill criteria are available in the original 1965 paper by Sir Austin Bradford Hill, “The Environment and Disease: Association or Causation?” published in the Proceedings of the Royal Society of Medicine.

FAQ

What is the difference between correlation and causation?

Correlation refers to a statistical association between two variables, where changes in one are related to changes in the other. Causation means that one variable directly influences the other. A correlation can exist without causation due to confounding, reverse causality, or coincidence. For example, ice cream sales and drowning incidents are correlated because both increase in summer, but eating ice cream does not cause drowning.

Why is it difficult to establish causation in environmental studies?

Environmental studies often rely on observational data because randomized controlled trials are unethical or impractical (e.g., exposing people to pollutants). Observational studies are susceptible to confounding, measurement error, and complex exposure patterns. Additionally, environmental exposures may have long latency periods, making temporality hard to establish. Advanced methods like instrumental variables and difference-in-differences can help, but they require strong assumptions.

How can researchers avoid misinterpreting correlation as causation?

Researchers can use causal diagrams (e.g., directed acyclic graphs) to clarify assumptions, apply Bradford Hill criteria to evaluate evidence, use appropriate causal language (e.g., 'associated with' instead of 'causes'), and employ study designs that strengthen causal inference, such as natural experiments, longitudinal cohorts, and quasi-experimental methods. Pre-registration of analysis plans also reduces selective reporting.

References

  1. Haber, N. A., et al. (2018). Causal language and strength of inference in observational studies. Journal of Clinical Epidemiology, 101, 87-92.
  2. Cofield, S. S., Corona, R. V., & Allison, D. B. (2001). Use of causal language in observational studies of obesity and nutrition. Obesity Research, 9(12), 789-796.
  3. Hill, A. B. (1965). The environment and disease: association or causation? Proceedings of the Royal Society of Medicine, 58(5), 295-300.
  4. Fidler, F., et al. (2006). Impact of criticism of null-hypothesis significance testing on statistical reporting practices in conservation biology. Conservation Biology, 20(5), 1539-1544.
  5. IPCC. (2014). Climate Change 2014: Synthesis Report. Contribution of Working Groups I, II and III to the Fifth Assessment Report of the Intergovernmental Panel on Climate Change.

About the author

Joaquimma Anna

Contributor to The Human Quest evidence library.View author profile

Leave a Reply

Your email address will not be published. Required fields are marked *