Correlation is Not Causation: A Plain-English Guide for Health Leaders
Four common methodological traps, a real-world example, and the four questions to ask before acting on any correlational claim.
Every summer, the same story appears: ice cream sales spike, and so do drowning deaths. Graph them side by side and the two lines rise and fall in near-perfect lockstep. Is ice cream killing swimmers? Of course not. Both go up in summer because warm weather drives both. The correlation is real. The causation is not.
That example sounds trivial, but the same logical mistake happens every day in health — with real money and real lives attached. When someone hands you a finding that says "X is associated with Y," the question that decides whether you should act on it is far deeper than it looks: did X actually cause Y, or is something else going on?
The stakes are not abstract. Biomedical research loses an estimated $170 billion a year to irreproducible and invalid studies (Chalmers & Glasziou, 2009), and roughly 76% of clinical trials require costly protocol amendments — up to $535,000 per amendment in Phase III (Getz et al., 2016 — Tufts CSDD). While that waste traces back to a spectrum of methodological design flaws, one of the most pervasive is building conclusions on correlation that was mistaken for causation.
This guide provides four common examples of methodological traps researchers need to be aware of, a real-world example, and the four questions to ask before you act on any correlational claim.
Correlation vs. causation in one sentence
A correlation means two things move together. A causation means one thing actually makes the other happen. In health, confusing the two is not an academic footnote. It sends funding after interventions that don't work, it delays the ones that do, and it shapes treatment decisions on findings that do not hold up.
The classic health example: coffee and longevity
One widely-cited study after another finds that coffee drinkers live longer than non-drinkers (Liu et al., 2022; Kim et al., 2019). The correlation shows up in dataset after dataset. So does coffee cause a longer life?
Coffee drinkers differ from non-drinkers in ways that are often unaccounted for — either because the datasets lack the adequate variables, or because the researchers did not conduct the correct causal analysis. For example, they may have different income levels, different stress profiles, different exercise habits, or different jobs. Any of those hidden differences could be the real driver of the longevity, with coffee merely along for the ride.
Below, we outline four methodological traps that contribute to this avoidable research waste (Ioannidis et al., 2014; Rothman, Greenland, & Lash, 2008).
Four methodological traps
1. Confounding — a third, invisible factor drives both things. Ice cream and drowning share "summer." Vegetable intake and longevity share "people who eat vegetables also tend to exercise." When a hidden factor drives both sides of the equation, the visible correlation is an illusion. Confounding is a pervasive and costly trap in health research.
2. Lack of Pre-defined Protocols (Analytical Flexibility) — building the plane while flying it. Conducting research without detailed written plans allows researchers to improvise choices to achieve desired results. This massive data dredging without preselected hypotheses—often called HARKing (hypothesizing after results are known)—is a driver of irreproducible and invalid science. The trap is not knowing that you will find a significant finding just by chance, thinking you can figure out your exact causal structure after the data is collected.
3. Selection bias (Collider Bias) — who is missing from your dataset changes the answer. If you only study patients admitted to a hospital, you are looking at a fundamentally distorted sample if you want to infer about a broader population. For instance, among hospitalized patients, a strong correlation often appears between two completely unrelated diseases simply because having either disease increases the chance of admission. This structural flaw—conditioning on a collider—forces a correlation where none exists in a population beyond hospital admitted persons. To avoid this trap, your sample must be genuinely representative of your target population. Otherwise, your inclusion criteria will mathematically manufacture a false result.
4. Chance — random variation mistaken for a real signal. If you look at enough subgroups, enough outcomes, enough time windows, random variation will eventually produce something that looks "statistically significant." This is why one isolated dramatic finding should move you less than a consistent pattern across many independent studies.
How careful science actually answers the question
The gold standard is the randomized controlled trial: you assign people to treatment or control by chance, which addresses confounding by design — randomization balances measured and unmeasured confounders between groups in expectation. If the groups are comparable except for treatment, then a difference in outcomes is attributable to the treatment.
But for many of the questions health leaders actually face, a trial is unethical, impractical, or impossible. You cannot randomize people to smoking, or to a pandemic response policy, or to childhood poverty. That is precisely why modern epidemiology developed observational study designs: to develop and test hypotheses and build evidence along the causal pathway — data collected without randomization.
Applied rigorously, observational epidemiology answers the question the honest way: it doesn't manufacture certainty, it removes the easy ways to be wrong. That means a protocol and analysis plan locked in before the data are analyzed, transparent adjustment for known confounders, samples that don't silently condition on a collider, and findings weighed as patterns across many studies rather than single flashes.
Guardrails exist. They are not magic — they are method.
Every trap in this guide is a checkable rule: the invisible confounder, the improvised protocol, the collider hiding in your sample, the chance finding wearing a signal's clothes. And every decision in a protocol is a branching logic point — too many branches for intuition alone.
That's why the answer is not better templates but a different way of building: Studio by Outcome Project — one ecosystem where the protocol is written as structured logic, not prose, and deterministic rule engines check each link of the methodological chain as you write, the way an IDE checks code. The four traps don't need memorizing; they need checking. The AI assists; the deterministic checks decide. Validation you can trace, not a black box you must trust — and you stay the scientist.
Four questions to ask before you act on any study
The next time someone presents you with a striking health finding, work through these before committing budget, policy, or attention:
- Was there a randomized trial? If yes, weigh it heavily. If no, don't dismiss the study — many causal questions can only be answered with observational data — but demand what a trial would have guaranteed: a protocol and analysis plan locked in before the data were analyzed, not improvised after the results (HARKing).
- Which confounders were accounted for? The study should tell you openly what it adjusted for — and just as honestly, what it could not.
- Who could not get into the data? Selection and conditioning can manufacture correlation from your sample alone — hospital admission, healthy-user effects, any collider. Ask who is missing before believing what is present.
- Does it repeat? Is this one dataset's flash, or a pattern that holds across independent populations and methods?
Rigor is a product, not a virtue
At Outcome Project, methodological rigor is the difference between research that impresses and research that informs — between a beautiful analysis and a decision you can defend.
- For researchers: build protocols inside Studio — try it free at studio.outcomeproject.com — and see your design checked as you write it.
- For institutions: see the environment applied to your own teams with a demo for your organization.
- For public health authorities: our Squad team runs specialized epidemiological services — surveillance and causal inference from real-world data — built on exactly these principles: outcomeproject.com/services/outcome-squad
Leadership decisions rest on causation, not coincidence. We help you get there.
References
- Chalmers I, Glasziou P. Avoidable waste in the production and reporting of research evidence. The Lancet. 2009;374(9683):86–89.
- Getz KA, Stergiopoulos S, Short M, et al. The impact of protocol amendments on clinical trial performance and cost. Therapeutic Innovation & Regulatory Science. 2016;50(4):436–441. (Tufts Center for the Study of Drug Development)
- Getz K, Smith Z, Botto E, Murphy E, Dauchy A. New benchmarks on protocol amendment practices, trends and their impact on clinical trial performance. Therapeutic Innovation & Regulatory Science. 2024;58(3):539–548.
- Ioannidis JPA, Greenland S, Hlatky MA, et al. Increasing value and reducing waste in research design, conduct, and analysis. The Lancet. 2014;383(9912):166–175.
- Kim Y, Je Y, Giovannucci E. Coffee consumption and all-cause and cause-specific mortality: a meta-analysis by potential modifiers. European Journal of Epidemiology. 2019;34:731–752.
- Liu D, Li ZH, Shen D, et al. Association of sugar-sweetened, artificially sweetened, and unsweetened coffee consumption with all-cause and cause-specific mortality. Annals of Internal Medicine. 2022;175(7):909–917.
- Rothman KJ, Greenland S, Lash TL. Modern Epidemiology. 3rd ed. Philadelphia: Lippincott Williams & Wilkins; 2008.