Correlation vs causation
Correlation is a measured tendency for two things to vary together, and treating it as proof of cause is the commonest way a claim passes untested. Causation is the stronger relation, one of the two things making the other happen, and a correlation on its own cannot tell you which of the two you are looking at. The distinction is the single most useful thing to know about reading numbers, because a correlation is easy to compute, easy to publish and easy to describe in a sentence that quietly upgrades it into a cause. The upgrade usually happens in the verb: two things are associated in the data and become linked, then affect, then cause, with no new evidence arriving between one word and the next.
The 4 alternatives to A caused B
4 alternatives stand between a correlation and a cause, and every one of them has to be ruled out before the causal claim is earned. Working through the list takes about a minute and it is the difference between reading a statistic and being read by one.
- Reverse direction. B causes A. Offices that install more coffee machines record longer working hours, and the plausible reading is that long hours got the machines installed, not that the machines produce the hours.
- A third factor causing both. C causes A and B independently, which leaves them moving together with no connection at all. Ice cream sales and swimming pool rescues rise together every year because hot weather drives both, and neither one touches the other.
- Chance. With enough variables compared against enough other variables, strong correlations appear by arithmetic necessity. A dataset with 100 columns contains 4,950 possible pairs, and some of them will track each other closely for no reason whatsoever.
- Selection. The correlation exists in the sample and not in the world, because of how the sample was assembled. A conservatory admitting students who are strong in either mathematics or music will find a negative correlation between the two abilities among its own students even if there is none in the wider population, since anyone weak at both was never admitted. Statisticians call this a selection effect, and the classic version is known as Berkson's paradox.
2 of these 4 are usually skipped. Reverse direction gets checked, because it is intuitive, and a third factor gets checked, because everyone has heard the word confounder. Chance and selection do the quiet damage, and they are the two that no amount of thinking about the mechanism will detect, because both are properties of how the data were gathered rather than of what the data describe.
Illusory correlation: the pattern that is not in the data at all
An illusory correlation is a perceived relationship between two variables that are in fact unrelated, and it is the failure mode that sits one step earlier than everything above. In the four alternatives the correlation is real and the cause is in doubt. Here there is nothing to explain, because the association exists only in the observer. The psychologists Loren Chapman and Jean Chapman documented the effect in the 1960s, showing that observers reported associations in test materials that contained none, and that the associations they reported were the ones they already expected to find.
Two mechanisms produce it. The first is expectancy: a pairing that fits an existing belief is noticed and stored, and the cases that fail to fit are processed as unremarkable and not stored at all. The second is the rarity of joint events. Two uncommon things occurring together are memorable precisely because both are uncommon, so a handful of striking co-occurrences leaves a much stronger impression than a long run of ordinary non-occurrences. The result is a confident report of a pattern that a tally would not support, which is why the tally is worth making. Write down all four cells, not just the one that stands out: the times both happened, the times only the first happened, the times only the second happened, and the times neither did. Most reported correlations in daily life are built from the first cell alone.
What would settle it
To settle whether a correlation reflects causation, look for evidence that closes off the four alternatives, in roughly this order of strength.
- Intervention. Change A deliberately, assign the change at random, and see whether B follows. Random assignment is what breaks the link between the groups and every third factor, including the ones nobody thought of.
- Temporal order. Establish that A reliably comes first. This is necessary and nowhere near sufficient, but it does dispose of reverse direction.
- Gradient. Check whether more A produces more B in a consistent way. A dose response pattern is hard for a third factor to imitate, though not impossible.
- Mechanism. Identify a route by which A could produce B, and test a prediction that route makes and the alternatives do not.
- Natural experiment. Find a case where something outside the system changed A for reasons unconnected to B, such as an administrative boundary or a scheduling accident, and compare across it.
Where an experiment is impossible, the causal claim is built by ruling out alternatives one at a time, which is slower and never quite finishes. That is not a reason to reject the claim. It is a reason to state it with the strength the evidence actually supports, and what counts as evidence at each of those stages has a settled answer that does not change according to how badly the conclusion is wanted.
Correlation is not nothing
A correlation is not proof of a cause, and it is also not noise to be waved away, which is the opposite error and the more fashionable one. 3 things a correlation does honestly:
- Predicts. A stable association forecasts without explaining, and forecasting is often all that is needed.
- Constrains. A causal hypothesis that predicts an association nobody can find is in trouble, so the absence of a correlation is informative even when its presence is not decisive.
- Points. An unexpected association is where investigations start, and demanding a mechanism before anyone is allowed to look would stop the process at the beginning.
The phrase correlation does not imply causation, repeated as a closing argument, has become a way of avoiding the work rather than doing it. It states the problem accurately and then leaves the four alternatives unexamined. Occam's razor is the more useful instrument at this point, and it is worth using correctly: when a plain explanation, such as a shared cause or the arithmetic of many comparisons, accounts for the observed association as well as an elaborate one does, prefer the plain one until the elaborate one predicts something the plain one cannot.
The test to run on any statistic you are shown
Run the test on any statistic you are shown by asking 4 questions, and ask them before deciding whether you like the conclusion.
- Name the direction. If the arrow ran the other way, would the same data look exactly like this?
- Hunt the third factor. What would cause both of these at once, and was it measured?
- Count the comparisons. How many relationships were examined before this one was reported?
- Ask who is in the sample. Who had to be excluded for this group to exist, and could the exclusion alone produce the pattern?
The third question is the one most often unanswerable from the source in front of you, which is itself the finding. That is where the chain leads back to the record: whether you are looking at the original analysis or at a description of it, and how many retellings sit between the two. Primary and secondary sources differ exactly here, because the conditions attached to a result are the first thing a summary drops and the confidence is the last.