Grain of Salt

Regression to the mean: why the treatment seemed to work

Regression to the mean

Struck diagram on assay stock: Regression to the mean

Regression to the mean is the tendency for an extreme measurement to be followed by one closer to average, purely because part of what made it extreme was chance. It is not a force, a correction or a rebound. It is what happens whenever a measurement contains any random component at all, and it is the single most reliable way for a useless intervention to look effective.

The mechanism takes one sentence. Any measurement is the underlying quantity plus noise. When you select a case because its measurement was extreme, you have selected for a large helping of both, and the noise does not repeat. The next measurement keeps the underlying quantity and gets a fresh draw of noise, which is on average zero. So the second reading moves toward the average, and anything you did in between gets the credit.

What regression to the mean is

Regression to the mean is a statistical consequence of selecting on an unreliable measurement, first described by Francis Galton, who called it regression towards mediocrity in his 1886 work on the heights of parents and their children. Galton found that unusually tall parents had children who were tall but less unusually so, and unusually short parents had children closer to average in the other direction. Nothing was pulling the children toward the middle. The extremes had simply been selected partly by luck, and luck does not run in families.

What distinguishes regression to the mean from every other effect on this site is that it needs no cause and permits no exceptions. It is present whenever the correlation between two measurements is less than perfect, and the weaker that correlation, the stronger the regression. If two readings of the same thing correlate at 0.5, a case measured 20 points above average on the first reading is expected to be about 10 points above average on the second, before anything at all is done to it.

The one-sentence version: extremes are partly luck, and luck does not repeat

In one sentence: whatever you select for being extreme will look less extreme next time, so any treatment applied to the worst cases will appear to work. That is the whole of it, and it is enough to explain a great many confident claims.

You are looking at regression to the mean whenever three conditions hold together: a group was chosen because it scored at one end of a distribution, something was done to that group, and the group was measured again. Those three conditions describe remedial programs for the lowest scorers, quality drives aimed at the worst-performing line, coaching for the salesperson having a bad quarter, and maintenance triggered by a spike in fault reports. In every one of them, improvement is the default expectation even if the intervention is inert.

The arithmetic worked through: twelve weeks on a production line

Work the arithmetic on twelve weeks of defect counts from a production line averaging exactly 100 defects a week. These are invented illustrative figures, generated to average 100 with ordinary week to week variation and no trend and no intervention of any kind. Nothing changes across the twelve weeks except chance.

WeekDefectsDefects the following week
186108
210884
384140
414092
59295
69578
778130
8130100
910088
1088112
1111287
1287not yet observed

The twelve counts total 1,200, so the mean is exactly 100. Now behave like a manager. Identify the three worst weeks and launch an improvement program the following Monday.

  • Worst three weeks: week 4 at 140, week 8 at 130, week 11 at 112.
  • Their average: 382 divided by 3, which is 127.3 defects.
  • The weeks that followed: week 5 at 92, week 9 at 100, week 12 at 87.
  • Their average: 279 divided by 3, which is 93.0 defects.
  • Apparent improvement: 127.3 minus 93.0, which is 34.3 fewer defects, a fall of 27 percent.

A 27 percent reduction in defects, delivered by a program that did not exist, on a line that did not change. Every number in that calculation is correct and the conclusion drawn from it is entirely false.

Run it in the other direction and the illusion reverses. Take the three best weeks: week 7 at 78, week 3 at 84, week 1 at 86, averaging 82.7. The weeks that followed were week 8 at 130, week 4 at 140 and week 2 at 108, averaging 126.0. The best performers got 52 percent worse. Reward the good weeks and the reward appears to backfire; punish the bad weeks and the punishment appears to work. Daniel Kahneman described exactly this trap among flight instructors who had concluded from experience that praising a good maneuver made the next one worse, and that criticizing a bad one made the next one better.

What a failure looks like: the intervention that took the credit

A regression failure looks like a before-and-after comparison with no control group in it. The structure is always the same: measure, select the extreme cases, intervene, measure again, report the difference. Because that structure guarantees improvement from the selection alone, the difference it reports is the sum of the real effect and the regression, and there is no way inside the design to separate them.

The answer is a control group: a second set of cases selected on exactly the same extreme criterion and left alone. Both groups regress by the same amount, so the difference between them is what the intervention actually did. If the twelve-week line above had a matched second line, also selected for its three worst weeks and given nothing, it would have shown a fall of about the same 27 percent, and the improvement program's true effect would have been visible as zero. This is why the control group is not a formality bolted onto the scientific method. It is the only part of the design that can tell regression from effect.

The common misreading: mistaking it for a law of averages

The common misreading is to treat regression to the mean as a corrective force that pushes values back toward the average. It is nothing of the kind, and the distinction matters because the corrective version is the gambler's fallacy. A run of extreme values is not owed a run of opposite values to balance it. The next value is simply drawn from the same distribution as always, and most of that distribution sits near the middle.

The related confusion is with the law of large numbers. Regression to the mean is about what the next single observation is expected to be after an extreme one. The law of large numbers is about what a long-run average converges to as observations accumulate. Regression to the mean applies to individual cases and appears immediately. The law of large numbers applies to aggregates and appears slowly. One more misreading is worth naming: regression is not a claim that everything ends up average, since new extremes keep arriving. It is a claim about which direction to expect, not about a destination.

The test to run when something appears to have worked

Run these four questions on any before-and-after result, and start with the second one, which does most of the work.

  1. Ask how the cases were chosen, and whether they were chosen for being extreme on the same measure now being re-tested.
  2. Ask what the group would have scored with no intervention at all.
  3. Ask whether a comparison group was selected the same way and left untreated.
  4. Ask how much of the measure is noise, since the noisier the measure, the larger the regression.

The definitions this sits next to: control group, placebo effect and the method

Three definitions from introductory psychology surround regression to the mean, and AP psychology questions frequently test the boundaries between them rather than the terms themselves.

  • Regression toward the mean is the movement of an extreme score toward the average on re-measurement, caused by the unreliability of the measure rather than by any treatment.
  • A control group is a set of participants selected under the same criteria as the treatment group and given no treatment, so that everything except the treatment is shared between the two.
  • The placebo effect is improvement observed after an intervention with no active ingredient, attributed to expectation on the part of the participant.
  • The scientific method, in this context, is the practice of specifying in advance what result would count against your hypothesis, then arranging a comparison capable of producing it.

The examinable point is that regression to the mean and the placebo effect are different explanations for the same observation, and a single-group before-and-after study cannot tell them apart, because it cannot tell either of them from a real effect. Reading numbers covers the wider habit of converting a reported improvement back into raw counts, and The null hypothesis sets out the assumption that a controlled comparison is built to test. Critical thinking exercises put the three-condition check into practice on cases where the answer is not announced in advance.

There is a calculator for regression, and it needs a number you have to measure first

A regression to the mean calculator does exist, and it is one line of arithmetic: the expected second measurement equals the mean, plus the correlation between the two measurements multiplied by the distance of the first measurement from the mean. Everything on this page follows from that formula, and it tells you in advance how much improvement to expect from doing nothing at all.

Work it on a measure with a mean of 100 and a standard deviation of 20, where a selected case scored 120, which is one standard deviation above average.

  • At a test-retest correlation of 0.5, the expected second score is 100 plus 0.5 times 20, which is 110. Half the 20 point excess evaporates on its own.
  • At a correlation of 0.8, the expected second score is 116, so only 4 points evaporate and a real effect is easier to see.
  • At a correlation of 0.2, the expected second score is 104, so 16 of the original 20 points disappear with no intervention whatever.

One input has to be decided before the formula can run, and it is the one nobody has: the correlation between the two measurements. It is not a property of the intervention and it cannot be looked up. It has to be measured on untreated cases, by taking the same measure twice on the same units and correlating the results, which is work that has to happen before the study rather than after it. Enter a guess and the calculator will return a precise expected value built on it. The mean matters too, and it must be the mean of the population the cases were selected from, not the mean of the selected group.

What the output does not settle is whether the intervention did anything. The formula tells you what to expect from regression alone, so a result that beats the expectation is interesting and a result that matches it is not. That is a useful screen and it is not a substitute for a control group, because the expected value is itself an estimate carrying its own error, and because the formula assumes the only thing acting between the two measurements is chance. A calculator can tell you how much of an improvement was free. It cannot tell you who earned the rest.

Where to go next

Proudly powered by WordPress | Theme: Amber Blog by Crimson Themes.