Control groups and blinding
A control group is the set of participants in an experiment who do not get the thing being tested, kept as similar as possible to the group that does, so that any difference between the two can be attributed to the thing itself rather than to everything else happening at the same time. Blinding is the second device: it keeps the people in the experiment, and often the people running it, from knowing who is in which group.
Together these two devices do most of the work in any test worth trusting. A study without them is not a smaller version of a good study. It is a different kind of object, one that can tell you what happened but not why.
What a control group rules out
A control group rules out four rival explanations at once, and each of them is capable of producing a convincing result on its own. Suppose a company tests a typing app that claims to cut keyboard errors, and 120 volunteers use it for six weeks. Their error rate falls by 18 percent. Without a control group, that number is uninterpretable, because all four of the following predict exactly the same drop.
- Practice and time. Six weeks of typing improves typing. Anything measured before and after a period of activity improves during that period.
- Regression to the mean. Francis Galton described the pattern in 1886: an extreme measurement tends to be followed by a less extreme one, simply because part of any extreme score is momentary noise. Volunteers who joined because they were typing badly that week will look better next month whatever you give them.
- Expectation. People who know they are being given a solution report and perceive improvement. That is the placebo effect, and it is a measurement problem before it is anything else.
- Being watched. Participants behave differently when observed, an idea named after the studies conducted at the Hawthorne Works of the Western Electric Company near Chicago between 1924 and 1932, though later reanalyses have disputed how large the original effect really was.
Give half the volunteers the app and half a plain typing exercise for the same six weeks, and all four explanations apply equally to both groups. Whatever is left over when you subtract one group from the other is the app.
Control group versus experimental group
The control group and the experimental group differ in exactly one respect by design, and that single difference is the whole point of the experiment. Everything else about them, including how they were recruited, what they were told, how often they were measured and by whom, is held identical on purpose.
| Feature | Experimental group | Control group |
|---|---|---|
| Receives the intervention | Yes | No, receives an inert or standard alternative |
| Assigned how | At random | At random, from the same pool of participants |
| Measured how | Same instrument, same schedule, same assessor | Same instrument, same schedule, same assessor |
| Knows which group it is in | No, in a blinded study | No, in a blinded study |
| Answers the question | What happened with the intervention | What would have happened anyway |
A control group is not always an untreated group. Where an accepted alternative already exists, the control is that alternative, and the experiment then answers a sharper question: is the new thing better than the current thing, not merely better than nothing. Historical comparisons, where this year's participants are measured against last year's records, are the weakest form, because a year changes far more than the intervention.
What blinding adds, and what single blind and double blind mean
Blinding adds protection against the one thing a control group cannot fix, which is knowledge of the assignment changing the behavior of the people in the experiment. A control group makes the comparison fair. Blinding keeps it fair while the experiment is running.
- Single blind: the participants do not know whether they are in the control group or the experimental group. The researchers do.
- Double blind: neither the participants nor the people who interact with them and record the results know who is in which group.
- Triple blind: the statisticians analyzing the data also work from coded group labels, so the analysis choices cannot be nudged toward the preferred answer.
So in a double blind study, who knows which participants are in the experimental group? Nobody in the room. The allocation is generated in advance, usually by a computer, and held by a third party, a pharmacist or an independent trials unit that has no contact with participants and no stake in the result. The code is broken only after the data are collected and locked, or earlier for a named safety reason recorded at the time. That answer matters because it defines what double blind actually rules out: not fraud, but the ordinary, unconscious, well meaning way an experimenter reads an ambiguous result generously when they know which group it came from.
The idea is older than the vocabulary. In 1784 a French royal commission appointed by Louis XVI, whose members included Benjamin Franklin and Antoine Lavoisier, investigated the claimed force of animal magnetism by blindfolding subjects and varying whether they were actually being treated. Subjects reacted when they believed they were being treated and not when they believed they were not, regardless of what was really happening. That is a blinded controlled experiment, conducted before either term existed.
How a double blind parallel group study runs
A double blind parallel group study runs two or more groups side by side through the same period of time, rather than putting one group through both conditions in sequence. Parallel means simultaneous, which removes the effect of the calendar. The order of operations is fixed before anyone is recruited.
- Define the outcome and the sample size in advance, in writing.
- Recruit participants who all meet the same stated criteria.
- Generate a random allocation sequence and give it to a third party who will not meet the participants.
- Prepare the intervention and the control so they are indistinguishable on inspection.
- Assign each participant to a group, concealing the next allocation from whoever enrolls them.
- Measure every group with the same instrument, on the same schedule, by assessors who do not know the assignments.
- Lock the data, then break the code, then run the analysis that was written down in step one.
The alternative design, a crossover study, gives each participant both conditions in a randomized order so that every person acts as their own control. It is more efficient with small numbers and useless when the first condition changes the participant permanently.
What randomization adds that a control group alone does not
Randomization adds control over the variables nobody thought of. A control group only works if the two groups are alike in every respect that matters, and the trouble is that you cannot list every respect that matters. If an experimenter assigns the keen volunteers to the app and the reluctant ones to the exercise, the groups differ in motivation before anything begins, and no amount of careful measurement afterward can unmix that. Random assignment distributes unknown differences across both groups by chance, and with enough participants the leftover imbalance becomes small and, importantly, calculable.
Ronald Fisher formalized this at Rothamsted Experimental Station and set it out in The Design of Experiments in 1935, working on agricultural plots where soil quality varied invisibly across a field. His answer, randomize the plots, is the same answer for people. Ignaz Semmelweis at the Vienna General Hospital in 1847 had a comparison but not a randomization: two maternity clinics with different staffing and sharply different outcomes, which let him identify a cause but left the door open to every objection about how the two groups differed to begin with. A comparison group without random assignment gives you a strong hint. Random assignment turns the hint into an estimate.
When blinding is impossible, and the question to ask instead
Blinding is impossible whenever the intervention is obvious to the person receiving it, such as a surgical procedure, a training course, an exercise regime or a redesigned user interface. Nobody can be unaware of which classroom they sat in. Three partial substitutes exist, and a good study names which one it used: blind the assessor even though the participant knows, use an objective outcome that cannot be talked up or down, and compare against an active alternative that feels equally like a real intervention rather than against nothing at all.
So the test to run on any study you meet, before you read the conclusion, is a short interrogation of its design. What was the control group, and who chose it? Who knew which participants were in the experimental group, and at what point were they told? Were participants assigned at random, or did they choose, or did someone choose for them? Was the outcome something you can count, or something somebody reported? Answers to those four questions tell you more about a result than the size of the effect does, which is the connection between critical thinking skills and the scientific method: the reasoning is the same, applied to a study instead of a sentence. When a claim keeps its confidence but sheds its control group, it is on the road toward pseudoscience.