Falsifiability
Falsifiability is the property a claim has when some possible observation would show it to be false. A falsifiable claim sticks its neck out: it rules something out in advance, and it names, at least implicitly, the result that would count against it. A claim that is compatible with every conceivable outcome is unfalsifiable, and it is not thereby wrong. It is something other than a testable statement about the world, and no amount of evidence gathered in its favor changes that, because nothing was ever at risk.
Karl Popper and the demarcation problem
Karl Popper, the Austrian and British philosopher who lived from 1902 to 1994, introduced falsifiability as a solution to the demarcation problem, which is the question of what separates scientific claims from claims of other kinds. He set it out in Logik der Forschung in 1934, published in English as The Logic of Scientific Discovery in 1959, and developed it further in Conjectures and Refutations in 1963.
His argument runs against the intuition most people start with. Confirmations are easy to obtain: a theory that predicts something common will be confirmed constantly, and a theory loose enough to fit anything will be confirmed by everything. Popper's move was to make the asymmetry between confirmation and refutation the centre of the method. No number of confirming observations establishes a universal claim, because the next observation is always outstanding, while a single genuine counterexample refutes it. On that account science does not proceed by proving theories true. It proceeds by proposing bold conjectures, deriving risky predictions from them, and trying seriously to knock them down. A theory that survives repeated honest attempts is corroborated, which means it has not yet failed, and Popper was careful that this is a weaker status than proven.
Three consequences follow, and they are the practical content of the idea.
- Prediction beats accommodation. A theory that forecasts a result nobody expected is worth more than one that explains a result after it arrived.
- Prohibition is information. The value of a theory lies in what it forbids. A theory that forbids nothing tells you nothing.
- Risk is a virtue. The more ways a claim could fail, the more you learn when it does not.
The practical test: name the observation that would sink it
To test any claim for falsifiability, ask the person making it a single question: what would you accept as evidence that this is wrong? The answer, or the absence of one, tells you what kind of claim you are dealing with before you have examined any evidence at all. Three responses are worth recognizing.
- A specific observation is named. The claim is falsifiable, and the conversation can proceed to whether that observation has been looked for.
- Nothing could count against it. The claim is unfalsifiable as stated, and the honest next step is to reformulate it into something that could fail, if a version like that exists.
- The answer changes once evidence appears. The claim was falsifiable and is being rescued, which is a different problem and the subject of the next section.
Compare two statements about the same subject. First: every swan that will ever be observed is white. Second: swans have a natural tendency toward whiteness, which expresses itself except where circumstances interfere. The first is falsifiable and was in fact falsified, decisively, by the observation of black swans in Australia in the seventeenth century. The second survives any observation whatsoever, because any exception is absorbed by the clause about circumstances. The second sounds more sophisticated and more cautious. It is worth less, because it forbids nothing.
Falsifiability in psychology
In psychology, falsifiability is enforced through two habits, and both exist because the subject matter makes vague formulations especially easy to write. The first is the operational definition: before data are collected, the researcher states exactly what will be counted as an instance of the thing being studied, in terms another team could apply without asking. The second is stating the hypothesis in a form that specifies a result which would count as a failure, usually by naming a null hypothesis that the data are then given a fair chance to leave standing.
The difference this makes is visible in the wording. A claim that people are driven by unconscious motives cannot fail, because any behavior at all is consistent with it, and so is the opposite behavior. A claim that participants who are shown a number before estimating a quantity will give estimates pulled toward that number can fail, and specifying the population, the size of the pull and the conditions makes it fail more easily. Preregistration, in which the prediction and the analysis plan are recorded before the data exist, exists to stop the second kind of claim quietly turning into the first. It is also the reason large replication efforts are informative: when the Open Science Collaboration repeated a set of published psychology studies and reported its results in 2015, fewer than half of the replications reproduced the original findings, and the exercise was only possible because those original claims had been stated in a falsifiable form.
How a claim escapes falsification
A claim escapes falsification in three ways, and none of them requires anyone to lie. The moves are usually made in good faith by people who are confident in a theory and are trying to save it, which is what makes them worth learning to spot in your own reasoning first.
| Escape route | What it looks like | What to ask |
|---|---|---|
| Auxiliary hypothesis | A new condition is added after the failure to explain why the test did not apply | Was this condition part of the theory before the result came in? |
| Moving the goalposts | The standard of evidence rises as soon as the evidence is produced | What was the standard stated in advance? |
| Retreat to vagueness | A specific prediction is restated as a general tendency | Does the revised claim still forbid anything? |
The first route is more respectable than the other two, because adding auxiliary assumptions is sometimes exactly right. This is the point Pierre Duhem and Willard Van Orman Quine made, and it is known as the Duhem-Quine thesis: a hypothesis is never tested by itself, but always together with assumptions about the instruments, the sample and the background conditions. When a prediction fails, logic alone does not say which element failed, so blame can always be assigned somewhere other than the main theory. Real discoveries have come from exactly that manoeuvre, when an anomaly was attributed to an unseen additional body rather than to the theory, and the body was later found. The difference between rescuing a theory and protecting it is whether the added assumption is itself testable and makes a new prediction of its own. Occam's razor is a reasonable guide here, since an assumption added only to absorb a failure is an entity multiplied beyond necessity.
The honest limit: the criterion is disputed
The honest limit is that the criterion is disputed: falsifiability is not the settled definition of science, and presenting it as one oversells it. Thomas Kuhn argued that working scientists do not abandon a theory at the first refuting result, and that they are right not to, because every theory has unresolved anomalies at all times and discarding them individually would leave nothing standing. Imre Lakatos proposed that the unit being judged is a research programme rather than a single claim, and that a programme should be assessed on whether it keeps generating new predictions that turn out to be correct, or has switched to explaining away its failures. Paul Feyerabend went further and denied that any single criterion captures scientific practice at all.
Two more limits are worth stating plainly. Falsifiability is a matter of degree rather than a yes or no, since claims can be more or less exposed depending on how sharply they are specified. And some legitimate scientific claims, particularly about the deep past or about very rare events, cannot be tested by direct intervention, so they are tested by the predictions they make about what evidence should still be found. What survives the criticism is the practical test rather than the grand definition. Asking what would show this to be false remains the fastest way to find out whether a claim is doing any work, and it is a question worth asking about the beliefs you hold most comfortably, since those are the ones least likely to have been asked it before. Everything downstream of it, including what counts as evidence and how belief becomes justified, is the province of epistemology, and this is the question that hands the inquiry over.