Grain of Salt

What counts as evidence, and what does not

What counts as evidence

Struck diagram on assay stock: What counts as evidence

Empirical evidence is information about the world obtained by observation or measurement and recorded in a form that another person can inspect, repeat or challenge. The word empirical describes where the information came from, not what it is about: a claim rests on empirical evidence when the reason to believe it is something that was seen, counted, weighed or timed, rather than something worked out from definitions or recalled after the fact. Almost every argument about whether a claim is supported turns out, on inspection, to be an argument about how well a particular record meets that one standard.

The three conditions a record has to meet before it counts

Information counts as empirical evidence when it satisfies 3 conditions, and a record that fails any one of them is doing something other than what it appears to do.

  1. Observation. The information came from contact with the thing being described, through the senses or through an instrument, and not from an assumption about how the thing must behave.
  2. Independence. The record exists apart from the memory and the interests of the person reporting it, so that a stranger could go and look at the same thing.
  3. Contingency. The observation could have come out differently. A measurement that would have been reported the same way whatever happened is not measuring anything.

The third condition is the one people forget, and it is the one that connects evidence to falsifiability. A claim that no possible observation could embarrass is not being supported by observation at all, however many observations are placed next to it. Psychology has a working habit that enforces the first two conditions: before a study runs, the researchers write down an operational definition, which states exactly what will be counted as an instance of the thing under study, in units anyone else could apply. Defining the measurement before collecting the data is what stops the definition from quietly adjusting itself to the result.

The ladder from a single observation to a replicated result

Observations climb a ladder of 5 rungs, and each rung supports a wider claim than the one below it while ruling out one more way of being fooled. Nothing on this ladder is worthless and nothing on it is final. What changes from rung to rung is the range of conclusions the evidence can carry.

RungWhat it can supportWhat it cannot support
Single anecdoteThat the thing described is possible at allAny statement about how often, how much, or for whom
Collected reportsA hypothesis worth testing, and a rough sense of varietyA rate, because nobody counted the cases that were never reported
Observational study with a comparison groupAn association, with a size attached to itA cause, because the groups differ in ways nobody measured
Controlled experiment with random assignmentA causal effect, within the conditions actually testedThat the effect holds outside those conditions
Independent replicationThat the result is not an artifact of one team, one sample or one analysisThat the underlying explanation is the right one

The jump from the third rung to the fourth is the one that does the real work, and the classic demonstration of why is small enough to fit on a page. In The Design of Experiments, published in 1935, the statistician Ronald Fisher described how to test a colleague's claim that she could taste whether the milk had gone into the cup before or after the tea. The design was 8 cups, 4 poured each way, presented in random order, with the number of correct identifications compared against what guessing alone would produce. Every element of a modern experiment is already there: a claim stated in advance, a fixed number of trials, randomization to break any pattern the taster could exploit, and a stated threshold for what would count as a result. Without the randomization there is a story about a talented taster. With it there is evidence.

Empirical evidence versus anecdotal evidence

Empirical evidence and anecdotal evidence are not opposites in the way the phrase suggests, and getting the relationship right matters more than winning the argument. An anecdote is an observation, so it is empirical in the literal sense. What it lacks is control: no denominator, no comparison group, no record made before the outcome was known, and no way to know how many similar cases went unmentioned because they ended undramatically.

The practical difference sits in 4 places:

  • Selection. Anecdotes arrive because somebody chose to tell them, and the reason they were worth telling is usually the reason they are unrepresentative.
  • Recall. Memory reconstructs, and it reconstructs in the direction of the story the teller already believes.
  • Denominator. A story reports the numerator and stays silent about the number of attempts underneath it.
  • Comparison. A single case has nothing to be compared with, so there is no way to know what would have happened anyway.

Readers looking for an antonym of empirical are usually reaching for one of three words, and they are not interchangeable. A priori means known independently of observation, as with a claim in arithmetic. Theoretical means derived from a model rather than measured. Anecdotal means observed but uncontrolled. The nearest synonyms for empirical, in the way the word is used in science, are observational, measured and data based. Only the first of the three antonyms is a genuine logical opposite. The third is a weaker relative in the same family.

Why evidence supports a claim and almost never proves it

Proof, in the strict sense, belongs to mathematics and formal logic, where conclusions follow from premises with nothing left over. Empirical evidence works the other way: it accumulates, it raises or lowers confidence, and it stays open to the next observation. A scientist who says the evidence is overwhelming is making a claim about weight, not about certainty, and the honest version of any empirical statement carries an implied margin. This is why the vocabulary of proof, once it leaves the seminar room, causes so much trouble. Asking for proof of an empirical claim sets a standard that no empirical claim has ever met, which is a convenient thing to demand of a claim you would rather not accept.

Two failure modes deserve names. The first is the leap from evidence to a conclusion the evidence does not reach, which is the shape of a non sequitur: the observations may be flawless and the conclusion still fail to follow from them. The second is the quiet addition of extra machinery to keep a favored explanation alive after the evidence has moved, which is the situation Occam's razor is meant for. The razor does not tell you which explanation is true. It tells you which to prefer when two explanations account for the same observations equally well, and the answer is the one carrying fewer assumptions. Both of those are downstream of a larger question, the one epistemology asks: what makes any belief justified in the first place.

Five questions to ask before you accept a piece of evidence

Ask 5 questions of any piece of evidence handed to you, in this order, and stop at the first one that has no answer.

  1. Count the cases. How many observations is this, and how many were made in total, including the ones that did not get reported?
  2. Find the comparison. Compared with what? If there is no group that did not get the thing, there is no way to know what would have happened anyway.
  3. Check the order of events. Was the prediction written down before the result was known, or was the pattern noticed afterwards in data collected for another purpose?
  4. Ask who could have seen it fail. What observation would the person offering this have accepted as a refutation, and did anyone go looking for it?
  5. Trace it back one step. Is this the record itself, or a description of a description? Every retelling loses conditions and gains confidence.

The fourth question is the sharpest of the five, and it is worth running on your own beliefs first. Abraham Wald, working for the Statistical Research Group at Columbia University during the Second World War, was asked where to add armor on aircraft, given the pattern of damage on the ones that came back. The observations were accurate. The inference everyone wanted to draw from them was backwards, because the aircraft that would have carried the most informative damage were the ones not available to inspect. The data were empirical and the sample selected itself.

Where empirical evidence is not the right tool

Empirical evidence answers questions about what is the case, and there are 3 kinds of question it cannot settle. Questions of definition are settled by stipulation, not by measurement, and arguing about whether a borderline case really counts is usually an argument about words. Questions of pure mathematics are settled by proof, and no amount of measuring can make a theorem more true. Questions of value are settled by argument about what matters, and evidence constrains the answer without determining it: it can tell you what a policy would do and it cannot tell you whether the result is worth having.

There is also one place where the weakest rung of the ladder is enough. A single well documented observation settles an existence claim outright, because a universal statement of the form nothing of this kind ever happens is refuted the moment one instance is produced. One counterexample beats a thousand confirmations, which is why the search for the disconfirming case is worth more than another round of examples that agree with you.

Where to go next

Proudly powered by WordPress | Theme: Amber Blog by Crimson Themes.