Grain of Salt

The scientific method: the steps, and what they are for

The scientific method

Struck diagram on assay stock: The scientific method

The scientific method is a procedure for testing an idea by giving it a fair chance to fail, so that whatever survives the test is more likely to be true than whatever you believed before you started. It is not a subject and it is not a ritual. It is an ordered set of habits that protects an investigation from the most reliable source of error inside it, which is the investigator.

Step counts vary. A classroom handout usually draws five steps in a circle, a college text often lists six or seven, and a working laboratory protocol can run to dozens of numbered operations. There is no official number of steps, because the steps are not the method. The purpose behind each step is the method, and a step performed without its purpose is decoration.

What the scientific method is, and why the step count keeps changing

The scientific method changes its step count between fields because different subjects need different amounts of machinery to reach one goal: a result that would have come out differently if the idea were wrong. Francis Bacon set out the modern ancestor of the method in Novum Organum in 1620, arguing that knowledge should be built upward from systematic observation rather than downward from authority, and that an investigator has a duty to hunt for the instances that contradict a rule rather than the ones that flatter it. Karl Popper moved the emphasis in Logik der Forschung in 1934, published in English as The Logic of Scientific Discovery in 1959. What marks a claim as scientific, on Popper's account, is not that evidence can be found for it, which is easy for almost any claim, but that it forbids something, so that one specific observation would sink it.

Both descriptions survive inside the diagram on a school worksheet, where five arrows loop from question to conclusion and back to a new question. The drawing is not wrong. It is silent about what each arrow is for, and that silence is where nearly every real failure of the scientific method happens.

The seven steps, and what each one is for

The seven steps below cover what most versions of the scientific method include, each paired with the specific failure that follows when it is skipped. Read the third column first, because it is the reason the step exists at all.

StepWhat it is forWhat goes wrong when it is skipped
Observe before explainingEstablishes that the thing you want to explain actually happensAn explanation gets built on an effect that was never there
State a precise questionNarrows a vague impression into something an answer could settleThe inquiry cannot end, because nothing counts as an answer
Form a hypothesisCommits you to one explanation in advance, in publicEvery result looks like confirmation, because the story is written afterward
Derive a prediction that could failNames ahead of time the observation that would refute the hypothesisThe idea becomes unfalsifiable and survives absolutely everything
Test against a controlSeparates the effect of the thing tested from everything else changing at onceA change gets credited to the treatment when time or expectation caused it
Analyze as plannedStops the analysis being chosen after the fact to favor the answer you likeEnough slicing produces an impressive looking finding out of noise
Report and invite replicationExposes the result to people who would be pleased to see it fallAn error stays in the record because nobody ever ran it again

Two of these are routinely collapsed into one, which causes trouble. A hypothesis is a proposed explanation, such as the claim that a fertilizer raises tomato yield by feeding nitrogen to the plant. A prediction is the observable consequence, such as the claim that fertilized plants in the same soil, light and water will out-yield unfertilized ones by a measurable margin in one season. You can only be wrong about the second, which is exactly why it has to be written down first.

Why observation is the first step, not the hypothesis

Observation comes first because a hypothesis formed before you look becomes the thing you look with. A cafe owner notices that sales rise on days when the front door is propped open, and concludes that passers-by can smell the coffee. That is a hypothesis arrived at after the fact, and once it exists, every busy open-door day is remembered and every busy closed-door day is forgotten. The rival explanation is duller and was available from the start: the door gets propped open when the weather is mild, and mild weather brings more people onto the street. The way to tell them apart is to prop the door open on cold days too and count.

The modern institutional version of this discipline is preregistration, where researchers publish the hypothesis, the sample size and the planned analysis before collecting a single data point. It exists because the alternative, deciding what the study was testing once the numbers are in, is the most common way a competent person produces a false result without ever lying. Several of the standard logical fallacies are simply this error wearing a formal name.

Science and the scientific method are not the same thing

Science and the scientific method differ in the way a library differs from the act of reading. Science is the accumulated body of tested knowledge together with the institutions that maintain it: universities, journals, peer review, funding bodies and a community of specialists. The scientific method is the procedure by which a single claim earns a place in that body. Three consequences follow, and each one surprises somebody.

  • Apply the method without being a scientist. An auditor tracing a discrepancy, a mechanic isolating an intermittent fault by swapping one part at a time, and a teacher testing two lesson orders on parallel classes are all running the procedure.
  • Publish through the institutions and still be wrong. Peer review checks whether a paper is competent, plausible and properly reported. It does not check whether the finding is true, and it was never designed to.
  • Study a subject scientifically without the subject being a natural science. The method attaches to conduct, not to content.

What the scientific method is protecting you from

The scientific method protects an investigation from four named errors, and the reason it feels excessive is that the errors do not feel like errors from the inside. Richard Feynman put the first principle plainly in his 1974 commencement address at Caltech, later published as Cargo Cult Science: "The first principle is that you must not fool yourself."

  • Expectation in the subject. People who believe they have received something active report improvement, which is the reason the placebo effect exists as a methodological problem rather than a curiosity.
  • Expectation in the experimenter. Someone who knows which group a participant is in measures that participant differently, usually without noticing, which is the reason blinding was invented.
  • Selection in the data. Abraham Wald, working for the Statistical Research Group at Columbia University in 1943, addressed the question of where to add armor to military aircraft. The damage recorded on returning aircraft shows where a plane can be hit and still come home, so the memoranda pointed attention toward the areas the surviving planes were not hit in. The sample you can measure is not the sample you care about.
  • Storytelling after the fact. A pattern found in data and then explained is not a tested pattern, it is a described one.

Strip these protections out but keep the vocabulary, the charts and the confident tone, and what remains is pseudoscience, which is a failure of method rather than a category of subject matter.

Does the scientific method make sociology a science?

Yes, to the extent that sociology uses it, and the same answer applies to psychology, economics and epidemiology. Emile Durkheim argued the case for treating social facts as objects of systematic study in The Rules of Sociological Method in 1895. The genuine difficulty is not the subject matter but the experiment: you frequently cannot randomly assign people to a childhood, a neighborhood or an income, so the field leans on quasi-experimental designs, natural experiments, statistical control and large samples instead. Those tools are weaker than random assignment, and honest social scientists say so in the paper.

The dividing line runs through practice, not through the department name. A discipline is doing science when it states claims specific enough to come out false, checks them against evidence collected for that purpose, and publishes when they fail. A discipline stops doing science the moment its central claims can absorb any result, and that is true of a physics laboratory as much as a sociology seminar.

Where the scientific method is weakest, and the test you can run tomorrow

The scientific method is weakest at the step it depends on most, which is replication, because almost nothing in the system rewards it. Journals prefer novel positive findings, so null results are written up less often and published less often still, which quietly biases the visible literature toward effects that are real, exaggerated or absent in unknown proportions. The Open Science Collaboration reported in Science in 2015 that when 100 studies from three psychology journals were repeated, only about a third of the replications produced a statistically significant result. John Ioannidis argued from the arithmetic of hypothesis testing in PLoS Medicine in 2005, in a paper titled Why Most Published Research Findings Are False, that in fields with small samples and many tested hypotheses a large share of published findings will not hold.

None of that makes the method optional. It makes the method the only available correction, applied more often. So when a claim arrives tomorrow, whether from a study, a salesperson or a confident relative, ask these four questions in order:

  1. What result would the person making this claim accept as showing they were wrong?
  2. What was this compared against, and who chose the comparison?
  3. Was the analysis decided before the data came in, or after?
  4. Has anyone who would be glad to see it fail tried to reproduce it?

If the first question has no answer, the other three do not matter yet.

Where to go next

Proudly powered by WordPress | Theme: Amber Blog by Crimson Themes.