Publication bias
P hacking is the practice of running analysis after analysis on the same data until one of them crosses the conventional significance threshold, then reporting that one as though it had been the plan from the start. It is not usually fraud and it rarely feels like cheating from the inside. It is a search procedure mistaken for a test, and it is one half of the reason the published research literature is not a fair sample of what researchers actually found.
Why p hacking works: the 0.05 threshold and researcher degrees of freedom
P hacking works because of an arithmetic property of the significance threshold. A p value below 0.05 is conventionally read as evidence against chance, and 0.05 means roughly a one in twenty chance of getting a result this extreme if nothing real is going on. Test one hypothesis and that risk is small. Test twenty independent variations on the same data and, purely by chance, you expect about one to clear the bar. Report only that one and you have manufactured a finding out of noise without writing down a single false number.
The choices that create those twenty tests are called researcher degrees of freedom, a term from a 2011 paper in Psychological Science on false positive psychology, which demonstrated that ordinary analytic flexibility pushes the false positive rate far above the nominal 5 percent. The flexible choices include:
- Stopping data collection when the result looks good rather than at a preset sample size.
- Dropping outliers under a rule chosen after seeing which points hurt.
- Adding covariates such as age or session order until the effect appears.
- Splitting the sample into subgroups and reporting the subgroup where the effect held.
- Choosing among several measures of the same construct after the fact.
- Reframing the hypothesis to match the result, a habit known as hypothesizing after the results are known.
Andrew Gelman and Eric Loken made the important refinement: this damage occurs even when a researcher runs only one analysis. If the analysis chosen would have been different had the data come out differently, the effective number of tests is still large. They call it the garden of forking paths, and it means p hacking does not require intent. It requires only that decisions be made after seeing the data.
Publication bias and the file drawer problem
Publication bias is the systematic tendency for studies with positive, novel or striking results to be published while studies finding nothing are not. The psychologist Robert Rosenthal named the mechanism the file drawer problem in 1979: null results end up in a drawer rather than in a journal, so the published literature over represents the successes of a research effort and hides its failures.
Three parties each contribute, and none of them has to act badly for the effect to appear. Journals prefer results that say something. Reviewers find null findings less interesting and more likely to be dismissed as underpowered. Authors, knowing all this, do not spend three months writing up a study nobody will take. The bias is a property of the system, not of the people in it.
The consequence is severe for anyone reading a research field rather than a single paper. If ten teams test the same idea and one gets a positive result by chance, the literature may contain one publication showing an effect and nine drawers. A meta analysis pooling only what was published will find that effect and report it with confidence. Funnel plot analysis is the standard detection tool: plot each study's effect size against its precision, and an unbiased field produces a symmetrical funnel, wide at the bottom where small studies scatter and narrowing toward the true value at the top. Missing small studies on the null side leave a visible notch in the funnel, and that asymmetry is evidence that something was filed rather than published.
How p hacking and publication bias produced a replication crisis
The replication crisis is the discovery, across several fields, that a substantial share of published findings do not hold when the study is run again by other people. P hacking supplies the false positives; publication bias ensures those are the ones that get printed and cited; and nobody notices, because direct replications have historically been hard to publish too.
Two pieces of work anchor this. In 2005, John Ioannidis published an essay in PLoS Medicine arguing that the probability a published finding is true depends on the prior plausibility of the hypothesis, the statistical power of the study, the amount of analytic flexibility and the number of teams chasing the same question. Under conditions common in many fields, that probability falls below one half. The argument is mathematical rather than empirical, and it was a prediction about what would be found later.
Then it was found. The Open Science Collaboration, coordinated through the Center for Open Science, ran a large program of direct replications of experimental and correlational studies published in leading psychology journals, and reported in Science in 2015 that fewer than half of the replications produced a statistically significant result, with replication effect sizes averaging roughly half the magnitude of the originals. Similar replication projects have since been run in experimental economics and in the social sciences more broadly, with the same broad shape.
What the replication crisis did to the field of psychology
The impact on psychology was disproportionate, and understanding why matters more than the headline. Psychology did not turn out to be worse than its neighbors. It turned out to be the field that looked, and the reason it could look is structural: its experiments are comparatively cheap, fast and repeatable, so a graduate student can attempt a direct replication in a semester. A field where one study costs eight years and a particle accelerator cannot audit itself the same way.
Social psychology absorbed the sharpest blow because several of its best known experimental effects, the sort that reached textbooks and popular science shelves, failed to replicate at the original size or at all. The field's response over the following decade is the part worth learning from, because it is what a discipline correcting itself actually looks like: larger samples, published data and code, multi laboratory studies that run the same protocol in twenty sites at once, journals adding replication sections, and a general downward revision of confidence in single study findings.
The lesson generalizes and it is not "psychology is unreliable". It is that a single experimental result, in any field, is a weak unit of evidence, and that the strength of science is in what survives repetition rather than in what was published first.
What pre-registration and Registered Reports fix
Pre-registration fixes the specific hole that p hacking exploits: it makes the analysis plan a matter of record before the data exist. Registered Reports go further and move peer review to the same point. The table sets out what each one blocks.
| Practice | What the researcher commits to, and when | What it blocks | What it does not fix |
|---|---|---|---|
| Conventional study | Nothing in advance; the paper is written after the results | Nothing | Everything below |
| Pre-registration | Hypothesis, sample size, exclusion rules and analysis, filed in a public registry before data collection | Undisclosed flexibility, silent stopping rules, hypothesizing after results are known | Bad measures, bad theory, and the file drawer if the study is still never written up |
| Registered Report | The same plan, submitted to a journal and peer reviewed before results exist, with acceptance granted in principle | All of the above, plus publication bias, because acceptance no longer depends on how the result comes out | Whether the question was worth asking, and whether the method measures what it claims |
Public registries such as those run by the Open Science Framework make pre-registration routine, and the Registered Report format has been adopted by well over a hundred journals since journals began offering it in the 2010s. Both remain a minority of what is published, and neither makes a study good. A pre-registered study can be badly designed; the registration only guarantees that it was badly designed on purpose and in advance, which at least makes the design inspectable.
Distinguishing the confirmatory part of a paper from the exploratory part is the whole benefit. Exploration is legitimate and necessary. Presenting it as confirmation is not, and pre-registration is simply the mechanism that keeps the two labeled.
Reading a single study with all of this in mind
Reading a single study now means asking a different first question. Not "is this significant" but "how many analyses could have produced this sentence". Four things tell you most of what you need: the sample size, whether the study was pre-registered, whether the data and code are available, and whether anyone has replicated it. A study reporting a surprising effect in 30 participants, with no registration and no replication, is a hypothesis. That is a useful thing to be, and it is not a finding.
Confirmation bias determines which studies you subject to that scrutiny, which is why the procedure has to be applied to the results you like. The chain by which a fragile finding becomes a confident headline is traced on Media literacy, the difference between the paper and the article about the paper is on Primary vs secondary sources, and the general procedure for weighing any source is on Evaluating sources.
The test, before you accept any single striking result: if this study had found nothing, would I ever have heard about it? When the answer is no, you are looking at the top of a funnel whose bottom you cannot see, and the honest reading of the finding is one notch weaker than the paper's own.