Cherry picking
Cherry picking is presenting the part of the evidence that supports a claim while leaving out the part that does not, so that every individual statement is true and the overall impression is false.
What cherry picking is
Cherry picking, in reasoning, is selective reporting: the selection is made after the results are known, and the basis for the selection is the conclusion the presenter wants. It is the most common way to mislead an audience without saying anything untrue, which is why it survives fact checking that only tests statements one at a time.
What separates it from the rest of the cognitive bias family is that cherry picking has two forms, one honest and one not. As a bias, it happens without intent: a person notices and remembers the four results that agreed and forgets the seven that did not. As a rhetorical move, it is deliberate and the missing data is known to the person who left it out. The reader usually cannot tell which they are looking at, and it rarely matters, because the repair is identical. Ask for the rest of it.
The one sentence version, and the three things that get picked
In one sentence, cherry picking is choosing the sample after seeing the results. Three things get picked in practice, and naming them makes the move easier to spot.
- Cases. Five customers who loved the product appear as testimonials, and the number of customers is never given.
- Time windows. A chart begins in the one year that makes the trend look strongest, and the years before it are cropped away.
- Outcomes. A trial measured nine things, one of them came out favorable, and the write up is about that one.
The third form is the hardest to detect from the outside, because the reader has no way of knowing that eight other measurements existed. It is also the one that has changed how serious research is now run.
The evidence: the file drawer, and what pre-registration fixed
The strongest evidence that cherry picking distorts entire fields is publication bias, and the psychologist Robert Rosenthal named its mechanism the file drawer problem in 1979. Studies that find an effect get written up and published. Studies that find nothing go into a drawer. A reader who then surveys the published literature is reading a cherry picked sample of all the research that was actually done, assembled by no one in particular and with no intent to deceive.
The statisticians Andrew Gelman and Eric Loken described a related mechanism they called the garden of forking paths: a researcher analyzing a dataset makes dozens of small, defensible choices about which cases to exclude, which groups to compare and which measure to use, and any of those paths can lead to a publishable result. No single choice is cheating, and the researcher may run only one analysis. The selection happened anyway, among the analyses that were never run because the first one worked.
The fix that emerged is pre-registration: the hypothesis, the sample size and the primary outcome are written down and time stamped before the data is collected. A pre-registered study that reports a different primary outcome than the one it registered is visibly doing something, and a reader can check. That is the whole value of the device. It does not make the study true; it makes the selection visible.
Where cherry picking shows up in things you read this week
Cherry picking appears wherever someone chooses what to show you, such as product comparison charts, testimonial pages, quarterly performance summaries, benchmark results published by the vendor of the thing being benchmarked, and survey write ups that quote three questions out of twenty.
Two worked examples, with the repair in each.
The training program
A company reports that staff who completed its new onboarding program hit full productivity in 6 weeks, against a historical average of 9 weeks. True as stated. The missing selection is who completed the program: the people who dropped out partway through are not in the 6 week figure, and they are exactly the people who were struggling. The repair is to report the outcome for everyone who started, which researchers call an intention to treat analysis, and which here just means counting the dropouts.
The performance chart
A chart shows a fund, a metric or a department's output rising steeply from a starting point five years ago. Move the start point two years earlier and the same series shows a fall followed by a recovery to roughly where it began. Neither chart lies. The one that gets published is the one whose start date was chosen after the shape was known, and a reader who never asks what came before the left edge has no way to see it.
What cherry picking is confused with: the legitimate exclusion of bad data
Cherry picking is confused, more than with anything else, with the legitimate exclusion of bad data, and the confusion is worth clearing up first. Choosing a subset of the data is not automatically cherry picking; choosing it after seeing the results is. That distinction saves a great deal of pointless argument, because analysts exclude data constantly for good reasons: a faulty sensor, a duplicated record, a period when the measurement instrument was changed. What makes an exclusion legitimate is that the rule for it was stated before the results were known, and that it would have been applied whichever way the results came out.
| Term | What it names | How it differs from cherry picking |
|---|---|---|
| Survivorship bias | Only the surviving cases are available to study | Nobody selected anything; the missing cases removed themselves |
| Texas sharpshooter fallacy | Drawing the target around the cluster after firing | The pattern is defined after the fact, not the sample |
| Confirmation bias | Noticing and recalling what fits a belief | Happens in the head of the reader, not in the presentation |
| Quote mining | Lifting a clause out of the sentence that qualifies it | The unit selected is words, not data points |
| Illustration | One example given to make an abstract point concrete | Declared as an example and not offered as the weight of evidence |
Note also what cherry picking is not, logically. It is not a non sequitur: the conclusion does follow from the evidence presented. The defect is upstream of the inference, in the evidence that was allowed into the room, which is why calling it a logical fallacy tends to confuse people who then go looking for a broken inference.
One question does the separating: was the rule for leaving this out written down before anyone saw the results? A rule set in advance and applied both ways is data cleaning, and a good analyst can produce it on request. A rule discovered afterward, that happens to remove the inconvenient cases and would never have been applied to the convenient ones, is cherry picking whatever it is called in the write up.
What actually reduces cherry picking
What reduces it is fixing the denominator before you look at the numerator. Five questions do most of the work, and all five can be asked by a reader who has no access to the raw data:
- Ask how many cases there were in total, not how many are shown.
- Ask what the series looks like if the start date moves back one year.
- Ask what else was measured, and where those results are.
- Ask whether the exclusion rule was written before or after the results.
- Ask who chose the comparison, and what they get if it comes out well.
Be honest about the limit. You will often not get an answer, and the absence of an answer is itself weak evidence rather than proof of bad faith. Plenty of organizations simply do not have the full dataset in a form anyone can send you, and the person presenting the chart may not know what was excluded upstream of them, which is the ordinary condition described on the Dunning Kruger effect page: the gap is invisible from inside it. The point of the questions is not to win; it is to know how much weight the number in front of you can carry, which is usually less than the presentation implies and more than zero.
The test to run on a chart or a claim
One question catches most of it: what would I have to see to conclude the opposite, and would that evidence have been shown to me? If the answer is that the opposite result would simply not have been published, printed or presented, then the evidence in front of you cannot distinguish between the claim being true and the claim being selected.
A second question for charts: where does the left edge come from? A start date with no reason behind it other than the shape it produces is the most common cherry pick in ordinary business reporting, and it takes ten seconds to check.
Cherry picking in psychology: a reporting practice rather than an effect
In psychology, cherry picking has no entry of its own, and the honest thing is to say so rather than to manufacture a literature for it. It is a description of what somebody presented, not a named effect measured in a laboratory. You will find it in style guides, statistics teaching and research methods courses; you will not find a body of experiments on the cherry picking effect, because there is not one.
Three neighboring things do have research behind them, and they are what a reader looking for the psychology usually wants:
- Confirmation bias, the selective search for and recall of supporting evidence, which is the closest mechanism inside one person's head.
- Selective exposure, the tendency to choose sources likely to agree, which selects the input rather than the output.
- Publication bias and researcher degrees of freedom, studied in metascience and statistics rather than in the psychology of judgment, which is where the file drawer problem and pre-registration belong.
So the accurate answer to the psychology question is a division of labor: the doing of cherry picking is a reporting practice, and it can be entirely deliberate, in which case no psychology is involved at all. The falling for it is where the psychology sits, and it is mostly the ordinary failure to ask for a denominator. Calling cherry picking a cognitive bias, or a logical fallacy, claims a mechanism that has not been demonstrated and points the reader at the wrong repair. The repair is not introspection. It is a request for the rest of the data.
The other meanings of cherry picking
The phrase has three unrelated everyday senses, and most searches for it are not about reasoning at all. In agriculture it means the literal harvest, and the u-pick cherry orchards of Door County in Wisconsin and Brentwood in California draw large numbers of visitors during a short season each summer. In basketball it means a player who stays near the opponent's basket instead of defending, waiting for an easy pass. In software it means taking a single commit from one branch and applying it to another, which is what the git cherry-pick command does.
All three share the same underlying image: taking the easy, ripe, convenient item and leaving the rest of the tree. This page is about the reasoning sense, where the rest of the tree is the part that decides whether the claim is true.