Reading numbers
Misleading statistics are true numbers arranged so that a reader draws a false conclusion, and six patterns account for most of them: a percentage with no denominator, a relative change quoted without the absolute one, an average that hides its own spread, a truncated graph, a date range chosen after the results were known, and a survey question that decides its own answer. None of the six requires anybody to lie. Every figure inside a misleading statistic can be arithmetically correct, which is why checking the arithmetic is not the check that matters.
Reading numbers is a short, repeatable procedure rather than a talent. You are not being asked to recompute anybody's percentage. You are being asked to find out what the number was measured on, what it was compared against, and what was left out of the frame. Those three questions catch most bad statistics before they reach the point of being argued about, and they take about a minute.
What separates a misleading statistic from a false one
A misleading statistic differs from a false one in that nothing in it is untrue. A false statistic reports a quantity that was never measured. A misleading statistic reports a quantity that was measured accurately and then presents it in a frame that makes it mean something else. The first is caught by an audit. The second is caught only by a reader who asks what the frame is doing.
This matters because the two failures have different repairs. You cannot fact-check your way out of a misleading statistic, because it passes the fact-check. A headline reading "warranty claims up 200 percent" is exactly right when claims moved from 3 to 9 in a year on a product with 40,000 units in the field. The 200 percent is real. The 9 is real. What the sentence conceals is that 9 claims out of 40,000 units is a failure rate of 0.0225 percent, and that a jump from 3 to 9 on those numbers is well inside the wobble you would expect from a quantity that small.
Six misleading statistics examples, with the arithmetic
These six misleading statistics examples are each worked through with invented but internally consistent figures, so you can see the mechanism rather than take the description on trust. The pattern is what transfers, not the numbers.
The percentage with no denominator
A percentage without its denominator is the cheapest misdirection available, because percentages of small numbers are enormous. Three complaints becoming nine is a 200 percent increase. Three becoming four is a 33 percent increase. Neither sentence is false and neither tells you whether anything happened. Ask for the raw counts every time a percentage change appears without them, and if the raw count is under about 30, treat the percentage as decoration.
The relative change without the absolute one
A relative change and an absolute change describe the same event at wildly different volumes. Suppose a payment processor reports that its fraud rate rose from 2 in every 10,000 transactions to 3 in every 10,000. That is a 50 percent increase in relative terms and one extra fraudulent transaction per 10,000 in absolute terms. Across a month of 4 million transactions it is a move from 800 to 1,200: 400 additional cases, which is worth knowing, and nothing like the alarm that "fraud up 50 percent" produces. Both numbers are correct. A report that gives you only the first has chosen the one that sounds larger.
The average that hides its own spread
An average conceals its own distribution, and the mean conceals it worst. Take ten employees: nine earn 30,000 and one earns 930,000. The total is 1,200,000, so the mean salary is 120,000 while the median salary is 30,000. Every word of "the average salary here is 120,000" is true and no employee earns anything close to it. Whenever an average is quoted for anything with a long tail, such as salaries, company sizes, download counts or repair costs, ask for the median and the range as well.
The graph that starts at 90
A graph that starts its vertical axis above zero multiplies the appearance of a difference without changing a single value. Two bars representing 92 and 96 on an axis running from 90 to 100 stand 2 units and 6 units tall, so the second bar carries three times the ink of the first while representing a value 4.3 percent higher. The truncated axis is the most common of the misleading graphs, and it is legitimate in some contexts and deceptive in others, which is why it deserves its own treatment rather than a blanket rule.
The date range chosen after the fact
A date range chosen after the results are known can reverse the direction of a trend without altering one data point. Imagine a series that reads 100 in 2015, peaks at 145 in 2019, and stands at 130 in 2025. Measured from 2015 the series is up 30 percent. Measured from 2019 it is down 10.3 percent. Both are honest arithmetic on the same series. The question that catches this one is not "is that number right" but "why does the chart begin in that year".
The survey question that decides its own answer
A survey question can determine its result through wording, response options or who was asked. A survey that invites respondents to name as many brands as they like will report that eight in ten professionals "recommend" a given brand, when the share who would choose it first might be two in ten. A question that asks whether you agree with a statement gets more agreement than one that offers you the opposing statement as well. Before accepting a survey figure, find the exact wording, the number of respondents, and how respondents were reached.
The six checks in one table
The six checks below turn each of the patterns above into a single question, phrased so you can ask it of any figure you meet tomorrow.
| Pattern | What it does to the number | The question that catches it |
|---|---|---|
| Percentage with no denominator | Inflates small counts into large-sounding changes | How many is that, out of how many? |
| Relative without absolute | Reports the larger of two true framings | What is the change in raw cases per thousand? |
| Mean without median | Lets one extreme value stand for the group | What is the median, and what is the range? |
| Truncated axis | Multiplies a small gap visually | Where does the vertical axis start? |
| Cherry-picked range | Reverses or exaggerates a trend | Why does the series begin and end there? |
| Loaded survey question | Produces the answer it was built to produce | What exactly was asked, of whom? |
Streaks, the gambler's fallacy and the law of large numbers
Streaks in data invite two opposite errors, and the difference between the gambler's fallacy and the law of large numbers is the cleanest way to see both. The gambler's fallacy is the belief that a run of one outcome makes the opposite outcome more likely next time. The law of large numbers is the true statement that as the number of trials grows, the observed proportion approaches the underlying probability. They sound like the same idea and they are not.
The reconciliation is arithmetic. Flip a fair coin ten times and get ten heads. The next flip is still 50 percent, because the coin has no memory of the previous ten. Flip 990 more times and expect roughly 495 heads, giving about 505 heads in 1,000 flips, or 50.5 percent. The early run was not cancelled by a compensating run of tails. It was diluted by the volume of later trials. The gambler's fallacy expects correction; the law of large numbers delivers dilution. The gambler's fallacy is treated in full on its own page.
The same distinction governs how you read a run in any dataset: a cluster of failures on one production line in one week, three complaints from the same postcode, four straight quarters of growth. The question is not whether the run is surprising, but how many opportunities there were for some run to occur somewhere. Testing that properly is what the null hypothesis is for, and a run that would arise by chance in a fifth of all datasets is not evidence of a mechanism.
When the attack lands on the source instead of the number
An attack on the source instead of the number is the point at which a statistical argument stops being statistical. Two moves account for nearly all of it. The straw man fallacy restates a figure as a stronger or stupider claim than anyone made, then refutes the restatement: someone reports that a defect rate rose from 0.02 to 0.03 percent, and the reply argues against the claim that the product is now dangerous, which nobody advanced. The ad hominem fallacy attacks whoever produced the figure rather than the figure: the survey was run by an interested party, so the survey is dismissed unread.
Both are misleading in the same specific way, and it is worth naming: they leave the number untouched. Funding and motive are legitimate reasons to check a methodology harder, but they are reasons to check it, not conclusions about what it found. The repair in both cases is the same. Restate the claim in the weakest form its author would accept, then test that. If you cannot find anything wrong with the number itself, saying who paid for it has not moved the argument.
The five questions to ask before you repeat a statistic
Ask these five questions of any statistic before you pass it on, in this order, and stop at the first one you cannot answer.
- Count the denominator: how many cases, out of how large a group, over what period?
- Convert the relative figure to an absolute one, in cases per thousand or per hundred thousand.
- Find the comparison: what is this being measured against, and who chose that baseline?
- Locate the missing group: who was measured, and who could not have been?
- Ask what result would have been reported if the effect were not real at all.
The fifth question is the one that does the most work and the one people skip. A statistic that would look identical whether or not the underlying effect exists is not evidence of anything, however precisely it is reported. That is the honest limit of this whole page: none of these checks tells you a claim is false. They tell you the number in front of you has not yet earned the conclusion attached to it, which is a smaller and far more useful thing to know.