Sample size
A sample size calculator turns three decisions into one number: how precise you want the answer, how confident you want to be in it, and how varied the thing you are measuring is. For the most common case, a survey estimating a percentage at 95 percent confidence, the whole calculator reduces to about 385 respondents for a margin of error of 5 points and about 1,068 for a margin of 3 points, whatever the size of the population behind them.
Sample size is the count of units actually measured: people surveyed, boards inspected, sessions logged. It controls one thing very well and another thing not at all. It controls how much random noise sits around your estimate. It does nothing whatever about a sample that was drawn from the wrong group, which is why a huge sample can be far more wrong than a small one.
What sample size is
Sample size is the number of observations used to estimate a property of a larger population, and its job is to shrink random error to a level you can live with. The relationship is not linear, and that is the single most useful fact about it: precision improves with the square root of the sample, so quadrupling the sample halves the margin of error.
What distinguishes sample size from the other levers in a study design is how quickly it stops paying. Going from 100 to 400 respondents takes the margin of error from roughly 10 points to roughly 5. Going from 400 to 1,600 takes it from 5 to 2.5. Going from 1,600 to 6,400 takes it from 2.5 to 1.25. Each step costs four times as much as the last and buys half as much precision as the step before, which is why national surveys cluster around 1,000 to 2,000 respondents rather than continuing upward.
The one-sentence version: what sample size buys and what it cannot buy
In one sentence: sample size buys precision around whatever number your sampling method is aiming at, and it cannot move the aim. Precision and accuracy are separate properties, and a sample size calculator only ever quotes you the first.
Picture a survey with 1,000,000 self-selected responses. The margin of error works out at roughly 0.1 percentage points, a figure precise enough to publish to one decimal place. If the people who chose to respond differ systematically from those who did not, that razor-sharp estimate is centered on the wrong value, and it stays wrong no matter how many more responses arrive. Adding respondents narrows the interval; it does not move it. Deciding who gets into the sample in the first place is a separate problem, treated under selection bias.
The arithmetic worked through: from a margin of error to a number of respondents
Work the arithmetic on the standard formula for estimating a proportion, which is what nearly every sample size calculator runs underneath. Required sample size n equals z squared, times p times one minus p, divided by e squared, where z is the confidence multiplier, p is the expected proportion, and e is the margin of error you will accept.
- Set z at 1.96, the multiplier for 95 percent confidence. Squared, that is 3.8416.
- Set p at 0.5, because 0.5 times 0.5 equals 0.25, the largest the term can be, which makes the answer safe for any proportion.
- Multiply: 3.8416 times 0.25 gives 0.9604. That numerator is fixed for every 95 percent calculation.
- Divide by e squared. For a 5 point margin, e is 0.05 and e squared is 0.0025.
- 0.9604 divided by 0.0025 equals 384.16, so 385 respondents after rounding up.
Change only the margin of error and the whole table falls out of the same division:
| Margin of error at 95 percent | e squared | 0.9604 divided by e squared | Respondents needed |
|---|---|---|---|
| Plus or minus 10 points | 0.01 | 96.04 | 97 |
| Plus or minus 5 points | 0.0025 | 384.16 | 385 |
| Plus or minus 3 points | 0.0009 | 1,067.1 | 1,068 |
| Plus or minus 2 points | 0.0004 | 2,401.0 | 2,401 |
| Plus or minus 1 point | 0.0001 | 9,604.0 | 9,604 |
Read the last two rows together. Halving the margin from 2 points to 1 point takes the requirement from 2,401 to 9,604, exactly four times as many. That is the square root relationship in plain arithmetic, and it is the reason precision beyond about 2 points is rarely bought.
Notice what is missing from the formula: the size of the population. A sample of 1,068 estimates a country of 300 million to within 3 points, and it estimates a town of 30,000 to within 3 points as well. The correction for a finite population takes 1,068 down to 1,032 for a population of 30,000, and leaves it essentially unchanged for anything above about a million. The instinct that a bigger population needs a proportionally bigger sample is wrong, and it is wrong by a wide margin.
What a failure looks like: the sample that was large and useless
A sample size failure looks like a study whose number of respondents is quoted as though it settled the question. Two versions turn up constantly. The first is the enormous convenience sample: 50,000 responses collected from whoever happened to see the link, reported with a margin of error of 0.4 points. The precision is real and the estimate is aimed at the population of people who saw the link and felt strongly enough to answer, which is not the population named in the headline.
The second is the subgroup split. A survey of 1,200 people has a margin of error near 3 points overall. Break it into twelve subgroups of about 100 each and every subgroup carries a margin of error near 10 points. A chart showing that group A sits at 46 percent and group B at 52 percent, each measured on 100 respondents, is showing a difference that the design cannot resolve. The overall sample size was adequate. The sample size for the claim being made was 100.
The common misreading: percentage of the population
The common misreading is that a sample must be a fixed percentage of the population, usually stated as 10 percent. It does not. The formula contains no population term until the population drops below roughly 20,000, and even then the correction is modest. A representative sample of 1,000 from 300 million is not a thousandth of a percent of a valid study. It is a complete one.
Two smaller misreadings ride along with it. The first is that a bigger sample makes a result more likely to be true; it makes the estimate more precise, and precision is only useful once the sampling method is sound. The second is that a non-significant result from a large sample proves there is no effect. That inference belongs to the null hypothesis and its interpretation, and the honest version is narrower: a large well-drawn sample that finds nothing has ruled out effects above a certain size, and has said nothing about effects below it.
The test to run before you accept a sample size
Run these five checks on any study that quotes its number of respondents, and run them in order.
- Find the margin of error, and if it is not stated, take 98 divided by the square root of n as a rough percentage.
- Identify the population the sample was meant to represent, in one sentence.
- Locate the people who could not have been sampled at all, given how the sample was recruited.
- Check the sample size for the specific claim, not for the whole study, whenever a subgroup is being discussed.
- Ask whether the difference being reported is larger than the margin of error around each figure.
The formula, its symbols and where the standard version stops working
The standard sample size formula, sometimes written as the sample size determination equation, applies to one situation: estimating a single proportion from a simple random sample. Four common departures change it, and a calculator that offers only the standard version will quietly get all four wrong.
- Estimating a mean rather than a proportion replaces the p term with an estimate of the standard deviation, so n equals z squared times sigma squared, divided by e squared, and requires a prior guess at how variable the measurements are.
- Comparing two groups rather than estimating one quantity requires a power calculation, which adds the smallest difference worth detecting and the tolerated false negative rate to the inputs.
- Sampling in clusters, such as selecting sites and then surveying everyone at each site, inflates the requirement by a design effect, because people within a cluster resemble each other.
- Sampling from a small population applies the finite population correction, which divides the raw n by one plus n minus one over the population size.
All four require a judgment made before the data exist, and that is the honest limit of any sample size calculator: it converts assumptions into a number, and hands back exactly the confidence you put in. The habit of translating that number back into raw counts before arguing about it is what Reading numbers is for, and the neighboring error of reading a run of results as meaningful in itself is the gambler's fallacy.