Grain of Salt

The Dunning Kruger effect, and what the study actually showed

The Dunning Kruger effect

Struck diagram on assay stock: The Dunning Kruger effect

The Dunning Kruger effect is a claim about calibration: people who perform worst on a test tend to overestimate their performance by the widest margin, while the best performers tend to underestimate theirs slightly. It comes from a 1999 paper by Justin Kruger and David Dunning, and it is a narrower claim than the version that circulates online, which is worth establishing before anything else, because this is a case where the popular account and the tested one have come apart. The finding concerns the gap between measured score and self rated score. It does not say that ignorant people are uniquely arrogant.

What the Dunning Kruger effect is

The Dunning Kruger effect describes a pattern in self assessment: across a test, self ratings are compressed toward the middle, so the bottom quarter of scorers place themselves far above where they landed and the top quarter place themselves a little below. Both halves of that sentence matter. Almost every popular retelling keeps the first half and drops the second, which turns a symmetrical finding about miscalibration into an insult aimed at other people.

What distinguishes it from its siblings in the cognitive bias family is its subject. Confirmation bias distorts your search for evidence about the world. The Dunning Kruger effect concerns your estimate of one specific quantity, your own competence, and it is measured against a hard number you can check, which is the score you actually got. That makes it unusually testable and also unusually easy to get wrong statistically, as the later argument about it shows.

The effect in one sentence, and the size of the gap

In one sentence: the less skilled you are at something, the less equipped you are to recognize that you are unskilled, because the knowledge required to do the task well is close to the knowledge required to judge whether it was done well. Kruger and Dunning called this the dual burden. A person with a weak grasp of grammar cannot reliably spot the sentences they got wrong, precisely because spotting them is the skill under test.

The gap in the original data was large. In the studies where students took a test and estimated their percentile rank against their peers, the bottom quarter scored near the bottom of the distribution and yet placed themselves above the midpoint, a difference on the order of fifty percentile points. The top quarter's error ran the other way and was much smaller. Dunning and Kruger attributed the top group's error to a different mechanism: strong performers assume the task was about as easy for everyone else, so they underrate how far ahead they are.

The 1999 study, and what Kruger and Dunning actually measured

The evidence behind the effect is a set of four studies reported by Justin Kruger and David Dunning in the Journal of Personality and Social Psychology in 1999, under the title "Unskilled and Unaware of It". Undergraduates at Cornell University took tests in three domains, humor, logical reasoning and English grammar, and then estimated both their raw score and their standing relative to other participants. The researchers sorted participants into quartiles by actual score and plotted perceived against actual performance for each quartile.

The origin story is worth telling because it is often mangled. The research was prompted by a Pittsburgh bank robbery in 1995 in which the man arrested had covered his face in lemon juice, apparently believing it would keep him off the security cameras. Dunning's interest was not the robbery itself but the question it raised: if you are bad enough at something, does the same deficit hide the deficit from you?

The statistical criticism, which is still a live argument

A serious methodological objection has been made to the standard chart, and it deserves a fair hearing rather than a dismissal. When you sort people by their score on a noisy measure and then look at anything else about them, the extreme groups move toward the average on the second measurement. That is regression to the mean, and it is not a psychological effect at all. Add a floor and a ceiling, since nobody can rate themselves below zero or above the hundredth percentile, and a plot resembling the published one can appear in data with no self assessment deficit in it whatever.

Gilles Gignac and Marcin Zajenkowski, writing in the journal Intelligence in 2020, argued that the effect is largely a statistical artifact of this kind. Edward Nuhfer and colleagues, in a series of papers in the journal Numeracy, showed that simulations using random numbers produce quartile plots that look much like the ones people cite as evidence. Defenders reply that the better than average tendency remains after several of these corrections, and that the dual burden claim can be tested in ways that do not depend on quartile plots. The honest summary is that the broad phenomenon of poor calibration is well supported, and the specific quartile chart is a weak way to demonstrate it.

Where the Dunning Kruger effect shows up

The effect shows up wherever self assessment is cheap and feedback is slow or absent. Four settings account for most of the real cases:

  • Early stage learning. A person three weeks into a new tool has learned enough to use it and not enough to know what it does badly.
  • Fields without a scoreboard. Where no test returns a number, such as interviewing, management or writing, nothing corrects the estimate.
  • Delayed feedback. A structural engineer finds out about a design fault years later, so the self rating floats free in the meantime.
  • Adjacent expertise. Genuine skill in one domain inflates the self rating in a neighboring one, which is the halo effect operating on yourself.

Notice what is not on that list: being a stupid person. The effect is about a task, not a personality. The same individual is well calibrated in the domain they have practiced for a decade and badly calibrated in the one they took up last month.

What the Dunning Kruger effect is confused with

The Dunning Kruger effect is confused with three other things, and each confusion has a single question that settles it: imposter syndrome, ordinary overconfidence, and regression to the mean.

Imposter syndrome, which is a felt state rather than a measured gap

The Dunning Kruger effect and imposter syndrome are not two ends of one scale, though they are constantly presented that way. The impostor phenomenon was described by Pauline Clance and Suzanne Imes in 1978 as a persistent internal experience of intellectual fraudulence in people whose achievements are objectively strong. It is a felt state, self reported, with no test score in the design.

The Dunning Kruger effectImposter syndrome
What is measuredEstimated score against actual scoreA reported feeling of not deserving one's position
Who shows itThe whole sample, in opposite directions by quartileOften high performers
SourceKruger and Dunning, 1999Clance and Imes, 1978
StatusA measured pattern, with a live statistical disputeA described experience, not a clinical diagnosis

Calling the effect the opposite of imposter syndrome is a category error: one compares two numbers, the other reports a feeling. A person can plausibly have both, feeling like a fraud while still overestimating a specific skill. The question that separates them: is there a test score in this claim, or only a self report?

Ordinary overconfidence, which is flat rather than steepest at the bottom

Overconfidence in general is the well documented tendency of people to rate themselves above the average, and it needs none of the machinery of the Dunning Kruger effect to explain it. The two make different predictions about shape. Plain overconfidence predicts roughly the same inflation at every level of skill. The Dunning Kruger claim is specifically that the inflation is widest among the weakest performers, because the deficit and the ability to notice the deficit are the same knowledge.

The question that separates them: does the size of the error change as measured skill rises, or does everyone overrate themselves by about the same amount? If the gap is flat, what you are looking at is ordinary overconfidence wearing a more interesting name.

Regression to the mean, which is a property of noisy measurement

The third confusion is the one that matters most, because it is the substance of the criticism set out above. Regression to the mean is not a psychological effect at all: it is what happens to extreme groups on any second measurement when the first measurement contains error. Readers routinely take the quartile chart as a demonstration of a mental deficit when a chart of that shape can be produced by a process containing no deficit whatever.

The question that separates them: would this pattern still appear if the self ratings had been generated at random? If the answer is yes, as the simulation work indicates it can be, then the chart is consistent with the effect but is not evidence for it, and something better than a quartile plot is needed to settle the matter.

What actually reduces the effect: training, not humility

To reduce the effect, improve the skill, because the fourth study in the original paper found that training changed the self estimate. Participants in the bottom quartile who were taught the logical reasoning task afterward rated their earlier performance more accurately. They did not become modest. They became able to see what they had missed, which is a different thing.

  1. Take the actual test and compare your predicted score with your real one, in writing, before you see the result.
  2. Learn the failure modes of the task, since knowing what a bad answer looks like is the calibration skill itself.
  3. Get feedback from someone who can outperform you, because feedback from a peer at your level cannot detect what neither of you can see.
  4. Track calibration over time, keeping a short log of predictions and outcomes rather than an impression of how you are doing.

Telling yourself to be humble does nothing here, in the same way that resolving to be open minded does nothing about confirmation bias. And beware the reverse move: continuing with a bad plan because of the effort already invested in it is the sunk cost fallacy, not calibration.

The test to run on your own estimate of your skill

Ask the question that separates a calibrated estimate from a comfortable one: can I name three specific ways this could be wrong, in the vocabulary of the field, and say which is most likely? Someone who can name them is describing the shape of their own ignorance, which requires having some. Someone who cannot may be an expert, but has no evidence of it available at that moment.

A second question works on other people's confidence without insulting anyone: what does this person get told when they are wrong, and how quickly? Confidence built where feedback is fast is worth more than confidence built where feedback never arrives.

The famous curve, and why it is not in the paper

The chart everybody shares, the one with a spike of confidence early on, a collapse into a valley, and a slow climb to a plateau, does not appear in the 1999 paper. It is a later illustration, drawn by people summarizing the idea, and its labels, the peak of a mountain and a valley of despair, are internet folklore rather than anything Kruger and Dunning measured. The graphic is memorable, funny and wrong about the data.

What the paper contains instead are quartile plots: two lines across four groups, one for actual performance and one for perceived performance, with the perceived line much flatter than the actual one. The lines cross near the upper quartiles. There is no curve, no peak, no valley, and no time axis at all, which is the detail that matters most: the popular chart shows confidence changing as a person learns, and the study measured no such thing. It photographed one moment, not a process running over months.

So the honest position is layered. The effect has not been debunked, and it has not been proved in the form the meme states. What survives is the modest, useful claim: self assessment is poorly correlated with performance, the poorest performers are the least able to detect it, and the chart you have seen is not evidence of anything.

Where to go next

Proudly powered by WordPress | Theme: Amber Blog by Crimson Themes.