Numbers & Logic

Simpson's Paradox

Berkeley looked biased, the surgery looked worse, the batter looked better — until someone split the data.

In the autumn of 1973, the University of California, Berkeley looked at its graduate admissions figures and saw a scandal in the making. Some 44 per cent of male applicants had been admitted, against roughly 35 per cent of women — a gap glaring enough, in the era of new anti-discrimination law, to demand an explanation. (A widely repeated version of the tale has the university being sued over the numbers; the record shows no lawsuit, only a nervous administration.) So the statistician Peter Bickel and his colleagues Eugene Hammel and J. W. O’Connell went through the data department by department — and found the scandal had vanished. In most departments, admission rates for men and women were about the same, and if anything the tilt ran slightly in women’s favour. Their 1975 paper in the journal Science became one of the most famous case studies in all of statistics.

How can every part of a university treat women fairly while the whole appears biased against them? The answer was hiding in where people applied. Men flocked to departments such as engineering, which admitted a large share of applicants. Women applied in far greater numbers to fiercely competitive departments — the paper’s authors pointed to fields like English — where nearly everyone, of either sex, was rejected. Each department could be even-handed, yet the university-wide average still punished the group that had queued at the harder doors. The aggregate number was not wrong; it was answering a different question.

This is Simpson’s paradox: a trend that appears in every subgroup of the data can weaken, vanish or fully reverse when the groups are combined. It is not a trick of small samples or sloppy measurement. It is honest arithmetic — the mathematics of weighted averages — colliding with a lurking third variable, which statisticians call a confounder.

The kidney stone trap

The medical classic comes from a 1986 study in the British Medical Journal, led by Charles Charig, comparing two treatments for kidney stones. Overall, keyhole removal (percutaneous nephrolithotomy) succeeded in 83 per cent of cases against 78 per cent for traditional open surgery — an apparently clear win for the modern technique. Then split the patients by stone size. For small stones, open surgery did better, 93 per cent to 87. For large stones, open surgery did better again, 73 per cent to 69. Open surgery won both categories and lost the total, because surgeons had steered the difficult, large-stone cases towards the operating theatre and the easy ones towards the keyhole. The old operation was handicapped by being trusted with the hard jobs. A patient choosing by the headline number would have chosen against the evidence.

Sport supplies the cleanest toy version. In the 1995 American baseball season, David Justice posted a higher batting average than Derek Jeter; in 1996, Justice edged him again. Combine the two seasons, and Jeter comes out ahead — because Jeter compiled vastly more at-bats in his strong year, while Justice’s better averages rested on lopsided sample sizes. Every pub argument about averages, league tables and pass rates is one weighted average away from the same ambush.

Who was Simpson, anyway?

The name is a fine example of Stigler’s law — that no scientific discovery is named after its original discoverer. The effect was described by the great Victorian statisticians Karl Pearson in 1899 and George Udny Yule in 1903, half a century before Edward H. Simpson, a British statistician and wartime Bletchley Park codebreaker, wrote the 1951 paper that examined it closely. The label “Simpson’s paradox” was only coined in 1972, by Colin Blyth, and purists still prefer the fairer “Yule–Simpson effect”. Simpson himself lived until 2019, long enough to see his name attached to one of the most cited curiosities in the discipline.

What should you actually do when the subgroup story and the aggregate story disagree? There is no mechanical rule; it depends on what causes what. In the kidney stone data, stone size influenced both the choice of treatment and the odds of success, so the honest comparison is within the size groups. In other datasets, slicing too finely creates the illusion instead, and the aggregate is the truthful number. Modern causal-inference researchers, most prominently Judea Pearl, treat the paradox as the founding puzzle of their field: the data alone can never tell you which reading is right, because the answer hangs on the causal story behind the numbers. That is also why the paradox is a favourite tool of the motivated: with a judicious choice of grouping, the same spreadsheet can be made to argue either side.

Simpson’s paradox is unsettling precisely because nobody is lying. The percentages are correct in every cell of the table; only our instinct that the whole must resemble its parts is at fault. Numbers never lie, the saying goes — but they answer only the exact question they were asked, and the whole art lies in noticing which question that was.

Quiz nuggets

  • In autumn 1973, Berkeley admitted about 44 per cent of male graduate applicants and 35 per cent of women, yet the 1975 Science paper by Bickel, Hammel and O’Connell found most departments showed no bias against women.
  • In the 1986 BMJ kidney stone study, keyhole treatment won overall (83 v 78 per cent) but open surgery won for both small stones (93 v 87) and large stones (73 v 69).
  • The paradox is named after Edward H. Simpson, a Bletchley Park codebreaker whose key paper appeared in 1951; Colin Blyth coined the term “Simpson’s paradox” in 1972.
  • Because Karl Pearson (1899) and George Udny Yule (1903) described the effect first, it is also called the Yule–Simpson effect.
  • David Justice out-hit Derek Jeter in both the 1995 and 1996 seasons, yet Jeter’s combined two-year batting average was higher.

Written from public sources and not individually checked — worth confirming before you stake a pint on it.