Numbers & Logic

Zipf's Law

The second commonest word appears half as often as the first — in every language, and in city sizes too.

Open any sufficiently long English text — a novel, a newspaper archive, the whole of Wikipedia — and count the words. "The" will come first, making up roughly 7 per cent of everything written. "Of" comes second, at about half the frequency of "the". "And" comes third, at roughly a third. Carry on down the list and the pattern holds with eerie discipline: the word at rank n appears about 1/n as often as the champion, so the hundredth word is about a hundredth as common. Plot rank against frequency on logarithmic axes and the points fall along a near-perfect straight line. Nobody planned this. No academy enforces it. Yet it holds for English, Spanish, Hindi, Mandarin and Latin — and, as far as anyone can tell, for texts that nobody has ever managed to read.

The pattern bears the name of George Kingsley Zipf, a Harvard linguist who documented it obsessively in The Psycho-Biology of Language in 1935 and, most famously, in Human Behavior and the Principle of Least Effort in 1949, a year before his death at just 48. Zipf was not actually first. Jean-Baptiste Estoup, a French stenographer, had noticed the regularity in the 1910s while analysing word frequencies to improve shorthand systems. But Zipf did what scientific eponyms usually reward: he measured relentlessly, generalised boldly, and proposed a grand explanation — that language is a compromise in a tug-of-war of laziness, with speakers preferring few, reusable words and listeners preferring many, precise ones.

From dictionaries to city maps

What lifted the law from linguistic curiosity to scientific obsession is that words are only the beginning. The German physicist Felix Auerbach had already reported in 1913 that city populations follow the same ladder: rank a country's cities by size and the second tends to hold roughly half the people of the first, the third roughly a third. Urban economists still call this the rank-size rule, and it fits many countries remarkably well — with famous exceptions, since "primate" capitals such as London and Paris are far larger than their nations' second cities, dwarfing what the rule predicts. Since then Zipf-like distributions have been reported for the sizes of firms, the traffic of websites, the populations of American metropolitan areas and the wealth of the richest — the last echoing Vilfredo Pareto's work on income concentration from the 1890s.

Why should one curve govern chatter and cities alike? Zipf's own "least effort" story has competitors. Herbert Simon showed in the 1950s that a simple rich-get-richer process — where popular words or big cities attract new arrivals in proportion to their existing size — generates the distribution automatically. Benoit Mandelbrot sharpened the formula into the Zipf–Mandelbrot law, which fits the top-ranked words better. More deflating still, statisticians have shown that even random typing — monkeys hitting keys with an occasional space bar — produces Zipf-like statistics, an argument associated with the psychologist George Miller. The debate over whether the law reveals something deep about human communication or is a statistical near-inevitability remains genuinely unsettled.

The manuscript test

The law has real forensic uses. The Voynich manuscript, a lavishly illustrated codex carbon-dated to the early fifteenth century and held at Yale's Beinecke Library, is written in a script no one has ever deciphered. Its word frequencies, however, follow Zipf's law rather faithfully — a point often cited as evidence that the text has language-like structure rather than being pure invented gibberish, though sceptics note that a clever hoax could produce the same signature. Researchers have looked for Zipfian statistics in bottlenose dolphin whistles as a possible marker of communication, and the idea has even been floated as a first filter for any signal from beyond Earth: before translating a message, check whether its symbols keep Zipf's books.

The law also explains a striking economy in everyday speech. Because frequency collapses so quickly down the ranks, a tiny elite of words does most of the work: in the Brown Corpus — the million-word sample of American English compiled at Brown University in the 1960s — around 135 word types account for half of the entire running text. Meanwhile, at the far end of the ladder, roughly half of a corpus's distinct words appear exactly once; linguists call these one-off visitors hapax legomena. Every conversation you have ever had was built mostly from a few dozen workhorses, trailed by a vast crowd of words used almost never.

Zipf believed he had glimpsed a universal law of human effort, and perhaps he had — or perhaps he had merely found the shape that ranking anything tends to take. Either way, the ledger balances with unreasonable precision. Language is a crowd of lazy choices, and crowds, it turns out, keep immaculate accounts.

Quiz nuggets

  • "The" is English's commonest word at roughly 7 per cent of running text; "of", at rank two, appears about half as often.
  • Harvard linguist George Kingsley Zipf published Human Behavior and the Principle of Least Effort in 1949.
  • French stenographer Jean-Baptiste Estoup spotted the rank-frequency pattern in the 1910s, before Zipf.
  • Felix Auerbach documented the matching rank-size rule for city populations in 1913.
  • In the million-word Brown Corpus, about 135 words account for half of all the running text.

Written from public sources and not individually checked — worth confirming before you stake a pint on it.