Everyone who has sat a pub quiz has a theory about what comes up. Capital cities. Football. Something about the periodic table. The theories are held confidently and tested rarely, because testing one means counting, and nobody has the questions to count.
Five public archives of quiz questions do exist, and together they hold 684,841 of them. This page is what came out of counting. The analysis is the one that orders the material inside QuizMaxx, so it was done for a practical reason rather than an interesting one; the interesting part turned out to be how easy it is to measure this badly.
What was counted
Five archives, of very different sizes and very different national character:
| archive | questions | character |
|---|---|---|
| Jeopardy! | 540,384 | American, televised, decades deep |
| QANTA | 118,941 | American academic quiz bowl |
| QuizQuestionsUK | 11,242 | British, written for pub quizzes |
| PubQuizHQ | 10,068 | British, written for pub quizzes |
| OpenTDB | 4,206 | Crowd-sourced, general |
From those, 127,498 distinct answer entities — the things questions are about, once "Which country…" and "In which nation…" have been reduced to the same answer.
Counting appearances gives you a frequency table. It does not tell you whether that table describes the quiz you are actually going to sit on a Tuesday night in a pub in England, and those are not the same question. So the counting needs something to be checked against.
The ground truth was 301 answers from real quiz nights
QuizMaxx's Culture pack is distilled from questions that actually came up at a particular run of British quiz nights — 301 answers, recorded as they were asked. It is a small sample and a partial one, but it has the property no archive has: it is the thing being predicted rather than a proxy for it.
76% of those 301 answers are asked about by at least one of the five archives. That is the encouraging headline, and it is worth stating plainly before the discouraging part: the corpora do broadly know what a British pub quiz asks about. The question is which of them knows best.
The obvious measurement is wrong, twice, in opposite directions
There are two natural ways to score an archive against those 301 answers, and corpus size confounds both.
Raw coverage — what share of the 301 does this archive ask about at all? — rewards being big. An archive with half a million questions will eventually mention almost anything.
Hits per thousand questions corrects for that and overcorrects. There are only 301 answers available to hit, so density necessarily collapses as an archive grows towards them. The same British archive scored 28 per thousand at a thousand questions and 9.7 per thousand at eleven thousand — the archive grew, which is the only thing that had changed, and its score fell by two thirds.
The fix is to sample every archive down to the size of the smallest, 4,206 questions, and ask how many of the 301 it hits from there. Read the last column:
| archive | questions | raw coverage | per 1,000 | matched at 4,206 |
|---|---|---|---|---|
| Jeopardy! | 540,384 | 73% | 0.4 | 50 |
| QANTA | 118,941 | 39% | 1.0 | 26 |
| QuizQuestionsUK | 11,242 | 36% | 9.6 | 56 |
| PubQuizHQ | 10,068 | 30% | 8.9 | 43 |
| OpenTDB | 4,206 | 17% | 12.1 | 39 |
By raw coverage, Jeopardy! looks four times better than OpenTDB. By hits per thousand, OpenTDB looks thirty times better than Jeopardy!. Matched at a common size, they are 50 and 39 — under 30% apart, in the opposite order to the per-thousand ranking.
And the archive that actually wins is neither: QuizQuestionsUK, at 11,242 questions, beats a corpus forty-eight times its size at predicting what a British pub quiz will ask. That is not a subtle effect, and both of the obvious metrics hide it — one by rewarding the big American archive, the other by rewarding the small one.
Where the questions actually are
Weighted by how often they are asked, this is where the volume sits. "Covered" is the share QuizMaxx could answer at the time of the analysis, which is included because a gap is only interesting next to what fills it.
| category | ask share | covered |
|---|---|---|
| Science | 122 | 29% |
| Geography | 117 | 75% |
| History | 102 | 26% |
| Literature | 91 | 27% |
| Music | 78 | 52% |
| Film | 67 | 31% |
| Popular culture | 52 | 10% |
| Sport | 49 | 33% |
| Words & language | 35 | 18% |
| Food & drink | 31 | 16% |
| Myth & religion | 29 | 27% |
| Art | 29 | 30% |
| Television | 27 | 12% |
| Ideas & society | 26 | 17% |
The share column is in arbitrary units and only means anything relative to itself. Science leads, narrowly, and is one of the worst covered. Geography is second and much the best covered, which is a fact about flashcard apps rather than about quizzes: capitals and flags are the easiest thing in the world to turn into cards, so that is what everyone turns into cards. The categories that resist the format — popular culture at 10%, television at 12% — are the ones where the asking goes unanswered.
An archive asks about answers, not about questions
This is the finding with the most practical consequence, and it took a while to see.
A frequency table is indexed by the answer. So when the analysis says France is among the most-asked entities in the whole corpus, that is a fact about the word France and not about any particular question. QuizMaxx holds eight cards whose answer is France — the flag, the country with the most time zones, two football tournaments, the first Winter Olympics, the country the town of Volvic is in — and the frequency data assigns all eight the same very high score.
Ordering study material by that number alone would put six of a twenty-question first session on one word, and would hand the strongest score in the dataset to the weakest question in it. The rate is a real signal and it is not a measure of a question's worth; conflating the two is the trap sitting directly underneath this kind of analysis.
What the rate is good for is choosing which answers are worth knowing. What it cannot do is tell you whether a particular question about that answer is one a quizmaster would ever ask.
What the corpora cannot see
Three limits worth stating, because a number with its limits omitted is worse than no number.
They are overwhelmingly American. The three non-British archives are 97% of the questions and about 89% American by volume, which is why the calibration above exists at all. An analysis that took them at face value would build a syllabus for a different country's quiz.
They are text. A text corpus cannot score a photograph or a song, so picture rounds, music rounds and map questions rate zero throughout — not because they are unimportant, but because they are invisible to this instrument. Anyone reading the tables above as a complete account of a quiz night is reading them wrong.
The extraction is imperfect, and visibly so. The two highest-rate "entities" in the raw output are Answers: and Answer:, at rates several times higher than Leonardo da Vinci. They are scraping artefacts — formatting from the source pages that survived parsing. They are left in the generated report rather than quietly filtered, because a pipeline that hides its own noise is one you stop checking.
What this changed
Concretely: the archives are blended by the matched column rather than by size, so the two British archives carry 45% of the weight between them despite being 3% of the questions. Previously unseen material is introduced in measured frequency order, so an early session is spent on the answers most likely to come up. A per-answer cap stops one word — France — taking six slots in a twenty-question session. And the expansion list is the reverse of the coverage table: science, history, literature and popular culture, in that order.
Ask-weighted coverage across the whole syllabus was 31% at the time of the analysis. That number is the honest one to end on. It is not a boast — most of what quizzes ask about is not in here, and a good deal of what quizzes ask about is unbounded in a way no syllabus closes. It is a measurement, which is more than the confident theories at the start of this page had.
Method, data and reuse
The analysis is a Python script in the QuizMaxx repository; it regenerates the report, the per-answer rates and the deck-tier suggestions in one pass, so the numbers here are reproducible rather than recalled. The archives themselves are third-party and copyrighted, and are used as analysis input only: what is published is aggregate statistics about them, never their question text, and they are not redistributed. One well-known archive is deliberately absent because its robots.txt asks not to be used this way, and the scraper honours crawl delays, which is why a full sweep is a three-hour job.
If you want the practical version of all this rather than the measurements, the guide to getting better at pub quizzes sets out what the memory literature says about how to use a list like this, and the comparison with Anki covers why item selection, not scheduling, is the binding constraint. The 364 short essays are free and need no software. QuizMaxx itself runs in the browser with no account and no installation.