Bayes' Theorem
- Posterior = prior × likelihood rationot yet tested
- Base-rate fallacy & priors that matternot yet tested
- Frequentist vs Bayesian inferencenot yet tested
- Bayes 1763 + the 20th-century revivalnot yet tested
In 1763, two years after his death, an English Presbyterian minister named Thomas Bayes had a paper on probability published by a friend who thought it deserved attention. The mathematics was modest. The conceptual move was not. Bayes argued that belief itself was a quantity — specifically, a ratio of evidence to prior expectation — and that one could update beliefs with formal precision as new data arrived. The world's first mathematical theory of how to change one's mind was published anonymously, ignored for fifty years, and then quietly absorbed into the foundations of every science that learns from data. It was Pierre-Simon Laplace, working independently a generation later, who gave the rule its general form and its first real teeth — estimating the mass of Saturn, the sex ratio at birth — so that what we now call Bayesian inference is as much his achievement as Bayes's.
Bayes' theorem says: the posterior odds (your belief after the evidence) equal the prior odds (what you thought before) times the likelihood ratio (how much more likely the evidence is under one hypothesis than the other). The mechanics are an afternoon's algebra. The discipline is a lifetime's: most reasoning errors are prior errors — beliefs that haven't been updated when they should have, or have been updated when they shouldn't. The theorem also explains the base-rate fallacy: a 99%-accurate cancer test is not 99% likely to mean you have cancer, because the prior probability of cancer is low. Work it through: if 1 in 100 people carry the disease, a test that catches 99% of true cases but also flags 5% of healthy ones will, in a town of 10,000, raise about 99 true alarms and some 495 false ones — so a positive result leaves you only about one chance in six of actually being ill. The arithmetic is unforgiving, and it is exactly the arithmetic intuition skips. Because the rule is iterative — today's posterior becomes tomorrow's prior — a Bayesian reasoner is never finished: every observation nudges the estimate, and a lifetime of small, honest updates can carry one a long way from the starting point. The catch it hands back is that your conclusion is only as good as your prior; feed it a confident, wrong starting belief and it will update too slowly to rescue you, which is why widening one's priors is part of the method, not a preliminary to it. The statistical wars of the twentieth century — frequentist vs. Bayesian — were largely a fight over whether prior beliefs could legitimately enter scientific reasoning. The Bayesians won the practical fight, partly because computers can finally do the integrals, and partly because the world is full of problems where you cannot collect more data and must reason from what you have.