Bayes from counts
A spam filter catches 90% of spam and wrongly flags only 5% of legitimate mail. It just flagged a message. How likely is it that the message is actually spam?
It depends on how much of the mail is spam to begin with, the base rate. In this inbox it's 2%. With that, the answer turns out to be about 27%. Most people guess 90%, which confuses "how often spam gets flagged" with "how often flagged mail is spam". The rest of the lesson makes that difference visible, starting from counts.
A thousand messages
Take 1,000 messages from this inbox. 2% are spam, so 20 pink dots and 980 grey ones.
\(20 \times 0.9 = 18\). Two spam messages slip through.
\(980 \times 0.05 = 49\). A small rate applied to a big group makes a big number. The ringed dots in the grid are the 67 flagged messages.
The evidence selects a group
You're told only one thing about the message: it was flagged.
Only the flagged ones. The other 933 are ruled out by the evidence, so they drop out of the picture. The grid above now gathers the 67 flagged messages into their own group. That group is the new "everything":
The 49 false alarms outnumber the 18 real catches because legitimate mail is 49 times more common.
Bayes' rule is that fraction, divided by 1,000
Divide every count by the population size and each one becomes a probability you've already met:
So the fraction \(18/(18+49)\) is
The denominator \(P(\text{flag})\) is the size of the flagged group: every way the evidence can occur. The two conditional probabilities point in opposite directions:
- \(P(\text{flag}\mid\text{spam}) = 0.90\) is the likelihood, a property of the filter. In ML it's the filter's recall.
- \(P(\text{spam}\mid\text{flag}) = 0.27\) is the posterior, what you believe after the evidence. In ML it's the filter's precision.
Recall doesn't depend on the base rate. Precision does. Change the inbox and watch:
The spam-rate slider is logarithmic, from 0.1% to 50%. Dot counts are rounded to whole messages; the percentages are exact.
Call the odds
Five cases. For each, state your probability. You're scored like lesson 1: the expected bits you lose against a perfect Bayesian, \(\mathrm{KL}(\text{truth}\,\|\,\text{you})\). Confident and wrong costs a lot; honest uncertainty costs little.
Evidence multiplies odds
Case 3 had two flags. You could redo the counts, but there's a shortcut. Write beliefs as odds, spam : legitimate. The prior odds are \(20 : 980 = 1 : 49\). Each piece of evidence multiplies the odds by its likelihood ratio:
\(1/49 \times 18 = 18/49\): exactly the 18 versus 49 counts. A second independent filter (recall 0.8, false-positive rate 0.1, so a ratio of 8) multiplies again: \(18/49 \times 8 = 144/49\), about 75%.
Take logs and multiplication becomes addition. Each piece of evidence adds its log-likelihood ratio to the log-odds:
That's naive Bayes. It's also the form of logistic regression, \(\sigma(b + \sum_i w_i x_i)\), where the intercept plays the prior and each weight plays a log-likelihood ratio. Naive Bayes derives the weights by assuming the features are independent given the class. Logistic regression fits them directly and so can discount correlated evidence that naive Bayes would double-count.
Apply it somewhere else
Downsampling negatives by \(w = 0.1\) multiplied the training odds by \(1/w = 10\). Undo it: true odds = model odds × \(w\) = \(1 \times 0.1\), so \(p = 0.1/1.1 \approx 0.091\). Equivalently \(q = p/(p + (1-p)/w)\), the correction He et al. (2014) use for ads click prediction. Ranking is unaffected, since the map is monotone. Anything that uses the probability's value, like auctions, thresholds or blending, is wrong until you correct it.
Recall and the false-positive rate are conditioned on the true class, so they carry over if the classes themselves look the same. Precision is a posterior, and it moves with the base rate. With 10× fewer fraud cases, the false alarms dominate: precision falls to roughly 45%. Always report precision at the deployment base rate.
Three impressions are weak evidence against a strong prior. With a Beta prior worth, say, 100 pseudo-impressions at 2%, the estimate is \((2 + 2)/(3 + 100) \approx 3.9\%\). The data moves it up from 2%, but not to 67%. This is empirical-Bayes smoothing, standard for cold-start click rates, and it's lesson 1's add-one smoothing with a prior fitted to your data.
Accuracy is dominated by the base rate. Look at precision and recall (or PR-AUC) for the positive class, and compare against the trivial baseline.
Write the updates
Plain editor, hidden tests. ⌘/Ctrl+Enter runs.
Exercise 1: posterior from a test result
Exercise 2: undo negative downsampling
Answer out loud, then check
- Likelihood \(P(\text{data}\mid\theta)\) is a function of \(\theta\) for fixed data. It doesn't sum to 1 over \(\theta\). The posterior \(P(\theta\mid\text{data}) \propto\) likelihood × prior does.
- MLE maximizes the likelihood. MAP maximizes the posterior, and a Gaussian prior on weights gives L2 regularization (a Laplace prior gives L1).
- In evaluation: recall is a likelihood-like quantity (conditioned on the true class). Precision is a posterior that depends on the base rate.
- Correct the odds by the sampling rate: \(q = p/(p + (1-p)/w)\), or add \(\log w\) to the logit.
- Then check calibration on unsampled holdout data (reliability plot, calibration by slice), because the model can be miscalibrated for other reasons.
- Ranking metrics (AUC, NDCG) are unaffected. Log loss, auctions, and thresholds are not.
- Beta-Binomial empirical Bayes: fit \(\mathrm{Beta}(\alpha,\beta)\) to the click rates of established items, then estimate \((c + \alpha)/(n + \alpha + \beta)\).
- Shrinkage strength comes from the data, and it can be done per segment (category, creator) for a better prior.
- For ranking under uncertainty, consider an upper confidence bound or Thompson sampling from the posterior to give new items exposure (explore/exploit).
- Both are linear in log-odds. Naive Bayes is generative: it estimates class-conditional feature distributions, assumes independence, and needs little data.
- Logistic regression is discriminative: it fits the weights directly, handles correlated features, and is usually better calibrated and more accurate with enough data.
- Ng and Jordan (2002): naive Bayes reaches its (higher) asymptotic error faster. It's a good baseline and a good fit for tiny data.
- Not 96%. A p-value is \(P(\text{data this extreme}\mid H_0)\), not \(P(H_0\mid\text{data})\), which is the same inversion as recall vs precision.
- The probability that a "significant" result is real depends on the prior share of ideas that work and on power. With low base rates and low power, many wins are false.
- A Bayesian analysis with an informed prior gives that probability directly. Lesson 7 covers tests and peeking.
Sources: He et al., Practical Lessons from Predicting Clicks on Ads at Facebook (ADKDD 2014), for the downsampling recalibration; Ng & Jordan, On Discriminative vs. Generative Classifiers (NeurIPS 2002).