← Atelier
Lesson 05 · Probability · ~30 min + code

Bayes from counts

A spam filter catches 90% of spam and wrongly flags only 5% of legitimate mail. It just flagged a message. How likely is it that the message is actually spam?

Your first instinct, before any calculation:

It depends on how much of the mail is spam to begin with, the base rate. In this inbox it's 2%. With that, the answer turns out to be about 27%. Most people guess 90%, which confuses "how often spam gets flagged" with "how often flagged mail is spam". The rest of the lesson makes that difference visible, starting from counts.

1 · Build the counts

A thousand messages

Take 1,000 messages from this inbox. 2% are spam, so 20 pink dots and 980 grey ones.

spamlegitimateflagged by the filter
The filter flags 90% of spam. How many of the 20 spam messages are flagged?

\(20 \times 0.9 = 18\). Two spam messages slip through.

It also flags 5% of legitimate mail. How many of the 980 legitimate messages are flagged?

\(980 \times 0.05 = 49\). A small rate applied to a big group makes a big number. The ringed dots in the grid are the 67 flagged messages.

2 · Condition

The evidence selects a group

You're told only one thing about the message: it was flagged.

Which messages could it be?

Only the flagged ones. The other 933 are ruled out by the evidence, so they drop out of the picture. The grid above now gathers the 67 flagged messages into their own group. That group is the new "everything":

\[ \begin{aligned} P(\text{spam} \mid \text{flagged}) &= \frac{\#\,\text{spam and flagged}}{\#\,\text{flagged}} \\ &= \frac{18}{18 + 49} = \frac{18}{67} \approx 27\% \end{aligned} \]

The 49 false alarms outnumber the 18 real catches because legitimate mail is 49 times more common.

3 · Name it

Bayes' rule is that fraction, divided by 1,000

Divide every count by the population size and each one becomes a probability you've already met:

\[ \begin{aligned} \tfrac{18}{1000} &= \underbrace{0.9}_{P(\text{flag}\mid\text{spam})} \times \underbrace{0.02}_{P(\text{spam})} \\[4pt] \tfrac{49}{1000} &= \underbrace{0.05}_{P(\text{flag}\mid\text{legit})} \times \underbrace{0.98}_{P(\text{legit})} \end{aligned} \]

So the fraction \(18/(18+49)\) is

\[ P(\text{spam}\mid\text{flag}) = \frac{P(\text{flag}\mid\text{spam})\,P(\text{spam})}{P(\text{flag})} \]

The denominator \(P(\text{flag})\) is the size of the flagged group: every way the evidence can occur. The two conditional probabilities point in opposite directions:

Recall doesn't depend on the base rate. Precision does. Change the inbox and watch:

The same filter in different inboxes

The spam-rate slider is logarithmic, from 0.1% to 50%. Dot counts are rounded to whole messages; the percentages are exact.

4 · Play it

Call the odds

Five cases. For each, state your probability. You're scored like lesson 1: the expected bits you lose against a perfect Bayesian, \(\mathrm{KL}(\text{truth}\,\|\,\text{you})\). Confident and wrong costs a lot; honest uncertainty costs little.

5 · Odds and evidence

Evidence multiplies odds

Case 3 had two flags. You could redo the counts, but there's a shortcut. Write beliefs as odds, spam : legitimate. The prior odds are \(20 : 980 = 1 : 49\). Each piece of evidence multiplies the odds by its likelihood ratio:

\[ \begin{aligned} \underbrace{\frac{P(\text{spam}\mid e)}{P(\text{legit}\mid e)}}_{\text{posterior odds}} &= \underbrace{\frac{P(\text{spam})}{P(\text{legit})}}_{\text{prior odds}} \\ &\quad\times \underbrace{\frac{P(e\mid\text{spam})}{P(e\mid\text{legit})}}_{\text{likelihood ratio}} \end{aligned} \]
Filter A's flag has likelihood ratio \(0.9/0.05 = 18\). Starting from odds 1 : 49, what are the odds after one flag?

\(1/49 \times 18 = 18/49\): exactly the 18 versus 49 counts. A second independent filter (recall 0.8, false-positive rate 0.1, so a ratio of 8) multiplies again: \(18/49 \times 8 = 144/49\), about 75%.

Take logs and multiplication becomes addition. Each piece of evidence adds its log-likelihood ratio to the log-odds:

\[ \begin{aligned} \log\text{odds}(\text{spam}\mid e_1,\dots,e_n) &= \log\text{odds}(\text{spam}) \\ &\quad+ \sum_i \log\frac{P(e_i\mid\text{spam})}{P(e_i\mid\text{legit})} \end{aligned} \]

That's naive Bayes. It's also the form of logistic regression, \(\sigma(b + \sum_i w_i x_i)\), where the intercept plays the prior and each weight plays a log-likelihood ratio. Naive Bayes derives the weights by assuming the features are independent given the class. Logistic regression fits them directly and so can discount correlated evidence that naive Bayes would double-count.

6 · New cases, no hints

Apply it somewhere else

A click model was trained on data where negatives were downsampled to 10% (all positives kept). For an impression it predicts 0.50. What's the calibrated click probability?

Downsampling negatives by \(w = 0.1\) multiplied the training odds by \(1/w = 10\). Undo it: true odds = model odds × \(w\) = \(1 \times 0.1\), so \(p = 0.1/1.1 \approx 0.091\). Equivalently \(q = p/(p + (1-p)/w)\), the correction He et al. (2014) use for ads click prediction. Ranking is unaffected, since the map is monotone. Anything that uses the probability's value, like auctions, thresholds or blending, is wrong until you correct it.

Your fraud model has 95% recall and 90% precision on a test set with 5% fraud. In production fraud is 0.5%. What happens?

Recall and the false-positive rate are conditioned on the true class, so they carry over if the classes themselves look the same. Precision is a posterior, and it moves with the base rate. With 10× fewer fraud cases, the false alarms dominate: precision falls to roughly 45%. Always report precision at the deployment base rate.

A new item has 2 clicks in 3 impressions. The average item's click rate is 2%. What's a sensible estimate of this item's click rate?

Three impressions are weak evidence against a strong prior. With a Beta prior worth, say, 100 pseudo-impressions at 2%, the estimate is \((2 + 2)/(3 + 100) \approx 3.9\%\). The data moves it up from 2%, but not to 67%. This is empirical-Bayes smoothing, standard for cold-start click rates, and it's lesson 1's add-one smoothing with a prior fitted to your data.

A classifier is 99% accurate on a dataset where 1% of examples are positive. How impressed should you be?

Accuracy is dominated by the base rate. Look at precision and recall (or PR-AUC) for the positive class, and compare against the trivial baseline.

7 · Implement, unaided

Write the updates

Plain editor, hidden tests. ⌘/Ctrl+Enter runs.

Exercise 1: posterior from a test result

Exercise 2: undo negative downsampling

8 · Staff-level follow-ups

Answer out loud, then check

Likelihood vs posterior: what's the difference, and where does it bite in ML?
  • Likelihood \(P(\text{data}\mid\theta)\) is a function of \(\theta\) for fixed data. It doesn't sum to 1 over \(\theta\). The posterior \(P(\theta\mid\text{data}) \propto\) likelihood × prior does.
  • MLE maximizes the likelihood. MAP maximizes the posterior, and a Gaussian prior on weights gives L2 regularization (a Laplace prior gives L1).
  • In evaluation: recall is a likelihood-like quantity (conditioned on the true class). Precision is a posterior that depends on the base rate.
You downsample negatives to train a CTR model. What must you do before using its outputs?
  • Correct the odds by the sampling rate: \(q = p/(p + (1-p)/w)\), or add \(\log w\) to the logit.
  • Then check calibration on unsampled holdout data (reliability plot, calibration by slice), because the model can be miscalibrated for other reasons.
  • Ranking metrics (AUC, NDCG) are unaffected. Log loss, auctions, and thresholds are not.
How would you estimate click rates for items with very few impressions?
  • Beta-Binomial empirical Bayes: fit \(\mathrm{Beta}(\alpha,\beta)\) to the click rates of established items, then estimate \((c + \alpha)/(n + \alpha + \beta)\).
  • Shrinkage strength comes from the data, and it can be done per segment (category, creator) for a better prior.
  • For ranking under uncertainty, consider an upper confidence bound or Thompson sampling from the posterior to give new items exposure (explore/exploit).
Naive Bayes vs logistic regression: when would you choose each?
  • Both are linear in log-odds. Naive Bayes is generative: it estimates class-conditional feature distributions, assumes independence, and needs little data.
  • Logistic regression is discriminative: it fits the weights directly, handles correlated features, and is usually better calibrated and more accurate with enough data.
  • Ng and Jordan (2002): naive Bayes reaches its (higher) asymptotic error faster. It's a good baseline and a good fit for tiny data.
An A/B test shows a 2% lift with p = 0.04. What's the probability the variant is actually better?
  • Not 96%. A p-value is \(P(\text{data this extreme}\mid H_0)\), not \(P(H_0\mid\text{data})\), which is the same inversion as recall vs precision.
  • The probability that a "significant" result is real depends on the prior share of ideas that work and on power. With low base rates and low power, many wins are false.
  • A Bayesian analysis with an informed prior gives that probability directly. Lesson 7 covers tests and peeking.

Sources: He et al., Practical Lessons from Predicting Clicks on Ads at Facebook (ADKDD 2014), for the downsampling recalibration; Ng & Jordan, On Discriminative vs. Generative Classifiers (NeurIPS 2002).