Machine learning, derived
Each lesson starts from one concrete question, builds the answer on a small example you can change, asks you to predict before it tells you, and ends with code you write yourself.
-
01
Cross-entropy and KL, from surprise
Why −log p, why the best model copies the frequencies, entropy as the floor under training loss, KL as the gap, the softmax gradient q − p, forward vs reverse KL. Game: bet on the next word.
-
02
From a reward to a gradient: REINFORCE, PPO, GRPO
The log-derivative trick, baselines and advantages, PPO's ratio and clipping, GRPO's group-relative advantages, per-token credit, length bias and the KL leash. Game: predict the push.
-
03
Backpropagation, by assigning blame
Real numbers forward into a loss, then blame traced back through every weight and ReLU gate, an update, and a gradient check against nudging. Why sigmoid pairs with cross-entropy. Game: the blame game.
-
04
Attention: score, normalize, mix
Dot products you can drag in 2-D, softmax as a temperature, outputs as convex mixes, why √d, causal masks, and the KV cache with a memory calculator. Game: be the query.
-
05
Bayes from counts
1,000 messages as dots: build the counts, let the evidence select the group, then name it. Likelihood vs posterior, recall vs precision, odds and log-odds, recalibrating downsampled models. Game: call the odds.
-
06
k-means, against Lloyd's algorithm
The objective, why nearest-center and mean updates, Lloyd as coordinate descent, local minima, k-means++ measured over 300 runs, feature scaling, empty clusters, and k-means inside ANN indexes. Game: place the centers.
-
07
A/B tests and the cost of peeking
p-values built from A/A tests, the standard error of a rate difference, power and sample size, sample-ratio mismatch, clustered units, multiple metrics, novelty, CUPED and sequential tests. Game: ship or not.
-
08
Optimizers: learning rate, momentum, Adam
The 2/λ stability limit, condition number, momentum's √κ speed-up, noise floors and learning-rate decay, Adam's bias correction and diagonal scaling, warmup, AdamW. Game: descend the valley.
-
09
Retrieval vs ranking
The recall ceiling, why two-tower models can't use cross features, in-batch negatives and the logQ correction, position bias, cold start, NDCG vs MRR vs recall, pointwise/pairwise/listwise losses. Game: budget the funnel.
-
10
PCA and SVD: what a projection keeps
Projection and the Pythagorean trade between kept variance and error, eigenvectors of the covariance, SVD and Eckart–Young, LoRA's rank bet, scree plots, scaling and outliers. Game: find the axis.
-
11
Leakage and the leaderboard that lies
What validation promises, prediction time and point-in-time features, target, temporal, group, preprocessing and selection leakage, bias vs variance, leakage vs drift. Game: validate, then deploy.