Course · Beginner
Probability & Statistics for Machine Learning
Uncertainty, likelihood and the maths of learning from noisy data.
Machine learning is learning from noisy data, and probability is the language of noise. This course turns the notation of the first course into working tools: Bayes' rule at scale, distributions and how to sample them, expectation and variance (and how they decide the way a network is initialised), the multivariate Gaussian behind diffusion models, the laws of large numbers, maximum likelihood, information theory, Bayesian inference and the statistics of honest evaluation. Each chamber runs the research loop, from discovery to derivation to real papers such as naive Bayes spam filters, Kaiming initialisation, DDPM and neural language models, with proofs to write and problems to code. Argus, the hundred-eyed watchman, waits at the centre.
- Chambers
- 9 + boss
- Time
- ~8 hours
- Questions
- 81
- Coding problems
- 27
- Concept cards
- 37
- XP available
- 4,960+
The path through the labyrinth
Three rings, 9 chambers, one guardian. You can wander ahead, but the thread works best in order: each chamber builds on the last.
The Outer Ring
The rules of chance- 1Bayes' Rule at Work: Odds and EvidenceYou are hereOdds, likelihood ratios and evidence that adds up, from the Monty Hall problem to a naive Bayes spam filter. 40 min 50 XP9 questions
- 2Distributions and How to Sample ThemPMFs, PDFs and CDFs, the named distributions of machine learning, and how to turn uniform noise into any distribution you like. 40 min 60 XP9 questions
- 3Expectation, Variance and InitialisationVariance of sums, covariance and correlation, Markov and Chebyshev, and the variance argument behind Xavier and Kaiming initialisation. 40 min 60 XP9 questions
The Middle Ring
From samples to models- 4The Multivariate GaussianCovariance matrices, ellipses of probability, sampling with Cholesky, and the fact about adding Gaussians that powers diffusion models. 45 min 60 XP9 questions
- 5Large Numbers and the Central Limit TheoremWhy averages settle, why they settle into a bell curve, and what that says about minibatches, error bars and Monte Carlo. 40 min 60 XP9 questions
- 6Maximum Likelihood: Fitting Models to DataChoose the parameters that make the data most probable, and find squared error and cross-entropy hiding inside. 45 min 60 XP9 questions
The Inner Ring
Inference and information- 7Entropy, Cross-Entropy and KL DivergenceMeasure surprise in bits: entropy, cross-entropy, KL divergence, mutual information and the evidence lower bound. 45 min 60 XP9 questions
- 8Bayesian Inference: Priors, Posteriors and BanditsTreat parameters as uncertain: conjugate priors, MAP estimates that turn out to be regularisation, and Thompson sampling. 45 min 60 XP9 questions
- 9Intervals, Tests and Honest EvaluationConfidence intervals, p-values, permutation tests, the bootstrap and multiple comparisons: how to tell if one model is really better. 45 min 60 XP9 questions