Home › Math for AI
∫ FIELD
Math for AI
Only the math you actually need to read AI papers, starting from what the symbols mean: linear algebra, calculus, probability, information theory, optimization.
Work through a volume in order: textbook → foundations → papers → lab.
② Foundations
Articles that assume nothing and build the ideas of the field, in order.
Linear Algebra
Vectors, matrices, eigenvalues, SVD — what a dimension really is
- Linear Algebra for AI — What Vectors and Matrices Are Actually Doing FREE
- The Linear Algebra Under LoRA and RAG — Eigenvalues, Low Rank and Vector Search, Hands On FREE
- A Tour of Matrix Decompositions — When to Reach for LU, QR, Cholesky, or SVD ★ MEMBER
- Tensors and Shape Manipulation — If You Can Read einsum, You Can Read Papers FREE
Calculus & Optimization
Partial derivatives, gradients, the chain rule, convexity, Lagrange
Probability & Statistics
Distributions, expectation, Bayes, MLE, sampling
- Probability and Statistics for AI — A Model's Output Is a Distribution ★ MEMBER
- Thinking Bayesian — A Working Feel for Priors, Likelihoods, and Posteriors ★ MEMBER
- Markov Chains from Scratch — The Process That Only Looks at Now ★ MEMBER
- A Field Guide to Probability Distributions — Where Normal, Poisson, and the Exponential Family Come From FREE
- Monte Carlo Methods from Scratch — Solving Integrals with Dice ★ MEMBER
- Hypothesis Testing and A/B Tests — How to Use a p-value, and How People Misuse It ★ MEMBER
Information Theory
Entropy, KL divergence, cross-entropy — where loss functions come from
③ Paper walkthroughs
Written from the papers themselves. Every piece links the paper page and its PDF.
- Compression Is Prediction Is Intelligence — LLMs Through Information Theory ★ MEMBER arXiv:2309.10668
- Symmetry and Equivariance — How Group Theory Shapes Network Design ★ MEMBER arXiv:1602.07576
- Kernel Methods and Gaussian Processes — The Champions Before Neural Nets ★ MEMBER arXiv:1206.2944
- Optimal Transport — The Mathematics of Moving Distributions ★ MEMBER arXiv:1306.0895
- Statistical Learning Theory — Why Does Learning Generalize? ★ MEMBER arXiv:1611.03530
- Mutual Information — Putting a Number on What You Know ★ MEMBER arXiv:1807.03748
- Beyond SGD — Adam, Second-Order Methods, and Constrained Optimization ★ MEMBER arXiv:1412.6980
- Bayes' Theorem in AI — Priors, Posteriors, and Uncertainty ★ MEMBER arXiv:1706.04599
- Convexity and Optimization — Why Deep Learning Works Even Though It Isn't Convex ★ MEMBER arXiv:1406.2572
- KL Divergence From Scratch — Measuring the Gap Between Two Distributions ★ MEMBER arXiv:1312.6114
- Singular Value Decomposition and Low-Rank Approximation — the Math Behind LoRA ★ MEMBER arXiv:2106.09685