JA EN
Home › Math for AI

∫ FIELD

Math for AI

Only the math you actually need to read AI papers, starting from what the symbols mean: linear algebra, calculus, probability, information theory, optimization.

0chapters 16Foundations 11Paper walkthroughs 0Interactive

Work through a volume in order: textbook → foundations → papers → lab.

② Foundations

Articles that assume nothing and build the ideas of the field, in order.

Linear Algebra

Vectors, matrices, eigenvalues, SVD — what a dimension really is

  1. Linear Algebra for AI — What Vectors and Matrices Are Actually Doing FREE
  2. The Linear Algebra Under LoRA and RAG — Eigenvalues, Low Rank and Vector Search, Hands On FREE
  3. A Tour of Matrix Decompositions — When to Reach for LU, QR, Cholesky, or SVD ★ MEMBER
  4. Tensors and Shape Manipulation — If You Can Read einsum, You Can Read Papers FREE

Calculus & Optimization

Partial derivatives, gradients, the chain rule, convexity, Lagrange

  1. Calculus for AI — The Gradient Is an Arrow Saying Which Way Is Better FREE
  2. Matrix Calculus from Scratch — Derive the Backward Pass Yourself ★ MEMBER
  3. Jacobians and Hessians — Multivariable Calculus, Drawn ★ MEMBER
  4. Calculus of Variations — What It Means to Differentiate a Function ★ MEMBER

Probability & Statistics

Distributions, expectation, Bayes, MLE, sampling

  1. Probability and Statistics for AI — A Model's Output Is a Distribution ★ MEMBER
  2. Thinking Bayesian — A Working Feel for Priors, Likelihoods, and Posteriors ★ MEMBER
  3. Markov Chains from Scratch — The Process That Only Looks at Now ★ MEMBER
  4. A Field Guide to Probability Distributions — Where Normal, Poisson, and the Exponential Family Come From FREE
  5. Monte Carlo Methods from Scratch — Solving Integrals with Dice ★ MEMBER
  6. Hypothesis Testing and A/B Tests — How to Use a p-value, and How People Misuse It ★ MEMBER

Information Theory

Entropy, KL divergence, cross-entropy — where loss functions come from

  1. Information Theory and AI — Where Cross-Entropy Loss Came From ★ MEMBER
  2. Entropy and Cross-Entropy — Where the Loss Function Comes From FREE

③ Paper walkthroughs

Written from the papers themselves. Every piece links the paper page and its PDF.

  1. Compression Is Prediction Is Intelligence — LLMs Through Information Theory ★ MEMBER arXiv:2309.10668
  2. Symmetry and Equivariance — How Group Theory Shapes Network Design ★ MEMBER arXiv:1602.07576
  3. Kernel Methods and Gaussian Processes — The Champions Before Neural Nets ★ MEMBER arXiv:1206.2944
  4. Optimal Transport — The Mathematics of Moving Distributions ★ MEMBER arXiv:1306.0895
  5. Statistical Learning Theory — Why Does Learning Generalize? ★ MEMBER arXiv:1611.03530
  6. Mutual Information — Putting a Number on What You Know ★ MEMBER arXiv:1807.03748
  7. Beyond SGD — Adam, Second-Order Methods, and Constrained Optimization ★ MEMBER arXiv:1412.6980
  8. Bayes' Theorem in AI — Priors, Posteriors, and Uncertainty ★ MEMBER arXiv:1706.04599
  9. Convexity and Optimization — Why Deep Learning Works Even Though It Isn't Convex ★ MEMBER arXiv:1406.2572
  10. KL Divergence From Scratch — Measuring the Gap Between Two Distributions ★ MEMBER arXiv:1312.6114
  11. Singular Value Decomposition and Low-Rank Approximation — the Math Behind LoRA ★ MEMBER arXiv:2106.09685