Learn
AI
From the basics of learning to CNNs, Transformers, VLMs and agents — following why each architecture took the shape it did.
Machine Learning Basics9What learning is: loss, overfitting, evaluation — the foundation for everything
- Step 1 What Machine Learning Really Is — Understanding “Learning” Without the Math
- Step 2 Loss Functions and Optimization — How a Model Learns From Being Wrong
- Step 3 Overfitting and Evaluation Design — Be Suspicious of 99% Accuracy
- Step 4 Decision Trees and Gradient Boosting — Still the Champion on Tabular Data
- Step 5 Unsupervised Learning from Scratch — Clustering and Dimensionality Reduction
- Step 6 Self-Supervised Learning — The Day Unlabeled Data Became an Asset
- Step 7 Data Leakage and Experiment Hygiene — When the Score Is Too Good, Suspect It
- Step 8 Imbalanced Data in Practice — What to Optimize When 99% Is Normal
- Step 9 ML System Design — The 90% Outside the Model
Deep Learning Basics7Neural nets, backprop, optimization, regularization
- Step 10 Neural Networks from Scratch — From One Neuron to Many Layers
- Step 11 Backpropagation from Scratch — It Is All Just the Chain Rule
- Step 12 A History of Normalization Layers — From BatchNorm to RMSNorm
- Step 13 Activation Functions from Scratch — Why Nonlinearity Is Non-Negotiable
- Step 14 Weight Initialization and Regularization — What Lets Training Start, and What Keeps It Going
- Step 15 Hyperparameter Search — Hunches, Grids, and Bayesian Optimization
- Step 16 Graph Neural Networks from Scratch — Learning from Connections
CNNs & Image Recognition9From the convolution to ResNet, efficiency, and detection
- Step 17 Image Classification from Scratch — The Invention of the Convolution
- Step 18 Object Detection from Scratch (from YOLO to DETR)
- Step 19 The CNN Family Tree — From AlexNet to ResNet and EfficientNet
- Step 20 The Autonomous Driving Perception Stack from Scratch — What Cameras and LiDAR Each Bring to the Table
- Step 21 Anomaly Detection from Scratch — Learning From Normal Alone
- Step 22 Medical Imaging AI — Validation Design Comes Before the Accuracy Number
- Step 23 The ImageNet Moment — The Day Deep Learning Won
- Step 24 Paper Walkthrough: TurboVLA — Kick the LLM Out of the Loop and Run a Robot Policy at 32 Hz on an RTX 4090 with Under 1 GB of VRAM
- Step 25 Paper Walkthrough: PhiZero — A World Model That Reasons in a Language of Physics Before It Renders
How Transformers Work13Attention, positional encoding, and the architecture dissected
- Step 26 Attention from Scratch — The Heart of the Transformer, Explained Visually
- Step 27 Positional Encoding from Scratch — From Absolute Positions to RoPE
- Step 28 Tokenizers from Scratch — The Unit an LLM Cuts the World Into
- Step 29 A Field Guide to Attention Variants — MQA, GQA, Sliding Windows, Linear Attention
- Step 30 The Transformer, End to End — One Token's Journey from Embedding to Output
- Step 31 FlashAttention from Scratch — The Paradox of Doing More Math to Go Faster
- Step 32 How Long-Context LLMs Work — From RoPE Interpolation to Ring Attention
- Step 33 Encoder or Decoder — The Fork in the Road Between BERT and GPT
- Step 34 Build Your Own BPE Tokenizer — Learning Merge Rules, and Getting Punished by Japanese
- Step 35 Build Your Own Mini GPT — A Language Model in 300 Lines
- Step 36 Paper Walkthrough: Designing Qwen3.8-Next — Accuracy, Efficiency and Stability as One Problem
- Step 37 Paper Explained: Why Gated DeltaNet Survives 4-Bit Quantization — NVFP4 W4A4 in a Hybrid 27B
- Step 38 Paper Walkthrough: Stop Anchoring to Frame One — Scal3R's Multi-Reference Relative Pose Query
Large Language Models14Scaling laws, prompting, and how LLMs behave inside
- Step 39 Scaling Laws from Scratch — Why Making Models Bigger Makes Them Smarter (and When It Doesn't)
- Step 40 The Science of Prompt Engineering — What Is Proven and What Is Folklore
- Step 41 Mamba and State Space Models — Handling Sequences Without Attention
- Step 42 Building a Pretraining Corpus — From Web Sludge to Textbook Quality
- Step 43 LLM Evaluation from Scratch — Reading Benchmarks and the Contamination Problem
- Step 44 Why Language Models Hallucinate — The Mechanics and What Actually Helps
- Step 45 Test-Time Scaling — How Models Get Better by Thinking Longer
- Step 46 Knowledge Distillation from Scratch — Copying a Big Model into a Small One
- Step 47 Alignment, Explained — From RLHF to Constitutional AI
- Step 48 Scaling Skepticism — A Genealogy of the "Just Make It Bigger" Critique
- Step 49 Paper Walkthrough: Can Anything Catch a Fake Crisis Video? — What RA-Bench Found
- Step 50 Paper Walkthrough: SA-MRPO — Stop Studying the Subject You've Already Aced
- Step 51 Paper Explained: Agentic Artifact Creation — Where Generation Ends and Construction Begins
- Step 52 Paper walkthrough: Puro-2B — pretraining a 2B model from scratch for $6.9K on consumer GPUs
VLMs & Multimodal6CLIP, ViT, and putting images and language in one space
- Step 53 Paper Deep Dive — ViT: Treating an Image Like a Sentence
- Step 54 Paper Deep Dive — CLIP: Putting Words and Images on One Map
- Step 55 How VLMs Came Together — Wiring a Vision Encoder into an LLM
- Step 56 BEV Representations From Scratch — Fusing Multiple Cameras Into One Top-Down Map
- Step 57 Document AI and OCR Today — How an LLM Ends Up Reading Your Invoices
- Step 58 Video Understanding from Scratch — From a Pile of Frames to a Sense of Time
Generative Models10Diffusion, VAE, GAN — models that create
- Step 59 Diffusion Models from the Ground Up — Add Noise, Then Subtract It
- Step 60 VAEs from Scratch — Stir Probability into "Compress and Restore" and You Get a Generator
- Step 61 The Rise and Fall of GANs — An Invention Trained by Rivalry, and Why Diffusion Won
- Step 62 Flow Matching from Scratch — What Came After Diffusion, and Why It Goes Straight
- Step 63 CFG and Samplers — What the "Strength" Knob in Generative AI Really Does
- Step 64 The Mathematics of Diffusion — Generation Seen Through Scores and SDEs
- Step 65 A Practical Map of Image Generation — SD, ControlNet, and Applying LoRA
- Step 66 Music and Audio Generation from Scratch — Sound as Tokens
- Step 67 The State of 3D Generation — From NeRF to Gaussian Splatting
- Step 68 Build Your Own Diffusion Model — Starting from MNIST
Training & Alignment13Pretraining, SFT, RLHF/DPO, fine-tuning
- Step 69 Instruction Tuning and RLHF from Scratch — How a Model Learns to Follow Orders
- Step 70 Learning Rate Schedules — Why Warmup and Why Cosine
- Step 71 Mixed Precision Training — Going Faster in fp16/bf16/fp8 Without Breaking
- Step 72 Building a Dataset in Practice — Collect, Clean, Blend
- Step 73 DPO and What Came After — The Lineage That Simplified RLHF
- Step 74 Continual Learning and Catastrophic Forgetting — Why Models Can't Just Keep Learning
- Step 75 Diagnosing Broken Training — Telling Divergence, NaN, and Plateaus Apart
- Step 76 Versioning Data and Models — An Experiment You Cannot Reproduce Never Happened
- Step 77 Paper Deep-Dive: ABSeeker — Training Long-Horizon Search Agents by Grading Each Step Backward from the Answer
- Step 78 PAWBench Explained — Can Video Generators Get the Odds Right, Not Just the Physics?
- Step 79 Paper Walkthrough: PaperGym — Turning One Paper Into a Graded Training Environment for Research Plans
- Step 80 Paper Walkthrough: StudentSim — Training a Simulator That Is Actually *That* Student
- Step 81 Paper Walkthrough: It Takes Two to Match — Co-Evolving Both Sides of Retrieval with RL
- +1 planned
Inference & Serving29KV cache, quantization, batching, serving
- Step 82 The KV Cache from Scratch — The Heart of Fast Inference
- Step 83 LLM Quantization from Scratch — Why Losing Precision Doesn't Break It
- Step 84 Speculative Decoding from Scratch — How a Tiny Draft Model Speeds Up an LLM Without Changing a Single Output
- Step 85 LLM Serving from Scratch — vLLM, Continuous Batching, and Not Letting the GPU Idle
- Step 86 Structured Output and Constrained Decoding — How to Stop an LLM from Breaking Your JSON
- Step 87 Surviving GPU Out-of-Memory — Every Cause, Every Fix
- Step 88 Cutting Inference Cost in Practice — What to Do First
- Step 89 Prompt Caching and Context Design — One Prefix Rule That Moves Your Bill by an Order of Magnitude
- Step 90 Paper Walkthrough: The Personalization Mirage — LLMs Invent a Version of You, and Their Self-Reports Point the Wrong Way
- Step 91 Paper Walkthrough: DAPD — Breaking the Teacher's "Cheat-Sheet Illusion" in Distillation with Dual Anchors
- Step 92 Paper Walkthrough: CodeNib — A Multi-View Data System That Serves Repository Context to Coding Agents
- Step 93 Paper explained: BDH-CQ — an AI that thinks without words. Recurrent memory plus latent reasoning resets ARC's cost frontier
- Step 94 Paper Deep-Dive: AgentOPSD — Finding the Turn That Won the Game with Recursive Bayesian Belief Updates
- Step 95 Paper Walkthrough: No Gold Answers, No Stronger Teacher — How u-OPSD Distills From Its Own Majority Vote
- Step 96 Paper Explained: Agentic ESOpt — Drop Backprop, Jiggle the Weights, and Train Long-Horizon LLM Agents
- Step 97 Paper Walkthrough — WarpSAC: When RL's Safety Rails Become Handcuffs
- Step 98 TTPO Explained: Training a Model Mid-Exam, With No Answer Key
- Step 99 Paper explained: Self-OPD — an image generator that distills itself, with no teacher
- Step 100 Paper walkthrough: CyberFactory — turning wild CVEs into runnable training problems
- Step 101 Paper Walkthrough: DART-SD — Training Tool-Calling Agents Without Flattening the Diamond
- Step 102 Paper Walkthrough — SMELT: Is Looping the Same Layers Twice Actually a Win When the Budget Is Matched?
- Step 103 Paper Walkthrough: Normalized Low-Rank Adaptation — Why Normalizing LoRA's Entry Matrix Works
- Step 104 Paper Walkthrough: From Production Traffic to Post-Training — Folding 200 Internal Apps Into One Self-Hosted LLM
- Step 105 Paper Explained: Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement
- Step 106 Paper Explained — LatentPress: Feeding Compressed Context Straight to a Frozen LLM, Neither as Text Nor as Pixels
- Step 107 Paper Walkthrough: Language Models Can Control Their Own Attention
- Step 108 Paper Explained: Compile by Training — Turning a Natural-Language Spec into a Function That Runs Locally
- Step 109 Paper Walkthrough: One Training Example Keeps On-Policy Distillation Improving for Hundreds of Steps
- Step 110 Paper Walkthrough: Random Attention — Throwing KV Cache Entries Away at Random Works Just as Well
- +1 planned
RAG & Retrieval10Embeddings, chunking, reranking, evaluation
- Step 111 RAG Fundamentals and Design Patterns — Embeddings, Chunking, Reranking, and Evaluation from Scratch
- Step 112 RAG vs Fine-Tuning — Which One, and When
- Step 113 Embeddings from Scratch — from word2vec Intuition to Contextual Embeddings
- Step 114 Recommenders and Embeddings — Same Math as RAG, Different Goal
- Step 115 Chunking Strategies — How You Split Decides What You Can Find
- Step 116 GraphRAG from Scratch — Where Knowledge Graphs Meet Retrieval
- Step 117 Evaluating RAG in Practice — Turning “Seems Better” Into a Number
- Step 118 Build Your Own Vector DB — From Brute Force to HNSW
- Step 119 Paper Walkthrough: WeMM-Embedding — Putting Text, Images and Video on One Ruler
- Step 120 Paper walkthrough: Hi-Q — splitting a question down to the granularity your corpus can actually retrieve
Agents55Tool use, planning, multi-agent systems
- Step 121 LLM Agents from Scratch — Designing the Tool-Use Loop
- Step 122 End-to-End Driving from Scratch — Perception to Control in a Single Network
- Step 123 MCP and Tool Protocols — The Standard That Connects an Agent's Hands
- Step 124 Multi-Agent Design Patterns — Division, Debate, Verification
- Step 125 Designing Agent Memory — Short-Term, Long-Term, Episodic
- Step 126 Evaluating Agents — How Benchmarks and Harnesses Are Built
- Step 127 AlphaGo from Scratch — The Marriage of Search and Learning
- Step 128 Build Your Own Agent Loop — The Minimal Shape of Tool Calling
- Step 129 Paper Walkthrough: SWE-Bench ProMax — Measuring What Coding Agents Can Really Do with Large-Scale, Multilingual Refactoring
- Step 130 Paper Walkthrough: Macaron-V1 — A Frozen Base plus a Mixture of LoRAs, Built to Keep Learning After Launch
- Step 131 Paper Walkthrough: EnvACE — Agents That Rehearse the World Instead of Calling It
- Step 132 Paper Explained: Video-DeepResearch — Agents That Watch a Video, Then Chase Down Every Lead
- Step 133 Paper Walkthrough: ToolArtist — Search, Draw, or Redraw? The Image Agent That Decides for Itself
- Step 134 Paper Deep-Dive: Recursive Synthesis — Extending Verified Tasks Into 40,000 Long-Horizon Terminal Problems
- Step 135 Paper Walkthrough: Qwen-UI-Agent — How Alibaba Built a GUI Agent That Works on Real Phones and PCs
- Step 136 Paper Explained: OSReward — Can You Trust the AI That Grades AI? Remeasuring Rewards for Computer-Use Agents
- Step 137 Paper Walkthrough: Metis — A 'Memory Foundation Model' That Moves Agent Memory Inside the Model
- Step 138 Paper Walkthrough: MerchantBench — Can an LLM Agent Run an Online Store for a Year? Why It Earns Only 27.3% of What Humans Do
- Step 139 Paper Walkthrough: Mental World Modeling — A World Model That Advances Minds, Not Just Physics
- Step 140 Paper Explained: LongHorizon-Harness — Long-Horizon Agent Tasks Are a State-Management Problem, Not an Execution Problem
- Step 141 Paper Deep-Dive: Frontis-MA1 — Training the AI That Builds AI: One Step Toward Recursive Self-Improvement in ML Engineering
- Step 142 Paper Walkthrough: ComBodied Agents — Moving an Agent's Target from Software and Matter to the Person
- Step 143 Paper Explained: Co-Evolution in Agentic Systems — Three Stages Toward Self-Directed Evolution
- Step 144 Paper walkthrough: Zetta ζ — a robot harness that repairs itself mid-execution, with the policy frozen
- Step 145 Paper walkthrough: StateM — 95.3% on Terminal-Bench 2.1 and a USD 15 run, without touching a single weight
- Step 146 Paper Explainer: SemaPLC — The Agent That Isn't Allowed to Say "Done"
- Step 147 Paper Walkthrough: OmniScientist — An AI Scientist That Actually Looks at the Raw Data
- Step 148 Paper Walkthrough — FACET: Grounding Instruction, Environment, Solution and Verifier in One Executable State
- Step 149 Paper Explained: EnvHarness — Reshaping an Agent's Training World Without Rebuilding It
- Step 150 Paper Explainer: Why Agent Skills Work — and Where They Break
- Step 151 Paper Explained: Co-RL — Reasoning Without Labels, Emerging From a Diverse Cohort
- Step 152 Paper Walkthrough: SWE-bench Science — Can Coding Agents Fix Scientific Code?
- Step 153 Paper Explained: FreeToken — Treating Your Own PC as a Single Elastic Inference Platform
- Step 154 Paper Walkthrough: Embodied-Navigator (TAMP-Nav) — Let the VLM Just Point, and Navigation Gets Both Faster and Better
- Step 155 Paper walkthrough: ASI-Bench — peeling away human guidance to measure what AI can do alone
- Step 156 Paper Walkthrough: FrontierChallenge — Grading Scientific Work on Whether It Was Actually Delivered
- Step 157 Paper Explained: JIT-Agent — A Model That Writes the Agent Harness On Demand
- Step 158 Paper Walkthrough: ZimaBlue — Turning 120,000 Hours of Egocentric Video into Robot Skill
- Step 159 Paper Explained: What Makes Good Agentic Data? The ACE Lens
- Step 160 Paper Walkthrough — UrbanGround: Where MLLM Agents Break Down on a Real Street
- Step 161 Paper Walkthrough: UI-Venus-2 — Taking Screen-Operating Agents From Benchmarks to Real Work
- Step 162 Paper Walkthrough: Training Agents to Evolve with Their Harness
- Step 163 Paper Explained: StarHarness — Evolving the Scaffold Instead of the Weights
- Step 164 Paper Walkthrough: SecOPD — Grading One Token at a Time to Cut Adaptive Prompt Injection by an Order of Magnitude
- Step 165 Paper Explained: Repo-To-Skill — Distilling GitHub Repositories Into Skills an AI Can Use
- Step 166 Paper Walkthrough: PILOT in the Loop — Fixing the Run While It Is Still Running
- Step 167 Paper Walkthrough: LoopArena — Benchmarking the Model That Steers a Coding Agent
- Step 168 Paper Walkthrough: Code World Model — Putting a Coding Agent in Charge of the World
- Step 169 Paper Walkthrough: Code as Worlds — An Agent That Writes the World Down as Runnable Code
- Step 170 Paper Walkthrough: AutoSaddler — Growing a Harness That Doesn't Break, from Agent Failure Logs
- Step 171 Paper Walkthrough: HarnessDev — Can an LLM Build and Maintain the System It Runs Inside?
- Step 172 Paper Walkthrough: EarlyEval — Making Agent Evaluation Cheaper by Stopping Early
- Step 173 Paper Walkthrough: Aspire — Can Models Self-Evolve from Vague Goals?
- Step 174 Paper Walkthrough: Terminal-Universe — Turning Agent Logs Back Into Reusable Execution Environments
- Step 175 CogEvol: What the Reward Cannot Measure, RL Will Quietly Destroy
- +6 planned
Audio & Speech9ASR, TTS, speech LLMs
- Step 176 Paper Deep-Dive: Why Whisper Is Robust — Large-Scale Weak Supervision
- Step 177 Speech Recognition from Scratch — From Waveform to Text
- Step 178 Speech Synthesis from Scratch — From Text to a Voice
- Step 179 Representing Sound — Mel Spectrograms and Audio Tokens
- Step 180 Paper Walkthrough: SwanTale — Designing Voices from Words Alone, with Speech and Sound in One Waveform
- Step 181 Paper Walkthrough: Interpretable MEG Decoding of Perceived Speech — Reading the Decoder's Weights as a Brain Map
- Step 182 Paper Walkthrough: VoiceMem — A Left Brain and a Right Brain for Voice Agents, at Zero Added Latency
- Step 183 Paper Walkthrough: Motion-Omni — Speaking and Moving in One Forward Pass
- Step 184 Paper Explained: Last Translation Benchmark — Measuring Translation with Breaking Examples and Verification Rules
Time Series5From RNN/LSTM to time-series foundation models
- Step 185 RNNs and LSTMs from Scratch — Why Learn Them in the Transformer Era
- Step 186 Time-Series Forecasting from Scratch — From Classical Methods to Foundation Models
- Step 187 Do Transformers Actually Work on Time Series? — The Argument and the Practical Answer
- Step 188 Time-Series Anomaly Detection — The Math Behind the Alerts
- Step 189 H3-World, Explained — Turning Language Understanding into World Control
Model Families7Gemma, Llama, Qwen and friends — lineage and when to use which
- Step 190 The Gemma Family from Scratch — Lineage, Inventions, and Where It Fits
- Step 191 The Llama Family from Scratch — The Main Line of Open LLMs
- Step 192 The Qwen Family from Scratch — Why It Tops Hugging Face's Download Charts
- Step 193 The DeepSeek Family from Scratch — Breaking In with MoE and Distillation
- Step 194 The GPT Lineage — Design Thinking from GPT-1 to Today
- Step 195 The Mistral Family from Scratch — Europe's Small-and-Strong Bet
- Step 196 Open vs Closed — The Economics of Releasing Weights, and the Safety Argument
Paper Deep-Dives13Landmark papers one by one, grounded in the source text
- Step 197 Paper Deep Dive — LoRA: Low-Rank Adaptation of Large Language Models: Why Low Rank Is Enough
- Step 198 When Proxies Stop Being Good Enough — Reading August 2026's Eight Autonomous Driving Papers Together
- Step 199 Paper Deep Dive — Attention Is All You Need: What Dropping Recurrence Actually Proved
- Step 200 Mixture of Experts (MoE) from Scratch — Routing and Load Balancing in the Switch Transformer
- Step 201 Paper Walkthrough: Alpamayo — NVIDIA's Reasoning Model for Autonomous Driving
- Step 202 Symbolic vs. Connectionist — Where a 60-Year Argument Stands Today
- Step 203 Paper Walkthrough: WorldClaw — Agents That Build Walkable, Editable 3D Open Worlds from a Single Sentence
- Step 204 Paper Deep Dive: AskChem — Changing the Unit of Search from Papers to Provenance-Carrying Claims
- Step 205 Paper Deep Dive — Large Discovery Models: giving an LLM a value signal for what to try next
- Step 206 Paper walkthrough: Apodex 1.1 — scaling agents around completed work
- Step 207 Paper Walkthrough: Turning Game Development into a Verifiable Trajectory Data Engine — RLHEV and AWoMo
- Step 208 Paper Walkthrough — J-Zero: Growing the Challenger, the Solver, and the Judge Together from Zero Data
- Step 209 Paper Walkthrough — RoboTok: Mining the Web for Demonstrations That Move Like Yours
Distillation & Compression9Moving capability from a large model into a small one — losses, data design and failure modes
- Step 210 The Math of Distillation — Why Soft Answers Teach More
- Step 211 A Field Guide to Distillation Recipes — logit, feature, attention, self
- Step 212 Designing Distillation Data — Deciding What to Ask the Teacher
- Step 213 When Distillation Fails — Capacity Gaps and Contagious Overconfidence
- Step 214 On-Policy Distillation — Learning From What the Student Actually Writes
- Step 215 Distillation vs. Quantization vs. Pruning — Three Roads to a Smaller Model
- Step 216 Distilling Agents — How to Compress a Long Trajectory
- Step 217 Evaluating Distilled Models — Is "Close to the Teacher" a Good Metric?
- Step 218 Build Your Own Distillation — Growing a Small Model in 100 Lines
Evaluation & Judging5How to measure AI — benchmarks, LLM judges, contamination and reward hacking
- Step 219 LLM-as-a-Judge from Scratch — How AI Grades AI, and Where It Breaks
- Step 220 A Map of Agent Benchmarks — What SWE-bench, GAIA, and OSWorld Actually Measure
- Step 221 Self-Improving AI — Self-Play, Co-Evolution, and Generated Curricula
- Step 222 Reward Hacking — Whatever You Measure Is Where It Breaks
- Step 223 Benchmark Contamination — How to Doubt a High Score
Semiconductors
How silicon pulled from sand comes to compute — from the physics of matter to arithmetic units, memory, power and fabrication.
Device Physics5Bands, doping, pn junctions, MOSFET, CMOS
- Step 224 MOSFETs from the Ground Up — A Sluice Gate Opened by Voltage, and the Reality of Leakage
- Step 225 Reading Chip Design as a Power Budget — The Physics of Leakage and Heat
- Step 226 Band Theory from the Ground Up — Why It Had to Be Silicon
- Step 227 From FinFET to GAA — Why the Transistor Had to Go Vertical
- Step 228 The Physics of NAND Flash — Remembering by Trapping Electrons
Computer Architecture5Logic, clocking, the memory hierarchy and the memory wall
- Step 229 The Memory Wall from Scratch — Why Moving Data Costs More Than Computing
- Step 230 The GPU Memory Hierarchy — HBM, SRAM, Registers, and Why Movement Wins
- Step 231 CPU Pipelines and Branch Prediction — The Factory Inside One Clock Tick
- Step 232 Systolic Arrays — Building the Heart of the TPU From Scratch
- Step 233 Interconnects — How NVLink, PCIe, and Light Set the Limits of Scale
Scaling & Power4The end of Dennard scaling, the physics of power, the cost of moving data
Fabrication & Packaging4Wafers, lithography, yield, chiplets, 3D integration
Accelerators3Number formats, accelerator design styles, edge deployment
Supply Chain6Materials, equipment, EDA and IP — who holds what upstream of the chip itself
- Step 245 Mapping the Semiconductor Supply Chain — From Sand to Chip, Who Holds What
- Step 246 Upstream of Semiconductors — Wafers, Photoresist, and Specialty Gases
- Step 247 Inside the Equipment Makers — What ASML, AMAT, TEL and Lam Actually Build
- Step 248 EDA Tools from Scratch — Chips Are Written in Software
- Step 249 IP Cores and the Fabless Model — How Arm Rules Silicon Without Making a Single Chip
- Step 250 The Geopolitics of Chips — Export Controls and Supply Chain Rewiring, Explained Technically
Algorithms
From complexity as a yardstick to search, dynamic programming, matrix computation and parallelism — the craft of computation underneath AI.
Complexity5Big-O, time-space tradeoffs, and where theory parts from the profiler
- Step 251 Complexity From Scratch — What Big-O Actually Measures
- Step 252 When Big-O and Your Benchmarks Disagree — Caches, Branches, and Memory Bandwidth
- Step 253 Randomized Algorithms — Why Rolling Dice Makes Things Faster
- Step 254 NP-Completeness from Scratch — Not Unsolvable, but Fast to Verify
- Step 255 Approximation Algorithms — Trading Exactness for a Guarantee
Data Structures5Arrays, trees, hashes, heaps — what each choice buys you
- Step 256 Choosing a Data Structure — Arrays, Hashes, Trees and Heaps
- Step 257 Hashing and Nearest-Neighbor Search — The Groundwork Under Vector Search
- Step 258 Cache-Friendly Code — Why Two O(n) Loops Can Differ by 10×
- Step 259 B-Trees and LSM-Trees — The Heart of Every Database
- Step 260 Probabilistic Data Structures — Counting Without Counting
Search & Optimization4Exhaustive, greedy, DP, branch and bound, approximation
- Step 261 Dynamic Programming From Scratch — On Remembering Subproblems
- Step 262 Graph Algorithms from Scratch — Shortest Paths and Where They Lead
- Step 263 Linear Programming from Scratch — The Workhorse of Optimization
- Step 264 Simulated Annealing and Genetic Algorithms — What to Do When Exact Solving Breaks Down
Numerical Computing6GEMM, decompositions, FFT, iterative methods — where AI's compute actually goes
- Step 265 The Cost of Matrix Multiplication — Where Almost All of AI's Compute Goes
- Step 266 The FFT from Scratch — Why Convolution Turns into Multiplication
- Step 267 Numerical Pitfalls — Cancellation, Rounding, and logsumexp
- Step 268 How Autodiff Actually Works — Unpacking the PyTorch Magic
- Step 269 Solving Systems of Equations — Direct Methods and Iterative Methods
- Step 270 Build Your Own Autograd — A Mini PyTorch in 100 Lines
Parallel & Distributed3Limits of parallelism, the GPU execution model, communication in distributed training
Math for AI
Only the math you actually need to read AI papers, starting from what the symbols mean: linear algebra, calculus, probability, information theory, optimization.
Linear Algebra6Vectors, matrices, eigenvalues, SVD — what a dimension really is
- Step 274 Linear Algebra for AI — What Vectors and Matrices Are Actually Doing
- Step 275 The Linear Algebra Under LoRA and RAG — Eigenvalues, Low Rank and Vector Search, Hands On
- Step 276 Singular Value Decomposition and Low-Rank Approximation — the Math Behind LoRA
- Step 277 A Tour of Matrix Decompositions — When to Reach for LU, QR, Cholesky, or SVD
- Step 278 Tensors and Shape Manipulation — If You Can Read einsum, You Can Read Papers
- Step 279 Symmetry and Equivariance — How Group Theory Shapes Network Design
Calculus & Optimization6Partial derivatives, gradients, the chain rule, convexity, Lagrange
- Step 280 Calculus for AI — The Gradient Is an Arrow Saying Which Way Is Better
- Step 281 Convexity and Optimization — Why Deep Learning Works Even Though It Isn't Convex
- Step 282 Beyond SGD — Adam, Second-Order Methods, and Constrained Optimization
- Step 283 Matrix Calculus from Scratch — Derive the Backward Pass Yourself
- Step 284 Jacobians and Hessians — Multivariable Calculus, Drawn
- Step 285 Calculus of Variations — What It Means to Differentiate a Function
Probability & Statistics10Distributions, expectation, Bayes, MLE, sampling
- Step 286 Probability and Statistics for AI — A Model's Output Is a Distribution
- Step 287 Bayes' Theorem in AI — Priors, Posteriors, and Uncertainty
- Step 288 Thinking Bayesian — A Working Feel for Priors, Likelihoods, and Posteriors
- Step 289 Markov Chains from Scratch — The Process That Only Looks at Now
- Step 290 A Field Guide to Probability Distributions — Where Normal, Poisson, and the Exponential Family Come From
- Step 291 Monte Carlo Methods from Scratch — Solving Integrals with Dice
- Step 292 Hypothesis Testing and A/B Tests — How to Use a p-value, and How People Misuse It
- Step 293 Statistical Learning Theory — Why Does Learning Generalize?
- Step 294 Optimal Transport — The Mathematics of Moving Distributions
- Step 295 Kernel Methods and Gaussian Processes — The Champions Before Neural Nets
Information Theory5Entropy, KL divergence, cross-entropy — where loss functions come from
- Step 296 Information Theory and AI — Where Cross-Entropy Loss Came From
- Step 297 KL Divergence From Scratch — Measuring the Gap Between Two Distributions
- Step 298 Entropy and Cross-Entropy — Where the Loss Function Comes From
- Step 299 Mutual Information — Putting a Number on What You Know
- Step 300 Compression Is Prediction Is Intelligence — LLMs Through Information Theory
Media & Compression
Why JPEG degrades and PNG does not — image, video and audio coding from both the information theory and the settings you actually ship.
Coding Theory3Lossless vs lossy, entropy coding, Huffman, arithmetic coding
Image Codecs3JPEG, PNG, WebP, AVIF — the DCT, quantization, and lossless pipelines
Video Codecs4MPEG, H.264, AV1 — motion compensation, GOP structure, rate control
Audio Codecs4MP3, AAC, Opus — perceptual masking as a way of discarding
Media in Production3Bitrate ladders, encoder settings, quality metrics — what the job actually asks
Systems
The floor AI runs on: operating systems, networks, databases, language runtimes, security and cloud — understood from their mechanisms.
OS & Runtime3Processes, memory, containers — the ground programs stand on
Networking3TCP/QUIC, DNS, CDN — how data actually arrives
Databases3Internals, transactions, distribution — why data survives
Compilers & Runtimes3Compilers, JITs, GC — why code runs fast
Security4Crypto, auth, LLM attacks — designing for defence
Cloud & Ops3Serverless, cost, observability — shipping without breaking
Engineering Process5Requirements, design, estimation — the upstream where projects are won or lost
- Step 337 Upstream Engineering from Scratch — Why Projects Are Won or Lost at Requirements
- Step 338 The Craft of Requirements — Why "We Built Exactly What They Asked For" Fails
- Step 339 Architecture Decisions — Telling Apart What You Can Undo From What You Cannot
- Step 340 Estimation and Scope — The Cone of Uncertainty and How to Negotiate
- Step 341 Upstream Work in the AI Era — When Code Gets Cheap, What Gets Valuable?