JA EN

Learn

AI

From the basics of learning to CNNs, Transformers, VLMs and agents — following why each architecture took the shape it did.

Machine Learning Basics9What learning is: loss, overfitting, evaluation — the foundation for everything

Open this genre →

  1. Step 1 What Machine Learning Really Is — Understanding “Learning” Without the Math FREE6 min read
  2. Step 2 Loss Functions and Optimization — How a Model Learns From Being Wrong FREE8 min read
  3. Step 3 Overfitting and Evaluation Design — Be Suspicious of 99% Accuracy ★ MEMBER8 min read
  4. Step 4 Decision Trees and Gradient Boosting — Still the Champion on Tabular Data FREE11 min read
  5. Step 5 Unsupervised Learning from Scratch — Clustering and Dimensionality Reduction ★ MEMBER11 min read
  6. Step 6 Self-Supervised Learning — The Day Unlabeled Data Became an Asset ★ MEMBER11 min read
  7. Step 7 Data Leakage and Experiment Hygiene — When the Score Is Too Good, Suspect It ★ MEMBER10 min read
  8. Step 8 Imbalanced Data in Practice — What to Optimize When 99% Is Normal ★ MEMBER10 min read
  9. Step 9 ML System Design — The 90% Outside the Model ★ MEMBER8 min read
Deep Learning Basics7Neural nets, backprop, optimization, regularization

Open this genre →

  1. Step 10 Neural Networks from Scratch — From One Neuron to Many Layers FREE7 min read
  2. Step 11 Backpropagation from Scratch — It Is All Just the Chain Rule ★ MEMBER9 min read
  3. Step 12 A History of Normalization Layers — From BatchNorm to RMSNorm ★ MEMBER10 min read
  4. Step 13 Activation Functions from Scratch — Why Nonlinearity Is Non-Negotiable FREE13 min read
  5. Step 14 Weight Initialization and Regularization — What Lets Training Start, and What Keeps It Going ★ MEMBER15 min read
  6. Step 15 Hyperparameter Search — Hunches, Grids, and Bayesian Optimization ★ MEMBER12 min read
  7. Step 16 Graph Neural Networks from Scratch — Learning from Connections ★ MEMBER12 min read
CNNs & Image Recognition9From the convolution to ResNet, efficiency, and detection

Open this genre →

  1. Step 17 Image Classification from Scratch — The Invention of the Convolution FREE8 min read
  2. Step 18 Object Detection from Scratch (from YOLO to DETR) ★ MEMBER9 min read
  3. Step 19 The CNN Family Tree — From AlexNet to ResNet and EfficientNet ★ MEMBER8 min read
  4. Step 20 The Autonomous Driving Perception Stack from Scratch — What Cameras and LiDAR Each Bring to the Table FREE9 min read
  5. Step 21 Anomaly Detection from Scratch — Learning From Normal Alone ★ MEMBER11 min read
  6. Step 22 Medical Imaging AI — Validation Design Comes Before the Accuracy Number ★ MEMBER9 min read
  7. Step 23 The ImageNet Moment — The Day Deep Learning Won FREE9 min read
  8. Step 24 Paper Walkthrough: TurboVLA — Kick the LLM Out of the Loop and Run a Robot Policy at 32 Hz on an RTX 4090 with Under 1 GB of VRAM ★ MEMBER8 min read
  9. Step 25 Paper Walkthrough: PhiZero — A World Model That Reasons in a Language of Physics Before It Renders ★ MEMBER10 min read
How Transformers Work13Attention, positional encoding, and the architecture dissected

Open this genre →

  1. Step 26 Attention from Scratch — The Heart of the Transformer, Explained Visually FREE8 min read
  2. Step 27 Positional Encoding from Scratch — From Absolute Positions to RoPE FREE10 min read
  3. Step 28 Tokenizers from Scratch — The Unit an LLM Cuts the World Into FREE10 min read
  4. Step 29 A Field Guide to Attention Variants — MQA, GQA, Sliding Windows, Linear Attention ★ MEMBER9 min read
  5. Step 30 The Transformer, End to End — One Token's Journey from Embedding to Output FREE10 min read
  6. Step 31 FlashAttention from Scratch — The Paradox of Doing More Math to Go Faster ★ MEMBER10 min read
  7. Step 32 How Long-Context LLMs Work — From RoPE Interpolation to Ring Attention ★ MEMBER11 min read
  8. Step 33 Encoder or Decoder — The Fork in the Road Between BERT and GPT ★ MEMBER9 min read
  9. Step 34 Build Your Own BPE Tokenizer — Learning Merge Rules, and Getting Punished by Japanese ★ MEMBER11 min read
  10. Step 35 Build Your Own Mini GPT — A Language Model in 300 Lines ★ MEMBER12 min read
  11. Step 36 Paper Walkthrough: Designing Qwen3.8-Next — Accuracy, Efficiency and Stability as One Problem ★ MEMBER16 min read
  12. Step 37 Paper Explained: Why Gated DeltaNet Survives 4-Bit Quantization — NVFP4 W4A4 in a Hybrid 27B ★ MEMBER10 min read
  13. Step 38 Paper Walkthrough: Stop Anchoring to Frame One — Scal3R's Multi-Reference Relative Pose Query ★ MEMBER10 min read
Large Language Models14Scaling laws, prompting, and how LLMs behave inside

Open this genre →

  1. Step 39 Scaling Laws from Scratch — Why Making Models Bigger Makes Them Smarter (and When It Doesn't) ★ MEMBER9 min read
  2. Step 40 The Science of Prompt Engineering — What Is Proven and What Is Folklore FREE7 min read
  3. Step 41 Mamba and State Space Models — Handling Sequences Without Attention FREE12 min read
  4. Step 42 Building a Pretraining Corpus — From Web Sludge to Textbook Quality ★ MEMBER10 min read
  5. Step 43 LLM Evaluation from Scratch — Reading Benchmarks and the Contamination Problem ★ MEMBER11 min read
  6. Step 44 Why Language Models Hallucinate — The Mechanics and What Actually Helps FREE9 min read
  7. Step 45 Test-Time Scaling — How Models Get Better by Thinking Longer ★ MEMBER13 min read
  8. Step 46 Knowledge Distillation from Scratch — Copying a Big Model into a Small One ★ MEMBER9 min read
  9. Step 47 Alignment, Explained — From RLHF to Constitutional AI ★ MEMBER8 min read
  10. Step 48 Scaling Skepticism — A Genealogy of the "Just Make It Bigger" Critique ★ MEMBER10 min read
  11. Step 49 Paper Walkthrough: Can Anything Catch a Fake Crisis Video? — What RA-Bench Found ★ MEMBER8 min read
  12. Step 50 Paper Walkthrough: SA-MRPO — Stop Studying the Subject You've Already Aced ★ MEMBER11 min read
  13. Step 51 Paper Explained: Agentic Artifact Creation — Where Generation Ends and Construction Begins ★ MEMBER9 min read
  14. Step 52 Paper walkthrough: Puro-2B — pretraining a 2B model from scratch for $6.9K on consumer GPUs ★ MEMBER13 min read
VLMs & Multimodal6CLIP, ViT, and putting images and language in one space

Open this genre →

  1. Step 53 Paper Deep Dive — ViT: Treating an Image Like a Sentence ★ MEMBER7 min read
  2. Step 54 Paper Deep Dive — CLIP: Putting Words and Images on One Map ★ MEMBER7 min read
  3. Step 55 How VLMs Came Together — Wiring a Vision Encoder into an LLM ★ MEMBER8 min read
  4. Step 56 BEV Representations From Scratch — Fusing Multiple Cameras Into One Top-Down Map ★ MEMBER8 min read
  5. Step 57 Document AI and OCR Today — How an LLM Ends Up Reading Your Invoices ★ MEMBER10 min read
  6. Step 58 Video Understanding from Scratch — From a Pile of Frames to a Sense of Time ★ MEMBER8 min read
Generative Models10Diffusion, VAE, GAN — models that create

Open this genre →

  1. Step 59 Diffusion Models from the Ground Up — Add Noise, Then Subtract It ★ MEMBER9 min read
  2. Step 60 VAEs from Scratch — Stir Probability into "Compress and Restore" and You Get a Generator FREE10 min read
  3. Step 61 The Rise and Fall of GANs — An Invention Trained by Rivalry, and Why Diffusion Won ★ MEMBER10 min read
  4. Step 62 Flow Matching from Scratch — What Came After Diffusion, and Why It Goes Straight ★ MEMBER10 min read
  5. Step 63 CFG and Samplers — What the "Strength" Knob in Generative AI Really Does ★ MEMBER8 min read
  6. Step 64 The Mathematics of Diffusion — Generation Seen Through Scores and SDEs ★ MEMBER11 min read
  7. Step 65 A Practical Map of Image Generation — SD, ControlNet, and Applying LoRA FREE11 min read
  8. Step 66 Music and Audio Generation from Scratch — Sound as Tokens ★ MEMBER9 min read
  9. Step 67 The State of 3D Generation — From NeRF to Gaussian Splatting ★ MEMBER10 min read
  10. Step 68 Build Your Own Diffusion Model — Starting from MNIST ★ MEMBER11 min read
Training & Alignment13Pretraining, SFT, RLHF/DPO, fine-tuning

Open this genre →

  1. Step 69 Instruction Tuning and RLHF from Scratch — How a Model Learns to Follow Orders ★ MEMBER9 min read
  2. Step 70 Learning Rate Schedules — Why Warmup and Why Cosine ★ MEMBER10 min read
  3. Step 71 Mixed Precision Training — Going Faster in fp16/bf16/fp8 Without Breaking ★ MEMBER10 min read
  4. Step 72 Building a Dataset in Practice — Collect, Clean, Blend ★ MEMBER11 min read
  5. Step 73 DPO and What Came After — The Lineage That Simplified RLHF ★ MEMBER13 min read
  6. Step 74 Continual Learning and Catastrophic Forgetting — Why Models Can't Just Keep Learning ★ MEMBER9 min read
  7. Step 75 Diagnosing Broken Training — Telling Divergence, NaN, and Plateaus Apart FREE11 min read
  8. Step 76 Versioning Data and Models — An Experiment You Cannot Reproduce Never Happened ★ MEMBER10 min read
  9. Step 77 Paper Deep-Dive: ABSeeker — Training Long-Horizon Search Agents by Grading Each Step Backward from the Answer ★ MEMBER11 min read
  10. Step 78 PAWBench Explained — Can Video Generators Get the Odds Right, Not Just the Physics? ★ MEMBER10 min read
  11. Step 79 Paper Walkthrough: PaperGym — Turning One Paper Into a Graded Training Environment for Research Plans ★ MEMBER10 min read
  12. Step 80 Paper Walkthrough: StudentSim — Training a Simulator That Is Actually *That* Student ★ MEMBER12 min read
  13. Step 81 Paper Walkthrough: It Takes Two to Match — Co-Evolving Both Sides of Retrieval with RL ★ MEMBER13 min read
  14. +1 planned
Inference & Serving29KV cache, quantization, batching, serving

Open this genre →

  1. Step 82 The KV Cache from Scratch — The Heart of Fast Inference FREE7 min read
  2. Step 83 LLM Quantization from Scratch — Why Losing Precision Doesn't Break It ★ MEMBER8 min read
  3. Step 84 Speculative Decoding from Scratch — How a Tiny Draft Model Speeds Up an LLM Without Changing a Single Output ★ MEMBER8 min read
  4. Step 85 LLM Serving from Scratch — vLLM, Continuous Batching, and Not Letting the GPU Idle FREE10 min read
  5. Step 86 Structured Output and Constrained Decoding — How to Stop an LLM from Breaking Your JSON ★ MEMBER11 min read
  6. Step 87 Surviving GPU Out-of-Memory — Every Cause, Every Fix ★ MEMBER11 min read
  7. Step 88 Cutting Inference Cost in Practice — What to Do First ★ MEMBER11 min read
  8. Step 89 Prompt Caching and Context Design — One Prefix Rule That Moves Your Bill by an Order of Magnitude ★ MEMBER10 min read
  9. Step 90 Paper Walkthrough: The Personalization Mirage — LLMs Invent a Version of You, and Their Self-Reports Point the Wrong Way ★ MEMBER9 min read
  10. Step 91 Paper Walkthrough: DAPD — Breaking the Teacher's "Cheat-Sheet Illusion" in Distillation with Dual Anchors ★ MEMBER8 min read
  11. Step 92 Paper Walkthrough: CodeNib — A Multi-View Data System That Serves Repository Context to Coding Agents ★ MEMBER8 min read
  12. Step 93 Paper explained: BDH-CQ — an AI that thinks without words. Recurrent memory plus latent reasoning resets ARC's cost frontier ★ MEMBER8 min read
  13. Step 94 Paper Deep-Dive: AgentOPSD — Finding the Turn That Won the Game with Recursive Bayesian Belief Updates ★ MEMBER9 min read
  14. Step 95 Paper Walkthrough: No Gold Answers, No Stronger Teacher — How u-OPSD Distills From Its Own Majority Vote ★ MEMBER14 min read
  15. Step 96 Paper Explained: Agentic ESOpt — Drop Backprop, Jiggle the Weights, and Train Long-Horizon LLM Agents ★ MEMBER15 min read
  16. Step 97 Paper Walkthrough — WarpSAC: When RL's Safety Rails Become Handcuffs ★ MEMBER8 min read
  17. Step 98 TTPO Explained: Training a Model Mid-Exam, With No Answer Key ★ MEMBER11 min read
  18. Step 99 Paper explained: Self-OPD — an image generator that distills itself, with no teacher ★ MEMBER13 min read
  19. Step 100 Paper walkthrough: CyberFactory — turning wild CVEs into runnable training problems ★ MEMBER11 min read
  20. Step 101 Paper Walkthrough: DART-SD — Training Tool-Calling Agents Without Flattening the Diamond ★ MEMBER12 min read
  21. Step 102 Paper Walkthrough — SMELT: Is Looping the Same Layers Twice Actually a Win When the Budget Is Matched? ★ MEMBER15 min read
  22. Step 103 Paper Walkthrough: Normalized Low-Rank Adaptation — Why Normalizing LoRA's Entry Matrix Works ★ MEMBER8 min read
  23. Step 104 Paper Walkthrough: From Production Traffic to Post-Training — Folding 200 Internal Apps Into One Self-Hosted LLM ★ MEMBER14 min read
  24. Step 105 Paper Explained: Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement ★ MEMBER11 min read
  25. Step 106 Paper Explained — LatentPress: Feeding Compressed Context Straight to a Frozen LLM, Neither as Text Nor as Pixels ★ MEMBER12 min read
  26. Step 107 Paper Walkthrough: Language Models Can Control Their Own Attention ★ MEMBER9 min read
  27. Step 108 Paper Explained: Compile by Training — Turning a Natural-Language Spec into a Function That Runs Locally ★ MEMBER10 min read
  28. Step 109 Paper Walkthrough: One Training Example Keeps On-Policy Distillation Improving for Hundreds of Steps ★ MEMBER11 min read
  29. Step 110 Paper Walkthrough: Random Attention — Throwing KV Cache Entries Away at Random Works Just as Well ★ MEMBER11 min read
  30. +1 planned
RAG & Retrieval10Embeddings, chunking, reranking, evaluation

Open this genre →

  1. Step 111 RAG Fundamentals and Design Patterns — Embeddings, Chunking, Reranking, and Evaluation from Scratch ★ MEMBER8 min read
  2. Step 112 RAG vs Fine-Tuning — Which One, and When FREE8 min read
  3. Step 113 Embeddings from Scratch — from word2vec Intuition to Contextual Embeddings FREE6 min read
  4. Step 114 Recommenders and Embeddings — Same Math as RAG, Different Goal ★ MEMBER10 min read
  5. Step 115 Chunking Strategies — How You Split Decides What You Can Find ★ MEMBER10 min read
  6. Step 116 GraphRAG from Scratch — Where Knowledge Graphs Meet Retrieval ★ MEMBER10 min read
  7. Step 117 Evaluating RAG in Practice — Turning “Seems Better” Into a Number ★ MEMBER13 min read
  8. Step 118 Build Your Own Vector DB — From Brute Force to HNSW ★ MEMBER13 min read
  9. Step 119 Paper Walkthrough: WeMM-Embedding — Putting Text, Images and Video on One Ruler ★ MEMBER12 min read
  10. Step 120 Paper walkthrough: Hi-Q — splitting a question down to the granularity your corpus can actually retrieve ★ MEMBER13 min read
Agents55Tool use, planning, multi-agent systems

Open this genre →

  1. Step 121 LLM Agents from Scratch — Designing the Tool-Use Loop ★ MEMBER8 min read
  2. Step 122 End-to-End Driving from Scratch — Perception to Control in a Single Network ★ MEMBER9 min read
  3. Step 123 MCP and Tool Protocols — The Standard That Connects an Agent's Hands FREE11 min read
  4. Step 124 Multi-Agent Design Patterns — Division, Debate, Verification ★ MEMBER10 min read
  5. Step 125 Designing Agent Memory — Short-Term, Long-Term, Episodic ★ MEMBER11 min read
  6. Step 126 Evaluating Agents — How Benchmarks and Harnesses Are Built ★ MEMBER11 min read
  7. Step 127 AlphaGo from Scratch — The Marriage of Search and Learning ★ MEMBER9 min read
  8. Step 128 Build Your Own Agent Loop — The Minimal Shape of Tool Calling ★ MEMBER8 min read
  9. Step 129 Paper Walkthrough: SWE-Bench ProMax — Measuring What Coding Agents Can Really Do with Large-Scale, Multilingual Refactoring ★ MEMBER7 min read
  10. Step 130 Paper Walkthrough: Macaron-V1 — A Frozen Base plus a Mixture of LoRAs, Built to Keep Learning After Launch ★ MEMBER8 min read
  11. Step 131 Paper Walkthrough: EnvACE — Agents That Rehearse the World Instead of Calling It ★ MEMBER8 min read
  12. Step 132 Paper Explained: Video-DeepResearch — Agents That Watch a Video, Then Chase Down Every Lead ★ MEMBER9 min read
  13. Step 133 Paper Walkthrough: ToolArtist — Search, Draw, or Redraw? The Image Agent That Decides for Itself ★ MEMBER8 min read
  14. Step 134 Paper Deep-Dive: Recursive Synthesis — Extending Verified Tasks Into 40,000 Long-Horizon Terminal Problems ★ MEMBER13 min read
  15. Step 135 Paper Walkthrough: Qwen-UI-Agent — How Alibaba Built a GUI Agent That Works on Real Phones and PCs ★ MEMBER10 min read
  16. Step 136 Paper Explained: OSReward — Can You Trust the AI That Grades AI? Remeasuring Rewards for Computer-Use Agents ★ MEMBER10 min read
  17. Step 137 Paper Walkthrough: Metis — A 'Memory Foundation Model' That Moves Agent Memory Inside the Model ★ MEMBER8 min read
  18. Step 138 Paper Walkthrough: MerchantBench — Can an LLM Agent Run an Online Store for a Year? Why It Earns Only 27.3% of What Humans Do ★ MEMBER10 min read
  19. Step 139 Paper Walkthrough: Mental World Modeling — A World Model That Advances Minds, Not Just Physics ★ MEMBER9 min read
  20. Step 140 Paper Explained: LongHorizon-Harness — Long-Horizon Agent Tasks Are a State-Management Problem, Not an Execution Problem ★ MEMBER8 min read
  21. Step 141 Paper Deep-Dive: Frontis-MA1 — Training the AI That Builds AI: One Step Toward Recursive Self-Improvement in ML Engineering ★ MEMBER11 min read
  22. Step 142 Paper Walkthrough: ComBodied Agents — Moving an Agent's Target from Software and Matter to the Person ★ MEMBER11 min read
  23. Step 143 Paper Explained: Co-Evolution in Agentic Systems — Three Stages Toward Self-Directed Evolution ★ MEMBER9 min read
  24. Step 144 Paper walkthrough: Zetta ζ — a robot harness that repairs itself mid-execution, with the policy frozen ★ MEMBER13 min read
  25. Step 145 Paper walkthrough: StateM — 95.3% on Terminal-Bench 2.1 and a USD 15 run, without touching a single weight ★ MEMBER11 min read
  26. Step 146 Paper Explainer: SemaPLC — The Agent That Isn't Allowed to Say "Done" ★ MEMBER13 min read
  27. Step 147 Paper Walkthrough: OmniScientist — An AI Scientist That Actually Looks at the Raw Data ★ MEMBER11 min read
  28. Step 148 Paper Walkthrough — FACET: Grounding Instruction, Environment, Solution and Verifier in One Executable State ★ MEMBER12 min read
  29. Step 149 Paper Explained: EnvHarness — Reshaping an Agent's Training World Without Rebuilding It ★ MEMBER9 min read
  30. Step 150 Paper Explainer: Why Agent Skills Work — and Where They Break ★ MEMBER9 min read
  31. Step 151 Paper Explained: Co-RL — Reasoning Without Labels, Emerging From a Diverse Cohort ★ MEMBER12 min read
  32. Step 152 Paper Walkthrough: SWE-bench Science — Can Coding Agents Fix Scientific Code? ★ MEMBER9 min read
  33. Step 153 Paper Explained: FreeToken — Treating Your Own PC as a Single Elastic Inference Platform ★ MEMBER13 min read
  34. Step 154 Paper Walkthrough: Embodied-Navigator (TAMP-Nav) — Let the VLM Just Point, and Navigation Gets Both Faster and Better ★ MEMBER12 min read
  35. Step 155 Paper walkthrough: ASI-Bench — peeling away human guidance to measure what AI can do alone ★ MEMBER10 min read
  36. Step 156 Paper Walkthrough: FrontierChallenge — Grading Scientific Work on Whether It Was Actually Delivered ★ MEMBER8 min read
  37. Step 157 Paper Explained: JIT-Agent — A Model That Writes the Agent Harness On Demand ★ MEMBER13 min read
  38. Step 158 Paper Walkthrough: ZimaBlue — Turning 120,000 Hours of Egocentric Video into Robot Skill ★ MEMBER12 min read
  39. Step 159 Paper Explained: What Makes Good Agentic Data? The ACE Lens ★ MEMBER11 min read
  40. Step 160 Paper Walkthrough — UrbanGround: Where MLLM Agents Break Down on a Real Street ★ MEMBER10 min read
  41. Step 161 Paper Walkthrough: UI-Venus-2 — Taking Screen-Operating Agents From Benchmarks to Real Work ★ MEMBER13 min read
  42. Step 162 Paper Walkthrough: Training Agents to Evolve with Their Harness ★ MEMBER15 min read
  43. Step 163 Paper Explained: StarHarness — Evolving the Scaffold Instead of the Weights ★ MEMBER13 min read
  44. Step 164 Paper Walkthrough: SecOPD — Grading One Token at a Time to Cut Adaptive Prompt Injection by an Order of Magnitude ★ MEMBER10 min read
  45. Step 165 Paper Explained: Repo-To-Skill — Distilling GitHub Repositories Into Skills an AI Can Use ★ MEMBER12 min read
  46. Step 166 Paper Walkthrough: PILOT in the Loop — Fixing the Run While It Is Still Running ★ MEMBER15 min read
  47. Step 167 Paper Walkthrough: LoopArena — Benchmarking the Model That Steers a Coding Agent ★ MEMBER11 min read
  48. Step 168 Paper Walkthrough: Code World Model — Putting a Coding Agent in Charge of the World ★ MEMBER14 min read
  49. Step 169 Paper Walkthrough: Code as Worlds — An Agent That Writes the World Down as Runnable Code ★ MEMBER13 min read
  50. Step 170 Paper Walkthrough: AutoSaddler — Growing a Harness That Doesn't Break, from Agent Failure Logs ★ MEMBER11 min read
  51. Step 171 Paper Walkthrough: HarnessDev — Can an LLM Build and Maintain the System It Runs Inside? ★ MEMBER12 min read
  52. Step 172 Paper Walkthrough: EarlyEval — Making Agent Evaluation Cheaper by Stopping Early ★ MEMBER12 min read
  53. Step 173 Paper Walkthrough: Aspire — Can Models Self-Evolve from Vague Goals? ★ MEMBER14 min read
  54. Step 174 Paper Walkthrough: Terminal-Universe — Turning Agent Logs Back Into Reusable Execution Environments ★ MEMBER12 min read
  55. Step 175 CogEvol: What the Reward Cannot Measure, RL Will Quietly Destroy ★ MEMBER9 min read
  56. +6 planned
Audio & Speech9ASR, TTS, speech LLMs

Open this genre →

  1. Step 176 Paper Deep-Dive: Why Whisper Is Robust — Large-Scale Weak Supervision ★ MEMBER8 min read
  2. Step 177 Speech Recognition from Scratch — From Waveform to Text FREE11 min read
  3. Step 178 Speech Synthesis from Scratch — From Text to a Voice ★ MEMBER12 min read
  4. Step 179 Representing Sound — Mel Spectrograms and Audio Tokens ★ MEMBER13 min read
  5. Step 180 Paper Walkthrough: SwanTale — Designing Voices from Words Alone, with Speech and Sound in One Waveform ★ MEMBER10 min read
  6. Step 181 Paper Walkthrough: Interpretable MEG Decoding of Perceived Speech — Reading the Decoder's Weights as a Brain Map ★ MEMBER9 min read
  7. Step 182 Paper Walkthrough: VoiceMem — A Left Brain and a Right Brain for Voice Agents, at Zero Added Latency ★ MEMBER14 min read
  8. Step 183 Paper Walkthrough: Motion-Omni — Speaking and Moving in One Forward Pass ★ MEMBER12 min read
  9. Step 184 Paper Explained: Last Translation Benchmark — Measuring Translation with Breaking Examples and Verification Rules ★ MEMBER9 min read
Time Series5From RNN/LSTM to time-series foundation models

Open this genre →

  1. Step 185 RNNs and LSTMs from Scratch — Why Learn Them in the Transformer Era FREE11 min read
  2. Step 186 Time-Series Forecasting from Scratch — From Classical Methods to Foundation Models ★ MEMBER9 min read
  3. Step 187 Do Transformers Actually Work on Time Series? — The Argument and the Practical Answer ★ MEMBER11 min read
  4. Step 188 Time-Series Anomaly Detection — The Math Behind the Alerts ★ MEMBER11 min read
  5. Step 189 H3-World, Explained — Turning Language Understanding into World Control ★ MEMBER12 min read
Model Families7Gemma, Llama, Qwen and friends — lineage and when to use which

Open this genre →

  1. Step 190 The Gemma Family from Scratch — Lineage, Inventions, and Where It Fits ★ MEMBER10 min read
  2. Step 191 The Llama Family from Scratch — The Main Line of Open LLMs ★ MEMBER9 min read
  3. Step 192 The Qwen Family from Scratch — Why It Tops Hugging Face's Download Charts ★ MEMBER7 min read
  4. Step 193 The DeepSeek Family from Scratch — Breaking In with MoE and Distillation FREE9 min read
  5. Step 194 The GPT Lineage — Design Thinking from GPT-1 to Today ★ MEMBER9 min read
  6. Step 195 The Mistral Family from Scratch — Europe's Small-and-Strong Bet ★ MEMBER9 min read
  7. Step 196 Open vs Closed — The Economics of Releasing Weights, and the Safety Argument ★ MEMBER11 min read
Paper Deep-Dives13Landmark papers one by one, grounded in the source text

Open this genre →

  1. Step 197 Paper Deep Dive — LoRA: Low-Rank Adaptation of Large Language Models: Why Low Rank Is Enough ★ MEMBER12 min read
  2. Step 198 When Proxies Stop Being Good Enough — Reading August 2026's Eight Autonomous Driving Papers Together FREE13 min read
  3. Step 199 Paper Deep Dive — Attention Is All You Need: What Dropping Recurrence Actually Proved ★ MEMBER12 min read
  4. Step 200 Mixture of Experts (MoE) from Scratch — Routing and Load Balancing in the Switch Transformer ★ MEMBER9 min read
  5. Step 201 Paper Walkthrough: Alpamayo — NVIDIA's Reasoning Model for Autonomous Driving ★ MEMBER8 min read
  6. Step 202 Symbolic vs. Connectionist — Where a 60-Year Argument Stands Today ★ MEMBER11 min read
  7. Step 203 Paper Walkthrough: WorldClaw — Agents That Build Walkable, Editable 3D Open Worlds from a Single Sentence ★ MEMBER11 min read
  8. Step 204 Paper Deep Dive: AskChem — Changing the Unit of Search from Papers to Provenance-Carrying Claims ★ MEMBER8 min read
  9. Step 205 Paper Deep Dive — Large Discovery Models: giving an LLM a value signal for what to try next ★ MEMBER13 min read
  10. Step 206 Paper walkthrough: Apodex 1.1 — scaling agents around completed work ★ MEMBER11 min read
  11. Step 207 Paper Walkthrough: Turning Game Development into a Verifiable Trajectory Data Engine — RLHEV and AWoMo ★ MEMBER13 min read
  12. Step 208 Paper Walkthrough — J-Zero: Growing the Challenger, the Solver, and the Judge Together from Zero Data ★ MEMBER12 min read
  13. Step 209 Paper Walkthrough — RoboTok: Mining the Web for Demonstrations That Move Like Yours ★ MEMBER11 min read
Distillation & Compression9Moving capability from a large model into a small one — losses, data design and failure modes

Open this genre →

  1. Step 210 The Math of Distillation — Why Soft Answers Teach More FREE8 min read
  2. Step 211 A Field Guide to Distillation Recipes — logit, feature, attention, self ★ MEMBER9 min read
  3. Step 212 Designing Distillation Data — Deciding What to Ask the Teacher ★ MEMBER8 min read
  4. Step 213 When Distillation Fails — Capacity Gaps and Contagious Overconfidence ★ MEMBER8 min read
  5. Step 214 On-Policy Distillation — Learning From What the Student Actually Writes ★ MEMBER9 min read
  6. Step 215 Distillation vs. Quantization vs. Pruning — Three Roads to a Smaller Model ★ MEMBER10 min read
  7. Step 216 Distilling Agents — How to Compress a Long Trajectory ★ MEMBER10 min read
  8. Step 217 Evaluating Distilled Models — Is "Close to the Teacher" a Good Metric? ★ MEMBER9 min read
  9. Step 218 Build Your Own Distillation — Growing a Small Model in 100 Lines ★ MEMBER11 min read
Evaluation & Judging5How to measure AI — benchmarks, LLM judges, contamination and reward hacking

Open this genre →

  1. Step 219 LLM-as-a-Judge from Scratch — How AI Grades AI, and Where It Breaks FREE10 min read
  2. Step 220 A Map of Agent Benchmarks — What SWE-bench, GAIA, and OSWorld Actually Measure FREE11 min read
  3. Step 221 Self-Improving AI — Self-Play, Co-Evolution, and Generated Curricula ★ MEMBER12 min read
  4. Step 222 Reward Hacking — Whatever You Measure Is Where It Breaks ★ MEMBER10 min read
  5. Step 223 Benchmark Contamination — How to Doubt a High Score ★ MEMBER10 min read

Semiconductors

How silicon pulled from sand comes to compute — from the physics of matter to arithmetic units, memory, power and fabrication.

Device Physics5Bands, doping, pn junctions, MOSFET, CMOS

Open this genre →

  1. Step 224 MOSFETs from the Ground Up — A Sluice Gate Opened by Voltage, and the Reality of Leakage FREE9 min read
  2. Step 225 Reading Chip Design as a Power Budget — The Physics of Leakage and Heat ★ MEMBER10 min read
  3. Step 226 Band Theory from the Ground Up — Why It Had to Be Silicon FREE11 min read
  4. Step 227 From FinFET to GAA — Why the Transistor Had to Go Vertical ★ MEMBER9 min read
  5. Step 228 The Physics of NAND Flash — Remembering by Trapping Electrons ★ MEMBER12 min read
Computer Architecture5Logic, clocking, the memory hierarchy and the memory wall

Open this genre →

  1. Step 229 The Memory Wall from Scratch — Why Moving Data Costs More Than Computing ★ MEMBER9 min read
  2. Step 230 The GPU Memory Hierarchy — HBM, SRAM, Registers, and Why Movement Wins ★ MEMBER9 min read
  3. Step 231 CPU Pipelines and Branch Prediction — The Factory Inside One Clock Tick ★ MEMBER10 min read
  4. Step 232 Systolic Arrays — Building the Heart of the TPU From Scratch ★ MEMBER10 min read
  5. Step 233 Interconnects — How NVLink, PCIe, and Light Set the Limits of Scale ★ MEMBER9 min read
Scaling & Power4The end of Dennard scaling, the physics of power, the cost of moving data

Open this genre →

  1. Step 234 The Physics of Power — Why Lowering Voltage Pays So Much ★ MEMBER9 min read
  2. Step 235 What Moore's Law Actually Says — What Ended, and What Is Still Going FREE10 min read
  3. Step 236 The Economics of Chiplets — We Split Dies Because We Cannot Build Them Big ★ MEMBER10 min read
  4. Step 237 Thermal Design from Scratch — The Wall in 3D Stacking Is Heat ★ MEMBER11 min read
Fabrication & Packaging4Wafers, lithography, yield, chiplets, 3D integration

Open this genre →

  1. Step 238 How Chips Are Made — From Wafer to Yield ★ MEMBER9 min read
  2. Step 239 Advanced Packaging — How CoWoS and HBM Stacking Became the Bottleneck for AI ★ MEMBER10 min read
  3. Step 240 EUV Lithography — The Madness of Making 13.5nm Light ★ MEMBER10 min read
  4. Step 241 Yield and Design — DFM, the Art of Giving Something Up ★ MEMBER11 min read
Accelerators3Number formats, accelerator design styles, edge deployment

Open this genre →

  1. Step 242 How AI Accelerators Are Designed — What Actually Separates GPUs, NPUs, and TPUs ★ MEMBER8 min read
  2. Step 243 How Numbers Are Represented — From FP32 to FP8 and INT4 FREE9 min read
  3. Step 244 The Inference Chip Wars — Inside the Design Philosophies of Groq, Cerebras, and the LPU ★ MEMBER9 min read
Supply Chain6Materials, equipment, EDA and IP — who holds what upstream of the chip itself

Open this genre →

  1. Step 245 Mapping the Semiconductor Supply Chain — From Sand to Chip, Who Holds What FREE8 min read
  2. Step 246 Upstream of Semiconductors — Wafers, Photoresist, and Specialty Gases ★ MEMBER13 min read
  3. Step 247 Inside the Equipment Makers — What ASML, AMAT, TEL and Lam Actually Build ★ MEMBER9 min read
  4. Step 248 EDA Tools from Scratch — Chips Are Written in Software ★ MEMBER10 min read
  5. Step 249 IP Cores and the Fabless Model — How Arm Rules Silicon Without Making a Single Chip ★ MEMBER11 min read
  6. Step 250 The Geopolitics of Chips — Export Controls and Supply Chain Rewiring, Explained Technically ★ MEMBER9 min read

Algorithms

From complexity as a yardstick to search, dynamic programming, matrix computation and parallelism — the craft of computation underneath AI.

Complexity5Big-O, time-space tradeoffs, and where theory parts from the profiler

Open this genre →

  1. Step 251 Complexity From Scratch — What Big-O Actually Measures FREE7 min read
  2. Step 252 When Big-O and Your Benchmarks Disagree — Caches, Branches, and Memory Bandwidth ★ MEMBER8 min read
  3. Step 253 Randomized Algorithms — Why Rolling Dice Makes Things Faster ★ MEMBER12 min read
  4. Step 254 NP-Completeness from Scratch — Not Unsolvable, but Fast to Verify FREE9 min read
  5. Step 255 Approximation Algorithms — Trading Exactness for a Guarantee ★ MEMBER12 min read
Data Structures5Arrays, trees, hashes, heaps — what each choice buys you

Open this genre →

  1. Step 256 Choosing a Data Structure — Arrays, Hashes, Trees and Heaps FREE7 min read
  2. Step 257 Hashing and Nearest-Neighbor Search — The Groundwork Under Vector Search ★ MEMBER9 min read
  3. Step 258 Cache-Friendly Code — Why Two O(n) Loops Can Differ by 10× ★ MEMBER11 min read
  4. Step 259 B-Trees and LSM-Trees — The Heart of Every Database ★ MEMBER10 min read
  5. Step 260 Probabilistic Data Structures — Counting Without Counting ★ MEMBER13 min read
Search & Optimization4Exhaustive, greedy, DP, branch and bound, approximation

Open this genre →

  1. Step 261 Dynamic Programming From Scratch — On Remembering Subproblems ★ MEMBER8 min read
  2. Step 262 Graph Algorithms from Scratch — Shortest Paths and Where They Lead ★ MEMBER8 min read
  3. Step 263 Linear Programming from Scratch — The Workhorse of Optimization ★ MEMBER11 min read
  4. Step 264 Simulated Annealing and Genetic Algorithms — What to Do When Exact Solving Breaks Down ★ MEMBER11 min read
Numerical Computing6GEMM, decompositions, FFT, iterative methods — where AI's compute actually goes

Open this genre →

  1. Step 265 The Cost of Matrix Multiplication — Where Almost All of AI's Compute Goes ★ MEMBER8 min read
  2. Step 266 The FFT from Scratch — Why Convolution Turns into Multiplication ★ MEMBER9 min read
  3. Step 267 Numerical Pitfalls — Cancellation, Rounding, and logsumexp ★ MEMBER10 min read
  4. Step 268 How Autodiff Actually Works — Unpacking the PyTorch Magic FREE10 min read
  5. Step 269 Solving Systems of Equations — Direct Methods and Iterative Methods ★ MEMBER13 min read
  6. Step 270 Build Your Own Autograd — A Mini PyTorch in 100 Lines ★ MEMBER12 min read
Parallel & Distributed3Limits of parallelism, the GPU execution model, communication in distributed training

Open this genre →

  1. Step 271 Why GPUs Are Fast — The Execution Model and the Limits of Parallelism ★ MEMBER10 min read
  2. Step 272 Distributed Training from Scratch — Data Parallel, Model Parallel, and When Communication Becomes the Bottleneck ★ MEMBER10 min read
  3. Step 273 Concurrency from Scratch — Locks, Atomics, and Memory Models ★ MEMBER9 min read

Math for AI

Only the math you actually need to read AI papers, starting from what the symbols mean: linear algebra, calculus, probability, information theory, optimization.

Linear Algebra6Vectors, matrices, eigenvalues, SVD — what a dimension really is

Open this genre →

  1. Step 274 Linear Algebra for AI — What Vectors and Matrices Are Actually Doing FREE9 min read
  2. Step 275 The Linear Algebra Under LoRA and RAG — Eigenvalues, Low Rank and Vector Search, Hands On FREE7 min read
  3. Step 276 Singular Value Decomposition and Low-Rank Approximation — the Math Behind LoRA ★ MEMBER10 min read
  4. Step 277 A Tour of Matrix Decompositions — When to Reach for LU, QR, Cholesky, or SVD ★ MEMBER13 min read
  5. Step 278 Tensors and Shape Manipulation — If You Can Read einsum, You Can Read Papers FREE10 min read
  6. Step 279 Symmetry and Equivariance — How Group Theory Shapes Network Design ★ MEMBER10 min read
Calculus & Optimization6Partial derivatives, gradients, the chain rule, convexity, Lagrange

Open this genre →

  1. Step 280 Calculus for AI — The Gradient Is an Arrow Saying Which Way Is Better FREE7 min read
  2. Step 281 Convexity and Optimization — Why Deep Learning Works Even Though It Isn't Convex ★ MEMBER9 min read
  3. Step 282 Beyond SGD — Adam, Second-Order Methods, and Constrained Optimization ★ MEMBER12 min read
  4. Step 283 Matrix Calculus from Scratch — Derive the Backward Pass Yourself ★ MEMBER11 min read
  5. Step 284 Jacobians and Hessians — Multivariable Calculus, Drawn ★ MEMBER11 min read
  6. Step 285 Calculus of Variations — What It Means to Differentiate a Function ★ MEMBER13 min read
Probability & Statistics10Distributions, expectation, Bayes, MLE, sampling

Open this genre →

  1. Step 286 Probability and Statistics for AI — A Model's Output Is a Distribution ★ MEMBER8 min read
  2. Step 287 Bayes' Theorem in AI — Priors, Posteriors, and Uncertainty ★ MEMBER10 min read
  3. Step 288 Thinking Bayesian — A Working Feel for Priors, Likelihoods, and Posteriors ★ MEMBER11 min read
  4. Step 289 Markov Chains from Scratch — The Process That Only Looks at Now ★ MEMBER11 min read
  5. Step 290 A Field Guide to Probability Distributions — Where Normal, Poisson, and the Exponential Family Come From FREE13 min read
  6. Step 291 Monte Carlo Methods from Scratch — Solving Integrals with Dice ★ MEMBER10 min read
  7. Step 292 Hypothesis Testing and A/B Tests — How to Use a p-value, and How People Misuse It ★ MEMBER10 min read
  8. Step 293 Statistical Learning Theory — Why Does Learning Generalize? ★ MEMBER11 min read
  9. Step 294 Optimal Transport — The Mathematics of Moving Distributions ★ MEMBER10 min read
  10. Step 295 Kernel Methods and Gaussian Processes — The Champions Before Neural Nets ★ MEMBER11 min read
Information Theory5Entropy, KL divergence, cross-entropy — where loss functions come from

Open this genre →

  1. Step 296 Information Theory and AI — Where Cross-Entropy Loss Came From ★ MEMBER9 min read
  2. Step 297 KL Divergence From Scratch — Measuring the Gap Between Two Distributions ★ MEMBER9 min read
  3. Step 298 Entropy and Cross-Entropy — Where the Loss Function Comes From FREE9 min read
  4. Step 299 Mutual Information — Putting a Number on What You Know ★ MEMBER11 min read
  5. Step 300 Compression Is Prediction Is Intelligence — LLMs Through Information Theory ★ MEMBER11 min read

Media & Compression

Why JPEG degrades and PNG does not — image, video and audio coding from both the information theory and the settings you actually ship.

Coding Theory3Lossless vs lossy, entropy coding, Huffman, arithmetic coding

Open this genre →

  1. Step 301 Entropy Coding from Scratch — From Huffman to Arithmetic Coding ★ MEMBER9 min read
  2. Step 302 Error Correction from Scratch — Sending Data on the Assumption It Will Break FREE11 min read
  3. Step 303 LDPC and Turbo Codes — The Error Correction Behind 5G and Deep Space ★ MEMBER11 min read
Image Codecs3JPEG, PNG, WebP, AVIF — the DCT, quantization, and lossless pipelines

Open this genre →

  1. Step 304 Why JPEG Degrades — The DCT and Quantization from Scratch ★ MEMBER9 min read
  2. Step 305 Why PNG Does Not Degrade — Prediction Filters and Deflate FREE8 min read
  3. Step 306 WebP, AVIF, JPEG XL — The Image Codec Changing of the Guard ★ MEMBER10 min read
Video Codecs4MPEG, H.264, AV1 — motion compensation, GOP structure, rate control

Open this genre →

  1. Step 307 Video Compression from Scratch — Motion Compensation and the GOP ★ MEMBER8 min read
  2. Step 308 From H.264 to AV1 — What a Codec Generation Change Really Involves ★ MEMBER7 min read
  3. Step 309 Motion Compensation from Scratch — Where 90% of Video Compression Happens ★ MEMBER10 min read
  4. Step 310 Neural Compression — The Codec That Learns ★ MEMBER11 min read
Audio Codecs4MP3, AAC, Opus — perceptual masking as a way of discarding

Open this genre →

  1. Step 311 Audio Compression from Scratch — The Art of Discarding What You Cannot Hear ★ MEMBER8 min read
  2. Step 312 MP3 and Psychoacoustics — The Science of Sounds You Cannot Hear ★ MEMBER10 min read
  3. Step 313 Why Opus Won — The Design of a Modern Audio Codec ★ MEMBER11 min read
  4. Step 314 Neural Audio Codecs — EnCodec and the Foundation Under Speech LLMs ★ MEMBER12 min read
Media in Production3Bitrate ladders, encoder settings, quality metrics — what the job actually asks

Open this genre →

  1. Step 315 Designing a Bitrate Ladder — The First Job You Get in Streaming ★ MEMBER9 min read
  2. Step 316 How to Read Encoder Settings — What CRF, Presets, and 2-Pass Actually Change ★ MEMBER7 min read
  3. Step 317 HLS and DASH — How Video Actually Reaches You FREE10 min read

Systems

The floor AI runs on: operating systems, networks, databases, language runtimes, security and cloud — understood from their mechanisms.

OS & Runtime3Processes, memory, containers — the ground programs stand on

Open this genre →

  1. Step 318 Processes and Memory from Scratch — What Is the OS Actually Protecting? FREE13 min read
  2. Step 319 Containers from Scratch — What namespaces and cgroups Actually Are ★ MEMBER12 min read
  3. Step 320 File Systems from Scratch — What Does "Saved" Actually Guarantee? ★ MEMBER9 min read
Networking3TCP/QUIC, DNS, CDN — how data actually arrives

Open this genre →

  1. Step 321 From TCP to QUIC — Reinventing Reliable Communication FREE11 min read
  2. Step 322 DNS and CDNs — What Happens Between Enter and Pixels ★ MEMBER11 min read
  3. Step 323 HTTP/1.1 → 2 → 3 — The Road to Multiplexing ★ MEMBER8 min read
Databases3Internals, transactions, distribution — why data survives

Open this genre →

  1. Step 324 Database Internals — What Happens Behind a Single Line of SQL FREE10 min read
  2. Step 325 Transactions and ACID — Concurrency Hell and the Isolation Levels ★ MEMBER9 min read
  3. Step 326 Distributed Databases — CAP, Replication and Consensus ★ MEMBER10 min read
Compilers & Runtimes3Compilers, JITs, GC — why code runs fast

Open this genre →

  1. Step 327 Compilers From Scratch — How Source Becomes Machine Code FREE11 min read
  2. Step 328 JIT and GC — Getting Faster While Running, Cleaning Up While Running ★ MEMBER10 min read
  3. Step 329 Why Python Is Slow — Said Precisely ★ MEMBER11 min read
Security4Crypto, auth, LLM attacks — designing for defence

Open this genre →

  1. Step 330 Cryptography from Scratch — Symmetric Keys, Public Keys, and Hashes FREE13 min read
  2. Step 331 TLS from Scratch — Key Exchange and Certificates ★ MEMBER10 min read
  3. Step 332 Authentication and Authorization — From Passwords to OAuth and Passkeys ★ MEMBER14 min read
  4. Step 333 LLM Security — Prompt Injection and How to Actually Defend Against It ★ MEMBER9 min read
Cloud & Ops3Serverless, cost, observability — shipping without breaking

Open this genre →

  1. Step 334 Serverless and Cost Design — Cloud That Won't Bankrupt You FREE10 min read
  2. Step 335 Observability — Logs, Metrics, and Traces in Practice ★ MEMBER10 min read
  3. Step 336 The Economics of GPU Cloud — Rent, Buy, or Commit ★ MEMBER10 min read
Engineering Process5Requirements, design, estimation — the upstream where projects are won or lost

Open this genre →

  1. Step 337 Upstream Engineering from Scratch — Why Projects Are Won or Lost at Requirements FREE8 min read
  2. Step 338 The Craft of Requirements — Why "We Built Exactly What They Asked For" Fails ★ MEMBER12 min read
  3. Step 339 Architecture Decisions — Telling Apart What You Can Undo From What You Cannot ★ MEMBER9 min read
  4. Step 340 Estimation and Scope — The Cone of Uncertainty and How to Negotiate ★ MEMBER10 min read
  5. Step 341 Upstream Work in the AI Era — When Code Gets Cheap, What Gets Valuable? ★ MEMBER9 min read