JA EN
Home › AI

◉ FIELD

AI

From the basics of learning to CNNs, Transformers, VLMs and agents — following why each architecture took the shape it did.

11chapters 32Foundations 191Paper walkthroughs 1Interactive

Work through a volume in order: textbook → foundations → papers → lab.

① Textbook

  1. 16Prehistory — Perceptrons and Two Winters
  2. 17The Invention of Convolution
  3. 182012: The AlexNet Shock
  4. 19The Race to Go Deeper — VGG to ResNet
  5. 20The Race to Go Lighter — the MobileNet Line
  6. 21Detection and Segmentation
  7. 22Attention and the Transformer
  8. 23ViT — Treating Images Like Words
  9. 24CLIP — Words and Images on One Map
  10. 25How VLMs Came About, and Where They Are
  11. 26The Reality of Running at the Edge

② Foundations

Articles that assume nothing and build the ideas of the field, in order.

Machine Learning Basics

What learning is: loss, overfitting, evaluation — the foundation for everything

  1. What Machine Learning Really Is — Understanding “Learning” Without the Math FREE
  2. Loss Functions and Optimization — How a Model Learns From Being Wrong FREE
  3. Overfitting and Evaluation Design — Be Suspicious of 99% Accuracy ★ MEMBER
  4. Data Leakage and Experiment Hygiene — When the Score Is Too Good, Suspect It ★ MEMBER
  5. Imbalanced Data in Practice — What to Optimize When 99% Is Normal ★ MEMBER
  6. ML System Design — The 90% Outside the Model ★ MEMBER

Deep Learning Basics

Neural nets, backprop, optimization, regularization

  1. Neural Networks from Scratch — From One Neuron to Many Layers FREE
  2. Backpropagation from Scratch — It Is All Just the Chain Rule ★ MEMBER

CNNs & Image Recognition

From the convolution to ResNet, efficiency, and detection

  1. Image Classification from Scratch — The Invention of the Convolution FREE
  2. The Autonomous Driving Perception Stack from Scratch — What Cameras and LiDAR Each Bring to the Table FREE
  3. Anomaly Detection from Scratch — Learning From Normal Alone ★ MEMBER
  4. Medical Imaging AI — Validation Design Comes Before the Accuracy Number ★ MEMBER

How Transformers Work

Attention, positional encoding, and the architecture dissected

  1. Attention from Scratch — The Heart of the Transformer, Explained Visually FREE

Large Language Models

Scaling laws, prompting, and how LLMs behave inside

Coming soon

VLMs & Multimodal

CLIP, ViT, and putting images and language in one space

Coming soon

Generative Models

Diffusion, VAE, GAN — models that create

  1. VAEs from Scratch — Stir Probability into "Compress and Restore" and You Get a Generator FREE

Training & Alignment

Pretraining, SFT, RLHF/DPO, fine-tuning

  1. Versioning Data and Models — An Experiment You Cannot Reproduce Never Happened ★ MEMBER

Inference & Serving

KV cache, quantization, batching, serving

  1. The KV Cache from Scratch — The Heart of Fast Inference FREE
  2. LLM Quantization from Scratch — Why Losing Precision Doesn't Break It ★ MEMBER
  3. Cutting Inference Cost in Practice — What to Do First ★ MEMBER
  4. Prompt Caching and Context Design — One Prefix Rule That Moves Your Bill by an Order of Magnitude ★ MEMBER

RAG & Retrieval

Embeddings, chunking, reranking, evaluation

  1. RAG Fundamentals and Design Patterns — Embeddings, Chunking, Reranking, and Evaluation from Scratch ★ MEMBER
  2. RAG vs Fine-Tuning — Which One, and When FREE
  3. Chunking Strategies — How You Split Decides What You Can Find ★ MEMBER

Agents

Tool use, planning, multi-agent systems

  1. MCP and Tool Protocols — The Standard That Connects an Agent's Hands FREE

Audio & Speech

ASR, TTS, speech LLMs

  1. Speech Recognition from Scratch — From Waveform to Text FREE

Time Series

From RNN/LSTM to time-series foundation models

  1. RNNs and LSTMs from Scratch — Why Learn Them in the Transformer Era FREE
  2. Time-Series Forecasting from Scratch — From Classical Methods to Foundation Models ★ MEMBER

Model Families

Gemma, Llama, Qwen and friends — lineage and when to use which

  1. The Llama Family from Scratch — The Main Line of Open LLMs ★ MEMBER
  2. The Qwen Family from Scratch — Why It Tops Hugging Face's Download Charts ★ MEMBER
  3. The DeepSeek Family from Scratch — Breaking In with MoE and Distillation FREE
  4. The GPT Lineage — Design Thinking from GPT-1 to Today ★ MEMBER
  5. The Mistral Family from Scratch — Europe's Small-and-Strong Bet ★ MEMBER
  6. Open vs Closed — The Economics of Releasing Weights, and the Safety Argument ★ MEMBER

Paper Deep-Dives

Landmark papers one by one, grounded in the source text

Coming soon

Distillation & Compression

Moving capability from a large model into a small one — losses, data design and failure modes

Coming soon

Evaluation & Judging

How to measure AI — benchmarks, LLM judges, contamination and reward hacking

Coming soon

③ Paper walkthroughs

Written from the papers themselves. Every piece links the paper page and its PDF.

  1. Paper Explained: Last Translation Benchmark — Measuring Translation with Breaking Examples and Verification Rules ★ MEMBER arXiv:2609.04173
  2. Paper Walkthrough: Motion-Omni — Speaking and Moving in One Forward Pass ★ MEMBER arXiv:2609.04250
  3. CogEvol: What the Reward Cannot Measure, RL Will Quietly Destroy ★ MEMBER arXiv:2608.30968
  4. Mamba and State Space Models — Handling Sequences Without Attention FREE arXiv:2111.00396
  5. Paper Walkthrough: Random Attention — Throwing KV Cache Entries Away at Random Works Just as Well ★ MEMBER arXiv:2609.03430
  6. Paper Walkthrough: One Training Example Keeps On-Policy Distillation Improving for Hundreds of Steps ★ MEMBER arXiv:2609.04172
  7. Paper Walkthrough: Terminal-Universe — Turning Agent Logs Back Into Reusable Execution Environments ★ MEMBER arXiv:2609.04148
  8. Paper Walkthrough — RoboTok: Mining the Web for Demonstrations That Move Like Yours ★ MEMBER arXiv:2609.03199
  9. Paper Walkthrough: Stop Anchoring to Frame One — Scal3R's Multi-Reference Relative Pose Query ★ MEMBER arXiv:2609.04201
  10. Paper Explained: Why Gated DeltaNet Survives 4-Bit Quantization — NVFP4 W4A4 in a Hybrid 27B ★ MEMBER arXiv:2609.04098
  11. Paper Explained: Compile by Training — Turning a Natural-Language Spec into a Function That Runs Locally ★ MEMBER arXiv:2609.04199
  12. Paper walkthrough: Hi-Q — splitting a question down to the granularity your corpus can actually retrieve ★ MEMBER arXiv:2608.30468
  13. Paper Walkthrough: Language Models Can Control Their Own Attention ★ MEMBER arXiv:2609.02737
  14. Paper Explained — LatentPress: Feeding Compressed Context Straight to a Frozen LLM, Neither as Text Nor as Pixels ★ MEMBER arXiv:2609.01507
  15. Paper Walkthrough: Aspire — Can Models Self-Evolve from Vague Goals? ★ MEMBER arXiv:2608.31111
  16. Paper Walkthrough: EarlyEval — Making Agent Evaluation Cheaper by Stopping Early ★ MEMBER arXiv:2609.02783
  17. Paper Walkthrough: HarnessDev — Can an LLM Build and Maintain the System It Runs Inside? ★ MEMBER arXiv:2609.01437
  18. Paper Walkthrough: It Takes Two to Match — Co-Evolving Both Sides of Retrieval with RL ★ MEMBER arXiv:2609.00638
  19. A Map of Agent Benchmarks — What SWE-bench, GAIA, and OSWorld Actually Measure FREE arXiv:2310.06770
  20. Paper Walkthrough: AutoSaddler — Growing a Harness That Doesn't Break, from Agent Failure Logs ★ MEMBER arXiv:2608.23041
  21. Paper Walkthrough: Code as Worlds — An Agent That Writes the World Down as Runnable Code ★ MEMBER arXiv:2608.27549
  22. Paper Walkthrough: Code World Model — Putting a Coding Agent in Charge of the World ★ MEMBER arXiv:2608.25927
  23. Paper Explained: Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement ★ MEMBER arXiv:2608.31046
  24. Benchmark Contamination — How to Doubt a High Score ★ MEMBER arXiv:2311.04850
  25. Paper Walkthrough: From Production Traffic to Post-Training — Folding 200 Internal Apps Into One Self-Hosted LLM ★ MEMBER arXiv:2609.01572
  26. H3-World, Explained — Turning Language Understanding into World Control ★ MEMBER arXiv:2609.01560
  27. LLM-as-a-Judge from Scratch — How AI Grades AI, and Where It Breaks FREE arXiv:2306.05685
  28. Paper Walkthrough: LoopArena — Benchmarking the Model That Steers a Coding Agent ★ MEMBER arXiv:2608.28281
  29. Paper Walkthrough: Normalized Low-Rank Adaptation — Why Normalizing LoRA's Entry Matrix Works ★ MEMBER arXiv:2608.31036
  30. Paper Walkthrough: Designing Qwen3.8-Next — Accuracy, Efficiency and Stability as One Problem ★ MEMBER "arXiv:2608.30320
  31. Paper Walkthrough: PILOT in the Loop — Fixing the Run While It Is Still Running ★ MEMBER arXiv:2608.26530
  32. Paper walkthrough: Puro-2B — pretraining a 2B model from scratch for $6.9K on consumer GPUs ★ MEMBER arXiv:2608.27370
  33. Paper Explained: Repo-To-Skill — Distilling GitHub Repositories Into Skills an AI Can Use ★ MEMBER arXiv:2609.02749
  34. Reward Hacking — Whatever You Measure Is Where It Breaks ★ MEMBER arXiv:1606.06565
  35. Paper Walkthrough: SecOPD — Grading One Token at a Time to Cut Adaptive Prompt Injection by an Order of Magnitude ★ MEMBER arXiv:2608.21500
  36. Self-Improving AI — Self-Play, Co-Evolution, and Generated Curricula ★ MEMBER arXiv:1712.01815
  37. Paper Walkthrough — SMELT: Is Looping the Same Layers Twice Actually a Win When the Budget Is Matched? ★ MEMBER arXiv:2609.01343
  38. Paper Explained: StarHarness — Evolving the Scaffold Instead of the Weights ★ MEMBER arXiv:2608.24804
  39. Paper Walkthrough: StudentSim — Training a Simulator That Is Actually *That* Student ★ MEMBER arXiv:2609.01591
  40. Paper Walkthrough: Training Agents to Evolve with Their Harness ★ MEMBER arXiv:2608.15763
  41. Paper Walkthrough: UI-Venus-2 — Taking Screen-Operating Agents From Benchmarks to Real Work ★ MEMBER arXiv:2609.00028
  42. Paper Walkthrough — UrbanGround: Where MLLM Agents Break Down on a Real Street ★ MEMBER arXiv:2608.27456
  43. Paper Explained: What Makes Good Agentic Data? The ACE Lens ★ MEMBER arXiv:2608.27260
  44. Paper Walkthrough: ZimaBlue — Turning 120,000 Hours of Egocentric Video into Robot Skill ★ MEMBER arXiv:2609.00188
  45. Paper Walkthrough: DART-SD — Training Tool-Calling Agents Without Flattening the Diamond ★ MEMBER arXiv:2608.18524
  46. Paper Walkthrough: PaperGym — Turning One Paper Into a Graded Training Environment for Research Plans ★ MEMBER arXiv:2608.31119
  47. Paper Explained: Agentic Artifact Creation — Where Generation Ends and Construction Begins ★ MEMBER "arXiv:2608.28122
  48. Paper walkthrough: CyberFactory — turning wild CVEs into runnable training problems ★ MEMBER arXiv:2608.23181
  49. Paper Walkthrough — J-Zero: Growing the Challenger, the Solver, and the Judge Together from Zero Data ★ MEMBER arXiv:2608.26582
  50. Paper explained: Self-OPD — an image generator that distills itself, with no teacher ★ MEMBER arXiv:2608.26872
  51. TTPO Explained: Training a Model Mid-Exam, With No Answer Key ★ MEMBER arXiv:2608.27448
  52. Paper Explained: JIT-Agent — A Model That Writes the Agent Harness On Demand ★ MEMBER arXiv:2608.25593
  53. PAWBench Explained — Can Video Generators Get the Odds Right, Not Just the Physics? ★ MEMBER arXiv:2608.27345
  54. When Proxies Stop Being Good Enough — Reading August 2026's Eight Autonomous Driving Papers Together FREE
  55. Paper Walkthrough: Turning Game Development into a Verifiable Trajectory Data Engine — RLHEV and AWoMo ★ MEMBER arXiv:2608.25518
  56. Build Your Own Distillation — Growing a Small Model in 100 Lines ★ MEMBER arXiv:1503.02531
  57. Designing Distillation Data — Deciding What to Ask the Teacher ★ MEMBER arXiv:2501.12948
  58. Evaluating Distilled Models — Is "Close to the Teacher" a Good Metric? ★ MEMBER arXiv:2305.15717
  59. When Distillation Fails — Capacity Gaps and Contagious Overconfidence ★ MEMBER arXiv:1503.02531
  60. Distilling Agents — How to Compress a Long Trajectory ★ MEMBER arXiv:2210.03629
  61. The Math of Distillation — Why Soft Answers Teach More FREE arXiv:1503.02531
  62. A Field Guide to Distillation Recipes — logit, feature, attention, self ★ MEMBER arXiv:1503.02531
  63. Distillation vs. Quantization vs. Pruning — Three Roads to a Smaller Model ★ MEMBER arXiv:1503.02531
  64. Paper Walkthrough: FrontierChallenge — Grading Scientific Work on Whether It Was Actually Delivered ★ MEMBER arXiv:2608.24979
  65. On-Policy Distillation — Learning From What the Student Actually Writes ★ MEMBER arXiv:2306.13649
  66. Paper Walkthrough — WarpSAC: When RL's Safety Rails Become Handcuffs ★ MEMBER arXiv:2608.24479
  67. Paper Walkthrough: VoiceMem — A Left Brain and a Right Brain for Voice Agents, at Zero Added Latency ★ MEMBER arXiv:2608.26005
  68. Evaluating Agents — How Benchmarks and Harnesses Are Built ★ MEMBER arXiv:2310.06770
  69. Designing Agent Memory — Short-Term, Long-Term, Episodic ★ MEMBER arXiv:2310.08560
  70. Alignment, Explained — From RLHF to Constitutional AI ★ MEMBER arXiv:2203.02155
  71. AlphaGo from Scratch — The Marriage of Search and Learning ★ MEMBER "doi:10.1038/nature16961
  72. Paper walkthrough: Apodex 1.1 — scaling agents around completed work ★ MEMBER arXiv:2608.23283
  73. Paper walkthrough: ASI-Bench — peeling away human guidance to measure what AI can do alone ★ MEMBER arXiv:2608.17271
  74. Music and Audio Generation from Scratch — Sound as Tokens ★ MEMBER arXiv:2306.05284
  75. Representing Sound — Mel Spectrograms and Audio Tokens ★ MEMBER arXiv:2106.07447
  76. Encoder or Decoder — The Fork in the Road Between BERT and GPT ★ MEMBER arXiv:1810.04805
  77. Build Your Own Agent Loop — The Minimal Shape of Tool Calling ★ MEMBER arXiv:2210.03629
  78. Build Your Own Diffusion Model — Starting from MNIST ★ MEMBER arXiv:2006.11239
  79. Build Your Own Mini GPT — A Language Model in 300 Lines ★ MEMBER arXiv:1706.03762
  80. Build Your Own BPE Tokenizer — Learning Merge Rules, and Getting Punished by Japanese ★ MEMBER arXiv:1508.07909
  81. Build Your Own Vector DB — From Brute Force to HNSW ★ MEMBER arXiv:1603.09320
  82. Continual Learning and Catastrophic Forgetting — Why Models Can't Just Keep Learning ★ MEMBER arXiv:1612.00796
  83. Unsupervised Learning from Scratch — Clustering and Dimensionality Reduction ★ MEMBER arXiv:1802.03426
  84. DPO and What Came After — The Lineage That Simplified RLHF ★ MEMBER arXiv:2305.18290
  85. Paper Walkthrough: Embodied-Navigator (TAMP-Nav) — Let the VLM Just Point, and Navigation Gets Both Faster and Better ★ MEMBER "arXiv:2608.17512
  86. FlashAttention from Scratch — The Paradox of Doing More Math to Go Faster ★ MEMBER arXiv:2205.14135
  87. Paper Explained: FreeToken — Treating Your Own PC as a Single Elastic Inference Platform ★ MEMBER arXiv:2608.16157
  88. Graph Neural Networks from Scratch — Learning from Connections ★ MEMBER arXiv:1609.02907
  89. Surviving GPU Out-of-Memory — Every Cause, Every Fix ★ MEMBER arXiv:1604.06174
  90. Why Language Models Hallucinate — The Mechanics and What Actually Helps FREE arXiv:2202.03629
  91. Hyperparameter Search — Hunches, Grids, and Bayesian Optimization ★ MEMBER JMLR 2012
  92. A Practical Map of Image Generation — SD, ControlNet, and Applying LoRA FREE arXiv:2112.10752
  93. The ImageNet Moment — The Day Deep Learning Won FREE arXiv:1409.0575
  94. Knowledge Distillation from Scratch — Copying a Big Model into a Small One ★ MEMBER arXiv:1503.02531
  95. LLM Evaluation from Scratch — Reading Benchmarks and the Contamination Problem ★ MEMBER arXiv:2009.03300
  96. LLM Serving from Scratch — vLLM, Continuous Batching, and Not Letting the GPU Idle FREE arXiv:2309.06180
  97. How Long-Context LLMs Work — From RoPE Interpolation to Ring Attention ★ MEMBER arXiv:2306.15595
  98. Multi-Agent Design Patterns — Division, Debate, Verification ★ MEMBER arXiv:2305.14325
  99. Evaluating RAG in Practice — Turning “Seems Better” Into a Number ★ MEMBER arXiv:2309.15217
  100. Scaling Skepticism — A Genealogy of the "Just Make It Bigger" Critique ★ MEMBER arXiv:2001.08361
  101. The Mathematics of Diffusion — Generation Seen Through Scores and SDEs ★ MEMBER arXiv:1907.05600
  102. Self-Supervised Learning — The Day Unlabeled Data Became an Asset ★ MEMBER arXiv:2002.05709
  103. Structured Output and Constrained Decoding — How to Stop an LLM from Breaking Your JSON ★ MEMBER arXiv:2307.09702
  104. Paper Walkthrough: SWE-bench Science — Can Coding Agents Fix Scientific Code? ★ MEMBER arXiv:2608.19799
  105. Symbolic vs. Connectionist — Where a 60-Year Argument Stands Today ★ MEMBER
  106. Test-Time Scaling — How Models Get Better by Thinking Longer ★ MEMBER arXiv:2201.11903
  107. The State of 3D Generation — From NeRF to Gaussian Splatting ★ MEMBER arXiv:2003.08934
  108. Diagnosing Broken Training — Telling Divergence, NaN, and Plateaus Apart FREE arXiv:1211.5063
  109. Video Understanding from Scratch — From a Pile of Frames to a Sense of Time ★ MEMBER arXiv:2102.05095
  110. Paper Walkthrough: WeMM-Embedding — Putting Text, Images and Video on One Ruler ★ MEMBER arXiv:2608.24053
  111. Building a Dataset in Practice — Collect, Clean, Blend ★ MEMBER arXiv:2212.10560
  112. Document AI and OCR Today — How an LLM Ends Up Reading Your Invoices ★ MEMBER arXiv:1912.13318
  113. GraphRAG from Scratch — Where Knowledge Graphs Meet Retrieval ★ MEMBER arXiv:2404.16130
  114. Building a Pretraining Corpus — From Web Sludge to Textbook Quality ★ MEMBER arXiv:2101.00027
  115. The Transformer, End to End — One Token's Journey from Embedding to Output FREE arXiv:1706.03762
  116. Speech Synthesis from Scratch — From Text to a Voice ★ MEMBER arXiv:1609.03499
  117. Activation Functions from Scratch — Why Nonlinearity Is Non-Negotiable FREE arXiv:1502.01852
  118. Decision Trees and Gradient Boosting — Still the Champion on Tabular Data FREE arXiv:1603.02754
  119. Time-Series Anomaly Detection — The Math Behind the Alerts ★ MEMBER arXiv:0710.3742
  120. Weight Initialization and Regularization — What Lets Training Start, and What Keeps It Going ★ MEMBER PMLR v9
  121. Paper Deep Dive — Large Discovery Models: giving an LLM a value signal for what to try next ★ MEMBER arXiv:2608.15669
  122. Paper Explained: Agentic ESOpt — Drop Backprop, Jiggle the Weights, and Train Long-Horizon LLM Agents ★ MEMBER arXiv:2608.17310
  123. A Field Guide to Attention Variants — MQA, GQA, Sliding Windows, Linear Attention ★ MEMBER arXiv:1911.02150
  124. Paper Explained: Co-RL — Reasoning Without Labels, Emerging From a Diverse Cohort ★ MEMBER arXiv:2608.17253
  125. Paper Explainer: Why Agent Skills Work — and Where They Break ★ MEMBER arXiv:2608.14036
  126. Paper Explained: EnvHarness — Reshaping an Agent's Training World Without Rebuilding It ★ MEMBER arXiv:2608.19880
  127. Paper Walkthrough — FACET: Grounding Instruction, Environment, Solution and Verifier in One Executable State ★ MEMBER arXiv:2608.18580
  128. Flow Matching from Scratch — What Came After Diffusion, and Why It Goes Straight ★ MEMBER arXiv:2210.02747
  129. The Rise and Fall of GANs — An Invention Trained by Rivalry, and Why Diffusion Won ★ MEMBER arXiv:1406.2661
  130. Learning Rate Schedules — Why Warmup and Why Cosine ★ MEMBER arXiv:1706.03762
  131. Mixed Precision Training — Going Faster in fp16/bf16/fp8 Without Breaking ★ MEMBER arXiv:1710.03740
  132. A History of Normalization Layers — From BatchNorm to RMSNorm ★ MEMBER arXiv:1502.03167
  133. Paper Walkthrough: OmniScientist — An AI Scientist That Actually Looks at the Raw Data ★ MEMBER arXiv:2608.13558
  134. Recommenders and Embeddings — Same Math as RAG, Different Goal ★ MEMBER arXiv:1205.2618
  135. CFG and Samplers — What the "Strength" Knob in Generative AI Really Does ★ MEMBER arXiv:2207.12598
  136. Paper Explainer: SemaPLC — The Agent That Isn't Allowed to Say "Done" ★ MEMBER "arXiv:2608.18565
  137. Paper walkthrough: StateM — 95.3% on Terminal-Bench 2.1 and a USD 15 run, without touching a single weight ★ MEMBER "arXiv:2608.15089
  138. Do Transformers Actually Work on Time Series? — The Argument and the Practical Answer ★ MEMBER arXiv:2205.13504
  139. Tokenizers from Scratch — The Unit an LLM Cuts the World Into FREE arXiv:1508.07909
  140. Paper walkthrough: Zetta ζ — a robot harness that repairs itself mid-execution, with the policy frozen ★ MEMBER arXiv:2608.16590
  141. Paper Walkthrough: SA-MRPO — Stop Studying the Subject You've Already Aced ★ MEMBER "arXiv:2608.16072
  142. Paper Walkthrough: Can Anything Catch a Fake Crisis Video? — What RA-Bench Found ★ MEMBER "arXiv:2608.14391
  143. Paper Deep-Dive: ABSeeker — Training Long-Horizon Search Agents by Grading Each Step Backward from the Answer ★ MEMBER arXiv:2608.05102
  144. The CNN Family Tree — From AlexNet to ResNet and EfficientNet ★ MEMBER arXiv:1409.1556
  145. Paper Explained: Co-Evolution in Agentic Systems — Three Stages Toward Self-Directed Evolution ★ MEMBER arXiv:2608.10299
  146. Paper Walkthrough: ComBodied Agents — Moving an Agent's Target from Software and Matter to the Person ★ MEMBER arXiv:2608.10915
  147. Embeddings from Scratch — from word2vec Intuition to Contextual Embeddings FREE arXiv:1301.3781
  148. End-to-End Driving from Scratch — Perception to Control in a Single Network ★ MEMBER arXiv:1604.07316
  149. Paper Deep-Dive: Frontis-MA1 — Training the AI That Builds AI: One Step Toward Recursive Self-Improvement in ML Engineering ★ MEMBER arXiv:2607.28568
  150. Paper Walkthrough: Interpretable MEG Decoding of Perceived Speech — Reading the Decoder's Weights as a Brain Map ★ MEMBER arXiv:2608.01481
  151. Paper Explained: LongHorizon-Harness — Long-Horizon Agent Tasks Are a State-Management Problem, Not an Execution Problem ★ MEMBER arXiv:2608.01964
  152. Paper Walkthrough: Mental World Modeling — A World Model That Advances Minds, Not Just Physics ★ MEMBER arXiv:2607.27201
  153. Paper Walkthrough: MerchantBench — Can an LLM Agent Run an Online Store for a Year? Why It Earns Only 27.3% of What Humans Do ★ MEMBER arXiv:2607.28956
  154. Paper Walkthrough: Metis — A 'Memory Foundation Model' That Moves Agent Memory Inside the Model ★ MEMBER arXiv:2607.26760
  155. Object Detection from Scratch (from YOLO to DETR) ★ MEMBER arXiv:2005.12872
  156. Paper Walkthrough: No Gold Answers, No Stronger Teacher — How u-OPSD Distills From Its Own Majority Vote ★ MEMBER arXiv:2608.06296
  157. Paper Explained: OSReward — Can You Trust the AI That Grades AI? Remeasuring Rewards for Computer-Use Agents ★ MEMBER arXiv:2607.28609
  158. Paper Walkthrough: PhiZero — A World Model That Reasons in a Language of Physics Before It Renders ★ MEMBER arXiv:2607.28624
  159. Paper Walkthrough: Qwen-UI-Agent — How Alibaba Built a GUI Agent That Works on Real Phones and PCs ★ MEMBER arXiv:2607.28227
  160. Paper Deep-Dive: Recursive Synthesis — Extending Verified Tasks Into 40,000 Long-Horizon Terminal Problems ★ MEMBER arXiv:2608.05466
  161. Scaling Laws from Scratch — Why Making Models Bigger Makes Them Smarter (and When It Doesn't) ★ MEMBER arXiv:2203.15556
  162. Paper Walkthrough: SwanTale — Designing Voices from Words Alone, with Speech and Sound in One Waveform ★ MEMBER arXiv:2608.02023
  163. Paper Walkthrough: ToolArtist — Search, Draw, or Redraw? The Image Agent That Decides for Itself ★ MEMBER arXiv:2608.04436
  164. Paper Walkthrough: TurboVLA — Kick the LLM Out of the Loop and Run a Robot Policy at 32 Hz on an RTX 4090 with Under 1 GB of VRAM ★ MEMBER arXiv:2607.27205
  165. Paper Explained: Video-DeepResearch — Agents That Watch a Video, Then Chase Down Every Lead ★ MEMBER arXiv:2608.03979
  166. How VLMs Came Together — Wiring a Vision Encoder into an LLM ★ MEMBER arXiv:2103.00020
  167. Paper Deep-Dive: AgentOPSD — Finding the Turn That Won the Game with Recursive Bayesian Belief Updates ★ MEMBER arXiv:2608.05987
  168. Paper Walkthrough: Alpamayo — NVIDIA's Reasoning Model for Autonomous Driving ★ MEMBER arXiv:2511.00088
  169. Paper Deep Dive: AskChem — Changing the Unit of Search from Papers to Provenance-Carrying Claims ★ MEMBER arXiv:2607.28618
  170. Paper explained: BDH-CQ — an AI that thinks without words. Recurrent memory plus latent reasoning resets ARC's cost frontier ★ MEMBER arXiv:2608.09888
  171. BEV Representations From Scratch — Fusing Multiple Cameras Into One Top-Down Map ★ MEMBER "arXiv:2008.05711
  172. Paper Walkthrough: CodeNib — A Multi-View Data System That Serves Repository Context to Coding Agents ★ MEMBER arXiv:2607.25431
  173. Paper Walkthrough: DAPD — Breaking the Teacher's "Cheat-Sheet Illusion" in Distillation with Dual Anchors ★ MEMBER arXiv:2608.01735
  174. Paper Walkthrough: EnvACE — Agents That Rehearse the World Instead of Calling It ★ MEMBER arXiv:2608.06197
  175. The Gemma Family from Scratch — Lineage, Inventions, and Where It Fits ★ MEMBER arXiv:2403.08295
  176. Paper Walkthrough: Macaron-V1 — A Frozen Base plus a Mixture of LoRAs, Built to Keep Learning After Launch ★ MEMBER arXiv:2608.09819
  177. Mixture of Experts (MoE) from Scratch — Routing and Load Balancing in the Switch Transformer ★ MEMBER arXiv:2101.03961
  178. Speculative Decoding from Scratch — How a Tiny Draft Model Speeds Up an LLM Without Changing a Single Output ★ MEMBER arXiv:2211.17192
  179. Paper Walkthrough: SWE-Bench ProMax — Measuring What Coding Agents Can Really Do with Large-Scale, Multilingual Refactoring ★ MEMBER arXiv:2608.09802
  180. Paper Walkthrough: The Personalization Mirage — LLMs Invent a Version of You, and Their Self-Reports Point the Wrong Way ★ MEMBER "arXiv:2608.04570
  181. Paper Walkthrough: WorldClaw — Agents That Build Walkable, Editable 3D Open Worlds from a Single Sentence ★ MEMBER arXiv:2608.05248
  182. Paper Deep Dive — CLIP: Putting Words and Images on One Map ★ MEMBER arXiv:2103.00020
  183. Diffusion Models from the Ground Up — Add Noise, Then Subtract It ★ MEMBER arXiv:2006.11239
  184. Instruction Tuning and RLHF from Scratch — How a Model Learns to Follow Orders ★ MEMBER arXiv:2203.02155
  185. LLM Agents from Scratch — Designing the Tool-Use Loop ★ MEMBER arXiv:2210.03629
  186. The Science of Prompt Engineering — What Is Proven and What Is Folklore FREE arXiv:2201.11903
  187. Paper Deep-Dive: Why Whisper Is Robust — Large-Scale Weak Supervision ★ MEMBER arXiv:2212.04356
  188. Paper Deep Dive — ViT: Treating an Image Like a Sentence ★ MEMBER arXiv:2010.11929
  189. Positional Encoding from Scratch — From Absolute Positions to RoPE FREE arXiv:1706.03762
  190. Paper Deep Dive — Attention Is All You Need: What Dropping Recurrence Actually Proved ★ MEMBER arXiv:1706.03762
  191. Paper Deep Dive — LoRA: Low-Rank Adaptation of Large Language Models: Why Low Rank Is Enough ★ MEMBER arXiv:2106.09685

④ Grab self-attention with your hands

Words sit around a round table and attention flies between them as light. Turn off the √d_k division and watch the distribution collapse; flip the causal mask and watch the future go dark.

FIG 1The math is exactly the scaled dot-product attention from the paper. Only the embeddings are toys: an 8-dimensional space designed by hand so the patterns are legible — not weights from a trained model.