JA EN

#mcts

1 articles

01 ·Agents·★ MEMBER·PAPER·9 min read AlphaGo from Scratch — The Marriage of Search and Learning Starting from why Go was considered unsolvable for so long, this piece unpacks how the policy network, the value network, Monte Carlo tree search and self-play each cover the others' weaknesses — with the formulas and the code. It closes with what this design handed down to inference-time compute in LLMs.