JA EN

#mmlu

1 articles

01 ·Large Language Models·★ MEMBER·PAPER·11 min read LLM Evaluation from Scratch — Reading Benchmarks and the Contamination Problem A guide to reading the bar charts in model release posts with the right kind of suspicion. Covers how the scoring method alone moves MMLU numbers, the error bar that comes from question count, why public benchmarks get contaminated structurally rather than accidentally, the Bradley-Terry model behind Chatbot Arena and where it breaks, and the three biases in LLM-as-a-judge.