Glossary › mt-bench
GLOSSARY
mt-bench
appears in 2 paper titles
Definition
A small, hand-written benchmark of multi-turn conversations, scored by a strong LLM judge, released by the group behind Chatbot Arena. Its point is to probe what single-turn benchmarks miss: whether a model keeps context across turns and follows up on earlier instructions. Questions span categories such as writing, reasoning, math, and coding. Being judge-scored, its numbers inherit the biases of LLM-as-a-judge evaluation.
Explainers using this term
- LLM-as-a-Judge from Scratch — How AI Grades AI, and Where It BreaksJudging LLM-as-a-Judge with MT-Bench and Chatbot Arena