JA EN
Glossary › mt-bench

GLOSSARY

mt-bench

appears in 2 paper titles

Definition

A small, hand-written benchmark of multi-turn conversations, scored by a strong LLM judge, released by the group behind Chatbot Arena. Its point is to probe what single-turn benchmarks miss: whether a model keeps context across turns and follows up on earlier instructions. Questions span categories such as writing, reasoning, math, and coding. Being judge-scored, its numbers inherit the biases of LLM-as-a-judge evaluation.

Explainers using this term