Glossary › judging
GLOSSARY
judging
appears in 2 paper titles
Definition
Scoring or ranking outputs. For open-ended tasks where no single reference answer exists, string-overlap metrics break down, so evaluation falls back on human raters or on a strong model acting as grader. Because the verdict depends entirely on what the judge was told to value, the credibility of any judging-based result rests on the rubric, the instructions given, and whether ties and refusals were handled consistently.
Explainers using this term
- LLM-as-a-Judge from Scratch — How AI Grades AI, and Where It BreaksJudging LLM-as-a-Judge with MT-Bench and Chatbot Arena