JA EN

#agreement

1 articles

01 ·Evaluation & Judging·FREE·PAPER·10 min read LLM-as-a-Judge from Scratch — How AI Grades AI, and Where It Breaks A ground-up guide to using one model to grade another. Covers reading a verdict as a probability distribution, the three recurring biases (position, verbosity, self-enhancement), why pairwise comparison cost grows quadratically, and how to validate the judge itself against human labels.