Glossary › evaluation
GLOSSARY
evaluation
appears in 10 paper titles
Definition
Measuring model quality against defined tasks and metrics. The credibility of an evaluation rests on unglamorous details: whether test data leaked into training, whether the metric tracks what you actually care about, and whether the gap between two systems exceeds run-to-run variance. Benchmark numbers alone rarely predict deployment usefulness, which is why human and task-grounded evaluations sit alongside them.
Explainers using this term
- Paper Walkthrough: EarlyEval — Making Agent Evaluation Cheaper by Stopping EarlyEarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction
- Paper Walkthrough: Designing Qwen3.8-Next — Accuracy, Efficiency and Stability as One ProblemOn the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability
- Paper Explained: Agentic Artifact Creation — Where Generation Ends and Construction BeginsAgentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities