Glossary › real-world
GLOSSARY
real-world
appears in 6 paper titles
Definition
Marks a claim about deployment conditions rather than curated benchmarks: messy inputs, distribution shift, missing data, latency budgets, adversarial users. It matters because benchmark gains routinely fail to transfer — the test set was clean and the deployment is not. When a paper claims real-world results, the details worth checking are scale, duration, and whether anything was measured beyond a controlled pilot.
Explainers using this term
- A Map of Agent Benchmarks — What SWE-bench, GAIA, and OSWorld Actually MeasureSWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Evaluating Agents — How Benchmarks and Harnesses Are BuiltSWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- LLM Security — Prompt Injection and How to Actually Defend Against ItNot What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection