JA EN

#autonomy

1 articles

01 ·Agents·★ MEMBER·PAPER·10 min read Paper walkthrough: ASI-Bench — peeling away human guidance to measure what AI can do alone ASI-Bench keeps the research goal, data and grading fixed while stripping away human methodological guidance one layer at a time. Average scores fall 50.91 → 29.10 → 26.62, and the place where things break is not method selection.