JA EN

#spatial-reasoning

2 articles

01 ·Agents·★ MEMBER·PAPER·10 min read Paper Walkthrough — UrbanGround: Where MLLM Agents Break Down on a Real Street Drop an MLLM agent into a real-scale replica of Hong Kong built from territory-wide 3D geospatial data. Visual recognition clears 90%, orientation sits near 40%, long-range navigation is close to 0%. A walkthrough of the benchmark that measures the gap between seeing and moving. 02 ·★ MEMBER·PAPER·8 min read Paper Walkthrough: GST-Bench — Can VLMs Build a Global Map of a Scene from Video? Show a VLM a walkthrough video of a house, then ask 'from where you're standing now, which way is the sofa?' — even the strongest model scores barely half of what humans do. A walkthrough of GST-Bench from ByteDance Seed: the shortcut-proof benchmark design, results across 22 models, and the training data that closed 27 points of the gap.