JA EN

#mllm

3 articles

01 ·★ MEMBER·PAPER·10 min read Paper Walkthrough — Beyond Retrieval: LatentStream Turns Retrieved Video Into Latent Memory For never-ending video streams, LatentStream stops appending retrieved evidence as extra context and instead internalizes it into fixed-length latent memory tokens. A ground-up walkthrough of its hierarchical memory, latent evolution, and confidence-driven test-time optimization. 02 ·Agents·★ MEMBER·PAPER·10 min read Paper Walkthrough — UrbanGround: Where MLLM Agents Break Down on a Real Street Drop an MLLM agent into a real-scale replica of Hong Kong built from territory-wide 3D geospatial data. Visual recognition clears 90%, orientation sits near 40%, long-range navigation is close to 0%. A walkthrough of the benchmark that measures the gap between seeing and moving. 03 ·Large Language Models·★ MEMBER·PAPER·8 min read Paper Walkthrough: Can Anything Catch a Fake Crisis Video? — What RA-Bench Found Sixteen thousand AI videos, each continuing from the real first frame of a genuine disaster or war clip, put against seven classical detectors, ten zero-shot multimodal models and two purpose-built fine-tunes. None of them generalized. One model turned out to be reading timestamps rather than pixels, and a lap through a social feed drops fake recall to 1.4%.