PaperLens
紙
Students
Professional
JA
EN
◐
Sign in with Google
Sign in
Read
Home
Close reading
New
Textbook
Go deeper
Learn
Lab
Landscape
Contributors
Glossary
You
Search
All-access
My Page
#mllm
3 articles
01
2026-09-08
·
★ MEMBER
·
PAPER
·
10 min read
Paper Walkthrough — Beyond Retrieval: LatentStream Turns Retrieved Video Into Latent Memory
For never-ending video streams, LatentStream stops appending retrieved evidence as extra context and instead internalizes it into fixed-length latent memory tokens. A ground-up walkthrough of its hierarchical memory, latent evolution, and confidence-driven test-time optimization.
02
2026-09-03
·
Agents
·
★ MEMBER
·
PAPER
·
10 min read
Paper Walkthrough — UrbanGround: Where MLLM Agents Break Down on a Real Street
Drop an MLLM agent into a real-scale replica of Hong Kong built from territory-wide 3D geospatial data. Visual recognition clears 90%, orientation sits near 40%, long-range navigation is close to 0%. A walkthrough of the benchmark that measures the gap between seeing and moving.
03
2026-08-20
·
Large Language Models
·
★ MEMBER
·
PAPER
·
8 min read
Paper Walkthrough: Can Anything Catch a Fake Crisis Video? — What RA-Bench Found
Sixteen thousand AI videos, each continuing from the real first frame of a genuine disaster or war clip, put against seven classical detectors, ten zero-shot multimodal models and two purpose-built fine-tunes. None of them generalized. One model turned out to be reading timestamps rather than pixels, and a lap through a social feed drops fake recall to 1.4%.