JA EN

#video-understanding

3 articles

01 ·★ MEMBER·PAPER·10 min read Paper Walkthrough — Beyond Retrieval: LatentStream Turns Retrieved Video Into Latent Memory For never-ending video streams, LatentStream stops appending retrieved evidence as extra context and instead internalizes it into fixed-length latent memory tokens. A ground-up walkthrough of its hierarchical memory, latent evolution, and confidence-driven test-time optimization. 02 ·★ MEMBER·PAPER·11 min read Paper Walkthrough: TLive-Omni — An Omni-Modal Model That Watches, Listens, and Tells You the Timestamp A from-scratch walkthrough of TLive-Omni, an omni-modal understanding model built for e-commerce live streaming: Per-vGrid timestamped audio-video interleaving, a three-stage SFT recipe, and Faithful-RFT, which rewards answer faithfulness instead of visible reasoning — grounded strictly in the paper's own numbers. 03 ·★ MEMBER·PAPER·8 min read Paper Walkthrough: GST-Bench — Can VLMs Build a Global Map of a Scene from Video? Show a VLM a walkthrough video of a house, then ask 'from where you're standing now, which way is the sofa?' — even the strongest model scores barely half of what humans do. A walkthrough of GST-Bench from ByteDance Seed: the shortcut-proof benchmark design, results across 22 models, and the training data that closed 27 points of the gap.