PaperLens
紙
Students
Professional
JA
EN
◐
Sign in with Google
Sign in
Read
Home
Close reading
New
Textbook
Go deeper
Learn
Lab
Landscape
Contributors
Glossary
You
Search
All-access
My Page
#diffusion-transformer
3 articles
01
2026-09-03
·
Agents
·
★ MEMBER
·
PAPER
·
12 min read
Paper Walkthrough: ZimaBlue — Turning 120,000 Hours of Egocentric Video into Robot Skill
A ground-up walkthrough of the World Action Model that converts 120,000 hours of action-free egocentric video into robot control: a three-stage curriculum, a 100-D unified action interface, and an asynchronous Slow-Fast pair that takes zero-shot success from 36.1% to 77.8% at a 33 ms control loop.
02
2026-09-02
·
★ MEMBER
·
PAPER
·
14 min read
Paper Walkthrough: DreamX-Creator — Making Sound and Picture Together in 7B, Then Finishing at 2K in One Step
A ground-up walkthrough of a 7B model that denoises audio and video inside one generative process: the gated cross-modal attention, the modality-aware reinforcement learning, the one-step 2K refiner — and the unusually heavy caveats the authors put on their own results.
03
2026-09-01
·
★ MEMBER
·
PAPER
·
13 min read
Paper Walkthrough: GameWAM — Generating the Next Frame and the Next Keystroke Together
A ground-up walkthrough of the first World–Action Model for native closed-loop game and GUI control: how it plans 16 actions but commits only 8, and how low-frequency noise in the sampled action source quietly spins the camera.