Glossary › timesformer
GLOSSARY
timesformer
appears in 1 paper titles
Definition
A convolution-free video model that extends the Transformer along the time axis. Its key ingredient, divided space-time attention, applies temporal attention and spatial attention in separate steps instead of attending over every patch in every frame at once, which keeps cost tractable as clips get longer. It was one of the early demonstrations that attention alone, without 3D convolutions, is sufficient for video recognition.
Explainers using this term
- Video Understanding from Scratch — From a Pile of Frames to a Sense of TimeIs Space-Time Attention All You Need for Video Understanding? (TimeSformer)