JA EN
Glossary › timesformer

GLOSSARY

timesformer

appears in 1 paper titles

Definition

A convolution-free video model that extends the Transformer along the time axis. Its key ingredient, divided space-time attention, applies temporal attention and spatial attention in separate steps instead of attending over every patch in every frame at once, which keeps cost tractable as clips get longer. It was one of the early demonstrations that attention alone, without 3D convolutions, is sufficient for video recognition.

Explainers using this term