JA EN
Glossary › bevformer

GLOSSARY

bevformer

appears in 1 paper titles

Definition

A Transformer that builds a bird's-eye-view representation from multiple vehicle cameras. Learnable BEV queries use spatial cross-attention to pull features from whichever camera regions project onto each grid location, and temporal attention lets each frame's BEV features attend to the previous frame's, which is where motion and occlusion handling come from. It fuses views and time without an explicit depth-prediction stage, and remains a standard camera-only baseline.