View article

[PDF] from thecvf.com

Mvitv2: Improved multiscale vision transformers for classification and detection

Authors

Yanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam, Bo Xiong, Jitendra Malik, Christoph Feichtenhofer

Publication date

2022

Conference

Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Pages

4804-4814

Description

In this paper, we study Multiscale Vision Transformers (MViTv2) as a unified architecture for image and video classification, as well as object detection. We present an improved version of MViT that incorporates decomposed relative positional embeddings and residual pooling connections. We instantiate this architecture in five sizes and evaluate it for ImageNet classification, COCO detection and Kinetics video recognition where it outperforms prior work. We further compare MViTv2s' pooling attention to window attention mechanisms where it outperforms the latter in accuracy/compute. Without bells-and-whistles, MViTv2 has state-of-the-art performance in 3 domains: 88.8% accuracy on ImageNet classification, 58.7 boxAP on COCO object detection as well as 86.1% on Kinetics-400 video classification. Code and models are available at https://github. com/facebookresearch/mvit.

Total citations

Cited by 609

20212022202320242 77 299 229

Scholar articles

Mvitv2: Improved multiscale vision transformers for classification and detection

Y Li, CY Wu, H Fan, K Mangalam, B Xiong, J Malik… - Proceedings of the IEEE/CVF conference on computer …, 2022