Example: confidence
Improved Multiscale Vision Transformers for Classification ...

Improved Multiscale Vision Transformers for Classification ...

Back to document page

Without bells-and-whistles, MViT has state-of-the-art per-formance in 3 domains: 88.8% accuracy on ImageNet clas-sification, 56.1 APbox on COCO object detection as well as 86.1% on Kinetics-400 video classification. Code and models will be made publicly available. 1. Introduction Designing architectures for different visual recognition

  Whistle

Download Improved Multiscale Vision Transformers for Classification ...


Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Other abuse

Advertisement

Related search queries