PDF4PRO ⚡AMP

Modern search engine that looking for books and documents around the web

Example: marketing

Multiscale Vision Transformers - arXiv

Back to document page

Multiscale Vision TransformersHaoqi Fan*, 1Bo Xiong*, 1Karttikeya Mangalam*, 1, 2Yanghao Li*, 1Zhicheng Yan1Jitendra Malik1, 2Christoph Feichtenhofer*, 11Facebook AI Research2UC BerkeleyAbstractWe present Multiscale Vision Transformers (MViT) forvideo and image recognition, by connecting the seminal ideaof Multiscale feature hierarchies with transformer Transformers have several channel-resolutionscale stages. Starting from the input resolution and a smallchannel dimension, the stages hierarchically expand thechannel capacity while reducing the spatial resolution. Thiscreates a Multiscale pyramid of features with early lay-ers operating at high spatial resolution to model simplelow-level visual information, and deeper layers at spatiallycoarse, but complex, high-dimensional features. We eval-uate this fundamental architectural prior for modeling thedense nature of visual signals for a variety of video recog-nition tasks where it outperforms concurrent Vision trans-formers that rely on large scale external pre-training andare 5-10 more costly in computation and parameters.

Multiscale Vision Transformers learn a hierarchy from dense (in space) and simple (in channels) to coarse and complex features. Several resolution-channel scale stages progressively increase the channel capacity of the intermediate latent sequence while reducing its length and thereby spatial resolution.

  Vision, Transformers, Multiscale, Multiscale vision transformers

Download Multiscale Vision Transformers - arXiv


Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Spam in document Broken preview Other abuse

Related search queries