Multiscale Vision Transformers - arXiv
Multiscale Vision TransformersHaoqi Fan*, 1Bo Xiong*, 1Karttikeya Mangalam*, 1, 2Yanghao Li*, 1Zhicheng Yan1Jitendra Malik1, 2Christoph Feichtenhofer*, 11Facebook AI Research2UC BerkeleyAbstractWe present Multiscale Vision Transformers (MViT) forvideo and image recognition, by connecting the seminal ideaof Multiscale feature hierarchies with transformer Transformers have several channel-resolutionscale stages. Starting from the input resolution and a smallchannel dimension, the stages hierarchically expand thechannel capacity while reducing the spatial resolution. Thiscreates a Multiscale pyramid of features with early lay-ers operating at high spatial resolution to model simplelow-level visual information, and deeper layers at spatiallycoarse, but complex, high-dimensional features.
models for computer vision. Based on their studies of cat and monkey visual cortex, Hubel and Wiesel [55] developed a hierarchical model of the visual pathway with neurons in lower areas such as V1 responding to features such as oriented edges and bars, and in higher areas to more spe-cific stimuli. Fukushima proposed the Neocognitron [32], a
Download Multiscale Vision Transformers - arXiv
Information
Domain:
Source:
Link to this page:
Please notify us if you found a problem with this document: