Transcription of Multiscale Vision Transformers - arXiv
{{id}} {{{paragraph}}}
Multiscale Vision TransformersHaoqi Fan*, 1Bo Xiong*, 1 Karttikeya Mangalam*, 1, 2 Yanghao Li*, 1 Zhicheng Yan1 Jitendra Malik1, 2 Christoph Feichtenhofer*, 11 Facebook AI Research2UC BerkeleyAbstractWe present Multiscale Vision Transformers (MViT) forvideo and image recognition, by connecting the seminal ideaof Multiscale feature hierarchies with transformer Transformers have several channel-resolutionscale stages. Starting from the input resolution and a smallchannel dimension, the stages hierarchically expand thechannel capacity while reducing the spatial resolution. Thiscreates a Multiscale pyramid of features with early lay-ers operating at high spatial resolution to model simplelow-level visual information, and deeper layers at spatiallycoarse, but complex, high-dimensional features.
models for computer vision. Based on their studies of cat and monkey visual cortex, Hubel and Wiesel [55] developed a hierarchical model of the visual pathway with neurons in lower areas such as V1 responding to features such as oriented edges and bars, and in higher areas to more spe-cific stimuli. Fukushima proposed the Neocognitron [32], a
Domain:
Source:
Link to this page:
Please notify us if you found a problem with this document:
{{id}} {{{paragraph}}}