Transcription of Training data-efficient image transformers & distillation ...
{{id}} {{{paragraph}}}
Training data-efficient image transformers & distillation through attentionHugo Touvron?, Matthieu Cord Matthijs Douze?Francisco Massa?Alexandre Sablayrolles?Herv e J egou??Facebook AI Sorbonne UniversityAbstractRecently, neural networks purely based on attention were shown to ad-dress image understanding tasks such as image classification. These high-performing vision transformers are pre-trained with hundreds of millionsof images using a large infrastructure, thereby limiting their this work, we produce competitive convolution-free transformers bytraining on Imagenet only.
normalized with a softmax function to obtain kweights. The output of the attention is the weighted sum of a set of kvalue vectors (packed into V 2Rk d). For a sequence of Nquery vectors (packed into Q2RN d), it produces an output matrix (of size N d):
Domain:
Source:
Link to this page:
Please notify us if you found a problem with this document:
{{id}} {{{paragraph}}}