Transcription of Training data-efficient image transformers & distillation ...
{{id}} {{{paragraph}}}
Training data-efficient image transformers & distillation through attentionHugo Touvron?, Matthieu Cord Matthijs Douze?Francisco Massa?Alexandre Sablayrolles?Herv e J egou??Facebook AI Sorbonne UniversityAbstractRecently, neural networks purely based on attention were shown to ad-dress image understanding tasks such as image classification. These high-performing vision transformers are pre-trained with hundreds of millionsof images using a large infrastructure, thereby limiting their this work, we produce competitive convolution-free transformers bytraining on Imagenet only. We train them on a single computer in less than3 days. Our reference vision transformer (86M parameters) achieves top-1accuracy of (single-crop) on ImageNet with no external importantly, we introduce a teacher-student strategy specific totransformers. It relies on a distillation token ensuring that the studentlearns from the teacher through attention.
image. For example, let us consider image with a “cat” label that represents a large landscape and a small cat in a corner. If the cat is no longer on the crop of the data augmentation it implicitly changes the label of the image. KD can transfer inductive biases [1] in a soft way in a student model using a teacher
Domain:
Source:
Link to this page:
Please notify us if you found a problem with this document:
{{id}} {{{paragraph}}}