PDF4PRO ⚡AMP

Modern search engine that looking for books and documents around the web

Example: confidence

Training data-efficient image transformers & distillation ...

Training data-efficient image transformers & distillation through attentionHugo Touvron?, Matthieu Cord Matthijs Douze?Francisco Massa?Alexandre Sablayrolles?Herv e J egou??Facebook AI Sorbonne UniversityAbstractRecently, neural networks purely based on attention were shown to ad-dress image understanding tasks such as image classification. These high-performing vision transformers are pre-trained with hundreds of millionsof images using a large infrastructure, thereby limiting their this work, we produce competitive convolution-free transformers bytraining on Imagenet only. We train them on a single computer in less than3 days. Our reference vision transformer (86M parameters) achieves top-1accuracy of (single-crop) on ImageNet with no external importantly, we introduce a teacher-student strategy specific totransformers.

Kernel [34] and Split-Attention Networks [61] exploit mechanism akin to trans-formers self-attention (SA) mechanism. Knowledge Distillation (KD), introduced by Hinton et al. [24], refers to the training paradigm in which a student model leverages “soft” labels coming from a strong teacher network. This is the output vector of the teacher ...

Loading..

Tags:

  Attention

Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Spam in document Broken preview Other abuse

Transcription of Training data-efficient image transformers & distillation ...

Related search queries