Example: bachelor of science
Training data-efficient image transformers & distillation ...

Training data-efficient image transformers & distillation ...

Back to document page

a pre-training phase on a large volume of curated data is required for the learned transformer to be effective. In our paper we achieve a strong perfor-mance without requiring a large training dataset, i.e., with Imagenet1k only. The Transformer architecture, introduced by Vaswani et al. [52] for machine

  Phases, Transformers

Download Training data-efficient image transformers & distillation ...


Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Other abuse

Advertisement

Related search queries