Example: barber
Training data-efficient image transformers & distillation ...

Training data-efficient image transformers & distillation ...

Back to document page

transformers “do not generalize well when trained on insufficient amounts of data”, and the training of these models involved extensive computing resources. In this paper, we train a vision transformer on a single 8-GPU node in two to three days (53 hours of pre-training, and optionally 20 hours of fine-tuning)

  Three, Transformers

Download Training data-efficient image transformers & distillation ...


Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Other abuse

Advertisement

Related search queries