Example: quiz answers

Video Swin Transformer

model pre-trained on a large-scale image dataset. With a model pre-trained on ImageNet-21K, we interestingly find that the learning rate of the backbone architecture needs to be smaller (e.g. 0.1 ) than that of the head, which is randomly initialized. As a …

Tags:

  Large, Scale, Imagenet

Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Other abuse

Advertisement

Related search queries