Example: barber
Training data-efficient image transformers & distillation ...

Training data-efficient image transformers & distillation ...

Back to document page

normalized with a softmax function to obtain kweights. The output of the attention is the weighted sum of a set of kvalue vectors (packed into V 2Rk d). For a sequence of Nquery vectors (packed into Q2RN d), it produces an output matrix (of size N d):

Download Training data-efficient image transformers & distillation ...


Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Other abuse

Advertisement

Related search queries