Transcription of COSFORMER : RETHINKING SOFTMAX IN ATTENTION
{{id}} {{{paragraph}}}
Published as a conference paper at ICLR 2022. COS F ORMER : R ETHINKING S OFTMAX IN ATTENTION . 1. Zhen Qin 1,3 Weixuan Sun 1,4 Hui Deng 3 Dongxu Li 1 Yunshen Wei 1 Baohong Lv 1. Junjie Yan 2,5 Lingpeng Kong 1,2 Yiran Zhong . 1 2 3. SenseTime Research Shanghai AI Laboratory Australian National University 4 5. Northwestern Polytechnical University The University of Hong Kong A BSTRACT. [ ] 17 Feb 2022. Transformer has shown great successes in natural language processing, computer vision, and audio processing. As one of its core components, the SOFTMAX atten- tion helps to capture long-range dependencies yet prohibits its scale-up due to the quadratic space and time complexity to the sequence length. Kernel methods are often adopted to reduce the complexity by approximating the SOFTMAX operator. Nevertheless, due to the approximation errors, their performances vary in differ- ent tasks/corpus and suffer crucial performance drops when compared with the vanilla SOFTMAX ATTENTION .
Performer Reformer cosFormer Sinkhorn Transformer Sparse Transformer Synthesizer Transformer Figure 1: Performance (yaxis), speed (xaxis), and memory footprint (circle sizes) of efficient transformers on the Long-Range Arena benchmark. The proposed COSFORMER achieves an all-around supremacy over competing methods in the top left quadrant.
Domain:
Source:
Link to this page:
Please notify us if you found a problem with this document:
{{id}} {{{paragraph}}}