Abstract - arXiv
Transformer Meets Tracker:Exploiting Temporal Context for Robust Visual TrackingNing Wang1Wengang Zhou1,2Jie Wang1,2Houqiang Li1,21CAS Key Laboratory of GIPAS, EEIS Department, University of Science and Technology of China (USTC)2Institute of Artificial Intelligence, Hefei Comprehensive National Science video object tracking, there exist rich temporal con-texts among successive frames, which have been largelyoverlooked in existing trackers. In this work, we bridge theindividual video frames and explore the temporal contextsacross them via a transformer architecture for robust objecttracking. Different from classic usage of the transformer innatural language processing tasks, we separate its encoderand decoder into two parallel branches and carefully designthem within the Siamese-like tracking pipelines.
the bottom branch classifies the current search patch. As shown in Figure1, we separate the transformer encoder and decoder into two branches within such a general Siamese-like structure. In the top branch, a set of template patches are fed to the transformer encoder to generate high-quality encoded features. In the bottom branch, the search ...
Download Abstract - arXiv
Information
Domain:
Source:
Link to this page:
Please notify us if you found a problem with this document: