Transcription of Florence: A New Foundation Model for Computer Vision
{{id}} {{{paragraph}}}
Florence: A New Foundation Model for Computer VisionLu Yuan1 Dongdong Chen* 1Yi-Ling Chen* 1 Noel Codella* 1 Xiyang Dai* 1 Jianfeng Gao* 2 Houdong Hu* 1 Xuedong Huang* 1 Boxin Li* 1 Chunyuan Li* 2Ce Liu* 1 Mengchen Liu* 1 Zicheng Liu* 1 Yumao Lu* 1Yu Shi* 1 Lijuan Wang* 1 Jianfeng Wang* 1 Bin Xiao* 1 Zhen Xiao* 1 Jianwei Yang* 2 Michael Zeng* 1 Luowei Zhou* 1 Pengchuan Zhang* 2 AbstractAutomated visual understanding of our diverseand open world demands Computer Vision modelsto generalize well with minimal customization forspecific tasks, similar to human Vision . Computervision Foundation models, which are trained ondiverse, large-scale dataset and can be adaptedto a wide range of downstream tasks, are criti-cal for this mission to solve real-world computervision applications. While existing Vision founda-tion models such as CLIP (Radford et al., 2021),ALIGN (Jia et al., 2021), and Wu Dao (Wud)focus mainly on mapping images and textual rep-resentations to a cross-modal shared representa-tion, we introduce a new Computer Vision foun-dation Model ,Florence, to expand the represen-tations from coarse (scene) to fine ( object ), fromstatic (images) to dynamic (videos), and fromRGB to multiple modalities (caption, depth).
cation) to fine-grained (e.g. object detection), 2) Time: from static (e.g. images) to dynamic (e.g. videos), and 3) Modal-ity: from RGB only to multiple senses (e.g. captioning and depth). Due to the diversity nature of visual understanding, we …
Domain:
Source:
Link to this page:
Please notify us if you found a problem with this document:
{{id}} {{{paragraph}}}