Transcription of Bilinear CNN Models for Fine-grained Visual Recognition
{{id}} {{{paragraph}}}
Bilinear CNN Models for Fine-grained Visual RecognitionTsung-Yu LinAruni RoyChowdhurySubhransu MajiUniversity of Massachusetts, propose Bilinear Models , a Recognition architecturethat consists of two feature extractors whose outputs aremultiplied using outer product at each location of the im-age and pooled to obtain an image descriptor. This archi-tecture can model local pairwise feature interactions in atranslationally invariant manner which is particularly use-ful for Fine-grained categorization. It also generalizes var-ious orderless texture descriptors such as the Fisher vec-tor, VLAD and O2P.
Bag-of-Visual-Words [8], VLAD [20], Fisher vector [28], and second-order pooling (O2P) [3]. Moreover, the archi-tecture can be easily trained end-to-end unlike these texture descriptions leading to significant improvements in perfor-mance. Although we don’t explore this connection further, our architecture is related to the two stream ...
Domain:
Source:
Link to this page:
Please notify us if you found a problem with this document:
{{id}} {{{paragraph}}}