Transcription of Bilinear CNN Models for Fine-grained Visual Recognition
{{id}} {{{paragraph}}}
Bilinear CNN Models for Fine-grained Visual RecognitionTsung-Yu LinAruni RoyChowdhurySubhransu MajiUniversity of Massachusetts, propose Bilinear Models , a Recognition architecturethat consists of two feature extractors whose outputs aremultiplied using outer product at each location of the im-age and pooled to obtain an image descriptor. This archi-tecture can model local pairwise feature interactions in atranslationally invariant manner which is particularly use-ful for Fine-grained categorization. It also generalizes var-ious orderless texture descriptors such as the Fisher vec-tor, VLAD and O2P.
“D-Net” of [32]. Out of the box these networks do remark-ably well, e.g., features from the penultimate layer of these networks achieve 52.7% and 61.0% accuracy on the CUB-200-2011 dataset [37] respectively. Fine-tuning improves the performance further to 58.8% and 70.4%. In compari-son a fine-tuned bilinear model consisting of a M-Net and
Domain:
Source:
Link to this page:
Please notify us if you found a problem with this document:
{{id}} {{{paragraph}}}