Transcription of Bilinear CNN Models for Fine-grained Visual Recognition
{{id}} {{{paragraph}}}
Bilinear CNN Models for Fine-grained Visual RecognitionTsung-Yu LinAruni RoyChowdhurySubhransu MajiUniversity of Massachusetts, propose Bilinear Models , a Recognition architecturethat consists of two feature extractors whose outputs aremultiplied using outer product at each location of the im-age and pooled to obtain an image descriptor. This archi-tecture can model local pairwise feature interactions in atranslationally invariant manner which is particularly use-ful for Fine-grained categorization. It also generalizes var-ious orderless texture descriptors such as the Fisher vec-tor, VLAD and O2P. We present experiments with bilinearmodels where the feature extractors are based on convolu-tional neural networks. The Bilinear form simplifies gra-dient computation and allows end-to-end training of bothnetworks using image labels only.
sance factors is to first localize various parts of the object and model the appearance conditioned on their detected locations. The parts are often defined manually and the part detectors are trained in a supervised manner. Recently variants of such models based on convolutional neural net-works (CNNs) [2, 38] have been shown to significantly
Domain:
Source:
Link to this page:
Please notify us if you found a problem with this document:
{{id}} {{{paragraph}}}