Transcription of ImageNet: A Large-Scale Hierarchical Image …
1 ImageNet: A Large-Scale Hierarchical Image DatabaseJia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li and Li Fei-FeiDept. of Computer Science, Princeton University, USA{jiadeng, wdong, rsocher, jial, li, explosion of Image data on the Internet has the po-tential to foster more sophisticated and robust models andalgorithms to index, retrieve, organize and interact with im-ages and multimedia data. But exactly how such data canbe harnessed and organized remains a critical problem. Weintroduce here a new database called ImageNet , a Large-Scale ontology of images built upon the backbone of theWordNet structure. ImageNet aims to populate the majorityof the 80,000 synsets of WordNet with an average of 500-1000 clean and full resolution images. This will result intens of millions of annotated images organized by the se-mantic hierarchy of WordNet. This paper offers a detailedanalysis of ImageNet in its current state: 12 subtrees with5247 synsets and million images in total.}
2 We show thatImageNet is much larger in scale and diversity and muchmore accurate than the current Image datasets. Construct-ing such a Large-Scale database is a challenging task. Wedescribe the data collection scheme with Amazon Mechan-ical Turk. Lastly, we illustrate the usefulness of ImageNetthrough three simple applications in object recognition, im-age classification and automatic object clustering. We hopethat the scale , accuracy, diversity and Hierarchical struc-ture of ImageNet can offer unparalleled opportunities to re-searchers in the computer vision community and IntroductionThe digital era has brought with it an enormous explo-sion of data. The latest estimations put a number of morethan3billion photos on Flickr, a similar number of videoclips on YouTube and an even larger number for images inthe Google Image Search database. More sophisticated androbust models and algorithms can be proposed by exploit-ing these images, resulting in better applications for usersto index, retrieve, organize and interact with these data.
3 Butexactly how such data can be utilized and organized is aproblem yet to be solved. In this paper, we introduce a newimage database called ImageNet , a Large-Scale ontologyof images. We believe thata Large-Scale ontology of imagesis a critical resource for developing advanced, large -scalecontent-based Image search and Image understanding algo-rithms, as well as for providing critical training and bench-marking data for such uses the Hierarchical structure of WordNet [9].Each meaningful concept in WordNet, possibly describedby multiple words or word phrases, is called a synonymset or synset . There are around80,000noun synsetsin WordNet. In ImageNet, we aim to provide on aver-age500-1000images to illustrate each synset. Images ofeach concept are quality-controlled and human-annotatedas described in Sec. ImageNet, therefore, will offertens of millions of cleanly sorted images.
4 In this paper,we report the current version of ImageNet, consisting of12 subtrees :mammal, bird, fish, reptile, amphibian, vehicle,furniture, musical instrument, geological formation, tool,flower, fruit. These subtrees contain5247synsets images. Fig. 1 shows a snapshot of two branches ofthe mammal and vehicle subtrees. The database is publiclyavailable rest of the paper is organized as follows:We firstshow that ImageNet is a Large-Scale , accurate and diverseimage database(Section 2). In Section 4, we present a fewsimple application examples by exploiting the current Ima-geNet, mostly the mammal and vehicle subtrees. Our goalis to show thatImageNet can serve as a useful resource forvisual recognition applications such as object recognition, Image classification and object addition, theconstruction of such a Large-Scale and high-quality databasecan no longer rely on traditional data collection 3 describes how ImageNet is constructed by leverag-ing Amazon Mechanical Properties of ImageNetImageNet is built upon the Hierarchical structure pro-vided by WordNet.
5 In its completion, ImageNet aims tocontain in the order of50million cleanly labeled full reso-lution images (500-1000per synset). At the time this paperis written, ImageNet consists of12subtrees. Most analysiswill be based on the mammal and vehicle aims to provide the most comprehensiveand diverse coverage of the Image world. The current12subtrees consist of a total cleanly annotated1mammalplacentalcarnivorecanine dogworking doghuskyvehiclecraftwatercraftsailing vesselsailboattrimaranFigure 1:A snapshot of two root-to-leaf branches of ImageNet: thetoprow is from the mammal subtree; thebottomrow is from thevehicle subtree. For each synset, 9 randomly sampled images are # images per synsetpercentageSubtree# SynsetsAvg. synset sizeTotal # imageMammal1170737862 KVehicle520610317 KGeoForm17643677 KFurniture197797157 KBird872809705 KMusicInstr164672110 KSummary of selected subtreesFigure 2: scale of curve: Histogram of numberof images per synset.
6 About20%of the synsets have very fewimages. Over50%synsets have more :Summary of selected subtrees. For complete and up-to-date statis-tics spread over5247categories (Fig. 2). On averageover600images are collected for each synset. Fig. 2 showsthe distributions of the number of images per synset for thecurrent ImageNet1. To our knowledge this is already thelargest clean Image dataset available to the vision researchcommunity, in terms of the total number of images, numberof images per category as well as the number of organizes the different classes ofimages in adensely populatedsemantic hierarchy. Themain asset of WordNet [9] lies in its semantic structure, ontology of concepts. Similarly to WordNet, synsets ofimages in ImageNet are interlinked by several types of re-lations, the IS-A relation being the most comprehensiveand useful.
7 Although one can map any dataset with cate-1 About20%of the synsets have very few images, because either thereare very few web images available, vespertilian bat , or the synset bydefinition is difficult to be illustrated by images, two-year-old horse .2It is claimed that the ESP game [25] has labeled a very large numberof images, but only a subset of 60K images are publicly Cattle SubtreeImagenet Cattle Subtree176 Imagenet Cat SubtreeESP Cat Subtree13773761830 Figure 3:Comparison of the cat and cattle subtrees betweenESP [25] and ImageNet. Within each tree, the size of a node isproportional to the number of images it contains. The number ofimages for the largest node is shown for each tree. Shared nodesbetween an ESP tree and an ImageNet tree are colored in labels into a semantic hierarchy by using WordNet, thedensity of ImageNet is unmatched by others.
8 For example,to our knowledge no existing vision dataset offers images of147dog categories. Fig. 3 compares the cat and cattle subtrees of ImageNet and the ESP dataset [25]. We observethat ImageNet offers much denser and larger would like to offer a clean dataset at alllevels of the WordNet hierarchy. Fig. 4 demonstrates thelabeling precision on a total of80synsets randomly sam-pled at different tree depths. An average is achieved on average. Achieving a high precision forall depths of the ImageNet tree is challenging because thelower in the hierarchy a synset is, the harder it is to classify, Siamese cat versus Burmese is constructed with the goal that ob-jects in images should have variable appearances, positions, depthFigure 4:Percent of clean images at different tree depth levels inImageNet. A total of80synsets are randomly sampled at everytree depth of the mammal and vehicle subtrees.
9 An independentgroup of subjects verified the correctness of each of the average of is achieved for each 1:Comparison of some of the properties of ImageNet ver-sus other existing datasets. ImageNet offers disambiguated la-bels (LabelDisam), clean annotations (Clean), a dense hierarchy(DenseHie), full resolution images (FullRes) and is publicly avail-able (PublicAvail). ImageNet currently does not provide segmen-tation points, poses as well as background clutter and occlu-sions. In an attempt to tackle the difficult problem of quan-tifying Image diversity, we compute the average Image ofeach synset and measure lossless JPG file size which reflectsthe amount of information in an Image . Our idea is that asynset containing diverse images will result in a blurrier av-erage Image , the extreme being a gray Image , whereas asynset with little diversity will result in a more structured,sharper average Image .
10 We therefore expect to see a smallerJPG file size of the average Image of a more diverse 5 compares the Image diversity in four randomly sam-pled synsets in Caltech101 [8]3and the mammal subtree ImageNet and Related DatasetsWe compare ImageNet with other datasets and summa-rize the differences in Table Image datasetsA number of well labeled smalldatasets (Caltech101/256 [8, 12], MSRC [22], PASCAL [7]etc.) have served as training and evaluation benchmarksfor most of today s computer vision algorithms. As com-puter vision research advances, larger and more challenging3We also compare with Caltech256 [12]. The result indicates the diver-sity of ImageNet is comparable, which is reassuring since Caltech256 wasspecifically designed to be more focus our comparisons on datasets of generic objects. Special pur-pose datasets, such as FERET faces [19], Labeled faces in the Wild [13]and the Mammal Benchmark by Fink and Ullman [11] are not are needed for the next generation of current ImageNet offers20 the number of categories,and100 the number of total images than these [24] is a dataset of80million32 32low resolution images, collected from the Inter-net by sending all words in WordNet as queries to imagesearch engines.