Example: biology

arXiv:1802.07856v1 [cs.CV] 22 Feb 2018

XView: Objects in Context in Overhead ImageryDarius Lam1 Richard Kuzma2 Kevin McGee3 Samuel Dooley4 Michael Laielli4 Matthew Klaric4 Yaroslav Bulatov5 Brendan McCord2 AbstractWe introduce a new large-scale dataset for theadvancement of object detection techniques andoverhead object detection research. This satel-lite imagery dataset enables research progress per-taining to four key computer vision a novel process for geospatial category de-tection and bounding box annotation with threestages of quality control. Our data is collectedfrom WorldView-3 satellites at ground sam-ple distance, providing higher resolution imagerythan most public satellite imagery datasets. Wecompare xView to other object detection datasetsin both natural and overhead imagery domains andthen provide a baseline analysis using the Sin-gle Shot MultiBox Detector.

with 1 million labeled objects covering over 1,400 km2 of the earth’s surface. The large chip sizes allow variability in pre-processing techniques; we discuss several options in section 4. xView has

Tags:

  Earth

Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Other abuse

Advertisement

Transcription of arXiv:1802.07856v1 [cs.CV] 22 Feb 2018

1 XView: Objects in Context in Overhead ImageryDarius Lam1 Richard Kuzma2 Kevin McGee3 Samuel Dooley4 Michael Laielli4 Matthew Klaric4 Yaroslav Bulatov5 Brendan McCord2 AbstractWe introduce a new large-scale dataset for theadvancement of object detection techniques andoverhead object detection research. This satel-lite imagery dataset enables research progress per-taining to four key computer vision a novel process for geospatial category de-tection and bounding box annotation with threestages of quality control. Our data is collectedfrom WorldView-3 satellites at ground sam-ple distance, providing higher resolution imagerythan most public satellite imagery datasets. Wecompare xView to other object detection datasetsin both natural and overhead imagery domains andthen provide a baseline analysis using the Sin-gle Shot MultiBox Detector.

2 XView is one of thelargest and most diverse publicly available object-detection datasets to date, with over 1 million ob-jects across 60 classes in over 1,400 km2of IntroductionThe abundance of overhead image data fromsatellites and the growing diversity and signifi-cance of real-world applications enabled by thatimagery provide impetus for creating more sophis-ticated and robust models and algorithms for ob-1D. Lam is at Harvard College, in support of DefenseInnovation Unit Experimental (DIUx)2R. Kuzma and B. McCord are at DIUx3K. McGee is at DigitalGlobe4S. Dooley, M. Laielli, and M. Klaric are at the NationalGeospatial-Intelligence Agency (NGA)5Y. Bulatov is in support of DIUxFigure 1: Four of the many views of xView. Im-agery comes from different geographical locationswith different levels of human use.

3 Imagery in thisfigure is from detection. We hope xView will become a cen-tral resource for a broad range of research in com-puter vision and overhead object vast majority of satellite information in thepublic domain is unlabeled. Azayev notes a lackof labeled imagery for developing deep learningmethods [2]. Hamid et al. also note that the cu-ration of high quality labeled data is central todeveloping remote sensor applications [11]. Ishiiet al., Chen et al., and Albert et al. all developmethods for detection or segmentation of build-1 [ ] 22 Feb 2018ings, all using different (and sometimes customcollected and labeled) datasets [13, 4, 1]. The uti-lization of different datasets makes it difficult tocompare results between different authors. We de-veloped xView as a general purpose object detec-tion dataset of satellite imagery so as to be famil-iar to the computer vision community and remotesensing community object detection datasets exist in thenatural imagery space, but there are few foroverhead satellite imagery.

4 The public overheaddatasets in existence typically suffer from low classcount, poor geographic diversity, few training in-stances, or too narrow class scope. xView reme-dies these gaps through a significant labeling effortinvolving the collection of imagery from a varietyof locations and the use of an ontology of parent-and child-level created xView with four computer visionfrontiers in mind:Improve Minimum Resolution and Multi-Scale Recognition:Computer vision algorithmsoften struggle with low-resolution objects [9, 5].For example, the YOLO architecture limits thenumber of predictable bounding boxes within aspatial region, making it difficult to detect smalland clustered objects [24]. Detecting objects onmultiple scales is an ongoing topic of research[19, 3]. Objects in xView vary in size from 3 meters(10 pixels at meters ground-sample distance[GSD]) to greater than 3,000 meters (10,000 pixelsat meters GSD).

5 The varying ground sampledistance of different satellites means that xViewhas significantly higher resolution than many pub-lic satellite imagery Learning Efficiency:In the realworld, objects are often not evenly distributedwithin images. There may be many thousandsmore cars in any given city than there are hos-pitals. Imbalanced classification and localizationon uneven datasets is important for real-world ap-plications. xView captures this property by in-cluding object classes with few instances as wellas classes with many instances (see Figure 4).Push the Limit of Discoverable ObjectClasses:xView includes 60 classes. For refer-ence, COCO includes 91 classes and SpaceNet in-cludes 2 classes. There is significant class diver-sity in xView, which contains both land-use andpedestrian classes with easily discretized objectssuch as cars and buildings as well object classesinvolving groupings of multiple object types suchas construction sites and vehicle Detection of Fine GrainedClasses:Fine grained object detection is neces-sary for practical applications.

6 Detecting a sail-boat gives different information than detecting an oil tanker , despite them both being maritimevessels . Fine grained object detection is a difficulttask and an ongoing area of research [30, 27, 14].Over 80% of classes in xView are fine grained, be-longing to 7 different parent classes. For example,xView contains 8 distinct truck child classes, in-cluding pickup truck , utility truck , and cargotruck .To create xView we designed a substantiveannotation and quality control satellite images delivered in RGB and 8-band multispectral format, annotators used QGIS(Q-Geographic Information System), an opensource tool, to load up and mark image an in-house plugin, annotators are able tocreate axis-aligned bounding boxes for individualobjects. Our dataset includes images prepared inways that are typical for satellite images includ-ing orthorectification, pan-sharpening, and atmo-spheric order to minimize biased image sampling,we define scene types that are relevant to manyoverhead applications and strive for a uniform dis-tribution of images across those scene types aswell as the ways in which those scenes may ap-pear.

7 Variability in scene type can come from aplace s function, while the visual appearance of ascene may vary according to a multitude of fac-tors. xView pulls from a wide range of geographiclocations (see Figure 11). Each location has itsown distinct features, including physical differ-ences (desert, forest, coastal, plains) and construc-tional differences (layout of houses, cities, roads).The variety of collection geometries possible withsatellite imagery produces images with multipleperspectives on objects within a given xView dataset contains 60 object categories2with 1 million labeled objects covering over 1,400km2of the earth s surface. The large chip sizesallow variability in pre-processing techniques; wediscuss several options in section 4. xView hasa similar number of instances and class counts asCOCO and substantially greater number of classesthan SpaceNet and Cars Overhead with Context[20, 21, 23].

8 2. Related WorkThe complexity of our world, combined with dif-ferences in collection geometry from space-basedimaging platforms, makes object recognition insatellite imagery a difficult con-tributes a large, multi-class, multi-location datasetin the object detection and satellite imagery space,built with the benchmark capabilities of PASCALVOC, the quality control methodologies of COCO,and the contributions of other overhead datasetsin mind. This combination opens up opportunitiesfor applied research in census mapping ( , corre-lating building count and inhabitance), economicreporting ( , predicting income level throughvehicle density), disaster response ( , identify-ing damaged regions), and task of object detection is properly iden-tifying an object within an image and localizingit, either through bounding boxes or segmenta-tion.

9 One such dataset, PASCAL VOC, has beenmaintained since 2005 and has grown to 20 objectclasses in 11,530 images containing 27,450 bound-ing boxes and 7,000 segmentations [8]. In thepast decade, the PASCAL VOC dataset has beenwidely used in the object detection space. A num-ber of object detection papers have used PASCALVOC as a benchmark [24, 20, 25, 7]. PASCALVOC, however, contained mostly iconic view scenes that are often non-representative of the realworld, an issue remedied by COCO. COCO con-tains 91 object classes across around 328,000 im-ages. The COCO dataset has on average morecategories per image at smaller sizes than PAS-CAL VOC [20]. The ImageNet detection datasetis a large-scale dataset containing 200 classesand around .5 million labeled instances (ILSVRC2014) [26].

10 Most recently, Google released Open-Images V2, a large-scale dataset containing images, around 4 million bounding boxes,and 600 object classes on natural imagery [15].The Cars Overhead with Context (COWC)dataset by the Lawrence Livermore National Lab-oratory is an overhead image dataset with around32,700 labeled cars [23]. COWC uses aerial im-age capture as opposed to satellite image cap-ture, so their images are of high resolution butcapture less area. COWC includes images fromsix locations. The SpaceNet dataset focuses onobject segmentation. SpaceNet has segmentationmasks for around 5 million buildings in 5 loca-tions [21]. Recently, the SpaceNet dataset hasexpanded to include roads. Both SpaceNet andCOWC have images taken at similar times of of these datasets contains few classes in fewgeographic regions, limiting their usability for gen-eral object recognition in overhead imagery.


Related search queries