Transcription of Scalability in Perception for Autonomous Driving: Waymo ...
1 Scalability in Perception for Autonomous Driving: Waymo Open DatasetPei Sun1, Henrik Kretzschmar1, Xerxes Dotiwalla1, Aur elien Chouard1, Vijaysai Patnaik1, Paul Tsui1,James Guo1, Yin Zhou1, Yuning Chai1, Benjamin Caine2, Vijay Vasudevan2, Wei Han2, Jiquan Ngiam2,Hang Zhao1, Aleksei Timofeev1, Scott Ettinger1, Maxim Krivokon1, Amy Gao1, Aditya Joshi1, YuZhang 1, Jonathon Shlens2, Zhifeng Chen2, and Dragomir Anguelov11 Waymo LLC2 Google LLCA bstractThe research community has increasing interest in au-tonomous driving research, despite the resource intensityof obtaining representative real world data. Existing self-driving datasets are limited in the scale and variation ofthe environments they capture, even though generalizationwithin and between operating regions is crucial to the over-all viability of the technology . In an effort to help align theresearch community s contributions with real-world self-driving problems, we introduce a new large-scale, highquality, diverse dataset.
2 Our new dataset consists of 1150scenes that each span 20 seconds, consisting of well syn-chronized and calibrated high quality LiDAR and cameradata captured across a range of urban and suburban ge-ographies. It is 15x more diverse than the largest cam-era+LiDAR dataset available based on our proposed geo-graphical coverage metric. We exhaustively annotated thisdata with 2D (camera image) and 3D (LiDAR) boundingboxes, with consistent identifiers across frames. Finally, weprovide strong baselines for 2D as well as 3D detectionand tracking tasks. We further study the effects of datasetsize and generalization across geographies on 3D detectionmethods. Find data, code and more up-to-date informationat IntroductionAutonomous driving technology is expected to enable awide range of applications that have the potential to savemany human lives, ranging from robotaxis to self-drivingtrucks. The availability of public large-scale datasets andbenchmarks has greatly accelerated progress in machineperception tasks, including image classification, object de-tection, object tracking, semantic segmentation as well as Work done while at Waymo segmentation [7,17,23,10].
3 To further accelerate the development of autonomousdriving technology , we present the largest and most diversemultimodal Autonomous driving dataset to date, comprisingof images recorded by multiple high-resolution cameras andsensor readings from multiple high-quality LiDAR scannersmounted on a fleet of self-driving vehicles. The geographi-cal area captured by our dataset is substantially larger thanthe area covered by any other comparable Autonomous driv-ing dataset, both in terms of absolute area coverage, andin distribution of that coverage across geographies. Datawas recorded across a range of conditions in multiple cities,namely San Francisco, Phoenix, and Mountain View, withlarge geographic coverage within each city. We demonstratethat the differences in these geographies lead to a pronounceddomain gap, enabling exciting research opportunities in thefield of domain proposed dataset contains a large number of high-quality, manually annotated 3D ground truth bounding boxesfor the LiDAR data, and 2D tightly fitting bounding boxesfor the camera images.
4 All ground truth boxes contain trackidentifiers to support object tracking. In addition, researcherscan extract 2D amodal camera boxes from the 3D LiDARboxes using our provided rolling shutter aware projectionlibrary. The multimodal ground truth facilitates research insensor fusion that leverages both the LiDAR and the cameraannotations. Our dataset contains around 12 million LiDARbox annotations and around 12 million camera box annota-tions, giving rise to around 113k LiDAR object tracks andaround 250k camera image tracks. All annotations werecreated and subsequently reviewed by trained labelers usingproduction-level labeling recorded all the sensor data of our dataset using anindustrial-strength sensor suite consisting of multiple high-resolution cameras and multiple high-quality LiDAR , we offer synchronization between the cameraand the LiDAR readings, which offers interesting opportu-12446nities for cross-domain learning and transfer.
5 We releaseour LiDAR sensor readings in the form of range images. Inaddition to sensor features such as elongation, we provideeach range image pixel with an accurate vehicle pose. Thisis the first dataset with such low-level, synchronized infor-mation available, making it easier to conduct research onLiDAR input representations other than the popular 3D pointset dataset currently consists of 1000 scenes for trainingand validation, and 150 scenes for testing, where each scenespans 20 s. Selecting the test set scenes from a geographicalholdout area allows us to evaluate how well models that weretrained on our dataset generalize to previously unseen present benchmark results of several state-of-the-art2D-and 3D object detection and tracking methods on Related WorkHigh-quality, large-scale datasets are crucial for au-tonomous driving research. There have been an increasingnumber of efforts in releasing datasets to the community inrecent Autonomous driving systems fuse sensor readingsfrom multiple sensors, including cameras, LiDAR, radar,GPS, wheel odometry, and IMUs.
6 Recently released au-tonomous driving datasets have included sensor readingsobtained by multiple sensors. Geigeret al. introduced themulti-sensor KITTI Dataset [9,8] in 2012, which providessynchronized stereo camera as well as LiDAR sensor datafor 22 sequences, enabling tasks such as 3D object detectionand tracking, visual odometry, and scene flow SemanticKITTI Dataset [2] provides annotations thatassociate each LiDAR point with one of 28 semantic classesin all 22 sequences of the KITTI ApolloScape Dataset [12], released in 2017, pro-vides per-pixel semantic annotations for 140k camera imagescaptured in various traffic conditions, ranging from simplescenes to more challenging scenes with many objects. Thedataset further provides pose information with respect tostatic background point clouds. The KAIST Multi-SpectralDataset [6] groups scenes recorded by multiple sensors, in-cluding a thermal imaging camera, by time slot, such asdaytime, nighttime, dusk, and dawn.
7 The Honda ResearchInstitute 3D Dataset (H3D) [19] is a 3D object detection andtracking dataset that provides 3D LiDAR sensor readingsrecorded in 160 crowded urban recently published datasets also include map infor-mation about the environment. For instance, in addition tomultiple sensors such as cameras, LiDAR, and radar, thenuScenes Dataset [4] provides rasterized top-down semanticmaps of the relevant areas that encode information aboutdriveable areas and sidewalks for 1k scenes. This dataset haslimited LiDAR sensor quality with 34K points per frame,KITTI NuScenes Argo OursScenes221000113 1150 Ann. Lidar 12M2D Boxes80K Points/Frame120K34K107K 177 KLiDAR Features1112 MapsNoYesYesNoVisited Area(km2) 1. Comparison of some popular datasets. The Argo Datasetrefers to their Tracking dataset only, not the Motion Forecastingdataset. 3D labels projected to 2D are not counted in the 2D Points/Frame is the number of points from all LiDAR returnscomputed on the released data.
8 Visited area is measured by dilutingtrajectories by 75 meters in radius and union all the diluted observations: 1. Our dataset has effective geographicalcoverage defined by the diversity area metric in 2. Ourdataset is larger than other camera+LiDAR datasets by differentmetrics. (Section2)TOPF,SL,SR,RVFOV[ , + ] [-90 , 30 ]Range (restricted)75 meters20 metersReturns/shot22 Table 2. LiDAR Data Specifications for Front (F), Right (R), Side-Left (SL), Side-Right (SR), and Top (TOP) sensors. The verticalfield of view (VFOV) is specified based on inclination ( ).limited geographical diversity covering an effective area of5km2(Table1).In addition to rasterized maps, the Argoverse Dataset [5]contributes detailed geometric and semantic maps of theenvironment comprising information about the ground heighttogether with a vector representation of road lanes and theirconnectivity. They further study the influence of the providedmap context on Autonomous driving tasks, including 3 Dtracking and trajectory prediction.
9 Argoverse has a verylimited amount raw sensor data Table1for a comparison of different Waymo Open Sensor SpecificationsThe data collection was conducted using five LiDAR sen-sors and five high-resolution pinhole cameras. We restrictthe range of the LiDAR data, and provide data for the firsttwo returns of each laser pulse. Table2contains detailedspecifications of our LiDAR data. The camera images arecaptured with rolling shutter scanning, where the exact scan-2447 FFL,FRSL,SRSize1920x1280 1920x1280 1920x1040 HFOV Table 3. Camera Specifications for Front (F), Front-Left (FL), Front-Right (FR), Side-Left (SL), Side-Right (SR) cameras. The imagesizes reflect the results of both cropping and downsampling theoriginal sensor data. The camera horizontal field of view (HFOV) isprovided as an angle range in the x-axis in the x-y plane of camerasensor frame (Figure1).Laser: FRONTL aser: REARV ehicleLaser: SIDE_LEFTL aser: SIDE_RIGHTL aser: TOPSIDE_LEFTFRONT_RIGHTC amerasFRONTFRONT_LEFTSIDE_RIGHTx-axisy-a xisz-axis is positive upwardsFigure 1.
10 Sensor layout and coordinate mode can vary from scene to scene. All camera imagesare downsampled and cropped from the raw images; Table3provides specifications of the camera images. See Figure1for the layout of sensors relevant to the Coordinate SystemsThis section describes the coordinate systems used inthe dataset. All of the coordinate systems follow the righthand rule, and the dataset contains all information needed totransform data between any two frames within a run frameis set prior to vehicle motion. It is anEast-North-Up coordinate system: Up (z) is aligned with thegravity vector, positive upwards; East (x) points directly eastalong the line of latitude; North (y) points towards the framemoves with the vehicle. Its x-axisis positive forwards, y-axis is positive to the left, z-axisis positive upwards. A vehicle pose is defined as a 4x4transform matrix from the vehicle frame to the global frame can be used as the proxy to transform betweendifferent vehicle frames.