Example: bachelor of science

A Benchmark Dataset and Evaluation Methodology for Video ...

A Benchmark Dataset and Evaluation Methodology forVideo Object SegmentationF. Perazzi1,2J. Pont-Tuset1B. McWilliams2L. Van Gool1M. Gross1,2A. Sorkine-Hornung21 ETH Zurich2 Disney ResearchAbstractOver the years, datasets and benchmarks have proventheir fundamental importance in computer vision research,enabling targeted progress and objective comparisons inmany fields. At the same time, legacy datasets may impendthe evolution of a field due to saturated algorithm perfor-mance and the lack of contemporary, high quality data. Inthis work we present a new Benchmark Dataset and evalu-ation Methodology for the area ofvideo object segmenta-tion. The Dataset , named DAVIS (Densely Annotated VIdeoSegmentation), consists of fifty high quality, Full HD videosequences, spanning multiple occurrences of common videoobject segmentation challenges such as occlusions, motion-blur and appearance changes. Each Video is accompaniedby densely annotated, pixel-accurate and per-frame groundtruth segmentation .

lenges encountered in realistic video object segmentation applications. Furthermore,theimagequalityisnotanymore representative of modern consumer devices, and due to the limited number of available video sequences, progress on this dataset plateaued. In [25] this dataset was extended with 8 additional sequences. While this is certainly an im-

Tags:

  Applications, Segmentation, Segmentation applications

Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Other abuse

Advertisement

Transcription of A Benchmark Dataset and Evaluation Methodology for Video ...

1 A Benchmark Dataset and Evaluation Methodology forVideo Object SegmentationF. Perazzi1,2J. Pont-Tuset1B. McWilliams2L. Van Gool1M. Gross1,2A. Sorkine-Hornung21 ETH Zurich2 Disney ResearchAbstractOver the years, datasets and benchmarks have proventheir fundamental importance in computer vision research,enabling targeted progress and objective comparisons inmany fields. At the same time, legacy datasets may impendthe evolution of a field due to saturated algorithm perfor-mance and the lack of contemporary, high quality data. Inthis work we present a new Benchmark Dataset and evalu-ation Methodology for the area ofvideo object segmenta-tion. The Dataset , named DAVIS (Densely Annotated VIdeoSegmentation), consists of fifty high quality, Full HD videosequences, spanning multiple occurrences of common videoobject segmentation challenges such as occlusions, motion-blur and appearance changes. Each Video is accompaniedby densely annotated, pixel-accurate and per-frame groundtruth segmentation .

2 In addition, we provide a comprehen-sive analysis of several state-of-the-art segmentation ap-proaches using three complementary metrics that measurethe spatial extent of the segmentation , the accuracy of thesilhouette contours and the temporal coherence. The resultsuncover strengths and weaknesses of current approaches,opening up promising directions for future IntroductionVideo object segmentation is a binary labeling prob-lem aiming to separate foreground object(s) from the back-ground region of a Video . A pixel-accurate, spatio-temporalbipartition of the Video is instrumental to several applica-tions including, among others, action recognition, objecttracking, Video summarization, and rotoscoping for videoediting. Despite remarkable progress in recent years, videoobject segmentation still remains a challenging problem andmost existing approaches still exhibit too severe limitationsin terms of quality and efficiency to be applicable in practi-cal applications , for processing large datasets, or videopost-production and editing in the visual effects is most striking is the performance gap amongstate-of-the-art Video object segmentation algorithms andclosely related methods focusing on image segmentationFigure 1: Sample sequences from our Dataset , with groundtruth segmentation masks overlayed.

3 Please refer to the sup-plemental material for the complete object recognition, which have experienced remark-able progress in the recent years. A key factor boot-strapping this progress has been the availability of largescale datasets and benchmarks [12,26,29,42]. This is instark contrast to Video object segmentation . While sev-eral datasets exists for various different Video segmentationtasks [1,4,5,15,20,21,25,38,41,44,46,47], none of themtargets the specific task of date, the most widely adopted Dataset is that of [47],which, however, was originally proposed for joint segmen-tation and tracking and only contains six low-resolutionvideo sequences, which are not representative anymore forthe image quality and resolution encountered in today svideo processing applications . As a consequence, evalua-tions performed on such datasets are likely to be overfit-ted, without reliable indicators regarding the differences be-tween individual Video segmentation approaches, and thereal performance on unseen, more contemporary data be-comes difficult to determine [6].

4 Despite the effort of someauthors to augment their Evaluation with additional datasets,a standardized and widely adopted Evaluation methodologyfor Video object segmentation does not yet this end, we introduce a new Dataset specifically de-signed for the task of Video object segmentation . The1724dataset, which will be made publicly available, containsfifty densely and professionally annotated high-resolutionFull HD Video sequences, with pixel-accurate ground-truthdata provided for every Video frame. The sequences havebeen carefully captured to cover multiple instances of ma-jor challenges typically faced in Video object Dataset is accompanied with a comprehensive evalua-tion of several state-of-the-art approaches [5,7,13,14,18,21,24,33,35,40,43,45]. To evaluate the performance weemploy three complementary metrics measuring the spa-tial accuracy of the segmentation , the quality of the sil-houette and its temporal coherence.

5 Furthermore, we anno-tated each Video with specific attributes such asocclusions,fast-motion,non-linear deformationandmotion-blur. Cor-related with the performance of the tested approaches, theseattributes enable a deeper understanding of the results andpoint towards promising avenues for future research. Thecomponents described above represent a complete bench-mark suite, providing researchers with the necessary toolsto facilitate the Evaluation of their methods and advance thefield of Video object Related WorksIn this section we provide an overview of datasets de-signed for different Video segmentation tasks, followed bya survey of techniques targeting Video object DatasetsThere exist several datasets for Video segmentation , butnone of them has been specifically designed for videoob-jectsegmentation, the task of pixel-accurate separation offoreground object(s) from the background Motion Segmentationdataset [5]MoSegis a popular Dataset for motion segmentation , regions with similar motion.

6 Despite being re-cently adopted by works focusing on Video object segmen-tation [35,45], the Dataset does not fulfill several importantrequirements. Most of the videos have low spatial resolu-tion, segmentation is only provided on a sparse subset of theframes, and the content is not sufficiently diverse to providea balanced distribution of challenging situations such as fastmotion and Video segmentation Dataset (BVSD) [44]comprises a total 100, higher resolution sequences. It wasoriginally meant to evaluate occlusions boundary detectionand later extended to over- and motion- segmentation tasks(VSB100 [19]). However, several sequences do not containa clear object. Furthermore, the ground-truth, available onlyfor a subset of the frames, is fragmented, with most of theobjects being covered by multiple manually annotated, dis-joint segments, and therefore this Dataset is not well suitedfor evaluating Video object [47] is a small Dataset composed of 6 denselyannotated videos of humans and animals.

7 It is designed tobe challenging with respect to background-foreground colorsimilarity, fast motion and complex shape deformation. Al-though it has been extensively used by several approaches,its content does not sufficiently span the variety of chal-lenges encountered in realistic Video object segmentationapplications. Furthermore, the image quality is not anymorerepresentative of modern consumer devices, and due to thelimited number of available Video sequences, progress onthis Dataset plateaued. In [25] this Dataset was extendedwith 8 additional sequences. While this is certainly an im-provement over the predecessor, it still suffers of the samelimitations. We refer the reader to the supplemental mate-rial for a comprehensive summary of the properties of theaforementioned datasets, including datasets exist, but they are mostly provided to sup-port specific findings and thus are either limited in terms oftotal number of frames, [8,21,25,47], or do not exhibit a suf-ficient variety in terms of content [1,4,5,15,17,20,41,46].

8 Others cover a broader range of content but do not provideenough ground-truth data for an accurate Evaluation of thesegmentation [21,38]. Video datasets designed to bench-mark tracking algorithms typically focus on surveillancescenarios with static cameras [9,16,32], and usually con-tain multiple instances of similar objects [50] ( a crowdof people), and annotation is typically provided only inthe form of axis-aligned bounding boxes, instead of pixel-accurate segmentation masks necessary to accurately eval-uate Video object segmentation . Importantly, none of theaforementioned methods includes contemporary high reso-lution videos, which is an absolute necessity to realisticallyevaluate the actual practical utility of such AlgorithmsWe categorize the body of literature related to Video ob-ject segmentation based on the level of supervision have historically targetedover- segmentation [21,51] or motion segmentation [5,18] and only recently automatic methods for foreground-background separation have been proposed [13,25,33,43,45,52].

9 These methods extend the concept of salient objectdetection [34] to videos. They do not require any manualannotation and do not assume any prior information on theobject to be segmented. Typically they are based on theassumption that object motion is dissimilar from the sur-roundings. Some of these methods generate several rankedsegmentation hypotheses [24]. While they are well suitedfor parsing large scale databases, they are bound to theirunderlying assumption and fail in cases it does not object segmentation methodspropagate a sparse manual labeling, generally given in theform of one or more annotated frames, to the entire video725 IDDescriptionBCBackground Clutter. The back- and foreground regions aroundthe object boundaries have similar colors ( 2over histograms).DEFD eformation. Object undergoes complex, non-rigid Blur. Object has fuzzy boundaries due to fast The average, per-frame object motion, computedas centroids Euclidean distance, is larger than fm=20 Resolution.

10 The ratio between the average objectbounding-box area and the image area is smaller thantlr= Object becomes partially or fully Object is partially clipped by the image The area ratio among any pair of bounding-boxes enclosing the target object is smaller than sv= Change. Noticeable appearance variation, dueto illumination changes and relative camera-object Ambiguity. Unreliable edge detection. The average ground-truth edge probability (using [11]) is smaller than e= Footage displays non-negligible Object. Object regions have distinct Objects. The target object is an ensemble of multiple,spatially-connected objects ( mother with stroller).DBDynamic Background. Background regions move or Complexity. The object has complex boundaries such asthin parts and 1: List of Video attributes and corresponding descrip-tion. We extend the annotations of [50] (top) with a comple-mentary set of attributes relevant to videoobjectsegmenta-tion (bottom).


Related search queries