Example: bachelor of science

Motion Representations for Articulated Animation

Motion Representations for Articulated AnimationAliaksandr Siarohin1 , Oliver J. Woodford , Jian Ren2, Menglei Chai2and Sergey Tulyakov21 DISI, University of Trento, Italy,2 Snap Inc., Santa Monica, Work done while at Snap propose novel Motion Representations for animatingarticulated objects consisting of distinct parts. In a com-pletely unsupervised manner, our method identifies objectparts, tracks them in a driving video, and infers their motionsby considering their principal axes. In contrast to the previ-ous keypoint-based works, our method extracts meaningfuland consistentregions, describing locations, shape, andpose.

Monkey-Net [31] learns a set of unsupervised keypoints to generate animations. Follow- ... being Z, and Mk(z) is the k-th heatmap weight at pixel z. Thus, the translation component of the affine transformation (which is the last column of Ak X R) can be …

Tags:

  Monkey

Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Other abuse

Advertisement

Transcription of Motion Representations for Articulated Animation

1 Motion Representations for Articulated AnimationAliaksandr Siarohin1 , Oliver J. Woodford , Jian Ren2, Menglei Chai2and Sergey Tulyakov21 DISI, University of Trento, Italy,2 Snap Inc., Santa Monica, Work done while at Snap propose novel Motion Representations for animatingarticulated objects consisting of distinct parts. In a com-pletely unsupervised manner, our method identifies objectparts, tracks them in a driving video, and infers their motionsby considering their principal axes. In contrast to the previ-ous keypoint-based works, our method extracts meaningfuland consistentregions, describing locations, shape, andpose.

2 The regions correspond to semantically relevant anddistinct object parts, that are more easily detected in framesof the driving video. To force decoupling of foreground frombackground, we model non-object related global Motion withan additional affine transformation. To facilitate animationand prevent the leakage of the shape of the driving object,we disentangle shape and pose of objects in the region model1can animate a variety of objects, surpassing pre-vious methods by a large margin on existing benchmarks. Wepresent a challenging new benchmark with high-resolutionvideos and show that the improvement is particularly pro-nounced when Articulated objects are considered, user preference vs.

3 The state of the IntroductionAnimation bringing static objects to life has broadapplications across education and entertainment. Animatedcharacters and objects, such as those in Fig. 1, increasethe creativity and appeal of content, improve the clarity ofmaterial through storytelling, and enhance user very recently, Animation techniques necessary forachieving such results required a trained professional, spe-cialized hardware, software, and a great deal of effort. Qual-ity results generally still do, but vision and graphics commu-nities have attempted to address some of these limitationsby training data-driven methods [39,6,27,11,10] on objectclasses for which prior knowledge of object shape and posecan be learned.

4 This, however, requires ground truth poseand shape data to be available during source code is publicly available at our project website for more qualitative 1: Our method animates still source images via unsu-pervised region detection (inset).Recent works have sought to avoid the need for groundtruth data throughunsupervisedmotion transfer [42,30,31].Significant progress has been made on several key chal-lenges, including training using image reconstruction as aloss [42,30,31], and disentangling Motion from appear-ance [20]. This has created the potential to animate a broaderrange of object categories, without any domain knowledgeor labelled data, requiring only videos of objects in motionduring training [30].

5 However, two key problems remainopen. The first is how to represent the parts of an articulatedor non-rigid moving object, including their shapes and second is given the object parts, how to animate themusing the sequence of motions in a driving attempts used end-to-end frameworks [42,31] tofirst extract unsupervised keypoints [20,17], then warp afeature embedding of a source image to align its keypointswith those of a driving video. Follow on work [30] furthermodelled the Motion around each keypoint with local, affinetransformations, and introduced a generation module thatboth composites warped source image regions and inpaintsoccluded regions, to render the final image.

6 This enabled avariety of creative applications,3for example needing onlyone source face image to generate a near photo-realisticanimation, driven by a video of a different a music video in which images are animated using prior work [30].1 [ ] 22 Apr 2021 However, the resulting unsupervised keypoints are de-tected on the boundary of the objects. While points onedges are easier to identify, tracking such keypoints betweenframes is problematic, as any point on the boundary is avalid candidate, making it hard to establish correspondencesbetween frames. A further problem is that the unsupervisedkeypoints do not correspond to semantically meaningful ob-ject parts, representing location and direction, but not to this limitation, animating Articulated objects, such asbodies, remains challenging.

7 Furthermore, these methodsassume static backgrounds, no camera Motion , leadingto leakage of background Motion information into one orseveral of the detected keypoints. Finally, absolute motiontransfer, as in [30], transfers the shape of the driving objectinto the generated sequence, decreasing the fidelity of thesource identity. These remaining deficiencies limit the scopeof previous works [30,31] to more trivial object categoriesand motions, especially when objects are work introduces three contributions to address thesechallenges. First, we redefine the underlying Motion rep-resentation, usingregionsfrom which first-order Motion ismeasured, rather than regressed.

8 This enables improvedconvergence, more stable, robust object and Motion repre-sentations, and also empirically captures the shape of theunderpinning object parts, leading to better Motion segmen-tation. This Motion representation is inspired by Hu mo-ments [8]. Fig. 3(a) contains several examples of region Motion , we explicitly model background or camera mo-tion between training frames by predicting the parametersof a global, affine transformation explaining non-object re-lated motions. This enables the model to focus solely on theforeground object, making the identified points more stable,and further improves convergence.

9 Finally, to prevent shapetransfer and improve Animation , we disentangle the shapeand pose of objects in the space of unsupervised regions. Ourframework is self-supervised, does not require any labels,and is optimized using reconstruction contributions further improve unsupervised motiontransfer methods, resulting in higher fidelity Animation of ar-ticulated objects in particular. To create a more challengingbenchmark for such objects, we present a newly collecteddataset of TED talk speakers. Our framework scales betterin the number of unsupervised regions, resulting in moredetailed Motion .

10 Our method outperforms previous unsuper-vised Animation methods on a variety of datasets, includingtalking faces, taichi videos and animated pixel art being pre-ferred by of independent raters when compared withthe state of the art [30] on our most challenging Related workImage Animation methods can be separated into super-vised, which require knowledge about the animated objectduring training, and unsupervised, which do not. Suchknowledge typically includes landmarks [4,44,26,12],semantic segmentations [24], and parametric 3D mod-els [11,35,9,22,19]. As a result, supervised methodsare limited to a small number of object categories for whicha lot of labelled data is available, such as faces and humanbodies.


Related search queries