Transferring human motion and appearance between videos of human actors remains one of the key challenges in Computer Vision. Despite the advances from recent image-to-image translation approaches, there are several transferring contexts where most end-to-end learning-based retargeting methods still perform poorly. Transferring human appearance from one actor to another is only ensured when a strict setup has been complied, which is generally built considering their training regime's specificities. In this work, we propose a shape-aware approach based on a hybrid image-based rendering technique that exhibits competitive visual retargeting quality compared to state-of-the-art neural rendering approaches. The formulation leverages the user body shape into the retargeting while considering physical constraints of the motion in 3D and the 2D image domain. We also present a new video retargeting benchmark dataset composed of different videos with annotated human motions to evaluate the task of synthesizing people's videos, which can be used as a common base to improve tracking the progress in the field. The dataset and its evaluation protocols are designed to evaluate retargeting methods in more general and challenging conditions. Our method is validated in several experiments, comprising publicly available videos of actors with different shapes, motion types, and camera setups.
Thiago L. Gomes, Renato Martins, João P. Ferreira, Rafael Azevedo, Guilherme Torres and Erickson R. Nascimento.
A Shape-Aware Retargeting Approach to Transfer Human Motion and Appearance in Monocular Videos, International Journal of Computer Vision, 2021.
PDF, BibTeX
Concurrently and independently from us, a number of groups have proposed closely related — and very interesting! — methods. Here is a partial list:
The authors thank CAPES, CNPq, and FAPEMIG for funding this work. We also thank NVIDIA for the donation of a Titan XP GPU used in this research. We use as template for this webpage the page of Speech2Gesture.