EDBT 2026 Demo / reviewers in the wild / expert
Zerong Zheng
dblp:218/6627
· DBLP profile ↗
38ranked-venue papers
6as first author
31since 2021 · last 2025
0000-0003-1339-2480ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 34 · 5 first-author · 28 since 2021Artificial intelligence and machine learning · 30 · 5 first-author · 23 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation ModelsabstractEnd-to-end human animation, such as audio-driven talking human generation, has undergone notable advancements in the recent few years. However, existing methods still struggle to scale up as large general video generation models, limiting their potential in real applications. In this paper, we propose OmniHuman, a Diffusion Transformer-based framework that scales up data by mixing motion-related conditions into the training phase. To this end, we introduce two training principles for these mixed conditions, along with the corresponding model architecture and inference strategy. These designs enable OmniHuman to fully leverage data-driven motion generation, ultimately achieving highly realistic human video generation. More importantly, OmniHuman supports various portrait contents (face close-up, portrait, half-body, full-body), supports both talking and singing, handles human-object interactions and challenging body poses, and accommodates different image styles. Compared to existing end-to-end audio-driven methods, OmniHuman not only produces more realistic videos, but also offers greater flexibility in inputs. It also supports multiple driving modalities (audio-driven, video-driven and combined driving signals). Video samples are provided on the ttfamily project page (https://omnihuman-lab.github.io) Gaojie Lin, Jianwen Jiang, Jiaqi Yang 0008, Zerong Zheng, Jingtuo Liu |
ICCV | 4 |
| 2025 | CyberHost: A One-stage Diffusion Framework for Audio-driven Talking Body GenerationabstractDiffusion-based video generation technology has advanced significantly, catalyzing a proliferation of research in human animation. While breakthroughs have been made in driving human animation through various modalities for portraits, most of current solutions for human body animation still focus on video-driven methods, leaving audio-driven taking body generation relatively underexplored. In this paper, we introduce CyberHost, a one-stage audio-driven talking body generation framework that addresses common synthesis degradations in half-body animation, including hand integrity, identity consistency, and natural motion.
CyberHost's key designs are twofold. Firstly, the Region Attention Module (RAM) maintains a set of learnable, implicit, identity-agnostic latent features and combines them with identity-specific local visual features to enhance the synthesis of critical local regions. Secondly, the Human-Prior-Guided Conditions introduce more human structural priors into the model, reducing uncertainty in generated motion patterns and thereby improving the stability of the generated videos.
To our knowledge, CyberHost is the first one-stage audio-driven human diffusion model capable of zero-shot video generation for the human body. Extensive experiments demonstrate that CyberHost surpasses previous works in both quantitative and qualitative aspects. CyberHost can also be extended to video-driven and audio-video hybrid-driven scenarios, achieving similarly satisfactory results. Gaojie Lin, Jianwen Jiang, Tianyun Zhong, Jiaqi Yang 0008, Zerong Zheng, Yanbo Zheng |
ICLR | 6 |
| 2025 | SpeechAct: Towards Generating Whole-Body Motion From SpeechabstractWhole-body motion generation from speech audio is crucial for computer graphics and immersive VR/AR. Prior methods struggle to produce natural and diverse whole-body motions from speech. In this paper, we introduce a novel method, named SpeechAct, based on a hybrid point representation and contrastive motion learning to boost realism and diversity in motion generation. Our hybrid point representation leverages the advantages of keypoint representation and surface points of 3D body model, which is easy to learn and helps to achieve smooth and natural motion generation from speech audio. We design a VQ-VAE to learn a motion codebook using our hybrid presentation, and then regress the motion from the input audio using a translation model. To boost diversity in motion generation, we propose a contrastive motion learning method according to the intuitive idea that the generated motion should be different from the motions of other audios and other speakers. We collect negative samples from other audio inputs and other speakers using our translation model. With these negative samples, we pull the current motion away from them using a contrastive loss to produce more distinctive representations. In addition, we compose a face generator to generate deterministic face motion due to the strong connection between the face movements and the speech audio. Experimental results validate the superior performance of our model. The code is available at http://cic.tju.edu.cn/faculty/likun/projects/SpeechAct/index.html. Minjie Zhu, Yuxiang Zhang 0006, Zerong Zheng, Yebin Liu, Kun Li 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | ProbIBR: Fast Image-Based Rendering With Learned Probability-Guided SamplingabstractWe present a general, fast, and practical solution for interpolating novel views of diverse real-world scenes given a sparse set of nearby views. Existing generic novel view synthesis methods rely on time-consuming scene geometry pre-computation or redundant sampling of the entire space for neural volumetric rendering, limiting the overall efficiency. Instead, we incorporate learned MVS priors into the neural volume rendering pipeline while improving the rendering efficiency by reducing sampling points under the guidance of depth probability distributions. Specifically, fewer but important points are sampled under the guidance of depth probability distributions extracted from the learned MVS architecture. Based on the learned probability-guided sampling, we develop a sophisticated neural volume rendering module that effectively integrates source view information with the learned scene structures. We further propose confidence-aware refinement to improve the rendering results in uncertain, occluded, and unreferenced regions. Moreover, we build a four-view camera system for holographic display and provide a real-time version of our framework for free-viewpoint experience, where novel view images of a spatial resolution of 512×512 can be rendered at around 20 fps on a single GTX 3090 GPU. Experiments show that our method achieves 15 to 40 times faster rendering compared to state-of-the-art baselines, with strong generalization capacity and comparable high-quality novel view synthesis performance. Yuemei Zhou, Tao Yu 0007, Zerong Zheng, Gaochang Wu, Guihua Zhao, Ying Fu 0001, Yebin Liu |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | RAM-Avatar: Real-time Photo-Realistic Avatar from Monocular Videos with Full-body ControlabstractThis paper focuses on advancing the applicability of hu-man avatar learning methods by proposing RAM-Avatar, which learns a Real-time, photo-realistic Avatar that sup-ports full-body control from Monocular videos. To achieve this goal, RAM-Avatar leverages two statistical templates responsible for modeling the facial expression and hand gesture variations, while a sparsely computed dual attention module is introduced upon another body template to facilitate high-fidelity texture rendering for the torsos and limbs. Building on this foundation, we deploy a lightweight yet powerful StyleUnet along with a temporal-aware dis-criminator to achieve real-time realistic rendering. To en-able robust animation for out-of-distribution poses, we pro-pose a Motion Distribution Align module to compensate for the discrepancies between the training and testing motion distribution. Results and extensive experiments conducted in various experimental settings demonstrate the superior-ity of our proposed method, and a real-time live system is proposed to further push research into applications. The training and testing code will be released for research pur-poses. Zerong Zheng, Yuxiang Zhang 0006, Jingxiang Sun, Lizhen Wang 0002, Yebin Liu |
CVPR | 2 |
| 2024 | Animatable Gaussians: Learning Pose-Dependent Gaussian Maps for High-Fidelity Human Avatar ModelingabstractModeling animatable human avatars from RGB videos is a longstanding and challenging problem. Recent works usually adopt MLP-based neural radiance fields (NeRF) to represent 3D humans, but it remains difficult for pure MLPs to regress pose-dependent garment details. To this end, we introduce Animatable Gaussians, a new avatar representation that leverages powerful 2D CNNs and 3D Gaussian splatting to create high-fidelity avatars. To associate 3D Gaussians with the animatable avatar, we learn a parametric template from the input videos, and then parameterize the template on two front & back canonical Gaussian maps where each pixel represents a 3D Gaussian. The learned template is adaptive to the wearing garments for modeling looser clothes like dresses. Such template-guided 2D parameterization enables us to employ a powerful StyleGAN-based CNN to learn the pose-dependent Gaussian maps for modeling detailed dynamic appearances. Furthermore, we introduce a pose projection strategy for better generalization given novel poses. Overall, our method can create lifelike avatars with dynamic, realistic and generalized appearances. Experiments show that our method outperforms other state-of-the-art approaches. Zhe Li 0027, Zerong Zheng, Lizhen Wang 0002, Yebin Liu |
CVPR | 2 |
| 2024 | Control4D: Efficient 4D Portrait Editing With TextabstractWe introduce Control4D, an innovative framework for editing dynamic 4D portraits using text instructions. Our method addresses the prevalent challenges in 4D editing, notably the inefficiencies of existing 4D representations and the inconsistent editing effect caused by diffusion-based editors. We first propose GaussianPlanes, a novel 4D representation that makes Gaussian Splatting more structured by applying plane-based decomposition in 3D space and time. This enhances both efficiency and robustness in 4D editing. Furthermore, we propose to leverage a 4D generator to learn a more continuous generation space from inconsistent edited images produced by the diffusion-based editor, which effectively improves the consistency and quality of 4D editing. Comprehensive evaluation demonstrates the superiority of Control4D, including significantly reduced training time, high-quality rendering, and spatial-temporal consistency in 4D portrait editing. The link to our project website is: https://control4darxiv.github.io/ Ruizhi Shao, Jingxiang Sun, Zerong Zheng, Boyao Zhou, Hongwen Zhang 0001, Yebin Liu |
CVPR | 4 |
| 2024 | DiffPerformer: Iterative Learning of Consistent Latent Guidance for Diffusion-Based Human Video GenerationabstractExisting diffusion models for pose-guided human video generation mostly suffer from temporal inconsistency in the generated appearance and poses due to the inherent randomization nature of the generation process. In this paper, we propose a novel framework, DiffPerformer, to synthesize high-fidelity and temporally consistent human video. Without complex architecture modification or costly training, DiffPerformer finetunes a pre-trained diffusion model on a single video of the target character and introduces an implicit video representation as a proxy to learn temporally consistent guidance for the diffusion model. The guidance is encoded into VAE latent space and an iterative optimization loop is constructed between the implicit video representation and the diffusion model, allowing to harness the smooth property of the implicit video representation and the generative capabilities of the diffusion model in a mutually beneficial way. Moreover, we propose 3D-aware human flow as a temporal constraint during the optimization to explicitly model the correspondence between driving poses and human appearance. This alleviates the mis-alignment between driving poses and target performer and therefore maintains the appearance coherence under various motions. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods. The code is available at https://github.com/aipixel/DiffPerformer. Chenyang Wang 0002, Zerong Zheng, Tao Yu 0007, Xiaoqian Lv, Bineng Zhong 0001, Shengping Zhang, Liqiang Nie |
CVPR | 2 |
| 2024 | Gaussian Head Avatar: Ultra High-Fidelity Head Avatar via Dynamic GaussiansabstractCreating high-fidelity 3D head avatars has always been a research hotspot, but there remains a great challenge under lightweight sparse view setups. In this paper, we propose Gaussian Head Avatar represented by controllable 3D Gaussians for high-fidelity head avatar modeling. We optimize the neutral 3D Gaussians and a fully learned MLP-based deformation field to capture complex expressions. The two parts benefit each other, thereby our method can model fine-grained dynamic details while ensuring expression accuracy. Furthermore, we devise a well-designed geometry-guided initialization strategy based on implicit SDF and Deep Marching Tetrahedra for the stability and convergence of the training procedure. Experiments show our approach outperforms other state-of-the-art sparse-view methods, achieving ultra high-fidelity rendering quality at 2K resolution even under exaggerated expressions. Project page: https://yuelangx.github.io/gaussianheadavatar. Yuelang Xu, Bengwang Chen, Zhe Li 0027, Hongwen Zhang 0001, Lizhen Wang 0002, Zerong Zheng, Yebin Liu |
CVPR | 6 |
| 2024 | MeshAvatar: Learning High-Quality Triangular Human Avatars from Multi-view Videos
Yushuo Chen 0001, Zerong Zheng, Zhe Li 0027, Yebin Liu |
ECCV (68) | 2 |
| 2024 | 3D Gaussian Parametric Head Model
Yuelang Xu, Lizhen Wang 0002, Zerong Zheng, Zhaoqi Su, Yebin Liu |
ECCV (35) | 3 |
| 2024 | Implicit Surface Representation Using Epanechnikov Mixture RegressionabstractWe propose a regression-based implicit surface representation using mixture-of-experts based on the Epanechnikov kernel (EK), a mathematical framework that does not depend on neural networks. The modeling method is implemented using signed distance fields (SDF), modeled using the expectation-maximization algorithm to iterate an optimal set of parameters of Epanechnikov mixture regression. The proposed pipeline achieves better reconstruction than the SDF itself and can be upsampled through mixture-of-experts-based interpolation without extra parameters and processing. Furthermore, the proposed method can efficiently realize data compression compared to meshes and SDF. As for the kernel theory, EK demonstrates a more accurate surface recovery than the Gaussian ones, which expands the applications for Epanechnikov-related theories and also shows potential for theoretical substitution for Gaussian-based modeling and representation. Boning Liu 0001, Zerong Zheng, Yebin Liu |
IEEE Signal Process. Lett. | 2 |
| 2024 | 360-degree Human Video Generation with 4D Diffusion TransformerabstractWe present a novel approach for generating 360-degree high-quality, spatiotemporally coherent human videos from a single image. Our framework combines the strengths of diffusion transformers for capturing global correlations across viewpoints and time, and CNNs for accurate condition injection. The core is a hierarchical 4D transformer architecture that factorizes self-attention across views, time steps, and spatial dimensions, enabling efficient modeling of the 4D space. Precise conditioning is achieved by injecting human identity, camera parameters, and temporal signals into the respective transformers. To train this model, we collect a multi-dimensional dataset spanning images, videos, multi-view data, and limited 4D footage, along with a tailored multi-dimensional training strategy. Our approach overcomes the limitations of previous methods based on generative adversarial networks or vanilla diffusion models, which struggle with complex motions, viewpoint changes, and generalization. Through extensive experiments, we demonstrate our method's ability to synthesize 360-degree realistic, coherent human motion videos, paving the way for advanced multimedia applications in areas such as virtual reality and animation. Ruizhi Shao, Youxin Pang, Zerong Zheng, Jingxiang Sun, Yebin Liu |
ACM Trans. Graph. | 3 |
| 2024 | HVTR++: Image and Pose Driven Human Avatars Using Hybrid Volumetric-Textural RenderingabstractRecent neural rendering methods have made great progress in generating photorealistic human avatars. However, these methods are generally conditioned only on low-dimensional driving signals (e.g., body poses), which are insufficient to encode the complete appearance of a clothed human. Hence they fail to generate faithful details. To address this problem, we exploit driving view images (e.g., in telepresence systems) as additional inputs. We propose a novel neural rendering pipeline, Hybrid Volumetric-Textural Rendering (HVTR++), which synthesizes 3D human avatars from arbitrary driving poses and views while staying faithful to appearance details efficiently and at high quality. First, we learn to encode the driving signals of pose and view image on a dense UV manifold of the human body surface and extract UV-aligned features, preserving the structure of a skeleton-based parametric model. To handle complicated motions (e.g., self-occlusions), we then leverage the UV-aligned features to construct a 3D volumetric representation based on a dynamic neural radiance field. While this allows us to represent 3D geometry with changing topology, volumetric rendering is computationally heavy. Hence we employ only a rough volumetric representation using a pose- and image-conditioned downsampled neural radiance field (PID-NeRF), which we can render efficiently at low resolutions. In addition, we learn 2D textural features that are fused with rendered volumetric features in image space. The key advantage of our approach is that we can then convert the fused features into a high-resolution, high-quality avatar by a fast GAN-based textural renderer. We demonstrate that hybrid rendering enables HVTR++ to handle complicated motions, render high-quality avatars under user-controlled poses/shapes, and most importantly, be efficient at inference time. Our experimental results also demonstrate state-of-the-art quantitative results. Tao Hu 0006, Linjie Luo, Tao Yu 0007, Zerong Zheng, He Zhang 0015, Yebin Liu, Matthias Zwicker |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | CloSET: Modeling Clothed Humans on Continuous Surface with Explicit Template DecompositionabstractCreating animatable avatars from static scans requires the modeling of clothing deformations in different poses. Existing learning-based methods typically add pose-dependent deformations upon a minimally-clothed mesh template or a learned implicit template, which have limitations in capturing details or hinder end-to-end learning. In this paper, we revisit point-based solutions and propose to decompose explicit garment-related templates and then add pose-dependent wrinkles to them. In this way, the clothing deformations are disentangled such that the pose-dependent wrinkles can be better learned and applied to unseen poses. Additionally, to tackle the seam artifact issues in recent state-of-the-art point-based methods, we propose to learn point features on a body surface, which establishes a continuous and compact feature space to capture the fine-grained and pose-dependent clothing geometry. To facilitate the research in this field, we also introduce a high-quality scan dataset of humans in real-world clothing. Our approach is validated on two existing datasets and our newly introduced dataset, showing better clothing deformation results in unseen poses. The project page with code and dataset can be found at https://www.liuyebin.com/closet. Hongwen Zhang 0001, Siyou Lin, Ruizhi Shao, Yuxiang Zhang 0006, Zerong Zheng, Han Huang 0005, Yandong Guo, Yebin Liu |
CVPR | 5 |
| 2023 | Tensor4D: Efficient Neural 4D Decomposition for High-Fidelity Dynamic Reconstruction and RenderingabstractWe present Tensor4D, an efficient yet effective approach to dynamic scene modeling. The key of our solution is an efficient 4D tensor decomposition method so that the dynamic scene can be directly represented as a 4D spatio-temporal tensor. To tackle the accompanying memory issue, we decompose the 4D tensor hierarchically by projecting it first into three time-aware volumes and then nine compact feature planes. In this way, spatial information over time can be simultaneously captured in a compact and memory-efficient manner. When applying Tensor4D for dynamic scene reconstruction and rendering, we further factorize the 4D fields to different scales in the sense that structural motions and dynamic detailed changes can be learned from coarse to fine. The effectiveness of our method is validated on both synthetic and real-world scenes. Extensive experiments show that our method is able to achieve high-quality dynamic reconstruction and rendering from sparse-view camera rigs or even a monocular camera. The code and dataset will be released at https://github.com/DSaurus/Tensor4D. Ruizhi Shao, Zerong Zheng, Hanzhang Tu, Boning Liu 0001, Hongwen Zhang 0001, Yebin Liu |
CVPR | 2 |
| 2023 | Leveraging Intrinsic Properties for Non-Rigid Garment AlignmentabstractWe address the problem of aligning real-world 3D data of garments, which benefits many applications such as texture learning, physical parameter estimation, generative modeling of garments, etc. Existing extrinsic methods typically perform non-rigid iterative closest point and struggle to align details due to incorrect closest matches and rigidity constraints. While intrinsic methods based on functional maps can produce high-quality correspondences, they work under isometric assumptions and become unreliable for garment deformations which are highly non-isometric. To achieve wrinkle-level as well as texture-level alignment, we present a novel coarse-to-fine two-stage method that leverages intrinsic manifold properties with two neural deformation fields, in the 3D space and the intrinsic space, respectively. The coarse stage performs a 3D fitting, where we leverage intrinsic manifold properties to define a manifold deformation field. The coarse fitting then induces a functional map that produces an alignment of intrinsic embeddings. We further refine the intrinsic alignment with a second neural deformation field for higher accuracy. We evaluate our method with our captured garment dataset, GarmCap. The method achieves accurate wrinkle-level and texture-level alignment and works for difficult garment types such as long coats. Our project page is https://jsnln.github.io/iccv2023intrinsic/index.html. Siyou Lin, Boyao Zhou, Zerong Zheng, Hongwen Zhang 0001, Yebin Liu |
ICCV | 3 |
| 2023 | AvatarReX: Real-time Expressive Full-body AvatarsabstractWe present AvatarReX, a new method for learning NeRF-based full-body avatars from video data. The learnt avatar not only provides expressive control of the body, hands and the face together, but also supports real-time animation and rendering. To this end, we propose a compositional avatar representation, where the body, hands and the face are separately modeled in a way that the structural prior from parametric mesh templates is properly utilized without compromising representation flexibility. Furthermore, we disentangle the geometry and appearance for each part. With these technical designs, we propose a dedicated deferred rendering pipeline, which can be executed at a real-time framerate to synthesize high-quality free-view images. The disentanglement of geometry and appearance also allows us to design a two-pass training strategy that combines volume rendering and surface rendering for network training. In this way, patch-level supervision can be applied to force the network to learn sharp appearance details on the basis of geometry estimation. Overall, our method enables automatic construction of expressive full-body avatars with real-time rendering capability, and can generate photo-realistic images with dynamic details for novel body motions and facial expressions. Zerong Zheng, Xiaochen Zhao, Hongwen Zhang 0001, Boning Liu 0001, Yebin Liu |
ACM Trans. Graph. | 1 |
| 2022 | HVTR: Hybrid Volumetric-Textural Rendering for Human AvatarsabstractWe propose a novel neural rendering pipeline, Hybrid Volumetric-Textural Rendering (HVTR), which synthesizes virtual human avatars from arbitrary poses efficiently and at high quality. First, we learn to encode articulated human motions on a dense UV manifold of the human body surface. To handle complicated motions (e.g., self-occlusions), we then leverage the encoded information on the UV manifold to construct a 3D volumetric representation based on a dynamic pose-conditioned neural radiance field. While this allows us to represent 3D geometry with changing topology, volumetric rendering is computationally heavy. Hence we employ only a rough volumetric representation using a pose-conditioned downsampled neural radiance field (PD-NeRF), which we can render efficiently at low resolutions. In addition, we learn 2D textural features that are fused with rendered volumetric features in image space. The key advantage of our approach is that we can then convert the fused features into a high-resolution, high-quality avatar by a fast GAN-based textural renderer. We demonstrate that hybrid rendering enables HVTR to handle complicated motions, render high-quality avatars under user-controlled poses/shapes and even loose clothing, and most importantly, be efficient at inference time. Our experimental results also demonstrate state-of-the-art quantitative results. More results are available at our project page: https://www.cs.umd.edu/~taohu/hvtr/ Tao Hu 0006, Tao Yu 0007, Zerong Zheng, He Zhang 0015, Yebin Liu, Matthias Zwicker |
3DV | 3 |
| 2022 | High-Fidelity Human Avatars from a Single RGB CameraabstractIn this paper, we propose a coarse-to-fine framework to reconstruct a personalized high-fidelity human avatar from a monocular video. To deal with the misalignment problem caused by the changed poses and shapes in different frames, we design a dynamic surface network to recover pose-dependent surface deformations, which help to decouple the shape and texture of the person. To cope with the complexity of textures and generate photo-realistic results, we propose a reference-based neural rendering network and exploit a bottom-up sharpening-guided fine-tuning strategy to obtain detailed textures. Our frame-work also enables photo-realistic novel view/pose syn-thesis and shape editing applications. Experimental re-sults on both the public dataset and our collected dataset demonstrate that our method outperforms the state-of-the-art methods. The code and dataset will be available at http://cic.tju.edu.cn/faculty/likun/projects/HF-Avatar. Yukun Lai, Zerong Zheng, Yingdi Xie, Yebin Liu, Kun Li 0001 |
CVPR | 4 |
| 2022 | Structured Local Radiance Fields for Human Avatar ModelingabstractIt is extremely challenging to create an animatable clothed human avatar from RGB videos, especially for loose clothes due to the difficulties in motion modeling. To address this problem, we introduce a novel representation on the basis of recent neural scene rendering techniques. The core of our representation is a set of structured local radiance fields, which are anchored to the pre-defined nodes sampled on a statistical human body template. These local radiance fields not only leverage the flexibility of implicit representation in shape and appearance modeling, but also factorize cloth deformations into skeleton motions, node residual translations and the dynamic detail variations inside each individual radiance field. To learn our representation from RGB data and facilitate pose generalization, we propose to learn the node translations and the detail variations in a conditional generative latent space. Overall, our method enables automatic construction of animatable human avatars for various types of clothes without the need for scanning subject-specific templates, and can generate realistic images with dynamic details for novel poses. Experiment show that our method outperforms state-of-the-art methods both qualitatively and quantitatively. Zerong Zheng, Han Huang 0005, Tao Yu 0007, Hongwen Zhang 0001, Yandong Guo, Yebin Liu |
CVPR | 1 |
| 2022 | AvatarCap: Animatable Avatar Conditioned Monocular Human Volumetric Capture
Zhe Li 0027, Zerong Zheng, Hongwen Zhang 0001, Chaonan Ji, Yebin Liu |
ECCV (1) | 2 |
| 2022 | Learning Implicit Templates for Point-Based Clothed Human Modeling
Siyou Lin, Hongwen Zhang 0001, Zerong Zheng, Ruizhi Shao, Yebin Liu |
ECCV (3) | 3 |
| 2022 | DiffuStereo: High Quality Human Reconstruction via Diffusion-Based Stereo Using Sparse Cameras
Ruizhi Shao, Zerong Zheng, Hongwen Zhang 0001, Jingxiang Sun, Yebin Liu |
ECCV (32) | 2 |
| 2022 | FloRen: Real-time High-quality Human Performance Rendering via Appearance Flow Using Sparse RGB CamerasabstractWe propose FloRen, a novel system for real-time, high-resolution free-view human synthesis. Our system runs at 15fps in 1K resolution with very sparse RGB cameras. In FloRen, a coarse-level implicit geometry is recovered at first as initialization, and then processed by a neural rendering framework based on appearance flow. Our appearance flow-based rendering framework consists of three steps, namely view-dependent depth refinement, appearance flow estimation and occlusion-aware color rendering. In this way, we resolve the view synthesis problem in the image plane, where 2D convolutional neural networks can be efficiently applied, contributing to high speed performance. For robust appearance flow estimation, we explicitly combine data-driven human prior knowledge with multiview geometric constraints. The accurate appearance flow enables precise color mapping from input view to novel view, which greatly facilitates high-resolution novel view generation. We demonstrate that our system achieves state-of-the-art performance and even outperforms many offline methods. Ruizhi Shao, Liliang Chen, Zerong Zheng, Hongwen Zhang 0001, Yuxiang Zhang 0006, Han Huang 0005, Yandong Guo, Yebin Liu |
SIGGRAPH Asia | 3 |
| 2022 | Robust and Accurate 3D Self-Portraits in SecondsabstractIn this paper, we propose an efficient method for robust and accurate 3D self-portraits using a single RGBD camera. Our method can generate detailed and realistic 3D self-portraits in seconds and shows the ability to handle subjects wearing extremely loose clothes. To achieve highly efficient and robust reconstruction, we propose PIFusion, which combines learning-based 3D recovery with volumetric non-rigid fusion to generate accurate sparse partial scans of the subject. Meanwhile, a non-rigid volumetric deformation method is proposed to continuously refine the learned shape prior. Moreover, a lightweight bundle adjustment algorithm is proposed to guarantee that all the partial scans can not only "loop" with each other but also remain consistent with the selected live key observations. Finally, to further generate realistic portraits, we propose non-rigid texture optimization to improve the texture quality. Additionally, we also contribute a benchmark for single-view 3D self-portrait reconstruction, an evaluation dataset that contains 10 single-view RGBD sequences of a self-rotating performer wearing various clothes and the corresponding ground-truth 3D models in the first frame of each sequence. The results and experiments based on this dataset show that the proposed method outperforms state-of-the-art methods on accuracy, efficiency, and generality. Zhe Li 0027, Tao Yu 0007, Zerong Zheng, Yebin Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | PaMIR: Parametric Model-Conditioned Implicit Representation for Image-Based Human ReconstructionabstractModeling 3D humans accurately and robustly from a single image is very challenging, and the key for such an ill-posed problem is the 3D representation of the human models. To overcome the limitations of regular 3D representations, we propose Parametric Model-Conditioned Implicit Representation (PaMIR), which combines the parametric body model with the free-form deep implicit function. In our PaMIR-based reconstruction framework, a novel deep neural network is proposed to regularize the free-form deep implicit function using the semantic features of the parametric model, which improves the generalization ability under the scenarios of challenging poses and various clothing topologies. Moreover, a novel depth-ambiguity-aware training loss is further integrated to resolve depth ambiguities and enable successful surface detail reconstruction with imperfect body reference. Finally, we propose a body reference optimization method to improve the parametric model estimation accuracy and to enhance the consistency between the parametric model and the implicit function. With the PaMIR representation, our framework can be easily extended to multi-image input scenarios without the need of multi-camera calibration and pose synchronization. Experimental results demonstrate that our method achieves state-of-the-art performance for image-based 3D human reconstruction in the cases of challenging poses and clothing types. Zerong Zheng, Tao Yu 0007, Yebin Liu, Qionghai Dai |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Function4D: Real-Time Human Volumetric Capture From Very Sparse Consumer RGBD SensorsabstractHuman volumetric capture is a long-standing topic in computer vision and computer graphics. Although high-quality results can be achieved using sophisticated off-line systems, real-time human volumetric capture of complex scenarios, especially using light-weight setups, remains challenging. In this paper, we propose a human volumetric capture method that combines temporal volumetric fusion and deep implicit functions. To achieve high-quality and temporal-continuous reconstruction, we propose dynamic sliding fusion to fuse neighboring depth observations together with topology consistency. Moreover, for detailed and complete surface generation, we propose detailpreserving deep implicit functions for RGBD input which can not only preserve the geometric details on the depth inputs but also generate more plausible texturing results. Results and experiments show that our method outperforms existing methods in terms of view sparsity, generalization capacity, reconstruction quality, and run-time efficiency. Tao Yu 0007, Zerong Zheng, Qionghai Dai, Yebin Liu |
CVPR | 2 |
| 2021 | POSEFusion: Pose-Guided Selective Fusion for Single-View Human Volumetric CaptureabstractWe propose POseguided SElective Fusion (POSEFu-sion), a single-view human volumetric capture method that leverages tracking-based methods and tracking-free inference to achieve high-fidelity and dynamic 3D reconstruction. By contributing a novel reconstruction framework which contains pose-guided keyframe selection and robust implicit surface fusion, our method fully utilizes the advantages of both tracking-based methods and tracking-free inference methods, and finally enables the high-fidelity recon-struction of dynamic surface details even in the invisible regions. We formulate the keyframe selection as a dynamic programming problem to guarantee the temporal continuity of the reconstructed sequence. Moreover, the novel robust implicit surface fusion involves an adaptive blending weight to preserve high-fidelity surface details and an automatic collision handling method to deal with the potential self-collisions. Overall, our method enables high-fidelity and dynamic capture in both visible and invisible regions from a single RGBD camera, and the results and experiments show that our method outperforms state-of-the-art methods. Zhe Li 0027, Tao Yu 0007, Zerong Zheng, Yebin Liu |
CVPR | 3 |
| 2021 | Deep Implicit Templates for 3D Shape RepresentationabstractDeep implicit functions (DIFs), as a kind of 3D shape representation, are becoming more and more popular in the 3D vision community due to their compactness and strong representation power. However, unlike polygon mesh-based templates, it remains a challenge to reason dense correspondences or other semantic relationships across shapes represented by DIFs, which limits its applications in texture transfer, shape analysis and so on. To overcome this limitation and also make DIFs more interpretable, we propose Deep Implicit Templates, a new 3D shape representation that supports explicit correspondence reasoning in deep implicit representations. Our key idea is to formulate DIFs as conditional deformations of a template implicit function. To this end, we propose Spatial Warping LSTM, which de-composes the conditional spatial transformation into multiple point-wise transformations and guarantees generalization capability. Moreover, the training loss is carefully designed in order to achieve high reconstruction accuracy while learning a plausible template with accurate correspondences in an unsupervised manner. Experiments show that our method can not only learn a common implicit tem-plate for a collection of shapes, but also establish dense correspondences across all the shapes simultaneously with-out any supervision. Zerong Zheng, Tao Yu 0007, Qionghai Dai, Yebin Liu |
CVPR | 1 |
| 2021 | DeepMultiCap: Performance Capture of Multiple Characters Using Sparse Multiview CamerasabstractWe propose DeepMultiCap, a novel method for multi-person performance capture using sparse multi-view cameras. Our method can capture time varying surface details without the need of using pre-scanned template models. To tackle with the serious occlusion challenge for close interacting scenes, we combine a recently proposed pixel-aligned implicit function with parametric model for robust reconstruction of the invisible surface areas. An effective attention-aware module is designed to obtain the fine-grained geometry details from multi-view images, where high-fidelity results can be generated. In addition to the spatial attention method, for video inputs, we further propose a novel temporal fusion method to alleviate the noise and temporal inconsistencies for moving character reconstruction. For quantitative evaluation, we contribute a high quality multi-person dataset, MultiHuman, which consists of 150 static scenes with different levels of occlusions and ground truth 3D human models. Experimental results demonstrate the state-of-the-art performance of our method and the well generalization to real multiview video data, which outperforms the prior works by a large margin. Ruizhi Shao, Yuxiang Zhang 0006, Tao Yu 0007, Zerong Zheng, Qionghai Dai, Yebin Liu |
ICCV | 5 |
| 2020 | Robust 3D Self-Portraits in SecondsabstractIn this paper, we propose an efficient method for robust 3D self-portraits using a single RGBD camera. Benefiting from the proposed PIFusion and lightweight bundle adjustment algorithm, our method can generate detailed 3D self-portraits in seconds and shows the ability to handle subjects wearing extremely loose clothes. To achieve highly efficient and robust reconstruction, we propose PIFusion, which combines learning-based 3D recovery with volumetric non-rigid fusion to generate accurate sparse partial scans of the subject. Moreover, a non-rigid volumetric deformation method is proposed to continuously refine the learned shape prior. Finally, a lightweight bundle adjustment algorithm is proposed to guarantee that all the partial scans can not only ``loop'' with each other but also remain consistent with the selected live key observations. The results and experiments show that the proposed method achieves more robust and efficient 3D self-portraits compared with state-of-the-art methods. Zhe Li 0027, Tao Yu 0007, Chuanyu Pan, Zerong Zheng, Yebin Liu |
CVPR | 4 |
| 2020 | RobustFusion: Human Volumetric Capture with Data-Driven Visual Cues Using a RGBD Camera
Zhuo Su 0006, Lan Xu 0003, Zerong Zheng, Tao Yu 0007, Yebin Liu, Lu Fang 0001 |
ECCV (4) | 3 |
| 2020 | DoubleFusion: Real-Time Capture of Human Performances with Inner Body Shapes from a Single Depth SensorabstractWe propose DoubleFusion, a new real-time system that combines volumetric non-rigid reconstruction with data-driven template fitting to simultaneously reconstruct detailed surface geometry, large non-rigid motion and the optimized human body shape from a single depth camera. One of the key contributions of this method is a double-layer representation consisting of a complete parametric body model inside, and a gradually fused detailed surface outside. A pre-defined node graph on the body parameterizes the non-rigid deformations near the body, and a free-form dynamically changing graph parameterizes the outer surface layer far from the body, which allows more general reconstruction. We further propose a joint motion tracking method based on the double-layer representation to enable robust and fast motion tracking performance. Moreover, the inner parametric body is optimized online and forced to fit inside the outer surface layer as well as the live depth input. Overall, our method enables increasingly denoised, detailed and complete surface reconstructions, fast motion tracking performance and plausible inner body shape reconstruction in real-time. Experiments and comparisons show improved fast motion tracking and loop closure performance on more challenging scenarios. Two extended applications including body measurement and shape retargeting show the potential of our system in terms of practical use. Tao Yu 0007, Jianhui Zhao 0002, Zerong Zheng, Qionghai Dai, Hao Li 0015, Gerard Pons-Moll, Yebin Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | SimulCap : Single-View Human Performance Capture With Cloth SimulationabstractThis paper proposes a new method for live free-viewpoint human performance capture with dynamic details (e.g., cloth wrinkles) using a single RGBD camera. Our main contributions are: (i) a multi-layer representation of garments and body, and (ii) a physics-based performance capture procedure. We first digitize the performer using multi-layer surface representation, which includes the undressed body surface and separate clothing meshes. For performance capture, we perform skeleton tracking, cloth simulation, and iterative depth fitting sequentially for the incoming frame. By incorporating cloth simulation into the performance capture pipeline, we can simulate plausible cloth dynamics and cloth-body interactions even in the occluded regions, which was not possible in previous capture methods. Moreover, by formulating depth fitting as a physical process, our system produces cloth tracking results consistent with the depth observation while still maintaining physical constraints. Results and evaluations show the effectiveness of our method. Our method also enables new types of applications such as cloth retargeting, free-viewpoint video rendering and animations. Tao Yu 0007, Zerong Zheng, Jianhui Zhao 0002, Qionghai Dai, Gerard Pons-Moll, Yebin Liu |
CVPR | 2 |
| 2019 | DeepHuman: 3D Human Reconstruction From a Single ImageabstractWe propose DeepHuman, an image-guided volume-to-volume translation CNN for 3D human reconstruction from a single RGB image. To reduce the ambiguities associated with the reconstruction of invisible areas, our method leverages a dense semantic representation generated from SMPL model as an additional input. One key feature of our network is that it fuses different scales of image features into the 3D space through volumetric feature transformation, which helps to recover accurate surface geometry. The surface details are further refined through a normal refinement network, which can be concatenated with the volume generation network using our proposed volumetric normal projection layer. We also contribute THuman, a 3D real-world human model dataset containing approximately 7000 models. The network is trained using training data generated from the dataset. Overall, due to the specific design of our network and the diversity in our dataset, our method enables 3D human model estimation given only a single image and outperforms state-of-the-art approaches. Zerong Zheng, Tao Yu 0007, Yixuan Wei, Qionghai Dai, Yebin Liu |
ICCV | 1 |
| 2018 | DoubleFusion: Real-Time Capture of Human Performances With Inner Body Shapes From a Single Depth SensorabstractWe propose DoubleFusion, a new real-time system that combines volumetric dynamic reconstruction with data-driven template fitting to simultaneously reconstruct detailed geometry, non-rigid motion and the inner human body shape from a single depth camera. One of the key contributions of this method is a double layer representation consisting of a complete parametric body shape inside, and a gradually fused outer surface layer. A pre-defined node graph on the body surface parameterizes the non-rigid deformations near the body, and a free-form dynamically changing graph parameterizes the outer surface layer far from the body, which allows more general reconstruction. We further propose a joint motion tracking method based on the double layer representation to enable robust and fast motion tracking performance. Moreover, the inner body shape is optimized online and forced to fit inside the outer surface layer. Overall, our method enables increasingly denoised, detailed and complete surface reconstructions, fast motion tracking performance and plausible inner body shape reconstruction in real-time. In particular, experiments show improved fast motion tracking and loop closure performance on more challenging scenarios. Tao Yu 0007, Zerong Zheng, Jianhui Zhao 0002, Qionghai Dai, Hao Li 0015, Gerard Pons-Moll, Yebin Liu |
CVPR | 2 |
| 2018 | HybridFusion: Real-Time Performance Capture Using a Single Depth Sensor and Sparse IMUs
Zerong Zheng, Tao Yu 0007, Hao Li 0015, Qionghai Dai, Lu Fang 0001, Yebin Liu |
ECCV (9) | 1 |