EDBT 2026 Demo / reviewers in the wild / expert
Yuxiang Zhang 0006
dblp:73/7697-6
· DBLP profile ↗
18ranked-venue papers
3as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 15 since 2021Artificial intelligence and machine learning · 12 · 3 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Consistent 3D Human Reconstruction From Monocular Video: Learning Correctable Appearance and Temporal Motion PriorsabstractRecent advancements in rendering dynamic humans using NeRF and 3D Gaussian splatting have made significant progress, leveraging implicit geometry learning and image appearance rendering to create digital humans. However, in monocular video rendering, there are still challenges in rendering subtle and complex motion from different viewpoints and states, primarily due to the imbalance of viewpoints. Additionally, ensuring continuity between adjacent frames when rendering from novel and free viewpoints remains a difficult task. To address these challenges, we first propose a pixel-level motion correction module that adjusts the errors in the learned representation between different viewpoints. We also introduce a temporal information-based model to improve motion continuity by leveraging adjacent frames. Experimental results on dynamic human rendering, using the NeuMan, ZJU-Mocap, and People-Snapshot datasets, demonstrate that our method outperforms state-of-the-art techniques both quantitatively and qualitatively. Cheng Shang, Liang An 0001, Jiajun Zhang 0012, Yuxiang Zhang 0006, Jidong Tian, Yebin Liu, Xubo Yang |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2026 | DanceAgent: Dance Movement Refinement With LLM AgentabstractRecent research on motion generation and text-to-motion synthesis focus on coarse-grained motion descriptions, neglecting fine-grained motion details and motion quality refinement. Additionally, current text-to-motion models, such as MotionGPT, lack multi-turn interaction capabilities, relying on single-turn and single-modality transformations, which limit their ability to integrate information from different modalities across interaction stages. These gaps leave critical questions, such as "How well is the motion performed" and "How can it be refined?" largely unaddressed. To address these issues, first, we introduce two fine-grained dance datasets-one focusing on jazz dance and the other on folk dance, which we have independently collected. Second, considering that dance motions are inherently complex and consist of long sequential actions, we introduce both global and local optimization during the motion encoding phase and employ Hidden Markov Model (HMM) temporal modeling to capture differential features between correct and incorrect movements, thereby optimizing the training process. Finally, we propose a multi-turn historical dialogue framework that enables three stages generation-motion assess, text instructions, and motion refinement-for input videos. This framework assists dance beginners by providing feedback on their movements, offering textual instructions, and delivering motion-based refinement. Experimental results on the jazz dance and folk dance datasets demonstrate that our method surpasses existing approaches in both quantitative and qualitative metrics, establishing a new benchmark for motion-text generation in the field of dance training. Cheng Shang, Liang An 0001, Jiajun Zhang 0012, Yuxiang Zhang 0006, Yebin Liu, Xubo Yang |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | Ins-HOI: Instance Aware Human-Object Interactions RecoveryabstractAccurately modeling detailed interactions between human/hand and object is an appealing yet challenging task. Current multi-view capture systems are only capable of reconstructing multiple subjects into a single, unified mesh, which fails to model the states of each instance individually during interactions. To address this, previous methods use template-based representations to track human/hand and object. However, the quality of the reconstructions is limited by the descriptive capabilities of the templates so these methods inherently struggle with geometric details, pressing deformations and invisible contact surfaces. In this work, we propose an end-to-end Instance-aware Human-Object Interactions recovery (Ins-HOI) framework by introducing an instance-level occupancy field representation. However, the real-captured data is presented as a holistic mesh, unable to provide instance-level supervision. To address this, we further propose a complementary training strategy that leverages synthetic data to introduce instance-level shape priors, enabling the disentanglement of occupancy fields for different instances. Specifically, synthetic data, created by randomly combining individual scans of humans/hands and objects, guides the network to learn a coarse prior of instances. Meanwhile, real-captured data helps in learning the overall geometry and restricting interpenetration in contact areas. As demonstrated in experiments, our method Ins-HOI supports instance-level reconstruction and provides reasonable and realistic invisible contact surfaces even in cases of extremely close interaction. To facilitate research on this task, we collect a large-scale, high-fidelity 3D scan dataset, including 5.2 k high-quality scans with real-world human-chair and hand-object interactions. The code and data will be public for research purposes. Jiajun Zhang 0012, Yuxiang Zhang 0006, Hongwen Zhang 0001, Xiao Zhou 0019, Boyao Zhou, Ruizhi Shao, Zonghai Hu, Yebin Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | SpeechAct: Towards Generating Whole-Body Motion From SpeechabstractWhole-body motion generation from speech audio is crucial for computer graphics and immersive VR/AR. Prior methods struggle to produce natural and diverse whole-body motions from speech. In this paper, we introduce a novel method, named SpeechAct, based on a hybrid point representation and contrastive motion learning to boost realism and diversity in motion generation. Our hybrid point representation leverages the advantages of keypoint representation and surface points of 3D body model, which is easy to learn and helps to achieve smooth and natural motion generation from speech audio. We design a VQ-VAE to learn a motion codebook using our hybrid presentation, and then regress the motion from the input audio using a translation model. To boost diversity in motion generation, we propose a contrastive motion learning method according to the intuitive idea that the generated motion should be different from the motions of other audios and other speakers. We collect negative samples from other audio inputs and other speakers using our translation model. With these negative samples, we pull the current motion away from them using a contrastive loss to produce more distinctive representations. In addition, we compose a face generator to generate deterministic face motion due to the strong connection between the face movements and the speech audio. Experimental results validate the superior performance of our model. The code is available at http://cic.tju.edu.cn/faculty/likun/projects/SpeechAct/index.html. Minjie Zhu, Yuxiang Zhang 0006, Zerong Zheng, Yebin Liu, Kun Li 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | ProxyCap: Real-Time Monocular Full-Body Capture in World Space via Human-Centric Proxy-to-Motion LearningabstractLearning-based approaches to monocular motion capture have recently shown promising results by learning to regress in a data-driven manner. However, due to the challenges in data collection and network designs, it remains challenging to achieve real-time full-body capture while being accurate in world space. In this work, we introduce ProxyCap, a human-centric proxy-to-motion learning scheme to learn world-space motions from a proxy dataset of 2D skeleton sequences and 3D rotational motions. Such proxy data enables us to build a learning-based network with accurate world-space supervision while also mitigating the generalization issues. For more accurate and physically plausible predictions in world space, our network is designed to learn human motions from a human-centric perspective, which enables the understanding of the same motion captured with different camera trajectories. Moreover, a contact-aware neural motion descent module is proposed to improve foot-ground contact and motion misalignment with the proxy observations. With the proposed learning-based solution, we demonstrate the first real-time monocular full-body capture system with plausible foot-ground contact in world space even using hand-held cameras. Yuxiang Zhang 0006, Hongwen Zhang 0001, Liangxiao Hu, Jiajun Zhang 0012, Hongwei Yi, Shengping Zhang, Yebin Liu |
CVPR | 1 |
| 2024 | RAM-Avatar: Real-time Photo-Realistic Avatar from Monocular Videos with Full-body ControlabstractThis paper focuses on advancing the applicability of hu-man avatar learning methods by proposing RAM-Avatar, which learns a Real-time, photo-realistic Avatar that sup-ports full-body control from Monocular videos. To achieve this goal, RAM-Avatar leverages two statistical templates responsible for modeling the facial expression and hand gesture variations, while a sparsely computed dual attention module is introduced upon another body template to facilitate high-fidelity texture rendering for the torsos and limbs. Building on this foundation, we deploy a lightweight yet powerful StyleUnet along with a temporal-aware dis-criminator to achieve real-time realistic rendering. To en-able robust animation for out-of-distribution poses, we pro-pose a Motion Distribution Align module to compensate for the discrepancies between the training and testing motion distribution. Results and extensive experiments conducted in various experimental settings demonstrate the superior-ity of our proposed method, and a real-time live system is proposed to further push research into applications. The training and testing code will be released for research pur-poses. Zerong Zheng, Yuxiang Zhang 0006, Jingxiang Sun, Lizhen Wang 0002, Yebin Liu |
CVPR | 3 |
| 2024 | GaussianAvatar: Towards Realistic Human Avatar Modeling from a Single Video via Animatable 3D GaussiansabstractWe present GaussianAvatar, an efficient approach to cre-ating realistic human avatars with dynamic 3D appear-ances from a single video. We start by introducing animat-able 3D Gaussians to explicitly represent humans in var-ious poses and clothing styles. Such an explicit and ani-matable representation can fuse 3D appearances more effi-ciently and consistently from 2D observations. Our repre-sentation is further augmented with dynamic properties to support pose-dependent appearance modeling, where a dy-namic appearance network along with an optimizable feature tensor is designed to learn the motion-to-appearance mapping. Moreover, by leveraging the differentiable motion condition, our method enables a joint optimization of motions and appearances during avatar modeling, which helps to tackle the long-standing issue of inaccurate motion esti-mation in monocular settings. The efficacy of GaussianA-vatar is validated on both the public dataset and our col-lected dataset, demonstrating its superior performances in terms of appearance quality and rendering efficiency. The code and dataset are available at https://github.com/aipixel/GaussianAvatar. Liangxiao Hu, Hongwen Zhang 0001, Yuxiang Zhang 0006, Boyao Zhou, Boning Liu 0001, Shengping Zhang, Liqiang Nie |
CVPR | 3 |
| 2024 | HHMR: Holistic Hand Mesh Recovery by Enhancing the Multimodal Controllability of Graph Diffusion ModelsabstractRecent years have witnessed a trend of the deep integration of the generation and reconstruction paradigms. In this paper, we extend the ability of controllable generative models for a more comprehensive hand mesh recovery task: di-rect hand mesh generation, inpainting, reconstruction, and fitting in a single framework, which we name as Holistic Hand Mesh Recovery (HHMR). Our key observation is that different kinds of hand mesh recovery tasks can be achieved by a single generative model with strong multimodal con-trollability, and in such a framework, realizing different tasks only requires giving different signals as conditions. To achieve this goal, we propose an all-in-one diffusion frame-work based on graph convolution and attention mechanisms for holistic hand mesh recovery. In order to achieve strong control generation capability while ensuring the decoupling of multimodal control signals, we map different modalities to a shared feature space and apply cross-scale random masking in both modality and feature levels. In this way, the correlation between different modalities can be fully exploited during the learning of hand priors. Furthermore, we propose Condition-aligned Gradient Guidance to enhance the alignment of the generated model with the control sig-nals, which significantly improves the accuracy of the hand mesh reconstruction and fitting. Experiments show that our novel framework can realize multiple hand mesh recovery tasks simultaneously and outperform the existing methods in different tasks, which provides more possibilities for sub-sequent downstream applications including gesture recognition, pose generation, mesh editing, and so on. Mengcheng Li, Hongwen Zhang 0001, Yuxiang Zhang 0006, Ruizhi Shao, Tao Yu 0007, Yebin Liu |
CVPR | 3 |
| 2024 | Lodge: A Coarse to Fine Diffusion Network for Long Dance Generation Guided by the Characteristic Dance PrimitivesabstractWe propose Lodge, a network capable of generating extremely long dance sequences conditioned on given music. We design Lodge as a two-stage coarse to fine diffusion architecture, and propose the characteristic dance primitives that possess significant expressiveness as intermediate representations between two diffusion models. The first stage is global diffusion, which focuses on comprehending the coarse-level music-dance correlation and production characteristic dance primitives. In contrast, the second-stage is the local diffusion, which parallelly generates detailed motion sequences under the guidance of the dance primitives and choreographic rules. In addition, we propose a Foot Refine Block to optimize the contact between the feet and the ground, enhancing the physical realism of the motion. Our approach can parallelly generate dance sequences of extremely long length, striking a balance between global choreographic patterns and local motion quality and expressiveness. Extensive experiments validate the efficacy of our method. Code, models, and demonstrative video results are available at: https://li-ronghui.github.io/lodge Ronghui Li, Yuxiang Zhang 0006, Yachao Zhang 0001, Hongwen Zhang 0001, Yan Zhang 0002, Yebin Liu, Xiu Li 0001 |
CVPR | 2 |
| 2024 | TACO: Benchmarking Generalizable Bimanual Tool-ACtion-Object UnderstandingabstractHumans commonly work with multiple objects in daily life and can intuitively transfer manipulation skills to novel objects by understanding object functional regularities. However, existing technical approaches for analyzing and synthesizing hand-object manipulation are mostly limited to handling a single hand and object due to the lack of data support. To address this, we construct TACO, an extensive bimanual hand-object-interaction dataset spanning a large variety of tool-action-object compositions for daily human activities. TACO contains 2.5K motion sequences paired with third-person and egocentric views, precise hand-object 3D meshes, and action labels. To rapidly expand the data scale, we present a fully automatic data acquisition pipeline combining multi-view sensing with an optical motion capture system. With the vast research fields provided by TACO, we benchmark three generalizable hand-object-interaction tasks: compositional action recognition, generalizable hand-object motion forecasting, and cooperative grasp synthesis. Extensive experiments re-veal new insights, challenges, and opportunities for advancing the studies of generalizable hand-object motion anal-ysis and synthesis. Our data and code are available at https://taco2024.github.io. Yun Liu 0018, Xu Si, Yuxiang Zhang 0006, Yebin Liu, Li Yi 0001 |
CVPR | 6 |
| 2023 | CloSET: Modeling Clothed Humans on Continuous Surface with Explicit Template DecompositionabstractCreating animatable avatars from static scans requires the modeling of clothing deformations in different poses. Existing learning-based methods typically add pose-dependent deformations upon a minimally-clothed mesh template or a learned implicit template, which have limitations in capturing details or hinder end-to-end learning. In this paper, we revisit point-based solutions and propose to decompose explicit garment-related templates and then add pose-dependent wrinkles to them. In this way, the clothing deformations are disentangled such that the pose-dependent wrinkles can be better learned and applied to unseen poses. Additionally, to tackle the seam artifact issues in recent state-of-the-art point-based methods, we propose to learn point features on a body surface, which establishes a continuous and compact feature space to capture the fine-grained and pose-dependent clothing geometry. To facilitate the research in this field, we also introduce a high-quality scan dataset of humans in real-world clothing. Our approach is validated on two existing datasets and our newly introduced dataset, showing better clothing deformation results in unseen poses. The project page with code and dataset can be found at https://www.liuyebin.com/closet. Hongwen Zhang 0001, Siyou Lin, Ruizhi Shao, Yuxiang Zhang 0006, Zerong Zheng, Han Huang 0005, Yandong Guo, Yebin Liu |
CVPR | 4 |
| 2023 | PyMAF-X: Towards Well-Aligned Full-Body Model Regression From Monocular ImagesabstractWe present PyMAF-X, a regression-based approach to recovering a parametric full-body model from a single image. This task is very challenging since minor parametric deviation may lead to noticeable misalignment between the estimated mesh and the input image. Moreover, when integrating part-specific estimations into the full-body model, existing solutions tend to either degrade the alignment or produce unnatural wrist poses. To address these issues, we propose a Pyramidal Mesh Alignment Feedback (PyMAF) loop in our regression network for well-aligned human mesh recovery and extend it as PyMAF-X for the recovery of expressive full-body models. The core idea of PyMAF is to leverage a feature pyramid and rectify the predicted parameters explicitly based on the mesh-image alignment status. Specifically, given the currently predicted parameters, mesh-aligned evidence will be extracted from finer-resolution features accordingly and fed back for parameter rectification. To enhance the alignment perception, an auxiliary dense supervision is employed to provide mesh-image correspondence guidance while spatial alignment attention is introduced to enable the awareness of the global contexts for our network. When extending PyMAF for full-body mesh recovery, an adaptive integration strategy is proposed in PyMAF-X to produce natural wrist poses while maintaining the well-aligned performance of the part-specific estimations. The efficacy of our approach is validated on several benchmark datasets for body, hand, face, and full-body mesh recovery, where PyMAF and PyMAF-X effectively improve the mesh-image alignment and achieve new The project page with code and video results can be found at https://www.liuyebin.com/pymaf-x. Hongwen Zhang 0001, Yating Tian, Yuxiang Zhang 0006, Mengcheng Li, Liang An 0001, Zhenan Sun, Yebin Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | FloRen: Real-time High-quality Human Performance Rendering via Appearance Flow Using Sparse RGB CamerasabstractWe propose FloRen, a novel system for real-time, high-resolution free-view human synthesis. Our system runs at 15fps in 1K resolution with very sparse RGB cameras. In FloRen, a coarse-level implicit geometry is recovered at first as initialization, and then processed by a neural rendering framework based on appearance flow. Our appearance flow-based rendering framework consists of three steps, namely view-dependent depth refinement, appearance flow estimation and occlusion-aware color rendering. In this way, we resolve the view synthesis problem in the image plane, where 2D convolutional neural networks can be efficiently applied, contributing to high speed performance. For robust appearance flow estimation, we explicitly combine data-driven human prior knowledge with multiview geometric constraints. The accurate appearance flow enables precise color mapping from input view to novel view, which greatly facilitates high-resolution novel view generation. We demonstrate that our system achieves state-of-the-art performance and even outperforms many offline methods. Ruizhi Shao, Liliang Chen, Zerong Zheng, Hongwen Zhang 0001, Yuxiang Zhang 0006, Han Huang 0005, Yandong Guo, Yebin Liu |
SIGGRAPH Asia | 5 |
| 2022 | Geometry-Guided Dense Perspective Network for Speech-Driven Facial AnimationabstractRealistic speech-driven 3D facial animation is a challenging problem due to the complex relationship between speech and face. In this paper, we propose a deep architecture, called Geometry-guided Dense Perspective Network (GDPnet), to achieve speaker-independent realistic 3D facial animation. The encoder is designed with dense connections to strengthen feature propagation and encourage the re-use of audio features, and the decoder is integrated with an attention mechanism to adaptively recalibrate point-wise feature responses by explicitly modeling interdependencies between different neuron units. We also introduce a non-linear face reconstruction representation as a guidance of latent space to obtain more accurate deformation, which helps solve the geometry-related deformation and is good for generalization across subjects. Huber and HSIC (Hilbert-Schmidt Independence Criterion) constraints are adopted to promote the robustness of our model and to better exploit the non-linear and high-order correlations. Experimental results on the public dataset and real scanned dataset validate the superiority of our proposed GDPnet compared with state-of-the-art model. The code is available for research purposes at http://cic.tju.edu.cn/faculty/likun/projects/GDPnet. Jingying Liu, Binyuan Hui, Kun Li 0001, Yunke Liu, Yukun Lai, Yuxiang Zhang 0006, Yebin Liu, Jing-Yu Yang 0002 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2021 | Lightweight Multi-person Total Motion Capture Using Sparse Multi-view CamerasabstractMulti-person total motion capture is extremely challenging when it comes to handle severe occlusions, different reconstruction granularities from body to face and hands, drastically changing observation scales and fast body movements. To overcome these challenges above, we contribute a lightweight total motion capture system for multi-person interactive scenarios using only sparse multi-view cameras. By contributing a novel hand and face bootstrapping algorithm, our method is capable of efficient localization and accurate association of the hands and faces even on severe occluded occasions. We leverage both pose regression and keypoints detection methods and further propose a unified two-stage parametric fitting method for achieving pixel-aligned accuracy. Moreover, for extremely self-occluded poses and close interactions, a novel feedback mechanism is proposed to propagate the pixel-aligned reconstructions into the next frame for more accurate association. Overall, we propose the first light-weight total capture system and achieves fast, robust and accurate multi-person total motion capture performance. The results and experiments show that our method achieves more accurate results than existing methods under sparse-view setups. Yuxiang Zhang 0006, Zhe Li 0027, Liang An 0001, Mengcheng Li, Tao Yu 0007, Yebin Liu |
ICCV | 1 |
| 2021 | DeepMultiCap: Performance Capture of Multiple Characters Using Sparse Multiview CamerasabstractWe propose DeepMultiCap, a novel method for multi-person performance capture using sparse multi-view cameras. Our method can capture time varying surface details without the need of using pre-scanned template models. To tackle with the serious occlusion challenge for close interacting scenes, we combine a recently proposed pixel-aligned implicit function with parametric model for robust reconstruction of the invisible surface areas. An effective attention-aware module is designed to obtain the fine-grained geometry details from multi-view images, where high-fidelity results can be generated. In addition to the spatial attention method, for video inputs, we further propose a novel temporal fusion method to alleviate the noise and temporal inconsistencies for moving character reconstruction. For quantitative evaluation, we contribute a high quality multi-person dataset, MultiHuman, which consists of 150 static scenes with different levels of occlusions and ground truth 3D human models. Experimental results demonstrate the state-of-the-art performance of our method and the well generalization to real multiview video data, which outperforms the prior works by a large margin. Ruizhi Shao, Yuxiang Zhang 0006, Tao Yu 0007, Zerong Zheng, Qionghai Dai, Yebin Liu |
ICCV | 3 |
| 2021 | Image-Guided Human Reconstruction via Multi-Scale Graph Transformation Networksabstract3D human reconstruction from a single image is a challenging problem. Existing methods have difficulties to infer 3D clothed human models with consistent topologies for various poses. In this paper, we propose an efficient and effective method using a hierarchical graph transformation network. To deal with large deformations and avoid distorted geometries, rather than using Euclidean coordinates directly, 3D human shapes are represented by a vertex-based deformation representation that effectively encodes the deformation and copes well with large deformations. To infer a 3D human mesh consistent with the input real image, we also use a perspective projection layer to incorporate perceptual image features into the deformation representation. Our model is easy to train and fast to converge with short test time. Besides, we present the$D^{2}Human$(Dynamic Detailed Human) dataset, including variously posed 3D human meshes with consistent topologies and rich geometry details, together with the captured color images and SMPL models, which is useful for training and evaluation of deep frameworks, particularly for graph neural networks. Experimental results demonstrate that our method achieves more plausible and complete 3D human reconstruction from a single image, compared with several state-of-the-art methods. The code and dataset are available for research purposes athttp://cic.tju.edu.cn/faculty/likun/projects/MGTnet. Kun Li 0001, Qiao Feng 0001, Yuxiang Zhang 0006, Xiongzheng Li, Cunkuan Yuan, Yukun Lai, Yebin Liu |
IEEE Trans. Image Process. | 4 |
| 2020 | 4D Association Graph for Realtime Multi-Person Motion Capture Using Multiple Video Camerasabstracthis paper contributes a novel realtime multi-person motion capture algorithm using multiview video inputs. Due to the heavy occlusions and closely interacting motions in each view, joint optimization on the multiview images and multiple temporal frames is indispensable, which brings up the essential challenge of realtime efficiency. To this end, for the first time, we unify per-view parsing, cross-view matching, and temporal tracking into a single optimization framework, i.e., a 4D association graph that each dimension (image space, viewpoint and time) can be treated equally and simultaneously. To solve the 4D association graph efficiently, we further contribute the idea of 4D limb bundle parsing based on heuristic searching, followed with limb bundle assembling by proposing a bundle Kruskal's algorithm. Our method enables a realtime motion capture system running at 30fps using 5 cameras on a 5-person scene. Benefiting from the unified parsing, matching and tracking constraints, our method is robust to noisy detection due to severe occlusions and close interacting motions, and achieves high-quality online pose reconstruction quality. The proposed method outperforms state-of-the-art methods quantitatively without using high-level appearance information. Yuxiang Zhang 0006, Liang An 0001, Tao Yu 0007, Xiu Li 0001, Kun Li 0001, Yebin Liu |
CVPR | 1 |