EDBT 2026 Demo / reviewers in the wild / expert
Wentao Zhu 0004
dblp:117/0354-4
· DBLP profile ↗
19ranked-venue papers
6as first author
18since 2021 · last 2026
0000-0002-5483-0259ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 6 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Action Counting with Dynamic Queries
Xiaoxuan Ma 0001, Zishi Li, Qiuyan Shang, Wentao Zhu 0004, Hai Ci, Yu Qiao 0001, Yizhou Wang 0001 |
Int. J. Comput. Vis. | 4 |
| 2025 | FreeCloth: Free-form Generation Enhances Challenging Clothed Human ModelingabstractAchieving realistic animated human avatars requires accurate modeling of pose-dependent clothing deformations. Existing learning-based methods heavily rely on the Linear Blend Skinning (LBS) of minimally-clothed human models like SMPL to model deformation. However, they struggle to handle loose clothing, such as long dresses, where the canonicalization process becomes ill-defined when the clothing is far from the body, leading to disjointed and fragmented results. To overcome this limitation, we propose FreeCloth, a novel hybrid framework to model challenging clothed humans. Our core idea is to use dedicated strategies to model different regions, depending on whether they are close to or distant from the body. Specifically, we segment the human body into three categories: unclothed, deformed, and generated. We simply replicate unclothed regions that require no deformation. For deformed regions close to the body, we leverage LBS to handle the deformation. As for the generated regions, which correspond to loose clothing areas, we introduce a novel free-form, part-aware generator to model them, as they are less affected by movements. This free-form generation paradigm brings enhanced flexibility and expressiveness to our hybrid framework, enabling it to capture the intricate geometric details of challenging loose clothing, such as skirts and dresses. Experimental results on the benchmark dataset featuring loose clothing demonstrate that FreeCloth achieves state-of-the-art performance with superior visual fidelity and realism, particularly in the most challenging cases. Hang Ye 0002, Xiaoxuan Ma 0001, Hai Ci, Wentao Zhu 0004, Yizhou Wang 0001 |
CVPR | 4 |
| 2025 | Embodied Representation Alignment with Mirror NeuronsabstractMirror neurons are a class of neurons that activate both when an individual observes an action and when they perform the same action. This mechanism reveals a fundamental interplay between action understanding and embodied execution, suggesting that these two abilities are inherently connected. Nonetheless, existing machine learning methods largely overlook this interplay, treating these abilities as separate tasks. In this study, we provide a unified perspective in modeling them through the lens of representation learning. We first observe that their intermediate representations spontaneously align. Inspired by mirror neurons, we further introduce an approach that explicitly aligns the representations of observed and executed actions. Specifically, we employ two linear layers to map the representations to a shared latent space, where contrastive learning enforces the alignment of corresponding representations, effectively maximizing their mutual information. Experiments demonstrate that this simple approach fosters mutual synergy between the two tasks, effectively improving representation quality and generalization. Wentao Zhu 0004, Zhining Zhang 0001, Yizhou Wang 0001 |
ICCV | 1 |
| 2025 | Aligning Human Motion Generation with Human PerceptionsabstractHuman motion generation is a critical task with a wide spectrum of applications. Achieving high realism in generated motions requires naturalness, smoothness, and plausibility. However, current evaluation metrics often rely on simple heuristics or distribution distances and do not align well with human perceptions. In this work, we propose a data-driven approach to bridge this gap by introducing a large-scale human perceptual evaluation dataset, MotionPercept, and a human motion critic model, MotionCritic, that capture human perceptual preferences. Our critic model offers a more accurate metric for assessing motion quality and could be readily integrated into the motion generation pipeline to enhance generation quality. Extensive experiments demonstrate the effectiveness of our approach in both evaluating and improving the quality of generated human motions by aligning with human perceptions. Code and data are publicly available at https://motioncritic.github.io/. Haoru Wang, Wentao Zhu 0004, Luyi Miao, Yishu Xu, Feng Gao 0014, Qi Tian 0001, Yizhou Wang 0001 |
ICLR | 2 |
| 2025 | PP-Motion: Physical-Perceptual Fidelity Evaluation for Human Motion Generation
Sihan Zhao, Zixuan Wang 0026, Tianyu Luan, Jia Jia 0001, Wentao Zhu 0004, Jiebo Luo 0001, Junsong Yuan 0001, Nan Xi |
ACM Multimedia | 5 |
| 2025 | VMarker-Pro: Probabilistic 3D Human Mesh Estimation From Virtual MarkersabstractMonocular 3D human mesh estimation faces challenges due to depth ambiguity and the complexity of mapping images to complex parameter spaces. Recent methods propose to use 3D poses as a proxy representation, which often lose crucial body shape information, leading to mediocre performance. Conversely, advanced motion capture systems, though accurate, are impractical for markerless wild images. Addressing these limitations, we introduce an innovative intermediate representation as virtual markers, which are learned from large-scale mocap data, mimicking the effects of physical markers. Building upon virtual markers, we propose VMarker, which detects virtual markers from wild images, and the intact mesh with realistic shapes can be obtained by simply interpolation from these markers. To address occlusions that obscure 3D virtual marker estimation, we further enhance our method with VMarker-Pro, a probabilistic framework that models the distribution of 3D virtual marker positions using diffusion models, enabling the generation of multiple plausible meshes aligned with images for robust 3D mesh estimation. Our approaches surpass existing methods on three benchmark datasets, particularly demonstrating significant improvements on the SURREAL dataset, which features diverse body shapes. Additionally, VMarker-Pro excels in accurately modeling data distributions, significantly enhancing performance in occluded scenarios. Xiaoxuan Ma 0001, Jiajun Su, Yuan Xu 0022, Wentao Zhu 0004, Chunyu Wang 0001, Yizhou Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | ScoreHypo: Probabilistic Human Mesh Estimation with Hypothesis ScoringabstractMonocular 3D human mesh estimation is an ill-posed problem, characterized by inherent ambiguity and occlusion. While recent probabilistic methods propose generating multiple solutions, little attention is paid to obtaining high-quality estimates from them. To address this limitation, we introduce ScoreHypo, a versatile framework by first leveraging our novel HypoNet to generate multiple hy-potheses, followed by employing a meticulously designed scorer, ScoreNet, to evaluate and select high-quality esti-mates. ScoreHypo formulates the estimation process as a re-verse denoising process, where HypoNet produces a diverse set of plausible estimates that effectively align with the im-age cues. Subsequently, ScoreNet is employed to rigorously evaluate and rank these estimates based on their quality and finally identify superior ones. Experimental results demon-strate that HypoNet outperforms existing state-of-the-art probabilistic methods as a multi-hypothesis mesh estimator. Moreover, the estimates selected by ScoreNet significantly outperform random generation or simple averaging. Notably, the trained ScoreNet exhibits generalizability, as it can effectively score existing methods and significantly reduce their errors by more than 15%. Code and models are available at ht tps: / /xy02- 05. gi thub. io/ScoreHypo. Yuan Xu 0022, Xiaoxuan Ma 0001, Jiajun Su, Wentao Zhu 0004, Yu Qiao 0001, Yizhou Wang 0001 |
CVPR | 4 |
| 2024 | Real-Time Holistic Robot Pose Estimation with Unknown States
Shikun Ban, Juling Fan, Xiaoxuan Ma 0001, Wentao Zhu 0004, Yu Qiao 0003, Yizhou Wang 0001 |
ECCV (49) | 4 |
| 2024 | Language Models Represent Beliefs of Self and OthersabstractUnderstanding and attributing mental states, known as Theory of Mind (ToM), emerges as a fundamental capability for human social reasoning. While Large Language Models (LLMs) appear to possess certain ToM abilities, the mechanisms underlying these capabilities remain elusive. In this study, we discover that it is possible to linearly decode the belief status from the perspectives of various agents through neural activations of language models, indicating the existence of internal representations of self and others’ beliefs. By manipulating these representations, we observe dramatic changes in the models’ ToM performance, underscoring their pivotal role in the social reasoning process. Additionally, our findings extend to diverse social reasoning tasks that involve different causal inference patterns, suggesting the potential generalizability of these representations. Wentao Zhu 0004, Zhining Zhang 0001, Yizhou Wang 0001 |
ICML | 1 |
| 2024 | Human Motion Generation: A SurveyabstractHuman motion generation aims to generate natural human pose sequences and shows immense potential for real-world applications. Substantial progress has been made recently in motion data collection technologies and generation methods, laying the foundation for increasing interest in human motion generation. Most research within this field focuses on generating human motions based on conditional signals, such as text, audio, and scene contexts. While significant advancements have been made in recent years, the task continues to pose challenges due to the intricate nature of human motion and its implicit relationship with conditional signals. In this survey, we present a comprehensive literature review of human motion generation, which, to the best of our knowledge, is the first of its kind in this field. We begin by introducing the background of human motion and generative models, followed by an examination of representative methods for three mainstream sub-tasks: text-conditioned, audio-conditioned, and scene-conditioned human motion generation. Additionally, we provide an overview of common datasets and evaluation metrics. Lastly, we discuss open problems and outline potential future research directions. We hope that this survey could provide the community with a comprehensive glimpse of this rapidly evolving field and inspire novel ideas that address the outstanding challenges. Wentao Zhu 0004, Xiaoxuan Ma 0001, Dongwoo Ro, Hai Ci, Jinlu Zhang 0001, Jiaxin Shi, Feng Gao 0014, Qi Tian 0001, Yizhou Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | GFPose: Learning 3D Human Pose Prior with Gradient FieldsabstractLearning 3D human pose prior is essential to human-centered AI. Here, we present GFPose, a versatile framework to model plausible 3D human poses for various applications. At the core of GFPose is a time-dependent score network, which estimates the gradient on each body joint and progressively denoises the perturbed 3D human pose to match a given task specification. During the denoising process, GFPose implicitly incorporates pose priors in gradients and unifies various discriminative and generative tasks in an elegant framework. Despite the simplicity, GFPose demonstrates great potential in several downstream tasks. Our experiments empirically show that 1) as a multi-hypothesis pose estimator, GFPose outperforms existing SOTAs by 20% on Human3.6M dataset. 2) as a single-hypothesis pose estimator, GFPose achieves comparable results to deterministic SOTAs, even with a vanilla backbone. 3) GFPose is able to produce diverse and realistic samples in pose denoising, completion and generation tasks.11Project page https://sites.google.com/view/gfpose/ Hai Ci, Mingdong Wu, Wentao Zhu 0004, Xiaoxuan Ma 0001, Hao Dong 0003, Fangwei Zhong, Yizhou Wang 0001 |
CVPR | 3 |
| 2023 | 3D Human Mesh Estimation from Virtual MarkersabstractInspired by the success of volumetric 3D pose estimation, some recent human mesh estimators propose to estimate 3D skeletons as intermediate representations, from which, the dense 3D meshes are regressed by exploiting the mesh topology. However, body shape information is lost in extracting skeletons, leading to mediocre performance. The advanced motion capture systems solve the problem by placing dense physical markers on the body surface, which allows to extract realistic meshes from their non-rigid motions. However, they cannot be applied to wild images without markers. In this work, we present an intermediate representation, named virtual markers, which learns 64 landmark keypoints on the body surface based on the large-scale mocap data in a generative style, mimicking the effects of physical markers. The virtual markers can be accurately detected from wild images and can reconstruct the intact meshes with realistic shapes by simple interpolation. Our approach outperforms the state-of-the-art methods on three datasets. In particular, it surpasses the existing methods by a notable margin on the SURREAL dataset, which has diverse body shapes. Code is available at https://github.com/ShirleyMaxx/VirtualMarker Xiaoxuan Ma 0001, Jiajun Su, Chunyu Wang 0001, Wentao Zhu 0004, Yizhou Wang 0001 |
CVPR | 4 |
| 2023 | MotionBERT: A Unified Perspective on Learning Human Motion RepresentationsabstractWe present a unified perspective on tackling various human-centric video tasks by learning human motion representations from large-scale and heterogeneous data resources. Specifically, we propose a pretraining stage in which a motion encoder is trained to recover the underlying 3D motion from noisy partial 2D observations. The motion representations acquired in this way incorporate geometric, kinematic, and physical knowledge about human motion, which can be easily transferred to multiple downstream tasks. We implement the motion encoder with a Dual-stream Spatio-temporal Transformer (DSTformer) neural network. It could capture long-range spatio-temporal relationships among the skeletal joints comprehensively and adaptively, exemplified by the lowest 3D pose estimation error so far when trained from scratch. Furthermore, our proposed framework achieves state-of-the-art performance on all three downstream tasks by simply finetuning the pretrained motion encoder with a simple regression head (1-2 layers), which demonstrates the versatility of the learned motion representations. Code and models are available at https://motionbert.github.io/ Wentao Zhu 0004, Xiaoxuan Ma 0001, Zhaoyang Liu 0001, Libin Liu 0002, Wayne Wu, Yizhou Wang 0001 |
ICCV | 1 |
| 2023 | ChimpACT: A Longitudinal Dataset for Understanding Chimpanzee BehaviorsabstractUnderstanding the behavior of non-human primates is crucial for improving animal welfare, modeling social behavior, and gaining insights into distinctively human and phylogenetically shared behaviors. However, the lack of datasets on non-human primate behavior hinders in-depth exploration of primate social interactions, posing challenges to research on our closest living relatives. To address these limitations, we present ChimpACT, a comprehensive dataset for quantifying the longitudinal behavior and social relations of chimpanzees within a social group. Spanning from 2015 to 2018, ChimpACT features videos of a group of over 20 chimpanzees residing at the Leipzig Zoo, Germany, with a particular focus on documenting the developmental trajectory of one young male, Azibo. ChimpACT is both comprehensive and challenging, consisting of 163 videos with a cumulative 160,500 frames, each richly annotated with detection, identification, pose estimation, and fine-grained spatiotemporal behavior labels. We benchmark representative methods of three tracks on ChimpACT: (i) tracking and identification, (ii) pose estimation, and (iii) spatiotemporal action detection of the chimpanzees. Our experiments reveal that ChimpACT offers ample opportunities for both devising new methods and adapting existing ones to solve fundamental computer vision tasks applied to chimpanzee groups, such as detection, pose estimation, and behavior analysis, ultimately deepening our comprehension of communication and sociality in non-human primates. Xiaoxuan Ma 0001, Stephan P. Kaufhold, Jiajun Su, Wentao Zhu 0004, Jack Terwilliger, Andres Meza 0001, Yixin Zhu 0001, Federico Rossano, Yizhou Wang 0001 |
NeurIPS | 4 |
| 2023 | Social Motion Prediction with Cognitive HierarchiesabstractHumans exhibit a remarkable capacity for anticipating the actions of others and planning their own actions accordingly. In this study, we strive to replicate this ability by addressing the social motion prediction problem. We introduce a new benchmark, a novel formulation, and a cognition-inspired framework. We present Wusi, a 3D multi-person motion dataset under the context of team sports, which features intense and strategic human interactions and diverse pose distributions. By reformulating the problem from a multi-agent reinforcement learning perspective, we incorporate behavioral cloning and generative adversarial imitation learning to boost learning efficiency and generalization. Furthermore, we take into account the cognitive aspects of the human social action planning process and develop a cognitive hierarchy framework to predict strategic human social interactions. We conduct comprehensive experiments to validate the effectiveness of our proposed dataset and approach. Wentao Zhu 0004, Jason Qin, Yuke Lou, Hang Ye 0002, Xiaoxuan Ma 0001, Hai Ci, Yizhou Wang 0001 |
NeurIPS | 1 |
| 2022 | MoCaNet: Motion Retargeting In-the-Wild via Canonicalization NetworksabstractWe present a novel framework that brings the 3D motion retargeting task from controlled environments to in-the-wild scenarios. In particular, our method is capable of retargeting body motion from a character in a 2D monocular video to a 3D character without using any motion capture system or 3D reconstruction procedure. It is designed to leverage massive online videos for unsupervised training, needless of 3D annotations or motion-body pairing information. The proposed method is built upon two novel canonicalization operations, structure canonicalization and view canonicalization. Trained with the canonicalization operations and the derived regularizations, our method learns to factorize a skeleton sequence into three independent semantic subspaces, i.e., motion, structure, and view angle. The disentangled representation enables motion retargeting from 2D to 3D with high precision. Our method achieves superior performance on motion transfer benchmarks with large body variations and challenging actions. Notably, the canonicalized skeleton sequence could serve as a disentangled and interpretable representation of human motion that benefits action analysis and motion retrieval. Wentao Zhu 0004, Zhuoqian Yang, Ziang Di, Wayne Wu, Yizhou Wang 0001, Chen Change Loy |
AAAI | 1 |
| 2022 | Faster VoxelPose: Real-time 3D Human Pose Estimation by Orthographic Projection
Hang Ye 0002, Wentao Zhu 0004, Chunyu Wang 0001, Rujie Wu, Yizhou Wang 0001 |
ECCV (6) | 2 |
| 2022 | CelebV-HQ: A Large-Scale Video Facial Attributes Dataset
Wayne Wu, Wentao Zhu 0004, Liming Jiang 0001, Siwei Tang, Li Zhang 0105, Ziwei Liu 0002, Chen Change Loy |
ECCV (7) | 3 |
| 2020 | TransMoMo: Invariance-Driven Unsupervised Video Motion RetargetingabstractWe present a lightweight video motion retargeting approach TransMoMo that is capable of transferring motion of a person in a source video realistically to another video of a target person. Without using any paired data for supervision, the proposed method can be trained in an unsupervised manner by exploiting invariance properties of three orthogonal factors of variation including motion, structure, and view-angle. Specifically, with loss functions carefully derived based on invariance, we train an auto-encoder to disentangle the latent representations of such factors given the source and target video clips. This allows us to selectively transfer motion extracted from the source video seamlessly to the target video in spite of structural and view-angle disparities between the source and the target. The relaxed assumption of paired data allows our method to be trained on a vast amount of videos needless of manual annotation of source-target pairing, leading to improved robustness against large structural variations and extreme motion in videos. We demonstrate the effectiveness of our method over the state-of-the-art methods. Code, model and data are publicly available on our project page (https://yzhq97.github.io/transmomo). Zhuoqian Yang, Wentao Zhu 0004, Wayne Wu, Chen Qian 0006, Qiang Zhou 0001, Bolei Zhou, Chen Change Loy |
CVPR | 2 |