Zhongcong Xu

dblp:232/3210 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0003-3511-8466ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
6 papers
Visual content generation and editing · 33% Rendering · 25% Computer animation and physical simulation · 25%
Artificial intelligence
3 papers
Generative modeling · 84% Video understanding and tracking · 8% 3D vision · 7%

Topics — the 16 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.522024
Exocentric-to-Egocentric Video Generation · NeurIPS 2024
MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model · CVPR 2024
Geometric modeling and processing › mesh processing
3d model processing
0.912025
MagicArticulate: Make Your 3D Models Articulation-Ready · CVPR 2025
Computer animation and physical simulation › character rigging
automatic rigging
0.912025
Puppeteer: Rig and Animate Your 3D Models · NeurIPS 2025
Computer animation and physical simulation › skinning
skinning weight prediction
0.912025
Puppeteer: Rig and Animate Your 3D Models · NeurIPS 2025
Machine learning › Generative modeling › diffusion model
video diffusion model
0.812024
Exocentric-to-Egocentric Video Generation · NeurIPS 2024
Visual content generation and editing › image animation
human image animation
0.812024
MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model · CVPR 2024
Visual content generation and editing
video generation
0.812024
Exocentric-to-Egocentric Video Generation · NeurIPS 2024
Visual content generation and editing
3d-aware generative model
0.712023
XAGen: 3D Expressive Human Avatars Generation · NeurIPS 2023
Visual content generation and editing
3d content generation
0.712023
XAGen: 3D Expressive Human Avatars Generation · NeurIPS 2023
Rendering › novel view synthesis
free-viewpoint rendering
0.712023
HOSNeRF: Dynamic Human-Object-Scene Neural Radiance Fields from a Single Video · ICCV 2023
Visual content generation and editing › avatar generation
human avatar synthesis
0.712023
XAGen: 3D Expressive Human Avatars Generation · NeurIPS 2023
Rendering
neural radiance fields
0.712023
HOSNeRF: Dynamic Human-Object-Scene Neural Radiance Fields from a Single Video · ICCV 2023
Rendering
neural rendering
0.712023
XAGen: 3D Expressive Human Avatars Generation · NeurIPS 2023
Rendering
novel view synthesis
0.712023
HOSNeRF: Dynamic Human-Object-Scene Neural Radiance Fields from a Single Video · ICCV 2023
Computer vision › Video understanding and tracking › multi-camera video analysis
multiview video understanding
0.212024
Exocentric-to-Egocentric Video Generation · NeurIPS 2024
Computer vision › 3D vision › 3d scene understanding › object relation reasoning
human-object interaction
0.212023
HOSNeRF: Dynamic Human-Object-Scene Neural Radiance Fields from a Single Video · ICCV 2023

Methods — techniques the papers use, named apart from their topics

autoregressive transformer · 1.7video fusion · 1.5video diffusion model · 1.5temporal attention · 1.5multi-view encoder · 1.5appearance encoder · 1.5volumetric geodesic distance · 0.9functional diffusion process · 0.9differentiable optimization · 0.9attention · 0.9view translation prior · 0.8object state embedding · 0.7object bones · 0.7
YearPublicationVenuePosition
2026 Unlocking the Video Prior for High-Fidelity Sparse Multi-View Image Synthesis
abstract
The development of multi-view image synthesis is constrained by the scarcity of training data. One promising solution is to finetune well-trained video generative models to synthesize 360-degree videos of objects. While these methods benefit from the strong generative priors inherited from the pretrained knowledge, they are limited by the high computational costs incurred by the large number of viewpoints. Existing methods commonly adopt temporal attention mechanism to address this. However, these methods suffer from undesirable artifacts such as 3D inconsistency and over-smoothing in the generated results. In this paper, we introduce a novel approach to unlock the video priors for multi-view synthesis by reducing generation into a sparser yet more precise process. Specifically, we introduce two strategies to achieve this: i) Condensing the video diffusion model to synthesize highly consistent sparse multiview images. ii) Extracting dense geometrical priors from the pretrained video diffusion models to enhance the generation stability. The combination of these two strategies formulates a novel framework for multi-view synthesis, which is capable of synthesizing highly consistent sparse multiview images with strong generalization ability. Extensive experiments demonstrate that our approach achieves superior efficiency, generalization, and consistency, outperforming state-of-the-art multi-view synthesis methods.
Fan Yang 0103, Jun Hao Liew, Chaoyue Song, Zhongcong Xu, Jiashi Feng, Guosheng Lin
3DV5
2025 MagicArticulate: Make Your 3D Models Articulation-Ready
abstract
With the explosive growth of 3D content creation, there is an increasing demand for automatically converting static 3D models into articulation-ready versions that support realistic animation. Traditional approaches rely heavily on manual annotation, which is both time-consuming and labor-intensive. Moreover, the lack of large-scale benchmarks has hindered the development of learning-based solutions. In this work, we present MagicArticulate, an effective framework that automatically transforms static 3D models into articulation-ready assets. Our key contributions are threefold. First, we introduce Articulation-XL, a large-scale benchmark containing over 33k 3D models with high-quality articulation annotations, carefully curated from Objaverse-XL. Second, we propose a novel skeleton generation method that formulates the task as a sequence modeling problem, leveraging an autoregressive transformer to naturally handle varying numbers of bones or joints within skeletons and their inherent dependencies across different 3D models. Third, we predict skinning weights using a functional diffusion process that incorporates volumetric geodesic distance priors between vertices and joints. Extensive experiments demonstrate that MagicArticulate significantly outperforms existing methods across diverse object categories, achieving high-quality articulation that enables realistic animation. Project page: https://chaoyuesong.github.io/MagicArticulate.
Chaoyue Song, Xiu Li 0001, Fan Yang 0103, Zhongcong Xu, Jun Hao Liew, Fayao Liu, Jiashi Feng, Guosheng Lin
CVPR6
2025 Puppeteer: Rig and Animate Your 3D Models
abstract
Modern interactive applications increasingly demand dynamic 3D content, yet the transformation of static 3D models into animated assets constitutes a significant bottleneck in content creation pipelines. While recent advances in generative AI have revolutionized static 3D model creation, rigging and animation continue to depend heavily on expert intervention. We present \textbf{Puppeteer}, a comprehensive framework that addresses both automatic rigging and animation for diverse 3D objects. Our system first predicts plausible skeletal structures via an auto-regressive transformer that introduces a joint-based tokenization strategy for compact representation and a hierarchical ordering methodology with stochastic perturbation that enhances bidirectional learning capabilities. It then infers skinning weights via an attention-based architecture incorporating topology-aware joint attention that explicitly encodes inter-joint relationships based on skeletal graph distances. Finally, we complement these rigging advances with a differentiable optimization-based animation pipeline that generates stable, high-fidelity animations while being computationally more efficient than existing approaches. Extensive evaluations across multiple benchmarks demonstrate that our method significantly outperforms state-of-the-art techniques in both skeletal prediction accuracy and skinning quality. The system robustly processes diverse 3D content, ranging from professionally designed game assets to AI-generated shapes, producing temporally coherent animations that eliminate the jittering issues common in existing methods.
Chaoyue Song, Xiu Li 0001, Fan Yang 0103, Zhongcong Xu, Jiacheng Wei, Fayao Liu, Jiashi Feng, Guosheng Lin
NeurIPS4
2024 MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model
abstract
This paper studies the human image animation task, which aims to generate a video of a certain reference iden-tity following a particular motion sequence. Existing an-imation works typically employ the frame-warping technique to animate the reference image towards the target motion. Despite achieving reasonable results, these approaches face challenges in maintaining temporal consistency throughout the animation due to the lack of temporal modeling and poor preservation of reference identity. In this work, we introduce Magic/snimate, a diffusion-based framework that aims at enhancing temporal consistency, preserving reference image faithfully, and improving animation fidelity. To achieve this, we first develop a video diffusion model to encode temporal information. Second, to maintain the appearance coherence across frames, we introduce a novel appearance encoder to retain the intricate details of the reference image. Leveraging these two inno-vations, we further employ a simple video fusion technique to encourage smooth transitions for long video animation. Empirical results demonstrate the superiority of our method over baseline approaches on two benchmarks. Notably, our approach outperforms the strongest baseline by over 38% in terms of video fidelity on the challenging TikTok dancing dataset. Code and model will be made available at https://showlab.github.io/magicanimate.
Zhongcong Xu, Jun Hao Liew, Hanshu Yan, Jia-Wei Liu, Jiashi Feng, Zheng Shou 0001
CVPR1
2024 Exocentric-to-Egocentric Video Generation
abstract
We introduce Exo2Ego-V, a novel exocentric-to-egocentric diffusion-based video generation method for daily-life skilled human activities where sparse 4-view exocentric viewpoints are configured 360° around the scene. This task is particularly challenging due to the significant variations between exocentric and egocentric viewpoints and high complexity of dynamic motions and real-world daily-life environments. To address these challenges, we first propose a new diffusion-based multi-view exocentric encoder to extract the dense multi-scale features from multi-view exocentric videos as the appearance conditions for egocentric video generation. Then, we design an exocentric-to-egocentric view translation prior to provide spatially aligned egocentric features as a concatenation guidance for the input of egocentric video diffusion model. Finally, we introduce the temporal attention layers into our egocentric video diffusion pipeline to improve the temporal consistency cross egocentric frames. Extensive experiments demonstrate that Exo2Ego-V significantly outperforms SOTA approaches on 5 categories from the Ego-Exo4D dataset with an average of 35% in terms of LPIPS. Our code and model will be made available on https://github.com/showlab/Exo2Ego-V.
Jia-Wei Liu, Weijia Mao, Zhongcong Xu, Jussi Keppo, Zheng Shou 0001
NeurIPS3
2023 HOSNeRF: Dynamic Human-Object-Scene Neural Radiance Fields from a Single Video
abstract
We introduce HOSNeRF, a novel 360° free-viewpoint rendering method that reconstructs neural radiance fields for dynamic human-object-scene from a single monocular in-the-wild video. Our method enables pausing the video at any frame and rendering all scene details (dynamic humans, objects, and backgrounds) from arbitrary viewpoints. The first challenge in this task is the complex object motions in human-object interactions, which we tackle by introducing the new object bones into the conventional human skeleton hierarchy to effectively estimate large object deformations in our dynamic human-object model. The second challenge is that humans interact with different objects at different times, for which we introduce two new learnable object state embeddings that can be used as conditions for learning our human-object representation and scene representation, respectively. Extensive experiments show that HOSNeRF significantly outperforms SOTA approaches on two challenging datasets by a large margin of 40%~50% in terms of LPIPS. The code, data, and compelling examples of 360° free-viewpoint renderings from single videos: https://showlab.github.io/HOSNeRF.
Jia-Wei Liu, Yan-Pei Cao 0001, Tianyuan Yang, Zhongcong Xu, Jussi Keppo, Ying Shan, Xiaohu Qie, Zheng Shou 0001
ICCV4
2023 XAGen: 3D Expressive Human Avatars Generation
abstract
Recent advances in 3D-aware GAN models have enabled the generation of realistic and controllable human body images. However, existing methods focus on the control of major body joints, neglecting the manipulation of expressive attributes, such as facial expressions, jaw poses, hand poses, and so on. In this work, we present XAGen, the first 3D generative model for human avatars capable of expressive control over body, face, and hands. To enhance the fidelity of small-scale regions like face and hands, we devise a multi-scale and multi-part 3D representation that models fine details. Based on this representation, we propose a multi-part rendering technique that disentangles the synthesis of body, face, and hands to ease model training and enhance geometric quality. Furthermore, we design multi-part discriminators that evaluate the quality of the generated avatars with respect to their appearance and fine-grained control capabilities. Experiments show that XAGen surpasses state-of-the-art methods in terms of realism, diversity, and expressive control abilities. Code and data will be made available at https://showlab.github.io/xagen.
Zhongcong Xu, Jun Hao Liew, Jiashi Feng, Zheng Shou 0001
NeurIPS1
2021 A Novel Illumination-Robust Hand Gesture Recognition System With Event-Based Neuromorphic Vision Sensor
abstract
The hand gesture recognition system is a noncontact and intuitive communication approach, which, in turn, allows for natural and efficient interaction. This work focuses on developing a novel and robust gesture recognition system, which is insensitive to environmental illumination and background variation. In the field of gesture recognition, standard vision sensors, such as CMOS cameras, are widely used as the sensing devices in state-of-the-art hand gesture recognition systems. However, such cameras depend on environmental constraints, such as lighting variability and the cluttered background, which significantly deteriorates their performances. In this work, we propose an event-based gesture recognition system to overcome the detriment constraints and enhance the robustness of the recognition performance. Our system relies on a biologically inspired neuromorphic vision sensor that has microsecond temporal resolution, high dynamic range, and low latency. The sensor output is a sequence of asynchronous events instead of discrete frames. To interpret the visual data, we utilize a wearable glove as an interaction device with five high-frequency (>100 Hz) active LED markers (ALMs), representing fingers and palm, which are tracked precisely in the temporal domain using a restricted spatiotemporal particle filter algorithm. The latency of the sensing pipeline is negligible compared with the dynamics of the environment as the sensor's temporal resolution allows us to distinguish high frequencies precisely. We design an encoding process to extract features and adopt a lightweight network to classify the hand gestures. The recognition accuracy of our system is comparable to the state-of-the-art methods. To study the robustness of the system, experiments considering illumination and background variations are performed, and the results show that our system is more robust than the state-of-the-art deep learning-based gesture recognition systems. Note to Practitioners-This article addresses the robustness of the hand gesture recognition system that is important for gesture recognition-based applications. Existing methods rely on either the large-volume data to train a deep learning model or to restrict the applied environments (e.g., an ideal environment without dynamic background). However, a vision-based deep learning model requires large computational resources, while the ideal environment limits the practicality of the system. In this work, we introduce a biologically inspired neuromorphic vision sensor and an ALM glove and build a novel gesture recognition system to tackle the above issue. The neuromorphic vision sensor has a microsecond temporal resolution and a high dynamic range. With these properties, the sensing system of our prototype operates in a very low-latency space, which, in turn, ensures that our gesture recognition system is robust to illumination variance and dynamic background. Thus, this work is valuable to the research of illumination-robust gesture recognition systems. Preliminary experiments suggest that our system prototype is feasible, but it has not yet been incorporated into an online gesture recognition system nor tested with complex gestures. In future work, we will concentrate on the improvement of the signal processing methods that advance the current system to complex and practical applications.
Guang Chen 0001, Zhongcong Xu, Zhijun Li 0001, Huajin Tang, Sanqing Qu, Kejia Ren, Alois C. Knoll
IEEE Trans Autom. Sci. Eng.2