EDBT 2026 Demo / reviewers in the wild / expert
Zhongcong Xu
dblp:232/3210
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0003-3511-8466ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
6 papers |
Visual content generation and editing · 33% Rendering · 25% Computer animation and physical simulation · 25% | |
| Artificial intelligence
3 papers |
Generative modeling · 84% Video understanding and tracking · 8% 3D vision · 7% |
Topics — the 16 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
1.5 | 2 | 2024 | Exocentric-to-Egocentric Video Generation · NeurIPS 2024 MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model · CVPR 2024 |
Geometric modeling and processing › mesh processing
3d model processing |
0.9 | 1 | 2025 | MagicArticulate: Make Your 3D Models Articulation-Ready · CVPR 2025 |
Computer animation and physical simulation › character rigging
automatic rigging |
0.9 | 1 | 2025 | Puppeteer: Rig and Animate Your 3D Models · NeurIPS 2025 |
Computer animation and physical simulation › skinning
skinning weight prediction |
0.9 | 1 | 2025 | Puppeteer: Rig and Animate Your 3D Models · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model
video diffusion model |
0.8 | 1 | 2024 | Exocentric-to-Egocentric Video Generation · NeurIPS 2024 |
Visual content generation and editing › image animation
human image animation |
0.8 | 1 | 2024 | MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model · CVPR 2024 |
Visual content generation and editing
video generation |
0.8 | 1 | 2024 | Exocentric-to-Egocentric Video Generation · NeurIPS 2024 |
Visual content generation and editing
3d-aware generative model |
0.7 | 1 | 2023 | XAGen: 3D Expressive Human Avatars Generation · NeurIPS 2023 |
Visual content generation and editing
3d content generation |
0.7 | 1 | 2023 | XAGen: 3D Expressive Human Avatars Generation · NeurIPS 2023 |
Rendering › novel view synthesis
free-viewpoint rendering |
0.7 | 1 | 2023 | HOSNeRF: Dynamic Human-Object-Scene Neural Radiance Fields from a Single Video · ICCV 2023 |
Visual content generation and editing › avatar generation
human avatar synthesis |
0.7 | 1 | 2023 | XAGen: 3D Expressive Human Avatars Generation · NeurIPS 2023 |
Rendering
neural radiance fields |
0.7 | 1 | 2023 | HOSNeRF: Dynamic Human-Object-Scene Neural Radiance Fields from a Single Video · ICCV 2023 |
Rendering
neural rendering |
0.7 | 1 | 2023 | XAGen: 3D Expressive Human Avatars Generation · NeurIPS 2023 |
Rendering
novel view synthesis |
0.7 | 1 | 2023 | HOSNeRF: Dynamic Human-Object-Scene Neural Radiance Fields from a Single Video · ICCV 2023 |
Computer vision › Video understanding and tracking › multi-camera video analysis
multiview video understanding |
0.2 | 1 | 2024 | Exocentric-to-Egocentric Video Generation · NeurIPS 2024 |
Computer vision › 3D vision › 3d scene understanding › object relation reasoning
human-object interaction |
0.2 | 1 | 2023 | HOSNeRF: Dynamic Human-Object-Scene Neural Radiance Fields from a Single Video · ICCV 2023 |
Methods — techniques the papers use, named apart from their topics
autoregressive transformer · 1.7video fusion · 1.5video diffusion model · 1.5temporal attention · 1.5multi-view encoder · 1.5appearance encoder · 1.5volumetric geodesic distance · 0.9functional diffusion process · 0.9differentiable optimization · 0.9attention · 0.9view translation prior · 0.8object state embedding · 0.7object bones · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unlocking the Video Prior for High-Fidelity Sparse Multi-View Image SynthesisabstractThe development of multi-view image synthesis is constrained by the scarcity of training data. One promising solution is to finetune well-trained video generative models to synthesize 360-degree videos of objects. While these methods benefit from the strong generative priors inherited from the pretrained knowledge, they are limited by the high computational costs incurred by the large number of viewpoints. Existing methods commonly adopt temporal attention mechanism to address this. However, these methods suffer from undesirable artifacts such as 3D inconsistency and over-smoothing in the generated results. In this paper, we introduce a novel approach to unlock the video priors for multi-view synthesis by reducing generation into a sparser yet more precise process. Specifically, we introduce two strategies to achieve this: i) Condensing the video diffusion model to synthesize highly consistent sparse multiview images. ii) Extracting dense geometrical priors from the pretrained video diffusion models to enhance the generation stability. The combination of these two strategies formulates a novel framework for multi-view synthesis, which is capable of synthesizing highly consistent sparse multiview images with strong generalization ability. Extensive experiments demonstrate that our approach achieves superior efficiency, generalization, and consistency, outperforming state-of-the-art multi-view synthesis methods. Fan Yang 0103, Jun Hao Liew, Chaoyue Song, Zhongcong Xu, Jiashi Feng, Guosheng Lin |
3DV | 5 |
| 2025 | MagicArticulate: Make Your 3D Models Articulation-ReadyabstractWith the explosive growth of 3D content creation, there is an increasing demand for automatically converting static 3D models into articulation-ready versions that support realistic animation. Traditional approaches rely heavily on manual annotation, which is both time-consuming and labor-intensive. Moreover, the lack of large-scale benchmarks has hindered the development of learning-based solutions. In this work, we present MagicArticulate, an effective framework that automatically transforms static 3D models into articulation-ready assets. Our key contributions are threefold. First, we introduce Articulation-XL, a large-scale benchmark containing over 33k 3D models with high-quality articulation annotations, carefully curated from Objaverse-XL. Second, we propose a novel skeleton generation method that formulates the task as a sequence modeling problem, leveraging an autoregressive transformer to naturally handle varying numbers of bones or joints within skeletons and their inherent dependencies across different 3D models. Third, we predict skinning weights using a functional diffusion process that incorporates volumetric geodesic distance priors between vertices and joints. Extensive experiments demonstrate that MagicArticulate significantly outperforms existing methods across diverse object categories, achieving high-quality articulation that enables realistic animation. Project page: https://chaoyuesong.github.io/MagicArticulate. Chaoyue Song, Xiu Li 0001, Fan Yang 0103, Zhongcong Xu, Jun Hao Liew, Fayao Liu, Jiashi Feng, Guosheng Lin |
CVPR | 6 |
| 2025 | Puppeteer: Rig and Animate Your 3D ModelsabstractModern interactive applications increasingly demand dynamic 3D content, yet the transformation of static 3D models into animated assets constitutes a significant bottleneck in content creation pipelines. While recent advances in generative AI have revolutionized static 3D model creation, rigging and animation continue to depend heavily on expert intervention. We present \textbf{Puppeteer}, a comprehensive framework that addresses both automatic rigging and animation for diverse 3D objects.
Our system first predicts plausible skeletal structures via an auto-regressive transformer that introduces a joint-based tokenization strategy for compact representation and a hierarchical ordering methodology with stochastic perturbation that enhances bidirectional learning capabilities. It then infers skinning weights via an attention-based architecture incorporating topology-aware joint attention that explicitly encodes inter-joint relationships based on skeletal graph distances.
Finally, we complement these rigging advances with a differentiable optimization-based animation pipeline that generates stable, high-fidelity animations while being computationally more efficient than existing approaches.
Extensive evaluations across multiple benchmarks demonstrate that our method significantly outperforms state-of-the-art techniques in both skeletal prediction accuracy and skinning quality. The system robustly processes diverse 3D content, ranging from professionally designed game assets to AI-generated shapes, producing temporally coherent animations that eliminate the jittering issues common in existing methods. Chaoyue Song, Xiu Li 0001, Fan Yang 0103, Zhongcong Xu, Jiacheng Wei, Fayao Liu, Jiashi Feng, Guosheng Lin |
NeurIPS | 4 |
| 2024 | MagicAnimate: Temporally Consistent Human Image Animation using Diffusion ModelabstractThis paper studies the human image animation task, which aims to generate a video of a certain reference iden-tity following a particular motion sequence. Existing an-imation works typically employ the frame-warping technique to animate the reference image towards the target motion. Despite achieving reasonable results, these approaches face challenges in maintaining temporal consistency throughout the animation due to the lack of temporal modeling and poor preservation of reference identity. In this work, we introduce Magic/snimate, a diffusion-based framework that aims at enhancing temporal consistency, preserving reference image faithfully, and improving animation fidelity. To achieve this, we first develop a video diffusion model to encode temporal information. Second, to maintain the appearance coherence across frames, we introduce a novel appearance encoder to retain the intricate details of the reference image. Leveraging these two inno-vations, we further employ a simple video fusion technique to encourage smooth transitions for long video animation. Empirical results demonstrate the superiority of our method over baseline approaches on two benchmarks. Notably, our approach outperforms the strongest baseline by over 38% in terms of video fidelity on the challenging TikTok dancing dataset. Code and model will be made available at https://showlab.github.io/magicanimate. Zhongcong Xu, Jun Hao Liew, Hanshu Yan, Jia-Wei Liu, Jiashi Feng, Zheng Shou 0001 |
CVPR | 1 |
| 2024 | Exocentric-to-Egocentric Video GenerationabstractWe introduce Exo2Ego-V, a novel exocentric-to-egocentric diffusion-based video generation method for daily-life skilled human activities where sparse 4-view exocentric viewpoints are configured 360° around the scene. This task is particularly challenging due to the significant variations between exocentric and egocentric viewpoints and high complexity of dynamic motions and real-world daily-life environments. To address these challenges, we first propose a new diffusion-based multi-view exocentric encoder to extract the dense multi-scale features from multi-view exocentric videos as the appearance conditions for egocentric video generation. Then, we design an exocentric-to-egocentric view translation prior to provide spatially aligned egocentric features as a concatenation guidance for the input of egocentric video diffusion model. Finally, we introduce the temporal attention layers into our egocentric video diffusion pipeline to improve the temporal consistency cross egocentric frames. Extensive experiments demonstrate that Exo2Ego-V significantly outperforms SOTA approaches on 5 categories from the Ego-Exo4D dataset with an average of 35% in terms of LPIPS. Our code and model will be made available on https://github.com/showlab/Exo2Ego-V. Jia-Wei Liu, Weijia Mao, Zhongcong Xu, Jussi Keppo, Zheng Shou 0001 |
NeurIPS | 3 |
| 2023 | HOSNeRF: Dynamic Human-Object-Scene Neural Radiance Fields from a Single VideoabstractWe introduce HOSNeRF, a novel 360° free-viewpoint rendering method that reconstructs neural radiance fields for dynamic human-object-scene from a single monocular in-the-wild video. Our method enables pausing the video at any frame and rendering all scene details (dynamic humans, objects, and backgrounds) from arbitrary viewpoints. The first challenge in this task is the complex object motions in human-object interactions, which we tackle by introducing the new object bones into the conventional human skeleton hierarchy to effectively estimate large object deformations in our dynamic human-object model. The second challenge is that humans interact with different objects at different times, for which we introduce two new learnable object state embeddings that can be used as conditions for learning our human-object representation and scene representation, respectively. Extensive experiments show that HOSNeRF significantly outperforms SOTA approaches on two challenging datasets by a large margin of 40%~50% in terms of LPIPS. The code, data, and compelling examples of 360° free-viewpoint renderings from single videos: https://showlab.github.io/HOSNeRF. Jia-Wei Liu, Yan-Pei Cao 0001, Tianyuan Yang, Zhongcong Xu, Jussi Keppo, Ying Shan, Xiaohu Qie, Zheng Shou 0001 |
ICCV | 4 |
| 2023 | XAGen: 3D Expressive Human Avatars GenerationabstractRecent advances in 3D-aware GAN models have enabled the generation of realistic and controllable human body images. However, existing methods focus on the control of major body joints, neglecting the manipulation of expressive attributes, such as facial expressions, jaw poses, hand poses, and so on. In this work, we present XAGen, the first 3D generative model for human avatars capable of expressive control over body, face, and hands. To enhance the fidelity of small-scale regions like face and hands, we devise a multi-scale and multi-part 3D representation that models fine details. Based on this representation, we propose a multi-part rendering technique that disentangles the synthesis of body, face, and hands to ease model training and enhance geometric quality. Furthermore, we design multi-part discriminators that evaluate the quality of the generated avatars with respect to their appearance and fine-grained control capabilities. Experiments show that XAGen surpasses state-of-the-art methods in terms of realism, diversity, and expressive control abilities. Code and data will be made available at https://showlab.github.io/xagen. Zhongcong Xu, Jun Hao Liew, Jiashi Feng, Zheng Shou 0001 |
NeurIPS | 1 |
| 2021 | A Novel Illumination-Robust Hand Gesture Recognition System With Event-Based Neuromorphic Vision SensorabstractThe hand gesture recognition system is a noncontact and intuitive communication approach, which, in turn, allows for natural and efficient interaction. This work focuses on developing a novel and robust gesture recognition system, which is insensitive to environmental illumination and background variation. In the field of gesture recognition, standard vision sensors, such as CMOS cameras, are widely used as the sensing devices in state-of-the-art hand gesture recognition systems. However, such cameras depend on environmental constraints, such as lighting variability and the cluttered background, which significantly deteriorates their performances. In this work, we propose an event-based gesture recognition system to overcome the detriment constraints and enhance the robustness of the recognition performance. Our system relies on a biologically inspired neuromorphic vision sensor that has microsecond temporal resolution, high dynamic range, and low latency. The sensor output is a sequence of asynchronous events instead of discrete frames. To interpret the visual data, we utilize a wearable glove as an interaction device with five high-frequency (>100 Hz) active LED markers (ALMs), representing fingers and palm, which are tracked precisely in the temporal domain using a restricted spatiotemporal particle filter algorithm. The latency of the sensing pipeline is negligible compared with the dynamics of the environment as the sensor's temporal resolution allows us to distinguish high frequencies precisely. We design an encoding process to extract features and adopt a lightweight network to classify the hand gestures. The recognition accuracy of our system is comparable to the state-of-the-art methods. To study the robustness of the system, experiments considering illumination and background variations are performed, and the results show that our system is more robust than the state-of-the-art deep learning-based gesture recognition systems. Note to Practitioners-This article addresses the robustness of the hand gesture recognition system that is important for gesture recognition-based applications. Existing methods rely on either the large-volume data to train a deep learning model or to restrict the applied environments (e.g., an ideal environment without dynamic background). However, a vision-based deep learning model requires large computational resources, while the ideal environment limits the practicality of the system. In this work, we introduce a biologically inspired neuromorphic vision sensor and an ALM glove and build a novel gesture recognition system to tackle the above issue. The neuromorphic vision sensor has a microsecond temporal resolution and a high dynamic range. With these properties, the sensing system of our prototype operates in a very low-latency space, which, in turn, ensures that our gesture recognition system is robust to illumination variance and dynamic background. Thus, this work is valuable to the research of illumination-robust gesture recognition systems. Preliminary experiments suggest that our system prototype is feasible, but it has not yet been incorporated into an online gesture recognition system nor tested with complex gestures. In future work, we will concentrate on the improvement of the signal processing methods that advance the current system to complex and practical applications. Guang Chen 0001, Zhongcong Xu, Zhijun Li 0001, Huajin Tang, Sanqing Qu, Kejia Ren, Alois C. Knoll |
IEEE Trans Autom. Sci. Eng. | 2 |