EDBT 2026 Demo / reviewers in the wild / expert
Chao Xu 0023
dblp:79/1442-23
· DBLP profile ↗
27ranked-venue papers
10as first author
25since 2021 · last 2026
0000-0001-9300-4497ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 9 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Knowledge is Power: Advancing Few-shot Action Recognition with Multimodal Semantics from MLLMs
Jiazheng Xing, Chao Xu 0023, Hangjie Yuan, Mengmeng Wang 0005, Jun Dan, Hangwei Qian, Yong Liu 0007 |
Int. J. Comput. Vis. | 2 |
| 2026 | MA-FSAR: Multimodal Adaptation of CLIP for few-shot action recognition
Jiazheng Xing, Jian Zhao 0006, Chao Xu 0023, Mengmeng Wang 0005, Guang Dai, Yong Liu 0007, Jingdong Wang 0001, Xuelong Li 0001 |
Pattern Recognit. | 3 |
| 2026 | Corrigendum to "MA-FSAR: Multimodal Adaptation of CLIP for few-shot action recognition" [Pattern Recognition 169 (2026) 111902]
Jiazheng Xing, Jian Zhao 0006, Chao Xu 0023, Mengmeng Wang 0005, Guang Dai, Yong Liu 0007, Jingdong Wang 0001, Xuelong Li 0001 |
Pattern Recognit. | 3 |
| 2026 | Corrigendum to "FaceChain-MMID: Generating highly identity-consistent realistic portraits via dividing & merging multi-modal representations" [Pattern Recognition 168 (2025) 111858]
Chao Xu 0023, Fei Wang 0032, Baigui Sun, Jian Zhao 0006 |
Pattern Recognit. | 1 |
| 2025 | Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual TransformersabstractWhile language models (LMs) paired with residual vector quantization (RVQ) tokenizers have shown promise in text-to-audio (T2A) generation, they still lag behind diffusion-based models by a non-trivial margin.We identify a critical dilemma underpinning this gap: incorporating more RVQ layers improves audio reconstruction fidelity but exceeds the generation capacity of conventional LMs.To address this, we first analyze RVQ dynamics and uncover two key limitations: 1) orthogonality of features across RVQ layers hinders effective LMs training, and 2) descending semantic richness in tokens from deeper RVQ layers exacerbates exposure bias during autoregressive decoding.Based on these insights, we propose Siren, a novel LM-based framework that employs multiple isolated transformers with causal conditioning and anti-causal alignment via reinforcement learning.Extensive experiments demonstrate that Siren outperforms both existing LM-based and diffusion-based T2A systems, achieving state-of-the-art results.By bridging the representational strengths of LMs with the fidelity demands of audio synthesis, our approach repositions LMs as competitive contenders against diffusion models in T2A tasks.Moreover, by aligning audio representations with linguistic structures, Siren facilitates a promising pathway toward unified multimodal generation frameworks.The code is released at https://github.com/wjc2830/Siren.git. Chao Xu 0023, Haoyu Xie 0002, Guoqi Yu |
EMNLP | 2 |
| 2025 | Beyond Talking - Generating Holistic 3D Human Dyadic Motion for Communication
Chao Xu 0023, Yang Liu 0356, Baigui Sun, Ruqi Huang |
Int. J. Comput. Vis. | 2 |
| 2025 | TryOn-Adapter: Efficient Fine-Grained Clothing Identity Adaptation for High-Fidelity Virtual Try-On
Jiazheng Xing, Chao Xu 0023, Yijie Qian, Yang Liu 0356, Guang Dai, Baigui Sun, Yong Liu 0007, Jingdong Wang 0001 |
Int. J. Comput. Vis. | 2 |
| 2025 | UniFace++: Revisiting a Unified Framework for Face Reenactment and Swapping via 3D Priors
Chao Xu 0023, Yijie Qian, Shaoting Zhu, Baigui Sun, Jian Zhao 0006, Yong Liu 0007, Xuelong Li 0001 |
Int. J. Comput. Vis. | 1 |
| 2025 | HybrIK-X: Hybrid Analytical-Neural Inverse Kinematics for Whole-Body Mesh RecoveryabstractRecovering whole-body mesh by inferring the abstract pose and shape parameters from visual content can obtain 3D bodies with realistic structures. However, the inferring process is highly non-linear and suffers from image-mesh misalignment, resulting in inaccurate reconstruction. In contrast, 3D keypoint estimation methods utilize the volumetric representation to achieve pixel-level accuracy but may predict unrealistic body structures. To address these issues, this paper presents a novel hybrid inverse kinematics solution, HybrIK, that integrates the merits of 3D keypoint estimation and body mesh recovery in a unified framework. HybrIK directly transforms accurate 3D joints to body-part rotations via twist-and-swing decomposition. The swing rotations are analytically solved with 3D joints, while the twist rotations are derived from visual cues through neural networks. To capture comprehensive whole-body details, we further develop a holistic framework, HybrIK-X, which enhances HybrIK with articulated hands and an expressive face. HybrIK-X is fast and accurate by solving the whole-body pose with a one-stage model. Experiments demonstrate that HybrIK and HybrIK-X preserve both the accuracy of 3D joints and the realistic structure of the parametric human model, leading to pixel-aligned whole-body mesh recovery. The proposed method significantly surpasses the state-of-the-art methods on various benchmarks for body-only, hand-only, and whole-body scenarios. Siyuan Bian, Chao Xu 0023, Zhicun Chen, Lixin Yang 0001, Cewu Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | FaceChain-MMID: Generating highly identity-consistent realistic portraits via dividing & merging multi-modal representations
Chao Xu 0023, Fei Wang 0032, Baigui Sun, Jian Zhao 0006 |
Pattern Recognit. | 1 |
| 2024 | FaceChain-ImagineID: Freely Crafting High-Fidelity Diverse Talking Faces from Disentangled AudioabstractIn this paper, we abstract the process of people hearing speech, extracting meaningful cues, and creating vari-ous dynamically audio-consistent talking faces, termed Lis-tening and Imagining, into the task of high-fidelity diverse talking faces generation from a single audio. Specifically, it involves two critical challenges: one is to effectively de-couple identity, content, and emotion from entangled au-dio, and the other is to maintain intra-video diversity and inter- video consistency. To tackle the issues, we first dig out the intricate relationships among facial factors and sim-plify the decoupling process, tailoring a Progressive Audio Disentanglement for accurate facial geometry and seman-tics learning, where each stage incorporates a customized training module responsible for a specific factor. Secondly, to achieve visually diverse and audio-synchronized animation solely from input audio within a single model, we intro-duce the Controllable Coherent Frame generation, which involves the flexible integration of three trainable adapters with frozen Latent Diffusion Models (LDMs) to focus on maintaining facial geometry and semantics, as well as texsture and temporal coherence between frames. In this way, we inherit high-quality diverse generation from LDMs while significantly improving their controllability at a low training cost. Extensive experiments demonstrate the flexibility and effectiveness of our method in handling this paradigm. The codes will be released at FaceChain. Chao Xu 0023, Yang Liu 0356, Jiazheng Xing, Weida Wang, Jun Dan, Tianxin Huang, Siyuan Li 0002, Zhi-Qi Cheng, Ying Tai, Baigui Sun |
CVPR | 1 |
| 2024 | Semi-supervised noise-resilient anomaly detection with feature autoencoder
Lina Liu 0010, Yuanlong Zhang, Chao Xu 0023, Jun Chen 0023 |
Knowl. Based Syst. | 6 |
| 2023 | Better "CMOS" Produces Clearer Images: Learning Space-Variant Blur Estimation for Blind Image Super-ResolutionabstractMost of the existing blind image Super-Resolution (SR) methods assume that the blur kernels are space-invariant. However, the blur involved in real applications are usually space-variant due to object motion, out-of-focus, etc., resulting in severe performance drop of the advanced SR methods. To address this problem, we firstly introduce two new datasets with out-of-focus blur, i.e., NYUv2-BSR and Cityscapes-BSR, to support further researches of blind SR with space-variant blur. Based on the datasets, we design a novel Cross-MOdal fuSion network (CMOS) that estimate both blur and semantics simultaneously, which leads to improved SR results. It involves a feature Grouping Interactive Attention (GIA) module to make the two modalities in-teract more effectively and avoid inconsistency. GIA can also be used for the interaction of other features because of the universality of its structure. Qualitative and quantitative experiments compared with state-of-the-art methods on above datasets and real-world images demonstrate the superiority of our method, e.g., obtaining PSNR/SSIM by +1.91↑/+0.0048↑ on NYUv2-BSR than MANet11https://github.com/ByChelsea/CMOS.git. Xuhai Chen, Jiangning Zhang, Chao Xu 0023, Yabiao Wang, Chengjie Wang 0001, Yong Liu 0007 |
CVPR | 3 |
| 2023 | High-Fidelity Generalized Emotional Talking Face Generation with Multi-Modal Emotion Space LearningabstractRecently, emotional talking face generation has received considerable attention. However, existing methods only adopt one-hot coding, image, or audio as emotion conditions, thus lacking flexible control in practical applications and failing to handle unseen emotion styles due to limited semantics. They either ignore the one-shot setting or the quality of generated faces. In this paper, we propose a more flexible and generalized framework. Specifically, we supplement the emotion style in text prompts and use an Aligned Multi-modal Emotion encoder to embed the text, image, and audio emotion modality into a unified space, which inherits rich semantic prior from CLIP. Consequently, effective multi-modal emotion space learning helps our method support arbitrary emotion modality during testing and could generalize to unseen emotion styles. Besides, an Emotion-aware Audio-to-3DMM Convertor is proposed to connect the emotion condition and the audio sequence to structural representation. A followed style-based High-fidelity Emotional Face generator is designed to generate arbitrary high-resolution realistic identities. Our texture generator hierarchically learns flow fields and animated faces in a residual manner. Extensive experiments demonstrate the flexibility and generalization of our method in emotion control and the effectiveness of high-quality face synthesis. Chao Xu 0023, Jiangning Zhang, Wenqing Chu, Ying Tai, Chengjie Wang 0001, Yong Liu 0007 |
CVPR | 1 |
| 2023 | Improving dynamic gesture recognition in untrimmed videos by an online lightweight framework and a new gesture dataset ZJUGesture
Chao Xu 0023, Mengmeng Wang 0005, Yong Liu 0007 |
Neurocomputing | 1 |
| 2023 | AlphaPose: Whole-Body Regional Multi-Person Pose Estimation and Tracking in Real-TimeabstractAccurate whole-body multi-person pose estimation and tracking is an important yet challenging topic in computer vision. To capture the subtle actions of humans for complex behavior analysis, whole-body pose estimation including the face, body, hand and foot is essential over conventional body-only pose estimation. In this article, we present AlphaPose, a system that can perform accurate whole-body pose estimation and tracking jointly while running in realtime. To this end, we propose several new techniques: Symmetric Integral Keypoint Regression (SIKR) for fast and fine localization, Parametric Pose Non-Maximum-Suppression (P-NMS) for eliminating redundant human detections and Pose Aware Identity Embedding for jointly pose estimation and tracking. During training, we resort to Part-Guided Proposal Generator (PGPG) and multi-domain knowledge distillation to further improve the accuracy. Our method is able to localize whole-body keypoints accurately and tracks humans simultaneously given inaccurate bounding boxes and redundant detections. We show a significant improvement over current state-of-the-art methods in both speed and accuracy on COCO-wholebody, COCO, PoseTrack, and our proposed Halpe-FullBody pose estimation dataset. Our model, source codes and dataset are made publicly available at https://github.com/MVIG-SJTU/AlphaPose. Haoshu Fang, Hongyang Tang, Chao Xu 0023, Haoyi Zhu, Yuliang Xiu, Yong-Lu Li 0001, Cewu Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | SCSNet: An Efficient Paradigm for Learning Simultaneously Image Colorization and Super-resolutionabstractIn the practical application of restoring low-resolution gray-scale images, we generally need to run three separate processes of image colorization, super-resolution, and dows-sampling operation for the target device. However, this pipeline is redundant and inefficient for the independent processes, and some inner features could have been shared. Therefore, we present an efficient paradigm to perform Simultaneously Image Colorization and Super-resolution (SCS) and propose an end-to-end SCSNet to achieve this goal. The proposed method consists of two parts: colorization branch for learning color information that employs the proposed plug-and-play Pyramid Valve Cross Attention (PVCAttn) module to aggregate feature maps between source and reference images; and super-resolution branch for integrating color and texture information to predict target images, which uses the designed Continuous Pixel Mapping (CPM) module to predict high-resolution images at continuous magnification. Furthermore, our SCSNet supports both automatic and referential modes that is more flexible for practical application. Abundant experiments demonstrate the superiority of our method for generating authentic images over state-of-the-art methods, e.g., averagely decreasing FID by 1.8 and 5.1 compared with current best scores for automatic and referential modes, respectively, while owning fewer parameters (more than x2) and faster running speed (more than x3). Jiangning Zhang, Chao Xu 0023, Jian Li 0062, Yabiao Wang, Ying Tai, Yong Liu 0007 |
AAAI | 2 |
| 2022 | Region-Aware Face SwappingabstractThis paper presents a novel Region-Aware Face Swapping (RAFSwap) network to achieve identity-consistent harmonious high-resolution face generation in a local-global manner: 1) Local Facial Region-Aware (FRA) branch augments local identity-relevant features by introducing the Transformer to effectively model misaligned crossscale semantic interaction. 2) Global Source Feature-Adaptive (SFA) branch further complements global identity-relevant cues for generating identity-consistent swapped faces. Besides, we propose a Face Mask Predictor (FMP) module incorporated with StyleGAN2 to predict identity-relevant soft facial masks in an unsupervised manner that is more practical for generating harmonious high-resolution faces. Abundant experiments qualitatively and quantitatively demonstrate the superiority of our method for generating more identity-consistent high-resolution swapped faces over SOTA methods, e.g., obtaining 96.70 ID retrieval that outperforms SOTA MegaFS by$5.87\uparrow$. Chao Xu 0023, Jiangning Zhang, Miao Hua, Zili Yi, Yong Liu 0007 |
CVPR | 1 |
| 2022 | D &D: Learning Human Dynamics from Dynamic Camera
Siyuan Bian, Chao Xu 0023, Gang Yu 0002, Cewu Lu |
ECCV (5) | 3 |
| 2022 | Designing One Unified Framework for High-Fidelity Face Reenactment and Swapping
Chao Xu 0023, Jiangning Zhang, Guanzhong Tian, Xianfang Zeng, Ying Tai, Yabiao Wang, Chengjie Wang 0001, Yong Liu 0007 |
ECCV (15) | 1 |
| 2022 | Real-Time Audio-Guided Multi-Face ReenactmentabstractAudio-guided face reenactment aims to generate authentic target faces that have matched facial expression of the input audio, and many learning-based methods have successfully achieved this. However, most methods can only reenact a particular person once trained or suffer from the low-quality generation of the target images. Also, nearly none of the current reenactment works consider the model size and running speed that are important for practical use. To solve the above challenges, we propose an efficientAudio-guidedMulti-face reenactment model namedAMNet, which can reenact target faces among multiple persons with corresponding source faces and drive signals as inputs. Concretely, we design aGeometric Controller(GC) module to inject the drive signals so that the model can be optimized in an end-to-end manner and generate more authentic images. Also, we adopt a lightweight network for our face reenactor so that the model can run in real-time on both CPU and GPU devices. Abundant experiments prove our approach’s superiority over existing methods,e.g., averagely decreasing FID by 0.12$\downarrow$and increasing SSIM by 0.031$\uparrow$than APB2Face, while owning fewer parameters ($\times 4 \downarrow$) and faster CPU speed ($\times 4 \uparrow$). Jiangning Zhang, Xianfang Zeng, Chao Xu 0023, Yong Liu 0007 |
IEEE Signal Process. Lett. | 3 |
| 2022 | Multilevel Spatial-Temporal Feature Aggregation for Video Object DetectionabstractVideo object detection (VOD) focuses on detecting objects for each frame in a video, which is a challenging task due to appearance deterioration in certain video frames. Recent works usually distill crucial information from multiple support frames to improve the reference features, but they only perform at frame level or proposal level that cannot integrate spatial-temporal features sufficiently. To deal with this challenge, we treat VOD as a spatial-temporal hierarchical features interacting process and introduce a Multi-level Spatial-Temporal (MST) feature aggregation framework to fully exploit frame-level, proposal-level, and instance-level information in a unified framework. Specifically, MST first measures context similarity in pixel space to enhance all frame-level features rather than only update reference features. The proposal-level feature aggregation then models object relation to augment reference object proposals. Furthermore, to filter out irrelevant information from other classes and backgrounds, we introduce an instance ID constraint to boost instance-level features by leveraging support object proposal features that belong to the same object. Besides, we propose a Deformable Feature Alignment (DAlign) module before MST to achieve a more accurate pixel-level spatial alignment for better feature aggregation. Extensive experiments are conducted on ImageNet VID and UAVDT datasets that demonstrate the superiority of our method over state-of-the-art (SOTA) methods. Our method achieves 83.3% and 62.1% with ResNet-101 on two datasets, outperforming SOTA MEGA by 0.4% and 2.7%. Chao Xu 0023, Jiangning Zhang, Mengmeng Wang 0005, Guanzhong Tian, Yong Liu 0007 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | HybrIK: A Hybrid Analytical-Neural Inverse Kinematics Solution for 3D Human Pose and Shape EstimationabstractModel-based 3D pose and shape estimation methods reconstruct a full 3D mesh for the human body by estimating several parameters. However, learning the abstract parameters is a highly non-linear process and suffers from image-model misalignment, leading to mediocre model performance. In contrast, 3D keypoint estimation methods combine deep CNN network with the volumetric representation to achieve pixel-level localization accuracy but may predict unrealistic body structure. In this paper, we address the above issues by bridging the gap between body mesh estimation and 3D keypoint estimation. We propose a novel hybrid inverse kinematics solution (HybrIK). HybrIK directly transforms accurate 3D joints to relative body-part rotations for 3D body mesh reconstruction, via the twistand-swing decomposition. The swing rotation is analytically solved with 3D joints, and the twist rotation is derived from the visual cues through the neural network. We show that HybrIK preserves both the accuracy of 3D pose and the realistic body structure of the parametric human model, leading to a pixel-aligned 3D body mesh and a more accurate 3D pose than the pure 3D keypoint estimation methods. Without bells and whistles, the proposed method surpasses the state-of-the-art methods by a large margin on various 3D human pose and shape benchmarks. As an illustrative example, HybrIK outperforms all the previous methods by 13.2 mm MPJPE and 21.9 mm PVE on 3DPW dataset. Our code is available at https://github.com/Jeff-sjtu/HybrIK. Chao Xu 0023, Zhicun Chen, Siyuan Bian, Lixin Yang 0001, Cewu Lu |
CVPR | 2 |
| 2021 | Analogous to Evolutionary Algorithm: Designing a Unified Sequence ModelabstractInspired by biological evolution, we explain the rationality of Vision Transformer by analogy with the proven practical Evolutionary Algorithm (EA) and derive that both of them have consistent mathematical representation. Analogous to the dynamic local population in EA, we improve the existing transformer structure and propose a more efficient EAT model, and design task-related heads to deal with different tasks more flexibly. Moreover, we introduce the spatial-filling curve into the current vision transformer to sequence image data into a uniform sequential format. Thus we can design a unified EAT framework to address multi-modal tasks, separating the network architecture from the data format adaptation. Our approach achieves state-of-the-art results on the ImageNet classification task compared with recent vision transformer works while having smaller parameters and greater throughput. We further conduct multi-modal tasks to demonstrate the superiority of the unified EAT, \eg, Text-Based Image Retrieval, and our approach improves the rank-1 by +3.7 points over the baseline on the CSS dataset. Jiangning Zhang, Chao Xu 0023, Jian Li 0062, Wenzhou Chen, Yabiao Wang, Ying Tai, Shuo Chen 0003, Chengjie Wang 0001, Feiyue Huang, Yong Liu 0007 |
NeurIPS | 2 |
| 2021 | Cross-modality online distillation for multi-view action recognition
Chao Xu 0023, Yachun Li, Yining Jin, Mengmeng Wang 0005, Yong Liu 0007 |
Neurocomputing | 1 |
| 2020 | DTVNet: Dynamic Time-Lapse Video Generation via Single Still Image
Jiangning Zhang, Chao Xu 0023, Liang Liu 0007, Mengmeng Wang 0005, Yong Liu 0007, Yunliang Jiang |
ECCV (5) | 2 |
| 2018 | Environment Upgrade Reinforcement Learning for Non-Differentiable Multi-Stage PipelinesabstractRecent advances in multi-stage algorithms have shown great promise, but two important problems still remain. First of all, at inference time, information can't feed back from downstream to upstream. Second, at training time, end-to-end training is not possible if the overall pipeline involves non-differentiable functions, and so different stages can't be jointly optimized. In this paper, we propose a novel environment upgrade reinforcement learning framework to solve the feedback and joint optimization problems. Our framework re-links the downstream stage to the upstream stage by a reinforcement learning agent. While training the agent to improve final performance by refining the upstream stage's output, we also upgrade the downstream stage (environment) according to the agent's policy. In this way, agent policy and environment are jointly optimized. We propose a training algorithm for this framework to address the different training demands of agent and environment. Experiments on instance segmentation and human pose estimation demonstrate the effectiveness of the proposed framework. Shuqin Xie, Zitian Chen, Chao Xu 0023, Cewu Lu |
CVPR | 3 |