Xiaozhong Ji

dblp:237/9544 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
11since 2021 · last 2025
0009-0000-1044-5853ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 9 since 2021
YearPublicationVenuePosition
2025 GroundingFace: Fine-grained Face Understanding via Pixel Grounding Multimodal Large Language Model
abstract
Multimodal Language Learning Models (MLLMs) have shown remarkable performance in image understanding, generation, and editing, with recent advancements achieving pixel-level grounding with reasoning. However, these models for common objects struggle with fine-grained face understanding. In this work, we introduce the FacePlayGround-240K dataset, the first pioneering large-scale, pixel-grounded face caption and question-answer (QA) dataset that includes 240K images, 47 mask categories, 5.4M mask annotations, and 7.3M grounded regions, meticulously curated for alignment pretraining and instruction-tuning. We present the GroundingFace framework, specifically designed to enhance fine-grained face understanding. This framework significantly augments the capabilities of existing grounding models in face part segmentation, face attribute comprehension, while preserving general scene understanding. Comprehensive experiments validate that our approach surpasses current state-of-the-art models in pixel-grounded face captioning/QA and various downstream tasks, including face captioning, referring segmentation, and zero-shot face attribute recognition.
Jiangning Zhang, Runze Hou, Xiaozhong Ji, Chuming Lin, Xiaobin Hu, Zhucun Xue, Yong Liu 0007
CVPR5
2025 Sonic: Shifting Focus to Global Audio Perception in Portrait Animation
abstract
The study of talking face generation mainly explores the intricacies of synchronizing facial movements and crafting visually appealing, temporally-coherent animations. However, due to the limited exploration of global audio perception, current approaches predominantly employ auxiliary visual and spatial knowledge to stabilize the movements, which often results in the deterioration of the naturalness and temporal inconsistencies. Considering the essence of audio-driven animation, the audio signal serves as the ideal and unique priors to adjust facial expressions and lip movements, without resorting to interference of any visual signals. Based on this motivation, we propose a novel paradigm, dubbed as Sonic, to shift focus on the exploration of global audio perception. To effectively leverage global audio knowledge, we disentangle it into intra-and inter-clip audio perception and collaborate with both aspects to enhance overall perception. For the intra-clip audio perception, 1). Context-enhanced audio learning, in which long-range intra-clip temporal audio knowledge is extracted to provide facial expression and lip motion priors implicitly expressed as the tone and speed of speech. 2). Motion-decoupled controller, in which the motion of the head and expression movement are disentangled and independently controlled by intra-audio clips. Most importantly, for inter-clip audio perception, as a bridge to connect the intra-clips to achieve the global perception, Time-aware position shift fusion, in which the global inter-clip audio information is considered and fused for long-audio inference via through consecutively time-aware shifted windows. Extensive experiments demonstrate that the novel audio-driven paradigm outperform existing SOTA methodologies in terms of video quality, temporally consistency, lip synchronization precision, and motion diversity.
Xiaozhong Ji, Xiaobin Hu, Chuming Lin, Qingdong He, Jiangning Zhang, Donghao Luo 0001, Qin Lin 0003, Qinglin Lu, Chengjie Wang 0001
CVPR1
2025 HunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animation
abstract
We introduce HunyuanPortrait, a diffusion-based condition control method that employs implicit representations for highly controllable and lifelike portrait animation. Given a single portrait image as an appearance reference and video clips as driving templates, HunyuanPortrait can animate the character in the reference image by the facial expression and head pose of the driving videos. In our framework, we utilize pre-trained encoders to achieve the decoupling of portrait motion information and identity in videos. To do so, implicit representation is adopted to encode motion information and is employed as control signals in the animation phase. By leveraging the power of stable video diffusion as the main building block, we carefully design adapter layers to inject control signals into the denoising unet through attention mechanisms. These bring spatial richness of details and temporal consistency. HunyuanPortrait also exhibits strong generalization performance, which can effectively disentangle appearance and motion under different image styles. Our framework outperforms existing methods, demonstrating superior temporal consistency and controllability. Our project is available at HunyuanPortrait.
Zunnan Xu, Zhentao Yu, Xiaoyu Jin, Fa-Ting Hong, Xiaozhong Ji, Chengfei Cai, Shiyu Tang, Qin Lin 0003, Xiu Li 0001, Qinglin Lu
CVPR7
2025 Disentangle Identity, Cooperate Emotion: Correlation-Aware Emotional Talking Portrait Generation
abstract
Recent advances in Talking Head Generation (THG) have achieved impressive lip synchronization and visual quality through diffusion models; yet existing methods struggle to generate emotionally expressive portraits while preserving speaker identity. We identify three critical limitations in current emotional talking head generation: insufficient utilization of audio's inherent emotional cues, identity leakage in emotion representations, and isolated learning of emotion correlations. To address these challenges, we propose a novel framework dubbed as DICE-Talk, following the idea of disentangling identity with emotion, and then cooperating emotions with similar characteristics. First, we develop a disentangled emotion embedder that jointly models audio-visual emotional cues through cross-modal attention, representing emotions as identity-agnostic Gaussian distributions. Second, we introduce a correlation-enhanced emotion conditioning module with learnable emotion banks that explicitly capture inter-emotion relationships through vector quantization and attention-based feature aggregation. Third, we design an emotion discrimination objective that enforces affective consistency during the diffusion process through latent-space classification. Extensive experiments on MEAD and HDTF datasets demonstrate our method's superiority, outperforming state-of-the-art approaches in emotion accuracy while maintaining competitive lip-sync performance. Qualitative results and user studies further confirm our method's ability to generate identity-preserving portraits with rich, correlated emotional expressions that naturally adapt to unseen identities.
Weipeng Tan, Chuming Lin, Chengming Xu 0001, FeiFan Xu, Xiaobin Hu, Xiaozhong Ji, Chengjie Wang 0001, Yanwei Fu 0001
ACM Multimedia6
2024 DiffuMatting: Synthesizing Arbitrary Objects with Matting-Level Annotation
Xiaobin Hu, Donghao Luo 0001, Xiaozhong Ji, Jinlong Peng, Zhengkai Jiang 0001, Jiangning Zhang, Taisong Jin, Chengjie Wang 0001, Rongrong Ji
ECCV (68)4
2024 UniM-OV3D: Uni-Modality Open-Vocabulary 3D Scene Understanding with Fine-Grained Feature Representation
Qingdong He, Jinlong Peng, Zhengkai Jiang 0001, Xiaozhong Ji, Jiangning Zhang, Yabiao Wang, Chengjie Wang 0001, Mingang Chen, Yunsheng Wu
IJCAI5
2022 Blind Face Restoration via Integrating Face Shape and Generative Priors
abstract
Blind face restoration, which aims to reconstruct high-quality images from low-quality inputs, can benefit many applications. Although existing generative-based methods achieve significant progress in producing high-quality images, they often fail to restore natural face shapes and high-fidelity facial details from severely-degraded inputs. In this work, we propose to integrate shape and generative priors to guide the challenging blind face restoration. Firstly, we set up a shape restoration module to recover reason-able facial geometry with 3D reconstruction. Secondly, a pretrained facial generator is adopted as decoder to generate photo-realistic high-resolution images. To ensure high-fidelity, hierarchical spatial features extracted from the low-quality inputs and rendered 3D images are inserted into the decoder with our proposed Adaptive Feature Fusion Block (AFFB). Moreover, we introduce hybrid-level losses to Jointly train the shape and generative priors together with other network parts such that these two priors better adapt to our blind face restoration task. The proposed Shape and Generative Prior integrated Network (SGPN) can re-store high-quality images with clear face shapes and real-istic facial details. Experimental results on synthetic and real-world datasets demonstrate SGPN performs favorably against state-of-the-art blind face restoration methods.
Feida Zhu 0002, Wenqing Chu, Xinyi Zhang 0005, Xiaozhong Ji, Chengjie Wang 0001, Ying Tai
CVPR5
2022 ColorFormer: Image Colorization via Color Memory Assisted Hybrid-Attention Transformer
Xiaozhong Ji, Boyuan Jiang, Donghao Luo 0001, Guangpin Tao, Wenqing Chu, Chengjie Wang 0001, Ying Tai
ECCV (16)1
2022 Efficient Reinforcement Learning for StarCraft by Abstract Forward Models and Transfer Learning
abstract
Injecting human knowledge is an effective way to accelerate reinforcement learning (RL). However, these methods are underexplored. This article presents our discovery that an abstract forward model [thought-game (TG)] combined with transfer learning is an effective way. We takeStarCraft IIas our study environment. With the help of a designed TG, the agent can learn a 99% win-rate on a 64×64 map against the Level-7 built-in AI, using only 1.08 h in a single commercial machine. We also show that the TG method is not as restrictive as it was thought to be. It can work with roughly designed TGs, and can also be useful when the environment changes. Comparing with previous model-based RL, we show TG is more effective. We also present a TG hypothesis that gives the influence of different fidelity levels of TG. For real games that have unequal state and action spaces, we proposed a novel XfrNet of which usefulness is validated while achieving a 90% win-rate against the cheating Level-10 AI. We argue that the TG method might shed light on further studies of efficient RL with human knowledge.
Ruo-Ze Liu, Xiaozhong Ji, Yang Yu 0001, Zhen-Jia Pang, Zitai Xiao, Yuzhou Wu, Tong Lu 0002
IEEE Trans. Games3
2021 Frequency Consistent Adaptation for Real World Super Resolution
abstract
Recent deep-learning based Super-Resolution (SR) methods have achieved remarkable performance on images with known degradation. However, these methods always fail in real-world scene, since the Low-Resolution (LR) images after the ideal degradation (e.g., bicubic down-sampling) deviate from real source domain. The domain gap between the LR images and the real-world images can be observed clearly on frequency density, which inspires us to explicitly narrow the undesired gap caused by incorrect degradation. From this point of view, we design a novel Frequency Consistent Adaptation (FCA) that ensures the frequency domain consistency when applying existing SR methods to the real scene. We estimate degradation kernels from unsupervised images and generate the corresponding LR images. To provide useful gradient information for kernel estimation, we propose Frequency Density Comparator (FDC) by distinguishing the frequency density of images on different scales. Based on the domain-consistent LR-HR pairs, we train easy-implemented Convolutional Neural Network (CNN) SR models. Extensive experiments show that the proposed FCA improves the performance of the SR model under real-world setting achieving state-of-the-art results with high fidelity and plausible perception, thus providing a novel effective framework for real-world SR application.
Xiaozhong Ji, Guangpin Tao, Yun Cao 0002, Ying Tai, Tong Lu 0002, Chengjie Wang 0001, Feiyue Huang
AAAI1
2021 Spectrum-to-Kernel Translation for Accurate Blind Image Super-Resolution
abstract
Deep-learning based Super-Resolution (SR) methods have exhibited promising performance under non-blind setting where blur kernel is known; however, blur kernels of Low-Resolution (LR) images in different practical applications are usually unknown. It may lead to a significant performance drop when degradation process of training images deviates from that of real images. In this paper, we propose a novel blind SR framework to super-resolve LR images degraded by arbitrary blur kernel with accurate kernel estimation in frequency domain. To our best knowledge, this is the first deep learning method which conducts blur kernel estimation in frequency domain. Specifically, we first demonstrate that feature representation in frequency domain is more conducive for blur kernel reconstruction than in spatial domain. Next, we present a Spectrum-to-Kernel (S$2$K) network to estimate general blur kernels in diverse forms. We use a conditional GAN (CGAN) combined with SR-oriented optimization target to learn the end-to-end translation from degraded images' spectra to unknown kernels. Extensive experiments on both synthetic and real-world images demonstrate that our proposed method sufficiently reduces blur kernel estimation error, thus enables the off-the-shelf non-blind SR methods to work under blind setting effectively, and achieves superior performance over state-of-the-art blind SR methods, averagely by 1.39dB, 0.48dB (Gaussian kernels) and 6.15dB, 4.57dB (motion kernels) for scales $2\times$ and $4\times$ respectively.
Guangpin Tao, Xiaozhong Ji, Wenzhuo Wang, Shuo Chen 0003, Chuming Lin, Yun Cao 0002, Tong Lu 0002, Donghao Luo 0001, Ying Tai
NeurIPS2
2020 AE TextSpotter: Learning Visual and Linguistic Representation for Ambiguous Text Spotting
Wenhai Wang, Xuebo Liu 0001, Xiaozhong Ji, Enze Xie, Ding Liang, Zhibo Yang 0003, Tong Lu 0002, Chunhua Shen, Ping Luo 0002
ECCV (14)3
2020 Context-Aware Residual Network with Promotion Gates for Single Image Super-Resolution
Xiaozhong Ji, Yirui Wu, Tong Lu 0002
MMM (2)1
2020 CASR: a context-aware residual network for single-image super-resolution
Yirui Wu, Xiaozhong Ji, Wanting Ji, Helen Zhou
Neural Comput. Appl.2
2019 A Novel Two-Factor Attention Encoder-Decoder Network through Combining Temporal and Prior Knowledge for Weather Forecasting
abstract
This paper proposes a novel two-factor attention based encoder-decoder model (TwoFactorEncoderDecoder) for multivariate weather prediction. The proposed model learns attention weights from two factors, namely, temporal information and prior knowledge inferred information. Here, temporal information contains change patterns hidden in observed time series data, while prior knowledge inferred information gives various types of meteorological observations in weather forecasting. Attention weights of the two factors are used to select the intermediate outputs of the encoder, and then combine the selected result with information inferred by prior knowledge for weather forecasting by a more effective way. In addition, this paper proposes a loss function for multivariate prediction. Compared with Mean Square Error (MSE) loss function, the proposed loss function can fit small variances more accurately in performing multivariate prediction. Compared with the attention model that only uses temporal information or the prior knowledge inferred information, the proposed TwoFactorEncoderDecoder model has encouraging improvements in prediction accuracy on the public weather forecasting dataset, namely, the MAPE of t2m is increased by 5.42%, the MAPE of rh2m is increased by 2.92%, and the MAPE of w2m is increased by 1.67%, which shows the effect of the two-factor attention mechanism. Source code for the complete system will be available at https://github.com/YuanMLer/TFAEncoderDecoder.
Minglei Yuan, Xiaozhong Ji, Hualu Zhang
IJCNN2