Xinjie Li 0002

dblp:67/950-2 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0002-6120-6892ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Co-Speech Gesture Video Generation with Implicit Motion-Audio Entanglement
abstract
Co-speech gestures are essential to non-verbal communication, enhancing both the naturalness and effectiveness of human interaction. Although recent methods have made progress in generating co-speech gesture videos, many rely on strong visual controls, such as pose images or TPS key-point movements, which often lead to artifacts like blurry hands and distorted fingers. In response to these challenges, we present the Implicit Motion-Audio Entanglement (IMAE) method for co-speech gesture video generation. IMAE strengthens audio control by entangling implicit motion parameters, including pose and expression, with audio inputs. Our method utilizes a two-branch framework that combines an audio-to-motion generation branch with a video diffusion branch, enabling realistic gesture generation without requiring additional inputs during inference. To improve training efficiency, we propose a two-stage slow-fast training strategy that balances memory constraints while facilitating the learning of meaningful gestures from long frame sequences. Extensive experimental results demonstrate that our method achieves state-of-the-art performance across multiple metrics. Project Page.
Xinjie Li 0002, Ziyi Chen 0005, Xinlu Yu, Iek-Heng Chu, Peng Chang 0002, Jing Xiao 0006
CVPR1
2025 DiffBody: Human Body Image Restoration with Generative Diffusion Prior
abstract
Human body image restoration is crucial for various applications but remains challenging due to the limitations of generative models: General image restoration methods built on generative models may generate unnatural textures, noticeable structural misalignments, and significant loss of fine details. To address these shortcomings, we present DiffBody, a novel human body-aware diffusion model that incorporates domain-specific knowledge to significantly enhance restoration quality. Our approach adopts a two-stage framework: (1) a multi-branch joint diffusion model generates preliminary priors, including normal and depth maps supported by a robust reconstruction pre-processing step; (2) a restoration stage refines the output using a body-prior ControlNet and a color adapter, ensuring structural accuracy and color consistency. Extensive quantitative evaluations, qualitative evaluations, and user studies validate the superior performance of DiffBody in producing perceptually high-quality human body restoration results. Code is available at https://github.com/yimingz1218/DiffBody.
Lionel Z. Wang, Sizhuo Ma, Xinjie Li 0002, Zhihang Zhong, Jian Wang 0100
ICCP4
2024 Repetitive Action Counting with Motion Feature Learning
abstract
Repetitive action counting aims to count the number of repetitive actions in a video. The critical challenge of this task is to uncover the periodic pattern between repetitive actions by computing feature similarity between frames. However, existing methods only rely on the RGB feature of each frame to compute the feature similarity while neglecting the background change of repetitive actions. The abrupt background change may cause feature discrepancies of the same action moment and lead to errors in counting. To this end, we propose a two-branch framework, i.e., RGB and motion branches, with the motion branch complementing the RGB branch to enhance the foreground motion feature learning. Specifically, foreground motion features are highlighted with flow-guided attention on frame features. In addition, to alleviate the noise from moving background distractors and reinforce the periodic pattern, we propose a temporal self-similarity matrix reconstruction loss to improve the temporal correspondence between the same motion feature from different frames. Lastly, to make the motion feature effectively supplement the RGB feature, we present a novel variance-prompted loss weights generation technique to automatically generate dynamic loss weights for two branches in collaborative training. Extensive experiments are conducted on the RepCount and UCFRep datasets to verify our proposed method with state-of-the-art performance. Our method also achieves the best performance on the cross-dataset generalization experiment.
Xinjie Li 0002, Huijuan Xu 0001
WACV1
2023 MEID: Mixture-of-Experts with Internal Distillation for Long-Tailed Video Recognition
abstract
The long-tailed video recognition problem is especially challenging, as videos tend to be long and untrimmed, and each video may contain multiple classes, causing frame-level class imbalance. The previous method tackles the long-tailed video recognition only through frame-level sampling for class re-balance without distinguishing the frame-level feature representation between head and tail classes. To improve the frame-level feature representation of tail classes, we modulate the frame-level features with an auxiliary distillation loss to reduce the distribution distance between head and tail classes. Moreover, we design a mixture-of-experts framework with two different expert designs, i.e., the first expert with an attention-based classification network handling the original long-tailed distribution, and the second expert dealing with the re-balanced distribution from class-balanced sampling. Notably, in the second expert, we specifically focus on the frames unsolved by the first expert through designing a complementary frame selection module, which inherits the attention weights from the first expert and selects frames with low attention weights, and we also enhance the motion feature representation for these selected frames. To highlight the multi-label challenge in long-tailed video recognition, we create two additional benchmarks based on Charades and CharadesEgo videos with the multi-label property, called CharadesLT and CharadesEgoLT. Extensive experiments are conducted on the existing long-tailed video benchmark VideoLT and the two new benchmarks to verify the effectiveness of our proposed method with state-of-the-art performance. The code and proposed benchmarks are released at https://github.com/VisionLanguageLab/MEID.
Xinjie Li 0002, Huijuan Xu 0001
AAAI1
2022 DANet: Dynamic Attention to Spoof Patterns for Face Anti-Spoofing
abstract
Face anti-spoofing is a vital part to protect the security of face recognition systems. Many existing face anti-spoofing methods rely on convolutional neural networks (CNNs) and achieve competitive performance. However, due to the power of CNNs, these methods will extract information that is irrelevant to spoof patterns, such as acquisition equipment and environmental characteristics, which makes the network vulnerable to changes of the illumination or camera. In this work, we propose a plug-and-play module called DyAttention, which can improve the robustness against environmental changes. Moreover, we build a network named DANet with DyAttention, which can accurately capture the spoof patterns from coarse to fine. DANet can dynamically capture the texture differences between live and spoof samples in the facial area. Specifically, we use the spatial attention mechanism to generate a mask of the facial area. Then, we extract the intrinsic texture patterns and piecewise enhance them via dynamic activation for clean representation, where the texture patterns are not affected by the environmental and domain factors. Through experiments on three benchmark datasets, our DANet achieves state-of-the-art intra-dataset accuracy on CASIA-MFSD, Replay-Attack, and OULU-NPU. Meanwhile, DANet can enhance the cross-dataset performance between CASIA-MFSD and Replay-Attack, improving the average HTER by 1.3%.
Chun-Yu Sun, Song-Lu Chen, Xinjie Li 0002, Feng Chen 0040, Xu-Cheng Yin
ICPR3
2022 SD-GAN: Semantic Decomposition for Face Image Synthesis with Discrete Attribute
abstract
Manipulating latent code in generative adversarial networks (GANs) for facial image synthesis mainly focuses on continuous attribute synthesis (e.g., age, pose and emotion), while discrete attribute synthesis (like face mask and eyeglasses) receives less attention. Directly applying existing works to facial discrete attributes may cause inaccurate results. In this work, we propose an innovative framework to tackle challenging facial discrete attribute synthesis via semantic decomposing, dubbed SD-GAN. To be concrete, we explicitly decompose the discrete attribute representation into two components, i.e. the semantic prior basis and offset latent representation. The semantic prior basis shows an initializing direction for manipulating face representation in the latent space. The offset latent presentation obtained by 3D-aware semantic fusion network is proposed to adjust prior basis. In addition, the fusion network integrates 3D embedding for better identity preservation and discrete attribute synthesis. The combination of prior basis and offset latent representation enable our method to synthesize photo-realistic face images with discrete attributes. Notably, we construct a large and valuable dataset MEGN (Face Mask and Eyeglasses images crawled from Google and Naver) for completing the lack of discrete attributes in the existing dataset. Extensive qualitative and quantitative experiments demonstrate the state-of-the-art performance of our method. Our code is available at an anonymous website: https://github.com/MontaEllis/SD-GAN.
Kangneng Zhou, Xiaobin Zhu 0001, Daiheng Gao, Kai Lee, Xinjie Li 0002, Xu-Cheng Yin
ACM Multimedia5
2022 Stereo CenterNet-based 3D object detection for autonomous driving
Yuguang Shi, Yu Guo 0001, Zhenqiang Mi, Xinjie Li 0002
Neurocomputing4
2020 Semantic Bilinear Pooling for Fine-Grained Recognition
abstract
Naturally, fine-grained recognition, e.g., vehicle identification or bird classification, has specific hierarchical labels, where fine categories are always harder to be classified than coarse categories. However, most of the recent deep learning based methods neglect the semantic structure of fine-grained objects and do not take advantage of the traditional fine-grained recognition techniques (e.g. coarse-to-fine classification). In this paper, we propose a novel framework with a two-branch network (coarse branch and fine branch), i.e., semantic bilinear pooling, for fine-grained recognition with a hierarchical label tree. This framework can adaptively learn the semantic information from the hierarchical levels. Specifically, we design a generalized cross-entropy loss for the training of the proposed framework to fully exploit the semantic priors via considering the relevance between adjacent levels and enlarge the distance between samples of different coarse classes. Furthermore, our method leverages only the fine branch when testing so that it adds no overhead to the testing time. Experimental results show that our proposed method achieves state-of-the-art performance on four public datasets.
Xinjie Li 0002, Song-Lu Chen, Chao Zhu 0003, Xu-Cheng Yin
ICPR1