Jun Xue 0001

dblp:39/2690-1 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
13since 2021 · last 2026
0009-0001-8465-011XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021
YearPublicationVenuePosition
2026 Masked autoencoders for spatio-temporal audio representations: Theory and optimization
Jiayu Xiong, Jing Wang 0049, Wanlong Wang, Xiaosen Lyu, Jianlong Kwan, Jun Xue 0001
Pattern Recognit.6
2025 Region-Based Optimization in Continual Learning for Audio Deepfake Detection
abstract
Rapid advancements in speech synthesis and voice conversion bring convenience but also new security risks, creating an urgent need for effective audio deepfake detection. Although current models perform well, their effectiveness diminishes when confronted with the diverse and evolving nature of real-world deepfakes. To address this issue, we propose a continual learning method named Region-Based Optimization (RegO) for audio deepfake detection. Specifically, we use the Fisher information matrix to measure important neuron regions for real and fake audio detection, dividing them into four regions. First, we directly fine-tune the less important regions to quickly adapt to new tasks. Next, we apply gradient optimization in parallel for regions important only to real audio detection, and in orthogonal directions for regions important only to fake audio detection. For regions that are important to both, we use sample proportion-based adaptive gradient optimization. This region-adaptive optimization ensures an appropriate trade-off between memory stability and learning plasticity. Additionally, to address the increase of redundant neurons from old tasks, we further introduce the Ebbinghaus forgetting mechanism to release them, thereby promoting the model’s ability to learn more generalized discriminative features. Experimental results show our method achieves a 21.3 percent improvement in EER over the state-of-the-art continual learning approach RWM for audio deepfake detection. Moreover, the effectiveness of RegO extends beyond the audio deepfake detection domain, showing potential significance in other tasks, such as image recognition.
Yujie Chen 0006, Jiangyan Yi, Cunhang Fan, Jianhua Tao 0001, Yong Ren 0006, Siding Zeng, Chu Yuan Zhang, Xinrui Yan, Jun Xue 0001, Chenglong Wang 0001, Zhao Lv, Xiaohui Zhang 0006
AAAI10
2025 Multi-Level Contrastive Learning: Hierarchical Alleviation of Heterogeneity in Multimodal Sentiment Analysis
abstract
Recently, multimodal fusion efforts have achieved remarkable success in Multimodal Sentiment Analysis (MSA). However, most of the existing methods are based on model-level fusion, and the challenge of heterogeneity between modalities is not well resolved. Heterogeneity lies in the different feature distributions and distinct representation spaces among different modalities. To mitigate this problem, we propose that fusion is a progressive process, and we introduce a novel multi-level contrastive learning and multi-layer convolution fusion (MCL-MCF) method for MSA. Due to the relationships among multimodal data, the fusion process that involves single-modal to single-modal, single-modal to bimodal or trimodal, and higher-level fused modality semantic consistency is divided into three levels. The first-level contrast learning alleviates heterogeneity between unimodal modalities at the early level of multimodal feature fusion. The second-level contrast learning mitigates heterogeneity between unimodal and fused modalities. At the third level, we introduce a tensor convolution fusion (TCF) module that extracts high-level semantic features from the fused modalities and mitigates heterogeneity at the higher feature level through contrastive learning. To simulate fusion as a progressive process, MCF is proposed to fuse shallow and deep features to model complex relationships among modalities. Experiments on three public datasets show our approach's state-of-the-art performance.
Cunhang Fan, Kang Zhu, Jianhua Tao 0001, Guofeng Yi, Jun Xue 0001, Zhao Lv
IEEE Trans. Affect. Comput.5
2024 Progressive Distillation Based on Masked Generation Feature Method for Knowledge Graph Completion
abstract
In recent years, knowledge graph completion (KGC) models based on pre-trained language model (PLM) have shown promising results. However, the large number of parameters and high computational cost of PLM models pose challenges for their application in downstream tasks. This paper proposes a progressive distillation method based on masked generation features for KGC task, aiming to significantly reduce the complexity of pre-trained models. Specifically, we perform pre-distillation on PLM to obtain high-quality teacher models, and compress the PLM network to obtain multi-grade student models. However, traditional feature distillation suffers from the limitation of having a single representation of information in teacher models. To solve this problem, we propose masked generation of teacher-student features, which contain richer representation information. Furthermore, there is a significant gap in representation ability between teacher and student. Therefore, we design a progressive distillation method to distill student models at each grade level, enabling efficient knowledge transfer from teachers to students. The experimental results demonstrate that the model in the pre-distillation stage surpasses the existing state-of-the-art methods. Furthermore, in the progressive distillation stage, the model significantly reduces the model parameters while maintaining a certain level of performance. Specifically, the model parameters of the lower-grade student model are reduced by 56.7\% compared to the baseline.
Cunhang Fan, Yujie Chen 0006, Jun Xue 0001, Yonghui Kong, Jianhua Tao 0001, Zhao Lv
AAAI3
2024 Dual-View Multimodal Interaction in Multimodal Sentiment Analysis
abstract
Outstanding performance in sentiment analysis not only relies on the design of sophisticated fusion methods but also on the crucial step of designing excellent modal interaction methods. To the best of our knowledge, there are few methods addressing the capture of multimodal spatial features. Majority of feature interactions have been primarily focused on temporal aspects, with less attention given to the combined spatiotemporal feature interaction (SFI). In this paper, we design a dual-view multimodal interaction method, named DVMI, primarily consisting of two parts. In the first part, a triangular convolutional module is proposed for ample temporal interaction between modalities, implicit local and global SFI, and capturing global spatial representations. Building upon the foundation laid in the first part, the second part employs an attention mechanism for explicit global SFI. To demonstrate the effectiveness of the DVMI framework,we conduct extensive experiments on three datasets, achieving state-of-the-art experimental results.
Kang Zhu, Cunhang Fan, Jianhua Tao 0001, Jun Xue 0001, Xuefei Liu, Zhengqi Wen, Zhao Lv
ICME4
2024 A Debiased Domain Adaptation Framework with Minimum Class Confusion for Motor Imagery Decoding
abstract
Recently, motor imagery decoding technology based on electroencephalogram (EEG) signals has made significant progress. However, there are still challenges in adapting to new sessions, mainly due to changes in data distribution between different sessions and confusion problems caused by similar oscillation patterns in various categories of EEG signals. To address these problems, this paper proposes a debiased domain adaptation framework with minimum class confusion to learn unbiased representations in motor imagery tasks. Specially, unlike the feature alignment and adversarial training methods, we explore the class predictions for domain adaptation, applying the minimum class confusion loss criterion in the target domain to reduce inter-class confusion. This approach aims to alleviate the bias issues inherent in classifiers trained on the source domain for making predictions in the target domain, achieving class-level alignment. Consequently, it enhances the model’s ability to adapt to the data distribution of the target domain. Experimental results on two public EEG datasets (BCI Competition IV datasets IIa and IIb) show that the method for cross-session decoding is significantly improved compared to the baseline, with average classification accuracy reaching 81.01% and 82.92%, respectively.
Cunhang Fan, Zhen Chen 0022, Xun Song, Jun Xue 0001, Ping Li 0020, Zhao Lv
IJCNN5
2024 RawBMamba: End-to-End Bidirectional State Space Model for Audio Deepfake Detection
abstract
Fake artefacts for discriminating between bonafide and fake audio can exist in both short-and long-range segments.Therefore, combining local and global feature information can effectively discriminate between bonafide and fake audio.This paper proposes an end-to-end bidirectional state space model, named RawBMamba, to capture both short-and long-range discriminative information for audio deepfake detection.Specifically, we use sinc Layer and multiple convolutional layers to capture short-range features, and then design a bidirectional Mamba to address Mamba's unidirectional modelling problem and further capture long-range feature information.Moreover, we develop a bidirectional fusion module to integrate embeddings, enhancing audio context representation and combining shortand long-range information.The results show that our proposed RawBMamba achieves a 34.1% improvement over Rawformer on ATSVspoof2021 LA dataset, and demonstrates competitive performance on other datasets.Codes will be released on https://github.com/cyjie429/RawBMamba.
Yujie Chen 0006, Jiangyan Yi, Jun Xue 0001, Chenglong Wang 0001, Xiaohui Zhang 0006, Shunbo Dong, Siding Zeng, Jianhua Tao 0001, Zhao Lv, Cunhang Fan
INTERSPEECH3
2024 Frequency-mix Knowledge Distillation for Fake Speech Detection
Cunhang Fan, Shunbo Dong, Jun Xue 0001, Yujie Chen 0006, Jiangyan Yi, Zhao Lv
INTERSPEECH3
2024 Spatial reconstructed local attention Res2Net with F0 subband for fake speech detection
Cunhang Fan, Jun Xue 0001, Jianhua Tao 0001, Jiangyan Yi, Chenglong Wang 0001, Chengshi Zheng, Zhao Lv
Neural Networks2
2024 DGSD: Dynamical graph self-distillation for EEG-based auditory spatial attention detection
Cunhang Fan, Jun Xue 0001, Jianhua Tao 0001, Jiangyan Yi, Zhao Lv, Xiaopei Wu
Neural Networks4
2024 Dynamic Ensemble Teacher-Student Distillation Framework for Light-Weight Fake Audio Detection
abstract
In recent years, fake audio detection (FAD) has made great progress, and lightweight is important to achieve fast and reliable audio authenticity verification on resource-limited devices. However, most of the researchers ignore lightweight when improving the performance of FAD. To develop the application of FAD for small-end devices, this paper proposes a novel light-weight network named Light-ECA2Net. Given that networks with different depths have different abilities in capturing fake speech artifacts, this paper proposes a dynamic ensemble teacher-student distillation framework to fully transfer distillation knowledge. The dynamic ensemble distillation is divided into two aspects. First, we adopt one-to-one feature mapping to perceive the multidimensional feature knowledge and dynamically adjust every dimension feature weight by using ground truth labels, which can enable students to receive feature knowledge efficiently. Secondly, different network layers also have their strengths of predicting, further dynamically predicting weight can improve the learning ability of the student. Experimental results on the ASVspoof 2019 LA and PA datasets show that compared to the baseline, our system further improves performance by reducing the model complexity by 45%.
Jun Xue 0001, Cunhang Fan, Jiangyan Yi, Jian Zhou 0006, Zhao Lv
IEEE Signal Process. Lett.1
2023 Learning From Yourself: A Self-Distillation Method For Fake Speech Detection
abstract
In this paper, we propose a novel self-distillation method for fake speech detection (FSD), which can significantly improve the performance of FSD without increasing the model complexity. For FSD, some fine-grained information is very important, such as spectrogram defects, mute segments, and so on, which are often perceived by shallow networks. However, shallow networks have much noise, which can not capture this very well. To address this problem, we propose using the deepest network instruct shallow network for enhancing shallow networks. Specifically, the networks of FSD are divided into several segments, the deepest network being used as the teacher model, and all shallow networks become multiple student models by adding classifiers. Meanwhile, the distillation path between the deepest network feature and shallow network features is used to reduce the feature difference. A series of experimental results on the ASVspoof 2019 LA and PA datasets show the effectiveness of the proposed method, with significant improvements compared to the baseline.
Jun Xue 0001, Cunhang Fan, Jiangyan Yi, Chenglong Wang 0001, Zhengqi Wen, Dan Zhang 0014, Zhao Lv
ICASSP1
2023 Subband fusion of complex spectrogram for fake speech detection
Cunhang Fan, Jun Xue 0001, Shunbo Dong, Mingming Ding, Jiangyan Yi, Jinpeng Li 0002, Zhao Lv
Speech Commun.2