Minqiang Xu

dblp:86/5698 · DBLP profile ↗
← Back
18ranked-venue papers
2as first author
16since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Integrating Spectro-Temporal Cross Aggregation and Multi-Scale Dynamic Learning for Audio Deepfake Detection
abstract
Audio deepfake refers to the technology of synthesizing speech using deep learning or large model algorithms. Compared to human voice, synthetic deepfake speech exhibits artifacts at global and local levels, which can be leveraged by audio deepfake detection (ADD) to distinguish real and fake speech. In this paper, we designed the spectro-temporal cross aggregation (STCA) module and the local multi-scale dynamic convolution (LMDC) module to extract global and local artifacts for detecting forged information, respectively. The STCA module utilizes a dual-branch structure with cross-attention, extracting global temporal and frequency features through the branches and aggregating mutual artifacts via cross-attention. The LMDC module uses multi-scale dynamic convolution for grouped features, to extract local information. Experimental results on multiple test sets demonstrate the effectiveness of our method. Specifically, we achieved an EER of 1.87% on the ASVspoof2021 DF evaluation set, surpassing the current state-of-the-art system by a relative 14.6%.
Yunqi Hao, Minqiang Xu
ICASSP2
2025 Sample-to-Sample Learning and inverted bottleneck for Speaker Verification
abstract
This paper presents the Sample-to-Sample Learning (S2SL) technique and the inverted bottleneck structure for speaker verification. Traditional softmax and margin-based softmax losses primarily optimize sample-to-prototype similarity while overlooking sample-to-sample relationships. To address this limitation and minimize intra-class variance, we propose S2SL, which integrates sample-to-prototype and sample-to-sample similarities through a linear combination. Additionally, we confirm the effectiveness of the inverted bottleneck in speaker verification. Our experiments demonstrate that S2SL is complementary to existing margin-based softmax losses and the inverted bottleneck. By incorporating both S2SL and the inverted bottleneck, we achieve state-of-the-art performance across multiple benchmark datasets.
Yunqi Hao, Minqiang Xu, Jianbo Zhan, Sian Fang
IJCNN4
2025 Wav2DF-TSL: Two-stage Learning with Efficient Pre-training and Hierarchical Experts Fusion for Robust Audio Deepfake Detection
abstract
In recent years, self-supervised learning (SSL) models have made significant progress in audio deepfake detection (ADD) tasks. However, existing SSL models mainly rely on large-scale real speech for pre-training and lack the learning of spoofed samples, which leads to susceptibility to domain bias during the fine-tuning process of the ADD task. To this end, we propose a two-stage learning strategy (Wav2DF-TSL) based on pretraining and hierarchical expert fusion for robust audio deepfake detection. In the pre-training stage, we use adapters to efficiently learn artifacts from 3000 hours of unlabelled spoofed speech, improving the adaptability of front-end features while mitigating catastrophic forgetting. In the fine-tuning stage, we propose the hierarchical adaptive mixture of experts (HA-MoE) method to dynamically fuse multi-level spoofing cues through multi-expert collaboration with gated routing. Experimental results show that the proposed method significantly outperforms the baseline system on all four benchmark datasets, especially on the cross-domain In-the-wild dataset, achieving a 27.5% relative improvement in equal error rate (EER), outperforming the existing state-of-the-art systems.
Yunqi Hao, Minqiang Xu, Jianbo Zhan, Sian Fang
IJCNN3
2025 Denoising GER: A Noise-Robust Generative Error Correction with LLM for Speech Recognition
abstract
In recent years, large language models (LLM) have made significant progress in the task of generation error correction (GER) for automatic speech recognition (ASR) post-processing. However, in complex noisy environments, they still face challenges such as poor adaptability and low information utilization, resulting in limited effectiveness of GER. To address these issues, this paper proposes a noise-robust multi-modal GER framework (Denoising GER). The framework enhances the model’s adaptability to different noisy scenarios through a noise-adaptive acoustic encoder and optimizes the integration of multi-modal information via a heterogeneous feature compensation dynamic fusion (HFCDF) mechanism, improving the LLM’s utilization of multi-modal information. Additionally, reinforcement learning (RL) training strategies are introduced to enhance the model’s predictive capabilities. Experimental results demonstrate that Denoising GER significantly improves accuracy and robustness in noisy environments and exhibits good generalization abilities in unseen noise scenarios.
Minqiang Xu, Sian Fang, Lin Liu 0017
IJCNN2
2025 Enhancing Self-Supervised Speaker Verification Using Similarity-Connected Graphs and GCN
abstract
With the continuous development of speech recognition technology, speaker verification (SV) has become an important method for identity authentication. Traditional SV methods rely on handcrafted feature extraction, while the introduction of deep learning has significantly improved system performance. However, the scarcity of labeled data still limits the widespread application of deep learning methods in SV. Self-supervised learning, by mining the latent information in massive unlabeled data, enhances the model’s generalization ability and has become a key technology to address this issue.DINO is an efficient self-supervised learning method that generates pseudo-labels from unlabeled speech data through clustering algorithms, providing support for subsequent training. However, the clustering process may produce noisy pseudolabels, which can reduce the overall recognition performance of the system and restrict further improvement of the model’s performance.To address this issue, this paper proposes an improved clustering framework based on similarity connection graphs and Graph Convolutional Networks (GCN). By leveraging GCN’s strength in modeling structured data and incorporating the relational information between nodes in the similarity connection graph, the clustering process is optimized, improving the accuracy of pseudo-labels and thereby enhancing the robustness and performance of the self-supervised speaker verification system. Experimental results show that this method can significantly improve system performance and provide a new approach for self-supervised speaker verification.
Zhaorui Sun, Minqiang Xu, Jianbo Zhan, Sian Fang
IJCNN4
2025 SAFE-AKT: Kazakh Image-Text Retrieval via Semantic-Agnostic Feature Enhancement and Adaptive Knowledge Transfer
abstract
Kazakh image-text retrieval is a challenging task with no dedicated research to date. Although existing multilingual vision-language pretraining models provide limited support for aligning Kazakh text with images, their performance remains poor due to the scarcity of annotated Kazakh resources and the complex expression patterns arising from its agglutinative linguistic nature, which hinder accurate modeling of text-image alignment. To address these challenges, we propose a new Kazakh image-text retrieval framework that integrates Semantic-Agnostic Feature Enhancement and Adaptive Knowledge Transfer (SAFE-AKT). SAFE employs an adversarial training strategy to generate semantic-agnostic features from Kazakh texts (e.g., expression patterns) and incorporates them into the kazakh text encoding process, enhancing the model’s robustness to diverse Kazakh expressions. AKT estimates sample-level alignment confidence by computing the entropy of the teacher distribution and weights the KL divergence loss accordingly, enabling effective transfer of high-quality English image-text alignment knowledge to the Kazakh representation space while mitigating overfitting to low-confidence samples, thereby improving Kazakh image-text alignment. Meanwhile, to address the scarcity of Kazakh data, we construct the first Kazakh image-text dataset, Flickr30k-kaz, based on machine translation and manual refinement. Experimental results on the Flickr30k-kaz and the Multi30K benchmark demonstrate that our method significantly outperforms existing state-of-the-art approaches and achieves new best performance.
Zhiqun Cao, Changle Yin, Minqiang Xu, Lumei Zhou
MMAsia4
2025 HyperSIGMA: Hyperspectral Intelligence Comprehension Foundation Model
abstract
Accurate hyperspectral image (HSI) interpretation is critical for providing valuable insights into various earth observation-related applications such as urban planning, precision agriculture, and environmental monitoring. However, existing HSI processing methods are predominantly task-specific and scene-dependent, which severely limits their ability to transfer knowledge across tasks and scenes, thereby reducing the practicality in real-world applications. To address these challenges, we present HyperSIGMA, a vision transformer-based foundation model that unifies HSI interpretation across tasks and scenes, scalable to over one billion parameters. To overcome the spectral and spatial redundancy inherent in HSIs, we introduce a novel sparse sampling attention (SSA) mechanism, which effectively promotes the learning of diverse contextual features and serves as the basic block of HyperSIGMA. HyperSIGMA integrates spatial and spectral features using a specially designed spectral enhancement module. In addition, we construct a large-scale hyperspectral dataset, HyperGlobal-450K, for pre-training, which contains about 450 K hyperspectral images, significantly surpassing existing datasets in scale. Extensive experiments on various high-level and low-level HSI tasks demonstrate HyperSIGMA's versatility and superior representational capability compared to current state-of-the-art methods. Moreover, HyperSIGMA shows significant advantages in scalability, robustness, cross-modal transferring capability, real-world applicability, and computational efficiency.
Di Wang 0023, Meiqi Hu, Yuchun Miao, Jiaqi Yang 0005, Yichu Xu, Xiaolei Qin, Jiaqi Ma 0002, Chenxing Li, Chuan Fu, Hongruixuan Chen, Chengxi Han, Naoto Yokoya, Jing Zhang 0037, Minqiang Xu, Lefei Zhang, Chen Wu 0003, Bo Du 0001, Dacheng Tao, Liangpei Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.16
2024 Multi-View Speaker Embedding Learning for Enhanced Stability and Discriminability
abstract
Deep neural network models based on x-vector have become the most popular framework for speaker recognition, and the quality of speaker features (embeddings) is important for open-set tasks such as speaker verification and speaker diarization. Currently, the most popular loss function is based on margin penalty, however, it only considers enlarging the inter-class distance while neglecting to reduce the intra-class feature differences. Therefore, we propose a multi-view learning approach that divides the training process into two views from the speaker embedding level. The classification view focuses on distinguishing the discriminability of different speakers, while the clustering view focuses on shrinking the feature boundaries of the same speaker, making intra-class differences smaller. The combined effect of the two perspectives achieves large inter-class distance and small intra-class distances, resulting in the extraction of more discriminative and stable speaker embeddings. We test the performance of the method on both speaker verification and speaker diarization tasks, and the results demonstrate the effectiveness of our approach.
Liang He 0003, Zhihua Fang, Zuoer Chen, Minqiang Xu
ICASSP4
2024 Scene Text Recognition Via k-NN Attention-Based Decoder and Margin-Based Softmax Loss
Minqiang Xu, Liang He 0003
PRCV (7)2
2024 Multi-satellite cooperative scheduling method for large-scale tasks based on hybrid graph neural network and metaheuristic algorithm
Xiaoen Feng, Minqiang Xu
Adv. Eng. Informatics3
2024 A conflict clique mitigation method for large-scale satellite mission planning based on heterogeneous graph learning
Xiaoen Feng, Minqiang Xu
Adv. Eng. Informatics2
2023 Dynamic Fully-Connected Layer for Large-Scale Speaker Verification
Zhida Song, Liang He 0003, Baowei Zhao, Minqiang Xu, Yu Zheng 0020
INTERSPEECH4
2023 SAMRS: Scaling-up Remote Sensing Segmentation Dataset with Segment Anything Model
abstract
The success of the Segment Anything Model (SAM) demonstrates the significance of data-centric machine learning. However, due to the difficulties and high costs associated with annotating Remote Sensing (RS) images, a large amount of valuable RS data remains unlabeled, particularly at the pixel level. In this study, we leverage SAM and existing RS object detection datasets to develop an efficient pipeline for generating a large-scale RS segmentation dataset, dubbed SAMRS. SAMRS totally possesses 105,090 images and 1,668,241 instances, surpassing existing high-resolution RS segmentation datasets in size by several orders of magnitude. It provides object category, location, and instance information that can be used for semantic segmentation, instance segmentation, and object detection, either individually or in combination. We also provide a comprehensive analysis of SAMRS from various aspects. Moreover, preliminary experiments highlight the importance of conducting segmentation pre-training with SAMRS to address task discrepancies and alleviate the limitations posed by limited training data during fine-tuning. The code and dataset will be available at https://github.com/ViTAE-Transformer/SAMRS
Di Wang 0023, Jing Zhang 0037, Bo Du 0001, Minqiang Xu, Dacheng Tao, Liangpei Zhang 0001
NeurIPS4
2023 Short-range air combat maneuver decision of UAV swarm based on multi-agent Transformer introducing virtual objects
Minqiang Xu, Hutao Cui
Eng. Appl. Artif. Intell.2
2023 Evaluating the application of using biological pulse sensor in aerobics
Libin Sun, Minqiang Xu, Yilun Gao, Haiyang Kou
Wirel. Networks2
2022 Multi-Query Multi-Head Attention Pooling and Inter-Topk Penalty for Speaker Verification
abstract
This paper describes the multi-query multi-head attention (MQMHA) pooling and inter-topK penalty methods which were first proposed in our submitted system description for VoxCeleb speaker recognition challenge (VoxSRC) 2021. Most multi-head attention pooling mechanisms either attend to the whole feature through multiple heads or attend to several split parts of the whole feature. Our proposed MQMHA combines both these two mechanisms and gain more diversified information. The margin-based softmax loss functions are commonly adopted to obtain discriminative speaker representations. To further enhance the inter-class discriminability, we propose a method that adds an extra inter-topK penalty on some confused speakers. By adopting both the MQMHA and inter-topK penalty, we achieved state-of-the-art performance in all of the public VoxCeleb test sets.
Miao Zhao, Yufeng Ma, Yiwei Ding, Yu Zheng 0020, Minqiang Xu
ICASSP6
2010 Extended Hierarchical Gaussianization for scene classification
abstract
In this paper, we propose a novel image representation for scene classification. Firstly, we model multiple order statistics of image patches via Gaussian Mixture Model(GMM) in a Bayesian framework. Secondly, we combine the information of mean and covariance of the GMM and represent it as a mean-covariance supervector through a new distance metric. Experimental results demonstrate that our new representation, by just using nearest centroid classifier, has significantly outperformed all existing methods on the fifteen scene category database.
Minqiang Xu, Zhen Li 0028, Beiqian Dai, Thomas S. Huang
ICIP1
2009 GMM kernel by Taylor series for speaker verification
Minqiang Xu, Beiqian Dai, Thomas S. Huang
INTERSPEECH1