Xingmei Wang 0002

dblp:34/486-2 · DBLP profile ↗
← Back
14ranked-venue papers
9as first author
13since 2021 · last 2026
0000-0002-0281-0336ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 8 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Dynamic residual multi-stage replay policy gradient method for multi-agent cooperation and competition
Xingmei Wang 0002, Junzheng Xu, Zining Yan
Eng. Appl. Artif. Intell.1
2026 IAGM-TAN: An interactive attention graph matching network for multi-target track association
Xingmei Wang 0002, Ziyan Zeng
Neurocomputing4
2026 SpeakerMatch: Matching reliable pseudo-labels in semi-supervised and self-supervised speaker recognition with confidence distribution
abstract
For speaker recognition, pseudo-labeling has shown advantages in alleviating the scarcity of labeled data. Inspired by image classification tasks, existing methods typically adopt threshold-based strategies to identify reliable pseudo-labels. However, compared to image classification, speaker recognition requires finer-grained class discrimination for open-set identity verification, and thus often adopts margin-based losses that amplify gradients near decision boundaries. While effective under full supervision, this design increases sensitivity to noisy or sparse pseudo-labels, limiting the effectiveness of threshold-based selection. In this work, we propose SpeakerMatch , a novel distribution-based framework for semi-supervised and self-supervised speaker recognition. SpeakerMatch models the confidence distribution to distinguish reliable pseudo-labels from noisy ones globally and selects those whose confidence values and confidence prediction behaviors closely align with high-quality signals. Systematic evaluation across five settings shows that our method outperforms existing approaches, achieving a 13.7% relative improvement over the best semi-supervised speaker recognition baseline, while also delivering a lower equal error rate (EER) and reduced training costs compared to self-supervised methods.
Jinghan Liu, Xingmei Wang 0002, Jiaxiang Meng, Boquan Li 0002
Signal Process.2
2025 Int*-Match: Balancing Intra-Class Compactness and Inter-Class Discrepancy for Semi-Supervised Speaker Recognition
abstract
Open-set speaker recognition is to identify whether the voices are from the same speaker. One challenge of speaker recognition is collecting large amounts of high-quality data. Based on the promising results of image classification, one intuitively feasible solution is semi-supervised learning (SSL) which uses confidence thresholds to assign pseudo labels for unlabeled data. However, we empirically demonstrated that applying SSL methods to speaker recognition is non-trivial. These methods focus solely on inter-class discrepancy as thresholds to select pseudo labels, overlooking intra-class compactness, which is particularly important for open-set speaker recognition tasks. Motivated by this, we propose Int*-Match, a semi-supervised speaker recognition method selecting reliable pseudo labels with intra-class compactness and inter-class discrepancy for speaker recognition. In particular, we use the inter-class discrepancy of labeled data as the threshold for pseudo-label selection and adjust the threshold based on the intra-class compactness of the pseudo labels dynamically and adaptively. Our systematic experiments demonstrate the superiority of Int*-Match, presenting an outstanding Equal Error Rate (EER) of 1.00% on the VoxCeleb1 original test set, which is merely 0.06% below the performance achieved by fully supervised learning.
Xingmei Wang 0002, Jinghan Liu, Jiaxiang Meng, Boquan Li 0002
AAAI1
2025 MIMTrack: In-Context Tracking via Masked Image Modeling
abstract
Current Siamese and Transformer trackers commonly use various subtask branches like regression and classification to predict object states. Despite the demonstrated success, these subtask branches might introduce location and scale offsets due to discrepancies and misalignment in the respective predictions. To address this, we propose a novel generative tracker, MIMTrack, which defines tracking as a Masked Image Modeling (MIM) process combined with in-context learning (ICL). MIMTrack begins with building the visual prompt image, which consists of a template, a search area, and two target images associated with them. The target image transforms the bounding box into a unified RGB image space as other tracking image. All states prediction are naturally aligned by pixels generation of search target image. In light of this, we perform a MIM process within the visual prompt to reconstruct a masked search target image using the context from other parts. MIM with ICL makes use of implicit cross-relations between template and search area. A singlestream generative framework reduces the offset in the estimation. Furthermore, a latent memory module is introduced as a plugin to enhance pixel generation by leveraging various target appearances over time. The advanced performance observed on leading benchmark datasets highlights the simplicity and effectiveness of our MIMTrack framework.
Xingmei Wang 0002, Guohao Nie, Jiaxiang Meng, Zining Yan
AAAI1
2025 Adaspeaker: Learning Discriminative Speaker Representations with Gradient-Aware Adaptive Scaling
abstract
Learning discriminative representations of different speakers is a key challenge in open-set speaker recognition. To mitigate the mismatch between closed-set training and open-set testing, margin-based losses have been widely adopted to directly optimize the cosine similarity between speaker representations and proxy class vectors. While recent studies have shown that enhancing the margin for hard samples can improve representation learning, we observe three key limitations: (1) the measurement of sample hardness fails to fully capture differences in speaker representations, (2) margin-based emphasis does not significantly increase the gradient magnitude, and (3) the potential performance degradation caused by emphasizing hard samples are rarely considered. To address these issues, we propose Adaspeaker, a novel loss framework that combines an Intra-Inter sample hardness coefficient (Int2H) with a gradient-aware adaptive scaling strategy. Specifically, Int2H jointly models inter-class and intra-class hardness to estimate sample importance, which is subsequently used to adaptively scale cosine similarities for enhancing the gradient contribution of important samples. Experiments conducted on five evaluation settings show that Adaspeaker outperforms existing loss functions. Moreover, Adaspeaker can be seamlessly integrated into margin-based losses, yielding an average performance improvement of 12.6%. Code is available at https://github.com/LiuJinghan2001/Adaspeaker.
Jinghan Liu, Xingmei Wang 0002, Jiaxiang Meng
ACM Multimedia2
2025 Multi-UAV intelligent decision-making method with layer delay dual-center MAPPO for air combat
Zhengkun Ding, Xingmei Wang 0002, Chengtao Cai, Luyu Jia
Appl. Intell.2
2024 Two-stage Semi-supervised Speaker Recognition with Gated Label Learning
Xingmei Wang 0002, Jiaxiang Meng, Kong-Aik Lee, Boquan Li 0002, Jinghan Liu
IJCAI1
2024 CATNet: Cross-modal fusion for audio-visual speech recognition
Xingmei Wang 0002, Jiachen Mi, Boquan Li 0002, Yixu Zhao, Jiaxiang Meng
Pattern Recognit. Lett.1
2023 Multi-intent autonomous decision-making for air combat with deep reinforcement learning
Luyu Jia, Chengtao Cai, Xingmei Wang 0002, Zhengkun Ding, Junzheng Xu, Kejun Wu
Appl. Intell.3
2023 Hierarchical memory-guided long-term tracking with meta transformer inquiry network
Xingmei Wang 0002, Guohao Nie, Boquan Li 0002, Minyang Kang
Knowl. Based Syst.1
2023 LDGC-Net: learnable descriptor graph convolutional network for image retrieval
Xingmei Wang 0002, Jinli Wang, Minyang Kang, Ze Feng
Vis. Comput.1
2022 RACP: A network with attention corrected prototype for few-shot speaker recognition using indefinite distance metric
Xingmei Wang 0002, Jiaxiang Meng, Fuzhao Xue
Neurocomputing1
2020 A network model of speaker identification with new feature extraction methods and asymmetric BLSTM
Xingmei Wang 0002, Fuzhao Xue, Wei Wang 0059, Anhua Liu
Neurocomputing1