Mingru Yang

dblp:240/5468 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2025
0009-0003-1656-7773ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Deep Enhancement Spotting Network for Low-complexity Keyword Spotting in Noisy Environments
abstract
Keyword Spotting (KWS) is crucial for hands-free voice-activated systems, requiring a balance between accuracy and complexity, especially in noisy environments. While Speech Enhancement (SE) can improve KWS accuracy, existing methods often lack the ability to effectively utilize the rich features produced during enhancement. In this paper, we design a low-complexity network to address the challenges of KWS in noisy environments. We integrate the tasks of both SE and KWS into a unified network that learns a shared representation from both tasks. The proposed network features two blocks: a Residual Full-band and Sub-band Fusion (RFSF) block, and a Deformable Transition (DT) block. Our dual-task network surpasses existing KWS models in accuracy with low complexity, making it suitable for deployment on edge devices.
Qianhua He, Yanxiong Li, Zunxian Liu, Mingru Yang, Jinxin Huang
ICASSP5
2025 An Efficient Sample Utilization Method for Deep Learning Based on Class Uncertainty
abstract
Deep learning has achieved success across many domains when sufficient training samples are available. However, the commonly used mini-batch stochastic gradient descent (SGD) training paradigm treats each sample equally, resulting in massive computational waste on samples that are easily identifiable. In contrast, low-quality samples, such as those with erroneous labels, can negatively impact the training process. To address these issues, we propose a training sample utilization method based on sample uncertainty. Once the model has acquired preliminary decision-making abilities, the class uncertainty for each sample can be evaluated within a training epoch. Subsequently, the samples are probabilistically selected based on their uncertainty for the next epoch. Experiments conducted on the GSC v2 and CIFAR-10 datasets demonstrate that the proposed method can reduce training time by over 32% and 58%, respectively, with only a loss of 1% performance. Additionally, the method has the capability to mitigate the adverse effects of samples with erroneous labels.
Jinxin Huang, Qianhua He, Jiezhi Xu, Sam Kwong, Mingru Yang
ICASSP5
2025 Cross-Domain Few-Shot Open-Set Keyword Spotting Using Keyword Adaptation and Prototype Reprojection
abstract
Personalized keyword spotting (KWS) with few enrollment utterances remains an important problem over years. KWS remains a challenging task due to the following factors, including the scarcity of enrollment samples, speech variation in the open-set scenarios, and distributional gap between source and target domains. In this paper, we formulate a KWS task of Cross-Domain Few-Shot Open-Set (CD-FSOS) and propose a dedicated framework Adapt-KWS to bridge the distribution gap between the source domain and target open-set domain with quite limited enrollment data. The proposed Adapt-KWS consists of a set of Custom-Keyword Adapters (CKAs) and a Prototype Reprojection Module (PRM). CKAs enable the efficient adaptation to new target tasks with limited training samples, aiming to improve cross-domain generalization. PRM reprojects the support prototypes into the query embedding space to enhance their alignment, mitigating the potential covariate shift between open-set queries and enrollments. Experimental results demonstrate the effectiveness of our framework and proposed modules on multiple datasets. Code will be available at: https://github.com/Raynaming/CD-FSOS-KWS.
Mingru Yang, Qianhua He, Jinxin Huang, Zunxian Liu, Yanxiong Li
ICASSP1
2025 Generalizable Audio Deepfake Detection via Hierarchical Structure Learning and Feature Whitening in Poincaré sphere
Mingru Yang, Yanmei Gu, Qianhua He, Yanxiong Li, Peirong Zhang 0001, Huijia Zhu, Weiqiang Wang 0002
INTERSPEECH1
2025 Generalizable Audio Deepfake Detection via Risk-Aware Style Alignment and Structural Empirical Risk Minimization
abstract
With the rapid advancement of AIGC technologies, audio deepfakes have become increasingly realistic, posing serious threats to information security and biometric authentication. Therefore, audio deepfake detection (ADD) has emerged as a critical and fast-evolving research area, particularly requiring superior generalization in out-of-domain scenarios. However, existing ADD methods suffer from constrained generalization and limited access to target data. To address these challenges, we propose Risk-Aware Style Alignment (RASA), a novel generalizable ADD framework that projects the style of any input feature into a shared style space through similarity-based projection. This alignment reduces both inter-domain and intra-source discrepancies without requiring target data during training. In addition, we adopt Structural Empirical Risk Minimization (SERM) in the Poincaré ball model to capture the hierarchical structure of the data and further minimize source risk. By jointly optimizing RASA and SERM, the proposed method effectively tightens the theoretical upper bound of target risk across three key dimensions: source risk, inter-domain divergence, and intra-source discrepancy. Extensive experiments demonstrate that our approach achieves superior generalization and outperforms existing state-of-the-art methods.
Mingru Yang, Yanmei Gu, Qianhua He, Peirong Zhang 0001, Haolin He, Huijia Zhu, Weiqiang Wang 0002
ACM Multimedia1
2025 Noise-robust feature extraction for keyword spotting based on supervised adversarial domain adaptation training strategies
Qianhua He, Zunxian Liu, Mingru Yang, Wenwu Wang 0001
Speech Commun.4