Yingying Gao

dblp:59/299 · DBLP profile ↗
← Back
31ranked-venue papers
8as first author
28since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 first-author · 21 since 2021Artificial intelligence and machine learning · 18 · 5 first-author · 16 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 CoCo-MILP: Inter-Variable Contrastive and Intra-Constraint Competitive MILP Solution Prediction
abstract
Mixed-Integer Linear Programming (MILP) is a cornerstone of combinatorial optimization, yet solving large-scale instances remains a significant computational challenge. Recently, Graph Neural Networks (GNNs) have shown promise in accelerating MILP solvers by predicting high-quality solutions. However, we identify that existing methods misalign with the intrinsic structure of MILP problems at two levels. At the leaning objective level, the Binary Cross-Entropy (BCE) loss treats variables independently, neglecting their relative priority and yielding plausible logits. At the model architecture level, standard GNN message passing inherently smooths the representations across variables, msking the natural competitive relationships within constraints. To address these challenges, we propose CoCo-MILP, which explicitly models inter-variable Contrast and intra-constraint Competition for advanced MILP solution prediction. At the objective level, CoCo-MILP introduces the Inter-Variable Contrastive Loss (VCL), which explicitly maximizes the embedding margin between variables assigned one versus zero. At the architectural level, we design an Intra-Constraint Competitive GNN layer that, instead of homogenizing features, learns to differentiate representations of competing variables within a constraint, capturing their exclusionary nature. Experimental results on standard benchmarks demonstrate that CoCo-MILP significantly outperforms existing learning-based approaches, reducing the solution gap by up to 68.12% compared to traditional solvers.
Tianle Pu, Yingying Gao, Zijie Geng, Haoyang Liu 0002, Chao Chen 0026, Changjun Fan
AAAI3
2026 DisCo_Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec
abstract
Tao Li, Wenshuo Ge, Zhichao Wang, Zihao Cui, Yong Ma, Yingying Gao, Chao Deng, Shilei Zhang, Junlan Feng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Wenshuo Ge, Zihao Cui, Yingying Gao, Chao Deng 0002, Shilei Zhang, Junlan Feng
ACL (1)6
2026 Navigating maritime emergencies with large models: A lifecycle-oriented review
abstract
As artificial intelligence advances toward cognitive capabilities, Large Model (LM) technologies have emerged as a novel paradigm for resolving the complexities of multi-source heterogeneous data in Maritime Emergency Management (MEM). This study presents a comprehensive survey of LMs throughout the MEM lifecycle, analyzing technical trajectories and application prospects. By characterizing the unstructured and sparse nature of maritime data, we elucidate the mechanisms of Large Language, Vision, and Multimodal Models, assessing their adaptability to maritime scenarios. The paper details specific applications across three critical stages: pre-event risk prevention and preparedness, during-event emergency response and decision support, and post-event accident investigation and analysis. Beyond application mapping, we critically interrogate bottlenecks in engineering deployment, specifically the “data silo” effect, the challenges of long-tail distribution, the risks associated with model hallucinations, and the constraints of shipboard computational resources. The review concludes by outlining future research frontiers, arguing that the development of domain-specific foundation models, deep semantic multimodal fusion, and the advancement of trustworthy and explainable AI (XAI) are pivotal for the profound integration of LMs into maritime safety systems.
Zhiwei Yang 0002, Yingying Gao, Ke-Wei Yang 0001, Jiang Jiang 0001
Eng. Appl. Artif. Intell.4
2026 NeuPath: A hybrid learning-based optimization approach for emergency search path planning
Yingying Gao, Tianle Pu, Zhiwei Yang 0002, Ke-Wei Yang 0001, Changjun Fan
Inf. Process. Manag.1
2025 DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles
abstract
Human speech exhibits rich and flexible prosodic variations. To address the one-to-many mapping problem from text to prosody in a reasonable and flexible manner, we propose DiffStyleTTS, a multi-speaker acoustic model based on a conditional diffusion module and an improved classifier-free guidance, which hierarchically models speech prosodic features, and controls different prosodic styles to guide prosody prediction. Experiments show that our method outperforms all baselines in naturalness and achieves superior synthesis speed compared to three diffusion-based baselines. Additionally, by adjusting the guiding scale, DiffStyleTTS effectively controls the guidance intensity of the synthetic prosody.
Zhaoci Liu, Yajun Hu, Yingying Gao, Shilei Zhang, Zhen-Hua Ling
COLING4
2025 Energy-based Model Guided Self-Supervised Learning for Speaker Verification
abstract
Self-supervised learning (SSL) has significantly advanced speaker verification, especially in scenarios with limited labeled data. This paper introduces Energy-based Confidence-Aware Distillation (EBCA-DINO), an SSL enhancement for speaker verification that integrates Energy-Based Models (EBMs) into the DINO (Distillation with No Labels) framework. EBMs use energy scores to assess data complexity and uncertainty, guiding label-free self-distillation. The adaptive temperature scaling tailors the learning process to data characteristics, allowing the teacher model to dynamically adjust the student model’s focus based on sample difficulty. This energy-aware distillation optimizes speaker verification performance. Experimental results demonstrate that EBCA-DINO improves speaker verification with relative performance gains of 4.3%, 4.9%, and 8.7% on the Vox1-O, E, and H test trials, respectively.
Yaqian Hao, Chenguang Hu, Chong Bian, Junlan Feng, Yingying Gao, Shilei Zhang
ICASSP5
2025 Codec-ASV: Exploring Neural Audio Codec For Speaker Representation Learning
abstract
Discrete speech representations have gained significant success in a variety of speech-related tasks. Among these, Neural Audio Codec (NAC), which serves as a compressed form of audio signals, have proven effective in speech AIGC applications. Moreover, we believe that the speaker information can be largely preserved in the compression process since the reconstructed voice is almost the same in human listening. In this paper, we explore various training strategies and codec types for NAC-based speaker representation learning. Using ECAPA-TDNN as the model backbone, our approach achieves state-of-the-art performance with a 2.08% EER in NAC-based speaker verification scenarios. To better retain speaker information in early, more compressed layers, we introduce mask-layer augmentation and embedding fusion techniques during the training process. Experimental results show the effectiveness of our methods, particularly when inferring with limited codec layers.
Yuke Lin, Fulin Zhang, Yingying Gao, Shilei Zhang, Ming Li 0026
ICASSP3
2025 Efficient Extreme Large-Scale Speaker Verification: Dynamic Active Sub Fully-Connected Layers for Faster Training and Memory Optimization
abstract
Using larger scale datasets in the training stage of speaker verification model usually leads to better performance. However, when the speaker number of the training dataset becomes extreme large (e.g., more than 1 million), the training speed and GPU memory demand will become bottlenecks which are mainly brought by the extreme large dimension of last fully-connected(FC) layer’s weight matrix. We propose dynamic active sub FC layers (DAS-FC) to tackle this problem. Firstly, all speakers are dynamically divided into speaker groups by clustering rows of last FC layer’s weight matrix. Then, sub FC layers are generated according to speaker groups for model training. We also introduce Mini-Batch K-means and speaker based dataloader to further reduce time and resource costing. Experiments on an extreme large dataset with 1,068,237 speakers show that compared to traditional FC layer, DAS-FC can save up to 87% training time and save 56% GPU memory occupancy with only a 4.2% drop in model performance.
Fulin Zhang, Chenguang Hu, Yingying Gao, Shilei Zhang, Junlan Feng
ICASSP4
2025 Privacy-Preserving Speaker Verification via End-to-End Secure Representation Learning
Chenguang Hu, Yaqian Hao, Fulin Zhang, Xiaoxue Luo, Yingying Gao, Chao Deng 0002, Shilei Zhang, Junlan Feng
INTERSPEECH6
2025 DCAPNet: A Contrast-Enhanced and Multi-scale Feature Fusion Network for Infrared Small Target Detection
Yingying Gao, Maoyong Li, Xuedong Guo, Mingli Dong, Lianqing Zhu
PRCV (18)1
2025 GDTFusion: Gated Dual-Branch Attention Transformer Network for Infrared and Visible Image Fusion
Xuedong Guo, Maoyong Li, Yingying Gao, Mingli Dong, Lianqing Zhu
PRCV (18)3
2025 Tri-guided Hybrid Attention Network with Adaptive Top-K Channel and Body-Edge Spatial Modeling for Infrared Small Target Detection
Maoyong Li, Yingying Gao, Xuedong Guo, Mingli Dong, Lianqing Zhu
PRCV (18)2
2025 Personalized federated learning with adaptive optimization of local model
Yingying Gao, Yajie Song, Haobin Shi
Knowl. Inf. Syst.1
2025 Personalized federated learning with adaptive aggregation within clusters
Shiyuan Ding, Haobin Shi, Yingying Gao
Pattern Anal. Appl.4
2024 MFSN: Multi-perspective Fusion Search Network For Pre-training Knowledge in Speech Emotion Recognition
Haiyang Sun 0004, Fulin Zhang, Yingying Gao, Shilei Zhang, Zheng Lian 0004, Junlan Feng
INTERSPEECH3
2024 GenDistiller: Distilling Pre-trained Language Models based on an Autoregressive Generative Model
Yingying Gao, Shilei Zhang, Chao Deng 0002, Junlan Feng
INTERSPEECH1
2024 Exploring Energy-Based Models for Out-of-Distribution Detection in Dialect Identification
Yaqian Hao, Chenguang Hu, Yingying Gao, Shilei Zhang, Junlan Feng
INTERSPEECH3
2024 On Calibration of Speech Classification Models: Insights from Energy-Based Model Investigations
Yaqian Hao, Chenguang Hu, Yingying Gao, Shilei Zhang, Junlan Feng
INTERSPEECH3
2024 VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark
Yuke Lin, Ming Cheng 0005, Fulin Zhang, Yingying Gao, Shilei Zhang, Ming Li 0026
INTERSPEECH4
2024 CEC: A Noisy Label Detection Method for Speaker Recognition
Yingying Gao, Yaqian Hao, Chenguang Hu, Fulin Zhang, Junlan Feng, Shilei Zhang
INTERSPEECH2
2023 Semi-Supervised Speech Enhancement Based On Speech Purity
abstract
We tend to assume most available speech corpora we use are either completely clean or completely noised. However, the reality is most of them are a mix of both. In this paper, we propose a semi-supervised speech enhancement framework to enhance such typical speech datasets. This framework includes an estimator to measure the speech purity. Utterances with high speech purity are considered clean, otherwise noised. For clean speech utterances, we follow the supervised learning mechanism to train a deep learning speech enhancement model. For noised speech, we update the model in an unsupervised manner. Hence, we design our training loss as a combination of the supervised loss and unsupervised loss. We refer to this framework as SemiEnhance. Experimental results show that SemiEnhance substantially improves the speech quality, and achieves new state-of-the-art results on benchmark datasets: 2022 DNS Challenge and NoiseX-92.
Zihao Cui, Shilei Zhang, Yingying Gao, Chao Deng 0002, Junlan Feng
ICASSP4
2023 VE-KWS: Visual Modality Enhanced End-to-End Keyword Spotting
abstract
The performance of the keyword spotting (KWS) system based on audio modality, commonly measured in false alarms and false rejects, degrades significantly under the far field and noisy conditions. Therefore, audio-visual keyword spotting, which leverages complementary relationships over multiple modalities, has recently gained much attention. However, current studies mainly focus on combining the exclusively learned representations of different modalities, instead of exploring the modal relationships during each respective modeling. In this paper, we propose a novel visual modality enhanced end-to-end KWS framework (VE-KWS), which fuses audio and visual modalities from two aspects. The first one is utilizing the speaker location information obtained from the lip region in videos to assist the training of multi-channel audio beamformer. By involving the beamformer as an audio enhancement module, the acoustic distortions, caused by the far field or noisy environments, could be significantly suppressed. The other one is conducting cross-attention between different modalities to capture the inter-modal relationships and help the representation learning of each modality. Experiments on the MSIP challenge corpus show that our proposed model achieves a 2.79% false rejection rate and a 2.95% false alarm rate on the Eval set, resulting in a new SOTA performance compared with the top-ranking systems in the ICASSP2022 MISP challenge.
He Wang 0022, Yihui Fu, Lei Xie 0001, Yingying Gao, Shilei Zhang, Junlan Feng
ICASSP6
2023 Cascaded Multi-task Adaptive Learning Based on Neural Architecture Search
Yingying Gao, Shilei Zhang, Zihao Cui, Chao Deng 0002, Junlan Feng
INTERSPEECH1
2023 Harmonic Attention for Monaural Speech Enhancement
abstract
To further improve the quality of the enhanced speech, it is appealing that more profound articulatory and auditory knowledge should be introduced into the speech enhancement model. Among these, harmonics seriously affect speech timbre and play a crucial role in speech intelligibility. Especially in the frequency domain, harmonics appear as the local maximum peaks of energy, which could be expected to serve as anchors to recover the distorted speech. In this paper, an explicit modeling method, harmonic attention, is presented, patching the harmonics with the help of residual ones. In order to maintain the spectral structure of speech during the processing and to enable the network to support harmonic modeling, a harmonic attention-based progressive enhancement network (HAPNet) is applied, which gradually approaches clean speech with stacked modules of harmonic attention. In addition, to make enhanced speech more consistent with hearing, a loss function based on the loudness power compression (LC-SNR) is used, which measures both magnitude and phase values with appropriate auditory effects. The experimental visualization indicates that the harmonic attention can capture and recover the harmonics of speech. And the objective evaluations show that the presented HAPNet and LC-SNR outperform the referenced methods. Furthermore, the presented model trained on 100 hours of data achieves competitive results with the referenced models trained on 3000+ hours of data, and one trained on 500 hours of data yields the state-of-the-art performance.
Tianrui Wang, Weibin Zhu, Yingying Gao, Shilei Zhang, Junlan Feng
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Harmonic Gated Compensation Network Plus for ICASSP 2022 DNS Challenge
abstract
The harmonic structure of speech is resistant to noise, but the harmonics may still be partially masked by noise. Therefore, we previously proposed a harmonic gated compensation network (HGCN) to predict the full harmonic locations based on the unmasked harmonics and process the result of a coarse enhancement module to recover the masked harmonics. In addition, the auditory loudness loss function is used to train the network. For the DNS Challenge, we update HGCN with the following aspects, resulting in HGCN+. First, a high-band module is employed to help the model handle full-band signals. Second, cosine is used to model the harmonic structure more accurately. Then, the dual-path encoder and dual-path rnn (DPRNN) are introduced to take full advantage of the features. Finally, a gated residual linear structure replaces the gated convolution in the compensation module to increase the receptive field of frequency. The experimental results show that each updated module brings performance improvement to the model. HGCN+ also outperforms the referenced models on both wide-band and full-band test sets.
Tianrui Wang, Weibin Zhu, Yingying Gao, Junlan Feng, Shilei Zhang
ICASSP3
2022 HGCN: Harmonic Gated Compensation Network for Speech Enhancement
abstract
Mask processing in the time-frequency (T-F) domain through the neural network has been one of the mainstreams for single-channel speech enhancement. However, it is hard for most models to handle the situation when harmonics are partially masked by noise. To tackle this challenge, we propose a harmonic gated compensation network (HGCN). We design a high-resolution harmonic integral spectrum to improve the accuracy of harmonic locations prediction. Then we add voice activity detection (VAD) and voiced region detection (VRD) to the convolutional recurrent network (CRN) to filter harmonic locations. Finally, the harmonic gating mechanism is used to guide the compensation model to adjust the coarse results from CRN to obtain the refinedly enhanced results. Our experiments show HGCN achieves substantial gain over a number of advanced approaches in the community.
Tianrui Wang, Weibin Zhu, Yingying Gao, Junlan Feng, Shilei Zhang
ICASSP3
2022 Meta Auxiliary Learning for Low-resource Spoken Language Understanding
abstract
Spoken language understanding (SLU) treats automatic speech recognition (ASR) and natural language understanding (NLU) as a unified task and usually suffers from data scarcity.We exploit an ASR and NLU joint training method based on meta auxiliary learning to improve the performance of low-resource SLU task by only taking advantage of abundant manual transcriptions of speech data.One obvious advantage of such method is that it provides a flexible framework to implement a lowresource SLU training task without requiring access to any further semantic annotations.In particular, a NLU model is taken as label generation network to predict intent and slot tags from texts; a multi-task network trains ASR task and SLU task synchronously from speech; and the predictions of label generation network are delivered to the multi-task network as semantic targets.The efficiency of the proposed algorithm is demonstrated with experiments on the public CATSLU dataset, which produces more suitable ASR hypotheses for the downstream NLU task.
Yingying Gao, Junlan Feng, Chao Deng 0002, Shilei Zhang
INTERSPEECH1
2021 Boundary and Context Aware Training for CIF-Based Non-Autoregressive End-to-End ASR
abstract
Continuous integrate-and-fire (CIF) based models, which use a soft and monotonic alignment mechanism, have been well applied in non-autoregressive (NAR) speech recognition with competitive performance compared with other NAR methods. However, such an alignment learning strategy may suffer from an erroneous acoustic boundary estimation, severely hindering the convergence speed as well as the system performance. In this paper, we propose a boundary and context aware training approach for CIF based NAR models. Firstly, the connectionist temporal classification (CTC) spike information is utilized to guide the learning of acoustic boundaries in the CIF. Besides, an additional contextual decoder is introduced behind the CIF decoder, aiming to capture the linguistic dependencies within a sentence. Finally, we adopt a recently proposed Conformer architecture to improve the capacity of acoustic modeling. Experiments on the open-source Mandarin AISHELL-1 corpus show that the proposed method achieves a comparable character error rates (CERs) of 4.9% with only 1/24 latency compared with a state-of-the-art autoregressive (AR) Conformer model. Futhermore, when evaluating on an internal 7500 hours Mandarin corpus, our model still outperforms other NAR methods and even reaches the AR Conformer model on a challenging real-world noisy test set.
Fan Yu 0002, Haoneng Luo, Yuhao Liang, Zhuoyuan Yao, Lei Xie 0001, Yingying Gao, Leijing Hou, Shilei Zhang
ASRU7
2019 Code-Switching Sentence Generation by Bert and Generative Adversarial Networks
Yingying Gao, Junlan Feng, Leijing Hou
INTERSPEECH1
2018 A Model-Based Architecture for Technological Management in Defense Acquisition
abstract
The military technology planning is a key driving force for improving combat effectiveness. With regard to technological management support, the research on the architecture has been relatively weak. This paper proposed a model based technological management framework for defense acquisition. By analyzing the core decision elements in the technological management process, we define the domain-specific meta-model of the framework to capture the different needs of stakeholders. And then, six products of the technology view (TV) and a specific development process are designed to organize the concepts and relationships into models to support decision-making from different perspectives. The data consistency of models will be ensured by the meta-model. As a consequence, technical analysis activities can be performed based on the data and models in this technology architecture. This model-based approach ensures the availability, integrity, shareability, and compatibility of the architecture. Ultimately, the technology of shipborne unmanned helicopters (SUH) was explored by the ModelLink to demonstrate the applicability and effectiveness of the architecture.
Minghao Li 0006, Ke-Wei Yang 0001, Yingying Gao
SMC5
2016 Detecting affective states from text based on a multi-component emotion model
Yingying Gao, Weibin Zhu
Comput. Speech Lang.1