EDBT 2026 Demo / reviewers in the wild / expert
Yonghe Wang
dblp:211/9454
· DBLP profile ↗
20ranked-venue papers
5as first author
15since 2021 · last 2025
0000-0003-1647-1539ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multilingual Parameter-Sharing Adapters: A Method for Optimizing Low-Resource Neural Machine TranslationabstractAdapter-based Multilingual Neural Machine Translation (MNMT) has become a significant approach in low-resource language translation by mitigating data imbalances between high-resource and low-resource language pairs and reducing training costs. However, existing adapter-based methods lack generalization in cross-lingual settings, particularly under low-resource conditions, where their scalability is limited. Additionally, current methods often introduce independent adapter modules for each language, leading to a linear increase in model parameters with the number of languages. To address these challenges, we propose a multilingual parameter-sharing adapter approach. Moreover, we introduce a neural architecture search (NAS)-based strategy to improve translation performance. Experimental results demonstrate that the multilingual parameter-sharing adapter exhibits competitive performance on both low-resource and high-resource datasets. The multilingual parameter-sharing adapter method has only 400K trainable parameters, which is 20× lower than the parameters of the traditional adapter method. Yonghe Wang, Xiangdong Su, Feilong Bao |
ICASSP | 3 |
| 2025 | Zero-Shot Speech Recognition from Text-Only Data through Synthesized Spectrogram Refinement Using Style Truncation and Contextual Alignment LossabstractUtilizing pseudo speech-label pairs synthesized via Text-to-Speech (TTS) systems as supplementary training data for automatic speech recognition (ASR) has shown significant benefits. However, the mismatch between synthesized and real speech make them unsuitable for direct use in zero-shot speech recognition tasks. In this paper, we propose a generative-adversarial model with a style truncation strategy that enhances mel-spectrograms by projecting style vectors into high-density regions. We also introduce a Contextual alignment loss function to align synthesized and real mel-spectrograms by computing global distances between high-dimensional features. Experimental results on zero-shot ASR tasks using the SLURP and LibriSpeech datasets show that our method achieves WER reductions of 4.3%, 11.1% (clean), and 6.9% (other) compared to the baseline mel-spectrograms enhanced model. Yonghe Wang, Zhenjie Gao, Feilong Bao |
MMAsia | 2 |
| 2025 | Mongolian Speech Recognition Based on Semi-supervised Learning and Syllable Subword Modeling Units
Yonghe Wang, Zhenjie Gao, Feilong Bao |
NLPCC (4) | 2 |
| 2025 | Exploring Representation-Efficient Transfer Learning Approaches for Speech Recognition and Translation Using Pre-trained Speech Models
Yonghe Wang, Feilong Bao |
NLPCC (1) | 2 |
| 2025 | Domain disentanglement and fusion based on hyperbolic neural networks for zero-shot sketch-based image retrieval
Xiangdong Su, Yonghe Wang, Feilong Bao, Guanglai Gao |
Inf. Process. Manag. | 4 |
| 2024 | Efficient Speech-to-Text Translation: Progressive Pruning for Accelerated Speech Pre-trained ModelabstractRecently, speech pre-trained models based on the Transformer architecture have become very popular for speech-to-text translation tasks. However, computing representation outputs of speech pre-trained models is highly time-consuming, primarily due to the length of speech sequences far exceeding the corresponding texts, leading to quadratic computation costs for the self-attention module. To address this issue, we propose a novel pruning method that progressively reduces representation sequence length layer by layer. We leverage attention scores to calculate importance scores for all tokens. Additionally, we introduce both fixed and scheduled pruning rate strategies to determine which tokens should be retained. Experiments demonstrate that our approach reduces the output tokens of the speech pre-trained model by 55%, with only a 0.7% performance decrease, and improves in practice encoding speed up to 1.76 ×. Our method also is an orthogonal and complementary direction to efficient speech pre-trained models. We release our code at https://github.com/myaxxxxx/pruning. Yonghe Wang, Xiangdong Su, Feilong Bao |
ICME | 2 |
| 2024 | Improving End-to-End Speech Recognition Through Conditional Cross-Modal Knowledge Distillation with Language ModelabstractRecently, cross-modal knowledge distillation methods for end-to-end automatic speech recognition (E2E-ASR) model training pointed out the potential help of text data for improving recognition performance. However, conventional optimization strategies can mislead student models to produce suboptimal performance due to erroneous predictions generated by teacher models. This paper addresses the issue by proposing a conditional cross-modal knowledge distillation strategy, a novel technique for selectively incorporating contextual linguistic information from language model into the E2E-ASR model for improving the recognition performance. We introduce a conditional selector to dynamically adjust the knowledge source of the student model to avoid knowledge distillation from erroneous predictions generated by teacher model. In pre-trained language model fine-tuning, we perform an analysis of the impact of unsupervised text data of varying scales on the quality of soft labels and the recognition performance of the E2E-ASR model. Our proposed method simultaneously improve two different non-autoregressive decoding approaches. Experiments on the Chinese speech datasets AISHELL-1 and AISHELL-2 show competitive performance. Yonghe Wang, Feilong Bao, Zhenjie Gao, Guanglai Gao |
IJCNN | 2 |
| 2024 | Parameter-Efficient Adapter Based on Pre-trained Models for Speech Translation
Yonghe Wang, Feilong Bao |
INTERSPEECH | 2 |
| 2024 | Knowledge-Preserving Pluggable Modules for Multilingual Speech Translation Tasks
Yonghe Wang, Feilong Bao |
INTERSPEECH | 2 |
| 2024 | Sign Value Constraint Decomposition for Efficient 1-Bit Quantization of Speech Translation Tasks
Yonghe Wang, Feilong Bao |
INTERSPEECH | 2 |
| 2023 | Noise-Separated Adaptive Feature Distillation for Robust Speech RecognitionabstractThis letter makes an improvement on feature-based knowledge distillation for robust speech recognition. The use of distillation techniques in speech recognition has been demonstrated to improve the robustness of the system. In this letter, we propose a noise-separated adaptive feature distillation method, including an adaptive distillation position selection strategy and a noise separation mechanism, assuming that there is a common network structure between the student and teacher. The proposed method has two improvements. First, distillation positions can be adaptively selected in each iteration by comparing loss values computed on intermediate representations of the student and the teacher, increasing the flexibility of knowledge transfer during distillation. Second, a noise separation module is proposed to constrain noise information elimination by explicitly separating the speech information and the noise information in noisy speech, which reduces the interference of noise information during distillation. Therefore, a better recognition performance is demonstrated with the proposed method compared to the standardized feature-based knowledge distillation method. Honglin Qu, Xiangdong Su, Yonghe Wang, Guanglai Gao |
IEEE Signal Process. Lett. | 3 |
| 2023 | A Comparative Study on Selecting Acoustic Modeling Units for WFST-based Mongolian Speech RecognitionabstractTraditional weighted finite-state transducer– (WFST) based Mongolian automatic speech recognition (ASR) systems use phonemes as pronunciation lexicon modeling units. However, Mongolian is an agglutinative, low-resource language, and building an ASR system based on the phoneme pronunciation lexicon remains a challenge for various reasons. First, the phoneme pronunciation lexicon manually constructed by Mongolian linguists is finite, which is usually used to build a grapheme-to-phoneme conversion (G2P) model to frequently expand new words. However, the data sparsity decreases the robustness of the G2P model and affects the performance of the final ASR system. Second, homophones and polysyllabic words are common in Mongolian, which has a certain impact on the construction of the Mongolian acoustic model. To address these problems, in this work, we first propose a grapheme-to-phoneme alignment model to obtain the mapping relationship between phonemes and subword units. Then, we construct an acoustic subword segmentation set to segment words directly instead of using the traditional G2P method to predict phoneme sequences to expand the pronunciation lexicon. Further, by analyzing the Mongolian encoding form, we also propose an acoustic subword modeling units construction method that removes control characters. Finally, we investigate various acoustic subword modeling units for pronunciation lexicon construction for the Mongolian ASR system. Experiments on a Mongolian dataset with 325 hours of training show that the pronunciation lexicon based on the acoustic subword modeling unit can effectively construct the WFST-based Mongolian ASR system. Further, removing the control characters when building the acoustic subword modeling unit can further improve the ASR system performance. Yonghe Wang, Feilong Bao, Guanglai Gao |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2022 | Alignment-Learning Based Single-Step Decoding for Accurate and Fast Non-Autoregressive Speech RecognitionabstractNon-autoregressive transformer (NAT) based speech recognition models have gained more and more attention since they perform faster inference speed compared with autoregressive counterparts, especially when the single-step decoding is applied. However, the single-step decoding process with length prediction will suffer from the decoding stability problem and limited improvement for inference speed. To address this, in this paper, we propose an alignment learning based NAT model, named AL-NAT. Our idea is inspired by the fact that the encoder CTC output and the target sequence are monotonically related. Specifically, we design an alignment cost matrix between the CTC output tokens and the target tokens and define a novel alignment loss to minimize the distance between the alignment cost matrix and the ground truth monotonic alignment path. By eliminating the length prediction mechanism, our AL-NAT model achieves remarkable improvements in recognition accuracy and decoding speed. To learn the contextual knowledge to improve the decoding accuracy, we further add lightweight language model on both the encoder and decoder side. Our proposed method achieves WERs of 2.8%/6.3% and RTF of 0.011 on Librispeech test clean/other sets with a lightweight 3-gram LM, and a CER of 5.3% and RTF of 0.005 on Aishell1 without LM, respectively. Yonghe Wang, Rui Liu 0008, Feilong Bao, Hui Zhang 0031, Guanglai Gao |
ICASSP | 1 |
| 2021 | Joint Alignment Learning-Attention Based Model for Grapheme-to-Phoneme ConversionabstractSequence-to-sequence attention-based models for grapheme-to-phoneme (G2P) conversion have gained significant interests. The attention-based encoder-decoder framework learns the mapping of input to output tokens by selectively focusing on relevant information, and has been shown well performance. However, the attention mechanism can result in non-monotonic alignments, resulting in poor G2P conversion performance. In this paper, we present a novel approach to optimize the G2P conversion model directly alignment grapheme-phoneme sequence by using alignment learning (AL) as the loss function. Besides, we propose a multi-task learning method that uses a joint alignment learning model and attention model to predict the proper alignments and thus improve the accuracy of G2P conversion. Evaluations on Mongolian and CMUDict tasks show that alignment learning as the loss function can effectively train G2P conversion model. Further, our multi-task method can significantly outperform both the alignment learning-based model and attention-based model. Yonghe Wang, Feilong Bao, Hui Zhang 0031, Guanglai Gao |
ICASSP | 1 |
| 2021 | Soft-BAC: Soft Bidirectional Alignment Cost for End-to-End Automatic Speech Recognition
Yonghe Wang, Hui Zhang 0031, Feilong Bao, Guanglai Gao |
PRICAI (2) | 1 |
| 2019 | Research on Khalkha Dialect Mongolian Speech Recognition Acoustic Model Based on Weight Transfer
Linyan Shi, Feilong Bao, Yonghe Wang, Guanglai Gao |
NLPCC (2) | 3 |
| 2018 | A LSTM Approach with Sub-Word Embeddings for Mongolian Phrase Break PredictionabstractIn this paper, we first utilize the word embedding that focuses on sub-word units to the Mongolian Phrase Break (PB) prediction task by using Long-Short-Term-Memory (LSTM) model. Mongolian is an agglutinative language. Each root can be followed by several suffixes to form probably millions of words, but the existing Mongolian corpus is not enough to build a robust entire word embedding, thus it suffers a serious data sparse problem and brings a great difficulty for Mongolian PB prediction. To solve this problem, we look at sub-word units in Mongolian word, and encode their information to a meaningful representation, then fed it to LSTM to decode the best corresponding PB label. Experimental results show that the proposed model significantly outperforms traditional CRF model using manually features and obtains 7.49% F-Measure gain. Rui Liu 0008, Feilong Bao, Guanglai Gao, Hui Zhang 0031, Yonghe Wang |
COLING | 5 |
| 2018 | Improving Mongolian Phrase Break Prediction by Using Syllable and Morphological Embeddings with BiLSTM Model
Rui Liu 0008, Feilong Bao, Guanglai Gao, Hui Zhang 0031, Yonghe Wang |
INTERSPEECH | 5 |
| 2018 | Phonologically Aware BiLSTM Model for Mongolian Phrase Break Prediction with Attention Mechanism
Rui Liu 0008, Feilong Bao, Guanglai Gao, Hui Zhang 0031, Yonghe Wang |
PRICAI (1) | 5 |
| 2017 | Research on Mongolian Speech Recognition Based on FSMN
Yonghe Wang, Feilong Bao, Guanglai Gao |
NLPCC | 1 |