Binbin Du

dblp:14/6889 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
YearPublicationVenuePosition
2025 BLR-MoE: Boosted Language-Routing Mixture of Experts for Domain-Robust Multilingual E2E ASR
abstract
Recently, the Mixture of Expert (MoE) architecture, such as LR-MoE, is often used to alleviate the impact of language confusion on the multilingual ASR (MASR) task. However, it still faces language confusion issues, especially in mismatched domain scenarios. In this paper, we decouple language confusion in LR-MoE into confusion in self-attention and router. To alleviate the language confusion in self-attention, based on LR-MoE, we propose to apply attention-MoE architecture for MASR. In our new architecture, MoE is utilized not only on feedforward network (FFN) but also on self-attention. In addition, to improve the robustness of the LID-based router on language confusion, we propose expert pruning and router augmentation methods. Combining the above, we get the boosted language-routing MoE (BLR-MoE) architecture. We verify the effectiveness of the proposed BLR-MoE in a 10,000-hour MASR dataset.
Lifeng Zhou 0003, Yuke Li 0002, Binbin Du
ICASSP6
2024 Learning from Back Chunks: Acquiring More Future Knowledge for Streaming ASR Models via Self Distillation
Yuke Li 0002, Binbin Du, Haoqi Zhu, Liang Ruan
INTERSPEECH4
2024 Enhancing Unified Streaming and Non-Streaming ASR Through Curriculum Learning With Easy-To-Hard Tasks
abstract
We expect a unified ASR model to deliver high performance in both streaming and non-streaming modes. However, a core challenge is that the lack of global contextual information in streaming ASR inherently hinders its performance from matching the non-streaming counterpart. Drawing inspiration from the human learning manner from easy concepts to difficult ones, we introduce a curriculum learning framework to enhance the training of unified ASR models. This framework strategically increases task complexity in a graduated, easy-to-hard order. Specifically, we develop a structured curriculum that begins with an elementary course focused on training a non-streaming model, progresses to an intermediate course for training an initial unified ASR model, and culminates in an advanced course designed to mutual promotion between these two modes via contrastive training. Experimental results on AISHELL-1 and AISHELL-2 show that our method achieves significant improvements in two modes.
Yuke Li 0002, Lifeng Zhou 0003, Binbin Du, Haoqi Zhu
SLT4
2023 LAE-ST-MOE: Boosted Language-Aware Encoder Using Speech Translation Auxiliary Task for E2E Code-Switching ASR
abstract
Recently, to mitigate the confusion between different languages in code-switching (CS) automatic speech recognition (ASR), the conditionally factorized models, such as the language-aware encoder (LAE), explicitly disregard the contextual information between different languages. However, this information may be helpful for ASR modeling. To alleviate this issue, we propose the LAE-ST-MoE framework. It incorporates speech translation (ST) tasks into LAE and utilizes ST to learn the contextual information between different languages. It introduces a task-based mixture of expert modules, employing separate feed-forward networks for the ASR and ST tasks. Experimental results on the ASRU 2019 Mandarin-English CS challenge dataset demonstrate that, compared to the LAE-based CTC, the LAE-ST-MoE model achieves a 9.26 % mix error reduction on the CS test with the same decoding parameter. Moreover, the well-trained LAE-ST-MoE model can perform ST tasks from CS speech to Mandarin or English text.
Yuke Li 0002, Binbin Du, Haoran Fu
ASRU5
2023 Improving CTC-Based ASR Models With Gated Interlayer Collaboration
abstract
The CTC-based automatic speech recognition (ASR) models without the external language model usually lack the capacity to model conditional dependencies and textual interactions. In this paper, we present a Gated Interlayer Collaboration (GIC) mechanism to improve the performance of CTC-based models, which introduces textual information into the model and thus relaxes the conditional independence assumption of CTC-based models. Specifically, we consider the weighted sum of token embeddings as the textual representation for each position, where the position-specific weights are the softmax probability distribution constructed via inter-layer auxiliary CTC losses. The textual representations are then fused with acoustic features by developing a gate unit. Experiments on AISHELL-1 [1], TEDLIUM2 [2], and AI-DATATANG [3] corpora show that the proposed method outperforms several strong baselines.
Yuke Li 0002, Binbin Du
ICASSP3
2023 Language-Routing Mixture of Experts for Multilingual and Code-Switching Speech Recognition
Yuke Li 0002, Binbin Du
INTERSPEECH4
2023 Enhancing the Unified Streaming and Non-streaming Model with Contrastive Learning
Yuke Li 0002, Binbin Du
INTERSPEECH3