EDBT 2026 Demo / reviewers in the wild / expert
Kaixun Huang
dblp:332/1577
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0001-7498-4397ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Contextualized Automatic Speech Recognition with Dynamic Vocabulary Prediction and Activation
Zhennan Lin, Kaixun Huang, Linju Yang, Lei Xie 0001 |
INTERSPEECH | 2 |
| 2024 | SSHR: Leveraging Self-supervised Hierarchical Representations for Multilingual Automatic Speech RecognitionabstractMultilingual automatic speech recognition (ASR) systems have garnered attention for their potential to extend language coverage globally. While self-supervised learning (SSL) models, like MMS, have demonstrated their effectiveness in multilingual ASR, it is worth noting that various layers’ representations potentially contain distinct information that has not been fully leveraged. In this study, we propose a novel method that leverages self-supervised hierarchical representations (SSHR) to fine-tune the MMS model. We first analyze the different layers of MMS and show that the middle layers capture language-related information, and the high layers encode content-related information, which gradually decreases in the final layers. Then, we extract a language-related frame from correlated middle layers and guide specific language extraction through self-attention mechanisms. Additionally, we steer the model toward acquiring more content-related information in the final layers using our proposed Cross-CTC. We evaluate SSHR on two multilingual datasets, Common Voice and ML-SUPERB, and the experimental results demonstrate that our method achieves state-of-the-art performance to the best of our knowledge. Hongfei Xue, Qijie Shao, Kaixun Huang, Peikun Chen, Jie Liu 0097, Lei Xie 0001 |
ICME | 3 |
| 2024 | SEQ-former: A context-enhanced and efficient automatic speech recognition framework
Kaixun Huang, Lei Xie 0001, Zongfeng Quan, Weihong Deng |
INTERSPEECH | 3 |
| 2024 | Towards Rehearsal-Free Multilingual ASR: A LoRA-based Case Study on Whisper
Kaixun Huang, Longtao Huang, Lei Xie 0001 |
INTERSPEECH | 2 |
| 2023 | Spike-Triggered Contextual Biasing for End-to-End Mandarin Speech RecognitionabstractThe attention-based deep contextual biasing method has been demonstrated to effectively improve the recognition performance of end-to-end automatic speech recognition (ASR) systems on given contextual phrases. However, unlike shallow fusion methods that directly bias the posterior of the ASR model, deep biasing methods implicitly integrate contextual information, making it challenging to control the degree of bias. In this study, we introduce a spike-triggered deep biasing method that simultaneously supports both explicit and implicit bias. Moreover, both bias approaches exhibit significant improvements and can be cascaded with shallow fusion methods for better results. Furthermore, we propose a context sampling enhancement strategy and improve the contextual phrase filtering algorithm. Experiments on the public WenetSpeech Mandarin biased-word dataset show a 32.0% relative CER reduction compared to the baseline model, with an impressively 68.6% relative CER reduction on contextual phrases. Kaixun Huang, Xingchen Song, Lei Xie 0001 |
ASRU | 1 |
| 2023 | U2-KWS: Unified Two-Pass Open-Vocabulary Keyword Spotting with Keyword BiasabstractOpen-vocabulary keyword spotting (KWS), which allows users to customize keywords, has attracted increasingly more interest. However, existing methods based on acoustic models and post-processing train the acoustic model with ASR training criteria to model all phonemes, making the acoustic model under-optimized for the KWS task. To solve this problem, we propose a novel unified two-pass open-vocabulary KWS (U2-KWS) framework inspired by the two-pass ASR model U2. Specifically, we employ the CTC branch as the first stage model to detect potential keyword candidates and the decoder branch as the second stage model to validate candidates. In order to enhance any customized keywords, we redesign the U2 training procedure for U2-KWS and add keyword information by audio and text cross-attention into both branches. We perform experiments on our internal dataset and Aishell-1. The results show that U2-KWS can achieve a significant relative wake-up rate improvement of 41 % compared to the traditional customized KWS systems when the false alarm rate is fixed to 0.5 times per hour. Kaixun Huang, Lei Xie 0001 |
ASRU | 3 |
| 2023 | Contextualized End-to-End Speech Recognition with Contextual Phrase Prediction Network
Kaixun Huang, Zhanheng Yang, Bingshen Mu, Lei Xie 0001 |
INTERSPEECH | 1 |
| 2023 | Adaptive Contextual Biasing for Transducer Based Streaming Speech Recognition
Zhanheng Yang, Kaixun Huang, Changru Chen, Lei Xie 0001 |
INTERSPEECH | 3 |