Xun Gong 0005

dblp:58/5901-5 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
15since 2021 · last 2025
0000-0002-3364-8407ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021
YearPublicationVenuePosition
2025 BR-ASR: Efficient and Scalable Bias Retrieval Framework for Contextual Biasing ASR in Speech LLM
Xun Gong 0005, Anqi Lv, Wangyou Zhang, Huijia Zhu, Yanmin Qian
INTERSPEECH1
2025 Ranking and Selection of Bias Words for Contextual Bias Speech Recognition
Haoxiang Hou, Xun Gong 0005, Wangyou Zhang, Wei Wang 0010, Yanmin Qian
INTERSPEECH2
2024 Contextual Biasing Speech Recognition in Speech-enhanced Large Language Model
Xun Gong 0005, Anqi Lv, Yanmin Qian
INTERSPEECH1
2024 DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
abstract
As a popular multilingual and multitask pre-trained speech model, Whisper has the problem of curse of multilinguality. To enhance multilingual capabilities in small Whisper models, we propose DQ-Whisper, a novel joint distillation and quantization framework to compress Whisper for efficient inference. Firstly, we propose a novel dynamic matching distillation strategy. Then, a quantization-aware distillation framework is introduced to integrate quantization with distillation. Experimental results on various multilingual datasets show that our suggested distillation approach can effectively enhance the multilingual capabilities of small Whisper models without increasing computational costs. Up to 5.18x reduction in model size is achieved with marginal performance degradation. In addition, quantization is compatible with distillation, which can result in a higher compression rate.
Hang Shao 0005, Bei Liu 0003, Wei Wang 0010, Xun Gong 0005, Yanmin Qian
SLT4
2024 Advanced Long-Content Speech Recognition With Factorized Neural Transducer
abstract
Long-form automatic speech recognition (ASR) has obtained increasing interest in recent years, as it captures the relationship among consecutive historical sentences while decoding the current sentence. In this paper, we propose two novel approaches, which integrate long-form information into the factorized neural transducer (FNT) based architecture in both non-streaming (referred to asLongFNT) and streaming (referred to asSLongFNT) scenarios. We first investigate whether long-form transcriptions can improve the vanilla conformer transducer (C-T) models. Our experiments indicate that the vanilla C-T models do not exhibit improved performance when utilizing long-form transcriptions, possibly due to the predictor network of C-T models not functioning as a pure language model. Instead, FNT shows its potential in utilizing long-form information, where we propose theLongFNTmodel and explore the impact of long-form information in both text (LongFNT-Text) and speech (LongFNT-Speech). The proposed LongFNT-Text and LongFNT-Speech models further complement each other to achieve better performance, with transcription history proving more valuable to the model. The effectiveness of our LongFNT approach is evaluated on LibriSpeech and GigaSpeech corpora, and obtains relative 19% and 12% word error rate reduction, respectively. Furthermore, we extend the LongFNT model to the streaming scenario, which is namedSLongFNT, consisting of SLongFNT-Text and SLongFNT-Speech approaches to utilize long-form text and speech information. Experiments show that the proposed SLongFNT model achieves relative 26% and 17% WER reduction on LibriSpeech and GigaSpeech respectively while keeping a good latency, compared to the FNT baseline. Overall, our proposedLongFNTandSLongFNThighlight the significance of considering long-form speech and transcription knowledge for improving both non-streaming and streaming speech recognition systems.
Xun Gong 0005, Yu Wu 0012, Jinyu Li 0001, Shujie Liu 0001, Rui Zhao 0017, Xie Chen 0001, Yanmin Qian
IEEE ACM Trans. Audio Speech Lang. Process.1
2024 SpeechLM: Enhanced Speech Pre-Training With Unpaired Textual Data
abstract
How to boost speech pre-training with textual data is an unsolved problem due to the fact that speech and text are very different modalities with distinct characteristics. In this paper, we propose a cross-modalSpeechandLanguageModel (SpeechLM) to explicitly align speech and text pre-training with a pre-defined unified discrete representation. Specifically, we introduce two alternative discrete tokenizers to bridge the speech and text modalities, including phoneme-unit and hidden-unit tokenizers, which can be trained using unpaired speech or a small amount of paired speech-text data. Based on the trained tokenizers, we convert the unlabeled speech and text data into tokens of phoneme units or hidden units. The pre-training objective is designed to unify the speech and the text into the same discrete semantic space with a unified Transformer network. We evaluate SpeechLM on various spoken language processing tasks including speech recognition, speech translation, and universal representation evaluation framework SUPERB, demonstrating significant improvements on content-related tasks. Code and models are available athttps://aka.ms/SpeechLM.
Sanyuan Chen, Yu Wu 0012, Shuo Ren 0002, Shujie Liu 0001, Zhuoyuan Yao, Xun Gong 0005, Li-Rong Dai 0001, Jinyu Li 0001, Furu Wei
IEEE ACM Trans. Audio Speech Lang. Process.8
2023 Efficient Text-Only Domain Adaptation For CTC-Based ASR
abstract
For connectionist temporal classification (CTC) based speech recognition (ASR) models, text-only domain adaptation still faces several challenges. In this study, we propose an efficient text-only domain adaptation method for CTC-based models. We introduce the assistant textual adapter (ATA) to learn textual features and transform them into the latent space of the acoustic encoder. With the help of the ATA module, the adaptation is achieved by fine-tuning the top layers of the acoustic encoder with the target domain text. Meanwhile, further improvement can be obtained by the integration with shallow fusion (SF). Adapted from LibriSpeech, experiments show that the proposed method can achieve averaged 29.7% relative WER reduction (WERR) compared with the un-adapted baseline on WSJ, and 10.5% WERR compared to SF as well. Moreover, it also shows 15.4∼37.1% WERR for 10 GigaSpeech target domains test sets compared to the un-adapted baseline, and also 6.5% WERR on average compared with SF.
Xun Gong 0005, Yanmin Qian
ASRU2
2023 LongFNT: Long-Form Speech Recognition with Factorized Neural Transducer
abstract
Traditional automatic speech recognition (ASR) systems usually focus on individual utterances, without considering long-form speech with useful historical information, which is more practical in real scenarios. Simply attending longer transcription history for a vanilla neural transducer model shows no much gain in our preliminary experiments, since the prediction network is not a pure language model. This motivates us to leverage the factorized neural transducer structure, containing a real language model, the vocabulary predictor. We propose the LongFNT-Text architecture, which fuses the sentence-level long-form features directly with the output of the vocabulary predictor and then embeds token-level long-form features inside the vocabulary predictor, with a pre-trained contextual encoder RoBERTa to further boost the performance. Moreover, we propose the LongFNT architecture by extending the long-form speech to the original speech input and achieve the best performance. The effectiveness of our LongFNT approach is validated on LibriSpeech and GigaSpeech corpora with 19% and 12% relative word error rate (WER) reduction, respectively.
Xun Gong 0005, Yu Wu 0012, Jinyu Li 0001, Shujie Liu 0001, Rui Zhao 0017, Xie Chen 0001, Yanmin Qian
ICASSP1
2023 Factorized AED: Factorized Attention-Based Encoder-Decoder for Text-Only Domain Adaptive ASR
abstract
End-to-end automatic speech recognition (ASR) systems have gained popularity given their simplified architecture and promising results. However, text-only domain adaptation remains a big challenge for E2E systems. Text-to-speech (TTS) based approaches fine-tune ASR models by synthesized speech with an auxiliary TTS model, thus increase deployment costs. Language model (LM) fusion based approaches can achieve good performance but are sensitive to interpolation parameters. In order to factorize out the language component in the AED model, we propose the factorized attention-based encoder-decoder (Factorized AED) model whose decoder takes as input the posterior probabilities of a jointly trained LM. Moreover, in the context of domain adaptation, the domain specific LM serves as a plug-and-play component for a well-trained factorized AED model. In-domain experiments on LibriSpeech and out-of-domain experiments adapting from LibriSpeech to a variety of domains in GigaSpeech are conducted to validate the effectiveness of our proposed methods. Results show 20% / 24% relative word error rate (WER) reduction for LibriSpeech test sets and 8 ∼34% relative WER reduction for 8 GigaSpeech target domains test sets compared to the AED baseline.
Xun Gong 0005, Wei Wang 0010, Hang Shao 0005, Xie Chen 0001, Yanmin Qian
ICASSP1
2023 Joint Discriminator and Transfer Based Fast Domain Adaptation For End-To-End Speech Recognition
abstract
Adapting End-to-End (E2E) models to unseen domains is still a big challenge since training E2E models requires lots of paired audio and text training data. We propose a novel domain adaptation framework for the E2E model, which only uses the text of the target domain. Moreover, the proposed methods can keep the performance on the source domain intact while greatly improving the performance on the target domain. The proposed framework consists of two parts: the discriminator and the transfer which were optimized separately. Finally, optimized discriminator and transfer were combined and evaluated on two domain adaption tasks. In the experiments of adapting the English Librispeech to Gigaspeech, we obtained an average relative 11.6% and 11.8% on word error rate (WER) reduction for the target domain dev and test sets, respectively, while almost without WER degradation on the source domain. For the inhouse Chinese corpus aviation and TV, the character error rate (CER) of the source domain increased within 5%, while the CER on the target domain achieved around relative 85% and 42% improvement, respectively. In addition, our approach is also more effective in the mixed domain scenarios in the evaluation.
Hang Shao 0005, Tian Tan 0002, Wei Wang 0010, Xun Gong 0005, Yanmin Qian
ICASSP4
2023 Text Only Domain Adaptation with Phoneme Guided Data Splicing for End-to-End Speech Recognition
Wei Wang 0010, Xun Gong 0005, Hang Shao 0005, Dongning Yang, Yanmin Qian
INTERSPEECH2
2022 The Sjtu System For Multimodal Information Based Speech Processing Challenge 2021
abstract
This paper describes the SJTU system for ICASSP Multi-modal Information based Speech Processing Challenge (MISP) 2021. To solve the speech recognition problem in real complex environments where time-synchronized near- and far-field signals are available for training an enhancement frontend. We build a joint system with speech enhancement frontend and speech recognition backend. These two modules are optimized jointly by both ASR and enhancement criteria. Audio-visual fusion is explored to further boost the ASR performance. ROVER and test time augmentation techniques are used to combine recognition results from multiple systems. The final system achieves Chinese character error rates (CCER) of 34.9% on dev set and 34.0% on test set, which achieved third place in the MISP challenge. The absolute CCER reduction compared with the official baseline system is 26.9% on dev set and 28.7% on test set.
Wei Wang 0010, Xun Gong 0005, Zhikai Zhou, Chenda Li, Wangyou Zhang, Bing Han 0008, Yanmin Qian
ICASSP2
2022 Knowledge Transfer and Distillation from Autoregressive to Non-Autoregessive Speech Recognition
Xun Gong 0005, Zhikai Zhou, Yanmin Qian
INTERSPEECH1
2022 Layer-Wise Fast Adaptation for End-to-End Multi-Accent Speech Recognition
abstract
The variety and complexity of accents pose a huge challenge to robust Automatic Speech Recognition (ASR). Some previous work has attempted to address such problems, however most of the current approaches either require prior knowledge about the target accent, or cannot handle unseen accents and accent-unspecific standard speech. In this work, we aim to improve multi-accent speech recognition in the end-to-end (E2E) framework with a novel layer-wise adaptation architecture. Firstly, we propose a robust deep accent representation learning architecture to obtain accurate accent embedding, and some advanced schemes are designed to further boost the quality of accent embeddings, including phone posteriorgram (PPG) feature, TTS based data augmentation in the training stage, test-time augmentation and multi-embedding fusion in the testing stage. Then, the layer-wise adaptation with accent embeddings is developed for fast accent adaptation in ASR, and two types of adapter layers are designed, including the gated adapter layer and multi-basis adapter layer. Compared to the usual two-pass adaptation, these adapter layers are injected between the ASR encoder layers to encode the accent information in ASR flexibly, and perform fast adaption on the corresponding speech accent. The experiments on Accent AESRC corpus show that the proposed deep accent representation learning can capture accurate accent knowledge, and get high performance on accent classification. The new layer-wise adaptation architecture with the accurate accent embedding outperforms the other traditional methods, and obtains consistent$\sim$15% relative word error rate (WER) reduction on all kinds of testing scenarios, including seen accents, unseen accents and accent-unspecific standard speech.
Yanmin Qian, Xun Gong 0005, Houjun Huang
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 Layer-Wise Fast Adaptation for End-to-End Multi-Accent Speech Recognition
abstract
Accent variability has posed a huge challenge to automatic speech recognition~(ASR) modeling. Although one-hot accent vector based adaptation systems are commonly used, they require prior knowledge about the target accent and cannot handle unseen accents. Furthermore, simply concatenating accent embeddings does not make good use of accent knowledge, which has limited improvements. In this work, we aim to tackle these problems with a novel layer-wise adaptation structure injected into the E2E ASR model encoder. The adapter layer encodes an arbitrary accent in the accent space and assists the ASR model in recognizing accented speech. Given an utterance, the adaptation structure extracts the corresponding accent information and transforms the input acoustic feature into an accent-related feature through the linear combination of all accent bases. We further explore the injection position of the adaptation layer, the number of accent bases, and different types of accent bases to achieve better accent adaptation. Experimental results show that the proposed adaptation structure brings 12\% and 10\% relative word error rate~(WER) reduction on the AESRC2020 accent dataset and the Librispeech dataset, respectively, compared to the baseline.
Xun Gong 0005, Yizhou Lu, Zhikai Zhou, Yanmin Qian
Interspeech1
2020 Text Adaptation for Speaker Verification with Speaker-Text Factorized Embeddings
abstract
Text mismatch between pre-collected data, either training data or enrollment data, and the actual test data can significantly hurt text-dependent speaker verification (SV) system performance. Although this problem can be solved by carefully collecting data with the target speech content, such data collection could be costly and inflexible. In this paper, we propose a novel text adaptation framework to address the text mismatch issue. Here, a speaker-text factorization network is proposed to factorize the input speech into speaker embeddings and text embeddings and then integrate them into a single representation in the later stage. Given a small amount of speaker-independent adaptation utterances, text embeddings of target speech content can be extracted and used to adapt the text-independent speaker embeddings to text-customized speaker embeddings. Experiments on RSR2015 show that text adaptation can significantly improve the performance of text mismatch conditions.
Yexin Yang, Shuai Wang 0016, Xun Gong 0005, Yanmin Qian, Kai Yu 0004
ICASSP3