Xucheng Wan

dblp:287/6420 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
10since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 CAMEL: Cross-Attention Enhanced Mixture-of-Experts and Language Bias for Code-Switching Speech Recognition
abstract
Code-switching automatic speech recognition (ASR) aims to transcribe speech that contains two or more languages accurately. To better capture language-specific speech representations and address language confusion in code-switching ASR, the mixture-of-experts (MoE) architecture and an additional language diarization (LD) decoder are commonly employed. However, most researches remain stagnant in simple operations like weighted summation or concatenation to fuse language-specific speech representations, leaving significant opportunities to explore the enhancement of integrating language bias information. In this paper, we introduce CAMEL, a cross-attention-based MoE and language bias approach for code-switching ASR. Specifically, after each MoE layer, we fuse language-specific speech representations with cross-attention, leveraging its strong contextual modeling abilities. Additionally, we design a source attention-based mechanism to incorporate the language information from the LD decoder output into text embeddings. Experimental results demonstrate that our approach achieves state-of-the-art performance on the SEAME, ASRU200, and ASRU700+LibriSpeech460 Mandarin-English code-switching ASR datasets.
He Wang 0022, Xucheng Wan, Naijun Zheng, Kai Liu 0053, Huan Zhou 0004, Guojian Li, Lei Xie 0001
ICASSP2
2025 SCDiar: a streaming diarization system based on speaker change detection and speech recognition
abstract
In hours-long meeting scenarios, real-time speech stream often struggles with achieving accurate speaker diarization, commonly leading to speaker identification and speaker count errors. To address this challenge, we propose SCDiar, a system that operates on speech segments, split at the token level by a speaker change detection (SCD) module. Building on these segments, we introduce several enhancements to efficiently select the best available segment for each speaker. These improvements lead to significant gains across various benchmarks. Notably, on real-world meeting data involving more than ten participants, SCDiar outperforms previous systems by up to 53.6% in accuracy, substantially narrowing the performance gap between online and offline systems.
Naijun Zheng, Xucheng Wan, Kai Liu 0053, Huan Zhou 0008
ICASSP2
2024 Real-time scheme for rapid extraction of speaker embeddings in challenging recording conditions
Kai Liu 0053, Ziqing Du, Huan Zhou 0008, Xucheng Wan, Naijun Zheng
INTERSPEECH4
2024 An efficient text augmentation approach for contextualized Mandarin speech recognition
Naijun Zheng, Xucheng Wan, Kai Liu 0053, Ziqing Du, Huan Zhou 0008
INTERSPEECH2
2024 Dynamic Resource Management in MEC Powered by Edge Intelligence for Smart City Internet of Things
Xucheng Wan
J. Grid Comput.1
2024 MMGER: Multi-Modal and Multi-Granularity Generative Error Correction With LLM for Joint Accent and Speech Recognition
abstract
Despite notable advancements in automatic speech recognition (ASR), performance tends to degrade when faced with adverse conditions. Generative error correction (GER) leverages the exceptional text comprehension capabilities of large language models (LLM), delivering impressive performance in ASR error correction, where N-best hypotheses provide valuable information for transcription prediction. However, GER encounters challenges such as fixed N-best hypotheses, insufficient utilization of acoustic information, and limited specificity to multi-accent scenarios. In this paper, we explore the application of GER in multi-accent scenarios. Accents represent deviations from standard pronunciation norms, and the multi-task learning framework for simultaneous ASR and accent recognition (AR) has effectively addressed the multi-accent scenarios, making it a prominent solution. In this work, we propose a unified ASR-AR GER model, named MMGER, leveraging multi-modal correction, and multi-granularity correction. Multi-task ASR-AR learning is employed to provide dynamic 1-best hypotheses and accent embeddings. Multi-modal correction accomplishes fine-grained frame-level correction by force-aligning the acoustic features of speech with the corresponding character-level 1-best hypothesis sequence. Multi-granularity correction supplements the global linguistic information by incorporating regular 1-best hypotheses atop fine-grained multi-modal correction to achieve coarse-grained utterance-level correction. MMGER effectively mitigates the limitations of GER and tailors LLM-based ASR error correction for the multi-accent scenarios. Experiments conducted on the multi-accent Mandarin KeSpeech dataset demonstrate the efficacy of MMGER, achieving a 26.72% relative improvement in AR accuracy and a 27.55% relative reduction in ASR character error rate, compared to a well-established standard baseline.
Bingshen Mu, Xucheng Wan, Naijun Zheng, Huan Zhou 0004, Lei Xie 0001
IEEE Signal Process. Lett.2
2023 BA-MoE: Boundary-Aware Mixture-of-Experts Adapter for Code-Switching Speech Recognition
abstract
Mixture-of-experts based models, which use language experts to extract language-specific representations effectively, have been well applied in code-switching automatic speech recognition. However, there is still substantial space to improve as similar pronunciation across languages may result in ineffective multi-language modeling and inaccurate language boundary estimation. To eliminate these drawbacks, we propose a cross-layer language adapter and a boundary-aware training method, namely Boundary-Aware Mixture-of-Experts (BA-MoE). Specifically, we introduce language-specific adapters to separate language-specific representations and a unified gating layer to fuse representations within each encoder layer. Second, we compute language adaptation loss of the mean output of each language-specific adapter to improve the adapter module’s language-specific representation learning. Besides, we utilize a boundary-aware predictor to learn boundary representations for dealing with language boundary confusion. Our approach achieves significant performance improvement, reducing the mixture error rate by 16.55% compared to the baseline on the ASRU 2019 Mandarin-English code-switching challenge dataset.
Peikun Chen, Fan Yu 0002, Yuhao Liang, Hongfei Xue, Xucheng Wan, Naijun Zheng, Huan Zhou 0004, Lei Xie 0001
ASRU5
2023 X-SEPFORMER: End-To-End Speaker Extraction Network with Explicit Optimization on Speaker Confusion
abstract
Target speech extraction (TSE) systems are designed to extract target speech from a multi-talker mixture. The popular training objective for most prior TSE networks is to enhance reconstruction performance of extracted speech waveform. However, it has been reported that a TSE system delivers high reconstruction performance may still suffer low-quality experience problems in practice. One such experience problem is wrong speaker extraction (called speaker confusion, SC), which leads to strong negative experience and hampers effective conversations. To mitigate the imperative SC issue, we reformulate the training objective and propose two novel loss schemes that explore the metric of reconstruction improvement performance defined at small chunk-level and leverage the metric associated distribution information. Both loss schemes aim to encourage a TSE network to pay attention to those SC chunks based on the said distribution information. On this basis, we present X-SepFormer, an end-to-end TSE model with proposed loss schemes and a backbone of SepFormer. Experimental results on the benchmark WSJ0-2mix dataset validate the effectiveness of our proposals, showing consistent improvements on SC errors (by 14.8% relative). Moreover, with SI-SDRi of 19.4 dB and PESQ of 3.81, our best system significantly outperforms the current SOTA systems and offers the top TSE results reported till date on the WSJ0-2mix.
Kai Liu 0053, Ziqing Du, Xucheng Wan, Huan Zhou 0008
ICASSP3
2021 Online Speaker Diarization Equipped with Discriminative Modeling and Guided Inference
Xucheng Wan, Kai Liu 0053, Huan Zhou 0008
Interspeech1
2021 Analysis and Exploration of Open Source Data in Traffic Network Based on Scheduling Model of Bike-Sharing
abstract
With the rapid development of sharing economy, bike-sharing becomes essential because of its zero emission, high flexibility and accessibility. The emergence of the public bicycle system not only alleviates the traffic pressure to a certain extent, but also contributes to solving the “last kilometer” problem of public transportation. However, due to the concentrated use of shared bikes, many shared bikes are left in disorder, which seriously affects the urban environment and causes traffic problems. How to manage the allocation of bike-sharing and improve the city’s shared cycling system have become a highly discussed issue. We, taking Beijing as an example, research on the allocation of shared bikes by using the open source data provided by Amap, Baidu Map and websites of shared bikes, which are used to analyze the allocation, and establish an optimizing comprehensive evaluation model to evaluate the required level. In the end, we look forward the future of bike-sharing market.
Xucheng Wan, Suyu Zhang
Int. J. Pattern Recognit. Artif. Intell.2