EDBT 2026 Demo / reviewers in the wild / expert
Naijun Zheng
dblp:163/6508
· DBLP profile ↗
14ranked-venue papers
8as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 7 first-author · 11 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CAMEL: Cross-Attention Enhanced Mixture-of-Experts and Language Bias for Code-Switching Speech RecognitionabstractCode-switching automatic speech recognition (ASR) aims to transcribe speech that contains two or more languages accurately. To better capture language-specific speech representations and address language confusion in code-switching ASR, the mixture-of-experts (MoE) architecture and an additional language diarization (LD) decoder are commonly employed. However, most researches remain stagnant in simple operations like weighted summation or concatenation to fuse language-specific speech representations, leaving significant opportunities to explore the enhancement of integrating language bias information. In this paper, we introduce CAMEL, a cross-attention-based MoE and language bias approach for code-switching ASR. Specifically, after each MoE layer, we fuse language-specific speech representations with cross-attention, leveraging its strong contextual modeling abilities. Additionally, we design a source attention-based mechanism to incorporate the language information from the LD decoder output into text embeddings. Experimental results demonstrate that our approach achieves state-of-the-art performance on the SEAME, ASRU200, and ASRU700+LibriSpeech460 Mandarin-English code-switching ASR datasets. He Wang 0022, Xucheng Wan, Naijun Zheng, Kai Liu 0053, Huan Zhou 0004, Guojian Li, Lei Xie 0001 |
ICASSP | 3 |
| 2025 | SCDiar: a streaming diarization system based on speaker change detection and speech recognitionabstractIn hours-long meeting scenarios, real-time speech stream often struggles with achieving accurate speaker diarization, commonly leading to speaker identification and speaker count errors. To address this challenge, we propose SCDiar, a system that operates on speech segments, split at the token level by a speaker change detection (SCD) module. Building on these segments, we introduce several enhancements to efficiently select the best available segment for each speaker. These improvements lead to significant gains across various benchmarks. Notably, on real-world meeting data involving more than ten participants, SCDiar outperforms previous systems by up to 53.6% in accuracy, substantially narrowing the performance gap between online and offline systems. Naijun Zheng, Xucheng Wan, Kai Liu 0053, Huan Zhou 0008 |
ICASSP | 1 |
| 2024 | Real-time scheme for rapid extraction of speaker embeddings in challenging recording conditions
Kai Liu 0053, Ziqing Du, Huan Zhou 0008, Xucheng Wan, Naijun Zheng |
INTERSPEECH | 5 |
| 2024 | An efficient text augmentation approach for contextualized Mandarin speech recognition
Naijun Zheng, Xucheng Wan, Kai Liu 0053, Ziqing Du, Huan Zhou 0008 |
INTERSPEECH | 1 |
| 2024 | MMGER: Multi-Modal and Multi-Granularity Generative Error Correction With LLM for Joint Accent and Speech RecognitionabstractDespite notable advancements in automatic speech recognition (ASR), performance tends to degrade when faced with adverse conditions. Generative error correction (GER) leverages the exceptional text comprehension capabilities of large language models (LLM), delivering impressive performance in ASR error correction, where N-best hypotheses provide valuable information for transcription prediction. However, GER encounters challenges such as fixed N-best hypotheses, insufficient utilization of acoustic information, and limited specificity to multi-accent scenarios. In this paper, we explore the application of GER in multi-accent scenarios. Accents represent deviations from standard pronunciation norms, and the multi-task learning framework for simultaneous ASR and accent recognition (AR) has effectively addressed the multi-accent scenarios, making it a prominent solution. In this work, we propose a unified ASR-AR GER model, named MMGER, leveraging multi-modal correction, and multi-granularity correction. Multi-task ASR-AR learning is employed to provide dynamic 1-best hypotheses and accent embeddings. Multi-modal correction accomplishes fine-grained frame-level correction by force-aligning the acoustic features of speech with the corresponding character-level 1-best hypothesis sequence. Multi-granularity correction supplements the global linguistic information by incorporating regular 1-best hypotheses atop fine-grained multi-modal correction to achieve coarse-grained utterance-level correction. MMGER effectively mitigates the limitations of GER and tailors LLM-based ASR error correction for the multi-accent scenarios. Experiments conducted on the multi-accent Mandarin KeSpeech dataset demonstrate the efficacy of MMGER, achieving a 26.72% relative improvement in AR accuracy and a 27.55% relative reduction in ASR character error rate, compared to a well-established standard baseline. Bingshen Mu, Xucheng Wan, Naijun Zheng, Huan Zhou 0004, Lei Xie 0001 |
IEEE Signal Process. Lett. | 3 |
| 2023 | BA-MoE: Boundary-Aware Mixture-of-Experts Adapter for Code-Switching Speech RecognitionabstractMixture-of-experts based models, which use language experts to extract language-specific representations effectively, have been well applied in code-switching automatic speech recognition. However, there is still substantial space to improve as similar pronunciation across languages may result in ineffective multi-language modeling and inaccurate language boundary estimation. To eliminate these drawbacks, we propose a cross-layer language adapter and a boundary-aware training method, namely Boundary-Aware Mixture-of-Experts (BA-MoE). Specifically, we introduce language-specific adapters to separate language-specific representations and a unified gating layer to fuse representations within each encoder layer. Second, we compute language adaptation loss of the mean output of each language-specific adapter to improve the adapter module’s language-specific representation learning. Besides, we utilize a boundary-aware predictor to learn boundary representations for dealing with language boundary confusion. Our approach achieves significant performance improvement, reducing the mixture error rate by 16.55% compared to the baseline on the ASRU 2019 Mandarin-English code-switching challenge dataset. Peikun Chen, Fan Yu 0002, Yuhao Liang, Hongfei Xue, Xucheng Wan, Naijun Zheng, Huan Zhou 0004, Lei Xie 0001 |
ASRU | 6 |
| 2023 | Integrated and Enhanced Pipeline System to Support Spoken Language Analytics for Screening Neurocognitive Disordersabstract24th Annual Conference of the International Speech Communication Association, INTERSPEECH 2023, Dublin, Ireland, August 20-24, 2023 Helen M. Meng, Brian Kan-Wing Mak, Man-Wai Mak, Helene H. Fung, Xianmin Gong, Timothy C. Y. Kwok, Xunying Liu, Vincent C. T. Mok, Patrick C. M. Wong, Jean Woo, Xixin Wu, Ka-Ho Wong, Sean Shensheng Xu, Naijun Zheng, Ranzo Huang, Jiawen Kang 0002, Xiaoquan Ke, Junan Li, Jinchao Li |
INTERSPEECH | 14 |
| 2022 | Partially Fake Audio Detection by Self-Attention-Based Fake Span DiscoveryabstractThe past few years have witnessed the significant advances of speech synthesis and voice conversion technologies. However, such technologies can undermine the robustness of broadly implemented biometric identification models and can be harnessed by in-the-wild attackers for illegal uses. The ASVspoof challenge mainly focuses on synthesized audios by advanced speech synthesis and voice conversion models, and replay attacks. Recently, the first Audio Deep Synthesis Detection challenge (ADD 2022) extends the attack scenarios into more aspects. Also, ADD 2022 is the first challenge to propose the partially fake audio detection task. Such brand new attacks are dangerous and how to tackle such attacks remains an open question. Thus, we propose a novel framework by introducing the question-answering (fake span discovery) strategy with the self-attention mechanism to detect partially fake audios. The proposed fake span detection module tasks the anti-spoofing model to predict the start and end positions of the fake clip within the partially fake audio, address the model’s attention into discovering the fake spans rather than other shortcuts with less generalization, and finally equips the model with the discrimination capacity between real and partially fake audios. Our submission ranked second in the partially fake audio detection track of ADD 2022. Heng-Cheng Kuo, Naijun Zheng, Kuo-Hsuan Hung, Hung-yi Lee, Yu Tsao 0001, Hsin-Min Wang, Helen M. Meng |
ICASSP | 3 |
| 2022 | The CUHK-Tencent Speaker Diarization System for the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription ChallengeabstractThis paper describes our speaker diarization system submitted to the Multi-channel Multi-party Meeting Transcription (M2MeT) challenge, where Mandarin meeting data were recorded in multi-channel format for diarization and automatic speech recognition (ASR) tasks. In these meeting scenarios, the uncertainty of the speaker number and the high ratio of overlapped speech present great challenges for diarization. Based on the assumption that there is valuable complementary information between acoustic features, spatial-related and speaker-related features, we propose a multi-level feature fusion mechanism based target-speaker voice activity detection (FFM-TS-VAD) system to improve the performance of the conventional TS-VAD system. Furthermore, we propose a data augmentation method during training to improve the system robustness when the angular difference between two speakers is relatively small. We provide comparisons for different sub-systems we used in M2MeT challenge. Our submission is a fusion of several sub-systems and ranks second in the diarization task. Naijun Zheng, Na Li 0012, Xixin Wu, Lingwei Meng, Jiawen Kang 0002, Chao Weng, Dan Su 0002, Helen M. Meng |
ICASSP | 1 |
| 2022 | Multi-Channel Speaker Diarization Using Spatial Features for MeetingsabstractSpeaker identification for overlapped speech presents a great challenge for speaker diarization tasks in meeting scenarios. In order to overcome such challenges, several overlap-aware resegmentation methods based on deep learning have been integrated into speaker diarization systems. In this paper we propose two multi-channel diarization systems which have enhanced capability in detecting overlapped speech and identify speakers via learning spatial features. The first system applies a multi-look strategy to train networks without given the speakers’ direction of arrival(DOA), and the other system estimates the DOA of target speakers based on existing diarization results. Both systems aim to estimate the voice activity of speakers in different directions to handle overlapped speech. Experimental results on the AMI corpus show that the relative improvements of both systems can reach 9.4% and 18.1% in term of diarization error rate (DER) against an overlap-aware single-channel system with a BeamformIt front-end. Naijun Zheng, Na Li 0012, Jianwei Yu 0001, Chao Weng, Dan Su 0002, Xunying Liu, Helen M. Meng |
ICASSP | 1 |
| 2021 | A Joint Training Framework of Multi-Look Separator and Speaker Embedding Extractor for Overlapped SpeechabstractIn multi-talker cases, overlapped speech degrades the speaker verification (SV) performance dramatically. To tackle this challenging problem, speech separation with multi-channel techniques can be adopted to extract each speaker’s signals to improve the SV performance. In this paper, a joint training framework of the front-end multi-look speech separator and the back-end speaker embedding extractor is proposed for multi-channel overlapped speech. To better leverage the complementarity between the speech separator and the speaker embedding extractor, several training strategies are proposed to jointly optimize the two modules. Experimental results show that the proposed joint training framework significantly outperforms the individual SV system by around 52% relative EER reduction. Additionally, the robustness of the proposed framework is further evaluated under different conditions. Naijun Zheng, Na Li 0012, Bo Wu 0011, Meng Yu 0003, Jianwei Yu 0001, Chao Weng, Dan Su 0002, Xunying Liu, Helen M. Meng |
ICASSP | 1 |
| 2020 | Speaker-Aware Linear Discriminant Analysis in Speaker Verification
Naijun Zheng, Xixin Wu, Jinghua Zhong, Xunying Liu, Helen M. Meng |
INTERSPEECH | 1 |
| 2019 | Phase-Aware Speech Enhancement Based on Deep Neural NetworksabstractShort-time frequency transform (STFT) is fundamental in speech processing. Because of the difficulty of processing highly unstructured STFT phase, most speech-processing algorithms only operate with STFT magnitude, leaving the STFT phase far from explored. However, with the recent development of deep neural network (DNN) based speech processing, e.g., speech enhancement and recognition, phase processing is becoming more important than ever before as a new growing point of DNN-based methods. In this paper, we propose a phase-aware speech enhancement algorithm based on DNN. Specifically, in the training stage, when incorporating phase as a target, our core idea is to transform an unstructured phase spectrogram to its derivative along the time axis, i.e., instantaneous frequency deviation (IFD), which has a similar structure with its corresponding magnitude spectrogram. We further propose to optimize both IFD and magnitude jointly in a multiobjective learning framework. In the test stage, we propose a postprocessing method to recover the phase spectrogram from the estimated IFD. Experimental results demonstrate the effectiveness of the proposed method. Naijun Zheng, Xiao-Lei Zhang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2017 | LDPC code design for Gaussian multiple-access channels using dynamic EXIT chart analysisabstractWe consider the degree distribution design of the low-density parity-check (LDPC) code ensembles for symmetric Gaussian multiple-access channels (GMAC). To characterize the probability density function (PDF) of the message passing in the process of joint decoding, we propose a new scheme to construct the associated Gaussian mixture (GM) distribution, where each GM component is assigned according to the corresponding signal group transmitted by the users. By tracking the variation of the GM components in the iterative decoding process, more accurate mutual information can be obtained for the extrinsic information transfer (EXIT) chart analysis. Simulation results show that the performance of our proposed LDPC codes is better than that of the existing methods. Naijun Zheng, Baoming Bai, Anthony Man-Cho So, Kehu Yang |
ICASSP | 1 |