Xiaobin Rong

dblp:359/7624 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0006-4373-8306ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 PASE: Leveraging the Phonological Prior of WavLM for Low-Hallucination Generative Speech Enhancement
abstract
Generative models have shown remarkable performance in speech enhancement (SE), achieving superior perceptual quality over traditional discriminative approaches. However, existing generative SE approaches often overlook the risk of hallucination under severe noise, leading to incorrect spoken content or inconsistent speaker characteristics, which we term linguistic and acoustic hallucinations, respectively. We argue that linguistic hallucination stems from models' failure to constrain valid phonological structures and it is a more fundamental challenge. While language models (LMs) are well-suited for capturing the underlying speech structure through modeling the distribution of discrete tokens, existing approaches are limited in learning from noise-corrupted representations, which can lead to contaminated priors and hallucinations. To overcome these limitations, we propose the Phonologically Anchored Speech Enhancer (PASE), a generative SE framework that leverages the robust phonological prior embedded in the pre-trained WavLM model to mitigate hallucinations. First, we adapt WavLM into a denoising expert via representation distillation to clean its final-layer features. Guided by the model's intrinsic phonological prior, this process enables robust denoising while minimizing linguistic hallucinations. To further reduce acoustic hallucinations, we train the vocoder with a dual-stream representation: the high-level phonetic representation provides clean linguistic content, while a low-level acoustic representation retains speaker identity and prosody. Experimental results demonstrate that PASE not only surpasses state-of-the-art discriminative models in perceptual quality, but also significantly outperforms prior generative models with substantially lower linguistic and acoustic hallucinations.
Xiaobin Rong, Qinwen Hu, Mansur Yesilbursa, Kamil Wójcicki
AAAI1
2026 Optimization of modular multi-speaker distant conversational speech recognition
Qinwen Hu, Tianchi Sun, Xiaobin Rong
Comput. Speech Lang.4
2026 DDSE: Efficient Neural Codec Language Models for speech enhancement with disentangled representations
Qinwen Hu, Xiaobin Rong, Mansur Yesilbursa, Kamil Wójcicki
Speech Commun.2
2026 CoFi-Lite: Pushing the Limits of Ultra-Lightweight Speech Enhancement
Leyan Yang, Dahan Wang, Xiaobin Rong, Jiadong Zhao
IEEE Signal Process. Lett.3
2025 TS-URGENet: A Three-stage Universal Robust and Generalizable Speech Enhancement Network
Xiaobin Rong, Dahan Wang, Qinwen Hu
INTERSPEECH1
2025 A Lightweight Hybrid Dual Channel Speech Enhancement System under Low-SNR Conditions
Xiaobin Rong, Tianchi Sun
INTERSPEECH2
2024 GTCRN: A Speech Enhancement Model Requiring Ultralow Computational Resources
abstract
While modern deep learning-based models have significantly outperformed traditional methods in the area of speech enhancement, they often necessitate a lot of parameters and extensive computational power, making them impractical to be deployed on edge devices in real-world applications. In this paper, we introduce Grouped Temporal Convolutional Recurrent Network (GTCRN), which incorporates grouped strategies to efficiently simplify a competitive model, DPCRN. Additionally, it leverages subband feature extraction modules and temporal recurrent attention modules to enhance its performance. Remarkably, the resulting model demands ultralow computational resources, featuring only 23.7 K parameters and 39.6 MMACs per second. Experimental results show that our proposed model not only surpasses RNNoise, a typical lightweight model with similar computational burden, but also achieves competitive performance when compared to recent baseline models with significantly higher computational resources requirements.
Xiaobin Rong, Tianchi Sun, Changbao Zhu
ICASSP1
2023 A Low-Latency Hybrid Multi-Channel Speech Enhancement System For Hearing Aids
abstract
This paper summarizes a hybrid multi-channel speech enhancement system for the ICASSP Signal Processing Grand Challenge: Clarity Challenge (Speech Enhancement for Hearing Aids) 2023. The system consists of a rule-based dereverberation module, a multi-channel enhancement module, and a post-processing module. Without using the head rotation information and the enrollment speech, the system can reach an average hearing aid speech perception index (HASPI) score of 0.696 and hearing aid speech quality index (HASQI) score of 0.320 on the official development set. The corresponding scores are 0.729 and 0.316 respectively on the Eval1 set for the challenge ranking.
Zhongshu Hou, Wanyu Yang, Tianchi Sun, Xiaobin Rong, Dahan Wang, Kai Chen 0029
ICASSP6