VLDB 2026 Research / reviewers in the wild / expert
Xiaobin Rong
dblp:359/7624
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0006-4373-8306ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PASE: Leveraging the Phonological Prior of WavLM for Low-Hallucination Generative Speech EnhancementabstractGenerative models have shown remarkable performance in speech enhancement (SE), achieving superior perceptual quality over traditional discriminative approaches. However, existing generative SE approaches often overlook the risk of hallucination under severe noise, leading to incorrect spoken content or inconsistent speaker characteristics, which we term linguistic and acoustic hallucinations, respectively. We argue that linguistic hallucination stems from models' failure to constrain valid phonological structures and it is a more fundamental challenge. While language models (LMs) are well-suited for capturing the underlying speech structure through modeling the distribution of discrete tokens, existing approaches are limited in learning from noise-corrupted representations, which can lead to contaminated priors and hallucinations. To overcome these limitations, we propose the Phonologically Anchored Speech Enhancer (PASE), a generative SE framework that leverages the robust phonological prior embedded in the pre-trained WavLM model to mitigate hallucinations. First, we adapt WavLM into a denoising expert via representation distillation to clean its final-layer features. Guided by the model's intrinsic phonological prior, this process enables robust denoising while minimizing linguistic hallucinations. To further reduce acoustic hallucinations, we train the vocoder with a dual-stream representation: the high-level phonetic representation provides clean linguistic content, while a low-level acoustic representation retains speaker identity and prosody. Experimental results demonstrate that PASE not only surpasses state-of-the-art discriminative models in perceptual quality, but also significantly outperforms prior generative models with substantially lower linguistic and acoustic hallucinations. Xiaobin Rong, Qinwen Hu, Mansur Yesilbursa, Kamil Wójcicki |
AAAI | 1 |
| 2026 | Optimization of modular multi-speaker distant conversational speech recognition
Qinwen Hu, Tianchi Sun, Xiaobin Rong |
Comput. Speech Lang. | 4 |
| 2026 | DDSE: Efficient Neural Codec Language Models for speech enhancement with disentangled representations
Qinwen Hu, Xiaobin Rong, Mansur Yesilbursa, Kamil Wójcicki |
Speech Commun. | 2 |
| 2026 | CoFi-Lite: Pushing the Limits of Ultra-Lightweight Speech Enhancement
Leyan Yang, Dahan Wang, Xiaobin Rong, Jiadong Zhao |
IEEE Signal Process. Lett. | 3 |
| 2025 | TS-URGENet: A Three-stage Universal Robust and Generalizable Speech Enhancement Network
Xiaobin Rong, Dahan Wang, Qinwen Hu |
INTERSPEECH | 1 |
| 2025 | A Lightweight Hybrid Dual Channel Speech Enhancement System under Low-SNR Conditions
Xiaobin Rong, Tianchi Sun |
INTERSPEECH | 2 |
| 2024 | GTCRN: A Speech Enhancement Model Requiring Ultralow Computational ResourcesabstractWhile modern deep learning-based models have significantly outperformed traditional methods in the area of speech enhancement, they often necessitate a lot of parameters and extensive computational power, making them impractical to be deployed on edge devices in real-world applications. In this paper, we introduce Grouped Temporal Convolutional Recurrent Network (GTCRN), which incorporates grouped strategies to efficiently simplify a competitive model, DPCRN. Additionally, it leverages subband feature extraction modules and temporal recurrent attention modules to enhance its performance. Remarkably, the resulting model demands ultralow computational resources, featuring only 23.7 K parameters and 39.6 MMACs per second. Experimental results show that our proposed model not only surpasses RNNoise, a typical lightweight model with similar computational burden, but also achieves competitive performance when compared to recent baseline models with significantly higher computational resources requirements. Xiaobin Rong, Tianchi Sun, Changbao Zhu |
ICASSP | 1 |
| 2023 | A Low-Latency Hybrid Multi-Channel Speech Enhancement System For Hearing AidsabstractThis paper summarizes a hybrid multi-channel speech enhancement system for the ICASSP Signal Processing Grand Challenge: Clarity Challenge (Speech Enhancement for Hearing Aids) 2023. The system consists of a rule-based dereverberation module, a multi-channel enhancement module, and a post-processing module. Without using the head rotation information and the enrollment speech, the system can reach an average hearing aid speech perception index (HASPI) score of 0.696 and hearing aid speech quality index (HASQI) score of 0.320 on the official development set. The corresponding scores are 0.729 and 0.316 respectively on the Eval1 set for the challenge ranking. Zhongshu Hou, Wanyu Yang, Tianchi Sun, Xiaobin Rong, Dahan Wang, Kai Chen 0029 |
ICASSP | 6 |