VLDB 2026 Research / reviewers in the wild / expert
Changbao Zhu
dblp:166/6494
· DBLP profile ↗
6ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0002-4438-8102ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking Flow and Diffusion Bridge Models for Speech EnhancementabstractFlow matching and diffusion bridge models have emerged as leading paradigms in generative speech enhancement, modeling stochastic processes between paired noisy and clean speech signals based on principles such as flow matching, score matching, and Schrödinger bridge. In this paper, we present a framework that unifies existing flow and diffusion bridge models by interpreting them as constructions of Gaussian probability paths with varying means and variances between paired data. Furthermore, we investigate the underlying consistency between the training/inference procedures of these generative models and conventional predictive models. Our analysis reveals that each sampling step of a well-trained flow or diffusion bridge model optimized with a data prediction loss is theoretically analogous to executing predictive speech enhancement. Motivated by this insight, we introduce an enhanced bridge model that integrates an effective probability path design with key elements from predictive paradigms, including improved network architecture, tailored loss functions, and optimized training strategies. Experiments on denoising and dereverberation tasks demonstrate that the proposed method outperforms existing flow and diffusion baselines with fewer parameters and reduced computational complexity. The results also highlight that the inherently predictive nature of this generative framework imposes limitations on its achievable upper-bound performance. Dahan Wang, Changbao Zhu, Kai Chen 0029 |
AAAI | 5 |
| 2024 | GTCRN: A Speech Enhancement Model Requiring Ultralow Computational ResourcesabstractWhile modern deep learning-based models have significantly outperformed traditional methods in the area of speech enhancement, they often necessitate a lot of parameters and extensive computational power, making them impractical to be deployed on edge devices in real-world applications. In this paper, we introduce Grouped Temporal Convolutional Recurrent Network (GTCRN), which incorporates grouped strategies to efficiently simplify a competitive model, DPCRN. Additionally, it leverages subband feature extraction modules and temporal recurrent attention modules to enhance its performance. Remarkably, the resulting model demands ultralow computational resources, featuring only 23.7 K parameters and 39.6 MMACs per second. Experimental results show that our proposed model not only surpasses RNNoise, a typical lightweight model with similar computational burden, but also achieves competitive performance when compared to recent baseline models with significantly higher computational resources requirements. Xiaobin Rong, Tianchi Sun, Changbao Zhu |
ICASSP | 5 |
| 2024 | A Lightweight Hybrid Multi-Channel Speech Extraction System with Directional Voice Activity DetectionabstractAlthough deep learning (DL) based end-to-end models have shown outstanding performance in multi-channel speech extraction, their practical applications on edge devices are restricted due to their high computational complexity. In this paper, we propose a hybrid system that can more effectively integrate the generalized sidelobe canceller (GSC) and a lightweight post-filtering model under the assistance of spatial speaker activity information provided by a directional voice activity detection (DVAD) module. In addition to guiding the update of the adaptive blocking matrix (ABM) and the adaptive interference canceller (AIC) used in GSC to alleviate the distortion of the desired speech, DVAD is also utilized as an auxiliary input to the post-filtering model to enhance its capability of interference suppression. The experimental results demonstrate that, with much lower computational costs, our method can achieve comparable performance with a current state-of-the-art end-to-end model on simulated data and generalize even better on real-world data. Tianchi Sun, Changbao Zhu |
ICASSP | 5 |
| 2023 | Convolutional Recurrent MetriCGAN With Spectral Dimension Compression For Full-Band Speech EnhancementabstractMetricGAN and its variations have been proven to be an effective wide-band speech enhancement model. In this paper, we expand it to full-band enhancement by combining our recently proposed learnable spectral dimension compression mapping strategy. The encoder-decoder structure with a time-frequency convolutional recurrent network is utilized as the generator. The proposed model is submitted to the ICASSP Signal Processing Grand Challenge: DNS-5 Challenge (2023). Without using the enrollment speech, it obtains a final score of 0.548 on Track-1 and 0.559 on Track-2. Zhongshu Hou, Qinwen Hu, Tianchi Sun, Changbao Zhu |
ICASSP | 5 |
| 2023 | TrimTail: Low-Latency Streaming ASR with Simple But Effective Spectrogram-Level Length PenaltyabstractIn this paper, we present TrimTail, a simple but effective emission regularization method to improve the latency of streaming ASR models. The core idea of TrimTail is to apply length penalty (i.e., by trimming trailing frames, see Fig. 1-(b)) directly on the spectrogram of input utterances, which does not require any alignment. We demonstrate that TrimTail is computationally cheap and can be applied online and optimized with any training loss or any model architecture on any dataset without any extra effort by applying it on various end-to-end streaming ASR networks either trained with CTC loss [1] or Transducer loss [2]. We achieve 100 ~ 200ms latency reduction with equal or even better accuracy on both Aishell-1 and Librispeech. Moreover, by using TrimTail, we can achieve a 400ms algorithmic improvement of User Sensitive Delay (USD) with an accuracy loss of less than 0.2. Xingchen Song, Di Wu 0061, Zhiyong Wu 0001, Yuekai Zhang, Zhendong Peng, Wenpeng Li, Fuping Pan, Changbao Zhu |
ICASSP | 9 |
| 2015 | Overview of the EVS codec architectureabstractThe recently standardized 3GPP codec for Enhanced Voice Services (EVS) offers new features and improvements for low-delay real-time communication systems. Based on a novel, switched low-delay speech/audio codec, the EVS codec contains various tools for better compression efficiency and higher quality for clean/noisy speech, mixed content and music, including support for wideband, super-wideband and full-band content. The EVS codec operates in a broad range of bitrates, is highly robust against packet loss and provides an AMR-WB interoperable mode for compatibility with existing systems. This paper gives an overview of the underlying architecture as well as the novel technologies in the EVS codec and presents listening test results showing the performance of the new codec in terms of compression and speech/audio quality. Martin Dietz, Markus Multrus, Vaclav Eksler, Vladimir Malenovsky, Erik Norvell, Harald Pobloth, Lei Miao 0004, Lasse Laaksonen, Adriana Vasilache, Yutaka Kamamoto, Kei Kikuiri, Stéphane Ragot, Julien Faure, Hiroyuki Ehara, Vivek Rajendran, Venkatraman Atti, Hosang Sung, Eunmi Oh, Changbao Zhu |
ICASSP | 21 |