Kai Chen 0029

dblp:181/2839-29 · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 9 since 2021Artificial intelligence and machine learning · 7 · 4 since 2021
YearPublicationVenuePosition
2026 Rethinking Flow and Diffusion Bridge Models for Speech Enhancement
abstract
Flow matching and diffusion bridge models have emerged as leading paradigms in generative speech enhancement, modeling stochastic processes between paired noisy and clean speech signals based on principles such as flow matching, score matching, and Schrödinger bridge. In this paper, we present a framework that unifies existing flow and diffusion bridge models by interpreting them as constructions of Gaussian probability paths with varying means and variances between paired data. Furthermore, we investigate the underlying consistency between the training/inference procedures of these generative models and conventional predictive models. Our analysis reveals that each sampling step of a well-trained flow or diffusion bridge model optimized with a data prediction loss is theoretically analogous to executing predictive speech enhancement. Motivated by this insight, we introduce an enhanced bridge model that integrates an effective probability path design with key elements from predictive paradigms, including improved network architecture, tailored loss functions, and optimized training strategies. Experiments on denoising and dereverberation tasks demonstrate that the proposed method outperforms existing flow and diffusion baselines with fewer parameters and reduced computational complexity. The results also highlight that the inherently predictive nature of this generative framework imposes limitations on its achievable upper-bound performance.
Dahan Wang, Changbao Zhu, Kai Chen 0029
AAAI6
2025 DistillW2N: A Lightweight One-Shot Whisper to Normal Voice Conversion Model Using Distillation of Self-Supervised Features
abstract
Whisper to Normal voice conversion (W2N) holds great promise for assistive communication and healthcare, making it an exciting area of research and development. Recent advancements in W2N are predominantly driven by self-supervised speech representation learning (SSL) techniques. While effective, SSL requires extensive parameters and high computational costs, making it impractical to be deployed in real-world applications. We propose DistillW2N, a lightweight one-shot W2N model. DistillW2N consists of a speech-to-unit (S2U) encoder, which seeks to close the gap between whispered and normal content units by distilling HuBERT-Soft representations from both normal speech and pseudo-whisper, and a unit-to-speech (U2S) decoder, which incorporates content units with timbre units by Style-Adaptive Layer Normalization (SALN) and leverages SoundStream decoder for lightweight high-quality speech synthesis. Moreover, we discover that the S2U encoder is able to learn from a VAD model to trim noise. Experiments show that DistillW2N significantly improves the intelligibility for whisper while preserving speaker similarity compared to prevailing SSL approaches. The resulting model demands only 10.91 M parameters and 1.63 GMACs per second and runs approximately 5 times faster than QuickVC. All samples and code are available at https://github.com/tan90xx/distillw2n.
Tianyi Tan, Haoxin Ruan, Kai Chen 0029
ICASSP4
2025 Leveraging Self-Supervised Learning Based Speaker Diarization for MISP 2025 AVSD Challenge
Zeyan Song, Tianchi Sun, Ronghui Hu, Kai Chen 0029
INTERSPEECH4
2025 Deep Learning-Based Approach for Identification and Compensation of Nonlinear Distortions in Parametric Array Loudspeakers
abstract
Compared to traditional electrodynamic loudspeakers, the parametric array loudspeaker (PAL) offers exceptional directivity for audio applications but suffers from significant nonlinear distortions due to its inherent intricate demodulation process. The Volterra filter-based approaches have been widely used to reduce these distortions, but the effectiveness is limited by its inverse filter's capability. Specifically, its$p$th-order inverse filter can only compensate for nonlinearities up to the$p$th order, while the higher-order nonlinearities it introduces continue to generate lower-order harmonics. In contrast, this paper introduces the modern deep learning methods for the first time to address nonlinear identification and compensation for PAL systems. Specifically, the WaveNet neural network, recognized for its success in audio nonlinear system modeling, is utilized to identify and compensate for distortions in a double sideband amplitude modulation-based PAL system. Experimental measurements from 250 Hz to 8 kHz demonstrate that our proposed approach significantly reduces both total harmonic distortion and intermodulation distortion of audio sound generated by PALs, achieving average reductions to 3.11% and 0.93%, respectively. This performance is notably superior to results obtained using the current state-of-the-art Volterra filter-based methods. Our work opens new possibilities for improving the sound reproduction performance of PALs.
Mengtong Li, Kai Chen 0029
IEEE Signal Process. Lett.3
2024 A Light-Weight State Detection Model for Kalman-Filter-Based Acoustic Feedback Cancellation with Rapid Recovery from Abrupt Path Changes
abstract
The partitioned block frequency domain Kalman Filter (PBFDKF) has been applied in acoustic feedback cancellation (AFC) due to its fast convergence and low steady-state misalignment. However, in cases where the feedback path experiences abrupt changes, the Kalman filter, once it reaches a steady state, might encounter the issue of deadlock and exhibit suboptimal tracking capabilities. In this paper, the Kalman filter with a light-weight state detection model (KF-SD) is proposed to effectively improve the robustness of AFC against abrupt path changes. The feedback return loss enhancement (FRLE) is proposed as the input to a state detection model with only 789 parameters to track the abrupt feedback path changes, and the state detection results are merged into the Kalman filter for a better re-convergence performance. A refined training label is proposed to ensure the robustness of model. Experimental results illustrate the superior performance of the proposed KF-SD algorithm, showcasing a high true positive rate, a low false alarm rate, and a short state detection latency. These advantages lead to faster re-convergence and enhanced sound quality when compared to the commonly used shadow filter strategy.
Haocheng Guo, Xiaohuai Le, Kai Chen 0029
ICASSP3
2023 A Low-Latency Hybrid Multi-Channel Speech Enhancement System For Hearing Aids
abstract
This paper summarizes a hybrid multi-channel speech enhancement system for the ICASSP Signal Processing Grand Challenge: Clarity Challenge (Speech Enhancement for Hearing Aids) 2023. The system consists of a rule-based dereverberation module, a multi-channel enhancement module, and a post-processing module. Without using the head rotation information and the enrollment speech, the system can reach an average hearing aid speech perception index (HASPI) score of 0.696 and hearing aid speech quality index (HASQI) score of 0.320 on the official development set. The corresponding scores are 0.729 and 0.316 respectively on the Eval1 set for the challenge ranking.
Zhongshu Hou, Wanyu Yang, Tianchi Sun, Xiaobin Rong, Dahan Wang, Kai Chen 0029
ICASSP8
2022 A Priori SNR Estimation for Speech Enhancement Based on PESQ-Induced Reinforcement Learning
abstract
Perceptual evaluation of speech quality (PESQ) is widely accepted as an effective objective metric closely related to the speech quality sensed by human listening perception. Due to its evaluation complexity and non-differentiability, PESQ is difficult to include in the cost function for deep learning-based speech enhancement. In this paper, we focus on introducing PESQ to improve Deep Xi, a recently proposed minimum mean square error (MMSE) based speech enhancement with a priori signal-to-ratio (SNR) estimated by a deep neural network. Regarding discrete a priori SNR as actions, we apply reinforcement learning (RL) to select the optimal SNR at the frame level through the reward function associated with PESQ. The experimental results show that the RL-trained network is able to achieve a better PESQ score, especially in low SNR conditions.
Haoxin Ruan, Kai Chen 0029
ICASSP3
2022 An Explicit Connection Between Independent Vector Analysis and Tensor Decomposition in Blind Source Separation
abstract
Independent vector analysis (IVA) and tensor decomposition are two types of effective algorithms for joint blind source separation (JBSS) with different statistical assumptions. Although IVA and tensor decomposition are intrinsically linked, their explicit connection has not been reported. In this letter, we reveal their explicit connection through a piecewise stationary multivariate complex Gaussian signal model. With this model, IVA can be explained as reconstructing the covariances of the mixtures in a similar manner as double coupled canonical polyadic decomposition (DC-CPD), a typical tensor-based algorithm, with the only difference being the distance metric used in the cost function. Numerical experiments show that IVA can achieve better separation performance but is highly dependent on how well thea priorimodel matches the actual signal, while DC-CPD is more robust to the model mismatch.
Haoxin Ruan, Kai Chen 0029
IEEE Signal Process. Lett.3
2022 Inference Skipping for More Efficient Real-Time Speech Enhancement With Parallel RNNs
abstract
Deep neural network (DNN) based speech enhancement models have attracted extensive attention due to their promising performance. However, it is difficult to deploy a powerful DNN in real-time applications because of its high computational cost. Typical compression methods such as pruning and quantization do not make good use of the data characteristics. In this paper, we introduce the Skip-RNN strategy into speech enhancement models with parallel RNNs. The states of the RNNs update intermittently without interrupting the update of the output mask, which leads to significant reduction of computational load without evident audio artifacts. To better leverage the difference between the voice and the noise, we further regularize the skipping strategy with voice activity detection (VAD) guidance, saving more computational load. Experiments on a high-performance speech enhancement model, dual-path convolutional recurrent network (DPCRN), show the superiority of our strategy over strategies like network pruning or directly training a smaller model. We also validate the generalization of the proposed strategy on two other competitive speech enhancement models.
Xiaohuai Le, Kai Chen 0029
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 DPCRN: Dual-Path Convolution Recurrent Network for Single Channel Speech Enhancement
abstract
The dual-path RNN (DPRNN) was proposed to more effectively model extremely long sequences for speech separation in the time domain.By splitting long sequences to smaller chunks and applying intra-chunk and inter-chunk RNNs, the DPRNN reached promising performance in speech separation with a limited model size.In this paper, we combine the DPRNN module with Convolution Recurrent Network (CRN) and design a model called Dual-Path Convolution Recurrent Network (DPCRN) for speech enhancement in the time-frequency domain.We replace the RNNs in the CRN with DPRNN modules, where the intra-chunk RNNs are used to model the spectrum pattern in a single frame and the inter-chunk RNNs are used to model the dependence between consecutive frames.With only 0.8M parameters, the submitted DPCRN model achieves an overall mean opinion score (MOS) of 3.57 in the wide band scenario track of the Interspeech 2021 Deep Noise Suppression (DNS) challenge.Evaluations on some other test sets also show the efficacy of our model.
Xiaohuai Le, Kai Chen 0029
Interspeech3
2020 Nonlinear Residual Echo Suppression Based on Multi-Stream Conv-TasNet
abstract
Acoustic echo cannot be entirely removed by linear adaptive filters due to the nonlinear relationship between the echo and far-end signal. Usually a post processing module is required to further suppress the echo. In this paper, we propose a residual echo suppression method based on the modification of fully convolutional time-domain audio separation network (Conv-TasNet). Both the residual signal of the linear acoustic echo cancellation system, and the output of the adaptive filter are adopted to form multiple streams for the Conv-TasNet, resulting in more effective echo suppression while keeping a lower latency of the whole system. Simulation results validate the efficacy of the proposed method in both single-talk and double-talk situations.
Teng Xiang, Kai Chen 0029
INTERSPEECH3
2020 U-Net Based Direct-Path Dominance Test for Robust Direction-of-Arrival Estimation
abstract
It has been noted that the identification of the time-frequency bins dominated by the contribution from the direct propagation of the target speaker can significantly improve the robustness of the direction-of-arrival estimation.However, the correct extraction of the direct-path sound is challenging especially in adverse environments.In this paper, a U-net based direct-path dominance test method is proposed.Exploiting the efficient segmentation capability of the U-net architecture, the directpath information can be effectively retrieved from a dedicated multi-task neural network.Moreover, the training and inference of the neural network only need the input of a single microphone, circumventing the problem of array-structure dependence faced by common end-to-end deep learning based methods.Simulations demonstrate that significantly higher estimation accuracy can be achieved in high reverberant and low signal-to-noise ratio environments.
Kai Chen 0029
INTERSPEECH2
2019 Speech Separation Using Independent Vector Analysis with an Amplitude Variable Gaussian Mixture Model
Zhaoyi Gu, Kai Chen 0029
INTERSPEECH3
2019 Effective Improvement of Under-Modeling Frequency-Domain Kalman Filter
abstract
The frequency-domain Kalman filter (FKF) has been utilized in many audio signal processing applications due to its fast convergence speed and robustness. However, the performance of the FKF in under-modeling situations has not been investigated. This letter presents an analysis of the steady-state behavior of the commonly used diagonalized FKF and reveals that it suffers from a biased solution in under-modeling scenarios. An effective improvement of the FKF is proposed, having the benefits of the guaranteed optimal steady-state behavior at the cost of a very limited increase of computational burden. The convergence behavior of the proposed algorithm is also analyzed. Computer simulations are conducted to validate the improved performance of the proposed method.
Wenzhi Fan, Kai Chen 0029, Jiancheng Tao
IEEE Signal Process. Lett.2
2017 Convergence analysis of the modified frequency-domain block LMS algorithm with guaranteed optimal steady state performance
Kai Chen 0029, Xiaojun Qiu
Signal Process.2