Xiaohuai Le

dblp:297/2930 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-6419-1825ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021
YearPublicationVenuePosition
2025 FAF-Filt: Frequency-aware Fourier Filter for Sound Event Detection
abstract
Capturing time-frequency patterns along the frequency axis, is crucial for the precision of sound event detection systems. Frequency dynamic convolution (FDY) and a series of its variants which incorporate frequency-adaptive kernels in standard 2D convolutions, have demonstrated remarkable performance, yet also suffered from high computational costs. To address the issue, we propose an efficient and light-weighted frequency-aware Fourier filter (FAF-Filt), which performs a 2D Fourier transform on features to the frequency domain and employs a learnable frequency-aware filter to process the transformed features, thereby integrating global information more effectively to extract decisive frequency components. In addition, frequency-adaptive convolution (FA-Conv) is adopted to further strengthen the representative ability of convolution, which incorporates the frequency-aware attention mechanism into the inputs and outputs of the convolutions. Experimental results exhibit superiority of the proposed method, achieving comparable performance with FDY-CRNN in terms of polyphonic sound event scores (PSDS) with a significantly 56% reduction in parameters.
Xiaohuai Le, Zhuangqi Chen, Xianjun Xia, Chuanzeng Huang
ICASSP2
2025 AF-Vocoder: Artifact-Free Neural Vocoder with Global Artifact Filter
Zhuangqi Chen, Xianjun Xia, Xiaohuai Le, Chuanzeng Huang
INTERSPEECH3
2025 Multistage Universal Speech Enhancement System for URGENT Challenge
Xiaohuai Le, Zhuangqi Chen, Xianjun Xia, Chuanzeng Huang
INTERSPEECH1
2025 CBA-Whisper: Curriculum Learning-Based AdaLoRA Fine-Tuning on Whisper for Low-Resource Dysarthric Speech Recognition
Tianyi Tan, Xiaohuai Le, Wenzhi Fan, Xianjun Xia, Chuanzeng Huang
INTERSPEECH3
2024 A Light-Weight State Detection Model for Kalman-Filter-Based Acoustic Feedback Cancellation with Rapid Recovery from Abrupt Path Changes
abstract
The partitioned block frequency domain Kalman Filter (PBFDKF) has been applied in acoustic feedback cancellation (AFC) due to its fast convergence and low steady-state misalignment. However, in cases where the feedback path experiences abrupt changes, the Kalman filter, once it reaches a steady state, might encounter the issue of deadlock and exhibit suboptimal tracking capabilities. In this paper, the Kalman filter with a light-weight state detection model (KF-SD) is proposed to effectively improve the robustness of AFC against abrupt path changes. The feedback return loss enhancement (FRLE) is proposed as the input to a state detection model with only 789 parameters to track the abrupt feedback path changes, and the state detection results are merged into the Kalman filter for a better re-convergence performance. A refined training label is proposed to ensure the robustness of model. Experimental results illustrate the superior performance of the proposed KF-SD algorithm, showcasing a high true positive rate, a low false alarm rate, and a short state detection latency. These advantages lead to faster re-convergence and enhanced sound quality when compared to the commonly used shadow filter strategy.
Haocheng Guo, Xiaohuai Le, Kai Chen 0029
ICASSP2
2023 Personalized Speech Enhancement Combining Band-Split RNN and Speaker Attentive Module
abstract
Target speaker information can be utilized in speech enhancement (SE) models to more effectively extract the desired speech. Previous works introduce the speaker embedding into speech enhancement models by means of concatenation or affine transformation. In this paper, we propose a speaker attentive module to calculate the attention scores between the speaker embedding and the intermediate features, which are used to rescale the features. By merging this module in the state-of-the-art SE model, we construct the personalized SE model for ICASSP Signal Processing Grand Challenge: DNS Challenge 5 (2023). Our system achieves a final score of 0.529 on the blind test set of track1 and 0.549 on track2.
Xiaohuai Le, Yiqing Guo, Xianjun Xia
ICASSP1
2023 Harmonic enhancement using learnable comb filter for light-weight full-band speech enhancement model
Xiaohuai Le, Yiqing Guo, Xianjun Xia, Hua Gao, Yijian Xiao, Piao Ding, Shenyi Song
INTERSPEECH1
2022 Inference Skipping for More Efficient Real-Time Speech Enhancement With Parallel RNNs
abstract
Deep neural network (DNN) based speech enhancement models have attracted extensive attention due to their promising performance. However, it is difficult to deploy a powerful DNN in real-time applications because of its high computational cost. Typical compression methods such as pruning and quantization do not make good use of the data characteristics. In this paper, we introduce the Skip-RNN strategy into speech enhancement models with parallel RNNs. The states of the RNNs update intermittently without interrupting the update of the output mask, which leads to significant reduction of computational load without evident audio artifacts. To better leverage the difference between the voice and the noise, we further regularize the skipping strategy with voice activity detection (VAD) guidance, saving more computational load. Experiments on a high-performance speech enhancement model, dual-path convolutional recurrent network (DPCRN), show the superiority of our strategy over strategies like network pruning or directly training a smaller model. We also validate the generalization of the proposed strategy on two other competitive speech enhancement models.
Xiaohuai Le, Kai Chen 0029
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 DPCRN: Dual-Path Convolution Recurrent Network for Single Channel Speech Enhancement
abstract
The dual-path RNN (DPRNN) was proposed to more effectively model extremely long sequences for speech separation in the time domain.By splitting long sequences to smaller chunks and applying intra-chunk and inter-chunk RNNs, the DPRNN reached promising performance in speech separation with a limited model size.In this paper, we combine the DPRNN module with Convolution Recurrent Network (CRN) and design a model called Dual-Path Convolution Recurrent Network (DPCRN) for speech enhancement in the time-frequency domain.We replace the RNNs in the CRN with DPRNN modules, where the intra-chunk RNNs are used to model the spectrum pattern in a single frame and the inter-chunk RNNs are used to model the dependence between consecutive frames.With only 0.8M parameters, the submitted DPCRN model achieves an overall mean opinion score (MOS) of 3.57 in the wide band scenario track of the Interspeech 2021 Deep Noise Suppression (DNS) challenge.Evaluations on some other test sets also show the efficacy of our model.
Xiaohuai Le, Kai Chen 0029
Interspeech1