Longbiao Cheng

dblp:284/2790 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0002-0635-1480ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Modulating State Space Model with SlowFast Framework for Compute-Efficient Ultra Low-Latency Speech Enhancement
abstract
Deep learning-based speech enhancement (SE) methods often face significant computational challenges when needing to meet low-latency requirements because of the increased number of frames to be processed. This paper introduces the SlowFast framework which aims to reduce computation costs specifically when low-latency enhancement is needed. The framework consists of a slow branch that analyzes the acoustic environment at a low frame rate, and a fast branch that performs SE in the time domain at the needed higher frame rate to match the required latency. Specifically, the fast branch employs a state space model where its state transition process is dynamically modulated by the slow branch. Experiments on a SE task with a 2 ms algorithmic latency requirement using the Voice Bank + Demand dataset show that our approach reduces computation cost by 70% compared to a baseline single-branch network with equivalent parameters, without compromising enhancement performance. Furthermore, by leveraging the SlowFast framework, we implemented a network that achieves an algorithmic latency of just 62.5 μs (one sample point at 16 kHz sample rate) with a computation cost of 100 M MACs/s, while scoring a PESQ-NB of 3.12 and SISNR of 16.62.
Longbiao Cheng, Ashutosh Pandey 0004, Buye Xu, Tobi Delbruck, Vamsi K. Ithapu, Shih-Chii Liu
ICASSP1
2025 CustomFit: Customizing Task-Optimized and Size-Specific Subnets from a Single Base Network During Runtime
abstract
We address the challenge of efficiently deploying networks across hardware platforms with varying resource constraints and task requirements. While conventional methods train a base network that can be split into subnets of different sizes for deployment, they typically focus on a single task. This paper introduces CustomFit, a framework that trains a base network from which different-sized subnets tailored to multiple tasks can be segmented during inference. CustomFit uses task-specific feature mapping to optimize the base network for each task without altering its weights. Subnet segmentation is guided by the task’s feature mapping process and the desired size. By training a base network with CustomFit on Speech Enhancement (SE) and Spoken Language Understanding (SLU) tasks, we show that 30 subnets of varying sizes can be segmented. Compared to state-of-the-art dynamic networks, CustomFit achieves similar SE scores and improves SLU accuracy by 2.75%, while using 2.9× fewer parameters in the base network. The parameter sizes of task-specific subnets range from 11k to 720k, enabling flexible deployment across a range of edge platforms from microcontroller units to mobile phones.
Longbiao Cheng, Shih-Chii Liu
ISCAS1
2024 Exploiting Symmetric Temporally Sparse BPTT for Efficient RNN Training
abstract
Recurrent Neural Networks (RNNs) are useful in temporal sequence tasks. However, training RNNs involves dense matrix multiplications which require hardware that can support a large number of arithmetic operations and memory accesses. Implementing online training of RNNs on the edge calls for optimized algorithms for an efficient deployment on hardware. Inspired by the spiking neuron model, the Delta RNN exploits temporal sparsity during inference by skipping over the update of hidden states from those inactivated neurons whose change of activation across two timesteps is below a defined threshold. This work describes a training algorithm for Delta RNNs that exploits temporal sparsity in the backward propagation phase to reduce computational requirements for training on the edge. Due to the symmetric computation graphs of forward and backward propagation during training, the gradient computation of inactivated neurons can be skipped. Results show a reduction of ∼80% in matrix operations for training a 56k parameter Delta LSTM on the Fluent Speech Commands dataset with negligible accuracy loss. Logic simulations of a hardware accelerator designed for the training algorithm show 2-10X speedup in matrix computations for an activation sparsity range of 50%-90%. Additionally, we show that the proposed Delta RNN training will be useful for online incremental learning on edge devices with limited computing resources.
Chang Gao 0002, Zuowen Wang, Longbiao Cheng, Shih-Chii Liu, Tobi Delbruck
AAAI4
2024 Regularized Parameter Uncertainty for Improving Generalization in Reinforcement Learning
abstract
In order for reinforcement learning (RL) agents to be deployed in real-world environments, they must be able to generalize to unseen environments. However, RL struggles with out-of-distribution generalization, often due to overfitting the particulars of the training environment. Although regularization techniques from supervised learning can be applied to avoid over-fitting, the differences between supervised learning and RL limit their application. To address this, we propose the Signal-to-Noise Ratio regulated Parameter Uncertainty Network (SNR PUN) for RL. We introduce SNR as a new measure of regularizing the parameter uncertainty of a network and provide a formal analysis explaining why SNR regularization works well for RL. We demonstrate the effectiveness of our proposed method to generalize in several simulated environments; and in a physical system showing the possibility of using SNR PUN for applying RL to real-world applications.
Pehuen Moure, Longbiao Cheng, Joachim Ott, Zuowen Wang, Shih-Chii Liu
CVPR2
2024 Dynamic Gated Recurrent Neural Network for Compute-efficient Speech Enhancement
abstract
This paper introduces a new Dynamic Gated Recurrent Neural Network (DG-RNN) for compute-efficient speech enhancement models running on resource-constrained hardware platforms.It leverages the slow evolution characteristic of RNN hidden states over steps, and updates only a selected set of neurons at each step by adding a newly proposed select gate to the RNN model.This select gate allows the computation cost of the conventional RNN to be reduced during network inference.As a realization of the DG-RNN, we further propose the Dynamic Gated Recurrent Unit (D-GRU) which does not require additional parameters.Test results obtained from several state-ofthe-art compute-efficient RNN-based speech enhancement architectures using the DNS challenge dataset, show that the D-GRU based model variants maintain similar speech intelligibility and quality metrics comparable to the baseline GRU based models even with an average 50% reduction in GRU computes.
Longbiao Cheng, Ashutosh Pandey 0004, Buye Xu, Tobi Delbruck, Shih-Chii Liu
INTERSPEECH1
2024 DeltaDEQ: Exploiting Heterogeneous Convergence for Accelerating Deep Equilibrium Iterations
abstract
Implicit neural networks including deep equilibrium models have achieved superior task performance with better parameter efficiency in various applications. However, it is often at the expense of higher computation costs during inference. In this work, we identify a phenomenon named $\textbf{heterogeneous convergence}$ that exists in deep equilibrium models and other iterative methods. We observe much faster convergence of state activations in certain dimensions therefore indicating the dimensionality of the underlying dynamics of the forward pass is much lower than the defined dimension of the states. We thereby propose to exploit heterogeneous convergence by storing past linear operation results (e.g., fully connected and convolutional layers) and only propagating the state activation when its change exceeds a threshold. Thus, for the already converged dimensions, the computations can be skipped. We verified our findings and reached 84\% FLOPs reduction on the implicit neural representation task, 73\% on the Sintel and 76\% on the KITTI datasets for the optical flow estimation task while keeping comparable task accuracy with the models that perform the full update.
Zuowen Wang, Longbiao Cheng, Pehuen Moure, Niklas Hahn 0002, Shih-Chii Liu
NeurIPS2
2023 The effect of source sparsity on independent vector analysis for blind source separation
Jianjun Gu 0005, Longbiao Cheng, Dingding Yao, Yonghong Yan 0002
Signal Process.2
2022 A Secondary Path-Decoupled Active Noise Control Algorithm Based on Deep Learning
abstract
Active noise control (ANC) systems are widely used to cancel unwanted noise. However, for high-level noise, the residual error signal cannot be fully eliminated because of the nonlinearity of the secondary path, resulting in the diverging of the adaptive filter. In this letter, we propose a secondary path-decoupled ANC (SPD-ANC) algorithm based on deep learning. Specifically, the secondary path decoupled module consisting of two time-domain convolutional recurrent networks, one for modeling the nonlinear secondary path and the other for modeling the reverse process, is employed to calculate the secondary path-decoupled (SPD) error signal. The control signal is then generated by an adaptive filter that is optimized towards minimizing the SPD error signal. Simulation results indicate that the proposed method outperforms the conventional ANC methods under different conditions.
Daocheng Chen, Longbiao Cheng, Dingding Yao, Yonghong Yan 0002
IEEE Signal Process. Lett.2
2021 Residual Echo and Noise Cancellation with Feature Attention Module and Multi-Domain Loss Function
Jianjun Gu 0005, Longbiao Cheng, Xingwei Sun, Yonghong Yan 0002
Interspeech2
2021 FSCNet: Feature-Specific Convolution Neural Network for Real-Time Speech Enhancement
abstract
In recent years, convolutional neural networks (CNNs) have been widely exploited in deep neural network (DNN)-based speech enhancement methods. However, the representation power of CNNs for speech modeling is limited because of the spatial-agnostic convolution kernels. This letter proposes a novel feature-specific convolution neural network (FSCNet) for real-time speech enhancement. In FSCNet, the encoder and decoder are adopted for forward and inverse feature space transformation, respectively. The denoising module based on the feature-specific convolution (FSC) is employed to enhance the generated deep features. Leveraging the long-term global contexts and considering the importance of each feature channel for speech modeling, the convolution kernels of FSC are dynamically parameterized in each time-frequency location. A function-constrained loss is further proposed to train the FSCNet, ensuring the encoder, denoising modules and decoder can function as expected. Experimental results show that the proposed FSCNet outperforms the state-of-the-art denoising algorithms in terms of five objective evaluation metrics and model size.
Longbiao Cheng, Yonghong Yan 0002
IEEE Signal Process. Lett.1
2021 Estimation Reliability Function Assisted Sound Source Localization With Enhanced Steering Vector Phase Difference
abstract
The performance of the traditional direction-of-arrival (DOA) estimation algorithms greatly degrades in noisy and reverberant environments. Recently, deep learning has been applied to sound source localization and provided the substantial improvement in robustness for DOA estimation. In this paper, we propose a sound source localization approach using the deep learning-based steering vector phase difference enhancement. The steering vectors and their estimation reliability functions (ERFs) are first estimated under the guidance of the time-frequency masks that are predicted using deep neural network (DNN). The phase difference of the steering vectors is further enhanced with a second DNN model, which is trained with the ERF-weighted mean square error (MSE) loss. The DOA of the sound source is finally determined by the ERF-weighted histogram analysis. Experimental results with various types and levels of noise and various reverberant conditions show that the proposed approach outperforms the state-of-the-art sound source localization algorithms in utterance and frame-level DOA estimation.
Longbiao Cheng, Xingwei Sun, Dingding Yao, Yonghong Yan 0002
IEEE ACM Trans. Audio Speech Lang. Process.1