Rong Chao

dblp:317/5074 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0001-6930-959XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Speech Enhancement with MAP-based Training for Robust ASR
abstract
To improve noise robustness in automatic speech recognition (ASR), a common strategy is to employ speech enhancement (SE) models as front-ends for ASR systems. However, SE models often introduce artifacts into enhanced signals, which can degrade ASR performance, particularly when the SE and ASR models are trained separately. Although various methods have been proposed to address this issue, they often come at the expense of increasing implementation complexity. Hence, this study proposes a maximum a posteriori (MAP) algorithm for training SE models by incorporating the posterior probability of clean speech, given the enhanced speech, into the loss function. Experimental results show that our method enhances the compatibility between SE and ASR models in both in-domain and out-of-domain testing scenarios, notably improving ASR performance. The proposed method does not require prior knowledge of ASR models or speech content during training or inference, nor does it involve additional post-processing steps.
You-Jin Li, Rong Chao, Borching Su, Yu Tsao 0001
ICASSP2
2025 MSEMG: Surface Electromyography Denoising with a Mamba-based Efficient Network
abstract
Surface electromyography (sEMG) recordings can be contaminated by electrocardiogram (ECG) signals when the monitored muscle is closed to the heart. Traditional signal processing-based approaches, such as high-pass filtering and template subtraction, have been used to remove ECG interference but are often limited in their effectiveness. Recently, neural network-based methods have shown greater promise for sEMG denoising, but they still struggle to balance both efficiency and effectiveness. In this study, we introduce MSEMG, a novel system that integrates the Mamba state space model with a convolutional neural network to serve as a lightweight sEMG denoising model. We evaluated MSEMG using sEMG data from the Non-Invasive Adaptive Prosthetics database and ECG signals from the MIT-BIH Normal Sinus Rhythm Database. The results show that MSEMG outperforms existing methods, generating higher-quality sEMG signals using fewer parameters.
Yu-Tung Liu, Kuan-Chen Wang, Rong Chao, Sabato Marco Siniscalchi, Ping-Cheng Yeh, Yu Tsao 0001
ICASSP3
2025 Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement
abstract
In multichannel speech enhancement, effectively capturing spatial and spectral information across different microphones is crucial for noise reduction. Traditional methods, such as CNN or LSTM, attempt to model the temporal dynamics of full-band and sub-band spectral and spatial features. However, these approaches face limitations in fully modeling complex temporal dependencies, especially in dynamic acoustic environments. To overcome these challenges, we modify the current advanced model McNet by introducing an improved version of Mamba, a state-space model, and further propose MCMamba. MCMamba has been completely reengineered to integrate full-band and narrow-band spatial information with sub-band and full-band spectral features, providing a more comprehensive approach to modeling spatial and spectral information. Our experimental results demonstrate that MCMamba significantly improves the modeling of spatial and spectral features in multichannel speech enhancement, outperforming McNet and achieving very promis- ing performance on the CHiME-3 dataset. Additionally, we find that Mamba performs exceptionally well in modeling spectral information.
Wenze Ren, Yi-Cheng Lin, Xuanjun Chen, Rong Chao, Kuo-Hsuan Hung, You-Jin Li, Wen-Yuan Ting, Hsin-Min Wang, Yu Tsao 0001
ICASSP5
2025 Universal Speech Enhancement with Regression and Generative Mamba
Rong Chao, Rauf Nasretdinov, Yu-Chiang Frank Wang, Ante Jukic, Szu-Wei Fu, Yu Tsao 0001
INTERSPEECH1
2024 Learning Efficient Interaction Anchor for HOI Detection
abstract
Human-object interaction (HOI) detection seeks complicated relationships between humans and objects, yet struggles persist in correctly associating multiple objects with a single human in complex interaction scenarios. In this paper, we tackle such an issue by introducing a novel interaction anchor that employs flexible strategies across different decoder layers with Barlow constraint and Interactivity-Instance Fusion. The proposed modules are both additive and easily implementable in existing approaches, offering computational efficiency within transformer-based models to compact cross-interactivity. Extensive experiments validate the effectiveness of our method, demonstrating comparable performance on HICO-DET and V-COCO for HOI detection.
Lirong Xue, Kang-Yang Huang, Rong Chao, Jhih-Ciang Wu, Hong-Han Shuai, Yung-Hui Li, Wen-Huang Cheng
ICME3
2024 An Investigation of Incorporating Mamba For Speech Enhancement
abstract
This work aims to investigate the use of a recently proposed, attention-free, scalable state-space model (SSM), Mamba, for the speech enhancement (SE) task. In particular, we employ Mamba to deploy different regression-based SE models (SEMamba) with different configurations, namely basic, advanced, causal, and non-causal. Furthermore, loss functions either based on signal-level distances or metric-oriented are considered. Experimental evidence shows that SEMamba attains a competitive PESQ of 3.55 on the VoiceBank-DEMAND dataset with the advanced, non-causal configuration. A new state-of-the-art PESQ of 3.69 is also reported when SEMamba is combined with Perceptual Contrast Stretching (PCS). Compared against Transformed-based equivalent SE solutions, a noticeable FLOPs reduction up to $\sim 12 \%$ is observed with the advanced non-causal configurations. Finally, SEMamba can be used as a pre-processing step before automatic speech recognition (ASR), showing competitive performance against recent SE solutions.
Rong Chao, Wen-Huang Cheng, Moreno La Quatra, Sabato Marco Siniscalchi, Chao-Han Huck Yang, Szu-Wei Fu, Yu Tsao 0001
SLT1
2024 An Explore-Exploit Workload-Bounded Strategy for Rare Event Detection in Massive Energy Sensor Time Series
abstract
With the rise of Internet-of-Things devices, the analysis of sensor-generated energy time series data has become increasingly important. This is especially crucial for detecting rare events like unusual electricity usage or water leakages in residential and commercial buildings, which is essential for optimizing energy efficiency and reducing costs. However, existing detection methods on large-scale data may fail to correctly detect rare events when they do not behave significantly differently from standard events or when their attributes are non-stationary. Additionally, the capacity of computational resources to analyze all time series data generated by an increasing number of sensors becomes a challenge. This situation creates an emergent demand for a workload-bounded strategy. To ensure both effectiveness and efficiency in detecting rare events in massive energy time series, we propose a heuristic-based framework called HALE . This framework utilizes an explore–exploit selection process that is specifically designed to recognize potential features of rare events in energy time series. HALE involves constructing an attribute-aware graph to preserve the attribute information of rare events. A heuristic-based random walk is then derived based on partial labels received at each time period to discover the non-stationarity of rare events. Potential rare event data are selected from the attribute-aware graph, and existing detection models are applied for final confirmation. Our study, which was conducted on three actual energy datasets, demonstrates that the HALE framework is both effective and efficient in its detection capabilities. This underscores its practicality in delivering cost-effective energy monitoring services.
Lo Pang-Yun Ting, Rong Chao, Chai-Shi Chang, Kun-Ta Chuang
ACM Trans. Intell. Syst. Technol.2
2022 Perceptual Contrast Stretching on Target Feature for Speech Enhancement
abstract
Speech enhancement (SE) performance has improved considerably owing to the use of deep learning models as a base function.Herein, we propose a perceptual contrast stretching (PCS) approach to further improve SE performance.The PCS is derived based on the critical band importance function and is applied to modify the targets of the SE model.Specifically, the contrast of target features is stretched based on perceptual importance, thereby improving the overall SE performance.Compared with post-processing-based implementations, incorporating PCS into the training phase preserves performance and reduces online computation.Notably, PCS can be combined with different SE model architectures and training criteria.Furthermore, PCS does not affect the causality or convergence of SE model training.Experimental results on the VoiceBank-DEMAND dataset show that the proposed method can achieve state-of-the-art performance on both causal (PESQ score = 3.07) and noncausal (PESQ score = 3.35) SE tasks.
Rong Chao, Szu-Wei Fu, Xugang Lu, Yu Tsao 0001
INTERSPEECH1