Yi Zhou 0014

dblp:01/1901-14 · DBLP profile ↗
← Back
26ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0001-7445-226XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 8 since 2021Systems, architecture and hardware · 5Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A novel sparse adaptive filter for suppressing impulsive disturbance in audio signals
Hongqing Liu 0002, Lu Gan 0002, Yi Zhou 0014, Maciej Niedzwiecki, Trieu-Kien Truong
Signal Process.4
2025 Scaling beyond Denoising: Submitted System and Findings in URGENT Challenge 2025
Zhihang Sun, Andong Li, Rilin Chen, Meng Yu 0003, Chengshi Zheng, Yi Zhou 0014, Dong Yu 0001
INTERSPEECH7
2025 EDTC: enhanced depth of text comprehension in automated audio captioning
abstract
Abstract Modality discrepancies have perpetually posed significant challenges within the realm of Automated Audio Captioning (AAC) and across all multimodal domains. Facilitating models in comprehending text information plays a pivotal role in establishing a seamless connection between the two modalities of text and audio. While recent research has focused on closing the gap between these two modalities through contrastive learning, it is challenging to bridge the difference between both modalities using only simple contrastive loss. This paper introduces enhanced depth of text comprehension, which enhances the model’s understanding of text information from three different perspectives. First, a combined Local-Global feature Fusion module is introduced to fuse heterogeneous audio features, enabling the extraction of high-level semantic information and the discovery of latent inter-sample relationships. Next, a novel representation module, TRANSLATOR, constructs a twin-branch structure based on the conventional dual-stream model, mapping features from both modalities into a shared high-dimensional audio-text space. Finally, contrastive learning is integrated with momentum-based weight updates, allowing the system to effectively capture shared high-level semantic representations across the audio and text modalities.
Liwen Tan, Yi Zhou 0014, Yin Cao
Comput. J.2
2025 Improved Encoder-Decoder Architecture With Human-Like Perception Attention for Monaural Speech Enhancement
Yi Zhou 0014, Zhenhua Cheng
IEEE Signal Process. Lett.2
2024 SMRU: Split-And-Merge Recurrent-Based UNet For Acoustic Echo Cancellation And Noise Suppression
abstract
The proliferation of deep neural networks has spawned the rapid development of acoustic echo cancellation and noise suppression, and plenty of prior arts have been proposed, which yield promising performance. Nevertheless, they rarely consider the deployment generality in different processing scenarios, such as edge devices, and cloud processing. To this end, this paper proposes a general model, termed SMRU, to cover different application scenarios. The novelty lies in two-fold. First, a multi-scale band split layer and band merge layer are proposed to effectively fuse local frequency bands for lower complexity modeling. Besides, by simulating the multi-resolution feature modeling characteristic of the classical UNet structure, a novel recurrent-dominated UNet is devised. It consists of multiple variable frame rate blocks, each of which involves the causal time down-/upsampling layer with varying compression ratios and the dualpath structure for inter- and intra-band modeling. The model is configured from $50 \mathrm{M} / \mathrm{s}$ to $6.8 \mathrm{G} / \mathrm{s}$ in terms of MACs, and the experimental results show that the proposed approach yields competitive or even better performance over existing baselines, and has the full potential to adapt to more general scenarios with varying complexity requirements.
Zhihang Sun, Andong Li, Rilin Chen, Hao Zhang 0112, Meng Yu 0003, Yi Zhou 0014, Dong Yu 0001
SLT6
2024 Cross Domain Optimization for Speech Enhancement: Parallel or Cascade?
abstract
This paper introduces five novel deep-learning architectures for speech enhancement. Existing methods typically use time-domain, time-frequency representations, or a hybrid approach. Recognizing the unique contributions of each domain to feature extraction and model design, this study investigates the integration of waveform and complex spectrogram models through cross-domain fusion to enhance speech feature learning and noise reduction, thereby improving speech quality. We examine both cascading and parallel configurations of waveform and complex spectrogram models to assess their effectiveness in speech enhancement. Additionally, we employ an orthogonal projection-based error decomposition technique and manage the inputs of individual sub-models to analyze factors affecting speech quality. The network is trained by optimizing three specific loss functions applied across all sub-models. Our experiments, using the DNS Challenge (ICASSP 2021) dataset, reveal that the proposed models surpass existing benchmarks in speech enhancement, offering superior speech quality and intelligibility. These results highlight the efficacy of our cross-domain fusion strategy.
Hongqing Liu 0001, Liming Shi, Yi Zhou 0014, Lu Gan 0002
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Time-frequency Domain Filter-and-sum Network for Multi-channel Speech Separation
Zhewen Deng, Yi Zhou 0014, Hongqing Liu 0002
INTERSPEECH2
2023 A Novel Earprint: Stimulus-Frequency Otoacoustic Emission for Biometric Recognition
abstract
Otoacoustic emission (OAE) biometrics are inherently robust to replay and falsification attacks. The widely studied transient-evoked OAE (TEOAE) is non-stationary and offers biometric value only in normal-hearing individuals since it is more susceptible to hearing loss. To address these issues, this paper presents a novel yet promising OAE biometric modality-stimulus-frequency OAE (SFOAE). Unlike TEOAE, SFOAE is a highly stationary signal whose fine structures are idiosyncratic to an individual, and relatively stable over time, making it easier to be a biometric without additional complex feature extraction. Moreover, SFOAE is even present in ears with 50 dB HL hearing loss, applicable to hearing-impaired users. In this paper, SFOAE spectra in response to three stimulus levels are fused in the feature level to consolidate different information, followed by a linear discriminant analysis or a multi-kernel convolutional neural network to further reduce the intra-subject variability and increase the inter-subject variability. Tested on a large cohort of subjects containing varying levels of deafness, the SFOAE-based biometric system yields an equal error rate of 0.541% and 1.364% in closed-set and open-set verification scenarios, respectively. In an identification mode, 99.43% and 97.37% accuracies are attained for closed-set and open-set protocols, respectively. In particular, we observe perfect performance in a population restricted to normal hearing in closed-set scenarios. The reason why the system performs well has been examined based on several comparative tests. Although there are implementation issues to be resolved before SFOAE can be applied in the field of biometric, this paper preliminarily demonstrates the basis and excellent potential of SFOAE as a biometric.
Borui Jiang, Hongqing Liu 0002, Fen Xiong, Yi Zhou 0014
IEEE Trans. Inf. Forensics Secur.6
2022 ICASSP 2022 L3DAS22 Challenge: Ensemble of Resnet-Conformers with Ambisonics Data Augmentation for Sound Event Localization and Detection
abstract
It remains a tough challenge to tackle sound event localization and detection (SELD) problem, especially when sound scene complexity increases and overlapping acoustic sources appear. To improve the SELD performance, we propose an ensemble system, which consists of a ResNet and Conformer backbone network (SELD-RCnet) and its two variants, SED-RCnet and SSL-RCnet. For SELD-RCnet and SSL-RCnet, we use short time Fourier transform (STFT) magnitude spectrogram, phase spectrogram, and active acoustic intensity vectors (IVs) as input features. For SSL-RCnet, an innovative predictive target is also developed and the performance is thus improved. For SED-RCnet, we use Log-Mel spectrogram as input features. To overcome the lack of training data, we adopt two novel approaches to first order Ambisonic (FOA) format dataset augmentation, namely audio channel swapping (ACS) and time-frequency masking (TFM). Finally, in the L3DAS22 Challenge, our submitted system achieves significant improvements over the baseline and ranks the second place for the Task2. Therefore, according to the competition rules, we submit this work to describe our system in details.
Yongjian Mao, Hongqing Liu 0002, Yi Zhou 0014
ICASSP5
2022 Acoustic Echo Cancellation and Noise Suppression with a Full Time-Frequency Cascaded Neural Network
abstract
With the developments of various multi-function communication services, acoustic echoes and background noises inevitably appear in hands-free calling occasions. Different from using the combination of neural network and traditional acoustic echo cancellation (AEC) method, this paper directly proposes a time-frequency complex cascaded neural network (TFCN) for echo cancellation and noise suppression. To that aim, in frequency domain, complex LSTM layers are employed to process the real and imaginary signals. After that, an end-to-end time domain network is designed using dilated convolution layers to further remove residual interferences. By adding rich delay information to the dataset and optimizing the model by a weighted loss function, the generalization ability of the model is also improved. The extensive experimental results show that the proposed frame-work is robust to blind test datasets, effectively removes echoes and noises, and achieves an excellent performance on AECMOS scores. The subjective mean score of the proposed method is 4.37, which is 0.50 higher than the INTERSPEECH2021 AEC-Challenge baseline.
Hongqing Liu 0001, Yi Zhou 0014, Lu Gan 0002
MMSP4
2022 Front-Wall Clutter Removal in Through-the-Wall Radar Based on Weighted Nuclear Norm Minimization
abstract
The front-wall clutter removal in the case of the through-the-wall radar (TWR) system is studied in this work. To remove the wall clutter, its low-rank property is utilized, and at the same time, the sparse property of the target returns is exploited to perform target reconstruction. To account for the unparalleled setting of the antenna and the wall, a weighted nuclear norm minimization (WNNM) is employed, and the resulting problem is solved in an alternating manner. In addition, different transmitted waveform signals, including monofrequency and stepped-frequency waveforms, are used to demonstrate their effects on the clutter suppression performances. The experimental results show that the proposed WNNM with stepped-frequency waveform outperforms other approaches.
Yi Zhou 0014, Hongqing Liu 0001, Dong Li 0007, Trieu-Kien Truong
IEEE Geosci. Remote. Sens. Lett.1
2021 A new diffusion variable spatial regularized LMS algorithm
Yijing Chu, S. C. Chan 0001, Yi Zhou 0014, Ming Wu 0005
Signal Process.3
2020 A Human Auditory Perception Loss Function Using Modified Bark Spectral Distortion for Speech Enhancement
Xiaofeng Shu, Yi Zhou 0014, Hongqing Liu 0002, Trieu-Kien Truong
Neural Process. Lett.2
2020 A New Diffusion Variable Spatial Regularized QRRLS Algorithm
abstract
This paper develops a framework for the design of diffusion adaptive algorithms, where a network of nodes aim to estimate system parameters from the collected distinct local data stream. We explore the time and spatial knowledge of system responses and model their evolution in both time and spatial domain. A weighted maximum a posteriori probability (MAP) is used to derive an adaptive estimator, where recent data has more influence on statistics via weighting factors. The resulting recursive least squares (RLS) local estimate can be implemented by the QR decomposition (QRD). To mediate the distinct spatial information incorporation within neighboring estimates, a variable spatial regularization (VSR) parameter is introduced. The estimation bias and variance of the proposed algorithm are analyzed. A new diffusion VSR QRRLS (Diff-VSR-QRRLS) algorithm is derived that balances the bias and variance terms. Simulations are carried out to illustrate the effectiveness of the theoretical analysis and evaluate the performance of the proposed algorithm.
Yijing Chu, S. C. Chan 0001, Yi Zhou 0014, Ming Wu 0005
IEEE Signal Process. Lett.3
2020 Clutter Reduction and Target Tracking in Through-the-Wall Radar
abstract
This article addresses the problem of tracking targets behind the wall using through-the-wall radar. To that end, the wall reflection, i.e., clutter, must be eliminated first because it interferes with the subsequent image formation operation. The low-rank of the clutter and sparseness of the useful signal are utilized to devise a joint low-rank and sparse framework to simultaneously suppress the clutter and recover the target returns, where alternating direction method of multipliers (ADMM) approach is developed to solve the corresponding optimization. Since then, an effective observation window scheme is proposed to locate the target and further to facilitate the tracking process. The tracking is finally provided by Kalman filter and particle filter. The numerical studies are provided to demonstrate that the performance of the proposed framework is superior to that of other methods in terms of clutter removal and tracking accuracy.
Hongqing Liu 0001, Lu Gan 0002, Yi Zhou 0014, Trieu-Kien Truong
IEEE Trans. Geosci. Remote. Sens.4
2019 Phase Time-Frequency Masking Based Speech Enhancement Algorithm Using Circular Microphone Array
abstract
A novel time-frequency masking approach for circular microphone array speech enhancement in the presence of competing interference and background noise is proposed in this paper. Multichannel speech enhancement systems can often be constructed by a concatenation of a beamformer and a single-channel postfilter, which rely on accurate estimation of steering vector and the residual interference plus noise power spectrum density (PSD), respectively. However, the performance of existing multiple microphone speech enhancement algorithm will degrade in the presence of competing interference. The proposed phase-based time-frequency masking approach can improve the estimation of the steering vector and residual interference plus noise PSD in the presence of competing interference and background noise. The experimental analysis verifies the advantages achieved by the proposed method, in comparison with the state-of-the-art multiple microphone speech enhancement methods.
Yi Zhou 0014, Hongqing Liu 0002
ICME2
2019 A Robust GSC Beamforming Method for Speech Enhancement using Linear Microphone Array
abstract
The speech enhancement problem is studied using an improved robust generalized sidelobe canceler (GSC) beamforming algorithm based on microphone array, in the cases of speaker noise and the music interferences. The conventional GSC algorithm based on variable step size and a priori signal-to-noise ratio (SNR) algorithm is not robust under the nonstationary noise because the solution of the signal-to-noise ratio (SNR) is not given. To enhance the robustness, in this paper, a improved GSC algorithm is developed, where adaptive filter coefficients are updated based on signal output power ratio (SPR). The numerical studies including speaker noise and music noise demonstrate that the improved algorithm outperforms the traditional GSC and the GSC based on variable step size technique.
Feng Ni, Yi Zhou 0014, Hongqing Liu 0002
MMSP2
2018 Simultaneous Radio Frequency and Wideband Interference Suppression in SAR Signals via Sparsity Exploitation in Time-Frequency Domain
abstract
This paper addresses the problem of recovering a synthetic aperture radar (SAR) signal that is corrupted by both radio frequency interference (RFI) and wideband interference (WBI). The time–frequency domain is utilized for both the SAR signal and interference in the form of sparse representations. By doing so, a unified framework that allows one to suppress both the RFI and WBI while recovering the SAR signal can be developed. The resulting framework is an optimization problem that is efficiently solved using a customized alternating direction method of multipliers approach. Finally, simulation results are provided to demonstrate that the performance of the joint estimation algorithm is superior to the performances of other methods in terms of both subjective and objective evaluation standards.
Hongqing Liu 0001, Dong Li 0007, Yi Zhou 0014, Trieu-Kien Truong
IEEE Trans. Geosci. Remote. Sens.3
2017 Image deblurring in the presence of salt-and-pepper noise
abstract
This work addresses image recovery problem in the presence of salt-and-pepper noise and image blur. The salt-and-pepper noise reviewed as the impulsive noise, in this paper, is modeled as a sparse signal because of its impulsiveness. To accurately reconstruct the clean image and the blur kernel, the framelet domains are exploited to sparsely represent the image and the blur kernel. From the reformulations conducted, a joint estimation is devised to simultaneously perform the image recovery, the salt-and-pepper noise suppression and the blur kernel estimation under a optimization framework. To solve the optimization problem, an efficient solver based on accelerated proximal gradient (APG) is developed to obtain the joint estimation solution. Numerical studies demonstrate the superior performance of the joint estimation algorithm compared with the state-of-the-art approaches in terms of both objective and subjective evaluation standards.
Liming Hou, Hongqing Liu 0002, Yi Zhou 0014, Trieu-Kien Truong
ICIP4
2017 Joint Wideband Interference Suppression and SAR Signal Recovery Based on Sparse Representations
abstract
The problem of synthetic aperture radar image recovery in the presence of wideband interference (WBI) is investigated. Delayed versions of a transmitted signal are utilized to construct a dictionary in which a signal of interest (SOI) has a sparse representation. In this letter, WBI is sparsely represented by the time-frequency domain. By utilizing the transform domains, a joint estimation approach is devised to simultaneously perform WBI suppression and SOI recovery within an optimization framework. Based on the separability property in the optimization, an alternating direction method of multipliers-based approach is developed to efficiently obtain a solution. Finally, simulation results are presented to demonstrate the superior performance of the joint estimation algorithm.
Hongqing Liu 0002, Dong Li 0007, Yi Zhou 0014, Trieu-Kien Truong
IEEE Geosci. Remote. Sens. Lett.3
2011 Two-channel post-filtering based on adaptive smoothing and noise properties
abstract
This paper studies the statistical properties of the gain functions, which are often used for two-channel post-filtering (TC-PF) algorithms. We reveal that the smoothing factor has a significant impact on both noise reduction and musical noise. When the smoothing factor increases, noise reduction can be improved and musical noise can be reduced simultaneously. However, the smoothing factor could not be too close to one because the system can only be assumed to be time-invariant for short durations. To solve this problem, this paper proposes an adaptive smoothing scheme by detecting the sudden change of the system. Moreover, the residual noise floor is adaptively chosen based on the structure of the noise power spectral density (NPSD) to further suppress the tonal noise components. Experimental results show the better performance of the proposed algorithm in terms of the segmental signal to-noise-ratio (SNR) and the PESQ improvements.
Chengshi Zheng, Yi Zhou 0014, Xiaohu Hu, Xiaodong Li 0002
ICASSP2
2009 On the Convergence Behavior of the Noise-constrained NLMS Algorithm
abstract
This paper studies the convergence behaviors of the noise-constrained normalized least mean squares (NCNLMS) algorithm recently proposed in the work of Chan et al. (2008). Like its LMS counterpart, the NCNLMS algorithm employs the prior knowledge of the additive noise to adjust its step-size. Following (Wei et al., 2001), the convergence behaviors of the NCLMS under the noise mismatch cases are firstly derived. Using a novel transformation approach and the small step-size properties of the NCNLMS algorithm at convergence, the mean and mean squares behaviors of this algorithm are derived. The validity of the proposed analysis is verified well by computer simulations and the relative merits of the NCLMS and NCNLMS algorithms are also compared.
S. C. Chan 0001, Y. J. Chu, Zhiguo Zhang 0001, Yi Zhou 0014
ISCAS4
2009 Convergence Behaviors of the Fast LMM/Newton Algorithm with Gaussian Inputs and Contaminated Gaussian Noise
abstract
This paper studies the convergence behaviors of the fast least mean M-estimate/Newton adaptive filtering algorithm proposed in (Y. Zhou et al.,2004), which is based on the fast LMS/Newton principle and the minimization of an M-estimate function using robust statistics for robust filtering in impulsive noise. By using the Price's theorem and its extension for contaminated Gaussian (CG) noise case, the convergence behaviors of the fast LMM/ Newton algorithm with Gaussian inputs and both Gaussian and CG noises are analyzed. Difference equations describing the mean and mean square behaviors of this algorithm and step size bound for ensuring stability are derived. These analytical results reveal the advantages of the fast LMM/Newton algorithm in combating impulsive noise, and they are in good agreement with computer simulation results.
S. C. Chan 0001, Yi Zhou 0014
ISCAS2
2009 Robust Linear Estimation using M-Estimation and Weighted L1 Regularization: Model Selection and Recursive Implementation
abstract
This paper studies an M-estimation-based method for linear estimation with weighted L1 regularization and its recursive implementation. Motivated by the sensitivity of conventional least-squares-based L1-regularized linear estimation (Lasso) in impulsive noise environment, an M-estimator-based Lasso (M-Lasso) method is introduced to restrain the outliers and an iterative re-weighted least-squares (IRLS) algorithm is proposed to solve this M-estimation problem. Moreover, instead of using the matrix inversion formula, QR decomposition (QRD) is employed in the M-Lasso for recursive implementation with a lower arithmetic complexity. Simulation results show that the M-estimation-based Lasso performs considerably better than the traditional LS-based Lasso in suppressing the impulsive noise, and its recursive QRD algorithm has a good performance in online processing.
Zhiguo Zhang 0001, S. C. Chan 0001, Yi Zhou 0014, Yong Hu 0003
ISCAS3
2006 Improved generalized-proportionate stepsize LMS algorithms and performance analysis
abstract
This paper analyzes the performance of the GP-NLMS algorithm, revealing the nature of its fast convergence as well as its deficiency of inducing bigger steady state error. Based on the analysis, a class of improved generalized-proportionate stepsize LMS (GPS-LMS) algorithms are proposed. With an efficient switching mechanism, the new algorithms can dynamically switch between the GP-NLMS and conventional LMS-type algorithms to achieve fast initial convergence and tracking speed and low steady state error. Computer simulations verified the superior performance of the proposed algorithms
S. C. Chan 0001, Yi Zhou 0014
ISCAS2
2006 A new adaptive Kalman filter-based subspace tracking algorithm and its application to DOA estimation
abstract
This paper presents a new Kalman filter-based subspace tracking algorithm and its application to directions of arrival (DOA) estimation. An autoregressive (AR) process is used to describe the dynamics of the subspace and a new adaptive Kalman filter with variable measurements (KFVM) algorithm is developed to estimate the time-varying subspace recursively from the state-space model and the given observations. For stationary subspace, the proposed algorithm will switch to the conventional PAST to lower the computational complexity. Simulation results show that the adaptive subspace tracking method has a better performance than conventional algorithms in DOA estimation for a wide variety of experimental condition
S. C. Chan 0001, Zhiguo Zhang 0001, Yi Zhou 0014
ISCAS3