Chengshi Zheng

dblp:88/580 · DBLP profile ↗
← Back
65ranked-venue papers
8as first author
53since 2021 · last 2026
0000-0001-5656-994XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 51 · 6 first-author · 41 since 2021Artificial intelligence and machine learning · 35 · 2 first-author · 32 since 2021
YearPublicationVenuePosition
2026 GOMPSNR: Reflourish the Signal-to-Noise Ratio Metric for Audio Generation Tasks
abstract
In the field of audio generation, signal-to-noise ratio (SNR) has long served as an objective metric for evaluating audio quality. Nevertheless, recent studies have shown that SNR and its variants are not always highly correlated with human perception, prompting us to raise the questions: Why does SNR fail in measuring audio quality? And how to improve its reliability as an objective metric? In this paper, we identify the inadequate measurement of phase distance as a pivotal factor and propose to reformulate SNR with specially designed phase-distance terms, yielding an improved metric named GOMPSNR. We further extend the newly proposed formulation to derive two novel categories of loss function, corresponding to magnitude-guided phase refinement and joint magnitude-phase optimization, respectively. Besides, extensive experiments are conducted for an optimal combination of different loss functions. Experimental results on advanced neural vocoders demonstrate that our proposed GOMPSNR exhibits more reliable error measurement than SNR. Meanwhile, our proposed loss functions yield substantial improvements in model performance, and our well-chosen combination of different loss functions further optimizes the overall model capability.
Lingling Dai, Andong Li, Yifan Liang, Xiaodong Li 0002, Chengshi Zheng
AAAI6
2026 DegVoC: Revisiting Neural Vocoder from a Degradation Perspective
abstract
Existing neural vocoders have demonstrated promising performance by leveraging Mel-spectrum as an acoustic feature for conditional audio generation. Nonetheless, they remain constrained by an inherent ``performance-cost'' dilemma that significantly hinders the development of this field. This paper revisits this foundational task from a novel degradation perspective, where Mel-spectrum is regarded as a special signal degradation process from the target spectrum. Drawing inspiration from traditional sparse signal recovery problems, we propose DegVoC, a GAN-based neural vocoder with a two-step solution procedure. First, by exploiting degradation priors, we attempt to retrieve the initial spectral structure from Mel-domain representations as an initial solution via a simple linear transformation. Based on that, we introduce a deep prior solver that accounts for the heterogeneous distribution of sub-bands in the time-frequency domain. A convolution-style attention module with a large kernel size is specially devised for efficient inter-frame and inter-band contextual modeling. With 3.89 M parameters and substantially reduced inference complexity, DegVoC achieves state-of-the-art performance across objective and subjective evaluations, outperforming existing GAN-, DDPM- and flow-matching-based baselines.
Andong Li, Lingling Dai, Rilin Chen, Meng Yu 0003, Xiaodong Li 0002, Dong Yu 0001, Chengshi Zheng
AAAI9
2026 SLD-L2S: Hierarchical Subspace Latent Diffusion for High-Fidelity Lip to Speech Synthesis
abstract
Although lip-to-speech synthesis (L2S) has achieved significant progress in recent years, current state-of-the-art methods typically rely on intermediate representations such as mel-spectrograms or discrete self-supervised learning (SSL) tokens. The potential of latent diffusion models (LDMs) in this task remains largely unexplored. In this paper, we introduce SLD-L2S, a novel L2S framework built upon a hierarchical subspace latent diffusion model. Our method aims to directly map visual lip movements to the continuous latent space of a pre-trained neural audio codec, thereby avoiding the information loss inherent in traditional intermediate representations. The core of our method is a hierarchical architecture that processes visual representations through multiple parallel subspaces, initiated by a subspace decomposition module. To efficiently enhance interactions within and between these subspaces, we design the diffusion convolution block (DiCB) as our network backbone. Furthermore, we employ a reparameterized flow matching technique to directly generate the target latent vectors. This enables a principled inclusion of speech language model (SLM) and semantic losses during training, moving beyond conventional flow matching objectives and improving synthesized speech quality. Our experiments show that SLD-L2S achieves state-of-the-art generation quality on multiple benchmark datasets, surpassing existing methods in both objective and subjective evaluations.
Yifan Liang, Andong Li, Guochen Yu, Fangkun Liu, Lingling Dai, Xiaodong Li 0002, Chengshi Zheng
AAAI8
2026 DBMIF: a deep balanced multimodal iterative fusion framework for air- and bone-conduction speech enhancement
Yilei Wu, Changyan Zheng, Yakun Zhang 0002, Chengshi Zheng, Ye Yan 0001, Erwei Yin
Appl. Intell.5
2026 A Physics-Informed Kolmogorov-Arnold Graph Attention Network for solid mechanics on arbitrary geometries
Zheng Duanmu, Ning Meng, Hao Gao 0002, Chengshi Zheng
Eng. Appl. Artif. Intell.6
2026 NaturalL2S: End-to-end high-quality multispeaker lip-to-speech synthesis with differential digital signal processing
Yifan Liang, Fangkun Liu, Andong Li, Xiaodong Li 0002, Chengyou Lei, Chengshi Zheng
Neural Networks6
2026 Rethinking the Joint Estimation of Magnitude and Phase for Time-Frequency Domain Neural Vocoders
abstract
Time-frequency (T-F) domain-based neural vocoders have shown promising results in synthesizing high-fidelity audio. Nevertheless, it remains unclear on the mechanism of effectively predicting magnitude and phase targets jointly. In this paper, we start from two representative T-F domain vocoders, namely Vocos and APNet2, which belong to the single-stream and dual-stream modes for magnitude and phase estimation, respectively. When evaluating their performance on a large-scale dataset, we accidentally observe severe performance collapse of APNet2. To stabilize its performance, in this paper, we introduce three simple yet effective strategies, each targeting the topological space, the source space, and the output space, respectively. Specifically, we modify the architectural topology for better information exchange in the topological space, introduce prior knowledge to facilitate the generation process in the source space, and optimize the backpropagation process for parameter updates with an improved output format in the output space. Experimental results demonstrate that our proposed method effectively facilitates the joint estimation of magnitude and phase in APNet2, thus bridging the performance disparities between the single-stream and dual-stream vocoders.
Lingling Dai, Andong Li, Xiaodong Li 0002, Chengshi Zheng
IEEE Signal Process. Lett.5
2026 OmniControl: Unified Audio Extraction and Elimination via Subband-Aware Separation
Yifan Liang, Andong Li, Xiaodong Li 0002, Chengshi Zheng
IEEE Signal Process. Lett.4
2025 BSDB-Net: Band-Split Dual-Branch Network with Selective State Spaces Mechanism for Monaural Speech Enhancement
abstract
Although the complex spectrum-based speech enhancement (SE) methods have achieved significant performance, coupling amplitude and phase can lead to a compensation effect, where amplitude information is sacrificed to compensate for the phase that is harmful to SE. In addition, to further improve the performance of SE, many modules are stacked onto SE, resulting in increased model complexity that limits the application of SE. To address these problems, we proposed a dual-path network based on compressed frequency using Mamba. First, we extract amplitude and phase information through parallel dual branches. This approach leverages structured complex spectra to implicitly capture phase information and solves the compensation effect by decoupling amplitude and phase, and the network incorporates an interaction module to suppress unnecessary parts and recover missing components from the other branch. Second, to reduce network complexity, the network introduces a band-split strategy to compress the frequency dimension. To further reduce complexity while maintaining good performance, we designed a Mamba-based module that models the time and frequency dimensions under linear complexity. Finally, compared to baselines, our model achieves an average 8.3 times reduction in computational complexity while maintaining superior performance. Furthermore, it achieves a 25 times reduction in complexity compared to transformer-based models.
Cunhang Fan, Enrui Liu, Andong Li, Jianhua Tao 0001, Jian Zhou 0006, Chengshi Zheng, Zhao Lv
AAAI7
2025 DSINet: Towards Real-Time Target Speaker Extraction with Dynamic Speaker Information Fusion
abstract
Target speaker extraction (TSE) aims to directly extract the desired speech given enrollment utterances of the target speaker. Despite significant progress in recent years, most existing methods remain non-causal and computationally intensive. This paper introduces DSINet, a real-time time-frequency (T-F) domain method that leverages the dynamic speaker information fusion mechanism to estimate the real and imaginary (RI) components of the target speech. This method incorporates the T-F band-split modeling as primary speaker extractor. Moreover, instead of explicitly calculating the target speaker embedding, a dynamic speaker information fusion mechanism is proposed for the efficient utilization of target speaker information within each mixture, guiding the backbone extractor towards the desired speech. Experimental results on the WSJ0-2mix and WHAMR! datasets confirm that the proposed method exhibits remarkable scalability and achieves comparable performance to prominent non-causal methods under different model sizes.
Fengyuan Hao, Andong Li, Xiaodong Li 0002, Chengshi Zheng
ICASSP4
2025 BiCG: Binaural Cue Generation from Unified HRTF Datasets
abstract
Head-related transfer functions (HRTFs) are important for spatial audio reproduction in immersive systems. Most existing data-driven methods focus on personalized HRTF estimation of monaural spectral factors. These methods ignore the importance of binaural cues, which are essential for binaural reproduction and perception. Moreover, the significant differences among various HRTF datasets in aspects such as measurement setup limit the potential of data-driven methods. This paper proposes a binaural cue generation method (BiCG), which utilizes an implicit neural network (INN) to estimate interaural level differences (ILDs) and interaural time differences (ITDs). Experimental results show that our method outperforms existing neural field methods in terms of binaural cue generation quality across datasets. We also evaluate various data preprocessing methods, and experimental results show that extreme smoothing improves binaural cue generation performance across datasets. The work provides new insights into enhancing HRTF modeling.
Xikun Lu, Jinqiu Sang, Chengshi Zheng
ICASSP4
2025 DeepPEM-AFC: An Improved Prediction-Error-Method-based Adaptive Feedback Cancellation with Deep Learning for Hearing Aids
abstract
Hearing assistive devices aim to compensate hearing loss for hearing-impaired listeners, and their maximum stable gain (MSG) is constrained because of the existence of the acoustic feedback between the receiver and microphone, resulting in their inefficiency for individuals with severe or profound hearing loss who require very large amplification gain. Adaptive feedback cancellation (AFC) is an effective method to reduce acoustic feedback and increase MSG but its performance often degrades because of the high correlation between the target and feedback signals. The prediction-error-method (PEM)-based AFC has shown its capability in reducing this degradation. This paper proposes a deep learning-based PEM-AFC dubbed DeepPEM-AFC to further improve the performance of traditional PEM-AFC by fully taking advantage of both deep learning in automatically finding the optimal step size when updating the filter coefficients and PEM in solving the abovementioned high-correlation problem. To improve generalization across different acoustic feedback paths, a path generation scheme is proposed for training purposes. Experimental results show that DeepPEM-AFC achieves superior tracking performance compared to state-of-the-art methods, including traditional methods and Neural-AFC. Moreover, combining DeepPEM-AFC with frequency shifting further improves the performance.
Xiaofan Zhan, Fengyuan Hao, Xiaodong Li 0002, Chengshi Zheng
ICASSP4
2025 Audiogram-Informed End-to-End Noise Reduction and Wide Dynamic Range Compression for Hearing Aids
abstract
Wide dynamic range compression (WDRC) provides level-dependent amplification, intended to make the output of a hearing aid fall between the hearing threshold and the highest comfortable level of the listener. Hearing aids often combine noise reduction with WDRC, applied sequentially. Unfortunately, this can result in across-source modulation, especially when fast-acting compression is used. However, fast-acting compression is theoretically preferable to compensate for the loss of compression in the cochlea. To apply fast-acting compression to speech while avoiding across-source modulation, we propose a deep-learning-based method that integrates noise reduction and source-independent WDRC in a two-stage low-complexity framework. The method applies fast-acting compression to speech and slow-acting compression to noise by an amount depending on the audiogram, with a controllable residual noise level. Objective measurements using simulated hearing-impaired listeners showed that the proposed method reduced negative interactions of speech and noise in highly non-stationary noise scenarios and generalized well across various hearing losses.
Huiyong Zhang, Brian C. J. Moore, Lingling Dai, Fengyuan Hao, Xiaodong Li 0002, Chengshi Zheng
ICASSP6
2025 BridgeVoC: Neural Vocoder with Schrödinger Bridge
abstract
While previous diffusion-based neural vocoders typically follow a noise-to-data generation pipe-line, the linear-degradation prior of the mel-spectrogram is often neglected, resulting in limited generation quality. By revisiting the vocoding task and excavating its connection with the signal restoration task, this paper proposes a time-frequency (T-F) domain-based neural vocoder with the Schrödinger Bridge, called BridgeVoC, which is the first to follow the data-to-data generation paradigm. Specifically, the mel-spectrogram can be projected into the target linear-scale domain and regarded as a degraded spectral representation with a deficient rank distribution. Based on this, the Schrödinger Bridge is leveraged to establish a connection between the degraded and target data distributions. During the inference stage, starting from the degraded representation, the target spectrum can be gradually restored rather than generated from a Gaussian noise process. Quantitative experiments on LJSpeech and LibriTTS show that BridgeVoC achieves faster inference and surpasses existing diffusion-based vocoder baselines, while also matching or exceeding non-diffusion state-of-the-art methods across evaluation metrics.
Rilin Chen, Meng Yu 0003, Chengshi Zheng, Dong Yu 0001, Andong Li
IJCAI6
2025 Learning Neural Vocoder from Range-Null Space Decomposition
abstract
Despite the rapid development of neural vocoders in recent years, they usually suffer from some intrinsic challenges like opaque modeling, and parameter-performance trade-off. In this study, we propose an innovative time-frequency (T-F) domain-based neural vocoder to resolve the above-mentioned challenges. To be specific, we bridge the connection between the classical signal range-null decomposition (RND) theory and vocoder task, and the reconstruction of target spectrogram can be decomposed into the superimposition between the range-space and null-space, where the former is enabled by a linear domain shift from the original mel-scale domain to the target linear-scale domain, and the latter is instantiated via a learnable network for further spectral detail generation. Accordingly, we propose a novel dual-path framework, where the spectrum is hierarchically encoded/decoded, and the cross- and narrow-band modules are elaborately devised for efficient sub-band and sequential modeling. Comprehensive experiments are conducted on the LJSpeech and LibriTTS benchmarks. Quantitative and qualitative results show that while enjoying lightweight network parameters, the proposed approach yields state-of-the-art performance among existing advanced methods. Our code and the pretrained model weights are available at https://github.com/Andong-Li-speech/RNDVoC.
Andong Li, Zhihang Sun, Rilin Chen, Erwei Yin, Xiaodong Li 0002, Chengshi Zheng
IJCAI7
2025 L3C-DeepMFC: Low-Latency Low-Complexity Deep Marginal Feedback Cancellation with Closed-Loop Fine Tuning for Hearing Aids
Fengyuan Hao, Brian C. J. Moore, Huiyong Zhang, Xiaodong Li 0002, Chengshi Zheng
INTERSPEECH5
2025 LightL2S: Ultra-Low Complexity Lip-to-Speech Synthesis for Multi-Speaker Scenarios
Yifan Liang, Fangkun Liu, Andong Li, Xiaodong Li 0002, Chengshi Zheng
INTERSPEECH6
2025 Scaling beyond Denoising: Submitted System and Findings in URGENT Challenge 2025
Zhihang Sun, Andong Li, Rilin Chen, Meng Yu 0003, Chengshi Zheng, Yi Zhou 0014, Dong Yu 0001
INTERSPEECH6
2025 BAPEN: Towards Versatile Audio Phase Retrieval
Lingling Dai, Andong Li, Chengshi Zheng, Xiaodong Li 0002
ACM Multimedia4
2025 Bridging semantics across modalities: Decoupled representation learning for audio-visual speech recognition
Linzhi Wu, Yakun Zhang 0002, Changyan Zheng, Tiejun Liu, Liang Xie 0012, Chengshi Zheng, Erwei Yin
Knowl. Based Syst.7
2024 All Neural Kronecker Product Beamforming for Speech Extraction with Large-Scale Microphone Arrays
abstract
Existing frame-wise neural beamformers for speech extraction can obtain promising performance in relatively high signal-to-noise ratio (SNR) scenarios using small microphone arrays, while they still suffer from performance degradation in relatively low SNR environments, e.g., SNR<-5 dB. As an attempt to solve this problem, this paper proposes an all-neural beamformer based on Kronecker product decomposition, denoted by NeuKP-BF, for large-scale microphone arrays. The core idea is to incorporate the high spatial resolution of large microphone arrays and the powerful non-linear modeling capability of deep neural networks to improve speech extraction performance in challenging environments. In this paper, to reduce the feature representation redundancy and improve the interpretability, we used the Kronecker product rule to decompose the original large-scale array into two small virtual subarrays, and beamformers for the two subarrays were then designed and merged finally. The whole system was designed to implement in an end-to-end manner. Experiments were conducted on both the synthesized data using the DNS-Challenge corpus. The results showed that the proposed approach outperformed existing advanced baselines in terms of multiple objective metrics.
Weixin Meng, Andong Li, Xiaodong Li 0002, Chengshi Zheng
ICASSP6
2024 BAE-Net: a Low Complexity and High Fidelity Bandwidth-Adaptive Neural Network for Speech Super-Resolution
abstract
Speech bandwidth extension (BWE) has demonstrated promising performance in enhancing the perceptual speech quality in real communication systems. Most existing BWE researches primarily focus on fixed upsampling ratios, disregarding the fact that the effective bandwidth of captured audio may fluctuate frequently due to various capturing devices and transmission conditions. In this paper, we propose a streaming adaptive bandwidth extension solution dubbed BAE-Net, which is suitable to handle the low-resolution speech with unknown and varying effective bandwidth. To address the challenges of recovering both the high-frequency magnitude and phase components of the speech content blindly, we devise a dual-stream architecture that incorporates the magnitude inpainting and phase refinement. For potential applications on edge devices, this paper also introduces BAE-NET-lite, which is a lightweight, streaming and efficient framework. Quantitative results demonstrate the superiority of BAE-Net in terms of performance and computational efficiency when compared with existing state-of-the-art BWE methods.
Guochen Yu, Xiguang Zheng, Runqiang Han, Chengshi Zheng
ICASSP5
2024 Spatial reconstructed local attention Res2Net with F0 subband for fake speech detection
Cunhang Fan, Jun Xue 0001, Jianhua Tao 0001, Jiangyan Yi, Chenglong Wang 0001, Chengshi Zheng, Zhao Lv
Neural Networks6
2024 Geometry Calibration for Deformable Linear Microphone Arrays With Bézier Curve Fitting
abstract
Geometry calibration of microphone arrays is an essential preprocessing step for many applications, e.g., beamforming in deformable arrays. However, most existing geometry calibration methods involve the use of non-convex cost functions, suffering from the local minimum problem and low stability. To overcome these drawbacks, we introduce a novel approach that leverages the geometry feature of deformable linear arrays (DLAs) as an additional constraint. The proposed method employs Bézier curve fitting, utilizing the characteristics of Bézier curves to model the geometry feature. Specifically, we first introduce the general form of the geometry calibration problem, and an alternative approach is then proposed for a specific scenario where quadratic Bézier curves are used to fit the array shape. Finally, an additional scale modification is adopted to improve the performance of the proposed method in real scenarios. Simulations and real experiments validate the effectiveness of the proposed method for geometry calibration of DLAs.
Yuhai Ge, Weixin Meng, Xiaodong Li 0002, Chengshi Zheng
IEEE Signal Process. Lett.4
2024 Deep Kronecker Product Beamforming for Large-Scale Microphone Arrays
abstract
Although deep learning based beamformers have achieved promising performance using small microphone arrays, they suffer from performance degradation in very challenging environments, such as extremely low Signal-to-Noise Ratio (SNR) environments, e.g., SNR$\le$−10 dB. A large-scale microphone array with dozens or hundreds of microphones can improve the performance of beamformers in these challenging scenarios because of its high spatial resolution. While a dramatic increase in the number of microphones leads to feature redundancy, causing difficulties in feature extraction and network training. As an attempt to improve the performance of deep beamformers for speech extraction in very challenging scenarios, this paper proposes a novel all neural Kronecker product beamforming denoted by ANKP-BF for large-scale microphone arrays by taking the following two aspects into account. Firstly, a larger microphone array can provide higher performance of spatial filtering when compared with a small microphone array, and deep neural networks are introduced for their powerful non-linear modeling capability in the speech extraction task. Secondly, the feature redundancy problem is solved by introducing the Kronecker product rule to decompose the original one high-dimension weight vector into the Kronecker product of two much lower-dimensional weight vectors. The proposed ANKP-BF is designed to operate in an end-to-end manner. Extensive experiments are conducted on simulated large-scale microphone-array signals using the DNS-Challenge corpus and WSJ0-SI84 corpus, and the real recordings in a semi-anechoic room and outdoor scenes are also used to evaluate and compare the performance of different methods. Quantitative results demonstrate that the proposed method outperforms existing advanced baselines in terms of multiple objective metrics, especially in very low SNR environments.
Weixin Meng, Andong Li, Xiaoxue Luo, Shefeng Yan, Xiaodong Li 0002, Chengshi Zheng
IEEE ACM Trans. Audio Speech Lang. Process.7
2023 Gesper: A Unified Framework for General Speech Restoration
abstract
This paper describes the legends-tencent team’s real-time General Speech Restoration (Gesper) system submitted to the ICASSP 2023 Speech Signal Improvement (SSI) Challenge. This newly proposed system is a two-stage architecture, in which the speech restoration is performed, and then followed by speech enhancement. We propose a complex spectral mapping-based generative adversarial network (CSM-GAN) as the speech restoration module for the first time. For noise suppression and dereverberation, the enhancement module is presented with fullband-wideband parallel processing. On the blind test set of ICASSP 2023 SSI Challenge, the proposed Gesper system, which satisfies the real-time condition, achieves 3.27 P.804 overall mean opinion score (MOS) and 3.35 P.835 overall MOS, ranked 1st in both track 1 and track 2.
Jun Chen 0024, Yupeng Shi, Wei Rao 0002, Shulin He, Andong Li, Yannan Wang, Zhiyong Wu 0001, Shidong Shang, Chengshi Zheng
ICASSP10
2023 TaylorBeamixer: Learning Taylor-Inspired All-Neural Multi-Channel Speech Enhancement from Beam-Space Dictionary Perspective
Andong Li, Weixin Meng, Guochen Yu, Xiaodong Li 0002, Chengshi Zheng
INTERSPEECH6
2023 Low-complexity Broadband Beampattern Synthesis using Array Response Control
Weixin Meng, Xiaodong Li 0002, Chengshi Zheng
INTERSPEECH5
2023 CompNet: Complementary network for single-channel speech enhancement
Cunhang Fan, Andong Li, Wang Xiang, Chengshi Zheng, Zhao Lv, Xiaopei Wu
Neural Networks5
2023 End-to-end neural speaker diarization with an iterative adaptive attractor estimation
Fengyuan Hao, Xiaodong Li 0002, Chengshi Zheng
Neural Networks3
2023 A General Unfolding Speech Enhancement Method Motivated by Taylor's Theorem
abstract
While deep neural networks have facilitated significant advancements in the field of speech enhancement, most existing methods are developed following either empirical or relatively blind criteria, lacking adequate guidelines in pipeline design. Inspired by Taylor's theorem, we propose a general unfolding framework for both single- and multi-channel speech enhancement tasks. Concretely, we formulate the complex spectrum recovery into the spectral magnitude mapping in the neighborhood space of the noisy mixture, in which an unknown sparse term is introduced and applied for phase modification in advance. Based on that, the mapping function is decomposed into the superimposition of the 0th-order and high-order polynomials in Taylor's series, where the former coarsely removes the interference in the magnitude domain and the latter progressively complements the remaining spectral detail in the complex spectrum domain. In addition, we study the relation between adjacent order terms and reveal that each high-order term can be recursively estimated with its lower-order term, and each high-order term is then proposed to evaluate using a surrogate function with trainable weights, so that the whole system can be trained in an end-to-end manner. Given that the proposed framework is devised with the motivation of Taylor's theorem, it possesses improved internal flexibility. Extensive experiments are conducted on WSJ0-SI84, DNS-Challenge, Voicebank+Demand, spatialized Librispeech, and L3DAS22 multi-channel speech enhancement challenge datasets. Quantitative results show that the proposed approach yields competitive performance over existing top-performing approaches in terms of multiple objective metrics.
Andong Li, Guochen Yu, Chengshi Zheng, Xiaodong Li 0002
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Embedding and Beamforming: All-Neural Causal Beamformer for Multichannel Speech Enhancement
abstract
Standing upon the intersection of traditional beamformers and deep neural networks, we propose a causal neural beamformer paradigm called Embedding and Beamforming, and two core modules are devised accordingly, namely EM and BM. For EM, instead of estimating spatial covariance matrix explicitly, the 3-D embedding tensor is learned with the network, where the spatial-spectral discriminative information can be implicitly represented. For BM, a network is directly leveraged to derive the beamforming weights so as to implement filter-and-sum operation. To further improve the speech quality, a post-processing module is introduced to further suppress the residual noise. Based on the DNS-Challenge dataset, we conduct the experiments for multichannel speech enhancement and the results show that the proposed system outperforms previous advanced baselines by a large margin in terms of multiple evaluation metrics.
Andong Li, Chengshi Zheng, Xiaodong Li 0002
ICASSP3
2022 Joint Magnitude Estimation and Phase Recovery Using Cycle-In-Cycle GAN for Non-Parallel Speech Enhancement
abstract
For the lack of adequate paired noisy-clean speech corpus in many real scenarios, non-parallel training is a promising task for DNN-based speech enhancement methods. However, because of the severe mismatch between input and target speeches, many previous studies only focus on the magnitude spectrum estimation and remain the phase unaltered, resulting in the degraded speech quality under low signal-to-noise ratio conditions. To tackle this problem, we decouple the difficult target w.r.t. original spectrum optimization into spectral magnitude and phase, and a novel Cycle-in-Cycle generative adversarial network (dubbed CinCGAN) is proposed to jointly estimate the spectral magnitude and phase information stage by stage under unpaired data. In the first stage, we pretrain a magnitude Cycle-GAN to coarsely estimate the spectral magnitude of clean speech. In the second stage, we incorporate the pretrained CycleGAN with a complex-valued CycleGAN as a cycle-in-cycle structure to simultaneously recover phase information and refine the overall spectrum. Experimental results demonstrate that the proposed approach significantly outperforms previous baselines under non-parallel training. The evaluation on training the models with standard paired data also shows that CinCGAN achieves remarkable performance especially in reducing background noise and speech distortion.
Guochen Yu, Andong Li, Yinuo Guo, Hui Wang 0070, Chengshi Zheng
ICASSP6
2022 Dual-Branch Attention-In-Attention Transformer for Single-Channel Speech Enhancement
abstract
Curriculum learning begins to thrive in the speech enhancement area, which decouples the original spectrum estimation task into multiple easier sub-tasks to achieve better performance. Motivated by that, we propose a dual-branch attention-in-attention transformer dubbed DB-AIAT to handle both coarse- and fine-grained regions of the spectrum in parallel. From a complementary perspective, a magnitude masking branch is proposed to coarsely estimate the overall magnitude spectrum, and simultaneously a complex refining branch is elaborately designed to compensate for the missing spectral details and implicitly derive phase information. Within each branch, we propose a novel attention-in-attention transformer-based module to replace the conventional RNNs and temporal convolutional networks for temporal sequence modeling. Specifically, the proposed attention-in-attention transformer consists of adaptive temporal-frequency attention transformer blocks and an adaptive hierarchical attention module, aiming to capture long-term temporal-frequency dependencies and further aggregate global hierarchical contextual information. Experimental results on Voice Bank + DEMAND demonstrate that DB-AIAT yields state-of-the-art performance (e.g., 3.31 PESQ, 95.6% STOI and 10.79dB SSNR) over previous advanced systems with a relatively small model size (2.81M).
Guochen Yu, Andong Li, Chengshi Zheng, Yinuo Guo, Hui Wang 0070
ICASSP3
2022 Taylor, Can You Hear Me Now? A Taylor-Unfolding Framework for Monaural Speech Enhancement
abstract
While the deep learning techniques promote the rapid development of the speech enhancement (SE) community, most schemes only pursue the performance in a black-box manner and lack adequate model interpretability. Inspired by Taylor's approximation theory, we propose an interpretable decoupling-style SE framework, which disentangles the complex spectrum recovery into two separate optimization problems i.e., magnitude and complex residual estimation. Specifically, serving as the 0th-order term in Taylor's series, a filter network is delicately devised to suppress the noise component only in the magnitude domain and obtain a coarse spectrum. To refine the phase distribution, we estimate the sparse complex residual, which is defined as the difference between target and coarse spectra, and measures the phase gap. In this study, we formulate the residual component as the combination of various high-order Taylor terms and propose a lightweight trainable module to replace the complicated derivative operator between adjacent terms. Finally, following Taylor's formula, we can reconstruct the target spectrum by the superimposition between 0th-order and high-order terms. Experimental results on two benchmark datasets show that our framework achieves state-of-the-art performance over previous competing baselines in various evaluation metrics. The source code is available at https://github.com/Andong-Li-speech/TaylorSENet.
Andong Li, Shan You, Guochen Yu, Chengshi Zheng, Xiaodong Li 0002
IJCAI4
2022 A deep complex multi-frame filtering network for stereophonic acoustic echo cancellation
abstract
In hands-free communication system, the coupling between loudspeaker and microphone generates echo signal, which can severely influence the quality of communication.Meanwhile, various types of noise in communication environments further reduce speech quality and intelligibility.It is difficult to extract the near-end signal from the microphone signal within one step, especially in low signal-to-noise ratio scenarios.In this paper, we propose a deep complex network approach to address this issue.Specially, we decompose the stereophonic acoustic echo cancellation into two stages, including linear stereophonic acoustic echo cancellation module and residual echo suppression module, where both modules are based on deep learning architectures.A multi-frame filtering strategy is introduced to benefit the estimation of linear echo by capturing more interframe information.Moreover, we decouple the complex spectral mapping into magnitude estimation and complex spectrum refinement.Experimental results demonstrate that our proposed approach achieves stage-of-the-art performance over previous advanced algorithms under various conditions.
Linjuan Cheng, Chengshi Zheng, Andong Li, Yuquan Wu, Renhua Peng, Xiaodong Li 0002
INTERSPEECH2
2022 TMGAN-PLC: Audio Packet Loss Concealment using Temporal Memory Generative Adversarial Network
abstract
Real-time communications in packet-switched networks have become widely used in daily communication, while they inevitably suffer from network delays and data losses in constrained real-time conditions.To solve these problems, audio packet loss concealment (PLC) algorithms have been developed to mitigate voice transmission failures by reconstructing the lost information.Limited by the transmission latency and device memory, it is still intractable for PLC to accomplish high-quality voice reconstruction using a relatively small packet buffer.In this paper, we propose a temporal memory generative adversarial network for audio PLC, dubbed TMGAN-PLC, which is comprised of a novel nested-UNet generator and the time-domain/frequency-domain discriminators.Specifically, a combination of the nested-UNet and temporal featurewise linear modulation is elaborately devised in the generator to finely adjust the intra-frame information and establish inter-frame temporal dependencies.To complement the missing speech content caused by longer loss bursts, we employ multistage gated vector quantizers to capture the correct content and reconstruct the near-real smooth audio.Extensive experiments on the PLC Challenge dataset demonstrate that the proposed method yields promising performance in terms of speech quality, intelligibility, and PLCMOS.
Yuansheng Guan, Guochen Yu, Andong Li, Chengshi Zheng
INTERSPEECH4
2022 TaylorBeamformer: Learning All-Neural Beamformer for Multi-Channel Speech Enhancement from Taylor's Approximation Theory
abstract
While existing end-to-end beamformers achieve impressive performance in various front-end speech processing tasks, they usually encapsulate the whole process into a black box and thus lack adequate interpretability. As an attempt to fill the blank, we propose a novel neural beamformer inspired by Taylor's approximation theory called TaylorBeamformer for multi-channel speech enhancement. The core idea is that the recovery process can be formulated as the spatial filtering in the neighborhood of the input mixture. Based on that, we decompose it into the superimposition of the 0th-order non-derivative and high-order derivative terms, where the former serves as the spatial filter and the latter is viewed as the residual noise canceller to further improve the speech quality. To enable end-to-end training, we replace the derivative operations with trainable networks and thus can learn from training data. Extensive experiments are conducted on the synthesized dataset based on LibriSpeech and results show that the proposed approach performs favorably against the previous advanced baselines.
Andong Li, Guochen Yu, Chengshi Zheng, Xiaodong Li 0002
INTERSPEECH3
2022 Bifurcation and Reunion: A Loss-Guided Two-Stage Approach for Monaural Speech Dereverberation
Xiaoxue Luo, Chengshi Zheng, Andong Li, Yuxuan Ke, Xiaodong Li 0002
INTERSPEECH2
2022 Fully Automatic Balance between Directivity Factor and White Noise Gain for Large-scale Microphone Arrays in Diffuse Noise Fields
Weixin Meng, Chengshi Zheng, Xiaodong Li 0002
INTERSPEECH2
2022 Analysis of trade-offs between magnitude and phase estimation in loss functions for speech denoising and dereverberation
Xiaoxue Luo, Chengshi Zheng, Andong Li, Yuxuan Ke, Xiaodong Li 0002
Speech Commun.2
2022 Filtering and Refining: A Collaborative-Style Framework for Single-Channel Speech Enhancement
abstract
In low signal-to-noise ratio (SNR) acoustic scenarios, it remains fairly challenging to extract the target speech from its noisy mixture. In this paper, we propose a collaborative-style framework, namely, filtering and refining network (FRNet) for single-channel speech enhancement, recovering the complex spectrum of the target speech from coarse and fine-grained perspectives. Specifically, we devise a two-branch structure dubbed filtering-refining module (FRM). In the filtering block, the phase impact is ignored, and we only focus on coarse filtering in the magnitude domain. In the refining block, instead of predicting the irregular phase distribution directly, we estimate the complex residual for phase modification and spectrum rehabilitation, which takes the harmonic structure but with rather sparse energy distribution. By cascading FRMs repeatedly, we can reconstruct the target spectrum progressively. Furthermore, we propose a two-stream feature encoder to extract the feature representation of magnitude and phase individually, and the utilization of feature recalibration layers can preserve the prominent information from multiple scales. Extensive experiments are conducted on the WSJ0-SI84, Voicebank+Demand, and DNS-Challenge corpora. Evaluation results show that the proposed system performs favorably against previous advanced systems and achieves overall state-of-the-art performance in PESQ, ESTOI, SDR, and DNSMOS metrics.
Andong Li, Chengshi Zheng, Guochen Yu, Juanjuan Cai, Xiaodong Li 0002
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 DBT-Net: Dual-Branch Federative Magnitude and Phase Estimation With Attention-in-Attention Transformer for Monaural Speech Enhancement
abstract
The decoupling-style concept begins to ignite in the speech enhancement area, which decouples the original complex spectrum estimation task into multiple easier sub-tasks (i.e., the magnitude-only recovery and residual complex spectrum estimation), resulting in better performance and easier interpretability. In this paper, we propose a dual-branch federative magnitude and phase estimation framework, dubbed DBT-Net, for monaural speech enhancement, aiming at recovering the coarse- and fine-grained regions of the overall spectrum in parallel. From the complementary perspective, the magnitude estimation branch is designed to filter out dominant noise components in the magnitude domain, while the complex spectrum purification branch is elaborately designed to inpaint the missing spectral details and implicitly estimate the phase information in the complex-valued spectral domain. To facilitate the information flow between each branch, interaction modules are introduced to leverage features learned from one branch, so as to suppress the undesired parts and recover the missing components of the other branch. Instead of adopting the conventional RNNs and temporal convolutional networks for sequence modeling, we employ a novel attention-in-attention transformer-based network within each branch for better feature learning. More specially, it is composed of several adaptive spectro-temporal attention transformer-based modules and an adaptive hierarchical attention module, aiming to capture long-term time-frequency dependencies and further aggregate intermediate hierarchical contextual information. Comprehensive evaluations on the WSJ0-SI84 + DNS-Challenge and VoiceBank + DEMAND dataset demonstrate that the proposed approach consistently outperforms previous advanced systems and yields state-of-the-art performance in terms of speech quality and intelligibility.
Guochen Yu, Andong Li, Hui Wang 0070, Yuxuan Ke, Chengshi Zheng
IEEE ACM Trans. Audio Speech Lang. Process.6
2021 ICASSP 2021 Deep Noise Suppression Challenge: Decoupling Magnitude and Phase Optimization with a Two-Stage Deep Network
abstract
It remains a tough challenge to recover the speech signals contaminated by various noises under real acoustic environments. To this end, we propose a novel system for denoising in the complicated applications, which is mainly comprised of two pipelines, namely a two-stage network and a post-processing module. The first pipeline is proposed to decouple the optimization problem w.r.t. magnitude and phase, i.e., only the magnitude is estimated in the first stage and both of them are further refined in the second stage. The second pipeline aims to further suppress the remaining unnatural distorted noise, which is demonstrated to sufficiently improve the subjective quality. In the ICASSP 2021 Deep Noise Suppression (DNS) Challenge, our submitted system ranked top-1 for the real-time track 1 in terms of Mean Opinion Score (MOS) with ITU-T P.808 framework.
Andong Li, Xiaoxue Luo, Chengshi Zheng, Xiaodong Li 0002
ICASSP4
2021 ICASSP 2021 Acoustic Echo Cancellation Challenge: Integrated Adaptive Echo Cancellation with Time Alignment and Deep Learning-Based Residual Echo Plus Noise Suppression
abstract
This paper describes a three-stage acoustic echo cancellation (AEC) and suppression framework for the ICASSP 2021 AEC Challenge. In the first stage, a partitioned block frequency domain adaptive filtering is implemented to cancel the linear echo components without introducing the near-end speech distortion, where we compensate the time delay between the far-end reference signal and the micro-phone signal beforehand. In the second stage, a deep complex U-Net integrated with gated recurrent unit is proposed to further suppress the residual echo components. In the last stage, an extremely tiny deep complex U-Net is trained to suppress non-speech residual components that have not been suppressed completely in the second stage, which can also further increase the echo return loss enhancement (ERLE) without increasing the computational complexity dramatically. Experimental results show that the proposed three-stage framework can get the ERLE higher than 50 dB in both single-talk and double-talk scenarios, and perceptual evaluation of speech quality can be improved about 0.75 in double-talk scenarios. The proposed framework outperforms the AEC-Challenge baseline ResRNN by 0.12 points in terms of the MOS.
Renhua Peng, Linjuan Cheng, Chengshi Zheng, Xiaodong Li 0002
ICASSP3
2021 A Simultaneous Denoising and Dereverberation Framework with Target Decoupling
abstract
Background noise and room reverberation are regarded as two major factors to degrade the subjective speech quality.In this paper, we propose an integrated framework to address simultaneous denoising and dereverberation under complicated scenario environments.It adopts a chain optimization strategy and designs four sub-stages accordingly.In the first two stages, we decouple the multi-task learning w.r.t.complex spectrum into magnitude and phase, and only implement noise and reverberation removal in the magnitude domain.Based on the estimated priors above, we further polish the spectrum in the third stage, where both magnitude and phase information are explicitly repaired with the residual learning.Due to the data mismatch and nonlinear effect of DNNs, the residual noise often exists in the DNN-processed spectrum.To resolve the problem, we adopt a light-weight algorithm as the post-processing module to capture and suppress the residual noise in the non-active regions.In the Interspeech 2021 Deep Noise Suppression (DNS) Challenge, our submitted system ranked top-1 for the real-time track in terms of Mean Opinion Score (MOS) with ITU-T P.835 framework.
Andong Li, Xiaoxue Luo, Guochen Yu, Chengshi Zheng, Xiaodong Li 0002
Interspeech5
2021 Know Your Enemy, Know Yourself: A Unified Two-Stage Framework for Speech Enhancement
Andong Li, Yuxuan Ke, Chengshi Zheng, Xiaodong Li 0002
Interspeech4
2021 Acoustic Echo Cancellation Using Deep Complex Neural Network with Nonlinear Magnitude Compression and Phase Information
abstract
This paper describes a two-stage acoustic echo cancellation (AEC) and suppression framework for the INTER-SPEECH2021 AEC Challenge.In the first stage, four parallel partitioned block frequency domain adaptive filters are used to cancel the linear echo components, where the far-end signal is delayed 0ms, 320ms, 640ms and 960ms for these four adaptive filters, respectively, thus a maximum 1280 ms time delay can be well handled in the blind test dataset.The error signal with minimum energy and its corresponding reference signal are chosen as the input for the second stage, where a gate complex convolutional recurrent neural network (GCCRN) is trained to further suppress the residual echo, late reverberation and environmental noise simultaneously.To improve the performance of GCCRN, we compress both the magnitude of the error signal and that of the far-end reference signal, and then the two compressed magnitudes are combined with the phase of the error signal to regenerate the complex spectra as the input features of GCCRN.Numerous experimental results show that the proposed framework is robust to the blind test dataset, and achieves a promising result with the P.808 evaluation.
Renhua Peng, Linjuan Cheng, Chengshi Zheng, Xiaodong Li 0002
Interspeech3
2021 Distributed node-specific block-diagonal LCMV beamforming in wireless acoustic sensor networks
Minmin Yuan, Yuxuan Ke, Chengshi Zheng, Xiaodong Li 0002
Signal Process.4
2021 Finite data performance analysis of one-bit MVDR and phase-only MVDR
Weixin Meng, Yuxuan Ke, Chengshi Zheng, Xiaodong Li 0002
Signal Process.4
2021 Corrigendum to 'Finite data performance analysis of one-bit MVDR and phase-only MVDR' [Signal Processing 183 (2021) Article 108018]
Weixin Meng, Yuxuan Ke, Chengshi Zheng, Xiaodong Li 0002
Signal Process.4
2021 A two-stage complex network using cycle-consistent generative adversarial networks for speech enhancement
Guochen Yu, Hui Wang 0070, Qin Zhang 0009, Chengshi Zheng
Speech Commun.5
2021 Two Heads are Better Than One: A Two-Stage Complex Spectral Mapping Approach for Monaural Speech Enhancement
abstract
For challenging acoustic scenarios as low signal-to-noise ratios, current speech enhancement systems usually suffer from performance bottleneck in extracting the target speech from the mixtures within one step. To address this issue, we propose a novel complex spectral mapping approach with a two-stage pipeline for monaural speech enhancement in the time-frequency domain. The proposed algorithm aims to decouple the primal problem into multiple sub-problems, which follows the classic proverb, “two heads are better than one”. More specifically, in the first stage, only magnitude is estimated, which is incorporated with the noisy phase to obtain a coarse complex spectrum estimation. To facilitate the previous estimation, in the second stage, an auxiliary network serves as the post-processing module, where residual noise is further suppressed and the phase information is effectively modified. The global residual connection strategy is adopted in the second stage to accelerate the training convergence speed. To alleviate the parameter burden caused by the multi-stage pipeline, we propose a light-weight temporal convolutional module, which substantially decreases the trainable parameters and obtains even better objective performance over the original version. We conduct extensive experiments on three standard corpora, including WSJ0-SI84, DNS Challenge dataset, and Voice Bank + DEMAND dataset. Objective test results demonstrate that our proposed approach achieves state-of-the-art performance over previous advanced systems under various conditions. Meanwhile, subjective listening test results further validate the superiority of our proposed method in terms of subjective quality.
Andong Li, Chengshi Zheng, Cunhang Fan, Xiaodong Li 0002
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 A Recursive Network with Dynamic Attention for Monaural Speech Enhancement
abstract
For continuous speech processing, dynamic attention is helpful in preferential processing, which has already been shown by the auditory dynamic attending theory.Accordingly, we propose a framework combining dynamic attention and recursive learning together for monaural speech enhancement.Apart from a major noise reduction network, we design a separated sub-network, which adaptively generates the attention distribution to control the information flow throughout the major network.Recursive learning is introduced to dynamically reduce the number of trainable parameters by reusing a network for multiple stages, where the intermediate output in each stage is corrected with a memory mechanism.By doing so, a more flexible and better estimation can be obtained.We conduct experiments on TIMIT corpus.Experimental results show that the proposed architecture obtains consistently better performance than recent state-of-the-art models in terms of both PESQ and STOI scores.The code is provided at https://github.com/Andong-Li-speech/DARCN.
Andong Li, Chengshi Zheng, Cunhang Fan, Renhua Peng, Xiaodong Li 0002
INTERSPEECH2
2020 Wideband sparse Bayesian learning for off-grid binaural sound source localization
Jiance Ding, Chengshi Zheng, Xiaodong Li 0002
Signal Process.3
2018 A perceptually motivated LP residual estimator in noisy and reverberant environments
Renhua Peng, Zheng-Hua Tan, Xiaodong Li 0002, Chengshi Zheng
Speech Commun.4
2018 Statistical Analysis of the Multichannel Wiener Filter Using a Bivariate Normal Distribution for Sample Covariance Matrices
abstract
This paper studies the statistical performance of the multichannel Wiener filter (MWF) when the weights are computed using estimates of the sample covariance matrices of the noisy and the noise signals. It is well known that the optimal weights of the minimum variance distortionless response beamformer are only determined by the noisy sample covariance matrix or the noise sample covariance matrix, while those of the MWF are determined by both of them. Therefore, the difficulty increases dramatically in statistically analyzing the MWF when compared to analyzing the MVDR, where the main reason is that expressing the general joint probability density function (p.d.f.) of the two sample covariance matrices presented a Hitherto unsolved problem, to the best of our knowledge. For a deeper insight into the statistical performance of the MWF, this paper first introduces a bivariate normal distribution to approximately model the joint p.d.f. of the noisy and the noise sample covariance matrices. Each sample covariance matrix is approximately modeled by a random scalar multiplied by its true covariance matrix. This approximation is designed to preserve both the bias and the mean squared error of the matrix with respect to a natural distance on covariance matrices. The correlation of the bivariate normal distribution, referred to as the sample covariance matrices intrinsic correlation coefficient, captures all second-order dependencies of the noisy and the noise sample covariance matrices. By using the proposed bivariate normal distribution, the performance of the MWF can be predicted from the derived analytical expressions and many interesting results are revealed. As an example, the theoretical analysis demonstrates that the MWF performance may degrade in terms of noise reduction and signal-to-noise-ratio improvement when using more sensors in some noise scenarios.
Chengshi Zheng, Antoine Deleforge, Xiaodong Li 0002, Walter Kellermann
IEEE ACM Trans. Audio Speech Lang. Process.1
2016 Analysis of Additional Stable Gain by Frequency Shifting for Acoustic Feedback Suppression using Statistical Room Acoustics
abstract
Currently, a quantitative prediction of the performance of the frequency shifting (FS) approach for stabilizing acoustic feedback loops of public address (PA) systems is only possible in case of negligibly small direct sound components in the feedback path. By employing a statistical room acoustics model, this letter derives a novel analytical expression for the additional stable gain (ASG) with the FS approach. This allows a quantitative prediction of the performance of FS in the presence of significant direct sound components for the first time. Simulation results confirm the theoretical analysis.
Chengshi Zheng, Christian Hofmann 0001, Xiaodong Li 0002, Walter Kellermann
IEEE Signal Process. Lett.1
2014 A Constrained MMSE LP Residual Estimator for Speech Dereverberation in Noisy Environments
abstract
After revealing that both late reverberation and noise are additive interference components in the residual domain, this paper proposes to suppress these additive interference components by using a constrained minimum mean square error linear prediction (LP) residual estimator, where the optimal filter can be obtained by the generalized singular value decomposition. We propose to estimate the LP residuals for both late reverberation and noise continuously, which is based on the non-VAD related noise power spectral density estimator and the incessant late reverberant spectral variance estimator. The non-intrusive objective measure and the PESQ show that the proposed algorithm is better than traditional LP residual-based algorithms and spectral subtraction-based algorithms.
Chengshi Zheng, Renhua Peng, Xiaodong Li 0002
IEEE Signal Process. Lett.1
2014 On Generalized Auto-Spectral Coherence Function and Its Applications to Signal Detection
abstract
Considering that spectral components of one random process are not necessarily independent for all types of signals, this paper defines a generalized auto-spectral coherence function (GAS-CF) to measure this spectral correlation. The GAS-CF is a generalization of the temporal coherence function and the spectral coherence function, where they have already been successfully applied to detect howling components and transient noise components, respectively. After defining the GAS-CF, this paper studies its statistical properties in detail. Simulation results show that the proposed GAS-CF can be applied to detect different types of signals, including transient noise, howling frequency and chirp signal, in a simple way.
Chengshi Zheng, Hefei Yang, Xiaodong Li 0002
IEEE Signal Process. Lett.1
2013 A Statistical Analysis of Two-Channel Post-Filter Estimators in Isotropic Noise Fields
abstract
This paper derives explicit expressions of the probability density functions of the two-channel post-filter estimators in isotropic noise fields to study their statistical properties. According to the analysis results, three methods are proposed to improve the performance of the noise filed coherence (NFC)-based post-filter estimator.
Chengshi Zheng, Renhua Peng, Xiaodong Li 0002
IEEE Trans. Speech Audio Process.1
2012 On second-order statistics of log-periodogram and cepstral coefficients for processes with mixed spectra
Chengshi Zheng
Signal Process.1
2011 Two-channel post-filtering based on adaptive smoothing and noise properties
abstract
This paper studies the statistical properties of the gain functions, which are often used for two-channel post-filtering (TC-PF) algorithms. We reveal that the smoothing factor has a significant impact on both noise reduction and musical noise. When the smoothing factor increases, noise reduction can be improved and musical noise can be reduced simultaneously. However, the smoothing factor could not be too close to one because the system can only be assumed to be time-invariant for short durations. To solve this problem, this paper proposes an adaptive smoothing scheme by detecting the sudden change of the system. Moreover, the residual noise floor is adaptively chosen based on the structure of the noise power spectral density (NPSD) to further suppress the tonal noise components. Experimental results show the better performance of the proposed algorithm in terms of the segmental signal to-noise-ratio (SNR) and the PESQ improvements.
Chengshi Zheng, Yi Zhou 0014, Xiaohu Hu, Xiaodong Li 0002
ICASSP1
2009 Acoustical Vehicle Detection Based on Bispectral Entropy
abstract
A vehicle detection algorithm based on Bispectral entropy is proposed in this paper. Based on Quadratic-Phase Coupling(QPC) analysis, Bispectral entropy, as the complexity measure of bispectra, is calculated from the vehicle acoustic signal database which is acquired in several real world experiments. Furthermore, an effective bispectral entropy-based algorithm is developed for vehicle detection. Experiments show that the proposed algorithm can achieve longer distance alert than other two detectors.
Ming Bao, Chengshi Zheng, Xiaodong Li 0002, Jun Yang 0004
IEEE Signal Process. Lett.2
2008 On the relationship of non-parametric methods for coherence function estimation
Chengshi Zheng, Mingyuan Zhou, Xiaodong Li 0002
Signal Process.1