Feiran Yang 0001

dblp:118/9346-1 · DBLP profile ↗
← Back
36ranked-venue papers
12as first author
22since 2021 · last 2026
0000-0002-1734-3785ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 7 first-author · 15 since 2021Artificial intelligence and machine learning · 13 · 5 first-author · 9 since 2021
YearPublicationVenuePosition
2026 A Schrödinger bridge-based two-stage model for bone-conducted speech restoration
Feiran Yang 0001, Chunguo Li
Speech Commun.3
2026 Convergence Behavior of Alternated-Constrained Partitioned-Block Frequency-Domain Adaptive Filters
Zhengqiang Luo, Weiqiang Tan, Feiran Yang 0001
IEEE Signal Process. Lett.3
2026 Cepstrum Enhancement-Based Nonnegative Matrix Factorization for Blind Source Separation
abstract
Nonnegative matrix factorization (NMF) has proven to be a powerful source spectrogram representation in many audio blind source separation (BSS) methods. However, NMFbased methods may still suffer from the well-known permutation problem especially when the number of NMF basis vectors is large. This letter presents a cepstrum enhancement-based NMF approach that effectively exploits the harmonic structure of audio signals. In oracle NMF-based methods, the source variances are estimated by directly applying NMF to the spectrograms of the separated signals. In the presented approach, however, the spectrogram of each separated source is firstly smoothed using a cepstrum thresholding method, and then NMF is applied to the enhanced spectrogram for source variance estimation. The cepstrum thresholding enhances the harmonic structure of the source signals while removing the interfering components from other sources, which helps alleviate the permutation problem. Also, the proposed cepstrum enhancement method is decoupled from the specific source model and can be easily incorporated into other NMF-based BSS methods as a plug-and-play module. Experimental results validate the effectiveness of the proposed method.
Mengchao Fan, Feiran Yang 0001, Jun Yang 0004
IEEE Signal Process. Lett.3
2025 SMRU-Lite: Efficient Low-Complexity Speech Enhancement Model with Uncertainty Estimation
abstract
Although neural network-based speech enhancement models perform much better than their traditional counterparts, their substantial computational demands make it challenging for real-time applications on edge devices. Moreover, compact models often exhibit weak generalization in complex and out-of-domain scenarios. In this paper, we propose an efficient model based on our previous work Split-and-Merge Recurrent-based UNet (SMRU). The proposed model achieves a significant reduction in computational load through the incorporation of Skip-RNN layers and an attention-based sub-band compression module. Moreover, the employment of a two-stage uncertainty-driven loss function for aleatoric uncertainty capture leads to enhanced generalization and denoising performance without increasing the computational complexity during inference. Experimental results demonstrate that our model not only surpasses the original SMRU but also outperforms recently proposed lightweight models with similar computational cost (approximately 200M MACs). Furthermore, our model exhibits strong generalization in cross-corpus test sets, making it a promising solution for real-time speech enhancement applications.
Zhihang Sun, Feiran Yang 0001, Rilin Chen, Chunguo Li
IJCNN4
2025 Online Blind Speech Separation Using Time-Varying All-Pole Source Model
abstract
Independent vector analysis (IVA) is a state-of-the-art blind source separation (BSS) method that utilizes higher-order correlation between frequency components in each source. However, online-IVA assumes a spherical multivariate distribution and it is not flexible. This letter presents an online BSS method based on the time-varying all-pole (TVAP) source model, which can better capture the spectral envelope of the source signal. The cost function is formulated under the maximum likelihood criterion, where the demixing matrix is updated using iterative projection algorithm and the TVAP model parameters is optimized based on a fixed-point iteration framework. Furthermore, the TVAP model with a low-order all-pole filter effectively solves the permutation problem. Experimental results show that the proposed algorithm outperforms the online-IVA methods for both simulated and live-recorded datasets.
Feiran Yang 0001, Jun Yang 0004
IEEE Signal Process. Lett.2
2024 Design of frequency-invariant uniform concentric circular arrays with first-order directional microphones
Feiran Yang 0001, Zhaoli Yan, Jun Yang 0004
Signal Process.2
2024 Restoration of Bone-Conducted Speech With U-Net-Like Model and Energy Distance Loss
abstract
Bone-conducted speech is less susceptible to ambient noise interference, but it suffers from poor speech quality due to the limited bandwidth. In this letter, we propose a U-Net-like network for the restoration of bone-conducted speech in the time domain. The proposed network consists of residual-connected one-dimensional convolutions and shifted window-based attention modules, which can model long-term dependencies crucial in speech processing. We find that the prevalent time-domain${{l}_{1}}$loss may be insufficient for the generation of high-frequency information absent in bone-conducted speech. To address this issue, we propose to utilize the generalized energy distance loss based on multi-scale Mel spectrograms as the objective function. Experimental results on the ESMB dataset validate the efficacy of our proposed method in restoration of bone-conducted speech. The proposed approach significantly outperforms two recent time-domain benchmarks, DPT-EGNet and EBEN, in terms of PESQ and STOI metrics.
Changtao Li, Feiran Yang 0001, Jun Yang 0004
IEEE Signal Process. Lett.2
2024 Selective-Memory Meta-Learning With Environment Representations for Sound Event Localization and Detection
abstract
Environment shifts and conflicts present significant challenges for learning-based sound event localization and detection (SELD) methods. SELD systems, when trained in particular acoustic settings, often show restricted generalization capabilities for diverse acoustic environments. Furthermore, obtaining annotated samples for spatial sound events is notably costly. Deploying a SELD system in a new environment requires extensive time for re-training and fine-tuning. To overcome these challenges, we propose environment-adaptive Meta-SELD, designed for efficient adaptation to new environments using minimal data. Our method specifically utilizes computationally synthesized spatial data and employs Model-Agnostic Meta-Learning (MAML) on a pre-trained, environment-independent model. The method then utilizes fast adaptation to unseen real-world environments using limited samples from the respective environments. Inspired by the Learning-to-Forget approach, we introduce the concept of selective memory as a strategy for resolving conflicts across environments. This approach involves selectively memorizing target-environment-relevant information and adapting to the new environments through the selective attenuation of model parameters. In addition, we introduce environment representations to characterize different acoustic settings, enhancing the adaptability of our attenuation approach to various environments. We evaluate our proposed method on the development set of the Sony-TAu Realistic Spatial Soundscapes 2023 (STARSS23) dataset and computationally synthesized scenes. Experimental results demonstrate the superior performance of the proposed method compared to conventional supervised learning methods, particularly in localization.
Jinbo Hu, Yin Cao, Ming Wu 0005, Qiuqiang Kong, Feiran Yang 0001, Mark D. Plumbley, Jun Yang 0004
IEEE ACM Trans. Audio Speech Lang. Process.5
2024 A Two-Stage Approach to Quality Restoration of Bone-Conducted Speech
abstract
Bone-conducted speech is not susceptible to background noise but suffers from poor speech quality and intelligibility due to the limited bandwidth. This paper proposes a two-stage approach to restore the quality of bone-conducted speech, namely, bandwidth extension and speech vocoder. In the first stage, a deep neural network is trained to learn mappings from a low-resolution representation of the bone-conducted speech, i.e., log Mel-scale spectrogram, to that of the air-conducted speech, which extends the bandwidth of the bone-conducted speech. In the second stage, a speech vocoder is employed to transform the extended log Mel-scale spectrogram of the bone-conducted speech back to time-domain waveforms. Due to the many-to-many correspondence between the air-conducted and bone-conducted speech, supervised learning may not be the best training protocol for the bone-conducted/air-conducted feature mapping. We thus propose to leverage adversarial training to further improve the bandwidth extension performance in the first stage. The two stages are decoupled and can be trained independently. The vocoder is trained on a large multi-speaker dataset and can generalize well to unknown speakers. Also, the vocoder can help to remedy the spectral artifacts introduced in the bandwidth extension stage. Objective and subjective evaluations on ESMB dataset show that the proposed two-stage system substantially outperforms existing bone-conducted speech enhancement systems.
Changtao Li, Feiran Yang 0001, Jun Yang 0004
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 Theoretical Analysis of Maclaurin Expansion Based Linear Differential Microphone Arrays and Improved Solutions
abstract
Linear differential microphone arrays (LDMAs) are becoming popular due to their potentially high directional gain and frequency-invariant beampattern. By increasing the number of microphones, the Maclaurin expansion-based LDMAs address the inherently poor robustness problem of the conventional LDMA at low frequencies. However, this method encounters severe beampattern distortion and the deep nulls problem in the white noise gain (WNG) and the directivity factor (DF) at high frequencies as the number of microphones increases. In this paper, we reveal that the severe beampattern distortion is attributed to the deviation term of the synthesized beampattern while the deep nulls problem in the WNG and the DF is attributed to the violation of the distortionless constraint in the desired direction. We then propose two new design methods to avoid the degraded performance of LDMAs. Compared to the Maclaurin series expansion-based method, the first method additionally imposes the distortionless constraint in the desired direction, and the deep nulls problem in the WNG and the DF can be avoided. The second method explicitly requires the response of the higher order spatial directivity pattern in the deviation term to be zero, and thus the beampattern distortion can be avoided. By choosing the frequency-wise parameter that determines the number of the considered higher order spatial directivity patterns, the second method enables a good trade-off between the WNG and the beampattern distortion. Simulations exemplify the superiority of the proposed method against existing methods in terms of the robustness and the beampattern distortion.
Feiran Yang 0001, Xiaoqing Hu, Jun Yang 0004
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 Multichannel Linear Prediction-Based Speech Dereverberation Considering Sparse and Low-Rank Priors
abstract
This article addresses the multi-channel linear prediction (MCLP)-based speech dereverberation problem by jointly considering the sparsity and low-rank priors of speech spectrograms. We utilize the complex generalized Gaussian (CGG) distribution as the source model and the generalized nonnegative matrix factorization (NMF) as the spectral model. The difference between the presented model and existing ones for MCLP is twofold. First, we adopt the CGG distribution with a time-frequency-variant scale parameter instead of that with a time-frequency-invariant scale parameter. Second, the time-frequency-varying scale parameter is approximated by NMF in a low-rank manner. Based on the maximum-likelihood criterion, speech dereverberation is formulated as an optimization problem that minimizes the prediction error weighted by the reciprocal of sparse and low-rank parameters. A convergence-guaranteed algorithm is derived to estimate the parameters using the majorization-minimization technology. The WPE, NMF-based WPE and CGG-based WPE can be treated as special cases of the proposed method with different shape and domain parameters. As a byproduct, the proposed method provides a simple and elegant way to derive the CGG-based WPE algorithm. A series of experiments show the superiority of the proposed method over WPE, NMF-based WPE and CGG-based WPE methods.
Taihui Wang, Feiran Yang 0001, Jun Yang 0004
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 A Perspective on Fully Steerable Differential Beamformers for Circular Arrays
abstract
Both the null-constrained method and the Jacobi-Anger expansion method are commonly adopted to design the fully steerable circular differential microphone arrays (CDMAs). So far, these two methods have been independently studied, and it is unclear why they exhibit pronounced performance difference and whether they have a potential connection. In this letter, we develop a unified framework for the design of the fully steerable CDMA, which jointly utilizes the null constraints and the Jacobi-Anger expansion. We show that the existing two methods can be treated as special cases of the proposed approach and they can indeed be connected. We then reveal that the null-constrained CDMA can be viewed as the regularized Jacobi-Anger expansion-based CDMA, which theoretically explains why the former can alleviate the deep nulls problem in white noise gain in certain cases and why it still suffers from the severe deep nulls in other cases. We further explain why the magnitude of the synthesized beampattern of the Jacobi-Anger expansion-based method in nulls' direction can be much larger than that of the null-constrained method. Simulations support the theoretical results well.
Feiran Yang 0001, Jun Yang 0004
IEEE Signal Process. Lett.2
2022 A Track-Wise Ensemble Event Independent Network for Polyphonic Sound Event Localization and Detection
abstract
Polyphonic sound event localization and detection (SELD) aims at detecting types of sound events with corresponding temporal activities and spatial locations. In this paper, a trackwise ensemble event independent network with a novel data augmentation method is proposed. The proposed model is based on our previous proposed Event-Independent Network V2 and is extended by conformer blocks and dense blocks. The track-wise ensemble model with track-wise output format is proposed to solve an ensemble model problem for track-wise output format that track permutation may occur among different models. The data augmentation approach contains several data augmentation chains, which are composed of random combinations of several data augmentation operations. The method also utilizes log-mel spectrograms, intensity vectors, and Spatial Cues-Augmented Log-Spectrogram (SALSA) for different models. We evaluate our proposed method in the Task of the L3DAS22 challenge and obtain the top ranking solution with a location-dependent F-score to be 0.699. Source code is released1.
Jinbo Hu, Yin Cao, Ming Wu 0005, Qiuqiang Kong, Feiran Yang 0001, Mark D. Plumbley, Jun Yang 0004
ICASSP5
2022 A Multi-Task Learning Method for Weakly Supervised Sound Event Detection
abstract
In weakly supervised sound event detection (SED), only coarse-grained labels are available, and thus the supervision information is quite limited. To fully utilize prior knowledge of the time-frequency masks of each sound event, we propose a novel multi-task learning (MTL) method that takes SED as the main task and source separation as the auxiliary task. For active events, we minimize the overlap of their masks as the segment loss to learn distinguishing features. For inactive events, the proposed method measures the activity of masks as silent loss to reduce the insertion error. The auxiliary source separation task calculates an extra penalty according to the shared masks, which can further incorporate prior knowledge in the form of regularization constraints. We demonstrated that the proposed method can effectively reduce the insertion error and achieve a better performance in SED task than single-task methods.
Sichen Liu 0002, Feiran Yang 0001, Fang Kang, Jun Yang 0004
ICASSP2
2022 The Role of Long-Term Dependency in Synthetic Speech Detection
abstract
Although much progress has been made in synthetic speech detection, there lacks comprehensive analysis of the essential differences between spoofed and genuine speech. We here utilize supervised contrastive loss originated from contrastive learning as an analytical tool to characterize the class similarity structure of ASVspoof 2019 logical access (LA) dataset, which shows that an ideal back-end classifier for synthetic speech detection should have the ability to capture long-term dependencies. Recently, Transformer has been found to have an excellent ability in learning long-term dependencies of input data. We hence propose a back-end classifier based on Transformer Encoder for synthetic speech detection. Convolution blocks are added before the Transformer Encoder, which leverages inductive biases to improve the generalization ability. Compared to two-dimensional convolution, one-dimensional convolution makes better architectural assumptions about the input speech features, which helps with modeling long-term dependencies and decreases the risk of overfitting. The proposed Transformer combined with one-dimensional convolution has fewer parameters than most existing back-end classifiers, and achieves an equal error rate of 1.06% and a minimum tandem detection cost function metric of 0.0345 when evaluated on ASVspoof 2019 LA dataset, which is one of the best models reported in the literature.
Changtao Li, Feiran Yang 0001, Jun Yang 0004
IEEE Signal Process. Lett.2
2022 Insights Into the MMSE-Based Frequency-Invariant Beamformers for Uniform Circular Arrays
abstract
This letter provides new insights into the minimum mean square error (MMSE)-based frequency-invariant beamformer with the uniform circular array (FIB-UCA). By exploiting the symmetry of the UCA, we derive an explicit form of the white noise gain (WNG), the directivity factor (DF) and the MSE, and also derive the corresponding approximate expressions at low frequencies. We then investigate the impact of the look direction on the performance of the MMSE-based FIB-UCA. Interestingly, we further prove that for a UCA with the fixed aperture, the MMSE-based FIB with 2$N$microphones achieves a similar WNG as that with 4$N$microphones at low frequencies if the$N$th-order beampattern is steered to microphone angles. We also derive a unified explicit form of weighting vector and reveal that the MMSE-based FIB-UCA reduces to Jacobi-Anger expansion-based differential beamformer with the UCA at low frequencies.
Feiran Yang 0001, Jun Yang 0004
IEEE Signal Process. Lett.2
2022 Analysis of Unconstrained Partitioned-Block Frequency-Domain Adaptive Filters
abstract
Previously, we have presented a performance analysis of the unconstrained partitioned-block frequency-domain adaptive filter (PBFDAF) with 50% overlap. It was found that the unconstrained PBFDAF always converges to a biased solution, and hence the mean-square deviation (MSD) learning curve does not match the mean-square error (MSE) learning curve. To deal with this problem, we present an alternative theoretical model of the unconstrained PBFDAF by introducing a modified time-domain weight vector. We show that the new weight vector can converge to the true system impulse response in the mean sense. We also derive the analytical expressions for the MSD and the MSE. The analysis here is equivalent to that in our previous work, but the former is much simpler and easier to handle. Also, the MSD learning curve calculated from the new weight vector is in good agreement with the MSE learning curve. Computer simulations support the analytical results well.
Feiran Yang 0001
IEEE Signal Process. Lett.1
2022 Convolutive Transfer Function-Based Multichannel Nonnegative Matrix Factorization for Overdetermined Blind Source Separation
abstract
Most multichannel blind source separation (BSS) approaches rely on a spatial model to encode the transfer functions from sources to microphones and a source model to encode the source power spectral density. The rank-1 spatial model has been widely exploited in independent component analysis (ICA), independent vector analysis (IVA), and independent low-rank matrix analysis (ILRMA). The full-rank spatial model is also considered in many BSS approaches, such as full-rank spatial covariance matrix analysis (FCA), multichannel nonnegative matrix factorization (MNMF), and FastMNMF, which can improve the separation performance in the case of long reverberation times. This paper proposes a new MNMF framework based on the convolutive transfer function (CTF) for overdetermined BSS. The time-domain convolutive mixture model is approximated by a frequency-wise convolutive mixture model instead of the widely adopted frequency-wise instantaneous mixture model. The iterative projection algorithm is adopted to estimate the demixing matrix, and the multiplicative update rule is employed to estimate nonnegative matrix factorization (NMF) parameters. Finally, the source image is reconstructed using a multichannel Wiener filter. The advantages of the proposed method are twofold. First, the CTF approximation enables us to use a short window to represent long impulse responses. Second, the full-rank spatial model can be derived based on the CTF approximation and slowly time-variant source variances, and close relationships between the proposed method and ILRMA, FCA, MNMF and FastMNMF are revealed. Extensive experiments show that the proposed algorithm achieves a higher separation performance than ILRMA and FastMNMF in reverberant environments.
Taihui Wang, Feiran Yang 0001, Jun Yang 0004
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Analysis of Deficient-Length Partitioned-Block Frequency-Domain Adaptive Filters
abstract
This paper studies the convergence behavior of the partitioned-block frequency-domain adaptive filters (PBFDAF) for under-modeling scenarios. We focus on a family of the overlap-save PBFDAF algorithms with 50% overlap, including both of the constrained and unconstrained versions. The stochastic analysis of the constrained and unconstrained algorithms is carried out individually due to their convergence differences. For each algorithm, the frequency-domain error vector and the update equations are transformed into the time-domain counterparts, so we can analyze their convergence behavior completely in the time domain. We present the mean and mean-square convergence behavior of the augmented weight-error vector, and we obtain the closed-form expressions for the learning curve and the steady-state solutions. Based on the solution of the steady-state weight-error vector, we analyze if each version of the PBFDAF algorithm converges to the true solution and the Wiener solution. The theoretical model gains new insights into the convergence behavior of the deficient-length PBFDAF algorithms. The computer simulations support the theoretical model very well.
Feiran Yang 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 Real-Time Independent Vector Analysis Using Semi-Supervised Nonnegative Matrix Factorization as a Source Model
Taihui Wang, Feiran Yang 0001, Jun Yang 0004
Interspeech2
2021 A Partitioned-Block Frequency-Domain Adaptive Kalman Filter for Stereophonic Acoustic Echo Cancellation
Feiran Yang 0001, Yuepeng Li, Shidong Shang
Interspeech2
2021 Real-Time Independent Vector Analysis with a Deep-Learning-Based Source Model
abstract
In this paper, we present a real-time blind source separation (BSS) algorithm, which unifies the independent vector analysis (IVA) as a spatial model and a deep neural network (DNN) as a source model. The auxiliary-function based IVA (Aux-IVA) is utilized to update the demixing matrix, and the required time-varying variance of the speech source is estimated by a DNN. The DNN could provide a more accurate source model, which then helps to optimize the spatial model. In addition, because the DNN is used to estimate the source variance instead of the source power spectrogram, the size of DNN can be reduced significantly. Experiment results show that the joint utilization of the model-based approach and the data-driven approach provides a more efficient solution than just alone in terms of convergence rate and source separation performance.
Fang Kang, Feiran Yang 0001, Jun Yang 0004
SLT2
2020 Convergence analysis of the conventional filtered-x affine projection algorithm for active noise control
Feiran Yang 0001, Jun Yang 0004
Signal Process.2
2020 Optimum step-size control for a variable step-size stereo acoustic echo canceller in the frequency domain
Zhenhai Yan, Feiran Yang 0001, Jun Yang 0004
Speech Commun.2
2020 Stochastic Analysis of the Filtered-x LMS Algorithm for Active Noise Control
abstract
The filtered-x least-mean-square (FxLMS) algorithm has been widely used for the active noise control. A fundamental analysis of the convergence behavior of the FxLMS algorithm, including the transient and steady-state performance, could provide some new insights into the algorithm and can be also helpful for its practical applications, e.g., the choice of the step size. Although many efforts have been devoted to the statistical analysis of the FxLMS algorithm, it was usually assumed that the reference signal is Gaussian or white. However, non-Gaussian and/or non-white processes could be very widespread in practice as well. Moreover, the step-size bound that guarantees both of the mean and mean-square stability of the FxLMS for an arbitrary reference signal and a general secondary path is still not available in the literature. To address these problems, this article presents a comprehensive statistical convergence analysis of the FxLMS algorithm without assuming a specific model for the reference signal. We formulate the mean weight behavior and the mean-square error (MSE) in terms of an augmented weight vector. The covariance matrix of the augmented weight-error vector is then evaluated using the vectorization operation, which makes the analysis easy to follow and suitable for arbitrary input distributions. The stability bound is derived based on the first-order and second-order moments analysis of the FxLMS. Computer simulations confirmed the effectiveness of the proposed theoretical model.
Feiran Yang 0001, Jun Yang 0004
IEEE ACM Trans. Audio Speech Lang. Process.1
2019 Mean-square performance of the modified frequency-domain block LMS algorithm
Feiran Yang 0001, Jun Yang 0004
Signal Process.1
2019 A low-complexity permutation alignment method for frequency-domain blind source separation
Fang Kang, Feiran Yang 0001, Jun Yang 0004
Speech Commun.2
2018 Multiband-structured Kalman filter
abstract
The broadband Kalman filter (BKF) and general Kalman filter (GKF) have been proposed for the application of acoustic system identification. Here, the authors present a multiband‐structured Kalman filter (MSKF) to speed up the convergence rate of BKF and GKF for highly correlated signal. A simplified version of MSKF (SMSKF) is also provided at the aim of reducing the complexity. It is shown that the BKF and GKF are the special cases of the proposed MSKF, and the SMSKF can be treated as the improved multiband‐structured subband adaptive filter algorithm with a variable regularisation matrix. The low‐complexity implementation of SMSKF, both in the fast filtering and matrix inversion operation, is discussed. Computer simulations confirm the performance advantage of the proposed algorithm.
Feiran Yang 0001, Jun Yang 0004
IET Signal Process.1
2017 Frequency-Domain Adaptive Kalman Filter With Fast Recovery of Abrupt Echo-Path Changes
abstract
The frequency-domain Kalman filter (FDKF) was successfully applied in echo-cancelation systems due to its fast initial convergence and robustness to double-talk. However, the reconvergence ability of the FDKF was not comprehensively resolved and analyzed in the literature. This letter presents the steady-state solutions of the predicted system distance and corresponding step size of the FDKF, which is found to be related to the signal-to-noise ratio and the transition parameter. On this basis, it is pointed out that a frequently observed slow reconvergence of the FDKF, when the sudden echo-path change occurs, can be mainly attributed to the over-estimated noise power spectral density. Finally, we propose a shadow filter approach to resolve the algorithm's lack of reconvergence. Computer simulations of acoustic echo cancelation confirm the improved performance of the proposed method.
Feiran Yang 0001, Gerald Enzner, Jun Yang 0004
IEEE Signal Process. Lett.1
2017 Statistical Convergence Analysis for Optimal Control of DFT-Domain Adaptive Echo Canceler
abstract
The frequency-domain adaptive filter (FDAF) is widely used in echo cancellation systems due to its low complexity and fast convergence rate. However, the FDAF algorithm with a fixed step size exhibits a tradeoff among the convergence rate, steady-state misalignment, tracking ability, and robustness to near-end speech interferences. Several variable step-size FDAF algorithms were presented to address this problem. However, the state-of-the-art variable step-size FDAF algorithms did not handle this problem comprehensively. This paper presents a new robust variable step-size control approach to the FDAF algorithm. Based on a statistical analysis of the FDAF algorithm, an optimal step size for each frequency bin is derived by minimizing the mean-square deviation (MSD) between the true weight vector and estimated weight vector at each frame. Calculation of the step size requires the system distance and the observation noise power spectral density (PSD). The system distance is estimated using the deterministic recursive equations of MSD and the noise PSD is computed using the magnitude squared coherence function between the far-end signal and error signal. Moreover, a close link between the proposed FDAF and the frequency-domain Kalman filter is revealed. Specifically, the work presented here can be understood as a means to adaptively monitor and control the underlying acoustic state space of the Kalman filter, including means for fast readaptation of the adaptive filter after abrupt echo path changes. Simulation results demonstrate that the proposed algorithm can achieve fast convergence and low steady-state misalignment. Furthermore, the algorithm is robust to the double-talk interferences, but it does not require an explicit double-talk detector.
Feiran Yang 0001, Gerald Enzner, Jun Yang 0004
IEEE ACM Trans. Audio Speech Lang. Process.1
2016 An RLS-Based Lattice-Form Complex Adaptive Notch Filter
abstract
This letter presents a new lattice-form complex adaptive IIR notch filter to estimate and track the frequency of a complex sinusoid signal. The IIR filter is a cascade of a direct-form all-pole prefilter and an adaptive lattice-form all-zero filter. A complex domain exponentially weighted recursive least square algorithm is adopted instead of the widely used least mean square algorithm to increase the convergence rate. The convergence property of this algorithm is investigated, and an expression for the steady-state asymptotic bias is derived. Analysis results indicate that the frequency estimate for a single complex sinusoid is unbiased. Simulation results demonstrate that the proposed method achieves faster convergence and better tracking performance than all traditional algorithms.
Feiran Yang 0001, Jun Yang 0004
IEEE Signal Process. Lett.2
2015 Fast implementation of a family of memory proportionate affine projection algorithm
abstract
Previously, a family of memory proportionate affine projection (MPAP) algorithms has been proposed by taking into account the “history” of the proportionate factors for sparse system identification. This paper presents a low-complexity implementation of this family of MPAP algorithms. Two most important ideas are used in the derivation of the fast algorithm. The first one is to update the auxiliary coefficient vector rather than the true coefficient vector. The second interesting idea is to calculate the error vector by using a recursive technique. Simulation results demonstrate the effectiveness of the fast algorithms.
Feiran Yang 0001, Jun Yang 0004
ICASSP1
2015 Transient and steady-state analyses of the improved multiband-structured subband adaptive filter algorithm
abstract
The authors previously proposed an improved multiband‐structured subband adaptive filter (IMSAF) algorithm. In this contribution, they first present two delayless structures of the IMSAF algorithm to remove delay in the signal path. Then, they study the transient and steady‐state behaviour of the IMSAF algorithms based on the energy conservation arguments and paraunitary condition imposed on the analysis and synthesis filter banks. The analysis does not require a model for the input signal. Simulation results show that the proposed delayless IMSAF algorithm has a faster convergence rate than the traditional delayless subband adaptive filtering algorithms. Theoretical analysis of the IMSAF algorithm is in good agreement with computer simulation results.
Feiran Yang 0001, Ming Wu 0005, Peifeng Ji, Zheng Kuang, Jun Yang 0004
IET Signal Process.1
2014 A fast exact filtering approach to a family of affine projection-type algorithms
Feiran Yang 0001, Ming Wu 0005, Jun Yang 0004, Zheng Kuang
Signal Process.1
2012 An Improved Multiband-Structured Subband Adaptive Filter Algorithm
abstract
Recently, a multiband-structured subband adaptive filter (MSAF) algorithm was proposed to speed up the convergence of the normalized least-mean-square (NLMS) algorithm. In this letter, we extend this work and propose an improved multiband-structured subband adaptive filter (IMSAF) algorithm to increase the convergence speed of the MSAF, which can also be regarded as a unifying framework for the NLMS, MSAF, and affine projection (AP) algorithms. The proposed optimization criterion is based on the principle of minimal disturbance, canceling the most recentPa posteriori errors in each of theNsubbands. The stability condition and the computational complexity are also analyzed. Computer simulations in the context of system identification demonstrate the effectiveness of the new algorithm.
Feiran Yang 0001, Ming Wu 0005, Peifeng Ji, Jun Yang 0004
IEEE Signal Process. Lett.1
2012 Stereophonic Acoustic Echo Suppression Based on Wiener Filter in the Short-Time Fourier Transform Domain
abstract
An open-loop stereophonic acoustic echo suppression (SAES) method without preprocessing is presented for teleconferencing systems, where the Wiener filter in the short-time Fourier transform (STFT) domain is employed. Instead of identifying the echo path impulse responses with adaptive filters, the proposed algorithm estimates the echo spectra from the stereo signals using two weighting functions. The spectral modification technique originally proposed for noise reduction is adopted to remove the echo from the microphone signal. Moreover, a priori signal-to-echo ratio (SER) based Wiener filter is used as the gain function to achieve a trade-off between musical noise reduction and computational load for real-time operations. Computer simulation shows the effectiveness and the robustness of the proposed method in several different scenarios.
Feiran Yang 0001, Ming Wu 0005, Jun Yang 0004
IEEE Signal Process. Lett.1