Jun Yang 0004

dblp:y/JunYang4 · DBLP profile ↗
← Back
58ranked-venue papers
2as first author
26since 2021 · last 2026
0000-0003-2807-6130ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 41 · 19 since 2021Artificial intelligence and machine learning · 17 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Cepstrum Enhancement-Based Nonnegative Matrix Factorization for Blind Source Separation
abstract
Nonnegative matrix factorization (NMF) has proven to be a powerful source spectrogram representation in many audio blind source separation (BSS) methods. However, NMFbased methods may still suffer from the well-known permutation problem especially when the number of NMF basis vectors is large. This letter presents a cepstrum enhancement-based NMF approach that effectively exploits the harmonic structure of audio signals. In oracle NMF-based methods, the source variances are estimated by directly applying NMF to the spectrograms of the separated signals. In the presented approach, however, the spectrogram of each separated source is firstly smoothed using a cepstrum thresholding method, and then NMF is applied to the enhanced spectrogram for source variance estimation. The cepstrum thresholding enhances the harmonic structure of the source signals while removing the interfering components from other sources, which helps alleviate the permutation problem. Also, the proposed cepstrum enhancement method is decoupled from the specific source model and can be easily incorporated into other NMF-based BSS methods as a plug-and-play module. Experimental results validate the effectiveness of the proposed method.
Mengchao Fan, Feiran Yang 0001, Jun Yang 0004
IEEE Signal Process. Lett.4
2025 DeepPreNet: A Deep Learning Pre-Processing Method for Speech Distortion Correction in Parametric Array Loudspeaker
abstract
The parametric array loudspeaker produces highly directional sound via a nonlinear process in air, which also introduces inherent baseband distortions. However, conventional recursive modulators employed to compensate for nonlinearity demand substantially increased bandwidth and are not optimized for speech applications. In this paper, we propose a deep learning method tailored for speech, called DeepPreNet. It contains two parts: a pre-processing network (PreNet) and a forward inference model (ForwModel). The ForwModel is a pre-trained network using real recorded speeches to model the actual nonlinear process, enhancing its reliability for PreNet training. The PreNet is trained to generate pre-processed signals, which are subsequently fed into the ForwModel to recover the distortion-free speech. By leveraging the harmonic-rich feature of speech, the proposed method incorporates distortions to reconstruct clean speech, thereby alleviating the bandwidth constraints imposed by the transducer. Experiments in both near- and far-field conditions demonstrate that the proposed method achieves remarkable performance compared to refined baseline techniques with the real transducer response.
Wenyao Ma, Yunxi Zhu, Fengyuan Hao, Liwen Qin, Fengyi Fan, Jun Yang 0004
ICASSP6
2025 Online secondary path modeling algorithm without auxiliary noise for narrowband active noise control
Ming Wu 0005, Jun Yang 0004
Signal Process.4
2025 Robust algorithms for spherical angle-of-arrival source localization
Pengxiao Teng, Jun Yang 0004
Signal Process.4
2025 Deep Preprocessing Method for Speech Restoration in Parametric Array Loudspeakers via Time-Frequency Domain Modeling
abstract
The parametric array loudspeaker inherently introduces baseband distortions in directional sound applications due to the nonlinear process in air. Recently, DNNs have been used to model this forward process and to generate preprocessed signals for distortion-free speech restoration. However, when trained on real-world audio, the preprocessing network can exploit weaknesses in the forward model, producing adversarial outputs. To address it, we propose a reorganization strategy for the two-stage framework, comprising a causal TF-GridNet for preprocessed signal generation and a modified time-frequency (T-F) domain differential Volterra Filter (DiffVF) as the forward model. The causal TF-GridNet estimates real and imaginary components using a T-F band-split mechanism. The modified forward model integrates the second-order difference and kernel convolution operations of the original time-domain version into the T-F domain, preserving interpretability while stabilizing training. A refined$N$th-order equalization, based on the T-F domain DiffVF model, is implemented as a competitive baseline. Simulated and real-world experiments demonstrate state-of-the-art reconstruction performance of the proposed method across various objective metrics.
Wenyao Ma, Yunxi Zhu, Jun Yang 0004
IEEE Signal Process. Lett.3
2025 Stochastic Analysis of FxLMS Algorithm for Feedback Active Noise Control
abstract
Feedback active noise control (ANC) systems are effective in reducing predictable noise, e.g. periodic, narrowband and colored noise. There are still few studies on the theoretical analysis of feedback ANC systems, and are limited to idealized signals such as sinusoidal or Gaussian signals. This paper presents the stochastic analysis of a feedback ANC system based on the filtered-x least mean square (FxLMS) algorithm, which is not relying on a specific noise model and perfect secondary path. The equations for the mean and mean-square convergence behavior are derived. Extensive simulations of sinusoidal, band-limited white noise, and hybrid signals illustrate the accuracy of the analysis.
Ming Wu 0005, Jun Yang 0004
IEEE Signal Process. Lett.4
2025 Online Blind Speech Separation Using Time-Varying All-Pole Source Model
abstract
Independent vector analysis (IVA) is a state-of-the-art blind source separation (BSS) method that utilizes higher-order correlation between frequency components in each source. However, online-IVA assumes a spherical multivariate distribution and it is not flexible. This letter presents an online BSS method based on the time-varying all-pole (TVAP) source model, which can better capture the spectral envelope of the source signal. The cost function is formulated under the maximum likelihood criterion, where the demixing matrix is updated using iterative projection algorithm and the TVAP model parameters is optimized based on a fixed-point iteration framework. Furthermore, the TVAP model with a low-order all-pole filter effectively solves the permutation problem. Experimental results show that the proposed algorithm outperforms the online-IVA methods for both simulated and live-recorded datasets.
Feiran Yang 0001, Jun Yang 0004
IEEE Signal Process. Lett.3
2024 CAGCN: Centrality-Aware Graph Convolution Network for Anomaly Detection in Industrial Control Systems
Jun Yang 0004, Yiqiang Sheng, Jinlin Wang 0001, Hong Ni
J. Comput. Sci. Technol.1
2024 Design of frequency-invariant uniform concentric circular arrays with first-order directional microphones
Feiran Yang 0001, Zhaoli Yan, Jun Yang 0004
Signal Process.4
2024 Restoration of Bone-Conducted Speech With U-Net-Like Model and Energy Distance Loss
abstract
Bone-conducted speech is less susceptible to ambient noise interference, but it suffers from poor speech quality due to the limited bandwidth. In this letter, we propose a U-Net-like network for the restoration of bone-conducted speech in the time domain. The proposed network consists of residual-connected one-dimensional convolutions and shifted window-based attention modules, which can model long-term dependencies crucial in speech processing. We find that the prevalent time-domain${{l}_{1}}$loss may be insufficient for the generation of high-frequency information absent in bone-conducted speech. To address this issue, we propose to utilize the generalized energy distance loss based on multi-scale Mel spectrograms as the objective function. Experimental results on the ESMB dataset validate the efficacy of our proposed method in restoration of bone-conducted speech. The proposed approach significantly outperforms two recent time-domain benchmarks, DPT-EGNet and EBEN, in terms of PESQ and STOI metrics.
Changtao Li, Feiran Yang 0001, Jun Yang 0004
IEEE Signal Process. Lett.3
2024 GPU Implementation of a Fast Multichannel Wiener Filter Algorithm for Active Noise Control
abstract
Most of the traditional active noise control (ANC) systems are implemented using digital signal processor (DSP), but the lack of computational power of DSP has been a limitation to the application and development of ANC technology. Currently, it is a feasible approach to utilize Graphics Processing Unit (GPU) to enhance the computational power of the ANC system. The advantage of GPU is parallel computation, but most of the traditional ANC algorithms can not be efficiently executed in parallel. In this letter, a parallelizable Fast Multi-channel Wiener Filter (FMWF) algorithm is proposed, and the feasibility of implementing the FMWF algorithm on GPU is verified through experiments, which show that the FMWF algorithm has obvious advantages in parallel execution on GPU. In addition, a DSP-CPU-GPU architecture for ANC systems is designed. In this architecture, each processor can make full use of its own advantages to enhance the computational capability of the system and guarantee the real-time processing of the signals at the same time.
Hongling Sun, Ming Wu 0005, Jun Yang 0004
IEEE Signal Process. Lett.4
2024 Selective-Memory Meta-Learning With Environment Representations for Sound Event Localization and Detection
abstract
Environment shifts and conflicts present significant challenges for learning-based sound event localization and detection (SELD) methods. SELD systems, when trained in particular acoustic settings, often show restricted generalization capabilities for diverse acoustic environments. Furthermore, obtaining annotated samples for spatial sound events is notably costly. Deploying a SELD system in a new environment requires extensive time for re-training and fine-tuning. To overcome these challenges, we propose environment-adaptive Meta-SELD, designed for efficient adaptation to new environments using minimal data. Our method specifically utilizes computationally synthesized spatial data and employs Model-Agnostic Meta-Learning (MAML) on a pre-trained, environment-independent model. The method then utilizes fast adaptation to unseen real-world environments using limited samples from the respective environments. Inspired by the Learning-to-Forget approach, we introduce the concept of selective memory as a strategy for resolving conflicts across environments. This approach involves selectively memorizing target-environment-relevant information and adapting to the new environments through the selective attenuation of model parameters. In addition, we introduce environment representations to characterize different acoustic settings, enhancing the adaptability of our attenuation approach to various environments. We evaluate our proposed method on the development set of the Sony-TAu Realistic Spatial Soundscapes 2023 (STARSS23) dataset and computationally synthesized scenes. Experimental results demonstrate the superior performance of the proposed method compared to conventional supervised learning methods, particularly in localization.
Jinbo Hu, Yin Cao, Ming Wu 0005, Qiuqiang Kong, Feiran Yang 0001, Mark D. Plumbley, Jun Yang 0004
IEEE ACM Trans. Audio Speech Lang. Process.7
2024 A Two-Stage Approach to Quality Restoration of Bone-Conducted Speech
abstract
Bone-conducted speech is not susceptible to background noise but suffers from poor speech quality and intelligibility due to the limited bandwidth. This paper proposes a two-stage approach to restore the quality of bone-conducted speech, namely, bandwidth extension and speech vocoder. In the first stage, a deep neural network is trained to learn mappings from a low-resolution representation of the bone-conducted speech, i.e., log Mel-scale spectrogram, to that of the air-conducted speech, which extends the bandwidth of the bone-conducted speech. In the second stage, a speech vocoder is employed to transform the extended log Mel-scale spectrogram of the bone-conducted speech back to time-domain waveforms. Due to the many-to-many correspondence between the air-conducted and bone-conducted speech, supervised learning may not be the best training protocol for the bone-conducted/air-conducted feature mapping. We thus propose to leverage adversarial training to further improve the bandwidth extension performance in the first stage. The two stages are decoupled and can be trained independently. The vocoder is trained on a large multi-speaker dataset and can generalize well to unknown speakers. Also, the vocoder can help to remedy the spectral artifacts introduced in the bandwidth extension stage. Objective and subjective evaluations on ESMB dataset show that the proposed two-stage system substantially outperforms existing bone-conducted speech enhancement systems.
Changtao Li, Feiran Yang 0001, Jun Yang 0004
IEEE ACM Trans. Audio Speech Lang. Process.3
2024 Theoretical Analysis of Maclaurin Expansion Based Linear Differential Microphone Arrays and Improved Solutions
abstract
Linear differential microphone arrays (LDMAs) are becoming popular due to their potentially high directional gain and frequency-invariant beampattern. By increasing the number of microphones, the Maclaurin expansion-based LDMAs address the inherently poor robustness problem of the conventional LDMA at low frequencies. However, this method encounters severe beampattern distortion and the deep nulls problem in the white noise gain (WNG) and the directivity factor (DF) at high frequencies as the number of microphones increases. In this paper, we reveal that the severe beampattern distortion is attributed to the deviation term of the synthesized beampattern while the deep nulls problem in the WNG and the DF is attributed to the violation of the distortionless constraint in the desired direction. We then propose two new design methods to avoid the degraded performance of LDMAs. Compared to the Maclaurin series expansion-based method, the first method additionally imposes the distortionless constraint in the desired direction, and the deep nulls problem in the WNG and the DF can be avoided. The second method explicitly requires the response of the higher order spatial directivity pattern in the deviation term to be zero, and thus the beampattern distortion can be avoided. By choosing the frequency-wise parameter that determines the number of the considered higher order spatial directivity patterns, the second method enables a good trade-off between the WNG and the beampattern distortion. Simulations exemplify the superiority of the proposed method against existing methods in terms of the robustness and the beampattern distortion.
Feiran Yang 0001, Xiaoqing Hu, Jun Yang 0004
IEEE ACM Trans. Audio Speech Lang. Process.4
2024 Multichannel Linear Prediction-Based Speech Dereverberation Considering Sparse and Low-Rank Priors
abstract
This article addresses the multi-channel linear prediction (MCLP)-based speech dereverberation problem by jointly considering the sparsity and low-rank priors of speech spectrograms. We utilize the complex generalized Gaussian (CGG) distribution as the source model and the generalized nonnegative matrix factorization (NMF) as the spectral model. The difference between the presented model and existing ones for MCLP is twofold. First, we adopt the CGG distribution with a time-frequency-variant scale parameter instead of that with a time-frequency-invariant scale parameter. Second, the time-frequency-varying scale parameter is approximated by NMF in a low-rank manner. Based on the maximum-likelihood criterion, speech dereverberation is formulated as an optimization problem that minimizes the prediction error weighted by the reciprocal of sparse and low-rank parameters. A convergence-guaranteed algorithm is derived to estimate the parameters using the majorization-minimization technology. The WPE, NMF-based WPE and CGG-based WPE can be treated as special cases of the proposed method with different shape and domain parameters. As a byproduct, the proposed method provides a simple and elegant way to derive the CGG-based WPE algorithm. A series of experiments show the superiority of the proposed method over WPE, NMF-based WPE and CGG-based WPE methods.
Taihui Wang, Feiran Yang 0001, Jun Yang 0004
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 A Perspective on Fully Steerable Differential Beamformers for Circular Arrays
abstract
Both the null-constrained method and the Jacobi-Anger expansion method are commonly adopted to design the fully steerable circular differential microphone arrays (CDMAs). So far, these two methods have been independently studied, and it is unclear why they exhibit pronounced performance difference and whether they have a potential connection. In this letter, we develop a unified framework for the design of the fully steerable CDMA, which jointly utilizes the null constraints and the Jacobi-Anger expansion. We show that the existing two methods can be treated as special cases of the proposed approach and they can indeed be connected. We then reveal that the null-constrained CDMA can be viewed as the regularized Jacobi-Anger expansion-based CDMA, which theoretically explains why the former can alleviate the deep nulls problem in white noise gain in certain cases and why it still suffers from the severe deep nulls in other cases. We further explain why the magnitude of the synthesized beampattern of the Jacobi-Anger expansion-based method in nulls' direction can be much larger than that of the null-constrained method. Simulations support the theoretical results well.
Feiran Yang 0001, Jun Yang 0004
IEEE Signal Process. Lett.3
2022 A Track-Wise Ensemble Event Independent Network for Polyphonic Sound Event Localization and Detection
abstract
Polyphonic sound event localization and detection (SELD) aims at detecting types of sound events with corresponding temporal activities and spatial locations. In this paper, a trackwise ensemble event independent network with a novel data augmentation method is proposed. The proposed model is based on our previous proposed Event-Independent Network V2 and is extended by conformer blocks and dense blocks. The track-wise ensemble model with track-wise output format is proposed to solve an ensemble model problem for track-wise output format that track permutation may occur among different models. The data augmentation approach contains several data augmentation chains, which are composed of random combinations of several data augmentation operations. The method also utilizes log-mel spectrograms, intensity vectors, and Spatial Cues-Augmented Log-Spectrogram (SALSA) for different models. We evaluate our proposed method in the Task of the L3DAS22 challenge and obtain the top ranking solution with a location-dependent F-score to be 0.699. Source code is released1.
Jinbo Hu, Yin Cao, Ming Wu 0005, Qiuqiang Kong, Feiran Yang 0001, Mark D. Plumbley, Jun Yang 0004
ICASSP7
2022 A Multi-Task Learning Method for Weakly Supervised Sound Event Detection
abstract
In weakly supervised sound event detection (SED), only coarse-grained labels are available, and thus the supervision information is quite limited. To fully utilize prior knowledge of the time-frequency masks of each sound event, we propose a novel multi-task learning (MTL) method that takes SED as the main task and source separation as the auxiliary task. For active events, we minimize the overlap of their masks as the segment loss to learn distinguishing features. For inactive events, the proposed method measures the activity of masks as silent loss to reduce the insertion error. The auxiliary source separation task calculates an extra penalty according to the shared masks, which can further incorporate prior knowledge in the form of regularization constraints. We demonstrated that the proposed method can effectively reduce the insertion error and achieve a better performance in SED task than single-task methods.
Sichen Liu 0002, Feiran Yang 0001, Fang Kang, Jun Yang 0004
ICASSP4
2022 Statistical analysis of multichannel FxLMS algorithm for narrowband active noise control
Ming Wu 0005, Jing Chen 0090, Zeqiang Zhang, Yin Cao, Jun Yang 0004
Signal Process.7
2022 Steady-State Performance Analysis of the Distributed FxLMS Algorithm for Narrowband ANC System With Frequency Mismatch
abstract
Distributed narrowband active noise control (NANC) systems using diffusion filtered-x least mean square (FxLMS) algorithm can effectively suppress the low-frequency periodic noise generated by rotating machinery. The computational burden is dispersed among the nodes over the acoustic sensor networks. The frequency of the reference signal is usually identified by a non-acoustic sensor in a NANC system. However, the frequency of the reference signal will be different from the primary noise frequency due to inevitable aging and fatigue accumulation and the control performance of NANC degrades considerably. This phenomenon is referred to as frequency mismatch (FM). In this letter, we analyze the performance degradation of the diffusion FxLMS algorithm due to FM in an NANC system. A theoretical model of the diffusion FxLMS algorithm in the presence of FM is derived based on the equivalent transfer function approach. Extensive simulations are performed to confirm the validity of the theoretical analysis.
Jing Chen 0090, Ming Wu 0005, Jun Yang 0004
IEEE Signal Process. Lett.5
2022 The Role of Long-Term Dependency in Synthetic Speech Detection
abstract
Although much progress has been made in synthetic speech detection, there lacks comprehensive analysis of the essential differences between spoofed and genuine speech. We here utilize supervised contrastive loss originated from contrastive learning as an analytical tool to characterize the class similarity structure of ASVspoof 2019 logical access (LA) dataset, which shows that an ideal back-end classifier for synthetic speech detection should have the ability to capture long-term dependencies. Recently, Transformer has been found to have an excellent ability in learning long-term dependencies of input data. We hence propose a back-end classifier based on Transformer Encoder for synthetic speech detection. Convolution blocks are added before the Transformer Encoder, which leverages inductive biases to improve the generalization ability. Compared to two-dimensional convolution, one-dimensional convolution makes better architectural assumptions about the input speech features, which helps with modeling long-term dependencies and decreases the risk of overfitting. The proposed Transformer combined with one-dimensional convolution has fewer parameters than most existing back-end classifiers, and achieves an equal error rate of 1.06% and a minimum tandem detection cost function metric of 0.0345 when evaluated on ASVspoof 2019 LA dataset, which is one of the best models reported in the literature.
Changtao Li, Feiran Yang 0001, Jun Yang 0004
IEEE Signal Process. Lett.3
2022 Insights Into the MMSE-Based Frequency-Invariant Beamformers for Uniform Circular Arrays
abstract
This letter provides new insights into the minimum mean square error (MMSE)-based frequency-invariant beamformer with the uniform circular array (FIB-UCA). By exploiting the symmetry of the UCA, we derive an explicit form of the white noise gain (WNG), the directivity factor (DF) and the MSE, and also derive the corresponding approximate expressions at low frequencies. We then investigate the impact of the look direction on the performance of the MMSE-based FIB-UCA. Interestingly, we further prove that for a UCA with the fixed aperture, the MMSE-based FIB with 2$N$microphones achieves a similar WNG as that with 4$N$microphones at low frequencies if the$N$th-order beampattern is steered to microphone angles. We also derive a unified explicit form of weighting vector and reveal that the MMSE-based FIB-UCA reduces to Jacobi-Anger expansion-based differential beamformer with the UCA at low frequencies.
Feiran Yang 0001, Jun Yang 0004
IEEE Signal Process. Lett.3
2022 Analysis of the Frequency Interference in the Narrowband Active Noise Control System
abstract
The narrowband active noise control (ANC) algorithm works well in the multi-tone noise reduction system. However, the performance of this algorithm degenerates significantly when the noisy frequencies become very close (e. g. frequency gap is 3 Hz where the sampling frequency is 5000 Hz), which can be seen as frequency interference effects. Up to now, however, this effect has not been researched in detail. In this work, the frequency interference effects are analyzed and discussed. The convergence process, the noise reduction value, the stability condition and the optimum step size are obtained by extensive math formulation derivation, provided two frequencies are close and fixed. Theoretical results show that, tones with large frequency difference can be approximated to independent convergence while the system with close frequencies suffers from poor convergence rate. The value of step size plays an important role in convergence, but the interference effects could never be eliminated by adjusting step size. Simulations are conducted to verify the analysis and the results match well with the analysis.
Hongling Sun, Ming Wu 0005, Jun Yang 0004
IEEE ACM Trans. Audio Speech Lang. Process.5
2022 Convolutive Transfer Function-Based Multichannel Nonnegative Matrix Factorization for Overdetermined Blind Source Separation
abstract
Most multichannel blind source separation (BSS) approaches rely on a spatial model to encode the transfer functions from sources to microphones and a source model to encode the source power spectral density. The rank-1 spatial model has been widely exploited in independent component analysis (ICA), independent vector analysis (IVA), and independent low-rank matrix analysis (ILRMA). The full-rank spatial model is also considered in many BSS approaches, such as full-rank spatial covariance matrix analysis (FCA), multichannel nonnegative matrix factorization (MNMF), and FastMNMF, which can improve the separation performance in the case of long reverberation times. This paper proposes a new MNMF framework based on the convolutive transfer function (CTF) for overdetermined BSS. The time-domain convolutive mixture model is approximated by a frequency-wise convolutive mixture model instead of the widely adopted frequency-wise instantaneous mixture model. The iterative projection algorithm is adopted to estimate the demixing matrix, and the multiplicative update rule is employed to estimate nonnegative matrix factorization (NMF) parameters. Finally, the source image is reconstructed using a multichannel Wiener filter. The advantages of the proposed method are twofold. First, the CTF approximation enables us to use a short window to represent long impulse responses. Second, the full-rank spatial model can be derived based on the CTF approximation and slowly time-variant source variances, and close relationships between the proposed method and ILRMA, FCA, MNMF and FastMNMF are revealed. Extensive experiments show that the proposed algorithm achieves a higher separation performance than ILRMA and FastMNMF in reverberant environments.
Taihui Wang, Feiran Yang 0001, Jun Yang 0004
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Real-Time Independent Vector Analysis Using Semi-Supervised Nonnegative Matrix Factorization as a Source Model
Taihui Wang, Feiran Yang 0001, Jun Yang 0004
Interspeech4
2021 Real-Time Independent Vector Analysis with a Deep-Learning-Based Source Model
abstract
In this paper, we present a real-time blind source separation (BSS) algorithm, which unifies the independent vector analysis (IVA) as a spatial model and a deep neural network (DNN) as a source model. The auxiliary-function based IVA (Aux-IVA) is utilized to update the demixing matrix, and the required time-varying variance of the speech source is estimated by a DNN. The DNN could provide a more accurate source model, which then helps to optimize the spatial model. In addition, because the DNN is used to estimate the source variance instead of the source power spectrogram, the size of DNN can be reduced significantly. Experiment results show that the joint utilization of the model-based approach and the data-driven approach provides a more efficient solution than just alone in terms of convergence rate and source separation performance.
Fang Kang, Feiran Yang 0001, Jun Yang 0004
SLT3
2020 Convergence analysis of the conventional filtered-x affine projection algorithm for active noise control
Feiran Yang 0001, Jun Yang 0004
Signal Process.3
2020 Optimum step-size control for a variable step-size stereo acoustic echo canceller in the frequency domain
Zhenhai Yan, Feiran Yang 0001, Jun Yang 0004
Speech Commun.3
2020 Stochastic Analysis of the Filtered-x LMS Algorithm for Active Noise Control
abstract
The filtered-x least-mean-square (FxLMS) algorithm has been widely used for the active noise control. A fundamental analysis of the convergence behavior of the FxLMS algorithm, including the transient and steady-state performance, could provide some new insights into the algorithm and can be also helpful for its practical applications, e.g., the choice of the step size. Although many efforts have been devoted to the statistical analysis of the FxLMS algorithm, it was usually assumed that the reference signal is Gaussian or white. However, non-Gaussian and/or non-white processes could be very widespread in practice as well. Moreover, the step-size bound that guarantees both of the mean and mean-square stability of the FxLMS for an arbitrary reference signal and a general secondary path is still not available in the literature. To address these problems, this article presents a comprehensive statistical convergence analysis of the FxLMS algorithm without assuming a specific model for the reference signal. We formulate the mean weight behavior and the mean-square error (MSE) in terms of an augmented weight vector. The covariance matrix of the augmented weight-error vector is then evaluated using the vectorization operation, which makes the analysis easy to follow and suitable for arbitrary input distributions. The stability bound is derived based on the first-order and second-order moments analysis of the FxLMS. Computer simulations confirmed the effectiveness of the proposed theoretical model.
Feiran Yang 0001, Jun Yang 0004
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 A narrowband active noise control system with a frequency estimation algorithm based on parallel adaptive notch filter
Hongling Sun, Yunping Sun, Ming Wu 0005, Jun Yang 0004
Signal Process.5
2019 Mean-square performance of the modified frequency-domain block LMS algorithm
Feiran Yang 0001, Jun Yang 0004
Signal Process.2
2019 A low-complexity permutation alignment method for frequency-domain blind source separation
Fang Kang, Feiran Yang 0001, Jun Yang 0004
Speech Commun.3
2019 Robust Personal Audio Geometry Optimization in the SVD-Based Modal Domain
abstract
Personal audio generates sound zones in a shared space to provide private and personalized listening experiences with minimized interference between consumers. Regularization has been commonly used to increase the robustness of such systems against potential perturbations in the sound reproduction. However, the performance is limited by the system geometry such as the number and location of the loudspeakers and controlled zones. This paper proposes a geometry optimization method to find the most geometrically robust approach for personal audio amongst all available candidate system placements. The proposed method aims to approach the most “natural” sound reproduction so that the solo control of the listening zone coincidently accompanies the preferred quiet zone. Being formulated in the SVD-based modal domain, the method is demonstrated by applications in three typical personal audio optimizations, i.e., the acoustic contrast control, the pressure matching, and the planarity control. Simulation results show that the proposed method can obtain the system geometry with better avoidance of “occlusion,” improved robustness to regularization, and improved broadband equalization.
Qiaoxi Zhu, Philip Coleman, Xiaojun Qiu, Ming Wu 0005, Jun Yang 0004, Ian S. Burnett
IEEE ACM Trans. Audio Speech Lang. Process.5
2018 Multiband-structured Kalman filter
abstract
The broadband Kalman filter (BKF) and general Kalman filter (GKF) have been proposed for the application of acoustic system identification. Here, the authors present a multiband‐structured Kalman filter (MSKF) to speed up the convergence rate of BKF and GKF for highly correlated signal. A simplified version of MSKF (SMSKF) is also provided at the aim of reducing the complexity. It is shown that the BKF and GKF are the special cases of the proposed MSKF, and the SMSKF can be treated as the improved multiband‐structured subband adaptive filter algorithm with a variable regularisation matrix. The low‐complexity implementation of SMSKF, both in the fast filtering and matrix inversion operation, is discussed. Computer simulations confirm the performance advantage of the proposed algorithm.
Feiran Yang 0001, Jun Yang 0004
IET Signal Process.2
2017 Intelligibilities of Mandarin Chinese Sentences with Spectral "Holes"
Yafan Chen, Jun Yang 0004
INTERSPEECH3
2017 Frequency-Domain Adaptive Kalman Filter With Fast Recovery of Abrupt Echo-Path Changes
abstract
The frequency-domain Kalman filter (FDKF) was successfully applied in echo-cancelation systems due to its fast initial convergence and robustness to double-talk. However, the reconvergence ability of the FDKF was not comprehensively resolved and analyzed in the literature. This letter presents the steady-state solutions of the predicted system distance and corresponding step size of the FDKF, which is found to be related to the signal-to-noise ratio and the transition parameter. On this basis, it is pointed out that a frequently observed slow reconvergence of the FDKF, when the sudden echo-path change occurs, can be mainly attributed to the over-estimated noise power spectral density. Finally, we propose a shadow filter approach to resolve the algorithm's lack of reconvergence. Computer simulations of acoustic echo cancelation confirm the improved performance of the proposed method.
Feiran Yang 0001, Gerald Enzner, Jun Yang 0004
IEEE Signal Process. Lett.3
2017 Statistical Convergence Analysis for Optimal Control of DFT-Domain Adaptive Echo Canceler
abstract
The frequency-domain adaptive filter (FDAF) is widely used in echo cancellation systems due to its low complexity and fast convergence rate. However, the FDAF algorithm with a fixed step size exhibits a tradeoff among the convergence rate, steady-state misalignment, tracking ability, and robustness to near-end speech interferences. Several variable step-size FDAF algorithms were presented to address this problem. However, the state-of-the-art variable step-size FDAF algorithms did not handle this problem comprehensively. This paper presents a new robust variable step-size control approach to the FDAF algorithm. Based on a statistical analysis of the FDAF algorithm, an optimal step size for each frequency bin is derived by minimizing the mean-square deviation (MSD) between the true weight vector and estimated weight vector at each frame. Calculation of the step size requires the system distance and the observation noise power spectral density (PSD). The system distance is estimated using the deterministic recursive equations of MSD and the noise PSD is computed using the magnitude squared coherence function between the far-end signal and error signal. Moreover, a close link between the proposed FDAF and the frequency-domain Kalman filter is revealed. Specifically, the work presented here can be understood as a means to adaptively monitor and control the underlying acoustic state space of the Kalman filter, including means for fast readaptation of the adaptive filter after abrupt echo path changes. Simulation results demonstrate that the proposed algorithm can achieve fast convergence and low steady-state misalignment. Furthermore, the algorithm is robust to the double-talk interferences, but it does not require an explicit double-talk detector.
Feiran Yang 0001, Gerald Enzner, Jun Yang 0004
IEEE ACM Trans. Audio Speech Lang. Process.3
2016 Estimating ear canal geometry and eardrum reflection coefficient from ear canal input impedance
abstract
Based on the signal model of ear canals, a novel method for solving the inverse problem of estimating the unique solution of the ear canal area function and the eardrum reflection coefficient given the acoustic input impedance at the entrance of an ear canal is presented. Up-sampling techniques to improve the accuracy of the estimates are also presented. The performance of this method and factors affecting the accuracy of the estimates are investigated via simulations. It is found that the accuracy of the estimates is limited by the measurement bandwidth of the given ear canal input impedance. In the audio frequency range, the estimates obtained approximate well to the true ones. To obtain more accurate estimates, a wider measurement bandwidth of the ear canal input impedance is required.
Huiqun Deng, Jun Yang 0004
ICASSP2
2016 Request distribution with pre-learning for distributed SSL reverse proxies
abstract
As network data security becoming more and more universalized, distributed Secure Sockets Layer (SSL) reverse proxies are often used in Web systems to offload CPU exhausting SSL operations from Web servers and improve the execution performance of the SSL protocol. The distribution strategy of user requests to the SSL reverse proxies is a significant factor affecting the system's performance in processing SSL operations. Aiming at improving the quality of request distribution decisions, this paper proposes a new approach for SSL reverse proxy load estimation, i.e. the family of algorithms called Load Estimation with Pre-Learning (LEPL), which estimates load using pre-learned machine learning models. Using LEPL, high accuracy of load estimation can be achieved, so that better request distribution decisions can be made. Our experimental results show that by using pre-learning, the SSL reverse proxy system's average response time can be shortened by about 30% - 50%.
Jun Yang 0004, Yiqiang Sheng
SNPD2
2016 A new efficient filtered-x affine projection sign algorithm for active control of impulsive noise
Longshuai Xiao, Ming Wu 0005, Jun Yang 0004
Signal Process.3
2016 An RLS-Based Lattice-Form Complex Adaptive Notch Filter
abstract
This letter presents a new lattice-form complex adaptive IIR notch filter to estimate and track the frequency of a complex sinusoid signal. The IIR filter is a cascade of a direct-form all-pole prefilter and an adaptive lattice-form all-zero filter. A complex domain exponentially weighted recursive least square algorithm is adopted instead of the widely used least mean square algorithm to increase the convergence rate. The convergence property of this algorithm is investigated, and an expression for the steady-state asymptotic bias is derived. Analysis results indicate that the frequency estimate for a single complex sinusoid is unbiased. Simulation results demonstrate that the proposed method achieves faster convergence and better tracking performance than all traditional algorithms.
Feiran Yang 0001, Jun Yang 0004
IEEE Signal Process. Lett.3
2015 Fast implementation of a family of memory proportionate affine projection algorithm
abstract
Previously, a family of memory proportionate affine projection (MPAP) algorithms has been proposed by taking into account the “history” of the proportionate factors for sparse system identification. This paper presents a low-complexity implementation of this family of MPAP algorithms. Two most important ideas are used in the derivation of the fast algorithm. The first one is to update the auxiliary coefficient vector rather than the true coefficient vector. The second interesting idea is to calculate the error vector by using a recursive technique. Simulation results demonstrate the effectiveness of the fast algorithms.
Feiran Yang 0001, Jun Yang 0004
ICASSP2
2015 Transient and steady-state analyses of the improved multiband-structured subband adaptive filter algorithm
abstract
The authors previously proposed an improved multiband‐structured subband adaptive filter (IMSAF) algorithm. In this contribution, they first present two delayless structures of the IMSAF algorithm to remove delay in the signal path. Then, they study the transient and steady‐state behaviour of the IMSAF algorithms based on the energy conservation arguments and paraunitary condition imposed on the analysis and synthesis filter banks. The analysis does not require a model for the input signal. Simulation results show that the proposed delayless IMSAF algorithm has a faster convergence rate than the traditional delayless subband adaptive filtering algorithms. Theoretical analysis of the IMSAF algorithm is in good agreement with computer simulation results.
Feiran Yang 0001, Ming Wu 0005, Peifeng Ji, Zheng Kuang, Jun Yang 0004
IET Signal Process.5
2015 A Step Size Control Method for Deficient Length FBLMS Algorithm
abstract
In many practical situations, where the length of the system impulse response is extremely long, the adaptive filter usually works in an under-modeling situation. In our previous work, we found that for deficient-length frequency-domain block least-mean-square (FBLMS) algorithm, the steady-state solution depends on the step size selected. The deficient length-FBLMS algorithm will converge to the Wiener solution only if the same step size is selected for each frequency bin. Numerous variable step-size methods for FBLMS algorithm have been proposed. However, almost all of them cannot converge to the Wiener solution in the under-modeling situation. In this letter, a step size control method for deficient-length FBLMS algorithm is proposed. Effectiveness of the proposed algorithm is demonstrated through computer simulations.
Ming Wu 0005, Jun Yang 0004
IEEE Signal Process. Lett.2
2014 A fast exact filtering approach to a family of affine projection-type algorithms
Feiran Yang 0001, Ming Wu 0005, Jun Yang 0004, Zheng Kuang
Signal Process.3
2014 Modeling and compensation for the distortion of parametric loudspeakers using a one-dimension Volterra filter
abstract
Recently, the general Volterra filter (VF) has been adopted for the modeling of parametric loudspeakers. However, the computation complexity of the VF is too high for real-time implementation. In this paper, a one-dimension Volterra filter (ODVF) with much lower complexity is introduced to model and compensate for the nonlinearity of parametric loudspeakers. A theoretical framework for ODVF model identification is established and a method of measuring the ODVF kernels using the exponential swept-sine signal is provided. The validity of modeling the nonlinearity of the parametric loudspeaker using the ODVF is verified theoretically and experimentally. Based on the established ODVF model, an inverse filter is designed to compensate for the 2nd harmonic distortion of the parametric loudspeakers. To further reduce the 3rd harmonic distortion, an improved compensation method is also proposed. Experimental results show that the performance of the ODVF-based compensation is comparable to that of the Volterra-filter based compensation.
Yongsheng Mu, Peifeng Ji, Ming Wu 0005, Jun Yang 0004
IEEE ACM Trans. Audio Speech Lang. Process.5
2013 Design of a time-domain acoustic contrast control for broadband input signals in personal audio systems
abstract
The acoustic contrast control (ACC) approach provides a simple strategy to focus sound on a designed target area in a personal audio system. All of the traditional ACC approaches have been designed in the frequency domain, and then transformed into the time domain through the application of an inverse Fourier transform. Therefore, the broadband contrast performance of the traditional ACC approach is not optimum. Especially, when the length of the control filter is short, the traditional approach may have poor contrast performance at non-control frequencies. This study proposes a novel method that can achieve the optimum broadband contrast performance, which also maintains a good frequency-response consistency by introducing the response variation constraints to improve sound quality. Experimental results demonstrate the effectiveness of the new method.
Yefeng Cai, Ming Wu 0005, Jun Yang 0004
ICASSP3
2012 An Improved Multiband-Structured Subband Adaptive Filter Algorithm
abstract
Recently, a multiband-structured subband adaptive filter (MSAF) algorithm was proposed to speed up the convergence of the normalized least-mean-square (NLMS) algorithm. In this letter, we extend this work and propose an improved multiband-structured subband adaptive filter (IMSAF) algorithm to increase the convergence speed of the MSAF, which can also be regarded as a unifying framework for the NLMS, MSAF, and affine projection (AP) algorithms. The proposed optimization criterion is based on the principle of minimal disturbance, canceling the most recentPa posteriori errors in each of theNsubbands. The stability condition and the computational complexity are also analyzed. Computer simulations in the context of system identification demonstrate the effectiveness of the new algorithm.
Feiran Yang 0001, Ming Wu 0005, Peifeng Ji, Jun Yang 0004
IEEE Signal Process. Lett.4
2012 Stereophonic Acoustic Echo Suppression Based on Wiener Filter in the Short-Time Fourier Transform Domain
abstract
An open-loop stereophonic acoustic echo suppression (SAES) method without preprocessing is presented for teleconferencing systems, where the Wiener filter in the short-time Fourier transform (STFT) domain is employed. Instead of identifying the echo path impulse responses with adaptive filters, the proposed algorithm estimates the echo spectra from the stereo signals using two weighting functions. The spectral modification technique originally proposed for noise reduction is adopted to remove the echo from the microphone signal. Moreover, a priori signal-to-echo ratio (SER) based Wiener filter is used as the gain function to achieve a trade-off between musical noise reduction and computational load for real-time operations. Computer simulation shows the effectiveness and the robustness of the proposed method in several different scenarios.
Feiran Yang 0001, Ming Wu 0005, Jun Yang 0004
IEEE Signal Process. Lett.3
2011 A Comparative Analysis of Preprocessing Methods for the Parametric Loudspeaker Based on the Khokhlov-Zabolotskaya-Kuznetsov Equation for Speech Reproduction
abstract
Based on the Berktay's farfield solution, various preprocessing methods were proposed to reduce the distortion of the highly directional audible signal in the parametric loudspeaker. However, the Berktay's farfield solution is an approximated model of nonlinear acoustic propagation. To determine the effectiveness of these methods, we analyze various preprocessing methods theoretically for directional speech reproduction using the Khokhlov-Zabolotskaya-Kuznetsov (KZK) equation, which provides a more accurate model of nonlinear acoustic propagation. In order to reduce the distortion effectively in the parametric loudspeaker with these preprocessing methods, the initial sound pressure level of the carrier frequency is found to be less than 132 dB according to the KZK equation. Unlike the Berktay' farfield solution that results in a +12 dB/octave gain slope, different gain slopes are derived using the KZK equation and appropriate equalizers are proposed to improve the frequency response of the parametric loudspeaker. The optimal preprocessing method for directional speech reproduction is established based on the KZK equation, which has a relatively flat frequency response of the desired speech signal and the best total harmonic distortion performance.
Peifeng Ji, Ee-Leng Tan, Woon-Seng Gan, Jun Yang 0004
IEEE ACM Trans. Audio Speech Lang. Process.4
2010 Performance analysis on recursive single-sideband amplitude modulation for parametric loudspeakers
abstract
A highly directional speech signal can be generated using parametric loudspeaker. The generation of highly directional sound beam is due to the nonlinear interaction of amplitude-modulated ultrasound waves in air. However, severe distortion is also generated during the reproduction of directional speech and several preprocessing techniques based on the Berktay's farfield model have been proposed by researchers to reduce the distortion. In this paper, we carried out a thorough investigation on the analytical performance of the recursive single-sideband amplitude modulation (RSSB-AM) technique, which has been found to perform well for directional speech reproduction. Several important characteristics of the performance of the RSSB-AM are observed and optimal parameters of the RSSB-AM are also presented.
Peifeng Ji, Woon-Seng Gan, Ee-Leng Tan, Jun Yang 0004
ICME4
2009 Acoustical Vehicle Detection Based on Bispectral Entropy
abstract
A vehicle detection algorithm based on Bispectral entropy is proposed in this paper. Based on Quadratic-Phase Coupling(QPC) analysis, Bispectral entropy, as the complexity measure of bispectra, is calculated from the vehicle acoustic signal database which is acquired in several real world experiments. Furthermore, an effective bispectral entropy-based algorithm is developed for vehicle detection. Experiments show that the proposed algorithm can achieve longer distance alert than other two detectors.
Ming Bao, Chengshi Zheng, Xiaodong Li 0002, Jun Yang 0004
IEEE Signal Process. Lett.4
2006 Orthogonal Relief Algorithm for Feature Selection
Jun Yang 0004, Yue-Peng Li
ICIC (1)1
2006 A digital beamsteerer for difference frequency in a parametric array
abstract
A steerable audio system can be realized using parametric array. However, the available steerable angle is often limited by the sampling interval used in the digital system. As such, the smallest steerable angle is large (/spl sim/26/spl deg/) for several hundred kilohertz of sampling frequency. Although there are some fractional delay or frequency domain algorithms can be used to improve the steering angle, most of the algorithms are either computational intensive or introduce error during the process. In this paper, an algorithm is proposed to rectify this problem by applying separate delays to the carrier and sideband frequencies. Different weighting functions also added to the carrier and sideband frequencies to control the difference frequency's beamwidth and sidelobe. Most importantly, the proposed system can steer the difference frequency to a small angle with minimal computation.
Woon-Seng Gan, Jun Yang 0004, Khim Sia Tan, Meng Hwa Er
IEEE Trans. Speech Audio Process.2
2005 Nonlinear least-square solution to flat-top pattern synthesis using arbitrary linear array
Yuan Wen, Woon-Seng Gan, Jun Yang 0004
Signal Process.3
2004 An efficient digital beamsteering system for difference frequency in parametric array
abstract
For a digital beamsteering system, the smallest time delay available is equal to the sampling period of the digital signal processing (DSP) board. As most of the time the sampling frequency is not high enough, the smallest steering angle available is large, which is undesirable. This limitation also occurs when performing beamsteering in a parametric array digitally. Although partial delay or frequency domain algorithms can be used to improve the steering angle, most of the algorithms are either computational intensive or introduce error during the process. In this paper, an algorithm is proposed to beamsteer the difference frequency in parametric array. The proposed system can be used to steer the difference frequency to a small angle, without the need to increase the sampling frequency or implement partial delay.
Khim Sia Tan, Woon-Seng Gan, Jun Yang 0004, Meng Hwa Er
ICASSP (2)3
2003 Constant beamwidth beamformer for difference frequency in parametric array
abstract
Sound reproduction in air by using a parametric acoustic array has been investigated for a few decades. Two inaudible ultrasonic frequencies are produced from the parametric array. Due to the nonlinearity of air, it is possible to produce an audible frequency with its frequency equal to the difference in the two ultrasonic frequencies. However, there is not much work done in controlling the beam pattern of the difference frequency generated by the primary waves. In this paper, an algorithm is proposed to control the sidelobe level of the difference frequency directivity. By making use of array signal processing techniques, the algorithm is also capable of producing a constant beamwidth for broadband difference frequency.
Khim Sia Tan, Woon-Seng Gan, Jun Yang 0004, Meng Hwa Er
ICASSP (5)3
2003 Constant beamwidth beamformer for difference frequency in parametric array
abstract
The sound reproduction in air by using a parametric acoustic array [P.J. Westervelt, 1963] has been reported for a few decades. Two inaudible ultrasonic frequencies are produced from the parametric array. Due to the nonlinearity of air, it is possible to produce an audible frequency with its frequency equal to the difference in the two ultrasonic frequencies. However, there is not much work done in controlling the beam pattern of the difference frequency generated by the primary waves. In this paper, an algorithm is proposed to control the sidelobe level of the difference frequency directivity. By making use of array signal processing techniques, the algorithm is also capable of producing a constant beamwidth for broadband difference frequency.
Khim Sia Tan, Woon-Seng Gan, Jun Yang 0004, Meng Hwa Er
ICME3