Jesper Rindom Jensen

dblp:93/9878 · DBLP profile ↗
← Back
58ranked-venue papers
16as first author
12since 2021 · last 2025
0000-0001-6023-8270ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 43 · 10 first-author · 12 since 2021Artificial intelligence and machine learning · 19 · 6 first-author · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 A Study of the Scale Invariant Signal to Distortion Ratio in Speech Separation with Noisy References*
abstract
This paper examines the implications of using the Scale-Invariant Signal-to-Distortion Ratio (SI-SDR) as both evaluation and training objective in supervised speech separation, when the training references contain noise, as is the case with the de facto benchmark WSJ0-2Mix. A derivation of the SI-SDR with noisy references reveals that noise limits the achievable SI-SDR, or leads to undesired noise in the separated outputs. To address this, a method is proposed to enhance references and augment the mixtures with WHAM!, aiming to train models that avoid learning noisy references. Two models trained on these enhanced datasets are evaluated with the non-intrusive NISQA.v2 metric. Results show reduced noise in separated speech but suggest that processing references may introduce artefacts, limiting overall quality gains. Negative correlation is found between SI-SDR and perceived noisiness across models on the WSJ0-2Mix and Libri2Mix test sets, underlining the conclusion from the derivation.
Simon Dahl Jepsen, Mads Græsbøll Christensen, Jesper Rindom Jensen
ASRU3
2025 Robust Fixed-Filter Sound Zone Control with Audio-Based Position Tracking
abstract
Performance of sound zone control (SZC) systems deployed in practical scenarios are highly sensitive to the location of the listener(s) and can degrade significantly when listener(s) are moving. This paper presents a robust SZC system that adapts to dynamic changes such as moving listeners and varying zone locations using a dictionary-based approach. The proposed system continuously monitors the environment and updates the fixed control filters by tracking the listener position using audio signals only. To test the effectiveness of the proposed SZC method, simulation studies are carried out using practically measured impulse responses. These studies show that SZC, when incorporated with the proposed audio-only position tracking scheme, achieves optimal performance when all listener positions are available in the dictionary. Moreover, even when not all listener positions are included in the dictionary, the method still provides good performance improvement compared to a traditional fixed filter SZC scheme.
Sankha Subhra Bhattacharjee, Andreas Jonas Fuglsig, Flemming Christensen, Jesper Rindom Jensen, Mads Græsbøll Christensen
ICASSP4
2025 Sound Zone Control Robust To Sound Speed Change
abstract
Sound zone control (SZC) implemented using static optimal filters is significantly affected by various perturbations in the acoustic environment, an important one being the fluctuation in the speed of sound, which is in turn influenced by changes in temperature and humidity (TH). This issue arises because control algorithms typically use pre-recorded, static impulse responses (IRs) to design the optimal control filters. The IRs, however, may change with time due to TH changes, which renders the derived control filters to become non-optimal. To address this challenge, we propose a straightforward model called sinc interpolation-compression/expansion-resampling (SICER), which adjusts the IRs to account for both sound speed reduction and increase. Using the proposed technique, IRs measured at a certain TH can be corrected for any TH change and control filters can be re-derived without the need of re-measuring the new IRs (which is impractical when SZC is deployed). We integrate the proposed SICER IR correction method with the recently introduced variable span trade-off (VAST) framework for SZC, and propose a SICER-corrected VAST method that is resilient to sound speed variations. Simulation studies show that the proposed SICER-corrected VAST approach significantly improves acoustic contrast and reduces signal distortion in the presence of sound speed changes.
Sankha Subhra Bhattacharjee, Jesper Rindom Jensen, Mads Græsbøll Christensen
ICASSP2
2025 Advances in Microphone Array Processing and Multichannel Speech Enhancement
abstract
This paper reviews pioneering works in microphone array processing and multichannel speech enhancement, highlighting historical achievements, technological evolution, commercialization aspects, and key challenges. It provides valuable insights into the progression and future direction of these areas. The paper examines foundational developments in microphone array design and optimization, showcasing innovations that improved sound acquisition and enhanced speech intelligibility in noisy and reverberant environments. It then introduces recent advancements and cutting-edge research in the field, particularly the integration of deep learning techniques such as all-neural beamformers. The paper also explores critical applications, discussing their evolution and current state-of-the-art technologies that significantly impact user experience. Finally, the paper outlines future research directions, identifying challenges and potential solutions that could drive further innovation in these fields. By providing a comprehensive overview and forward-looking perspective, this paper aims to inspire ongoing research and contribute to the sustained growth and development of microphone arrays and multichannel speech enhancement.
Gongping Huang, Jesper Rindom Jensen, Jingdong Chen, Jacob Benesty, Mads Græsbøll Christensen, Akihiko Sugiyama, Gary W. Elko, Tomas Gänsler
ICASSP2
2024 Broadband Personal Sound Zone Control in the Presence of Nonlinearities
abstract
Existing literature on sound zone control generally consider the signal model to be linear. However, this is seldom true in practice owing to nonlinear distortions arising from the loudspeakers, especially in consumer applications. In this paper, we propose a new signal model for personal sound zone control that takes into consideration any nonlinear behaviour that may arise from the loudspeakers. Following the proposed signal model, the optimization problem is formulated such that it inherently ensures the reduction of nonlinear distortion effects in both the bright and dark zones. In addition, a broadband nonlinear acoustic contrast control - pressure matching approach is proposed for the new signal model. Simulation results on practical data show that our proposed approach can provide improvement in the acoustic contrast and/or the signal distortion performance compared to the traditional linear solution, in the presence of nonlinear distortions. Moreover, important observations are made for the study of nonlinear effects on sound zone control.
Sankha Subhra Bhattacharjee, Srikanth Burra, Jesper Rindom Jensen, Liming Shi, Guoli Ping, Jingkai Weng, Mads Græsbøll Christensen
ICASSP3
2024 Nonlinear acoustic echo cancellation using low-complexity low-rank recursive least-squares algorithms
Vinal Patel, Sankha Subhra Bhattacharjee, Jesper Rindom Jensen, Mads Græsbøll Christensen, Jacob Benesty
Signal Process.3
2023 Sparse Bayesian Learning Based Three-Dimensional Imaging for Antenna Array Radar
abstract
In recent years, the development of compressed sensing and sparse representation provide us with a broader perspective of three-dimensional (3-D) imaging. In this work, we propose a 3-D imaging method based on a sparse Bayesian learning(SBL) framework for antenna array radar. It solves the problem of long-term accumulation and complicated motion compensation problem that occurs with interferometric inverse synthetic aperture radar (InISAR). Using the framework, the proposed method can automatically learn optimal hyper-parameters from the data at a low computational cost. Experimental results show that the proposed method has advantages in terms of 3-D imaging accuracy and computational efficiency compared to existing methods.
Yuhan Li 0002, Jesper Rindom Jensen, Maozhong Fu, Zhenmiao Deng, Mads Græsbøll Christensen
ICASSP2
2023 Frequency Bin-Wise Single Channel Speech Presence Probability Estimation Using Multiple DNNS
abstract
In this work, we propose a frequency bin-wise method to estimate the single-channel speech presence probability (SPP) with multiple deep neural networks (DNNs) in the short-time Fourier transform domain. Since all frequency bins are typically considered simultaneously as input features for conventional DNN-based SPP estimators, high model complexity is inevitable. To reduce the model complexity and the requirements on the training data, we take a single frequency bin and some of its neighboring frequency bins into account to train separate gate recurrent units. In addition, the noisy speech and the a posteriori probability SPP representation are used to train our model. The experiments were performed on the Deep Noise Suppression challenge dataset. The experimental results show that the speech detection accuracy can be improved when we employ the frequency bin-wise model. Finally, we also demonstrate that our proposed method outperforms most of the state-of-the-art SPP estimation methods in terms of speech detection accuracy and model complexity.
Shuai Tao, Himavanth Reddy, Jesper Rindom Jensen, Mads Græsbøll Christensen
ICASSP3
2023 Recursive least-squares algorithm based on a third-order tensor decomposition for low-rank system identification
Constantin Paleologu, Jacob Benesty, Cristian Lucian Stanciu, Jesper Rindom Jensen, Mads Græsbøll Christensen, Silviu Ciochina
Signal Process.4
2022 Sparse Modeling of The Early Part of Noisy Room Impulse Responses with Sparse Bayesian Learning
abstract
A model of a room impulse response (RIR) is useful for a wide range of applications. Typically, the early part of a RIR is sparse, and its sparse structure allows for accurate and simple modeling of the RIR. The existing ℓp(0 < p ≤ 1)-norm-based methods suffer from the sensitivity to the user-selected regularization parameters or a high computational burden. In this work, we propose to reconstruct the sparse model for the early part of RIRs with sparse Bayesian learning (SBL). Under the framework of SBL, the proposed method can adaptively learn the optimal hyper-parameters from data at a low computational cost. Experiment results show that the proposed method has advantages in terms of noise robustness, reconstruction sparsity, and computational efficiency compared to the existing methods.
Maozhong Fu, Jesper Rindom Jensen, Yuhan Li 0002, Mads Græsbøll Christensen
ICASSP2
2021 Detecting Acoustic Reflectors Using A Robot's Ego-Noise
abstract
In this paper, we propose a method to estimate the proximity of an acoustic reflector, e.g., a wall, using ego-noise, i.e., the noise produced by the moving parts of a listening robot. This is achieved by estimating the times of arrival of acoustic echoes reflected from the surface. Simulated experiments show that the proposed non-intrusive approach is capable of accurately estimating the distance of a reflector up to 1 meter and outperforms a previously proposed intrusive approach under loud ego-noise conditions. The proposed method is helped by a probabilistic echo detector that estimates whether or not an acoustic reflector is within a short range of the robotic platform. This preliminary investigation paves the way towards a new kind of collision avoidance system that would purely rely on audio sensors rather than conventional proximity sensors.
Usama Saqib, Antoine Deleforge, Jesper Rindom Jensen
ICASSP3
2021 Automatic quality control and enhancement for voice-based remote Parkinson's disease detection
Amir Hossein Poorjam, Mathew Shaji Kavalekalam, Liming Shi, Yordan P. Raykov, Jesper Rindom Jensen, Max A. Little, Mads Græsbøll Christensen
Speech Commun.5
2020 A Model-based Approach to Acoustic Reflector Localization with a Robotic Platform
abstract
Constructing a spatial map of an indoor environment, e.g., a typical office environment with glass surfaces, is a difficult and challenging task. Current state-of-the-art, e.g., camera- and laser-based approaches are unsuitable for detecting transparent surfaces. Hence, the spatial map generated with these approaches are often inaccurate. In this paper, a method that utilizes echolocation with sound in the audible frequency range is proposed to robustly localize the position of an acoustic reflector, e.g., walls, glass surfaces etc., which could be used to construct a spatial map of an indoor environment as the robot moves. The proposed method estimate the acoustic reflector's position, using only a single microphone and a loudspeaker that are present on many socially assistive robot platforms such as the NAO robot. The experimental results show that the proposed method could robustly detect an acoustic reflector up to a distance of 1.5 m in more than 60% of the trials and works efficiently even under low SNRs. To test the proposed method, a proof-of-concept robotic platform was build to construct a spatial map of an indoor environment.
Usama Saqib, Jesper Rindom Jensen
IROS2
2020 Harmonic beamformers for speech enhancement and dereverberation in the time domain
Jesper Rindom Jensen, Sam Karimian-Azari, Mads Græsbøll Christensen, Jacob Benesty
Speech Commun.1
2019 Quality Control of Voice Recordings in Remote Parkinson's Disease Monitoring Using the Infinite Hidden Markov Model
abstract
The performance of voice-based systems for remote monitoring of Parkinson's disease is highly dependent on the degree of adherence of the recordings to the test protocols, which probe for specific symptoms. Identifying segments of the signal that adhere to the protocol assumptions is typically performed manually by experts. This process is costly, time consuming, and often infeasible for large-scale data sets. In this paper, we propose a method to automatically identify the segments of signals that violate the test protocol with a high accuracy. In our approach, the signal is first split into variable duration segments by fitting an infinite hidden Markov model (iHMM) to the frames of the signals in the mel-frequency cepstral domain. The complexity of the iHMM is capable of growing jointly with the data allowing us to infer a potentially large (asymptotically infinite) number of different phenomena segmented into different hidden states. Then, we identify the segments that adhere to the test protocol by applying a multinomial naive Bayes classifier to the state indicators of segments. The experimental results show that even by using a small amount of training data, we can achieve around 96% accuracy in identifying short-term protocol violations with a 0.2 s resolution.
Amir Hossein Poorjam, Yordan P. Raykov, Reham Badawy, Jesper Rindom Jensen, Mads Græsbøll Christensen, Max A. Little
ICASSP4
2019 Estimation of Fundamental Frequencies in Stereophonic Music Mixtures
abstract
In this paper, a method for multi-pitch estimation of stereophonic mixtures of harmonic signals, e.g., instrument recordings, is presented. The proposed method is based on a signal model that includes the panning parameters of the sources in a stereophonic mixture, such as those applied artificially in a recording studio. If the sources in a mixture have different panning parameters, this diversity can be used to simplify the pitch estimation problem. The mixing parameters of the sources might be shared, resulting in a multi-pitch estimation problem, which is solved using an approach based on an expectation-maximization algorithm for Gaussian sources, where the fundamental frequencies and model orders are estimated jointly. The fundamental frequencies may be related, resulting in overlapping harmonics, complicating the estimation of the parameters. A codebook of harmonic amplitude vectors is trained on recordings of instruments playing single notes, and used when estimating the amplitudes of the mixture components. The proposed method is evaluated using stereophonic mixtures of instrument recordings and is compared to state-of-the-art transcription and multi-pitch estimation methods. Experiments show an increase in performance when knowledge about the panning parameters is taken into account. The proposed method provides a full parameterization of the components of the observed signal. Possible applications include instrument tuning, audio editing tools, modification of harmonic mixture components, and audio effects.
Martin Weiss Hansen, Jesper Rindom Jensen, Mads Græsbøll Christensen
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 Robust Bayesian Pitch Tracking Based on the Harmonic Model
abstract
Fundamental frequency is one of the most important characteristics of speech and audio signals. Harmonic model-based fundamental frequency estimators offer a higher estimation accuracy and robustness against noise than the widely used autocorrelation-based methods. However, the traditional harmonic model-based estimators do not take the temporal smoothness of the fundamental frequency, the model order, and the voicing into account as they process each data segment independently. In this paper, a fully Bayesian fundamental frequency tracking algorithm based on the harmonic model and a first-order Markov process model is proposed. Smoothness priors are imposed on the fundamental frequencies, model orders, and voicing using first-order Markov process models. Using these Markov models, fundamental frequency estimation and voicing detection errors can be reduced. Using the harmonic model, the proposed fundamental frequency tracker has an improved robustness to noise. An analytical form of the likelihood function, which can be computed efficiently, is derived. Compared to the state-of-the-art neural network and nonparametric approaches, the proposed fundamental frequency tracking algorithm has superior performance in almost all investigated scenarios, especially in noisy conditions. For example, under 0 dB white Gaussian noise, the proposed algorithm reduces the mean absolute errors and gross errors by 15% and 20% on the Keele pitch database and 36% and 26% on sustained /a/ sounds from a database of Parkinson's disease voices. A MATLAB version of the proposed algorithm is made freely available for reproduction of the results.11An implementation of the proposed algorithm using MATLAB may be found in https://tinyurl.com/yxn4a543.
Liming Shi, Jesper Kjær Nielsen, Jesper Rindom Jensen, Max A. Little, Mads Græsbøll Christensen
IEEE ACM Trans. Audio Speech Lang. Process.3
2018 A Unified Approach to Generating Sound Zones Using Variable Span Linear Filters
abstract
Sound zones are typically created using Acoustic Contrast Control (ACC), Pressure Matching (PM), or variations of the two. ACC maximizes the acoustic potential energy contrast between a listening zone and a quiet zone. Although the contrast is maximized, the phase is not controlled. To control both the amplitude and the phase, PM instead minimizes the difference between the reproduced sound field and the desired sound field in all zones. On the surface, ACC and PM seem to control sound fields differently, but we here demonstrate they are actually extreme special cases of a much more general framework. The framework is inspired by the variable span linear filtering framework for speech enhancement. Using this framework, we demonstrate that 1) ACC gives the best contrast, but the highest signal distortion in the bright zone, and 2) PM gives the smallest signal distortion in the bright zone, but the worst contrast. Aside from showing this mathematically, we also demonstrate this via a small toy example.
Taewoong Lee, Jesper Kjær Nielsen, Jesper Rindom Jensen, Mads Græsbøll Christensen
ICASSP3
2018 A Parametric Approach for Classification of Distortions in Pathological Voices
abstract
In biomedical acoustics, distortion in voice signals, commonly present during acquisition and transmission, adversely affects acoustic features extracted from pathological voice. Information on the type of distortion can help in compensating for its effects. This paper proposes a new approach to detecting four major types of commonly encountered distortion in remote analysis of pathological voice, namely background noise, reverberation, clipping and coding. In this approach, by applying factor analysis to Gaussian mixture model mean supervectors, distortions in variable-duration recordings are modeled by fixed-length, low-dimensional channel vectors. Then, linear discriminant analysis (LDA) is used to remove the remaining nuisance effects in the channel vectors. Finally, two different classifiers, namely support vector machines and probabilistic LDA classify the different types of distortion. Experimental results obtained using Parkinson's voices, as an example of pathological voice, show 11.4% relative improvement in performance over systems which directly use acoustic features for distortion classification.
Amir Hossein Poorjam, Max A. Little, Jesper Rindom Jensen, Mads Græsbøll Christensen
ICASSP3
2018 A Supervised Approach to Global Signal-to-Noise Ratio Estimation for Whispered and Pathological Voices
abstract
The presence of background noise in signals adversely affects the performance of many speech-based algorithms. Accurate estimation of signal-to-noise-ratio (SNR), as a measure of noise level in a signal, can help in compensating for noise effects. Most existing SNR estimation methods have been developed for normal speech and might not provide accurate estimation for special speech types such as whispered or disordered voices, particularly, when they are corrupted by non-stationary noises. In this paper, we first investigate the impact of stationary and non-stationary noise on the behavior of mel-frequency cepstral coefficients (MFCCs) extracted from normal, whispered and pathological voices. We demonstrate that, regardless of the speech type, the mean and the covariance of MFCCs are predictably modified by additive noise and the amount of change is related to the noise level. Then, we propose a new supervised method for SNR estimation which is based on a regression model trained on MFCCs of the noisy signals. Experimental results show that the proposed approach provides accurate estimation and consistent performance for various speech types under different noise conditions.
Amir Hossein Poorjam, Max A. Little, Jesper Rindom Jensen, Mads Græsbøll Christensen
ICASSP3
2018 Multipitch Estimation Using Block Sparse Bayesian Learning and Intra-Block Clustering
abstract
Pitch estimation is an important task in speech and audio analysis. In this paper, we present a multi-pitch estimation algorithm based on block sparse Bayesian learning and intra-block clustering for speech analysis. A statistical hierarchical model is formulated based on a pitch dictionary with a fixed maximum number of harmonics for all the candidate pitches. Block sparse Bayesian learning is proposed for estimating the complex amplitudes. To deal with the problem of unknown harmonic orders and subharmonic errors, intra-block clustering structured sparsity prior is also introduced. The statistical update formulas are obtained by the variational Bayesian inference. Compared with the conventional group LASSO-type algorithms for multi-pitch estimation, experimental results indicate robustness against noise and improved estimation accuracy of the proposed method.
Liming Shi, Jesper Rindom Jensen, Jesper Kjær Nielsen, Mads Græsbøll Christensen
ICASSP2
2017 Estimation of multiple pitches in stereophonic mixtures using a codebook-based approach
abstract
In this paper, a method for multi-pitch estimation of stereophonic mixtures of multiple harmonic signals is presented. The method is based on a signal model which takes the amplitude and delay panning parameters of the sources in a stereophonic mixture into account. Furthermore, the method is based on the extended invariance principle (EXIP), and a codebook of realistic amplitude vectors. For each fundamental frequency candidate in each of the sources, the amplitude estimates are mapped to entries in the codebook, and the pitch and model order are estimated jointly. The performance of the proposed method is evaluated using mixtures of real signals. Experiments show an increase in performance when knowledge about the panning parameters is utilized together with the codebook of magnitude amplitudes when compared to a state-of-the-art transcription method.
Martin Weiss Hansen, Jesper Rindom Jensen, Mads Græsbøll Christensen
ICASSP2
2017 Harmonic minimum mean squared error filters for multichannel speech enhancement
abstract
Many state-of-the-art multichannel speech enhancement methods rely on second-order statistics of the desired speech signal, the noise signal, or both. Estimation of those are difficult in practice, resulting in a practical performance that is typically much lower than their potential theoretical performance. We propose two multichannel enhancement techniques that instead rely on a model for voiced speech. That is, the proposed methods are driven by the signals' fundamental frequencies, which may be accurately estimated even in noisy scenarios. The first method is designed independently of the microphone array geometry and source position, whereas these are utilized in the second approach. Thereby, we can investigate when to exploit such information in the case of localization errors and violations of the spatial assumptions. Numerical results show that the proposed method is able to outperform competing methods in terms of both output SNRs and PESQ scores.
Jesper Rindom Jensen, Mads Græsbøll Christensen, Andreas Jakobsson
ICASSP1
2017 Fast harmonic chirp summation
abstract
The harmonic chirp signal model has only very recently been introduced for modelling approximately periodic signals with a time-varying fundamental frequency. A number of estimators for the parameters of this model have already been proposed, but they are either inaccurate, non-robust to noise, or very computationally intensive. In this paper, we propose a fast algorithm for the harmonic chirp summation method which has been demonstrated in the literature to be accurate and robust to noise. The proposed algorithm is orders of magnitudes faster than previous algorithms which is also demonstrated via timing studies.
Jesper Kjær Nielsen, Tobias Lindstrøm Jensen, Jesper Rindom Jensen, Mads Græsbøll Christensen, Søren Holdt Jensen
ICASSP3
2017 Least 1-norm pole-zero modeling with sparse deconvolution for speech analysis
abstract
In this paper, we present a speech analysis method based on sparse pole-zero modeling of speech. Instead of using the all-pole model to approximate the speech production filter, a pole-zero model is used for the combined effect of the vocal tract; radiation at the lips and the glottal pulse shape. Moreover, to consider the spiky excitation form of the pulse train during voiced speech, the modeling parameters and sparse residuals are estimated in an iterative fashion using a least 1-norm pole-zero with sparse deconvolution algorithm. Compared with the conventional two-stage least squares pole-zero, linear prediction and sparse linear prediction methods, experimental results show that the proposed speech analysis method has lower spectral distortion, higher reconstruction SNR and sparser residuals.
Liming Shi, Jesper Rindom Jensen, Mads Græsbøll Christensen
ICASSP2
2017 Distributed max-SINR speech enhancement with ad hoc microphone arrays
abstract
In recent years, signal processing with ad hoc microphone arrays has attracted a lot of attention. Speech enhancement in noisy, interfered, and reverberant environments is one of the problems targeted by ad hoc microphone arrays. Most of the proposed solutions require knowledge of fingerprints, such as acoustic transfer functions, which may not be known as accurately as required in practical situations. In this paper, a distributed signal subspace filtering method is proposed which is not restricted to a special graph topology. Here, the maximum signal to interference-plus-noise ratio (max-SINR) criterion is used with the primal-dual method of multipliers for distributed filtering. The paper investigates the convergence of the algorithm in both synchronous and asynchronous schemes, and also discusses some practical pros and cons. The applicability of the proposed method is demonstrated by means of simulation results.
Vincent Mohammad Tavakoli, Jesper Rindom Jensen, Richard Heusdens, Jacob Benesty, Mads Græsbøll Christensen
ICASSP2
2017 Dominant Distortion Classification for Pre-Processing of Vowels in Remote Biomedical Voice Analysis
abstract
Advances in speech signal analysis facilitate the development of techniques for remote biomedical voice assessment. However, the performance of these techniques is affected by noise and distortion in signals. In this paper, we focus on the vowel /a/ as the most widely-used voice signal for pathological voice assessments and investigate the impact of four major types of distortion that are commonly present during recording or transmission in voice analysis, namely: background noise, reverberation, clipping and compression, on Mel-frequency cepstral coefficients (MFCCs) - the most widely-used features in biomedical voice analysis. Then, we propose a new distortion classification approach to detect the most dominant distortion in such voice signals. The proposed method involves MFCCs as frame-level features and a support vector machine as classifier to detect the presence and type of distortion in frames of a given voice signal. Experimental results obtained from the healthy and Parkinson's voices show the effectiveness of the proposed approach in distortion detection and classification.
Amir Hossein Poorjam, Jesper Rindom Jensen, Max A. Little, Mads Græsbøll Christensen
INTERSPEECH2
2017 Fast fundamental frequency estimation: Making a statistically efficient estimator computationally efficient
Jesper Kjær Nielsen, Tobias Lindstrøm Jensen, Jesper Rindom Jensen, Mads Græsbøll Christensen, Søren Holdt Jensen
Signal Process.3
2017 Correction to "Maximum Likelihood PSD Estimation for Speech Enhancement in Reverberation and Noise"
abstract
Presents corrections to the paper, "A Novel Approach Based on Marine Radar Data Analysis for High-Resolution Bathymetry Map Generation."
Adam Kuklasinski, Simon Doclo, Søren Holdt Jensen, Jesper Rindom Jensen
IEEE ACM Trans. Audio Speech Lang. Process.4
2016 Variable span filters for speech enhancement
abstract
In this work, we consider enhancement of multichannel speech recordings. Linear filtering and subspace approaches have been considered previously for solving the problem. The current linear filtering methods, although many variants exist, have limited control of noise reduction and speech distortion. Subspace approaches, on the other hand, can potentially yield better control by filtering in the eigen-domain, but traditionally these approaches have not been optimized explicitly for traditional noise reduction and signal distortion measures. Herein, we combine these approaches by deriving optimal filters using a joint diagonalization as a basis. This gives excellent control over the performance, as we can optimize for noise reduction or signal distortion performance. Results from real data experiments show that the proposed variable span filters can achieve better performance than existing filters. In terms of output SNR, the gain was more than 8 dB, and more than 0.1 in mean opinion score in the conducted experiments.
Jesper Rindom Jensen, Jacob Benesty, Mads Græsbøll Christensen
ICASSP1
2016 DOA estimation of audio sources in reverberant environments
abstract
Reverberation is well-known to have a detrimental impact on many localization methods for audio sources. We address this problem by imposing a model for the early reflections as well as a model for the audio source itself. Using these models, we propose two iterative localization methods that estimate the direction-of-arrival (DOA) of both the direct path of the audio source and the early reflections. In these methods, the contribution of the early reflections is essentially subtracted from the signal observations before localization of the direct path component, which may reduce the estimation bias. Our simulation results show that we can estimate the DOA of the desired signal more accurately with this procedure compared to state-of-the-art estimator in both synthetic and real data experiments with reverberation.
Jesper Rindom Jensen, Jesper Kjær Nielsen, Richard Heusdens, Mads Græsbøll Christensen
ICASSP1
2016 Fast and statistically efficient fundamental frequency estimation
abstract
Fundamental frequency estimation is a very important task in many applications involving periodic signals. For computational reasons, fast autocorrelation-based estimation methods are often used despite parametric estimation methods having superior estimation accuracy. However, these parametric methods are much more costly to run. In this paper, we propose an algorithm which significantly reduces the computational cost of an accurate maximum likelihood-based estimator for real-valued data. The computational cost is reduced by exploiting the matrix structure of the problem and by using a recursive solver. Via benchmarks, we demonstrate that the computation time is reduced by approximately two orders of magnitude. The proposed fast algorithm is available for download online.
Jesper Kjær Nielsen, Tobias Lindstrøm Jensen, Jesper Rindom Jensen, Mads Græsbøll Christensen, Søren Holdt Jensen
ICASSP3
2016 A partitioned approach to signal separation with microphone ad hoc arrays
abstract
In this paper, a blind algorithm is proposed for speech enhancement in multi-speaker scenarios, in which interference rejection is the main objective. Here, the ad hoc array is broken into microphone duples which are used to partition the array into local sub-arrays. The core algorithm takes advantage of differences in signal structure in each duple. A geometric mean filter is then used to merge the output signals obtained with different duples, and to form a global broadband maximum signal-to-interference ratio (SIR) enhancement apparatus. The resulting filter outputs are enhanced acoustic signals in terms of SIR, as shown with experiments.
Vincent Mohammad Tavakoli, Jesper Rindom Jensen, Jacob Benesty, Mads Græsbøll Christensen
ICASSP2
2016 Noise Reduction with Optimal Variable Span Linear Filters
abstract
In this paper, the problem of noise reduction is addressed as a linear filtering problem in a novel way by using concepts from subspace-based enhancement methods, resulting in variable span linear filters. This is done by forming the filter coefficients as linear combinations of a number of eigenvectors stemming from a joint diagonalization of the covariance matrices of the signal of interest and the noise. The resulting filters are flexible in that it is possible to trade off distortion of the desired signal for improved noise reduction. This tradeoff is controlled by the number of eigenvectors included in forming the filter. Using these concepts, a number of different filter designs are considered, like minimum distortion, Wiener, maximum SNR, and tradeoff filters. Interestingly, all these can be expressed as special cases of variable span filters. We also derive expressions for the speech distortion and noise reduction of the various filter designs. Moreover, we consider an alternative approach, wherein the filter is designed for extracting an estimate of the noise signal, which can then be extracted from the observed signals, which is referred to as the indirect approach. Simulations demonstrate the advantages and properties of the variable span filter designs, and their potential performance gain compared to widely used speech enhancement methods.
Jesper Rindom Jensen, Jacob Benesty, Mads Græsbøll Christensen
IEEE ACM Trans. Audio Speech Lang. Process.1
2016 Computationally Efficient and Noise Robust DOA and Pitch Estimation
abstract
Many natural signals, such as voiced speech and some musical instruments, are approximately periodic over short intervals. These signals are often described in mathematics by the sum of sinusoids (harmonics) with frequencies that are proportional to the fundamental frequency, or pitch. In sensor (microphone) array signal processing, the periodic signals are estimated from spatio-temporal samples regarding to the direction of arrival (DOA) of the signal of interest. In this paper, we consider the problem of pitch and DOA estimation of quasi-periodic audio signals. In real-life scenarios, recorded signals are often contaminated by different types of noise, which challenges the assumption of white Gaussian noise in most state-of-the-art methods. We establish filtering methods based on noise statistics to apply to nonparametric spectral and spatial parameter estimates of the harmonics. We design minimum variance solutions with distortionless constraints to estimate the pitch from the frequency estimates, and to estimate the DOA from multichannel phase estimates of the harmonics. Applying this filtering method as the sum of weighted frequency and DOA estimates of the harmonics, we also design a joint DOA and pitch estimator. In white Gaussian noise, we derive even more computationally efficient solutions which are designed using the narrowband power spectrum of the harmonics. Numerical results reveal the performance of the estimators in colored noise compared with the Cramér-Rao lower bound. Experiments on real-life signals indicate the applicability of the methods in practical low local signal-to-noise ratios.
Sam Karimian-Azari, Jesper Rindom Jensen, Mads Græsbøll Christensen
IEEE ACM Trans. Audio Speech Lang. Process.2
2016 Enhancement and Noise Statistics Estimation for Non-Stationary Voiced Speech
abstract
In this paper, single channel speech enhancement in the time domain is considered. We address the problem of modelling non-stationary speech by describing the voiced speech parts by a harmonic linear chirp model instead of using the traditional harmonic model. This means that the speech signal is not assumed stationary, instead the fundamental frequency can vary linearly within each frame. The linearly constrained minimum variance (LCMV) filter and the amplitude and phase estimation (APES) filter are derived in this framework and compared to the harmonic versions of the same filters. It is shown through simulations on synthetic and speech signals, that the chirp versions of the filters perform better than their harmonic counterparts in terms of output signal-to-noise ratio (SNR) and signal reduction factor. For synthetic signals, the output SNR for the harmonic chirp APES based filter is increased 3 dB compared to the harmonic APES based filter at an input SNR of 10 dB, and at the same time the signal reduction factor is decreased. For speech signals, the increase is 1.5 dB along with a decrease in the signal reduction factor of 0.7. As an implicit part of the APES filter, a noise covariance matrix estimate is obtained. We suggest using this estimate in combination with other filters such as the Wiener filter. The performance of the Wiener filter and LCMV filter are compared using the APES noise covariance matrix estimate and a power spectral density (PSD) based noise covariance matrix estimate. It is shown that the APES covariance matrix works well in combination with the Wiener filter, and the PSD based covariance matrix works well in combination with the LCMV filter.
Sidsel Marie Nørholm, Jesper Rindom Jensen, Mads Græsbøll Christensen
IEEE ACM Trans. Audio Speech Lang. Process.2
2016 Instantaneous Fundamental Frequency Estimation With Optimal Segmentation for Nonstationary Voiced Speech
abstract
In speech processing, the speech is often considered stationary within segments of 20-30 ms even though it is well known not to be true. In this paper, we take the nonstationarity of voiced speech into account by using a linear chirp model to describe the speech signal. We propose a maximum likelihood estimator of the fundamental frequency and chirp rate of this model, and show that it reaches the Cramer-Rao lower bound. Since the speech varies over time, a fixed segment length is not optimal, and we propose making a segmentation of the signal based on the maximum a posteriori criterion. Using this segmentation method, the segments are on average longer for the chirp model compared to the traditional harmonic model. For the signal under test, the average segment length is 24.4 and 17.1 ms for the chirp model and traditional harmonic model, respectively. This suggests a better fit of the chirp model than the harmonic model to the speech signal. The methods are based on an assumption of white Gaussian noise, and, therefore, two prewhitening filters are also proposed.
Sidsel Marie Nørholm, Jesper Rindom Jensen, Mads Græsbøll Christensen
IEEE ACM Trans. Audio Speech Lang. Process.2
2016 A Framework for Speech Enhancement With Ad Hoc Microphone Arrays
abstract
Speech enhancement is vital for improved listening practices. Ad hoc microphone arrays are promising assets for this purpose. Most well-established enhancement techniques with conventional arrays can be adapted into ad hoc scenarios. Despite recent efforts to introduce various ad hoc speech enhancement apparatus, a common framework for integration of conventional methods into this new scheme is still missing. This paper establishes such an abstraction based on inter and intra subarray speech coherencies. Along with measures for signal quality at the input of subarrays, a measure of coherency is proposed both for subarray selection in local enhancement approaches, and also for selecting a proper global reference when more than one subarray are used. Proposed methods within this framework are evaluated with regard to quantitative and qualitative measures, including array gains, the speech distortion ratio, the PESQ measure, and the STOI intelligibility measure. Major findings in this work are the observed changes in the superiority of different methods for certain conditions. When perceptual quality or intelligibility of the speech are the ultimate goals, there are turning points where the MVDR and the LCMV are superior to Wiener-based methods. Also, for certain scenarios, local approaches may be preferred to global ones.
Vincent Mohammad Tavakoli, Jesper Rindom Jensen, Mads Græsbøll Christensen, Jacob Benesty
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 Pitch and TDOA-based localization of acoustic sources with distributed arrays
abstract
In this paper, a method for acoustic source localization using distributed microphone arrays based on time-differences of arrival (TDOAs) is presented. The TDOAs are used to estimate the location of an acoustic source using a recently proposed method, based on a 4D parameter space defined by the 3D location of the source, and the TDOAs. The performance of the proposed method for acoustic source localization is compared to the performance of a method based on generalized cross-correlation with phase transform (GCC-PHAT) using synthetic and speech signals with varying source position. Results show a decrease in the error of the estimated position when the proposed method is used.
Martin Weiss Hansen, Jesper Rindom Jensen, Mads Græsbøll Christensen
ICASSP2
2015 A joint audio-visual approach to audio localization
abstract
Localization of audio sources is an important research problem, e.g., to facilitate noise reduction. In the recent years, the problem has been tackled using distributed microphone arrays (DMA). A common approach is to apply direction-of-arrival (DOA) estimation on each array (denoted as nodes), and then map the DOA estimates to a location. In practice, however, the individual nodes contain few microphones, limiting the DOA estimation accuracy and, thereby, also the localization performance. We investigate a new approach, where range estimates are also obtained and utilized from each node, e.g., using time-of-flight cameras. Moreover, we propose an optimal method for weighting such DOA and range information for audio localization. Our experiments on both synthetic and real data show that there is a clear, potential advantage of using the joint audio-visual localization framework.
Jesper Rindom Jensen, Mads Græsbøll Christensen
ICASSP1
2015 On frequency domain models for TDOA estimation
abstract
Time-difference-of-arrival (TDOA) estimation is an important problem in many microphone signal processing applications. Traditionally, this problem is solved by using a cross-correlation method, but in this paper we show that the cross-correlation method is actually a restricted special case of a much more general method. In this connection, we establish the conditions under which the crosscorrelation method is a statistically efficient estimator. One of the conditions is that the source signal is periodic with a known fundamental frequency of 2π/N radians per sample, where N is the number of data points, and a known number of harmonics. The more general method only relies on that the source signal is periodic and is, therefore, able to outperform the cross-correlation method in terms of estimation accuracy on both synthetic data and artificially delayed speech data. The simulation code is available online.
Jesper Rindom Jensen, Jesper Kjær Nielsen, Mads Græsbøll Christensen, Søren Holdt Jensen
ICASSP1
2015 Pitch estimation and tracking with harmonic emphasis on the acoustic spectrum
abstract
In this paper, we use unconstrained frequency estimates (UFEs) from a noisy harmonic signal and propose two methods to estimate and track the pitch over time. We assume that the UFEs are multivariate-normally-distributed random variables, and derive a maximum likelihood (ML) pitch estimator by maximizing the likelihood of the UFEs over short time-intervals. As the main contribution of this paper, we propose two state-space representations to model the pitch continuity, and, accordingly, we propose two Bayesian methods, namely a hidden Markov model and a Kalman filter. These methods are designed to optimally use the correlations in the consecutive pitch values, where the past pitch estimates are used to recursively update the prior distribution for the pitch variable. We perform experiments using synthetic data as well as a noisy speech recording, and show that the Bayesian methods provide more accurate estimates than the corresponding ML methods.
Sam Karimian-Azari, Nasser Mohammadiha, Jesper Rindom Jensen, Mads Græsbøll Christensen
ICASSP3
2015 Pseudo-coherence-based MVDR beamformer for speech enhancement with ad hoc microphone arrays
abstract
Speech enhancement with distributed arrays has been met with various methods. On the one hand, data independent methods require information about the position of sensors, so they are not suitable for dynamic geometries. On the other hand, Wiener-based methods cannot assure a distortionless output. This paper proposes minimum variance distortionless response filtering based on multichannel pseudo-coherence for speech enhancement with ad hoc microphone arrays. This method requires neither position information nor control of the trade-off used in the distortion weighted methods. Furthermore, certain performance criteria are derived in terms of the pseudo-coherence vector, and the method is compared with the multichannel Wiener filter. Evaluation shows the suitability of the proposed method in terms of noise reduction with minimum distortion in ad hoc scenarios.
Vincent Mohammad Tavakoli, Jesper Rindom Jensen, Mads Græsbøll Christensen, Jacob Benesty
ICASSP2
2015 Enhancement of non-stationary speech using harmonic chirp filters
abstract
In this paper, the issue of single channel speech enhancement of non-stationary voiced speech is addressed. The non-stationarity of speech is well known, but state of the art speech enhancement methods assume stationarity within frames of 20–30 ms. We derive optimal distortionless filters that take the non-stationarity nature of voiced speech into account via linear constraints. This is facilitated by imposing a harmonic chirp model on the speech signal. As an implicit part of the filter design, the noise statistics are also estimated based on the observed signal and parameters of the harmonic chirp model. Simulations on real speech show that the chirp based filters perform better than their harmonic counterparts. Further, it is seen that the gain of using the chirp model increases when the estimated chirp parameter is big corresponding to periods in the signal where the instantaneous fundamental frequency changes fast.
Sidsel Marie Nørholm, Jesper Rindom Jensen, Mads Græsbøll Christensen
INTERSPEECH2
2015 Least squares estimate of the initial phases in STFT based speech enhancement
abstract
In this paper, we consider single-channel speech enhancement in the short time Fourier transform (STFT) domain. We suggest to improve an STFT phase estimate by estimating the initial phases. The method is based on the harmonic model and a model for the phase evolution over time. The initial phases are estimated by setting up a least squares problem between the noisy phase and the model for phase evolution. Simulations on synthetic and speech signals show a decreased error on the phase when an estimate of the initial phase is included compared to using the noisy phase as an initialisation. The error on the phase is decreased at input SNRs from -10 to 10 dB. Reconstructing the signal using the clean amplitude, the mean squared error is decreased and the PESQ score is increased.
Sidsel Marie Nørholm, Martin Krawczyk-Becker, Timo Gerkmann, Steven van de Par, Jesper Rindom Jensen, Mads Græsbøll Christensen
INTERSPEECH5
2015 Joint Spatio-Temporal Filtering Methods for DOA and Fundamental Frequency Estimation
abstract
In this paper, spatio-temporal filtering methods are proposed for estimating the direction-of-arrival (DOA) and fundamental frequency of periodic signals, like those produced by the speech production system and many musical instruments using microphone arrays. This topic has quite recently received some attention in the community and is quite promising for several applications. The proposed methods are based on optimal, adaptive filters that leave the desired signal, having a certain DOA and fundamental frequency, undistorted and suppress everything else. The filtering methods simultaneously operate in space and time, whereby it is possible resolve cases that are otherwise problematic for pitch estimators or DOA estimators based on beamforming. Several special cases and improvements are considered, including a method for estimating the covariance matrix based on the recently proposed iterative adaptive approach (IAA). Experiments demonstrate the improved performance of the proposed methods under adverse conditions compared to the state of the art using both synthetic signals and real signals, as well as illustrate the properties of the methods and the filters.
Jesper Rindom Jensen, Mads Græsbøll Christensen, Jacob Benesty, Søren Holdt Jensen
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 Joint Pitch and DOA Estimation Using the ESPRIT Method
abstract
In this paper, the problem of joint multi-pitch and direction-of-arrival (DOA) estimation for multichannel harmonic sinusoidal signals is considered. A spatio-temporal matrix signal model for a uniform linear array is defined, and then the ESPRIT method based on subspace techniques that exploits the invariance property in the time domain is first used to estimate the multi pitch frequencies of multiple harmonic signals. Followed by the estimated pitch frequencies, the DOA estimations based on the ESPRIT method are also presented by using the shift invariance structure in the spatial domain. Compared to the existing state-of-the-art algorithms, the proposed method based on ESPRIT without 2-D searching is computationally more efficient but performs similarly. An asymptotic performance analysis of the DOA and pitch estimation of the proposed method are also presented. Finally, the effectiveness of the proposed method is illustrated on a synthetic signal as well as real-life recorded data.
Yuntao Wu, Amir Leshem, Jesper Rindom Jensen, Guisheng Liao
IEEE ACM Trans. Audio Speech Lang. Process.3
2014 Fundamental frequency and model order estimation using spatial filtering
abstract
In signal processing applications of harmonic-structured signals, estimates of the fundamental frequency and number of harmonics are often necessary. In real scenarios, a desired signal is contaminated by different levels of noise and interferers, which complicate the estimation of the signal parameters. In this paper, we present an estimation procedure for harmonic-structured signals in situations with strong interference using spatial filtering, or beamforming. We jointly estimate the fundamental frequency and the constrained model order through the output of the beamformers. Besides that, we extend this procedure to account for inharmonicity using unconstrained model order estimation. The simulations show that beamforming improves the performance of the joint estimates of fundamental frequency and the number of harmonics in low signal to interference (SIR) levels, and an experiment on a trumpet signal show the applicability on real signals.
Sam Karimian-Azari, Jesper Rindom Jensen, Mads Græsbøll Christensen
ICASSP2
2014 Noise reduction in the time domain using joint diagonalization
abstract
A new filter design based on joint diagonalization of the clean speech and noise covariance matrices is proposed. First, an estimate of the noise is found by filtering the observed signal. The filter for this is generated by a weighted sum of the eigenvectors from the joint diagonalization. Second, an estimate of the desired signal is found by subtraction of the noise estimate from the observed signal. The filter can be designed to obtain a desired trade-off between noise reduction and signal distortion, depending on the number of eigenvectors included in the filter design. This is explored through simulations using a speech signal corrupted by car noise, and the results confirm that the output signal-to-noise ratio and speech distortion index both increase when more eigenvectors are included in the filter design.
Sidsel Marie Nørholm, Jacob Benesty, Jesper Rindom Jensen, Mads Græsbøll Christensen
ICASSP3
2013 Multichannel signal enhancement using non-causal, time-domain filters
abstract
In the vast amount of time-domain filtering methods for speech enhancement, the filters are designed to be causal. Recently, however, it was shown that the noise reduction and signal distortion capabilities of such single-channel filters can be improved by allowing the filters to be non-causal. While non-causal filters require knowledge of the future, they can be implemented in practice by introducing a short delay. In this paper, we generalize the idea of exploiting non-causality in optimal filter designs to the multichannel scenario. More specifically, a set of optimal, non-causal, multichannel filters for enhancement based on an orthogonal decomposition is proposed. The evaluation shows that there is a potential gain in noise reduction and signal distortion by introducing non-causality. Moreover, experiments on real-life speech show that we can improve the perceptual quality.
Jesper Rindom Jensen, Mads Græsbøll Christensen, Jacob Benesty
ICASSP1
2013 Statistically efficient methods for pitch and DOA estimation
abstract
Traditionally, direction-of-arrival (DOA) and pitch estimation of multichannel, periodic sources have been considered as two separate problems. Separate estimation may render the task of resolving sources with similar DOA or pitch impossible, and it may decrease the estimation accuracy. Therefore, it was recently considered to estimate the DOA and pitch jointly. In this paper, we propose two novel methods for DOA and pitch estimation. They both yield maximum-likelihood estimates in white Gaussian noise scenarios, where the SNR may be different across channels, as opposed to state-of-the-art methods. The first method is a joint estimator, whereas the latter use a cascaded approach, but with a much lower computational complexity. The simulation results confirm that the proposed methods outperform state-of-the-art methods in terms of estimation accuracy in both synthetic and real-life signal scenarios.
Jesper Rindom Jensen, Mads Græsbøll Christensen, Søren Holdt Jensen
ICASSP1
2013 Joint DOA and fundamental frequency estimation based on relaxed iterative adaptive approach and optimal filtering
abstract
In this work, the problem of joint direction-of-arrival and fundamental frequency estimation for multi-channel harmonic sinusoidal signals is addressed. Different from the conventional optimal filtering method, we estimate the covariance matrix with the 2-D iterative adaptive approach, which is based on a single snapshot. In addition, to improve the estimation accuracy for the off-grid sources, a relaxation technique is utilized. Then, joint estimation is conducted on this covariance matrix estimate with the optimal filtering method. As a result, the relaxed iterative adaptive approach - optimal filtering method is devised. Statistical evaluation with synthetic signals shows the accurate performance of the proposed method compared with the Cramér-Rao lower bound.
Zhenhua Zhou, Mads Græsbøll Christensen, Jesper Rindom Jensen, Hing-Cheung So
ICASSP3
2013 A Class of Optimal Rectangular Filtering Matrices for Single-Channel Signal Enhancement in the Time Domain
abstract
In this paper, we introduce a new class of optimal rectangular filtering matrices for single-channel speech enhancement. The new class of filters exploits the fact that the dimension of the signal subspace is lower than that of the full space. By doing this, extra degrees of freedom in the filters, that are otherwise reserved for preserving the signal subspace, can be used for achieving an improved output signal-to-noise ratio (SNR). Moreover, the filters allow for explicit control of the tradeoff between noise reduction and speech distortion via the chosen rank of the signal subspace. An interesting aspect is that the framework in which the filters are derived unifies the ideas of optimal filtering and subspace methods. A number of different optimal filter designs are derived in this framework, and the properties and performance of these are studied using both synthetic, periodic signals and real signals. The results show a number of interesting things. Firstly, they show how speech distortion can be traded for noise reduction and vice versa in a seamless manner. Moreover, the introduced filter designs are capable of achieving both the upper and lower bounds for the output SNR via the choice of a single parameter.
Jesper Rindom Jensen, Jacob Benesty, Mads Græsbøll Christensen, Jingdong Chen
IEEE ACM Trans. Audio Speech Lang. Process.1
2013 Nonlinear Least Squares Methods for Joint DOA and Pitch Estimation
abstract
In this paper, we consider the problem of joint direction-of-arrival (DOA) and fundamental frequency estimation. Joint estimation enables robust estimation of these parameters in multi-source scenarios where separate estimators may fail. First, we derive the exact and asymptotic Cramér-Rao bounds for the joint estimation problem. Then, we propose a nonlinear least squares (NLS) and an approximate NLS (aNLS) estimator for joint DOA and fundamental frequency estimation. The proposed estimators are maximum likelihood estimators when: 1) the noise is white Gaussian, 2) the environment is anechoic, and 3) the source of interest is in the far-field. Otherwise, the methods still approximately yield maximum likelihood estimates. Simulations on synthetic data show that the proposed methods have similar or better performance than state-of-the-art methods for DOA and fundamental frequency estimation. Moreover, simulations on real-life data indicate that the NLS and aNLS methods are applicable even when reverberation is present and the noise is not white Gaussian.
Jesper Rindom Jensen, Mads Græsbøll Christensen, Søren Holdt Jensen
IEEE Trans. Speech Audio Process.1
2012 Non-Causal Time-Domain Filters for Single-Channel Noise Reduction
abstract
In many existing time-domain filtering methods for noise reduction in, e.g., speech processing, the filters are causal. Such causal filters can be implemented directly in practice. However, it is possible to improve the performance of such noise reduction filtering methods in terms of both noise suppression and signal distortion by allowing the filters to be non-causal. Non-causal time-domain filters require knowledge of the future, and are therefore not directly implementable. If the observed signal is processed in blocks, however, the non-causal filters are implementable. In this paper, we propose such non-causal time-domain filters for noise reduction in speech applications. We also propose some performance measures that enable us to evaluate the performance of non-causal filters. Moreover, it is shown how some of the filters can be updated recursively. Using the recursive expressions, it is also shown that the output SNRs of the filters always increase as we increase the length of the filter when the desired signal is stationary. From both the theoretical and practical evaluations of the filters, it is clearly shown that the performance of time-domain filtering methods for noise reduction can be improved by introducing non-causality.
Jesper Rindom Jensen, Jacob Benesty, Mads Græsbøll Christensen, Søren Holdt Jensen
IEEE Trans. Speech Audio Process.1
2012 Enhancement of Single-Channel Periodic Signals in the Time-Domain
abstract
Most state-of-the-art filtering methods for speech enhancement require an estimate of the noise statistics, but the noise statistics are difficult to estimate in practice when speech is present. Thus, nonstationary noise will have a detrimental impact on the performance of most speech enhancement filters. The impact of such noise can be reduced by using the signal statistics rather than the noise statistics in the filter design. For example, this is possible by assuming a harmonic model for the desired signal; while this model fits well for voiced speech, it will not be appropriate for unvoiced speech. That is, signal-dependent methods based on the signal statistics will introduce undesired distortion for some parts of speech compared to signal-independent methods based on the noise statistics. Since both the signal-independent and signal-dependent approaches to speech enhancement have advantages, it is relevant to combine them to reduce the impact of their individual disadvantages. In this paper, we give theoretical insights into the relationship between these different approaches, and these reveal a close relationship between the two approaches. This justifies joint use of such filtering methods which can be beneficial from a practical point of view. Our experimental results confirm that both signal-independent and signal-dependent approaches have advantages and that they are closely-related. Moreover, as a part of our experiments, we illustrate the practical usefulness of combining signal-independent and signal-dependent enhancement methods by applying such methods jointly on real-life speech.
Jesper Rindom Jensen, Jacob Benesty, Mads Græsbøll Christensen, Søren Holdt Jensen
IEEE Trans. Speech Audio Process.1
2011 A single snapshot optimal filtering method for fundamental frequency estimation
abstract
Recently, optimal linearly constrained minimum variance (LCMV) filtering methods have been applied for fundamental frequency estimation. Like many other fundamental frequency estimators, these methods utilize the inverse covariance matrix. Therefore, the covariance matrix needs to be invertible which is typically ensured by using the sample covariance matrix involving data partitioning. The partitioning adversely affects the spectral resolution. We propose a novel optimal filtering method which utilizes the LCMV principle in conjunction with the iterative adaptive approach (IAA). The IAA enables us to estimate the covariance matrix from a single snapshot, i.e., without data partitioning. The experimental results show, that the performance of the proposed method is comparable or better than that of other competing methods in terms of spectral resolution.
Jesper Rindom Jensen, Mads Græsbøll Christensen, Søren Holdt Jensen
ICASSP1
2009 Robust Parametric Audio Coding Using Multiple Description Coding
abstract
We propose a new multiple description spherical quantization with repetitively coded amplitudes (MDSQRA) scheme suited for quantization of sinusoidal parameters. The quantization scheme is constituted by a set of spherical quantizers inspired by the multiple description spherical trellis-coded quantization (MDSTCQ) scheme. In this scheme, we apply repetitive coding on the amplitudes, while multiple description coding are applied on the phases and frequencies. Thereby, MDSQRA becomes directly implementable, as opposed to MDSTCQ, since the phase and frequency quantizers depend on the amplitudes which have dissimilar descriptions in MDSTCQ. Furthermore, we implement MDSQRA into a perceptual matching pursuit based sinusoidal audio coder. Finally, we evaluate MDSQRA through perceptual distortion measurements and MUSHRA listening tests. The tests show that MDSQRA outperforms MDSTCQ with respect to a expected perceptual distortion measure. The same results are obtained through the MUSHRA tests performed on sound clips coded using MDSQRA and MDSTCQ.
Jesper Rindom Jensen, Mads Græsbøll Christensen, Morten Holm Jensens, Søren Holdt Jensen, Torben Larsen
IEEE Signal Process. Lett.1