Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Søren Holdt Jensen

dblp:93/2935 · DBLP profile ↗
← Back
121ranked-venue papers
1as first author
2since 2021 · last 2024
0000-0002-1050-470XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 81Artificial intelligence and machine learning · 44 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 7Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 2Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
28 papers
Audio and music processing · 97% Image and video coding · 2% Multimedia analysis and retrieval · 1%
Artificial intelligence
4 papers
Speech recognition and synthesis · 96% Probabilistic and Bayesian machine learning · 4%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%

Topics — the 30 heaviest of 46, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing
speech enhancement
1.372020
On Loss Functions for Supervised Monaural Time-Domain Speech Enhancement · IEEE ACM Trans. Audio Speech Lang. Process. 2020
Single-Channel Online Enhancement of Speech Corrupted by Reverberation and Noise · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Maximum Likelihood PSD Estimation for Speech Enhancement in Reverberation and Noise · IEEE ACM Trans. Audio Speech Lang. Process. 2016
Audio and music processing › speech quality assessment
speech intelligibility prediction
0.812024
Data-Driven Non-Intrusive Speech Intelligibility Prediction Using Speech Presence Probability · IEEE ACM Trans. Audio Speech Lang. Process. 2024
Audio and music processing
speech processing
0.742014
A frequency-domain adaptive filter (FDAF) prediction error method (PEM) framework for double-talk-robust acoustic echo cancellation · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Stable 1-Norm Error Minimization Based Linear Predictors for Speech Modeling · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Nonlinear Acoustic Echo Cancellation Based on a Sliding-Window Leaky Kernel Affine Projection Algorithm · IEEE Trans. Speech Audio Process. 2013
Audio and music processing › speech enhancement
noise reduction
0.642017
Single-Channel Online Enhancement of Speech Corrupted by Reverberation and Noise · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Enhancement of Single-Channel Periodic Signals in the Time-Domain · IEEE Trans. Speech Audio Process. 2012
Non-Causal Time-Domain Filters for Single-Channel Noise Reduction · IEEE Trans. Speech Audio Process. 2012
Audio and music processing › speech enhancement
dereverberation
0.522017
Single-Channel Online Enhancement of Speech Corrupted by Reverberation and Noise · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Maximum Likelihood PSD Estimation for Speech Enhancement in Reverberation and Noise · IEEE ACM Trans. Audio Speech Lang. Process. 2016
Audio and music processing › speech analysis
fundamental frequency estimation
0.552015
Default Bayesian Estimation of the Fundamental Frequency · IEEE Trans. Speech Audio Process. 2013
A Robust and Computationally Efficient Subspace-Based Fundamental Frequency Estimator · IEEE Trans. Speech Audio Process. 2010
Joint High-Resolution Fundamental Frequency and Order Estimation · IEEE Trans. Speech Audio Process. 2007
Natural language and speech › Speech recognition and synthesis › speech enhancement
active noise control
0.432013
Binaural Integrated Active Noise Control and Noise Reduction in Hearing Aids · IEEE Trans. Speech Audio Process. 2013
A Zone-of-Quiet Based Approach to Integrated Active Noise Control and Noise Reduction for Speech Enhancement in Hearing Aids · IEEE Trans. Speech Audio Process. 2012
Integrated Active Noise Control and Noise Reduction in Hearing Aids · IEEE Trans. Speech Audio Process. 2010
Natural language and speech › Speech recognition and synthesis › speech enhancement
noise reduction
0.432013
Binaural Integrated Active Noise Control and Noise Reduction in Hearing Aids · IEEE Trans. Speech Audio Process. 2013
A Zone-of-Quiet Based Approach to Integrated Active Noise Control and Noise Reduction for Speech Enhancement in Hearing Aids · IEEE Trans. Speech Audio Process. 2012
Integrated Active Noise Control and Noise Reduction in Hearing Aids · IEEE Trans. Speech Audio Process. 2010
Natural language and speech › Speech recognition and synthesis
speech enhancement
0.432013
Binaural Integrated Active Noise Control and Noise Reduction in Hearing Aids · IEEE Trans. Speech Audio Process. 2013
A Zone-of-Quiet Based Approach to Integrated Active Noise Control and Noise Reduction for Speech Enhancement in Hearing Aids · IEEE Trans. Speech Audio Process. 2012
Integrated Active Noise Control and Noise Reduction in Hearing Aids · IEEE Trans. Speech Audio Process. 2010
Audio and music processing › sound source localization
direction-of-arrival estimation
0.422015
Joint Spatio-Temporal Filtering Methods for DOA and Fundamental Frequency Estimation · IEEE ACM Trans. Audio Speech Lang. Process. 2015
Nonlinear Least Squares Methods for Joint DOA and Pitch Estimation · IEEE Trans. Speech Audio Process. 2013
Audio and music processing
acoustic echo cancellation
0.422014
A frequency-domain adaptive filter (FDAF) prediction error method (PEM) framework for double-talk-robust acoustic echo cancellation · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Nonlinear Acoustic Echo Cancellation Based on a Sliding-Window Leaky Kernel Affine Projection Algorithm · IEEE Trans. Speech Audio Process. 2013
Audio and music processing
speech coding
0.432014
Stable 1-Norm Error Minimization Based Linear Predictors for Speech Modeling · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Sparse Linear Prediction and Its Applications to Speech Processing · IEEE Trans. Speech Audio Process. 2012
Hidden Markov model-based packet loss concealment for voice over IP · IEEE Trans. Speech Audio Process. 2006
Audio and music processing
linear prediction
0.322014
Stable 1-Norm Error Minimization Based Linear Predictors for Speech Modeling · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Sparse Linear Prediction and Its Applications to Speech Processing · IEEE Trans. Speech Audio Process. 2012
Audio and music processing › audio representation
sinusoidal modeling
0.352011
Bayesian Interpolation and Parameter Estimation in a Dynamic Sinusoidal Model · IEEE Trans. Speech Audio Process. 2011
New Results on Single-Channel Speech Separation Using Sinusoidal Modeling · IEEE Trans. Speech Audio Process. 2011
New Results on Perceptual Distortion Minimization and Nonlinear Least-Squares Frequency Estimation · IEEE Trans. Speech Audio Process. 2011
Audio and music processing
acoustic signal processing
0.312017
A Scalable Algorithm for Physically Motivated and Sparse Approximation of Room Impulse Responses With Orthonormal Basis Functions · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Audio and music processing
room acoustics
0.312017
A Scalable Algorithm for Physically Motivated and Sparse Approximation of Room Impulse Responses With Orthonormal Basis Functions · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Audio and music processing › room acoustics
room impulse response modeling
0.312017
A Scalable Algorithm for Physically Motivated and Sparse Approximation of Room Impulse Responses With Orthonormal Basis Functions · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Audio and music processing › source separation › speech separation
single-channel speech separation
0.322012
A Joint Approach for Single-Channel Speaker Identification and Speech Separation · IEEE Trans. Speech Audio Process. 2012
New Results on Single-Channel Speech Separation Using Sinusoidal Modeling · IEEE Trans. Speech Audio Process. 2011
Audio and music processing › source separation
speech separation
0.322012
A Joint Approach for Single-Channel Speaker Identification and Speech Separation · IEEE Trans. Speech Audio Process. 2012
New Results on Single-Channel Speech Separation Using Sinusoidal Modeling · IEEE Trans. Speech Audio Process. 2011
Natural language and speech › Speech recognition and synthesis › speaker recognition
i-vector
0.212016
Total Variability Modeling Using Source-Specific Priors · IEEE ACM Trans. Audio Speech Lang. Process. 2016
Natural language and speech › Speech recognition and synthesis › speaker recognition
speaker embedding
0.212016
Total Variability Modeling Using Source-Specific Priors · IEEE ACM Trans. Audio Speech Lang. Process. 2016
Natural language and speech › Speech recognition and synthesis
speaker recognition
0.212016
Total Variability Modeling Using Source-Specific Priors · IEEE ACM Trans. Audio Speech Lang. Process. 2016
Audio and music processing › speech enhancement
multichannel wiener filter
0.212016
Maximum Likelihood PSD Estimation for Speech Enhancement in Reverberation and Noise · IEEE ACM Trans. Audio Speech Lang. Process. 2016
Audio and music processing › speech enhancement
power spectral density estimation
0.212016
Maximum Likelihood PSD Estimation for Speech Enhancement in Reverberation and Noise · IEEE ACM Trans. Audio Speech Lang. Process. 2016
Recommender systems
content-based recommendation
0.212014
Using Audio-Derived Affective Offset to Enhance TV Recommendation · IEEE Trans. Multim. 2014
Recommender systems › video recommendation
TV show recommendation
0.212014
Using Audio-Derived Affective Offset to Enhance TV Recommendation · IEEE Trans. Multim. 2014
Audio and music processing
emotion recognition
0.212014
Using Audio-Derived Affective Offset to Enhance TV Recommendation · IEEE Trans. Multim. 2014
Audio and music processing › audio coding
perceptual audio coding
0.222011
New Results on Perceptual Distortion Minimization and Nonlinear Least-Squares Frequency Estimation · IEEE Trans. Speech Audio Process. 2011
On perceptual distortion minimization and nonlinear least-squares frequency estimation · IEEE Trans. Speech Audio Process. 2006
Audio and music processing › speech processing
pitch estimation
0.222010
A Robust and Computationally Efficient Subspace-Based Fundamental Frequency Estimator · IEEE Trans. Speech Audio Process. 2010
Joint High-Resolution Fundamental Frequency and Order Estimation · IEEE Trans. Speech Audio Process. 2007
Image and video coding › error resilience
error concealment
0.212013
Sequential Error Concealment for Video/Images by Sparse Linear Prediction · IEEE Trans. Multim. 2013

Methods — techniques the papers use, named apart from their topics

speech presence probability · 0.8neural network · 0.8scale-invariant signal-to-distortion ratio · 0.4mean-square error · 0.4cramér-rao lower bound · 0.4hidden markov model · 0.3binaural processing · 0.3zone-of-quiet optimization · 0.3mean squared error optimization · 0.3wave equation approximation · 0.3orthonormal basis functions · 0.3bayesian filtering · 0.3autoregressive model · 0.3minimum divergence criterion · 0.2informative priors · 0.2factor analysis · 0.2filtered-x multichannel wiener filter · 0.2cascaded filtering · 0.2
YearPublicationVenuePosition
2024 Data-Driven Non-Intrusive Speech Intelligibility Prediction Using Speech Presence Probability
abstract
Time consuming Speech Intelligibility (SI) listening tests with human subjects can be replaced by algorithmic SI predictors. In recent years, data-driven SI predictors have been showing promising results. A major limiting factor in the advancement of data-driven SI prediction is that there is a scarcity of SI listening test data available to train the data-driven methods. In this article we propose a data-driven SI predictor that does not require access to an underlying noise-free reference signal, i.e.,non-intrusive, and which does not require listening test data for training. Instead, the proposed method exploits a hypothesized link between SI and Speech Presence Probability (SPP). We show that a neural network can be trained on easily obtainable speech in additive noise data to estimate SPP, and that a simple post-processing stage can be applied in order to map the estimated SPP to SI predictions with high accuracy. The proposed method is evaluated and compared to other state-of-the art non-intrusive SI predictors, and achieves the highest performance even in the presence of processed noisy speech, which the SPP estimator has not been trained on.
Mathias Bach Pedersen, Søren Holdt Jensen, Zheng-Hua Tan, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Modeling and Control Design for a Bidirectional DC-DC Converter System for Cyclic Operation of a Reversible Solid Oxide Electrolysis Cell Stack
abstract
This paper presents a design of a bidirectional DC- DC power electronic converter system enabling cyclic operation for a Reversible Solid Oxide Electrolysis Cell (RSO EC) stack for steam electrolysis. The cyclic operation of the RSOEC stack is investigated, and two different equivalent circuit models are presented for the mathematical representation of the stack's electrical dynamics: The well-established steady-state Resistive model and a novel Voigt model. From these, two combined mathematical models of the bidirectional Buck-Boost converter supplying the RSOEC stack are derived using the small-signal averaging technique. For tracking the cyclic output current reference to an RSOEC stack, two Proportional Integral Derivative with derivative Filter (PIDF) controllers are designed using the two combined mathematical models derived. Finally, the performances of the two PIDF controllers for the bidirectional DC-DC converter systems are compared and validated through simulations. The simulation results confirm the bidirectional Buck-Boost converter's ability to deliver cyclic bidirectional output current and demonstrate that the control tuned based on the Voigt mathematical representation of the RSOEC stack yields superior closed-loop performance in accordance with the control design requirements.
Kasper Jessen, Mohsen Soltani, Amin Hajizadeh, Søren Holdt Jensen, Erik Schaltz
IECON4
2020 A Neural Network for Monaural Intrusive Speech Intelligibility Prediction
abstract
Monaural intrusive speech intelligibility prediction (SIP) methods aim to predict the speech intelligibility (SI) of a single-microphone noisy and/or processed speech signal using the underlying clean speech signal. In the present work, we propose a neural network for monaural intrusive SIP. The proposed network is trained on data from multiple listening tests to predict SI. In the interest of using the available listening test data as efficiently as possible and to facilitate SI prediction of short duration speech signals, training is based on a local-time intelligibility curve derived from the listening test data. The trained neural network is evaluated, in terms of rank order correlation, against the classical monaural intrusive predictors STOI and ESTOI. The network is found to perform the best overall with a Kendall's tau of 0.825 measured over long duration, i.e. speech signals up to several minutes in duration. For short-term prediction using short speech signals of 1 - 10 seconds the network also shows better performance and smaller prediction variance.
Mathias Bach Pedersen, Asger Heidemann Andersen, Søren Holdt Jensen, Jesper Jensen 0001
ICASSP3
2020 End-to-End Speech Intelligibility Prediction Using Time-Domain Fully Convolutional Neural Networks
abstract
Data-driven speech intelligibility prediction has been slow totake off. Datasets of measured speech intelligibility are scarce,and so current models are relatively small and rely on hand-picked features. Classical predictors based on psychoacousticmodels and heuristics are still the state-of-the-art. This workproposes a U-Net inspired fully convolutional neural networkarchitecture, NSIP, trained and tested on ten datasets to pre-dict intelligibility of time-domain speech. The architecture iscompared to a frequency domain data-driven predictor and tothe classical state-of-the-art predictors STOI, ESTOI, HASPIand SIIB. The performance of NSIP is found to be superior fordatasets seen in the training phase. On unseen datasets NSIPreaches performance comparable to classical predictors.
Mathias Bach Pedersen, Morten Kolbæk, Asger Heidemann Andersen, Søren Holdt Jensen, Jesper Jensen 0001
INTERSPEECH4
2020 On Loss Functions for Supervised Monaural Time-Domain Speech Enhancement
abstract
Many deep learning-based speech enhancement algorithms are designed to minimize the mean-square error (MSE) in some transform domain between a predicted and a target speech signal. However, optimizing for MSE does not necessarily guarantee high speech quality or intelligibility, which is the ultimate goal of many speech enhancement algorithms. Additionally, only little is known about the impact of the loss function on the emerging class of time-domain deep learning-based speech enhancement systems. We study how popular loss functions influence the performance of time-domain deep learning-based speech enhancement systems. First, we demonstrate that perceptually inspired loss functions might be advantageous over classical loss functions like MSE. Furthermore, we show that the learning rate is a crucial design parameter even for adaptive gradient-based optimizers, which has been generally overlooked in the literature. Also, we found that waveform matching performance metrics must be used with caution as they in certain situations can fail completely. Finally, we show that a loss function based on scale-invariant signal-to-distortion ratio (SI-SDR) achieves good general performance across a range of popular speech enhancement evaluation metrics, which suggests that SI-SDR is a good candidate as a general-purpose loss function for speech enhancement systems.
Morten Kolbæk, Zheng-Hua Tan, Søren Holdt Jensen, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 Mean square performance evaluation in frequency domain for an improved adaptive feedback cancellation in hearing aids
Asutosh Kar, Jan Østergaard, Søren Holdt Jensen, M. N. S. Swamy 0001
Signal Process.4
2018 Audio-Based Granularity-Adapted Emotion Classification
abstract
This paper introduces a novel framework for combining the strengths of machine-based and human-based emotion classification. Peoples' ability to tell similar emotions apart is known as emotional granularity, which can be high or low, and is measurable. This paper proposes granularity-adapted classification that can be used as a front-end to drive a recommender, based on emotions from speech. In this context, incorrectly predicted peoples' emotions could lead to poor recommendations, reducing user satisfaction. Instead of identifying a single emotion class, an adapted class is proposed, and is an aggregate of underlying emotion classes chosen based on granularity. In the recommendation context, the adapted class maps to a larger region in valence-arousal space, from which a list of potentially more similar content items is drawn, and recommended to the user. To determine the effectiveness of adapted classes, we measured the emotional granularity of subjects, and for each subject, used their pairwise similarity judgments of emotion to compare the effectiveness of adapted classes versus single emotion classes taken from a baseline system. A customized Euclidean-based similarity metric is used to measure the relative proximity of emotion classes. Results show that granularity-adapted classification can improve the potential similarity by up to 9.6 percent.
Sven Ewan Shepstone, Zheng-Hua Tan, Søren Holdt Jensen
IEEE Trans. Affect. Comput.3
2017 Fast harmonic chirp summation
abstract
The harmonic chirp signal model has only very recently been introduced for modelling approximately periodic signals with a time-varying fundamental frequency. A number of estimators for the parameters of this model have already been proposed, but they are either inaccurate, non-robust to noise, or very computationally intensive. In this paper, we propose a fast algorithm for the harmonic chirp summation method which has been demonstrated in the literature to be accurate and robust to noise. The proposed algorithm is orders of magnitudes faster than previous algorithms which is also demonstrated via timing studies.
Jesper Kjær Nielsen, Tobias Lindstrøm Jensen, Jesper Rindom Jensen, Mads Græsbøll Christensen, Søren Holdt Jensen
ICASSP5
2017 Weighted Score Based Fast Converging CO-training with Application to Audio-Visual Person Identification
abstract
One potential problem in real classification applications is that the amount of labeled training data is insufficient since it is usually time-consuming to label data manually. When multiple modalities are available, it is possible to train an initial classifier for each modality using a small amount of labeled data, and then re-train each classifier using unlabeled data associated with the labels generated from the other modalities. This can be achieved by the well-known CO-training algorithm. Assuming that two modalities are available, it only takes the information from the other modality but not that from the self modality into account when choosing data, which usually results in slow convergence of classification accuracy. This may make the CO-training procedure time-consuming. To overcome this, we present a novel modification to the original CO-training algorithm, which is concerned with how new samples are chosen at each iteration to re-train the classifiers in order to improve the convergence of classification accuracy. In our method, the new data is chosen based on the weighted scores which are generated from both modalities instead of only the scores from the other modality as in the original CO-training. We apply both the modified and original CO-training methods on multi-modal person identification task using speech and vision. Experiments on a publicly available database show that our method outperforms the original CO-training by a large margin, in terms of convergence of classification accuracy on a separate testing data set.
Xiaodong Duan, Nicolai Bæk Thomsen, Zheng-Hua Tan, Børge Lindberg, Søren Holdt Jensen
ICTAI5
2017 Fast fundamental frequency estimation: Making a statistically efficient estimator computationally efficient
Jesper Kjær Nielsen, Tobias Lindstrøm Jensen, Jesper Rindom Jensen, Mads Græsbøll Christensen, Søren Holdt Jensen
Signal Process.5
2017 Real-Time Perceptual Model for Distraction in Interfering Audio-on-Audio Scenarios
abstract
This letter proposes a real-time perceptual model predicting the experienced distraction occurring in interfering audio-on-audio situations. The proposed model improves the computational efficiency of a previous distraction model, which cannot provide predictions in real time. The chosen approach was to utilize similar features as the previous model, but to use faster underlying algorithms to calculate these features. The results show that the proposed model has a root mean squared error of 11.9%, compared to the previous model's 11.0%, while only taking 0.04% of the computational time of the previous model. Thus, while providing similar accuracy as the previous model, the proposed model can be run in real time. The proposed distraction model can be used as a tool for evaluating and optimizing sound-zone systems. Furthermore, the real-time capability of the model introduces new possibilities, such as adaptive sound-zone systems.
Jussi Rämö, Soren Bech, Søren Holdt Jensen
IEEE Signal Process. Lett.3
2017 Single-Channel Online Enhancement of Speech Corrupted by Reverberation and Noise
abstract
This paper proposes an online single-channel speech enhancement method designed to improve the quality of speech degraded by reverberation and noise. Based on an autoregressive model for the reverberation power and on a hidden Markov model for clean speech production, a Bayesian filtering formulation of the problem is derived and online joint estimation of the acoustic parameters and mean speech, reverberation, and noise powers is obtained in mel-frequency bands. From these estimates, a real-valued spectral gain is derived and spectral enhancement is applied in the short-time Fourier transform (STFT) domain. The method yields state-of-the-art performance and greatly reduces the effects of reverberation and noise while improving speech quality and preserving speech intelligibility in challenging acoustic environments.
Clement S. J. Doire, Mike Brookes, Patrick A. Naylor, Christopher M. Hicks, Dave Betts, Mohammad A. Dmour, Søren Holdt Jensen
IEEE ACM Trans. Audio Speech Lang. Process.7
2017 Correction to "Maximum Likelihood PSD Estimation for Speech Enhancement in Reverberation and Noise"
abstract
Presents corrections to the paper, "A Novel Approach Based on Marine Radar Data Analysis for High-Resolution Bathymetry Map Generation."
Adam Kuklasinski, Simon Doclo, Søren Holdt Jensen, Jesper Rindom Jensen
IEEE ACM Trans. Audio Speech Lang. Process.3
2017 A Scalable Algorithm for Physically Motivated and Sparse Approximation of Room Impulse Responses With Orthonormal Basis Functions
abstract
Parametric modeling of room acoustics aims at representing room transfer functions by means of digital filters and finds application in many acoustic signal enhancement algorithms. In previous work by other authors, the use of orthonormal basis functions (OBFs) for modeling room acoustics has been proposed. Some advantages of OBF models over all-zero and pole-zero models have been illustrated, mainly focusing on the fact that OBF models typically require less model parameters to provide the same model accuracy. In this paper, it is shown that the orthogonality of the OBF model brings several additional advantages, which can be exploited if a suitable algorithm for identifying the OBF model parameters is applied. Specifically, the orthogonality of OBF models does not only lead to improved model efficiency (as pointed out in previous work), but also leads to improved model scalability and model stability. Its appealing scalability property derives from a previously unexplored interpretation of the OBF model as an approximation to a solution of the inhomogeneous acoustic wave equation. Following this interpretation, a novel identification algorithm is proposed that takes advantage of the OBF model orthogonality to deliver efficient, scalable, and stable OBF model estimates, which is not necessarily the case for nonlinear estimation techniques that are normally applied.
Giacomo Vairetti, Enzo De Sena, Michael Catrysse, Søren Holdt Jensen, Marc Moonen, Toon van Waterschoot
IEEE ACM Trans. Audio Speech Lang. Process.4
2016 On Perceptual Audio Compression with Side Information at the Decoder
abstract
Due to the distributed structure of many modern audio transmission setups, it is likely to have an observation at the receiver which is correlated with the desired source at the transmitter. This observation could be used as side information to reduce the transmission rate using distributed source coding. How to integrate distributed source coding into the perceptual audio compression procedure is thus a fundamental question. In this paper, we take a completely analytical approach to this problem, in particular to the rate-distortion trade-off and the corresponding coding schemes. We then interpret the results from an audio coding perspective. The main result is that, to upgrade a regular perceptual audio coder to a distributed coder, one needs to revise the perceptual masking curve. The revised masking curve models the availability of the side information as an extra masking effect, yielding lower rates. Interestingly, this means that at least conceptually, the distributed coding scenario could be integrated into the audio coder with minor changes, and without destructing the original coder.
Adel Zahedi, Jan Østergaard, Søren Holdt Jensen, Patrick A. Naylor, Soren Bech
DCC3
2016 Fast and statistically efficient fundamental frequency estimation
abstract
Fundamental frequency estimation is a very important task in many applications involving periodic signals. For computational reasons, fast autocorrelation-based estimation methods are often used despite parametric estimation methods having superior estimation accuracy. However, these parametric methods are much more costly to run. In this paper, we propose an algorithm which significantly reduces the computational cost of an accurate maximum likelihood-based estimator for real-valued data. The computational cost is reduced by exploiting the matrix structure of the problem and by using a recursive solver. Via benchmarks, we demonstrate that the computation time is reduced by approximately two orders of magnitude. The proposed fast algorithm is available for download online.
Jesper Kjær Nielsen, Tobias Lindstrøm Jensen, Jesper Rindom Jensen, Mads Græsbøll Christensen, Søren Holdt Jensen
ICASSP5
2016 Multichannel identification of room acoustic systems with adaptive filters based on orthonormal basis functions
abstract
Many acoustic signal enhancement applications require adaptive filters with a long impulse response, but with a small number of filter parameters. Fixed-poles infinite impulse response (IIR) adaptive filters based on orthonormal basis functions (OBFs) present advantages over finite impulse response filters and other IIR filters, assuring stability and fast global convergence in the adaptation of the filter parameters. A scalable algorithm is introduced for the estimation of the poles of an adaptive OBF filter from multichannel input-output data. The set of poles, common to all the acoustic channels considered, is estimated in parallel to the adaptation of the linear filter parameters. It will be shown that the result of the identification with common poles is quite robust to variations in the room transfer function, suggesting the possibility that poles may be kept fixed after estimation.
Giacomo Vairetti, Søren Holdt Jensen, Enzo De Sena, Marc Moonen, Michael Catrysse, Toon van Waterschoot
ICASSP2
2016 Speaker-Dependent Dictionary-Based Speech Enhancement for Text-Dependent Speaker Verification
abstract
The problem of text-dependent speaker verification under noisy conditions is becoming ever more relevant, due to increased usage for authentication in real-world applications. Classical methods for noise reduction such as spectral subtraction and Wiener filtering introduce distortion and do not perform well in this setting. In this work we compare the performance of different noise reduction methods under different noise conditions in terms of speaker verification when the text is known and the system is trained on clean data (mis-matched conditions). We furthermore propose a new approach based on dictionary-based noise reduction and compare it to the baseline methods.
Nicolai Bæk Thomsen, Dennis Alexander Lehmann Thomsen, Zheng-Hua Tan, Børge Lindberg, Søren Holdt Jensen
INTERSPEECH5
2016 Maximum Likelihood PSD Estimation for Speech Enhancement in Reverberation and Noise
abstract
In this contribution, we focus on the problem of power spectral density (PSD) estimation from multiple microphone signals in reverberant and noisy environments. The PSD estimation method proposed in this paper is based on the maximum likelihood (ML) methodology. In particular, we derive a novel ML PSD estimation scheme that is suitable for sound scenes which besides speech and reverberation consists of an additional noise component whose second-order statistics are known. The proposed algorithm is shown to outperform an existing similar algorithm in terms of PSD estimation accuracy. Moreover, it is shown numerically that the mean-squared estimation error achieved by the proposed method is near the limit set by the corresponding Cramér-Rao lower bound. The speech dereverberation performance of a multichannel Wiener filter based on the proposed PSD estimators is measured using several instrumental measures and is shown to be higher than when the competing estimator is used. Moreover, we perform a speech intelligibility test where we demonstrate that both the proposed and the competing PSD estimators lead to similar intelligibility improvements.
Adam Kuklasinski, Simon Doclo, Søren Holdt Jensen, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2016 Total Variability Modeling Using Source-Specific Priors
abstract
In total variability modeling, variable length speech utterances are mapped to fixed low-dimensional i-vectors. Central to computing the total variability matrix and i-vector extraction, is the computation of the posterior distribution for a latent variable conditioned on an observed feature sequence of an utterance. In both cases the prior for the latent variable is assumed to be non-informative, since for homogeneous datasets there is no gain in generality in using an informative prior. This work shows in the heterogeneous case, that using informative priors for computing the posterior, can lead to favorable results. We focus on modeling the priors using minimum divergence criterion or factor analysis techniques. Tests on the NIST 2008 and 2010 Speaker Recognition Evaluation (SRE) dataset show that our proposed method beats four baselines: For i-vector extraction using an already trained matrix, for the short2-short3 task in SRE'08, five out of eight female and four out of eight male common conditions, were improved. For the core-extended task in SRE'10, four out of nine female and six out of nine male common conditions were improved. When incorporating prior information into the training of the T matrix itself, the proposed method beats the baselines for six out of eight female and five out of eight male common conditions, for SRE'08, and five and six out of nine conditions, for the male and female case, respectively, for SRE'10. Tests using factor analysis for estimating priors show that two priors do not offer much improvement, but in the case of three separate priors (sparse data), considerable improvements were gained.
Sven Ewan Shepstone, Kong-Aik Lee, Haizhou Li 0001, Zheng-Hua Tan, Søren Holdt Jensen
IEEE ACM Trans. Audio Speech Lang. Process.5
2015 Coding and Enhancement in Wireless Acoustic Sensor Networks
abstract
We formulate a new problem which bridges between source coding and enhancement in wireless acoustic sensor networks. We consider a network of wireless microphones, each of which encoding its own measurement under a covariance matrix distortion constraint and sending it to a fusion center. To process the data at the center, we use a recent spatio-temporal prediction filter. We assume that a weighted sum-rate for the network is specified. The problem is to allocate optimal distortion matrices to the nodes in order to achieve a maximum output SNR at the fusion center after processing the received data, while the weighted sum-rate for the network is no more than the specified value. We formulate this problem as an optimization problem for which we derive a set of equalities imposed on the solution by studying the KKT conditions. In particular, for the special case of scalar sources with two microphones and a sum-rate constraint, we derive the distortion allocation in closed form and will show that if the given sum-rate is higher than a critical value, the stationary points from the KKT conditions lead to distortion allocations which maximize the output SNR of the filter.
Adel Zahedi, Jan Østergaard, Søren Holdt Jensen, Patrick A. Naylor, Soren Bech
DCC3
2015 Single-channel blind estimation of reverberation parameters
abstract
The reverberation of an acoustic channel can be characterised by two frequency-dependent parameters: the reverberation time and the direct-to-reverberant energy ratio. This paper presents an algorithm for blindly determining these parameters from a single-channel speech signal. The algorithm uses an extended Kalman filter to estimate the parameters together with a hidden semi-Markov model to identify intervals of speech activity.
Clement S. J. Doire, Mike Brookes, Patrick A. Naylor, Dave Betts, Christopher M. Hicks, Mohammad A. Dmour, Søren Holdt Jensen
ICASSP7
2015 On frequency domain models for TDOA estimation
abstract
Time-difference-of-arrival (TDOA) estimation is an important problem in many microphone signal processing applications. Traditionally, this problem is solved by using a cross-correlation method, but in this paper we show that the cross-correlation method is actually a restricted special case of a much more general method. In this connection, we establish the conditions under which the crosscorrelation method is a statistically efficient estimator. One of the conditions is that the source signal is periodic with a known fundamental frequency of 2π/N radians per sample, where N is the number of data points, and a known number of harmonics. The more general method only relies on that the source signal is periodic and is, therefore, able to outperform the cross-correlation method in terms of estimation accuracy on both synthetic data and artificially delayed speech data. The simulation code is available online.
Jesper Rindom Jensen, Jesper Kjær Nielsen, Mads Græsbøll Christensen, Søren Holdt Jensen
ICASSP4
2015 Multi-channel PSD estimators for speech dereverberation - A theoretical and experimental comparison
abstract
In this paper we perform an extensive theoretical and experimental comparison of two recently proposed multi-channel speech dereverberation algorithms. Both of them are based on the multi-channel Wiener filter but they use different estimators of the speech and reverberation power spectral densities (PSDs). We first derive closedform expressions for the mean square error (MSE) of both PSD estimators and then show that one estimator - previously used for speech dereverberation by the authors - always yields a better MSE. Only in the case of a two microphone array or for special spatial distributions of the interference both estimators yield the same MSE. The theoretically derived MSE values are in good agreement with numerical simulation results and with instrumental speech quality measures in a realistic speech dereverberation task for binaural hearing aids.
Adam Kuklasinski, Simon Doclo, Timo Gerkmann, Søren Holdt Jensen, Jesper Jensen 0001
ICASSP4
2015 Source-specific informative prior for i-vector extraction
abstract
An i-vector is a low-dimensional fixed-length representation of a variable-length speech utterance, and is defined as the posterior mean of a latent variable conditioned on the observed feature sequence of an utterance. The assumption is that the prior for the latent variable is non-informative, since for homogeneous datasets there is no gain in generality in using an informative prior. This work shows that extracting i-vectors for a heterogeneous dataset, containing speech samples recorded from multiple sources, using informative priors instead is applicable, and leads to favorable results. Tests carried out on the NIST 2008 and 2010 Speaker Recognition Evaluation (SRE) dataset show that our proposed method beats three baselines: For the short2-short3 core-task in SRE'08, for the female and male cases, five and six respectively, out of eight common conditions were beaten, and for the core-core task in SRE'10, for both genders, five out of nine common conditions were beaten.
Sven Ewan Shepstone, Kong-Aik Lee, Haizhou Li 0001, Zheng-Hua Tan, Søren Holdt Jensen
ICASSP5
2015 A heuristic approach for a social robot to navigate to a person based on audio and range information
abstract
The use of social robots for elderly care is becoming ever more relevant, thus introducing new challenges which need to be solved to achieve acceptable performance. One fundamental task for a social robot is to move to the person of interest in order to start interacting or perform a service. In this paper we address the task of a robot having to navigate to a possibly occluded person, which needs assistance, based only on audio and range information. Our approach is based on forming a heuristic cost function which is based on combining two audio features, and then moving to the optimum position indicated by this cost function after every interaction with the person. The method is compared to a greedy approach in 20 different tasks using a loud speaker playing at approximately 60dB sound pressure level (SPL) to mimic a human speaker, and the proposed method shows superior performance. A second experiment with an increase of 13dB SPL of the loud speaker is conducted and the proposed method is able to handle this.
Nicolai Bæk Thomsen, Zheng-Hua Tan, Børge Lindberg, Søren Holdt Jensen
IROS4
2015 Audio coding in wireless acoustic sensor networks
Adel Zahedi, Jan Østergaard, Søren Holdt Jensen, Soren Bech, Patrick A. Naylor
Signal Process.3
2015 Joint Spatio-Temporal Filtering Methods for DOA and Fundamental Frequency Estimation
abstract
In this paper, spatio-temporal filtering methods are proposed for estimating the direction-of-arrival (DOA) and fundamental frequency of periodic signals, like those produced by the speech production system and many musical instruments using microphone arrays. This topic has quite recently received some attention in the community and is quite promising for several applications. The proposed methods are based on optimal, adaptive filters that leave the desired signal, having a certain DOA and fundamental frequency, undistorted and suppress everything else. The filtering methods simultaneously operate in space and time, whereby it is possible resolve cases that are otherwise problematic for pitch estimators or DOA estimators based on beamforming. Several special cases and improvements are considered, including a method for estimating the covariance matrix based on the recently proposed iterative adaptive approach (IAA). Experiments demonstrate the improved performance of the proposed methods under adverse conditions compared to the state of the art using both synthetic signals and real signals, as well as illustrate the properties of the methods and the filters.
Jesper Rindom Jensen, Mads Græsbøll Christensen, Jacob Benesty, Søren Holdt Jensen
IEEE ACM Trans. Audio Speech Lang. Process.4
2014 Distributed Remote Vector Gaussian Source Coding for Wireless Acoustic Sensor Networks
abstract
In this paper, we consider the problem of remote vector Gaussian source coding for a wireless acoustic sensor network. Each node receives messages from multiple nodes in the network and decodes these messages using its own measurement of the sound field as side information. The node's measurement and the estimates of the source resulting from decoding the received messages are then jointly encoded and transmitted to a neighbouring node in the network. We show that for this distributed source coding scenario, one can encode a so-called conditional sufficient statistic of the sources instead of jointly encoding multiple sources. We focus on the case where node measurements are in form of noisy linearly mixed combinations of the sources and the acoustic channel mixing matrices are invertible. For this problem, we derive the rate-distortion function for vector Gaussian sources and under covariance distortion constraints.
Adel Zahedi, Jan Østergaard, Søren Holdt Jensen, Patrick A. Naylor, Soren Bech
DCC3
2014 Joint sparsity and frequency estimation for spectral compressive sensing
abstract
Parameter estimation from compressively sensed signals has recently received some attention. We here also consider this problem in the context of frequency sparse signals which are encountered in many application. Existing methods perform the estimation using finite dictionaries or incorporate various interpolation techniques to estimate the continuous frequency parameters. In this paper, we show that solving the problem in a probabilistic framework instead produces an asymptotically efficient estimator which outperforms existing methods in terms of estimation accuracy while still having a low computational complexity. Moreover, the proposed algorithm is also able to make inference about the sparsity level of the measured signal. The simulation code is available online.
Jesper Kjær Nielsen, Mads Græsbøll Christensen, Søren Holdt Jensen
ICASSP3
2014 Model selection and comparison for independents sinusoids
abstract
In the signal processing literature, many methods have been proposed for estimating the number of sinusoidal basis functions from a noisy data set. The most popular method is the asymptotic MAP criterion, which is sometimes also referred to as the BIC. In this paper, we extend and improve this method by considering the problem in a full Bayesian framework instead of the approximate formulation, on which the asymptotic MAP criterion is based. This leads to a new model selection and comparison method, the lp-BIC, whose computational complexity is of the same order as the asymptotic MAP criterion. Through simulations, we demonstrate that the lp-BIC outperforms the asymptotic MAP criterion and other state of the art methods in terms of model selection, de-noising and prediction performance. The simulation code is available online.
Jesper Kjær Nielsen, Mads Græsbøll Christensen, Søren Holdt Jensen
ICASSP3
2014 Distributed remote vector gaussian source coding with covariance distortion constraints
abstract
In this paper, we consider a distributed remote source coding problem, where a sequence of observations of source vectors is available at the encoder. The problem is to specify the optimal rate for encoding the observations subject to a covariance matrix distortion constraint and in the presence of side information at the decoder. For this problem, we derive lower and upper bounds on the rate-distortion function (RDF) for the Gaussian case, which in general do not coincide. We then provide some cases, where the RDF can be derived exactly. We also show that previous results on specific instances of this problem can be generalized using our results. We finally show that if the distortion measure is the mean squared error, or if it is replaced by a certain mutual information constraint, the optimal rate can be derived from our main result.
Adel Zahedi, Jan Østergaard, Søren Holdt Jensen, Patrick A. Naylor, Soren Bech
ISIT3
2014 Wiener variable step size and gradient spectral variance smoothing for double-talk-robust acoustic echo cancellation and acoustic feedback cancellation
Jose Manuel Gil-Cacho, Toon van Waterschoot, Marc Moonen, Søren Holdt Jensen
Signal Process.4
2014 Stable 1-Norm Error Minimization Based Linear Predictors for Speech Modeling
abstract
In linear prediction of speech, the 1-norm error minimization criterion has been shown to provide a valid alternative to the 2-norm minimization criterion. However, unlike 2-norm minimization, 1-norm minimization does not guarantee the stability of the corresponding all-pole filter and can generate saturations when this is used to synthesize speech. In this paper, we introduce two new methods to obtain intrinsically stable predictors with the 1-norm minimization. The first method is based on constraining the roots of the predictor to lie within the unit circle by reducing the numerical range of the shift operator associated with the particular prediction problem considered. The second method uses the alternative Cauchy bound to impose a convex constraint on the predictor in the 1-norm error minimization. These methods are compared with two existing methods: the Burg method, based on the 1-norm minimization of the forward and backward prediction error, and the iteratively reweighted 2-norm minimization known to converge to the 1-norm minimization with an appropriate selection of weights. The evaluation gives proof of the effectiveness of the new methods, performing as well as unconstrained 1-norm based linear prediction for modeling and coding of speech.
Daniele Giacobello, Mads Græsbøll Christensen, Tobias Lindstrøm Jensen, Manohar N. Murthi, Søren Holdt Jensen, Marc Moonen
IEEE ACM Trans. Audio Speech Lang. Process.5
2014 A frequency-domain adaptive filter (FDAF) prediction error method (PEM) framework for double-talk-robust acoustic echo cancellation
abstract
In this paper, we propose a new framework to tackle the double-talk (DT) problem in acoustic echo cancellation (AEC). It is based on a frequency-domain adaptive filter (FDAF) implementation of the so-called prediction error method adaptive filtering using row operations (PEM-AFROW) leading to the FDAF-PEM-AFROW algorithm. We show that FDAF-PEM-AFROW is by construction related to the best linear unbiased estimate (BLUE) of the echo path. We depart from this framework to show an improvement in performance with respect to other adaptive filters minimizing the BLUE criterion, namely the PEM-AFROW and the FDAF-NLMS with near-end signal normalization. One of the contributions is to propose the instantaneous pseudo-correlation (IPC) measure between the near-end signal and the loudspeaker signal. The IPC measure serves as an indication of the effect of a DT situation occurring during adaptation. We motivate the choice of FDAF-PEM-AFROW over PEM-AFROW and FDAF-NLMS with near-end signal normalization, based on performance, computational complexity and related IPC measure values. Moreover, we use the FDAF-PEM-AFROW framework to improve several state-of-the-art variable step-size (VSS) and variable regularization (VR) algorithms. The FDAF-PEM-AFROW versions significantly outperform the original versions in every simulation. In terms of computational complexity, the FDAF-PEM-AFROW versions are themselves about two orders of magnitude cheaper than the original versions.
Jose Manuel Gil-Cacho, Toon van Waterschoot, Marc Moonen, Søren Holdt Jensen
IEEE ACM Trans. Audio Speech Lang. Process.4
2014 Using Audio-Derived Affective Offset to Enhance TV Recommendation
abstract
This paper introduces the concept of affective offset, which is the difference between a user's perceived affective state and the affective annotation of the content they wish to see. We show how this affective offset can be used within a framework for providing recommendations for TV programs. First a user's mood profile is determined using 12-class audio-based emotion classifications . An initial TV content item is then displayed to the user based on the extracted mood profile. The user has the option to either accept the recommendation, or to critique the item once or several times, by navigating the emotion space to request an alternative match. The final match is then compared to the initial match, in terms of the difference in the items' affective parameterization . This offset is then utilized in future recommendation sessions. The system was evaluated by eliciting three different moods in 22 separate users and examining the influence of applying affective offset to the users' sessions. Results show that, in the case when affective offset was applied, better user satisfaction was achieved: the average ratings went from 7.80 up to 8.65, with an average decrease in the number of critiquing cycles which went from 29.53 down to 14.39.
Sven Ewan Shepstone, Zheng-Hua Tan, Søren Holdt Jensen
IEEE Trans. Multim.3
2013 Analysis of closed-loop acoustic feedback cancellation systems
abstract
In a previous study, the performance of an acoustic feedback/echo cancellation system was analyzed using a power transfer function method. Whereas the analysis result provides very accurate performance predictions in open-loop acoustic echo cancellation systems, it is less accurate in closed-loop acoustic feedback cancellation systems if there is a strong correlation between the loudspeaker signal and the signals entering the microphones. This work extends the performance analysis to include the effects of the nonzero correlation on the adaptive filters. Simulation results verify that this extension provides much more accurate performance predictions in closed-loop acoustic feedback cancellation systems.
Meng Guo 0001, Søren Holdt Jensen, Jesper Jensen 0001, Steven L. Grant
ICASSP2
2013 Statistically efficient methods for pitch and DOA estimation
abstract
Traditionally, direction-of-arrival (DOA) and pitch estimation of multichannel, periodic sources have been considered as two separate problems. Separate estimation may render the task of resolving sources with similar DOA or pitch impossible, and it may decrease the estimation accuracy. Therefore, it was recently considered to estimate the DOA and pitch jointly. In this paper, we propose two novel methods for DOA and pitch estimation. They both yield maximum-likelihood estimates in white Gaussian noise scenarios, where the SNR may be different across channels, as opposed to state-of-the-art methods. The first method is a joint estimator, whereas the latter use a cascaded approach, but with a much lower computational complexity. The simulation results confirm that the proposed methods outperform state-of-the-art methods in terms of estimation accuracy in both synthetic and real-life signal scenarios.
Jesper Rindom Jensen, Mads Græsbøll Christensen, Søren Holdt Jensen
ICASSP3
2013 Real-time implementations of sparse linear prediction for speech processing
abstract
Employing sparsity criteria in linear prediction of speech has been proven successful for several analysis and coding purposes. However, sparse linear prediction comes at the expenses of a much higher computational burden and numerical sensitivity compared to the traditional minimum variance approach. This makes sparse linear prediction difficult to deploy in real-time systems. In this paper, we present a step towards real-time implementation of the sparse linear prediction problem using hand-tailored interior-point methods. Using compiled implementations the sparse linear prediction problems corresponding to a frame size of 20ms can be solved on a standard PC in approximately 2ms and orders faster than with general purpose software.
Tobias Lindstrøm Jensen, Daniele Giacobello, Mads Græsbøll Christensen, Søren Holdt Jensen, Marc Moonen
ICASSP4
2013 Bayesian model comparison and the BIC for regression models
abstract
In the signal processing literature, many methods have been proposed for solving the important model comparison and selection problem. However, most of these methods only find the most likely model or only work well under particular circumstances such as a large number of data points or a high signal-to-noise ratio (SNR). One of the most successful classes of methods is the Bayesian information criteria (BIC) and in this paper, we extend some of the recent work on the BIC. In particular, we develop methods in a full Bayesian framework which work well across a large/small number of data points and high/low SNR for either real- or complex-valued data originating from a regression model. Aside from selecting the most probable model, these rules can also be used for model averaging as they assign a probability to each candidate model. Through simulations on a polynomial trend model, we demonstrate that the proposed rules outperform other rules in terms of detecting the true model order, de-noising the noisy signal, and making predictions of unobserved data points. The simulation code is available online.
Jesper Kjær Nielsen, Mads Græsbøll Christensen, Søren Holdt Jensen
ICASSP3
2013 Demographic recommendation by means of group profile elicitation using speaker age and gender recognition
abstract
In this paper we show a new method of using automatic age and gender recognition to recommend a sequence of multimedia items to a home TV audience comprising multiple viewers. Instead of relying on explicitly provided demographic data for each user, we define an audio-based demographic group profile that captures the age and gender for all members of the audience. A 7-class age and gender classifier employing a fusion of acoustic and prosodic features determines the probability of each speaker belonging to each class. The information for all speakers is then combined to form the group profile, which itself is the input to a recommender system. The recommender system finds the content items whose demographics best match the group profile. We tested the effectiveness of the system for several typical home audience configurations. In a survey, users were given a configuration and asked to rate a set of advertisements on how well each advertisement matched the configuration. Unbeknown to the subjects, half of the adverts were recommended using the derived audio demographics and the other half were randomly chosen. The recommended adverts received a significantly higher median rating of 7.75, as opposed to 4.25 for the randomly selected adverts.
Sven Ewan Shepstone, Zheng-Hua Tan, Søren Holdt Jensen
INTERSPEECH3
2013 Improved prediction error filters for adaptive feedback cancellation in hearing aids
Kim Ngo, Toon van Waterschoot, Mads Græsbøll Christensen, Marc Moonen, Søren Holdt Jensen
Signal Process.5
2013 A speech distortion weighting based approach to integrated active noise control and noise reduction in hearing aids
Romain Serizel, Marc Moonen, Jan Wouters, Søren Holdt Jensen
Signal Process.4
2013 Nonlinear Acoustic Echo Cancellation Based on a Sliding-Window Leaky Kernel Affine Projection Algorithm
abstract
Acoustic echo cancellation (AEC) is used in speech communication systems where the existence of echoes degrades the speech intelligibility. Standard approaches to AEC rely on the assumption that the echo path to be identified can be modeled by a linear filter. However, some elements introduce nonlinear distortion and must be modeled as nonlinear systems. Several nonlinear models have been used with more or less success. The kernel affine projection algorithm (KAPA) has been successfully applied to many areas in signal processing but not yet to nonlinear AEC (NLAEC). The contribution of this paper is three-fold: (1) to apply KAPA to the NLAEC problem, (2) to develop a sliding-window leaky KAPA (SWL-KAPA) that is well suited for NLAEC applications, and (3) to propose a kernel function, consisting of a weighted sum of a linear and a Gaussian kernel. In our experiment set-up, the proposed SWL-KAPA for NLAEC consistently outperforms the linear APA, resulting in up to 12 dB of improvement in ERLE at a computational cost that is only 4.6 times higher. Moreover, it is shown that the SWL-KAPA outperforms, by 4-6 dB, a Volterra-based NLAEC, which itself has a much higher 413 times computational cost than the linear APA.
Jose Manuel Gil-Cacho, Marco Signoretto, Toon van Waterschoot, Marc Moonen, Søren Holdt Jensen
IEEE Trans. Speech Audio Process.5
2013 Nonlinear Least Squares Methods for Joint DOA and Pitch Estimation
abstract
In this paper, we consider the problem of joint direction-of-arrival (DOA) and fundamental frequency estimation. Joint estimation enables robust estimation of these parameters in multi-source scenarios where separate estimators may fail. First, we derive the exact and asymptotic Cramér-Rao bounds for the joint estimation problem. Then, we propose a nonlinear least squares (NLS) and an approximate NLS (aNLS) estimator for joint DOA and fundamental frequency estimation. The proposed estimators are maximum likelihood estimators when: 1) the noise is white Gaussian, 2) the environment is anechoic, and 3) the source of interest is in the far-field. Otherwise, the methods still approximately yield maximum likelihood estimates. Simulations on synthetic data show that the proposed methods have similar or better performance than state-of-the-art methods for DOA and fundamental frequency estimation. Moreover, simulations on real-life data indicate that the NLS and aNLS methods are applicable even when reverberation is present and the noise is not white Gaussian.
Jesper Rindom Jensen, Mads Græsbøll Christensen, Søren Holdt Jensen
IEEE Trans. Speech Audio Process.3
2013 Default Bayesian Estimation of the Fundamental Frequency
abstract
Joint fundamental frequency and model order estimation is an important problem in several applications. In this paper, a default estimation algorithm based on a minimum of prior information is presented. The algorithm is developed in a Bayesian framework, and it can be applied to both real- and complex-valued discrete-time signals which may have missing samples or may have been sampled at a non-uniform sampling frequency. The observation model and prior distributions corresponding to the prior information are derived in a consistent fashion using maximum entropy and invariance arguments. Moreover, several approximations of the posterior distributions on the fundamental frequency and the model order are derived, and one of the state-of-the-art joint fundamental frequency and model order estimators is demonstrated to be a special case of one of these approximations. The performance of the approximations are evaluated in a small-scale simulation study on both synthetic and real world signals. The simulations indicate that the proposed algorithm yields more accurate results than previous algorithms. The simulation code is available online.
Jesper Kjær Nielsen, Mads Græsbøll Christensen, Søren Holdt Jensen
IEEE Trans. Speech Audio Process.3
2013 Binaural Integrated Active Noise Control and Noise Reduction in Hearing Aids
abstract
This paper presents a binaural approach to integrated active noise control and noise reduction in hearing aids and aims at demonstrating that a binaural setup indeed provides significant advantages in terms of the number of noise sources that can be compensated for and in terms of the causality margins.
Romain Serizel, Marc Moonen, Jan Wouters, Søren Holdt Jensen
IEEE Trans. Speech Audio Process.4
2013 Sequential Error Concealment for Video/Images by Sparse Linear Prediction
abstract
In this paper, we propose a novel sequential error concealment algorithm for video and images based on sparse linear prediction. Block-based coding schemes in packet loss environments are considered. Images are modelled by means of linear prediction, and missing macroblocks are sequentially reconstructed using the available groups of pixels. The optimal predictor coefficients are computed by applying a missing data regression imputation procedure with a sparsity constraint. Moreover, an efficient procedure for the computation of these coefficients based on an exponential approximation is also proposed. Both techniques provide high-quality reconstructions and outperform the state-of-the-art algorithms both in terms of PSNR and MS-SSIM.
Ján Koloda, Jan Østergaard, Søren Holdt Jensen, Victoria E. Sánchez, Antonio M. Peinado
IEEE Trans. Multim.3
2013 Compressive Sensing for Spread Spectrum Receivers
abstract
With the advent of ubiquitous computing there are two design parameters of wireless communication devices that become very important: power efficiency and production cost. Compressive sensing enables the receiver in such devices to sample below the Shannon-Nyquist sampling rate, which may lead to a decrease in the two design parameters. This paper investigates the use of Compressive Sensing (CS) in a general Code Division Multiple Access (CDMA) receiver. We show that when using spread spectrum codes in the signal domain, the CS measurement matrix may be simplified. This measurement scheme, named Compressive Spread Spectrum (CSS), allows for a simple, effective receiver design. Furthermore, we numerically evaluate the proposed receiver in terms of bit error rate under different signal to noise ratio conditions and compare it with other receiver structures. These numerical experiments show that though the bit error rate performance is degraded by the subsampling in the CS-enabled receivers, this may be remedied by including quantization in the receiver model. We also study the computational complexity of the proposed receiver design under different sparsity and measurement ratios. Our work shows that it is possible to subsample a CDMA signal using CSS and that in one example the CSS receiver outperforms the classical receiver.
Karsten Fyhn Nielsen, Tobias Lindstrøm Jensen, Torben Larsen, Søren Holdt Jensen
IEEE Trans. Wirel. Commun.4
2012 Sequential Error Concealment for Video/Images by Weighted Template Matching
abstract
In this paper we propose a novel spatial error concealment algorithm for video and images based on convex optimization. Block-based coding schemes in packet loss environment are considered. Missing macro blocks are sequentially reconstructed by filling them with a weighted set of templates extracted from the available neighbourhood. Moreover, a fast approximation of the optimization method is proposed. The technique produces high quality reconstructions that outperforms the state-of-the-art algorithms both in terms of PSNR and MS-SSIM.
Ján Koloda, Jan Østergaard, Søren Holdt Jensen, Antonio M. Peinado, Victoria E. Sánchez
DCC3
2012 Nonlinear acoustic echo cancellation based on a parallel-cascade kernel affine projection algorithm
abstract
In acoustic echo cancellation (AEC) applications, oftentimes an acoustic path from a loudspeaker to a microphone is estimated by means of a linear adaptive filter. However, loudspeakers introduce nonlinear distortions which may strongly degrade the adaptive filter performance, thus nonlinear filters have to be considered. This paper proposes two adaptive algorithms namely the parallel and cascade sliding-window kernel based affine projection algorithm (PSW-KAPA and CSW-KAPA) to solve the problem of nonlinear AEC (NLAEC) while keeping the computational complexity low. They are based on a leaky KAPA which employs the theory and algorithms of kernel methods. The basic concept is to perform adaptive filtering in a linear space that is nonlinearly related to the original input space. A kernel specifically designed for acoustic applications is proposed, which consists in a weighted sum of the linear and the Gaussian kernels. The motivation is basically to separate the problem into linear and nonlinear subproblems. The weights in the kernel also impose different forgetting mechanisms in the sliding window which in turn translates to a more flexible regularization. Simulation results show that PSW-KAPA and CSW-KAPA consistently outperform the linear NLMS, and generalize well both in high and low linear to nonlinear ratio (LNLR).
Jose Manuel Gil-Cacho, Toon van Waterschoot, Marc Moonen, Søren Holdt Jensen
ICASSP4
2012 On compressed sensing and the estimation of continuous parameters from noisy observations
abstract
Compressed sensing (CS) has in recent years become a very popular way of sampling sparse signals. This sparsity is measured with respect to some known dictionary consisting of a finite number of atoms. Most models for real world signals, however, are parametrised by continuous parameters corresponding to a dictionary with an infinite number of atoms. Examples of such parameters are the temporal and spatial frequency. In this paper, we analyse how CS affects the estimation performance of any unbiased estimator when we assume such infinite dictionaries. We base our analysis on the Cramer-Rao lower bound (CRLB) which is frequently used for benchmarking the estimation accuracy of unbiased estimators. For the popular sensing matrices such as the Gaussian sensing matrix, our analysis shows that compressed sensing on average degrades the estimation accuracy by at least the down-sample factor.
Jesper Kjær Nielsen, Mads Græsbøll Christensen, Søren Holdt Jensen
ICASSP3
2012 An approximate Bayesian fundamental frequency estimator
abstract
Joint fundamental frequency and model order estimation is an important problem in several applications such as speech and music processing. In this paper, we develop an approximate estimation algorithm of these quantities using Bayesian inference. The inference about the fundamental frequency and the model order is based on a probability model which corresponds to a minimum of prior information. From this probability model, we give the exact posterior distributions on the fundamental frequency and the model order, and we also present analytical approximations of these distributions which lower the computational load of the algorithm. By use of simulations on both a synthetic signal and a speech signal, the algorithm is demonstrated to be more accurate than a state-of-the-art maximum likelihood-based method.
Jesper Kjær Nielsen, Mads Græsbøll Christensen, Søren Holdt Jensen
ICASSP3
2012 A combined multi-channel Wiener filter-based noise reduction and dynamic range compression in hearing aids
Kim Ngo, Ann Spriet, Marc Moonen, Jan Wouters, Søren Holdt Jensen
Signal Process.5
2012 On Acoustic Feedback Cancellation Using Probe Noise in Multiple-Microphone and Single-Loudspeaker Systems
abstract
A probe noise signal can be used in an acoustic feedback cancellation system to prevent biased adaptive estimation of acoustic feedback paths. However, practical experiences and simulation results indicate that whenever a low-level and inaudible probe noise signal is used, the convergence rate of the adaptive estimation is significantly decreased when keeping the steady-state error unchanged. The goal of this work is to derive analytic expressions for the system behavior such as convergence rate and steady-state error for a multiple-microphone and single-loudspeaker audio system, where the acoustic feedback cancellation is carried out using a probe noise signal. The derived results show how different system parameters and signal properties affect the cancellation performance, and the results explain theoretically the decreased convergence rate. Understanding this is important for making further improvements in the existing probe noise approach.
Meng Guo 0001, Thomas Bo Elmedyb, Søren Holdt Jensen, Jesper Jensen 0001
IEEE Signal Process. Lett.3
2012 Sparse Linear Prediction and Its Applications to Speech Processing
abstract
The aim of this paper is to provide an overview of Sparse Linear Prediction, a set of speech processing tools created by introducing sparsity constraints into the linear prediction framework. These tools have shown to be effective in several issues related to modeling and coding of speech signals. For speech analysis, we provide predictors that are accurate in modeling the speech production process and overcome problems related to traditional linear prediction. In particular, the predictors obtained offer a more effective decoupling of the vocal tract transfer function and its underlying excitation, making it a very efficient method for the analysis of voiced speech. For speech coding, we provide predictors that shape the residual according to the characteristics of the sparse encoding techniques resulting in more straightforward coding strategies. Furthermore, encouraged by the promising application of compressed sensing in signal compression, we investigate its formulation and application to sparse linear predictive coding. The proposed estimators are all solutions to convex optimization problems, which can be solved efficiently and reliably using, e.g., interior-point methods. Extensive experimental results are provided to support the effectiveness of the proposed methods, showing the improvements over traditional linear prediction in both speech analysis and coding.
Daniele Giacobello, Mads Græsbøll Christensen, Manohar N. Murthi, Søren Holdt Jensen, Marc Moonen
IEEE Trans. Speech Audio Process.4
2012 Novel Acoustic Feedback Cancellation Approaches in Hearing Aid Applications Using Probe Noise and Probe Noise Enhancement
abstract
Adaptive filters are widely used in acoustic feedback cancellation systems and have evolved to be state-of-the-art. One major challenge remaining is that the adaptive filter estimates are biased due to the nonzero correlation between the loudspeaker signals and the signals entering the audio system. In many cases, this bias problem causes the cancellation system to fail. The traditional probe noise approach, where a noise signal is added to the loudspeaker signal can, in theory, prevent the bias. However, in practice, the probe noise level must often be so high that the noise is clearly audible and annoying; this makes the traditional probe noise approach less useful in practical applications. In this work, we explain theoretically the decreased convergence rate when using low-level probe noise in the traditional approach, before we propose and study analytically two new probe noise approaches utilizing a combination of specifically designed probe noise signals and probe noise enhancement. Despite using low-level and inaudible probe noise signals, both approaches significantly improve the convergence behavior of the cancellation system compared to the traditional probe noise approach. This makes the proposed approaches much more attractive in practical applications. We demonstrate this through a simulation experiment with audio signals in a hearing aid acoustic feedback cancellation system, where the convergence rate is improved by as much as a factor of 10.
Meng Guo 0001, Søren Holdt Jensen, Jesper Jensen 0001
IEEE Trans. Speech Audio Process.2
2012 Non-Causal Time-Domain Filters for Single-Channel Noise Reduction
abstract
In many existing time-domain filtering methods for noise reduction in, e.g., speech processing, the filters are causal. Such causal filters can be implemented directly in practice. However, it is possible to improve the performance of such noise reduction filtering methods in terms of both noise suppression and signal distortion by allowing the filters to be non-causal. Non-causal time-domain filters require knowledge of the future, and are therefore not directly implementable. If the observed signal is processed in blocks, however, the non-causal filters are implementable. In this paper, we propose such non-causal time-domain filters for noise reduction in speech applications. We also propose some performance measures that enable us to evaluate the performance of non-causal filters. Moreover, it is shown how some of the filters can be updated recursively. Using the recursive expressions, it is also shown that the output SNRs of the filters always increase as we increase the length of the filter when the desired signal is stationary. From both the theoretical and practical evaluations of the filters, it is clearly shown that the performance of time-domain filtering methods for noise reduction can be improved by introducing non-causality.
Jesper Rindom Jensen, Jacob Benesty, Mads Græsbøll Christensen, Søren Holdt Jensen
IEEE Trans. Speech Audio Process.4
2012 Enhancement of Single-Channel Periodic Signals in the Time-Domain
abstract
Most state-of-the-art filtering methods for speech enhancement require an estimate of the noise statistics, but the noise statistics are difficult to estimate in practice when speech is present. Thus, nonstationary noise will have a detrimental impact on the performance of most speech enhancement filters. The impact of such noise can be reduced by using the signal statistics rather than the noise statistics in the filter design. For example, this is possible by assuming a harmonic model for the desired signal; while this model fits well for voiced speech, it will not be appropriate for unvoiced speech. That is, signal-dependent methods based on the signal statistics will introduce undesired distortion for some parts of speech compared to signal-independent methods based on the noise statistics. Since both the signal-independent and signal-dependent approaches to speech enhancement have advantages, it is relevant to combine them to reduce the impact of their individual disadvantages. In this paper, we give theoretical insights into the relationship between these different approaches, and these reveal a close relationship between the two approaches. This justifies joint use of such filtering methods which can be beneficial from a practical point of view. Our experimental results confirm that both signal-independent and signal-dependent approaches have advantages and that they are closely-related. Moreover, as a part of our experiments, we illustrate the practical usefulness of combining signal-independent and signal-dependent enhancement methods by applying such methods jointly on real-life speech.
Jesper Rindom Jensen, Jacob Benesty, Mads Græsbøll Christensen, Søren Holdt Jensen
IEEE Trans. Speech Audio Process.4
2012 A Joint Approach for Single-Channel Speaker Identification and Speech Separation
abstract
In this paper, we present a novel system for joint speaker identification and speech separation. For speaker identification a single-channel speaker identification algorithm is proposed which provides an estimate of signal-to-signal ratio (SSR) as a by-product. For speech separation, we propose a sinusoidal model-based algorithm. The speech separation algorithm consists of a double-talk/single-talk detector followed by a minimum mean square error estimator of sinusoidal parameters for finding optimal codevectors from pre-trained speaker codebooks. In evaluating the proposed system, we start from a situation where we have prior information of codebook indices, speaker identities and SSR-level, and then, by relaxing these assumptions one by one, we demonstrate the efficiency of the proposed fully blind system. In contrast to previous studies that mostly focus on automatic speech recognition (ASR) accuracy, here, we report the objective and subjective results as well. The results show that the proposed system performs as well as the best of the state-of-the-art in terms of perceived quality while its performance in terms of speaker identification and automatic speech recognition results are generally lower. It outperforms the state-of-the-art in terms of intelligibility showing that the ASR results are not conclusive. The proposed method achieves on average, 52.3% ASR accuracy, 41.2 points in MUSHRA and 85.9% in speech intelligibility.
Pejman Mowlaee, Rahim Saeidi, Mads Græsbøll Christensen, Zheng-Hua Tan, Tomi Kinnunen, Pasi Fränti, Søren Holdt Jensen
IEEE Trans. Speech Audio Process.7
2012 A Zone-of-Quiet Based Approach to Integrated Active Noise Control and Noise Reduction for Speech Enhancement in Hearing Aids
abstract
This paper focuses on speech enhancement in hearing aids and presents an integrated approach to active noise control and noise reduction which is based on an optimization over a zone-of-quiet generated by the active noise control. A basic integrated active noise control and noise reduction scheme has been introduced previously to tackle secondary path effects and effects of noise leakage through an open fitting. This scheme however, only takes the sound pressure at the ear canal microphone into account. For an integrated active noise control and noise reduction scheme to be efficient, it is desired to achieve active noise control at the eardrum which in practice is away from the ear canal microphone. In some cases, it can also be desired to achieve noise control over a zone not limited to a single point. Two different schemes are presented. The first scheme is based on a mean squared error criterion expressed at a remote point (RP) away from the ear canal microphone and the second scheme is based on an average mean squared error criterion over a desired zone-of-quiet. They are both compared experimentally with the original scheme for both active noise control and integrated active noise control and noise reduction, respectively. The remote-point approach then allows to restore the performance of the original scheme at the desired remote point while the zone-of-quiet approach allows to increase performance up to 3 dB on the desired zone-of-quiet.
Romain Serizel, Marc Moonen, Jan Wouters, Søren Holdt Jensen
IEEE Trans. Speech Audio Process.4
2011 Analysis of adaptive feedback and echo cancelation algorithms in a general multiple-microphone and single-loudspeaker system
abstract
In this paper, we analyze a general multiple-microphone and single-loudspeaker system, where an adaptive algorithm is used to cancel acoustic feedback/echo and a beamformer processes the feedback/echo canceled signals. This system can be viewed as part of a typical hearing aid system and/or a traditional acoustic echo cancellation system. We introduce and derive an approximation of a useful frequency domain measure - the power transfer function - and show how to predict the system stability bound, convergence rate and the steady-state behavior across time and frequency. Furthermore, we show how the derived expressions can be used to determine e.g. the step size parameter in the adaptive algorithms to achieve a desired system property e.g. convergence rate at a specific frequency.
Meng Guo 0001, Thomas Bo Elmedyb, Søren Holdt Jensen, Jesper Jensen 0001
ICASSP3
2011 A single snapshot optimal filtering method for fundamental frequency estimation
abstract
Recently, optimal linearly constrained minimum variance (LCMV) filtering methods have been applied for fundamental frequency estimation. Like many other fundamental frequency estimators, these methods utilize the inverse covariance matrix. Therefore, the covariance matrix needs to be invertible which is typically ensured by using the sample covariance matrix involving data partitioning. The partitioning adversely affects the spectral resolution. We propose a novel optimal filtering method which utilizes the LCMV principle in conjunction with the iterative adaptive approach (IAA). The IAA enables us to estimate the covariance matrix from a single snapshot, i.e., without data partitioning. The experimental results show, that the performance of the proposed method is comparable or better than that of other competing methods in terms of spectral resolution.
Jesper Rindom Jensen, Mads Græsbøll Christensen, Søren Holdt Jensen
ICASSP3
2011 A flexible Speech Distortion Weighted Multi-channel Wiener Filter for noise reduction in hearing aids
abstract
In this paper, a multi-channel noise reduction algorithm is presented based on a Speech Distortion Weighted Multi-channel Wiener Filter (SDW-MWF) approach that incorporates a flexible weighting factor. A typical SDW-MWF uses a fixed weighting factor to trade-off between noise reduction and speech distortion without taking speech presence or speech absence into account. Consequently, the improvement in noise reduction comes at the cost of a higher speech distortion since the speech dominant segments and the noise dominant segments are weighted equally. Based on a two-state speech model with a noise-only and a speech+noise state, a solution is introduced that allows for a more flexible trade-off between noise reduction and speech distortion. Experimental results with hearing aid scenarios demonstrate that the proposed SDW-MWF incorporating the flexible weighting factor improves the signal-to-noise-ratio with lower speech distortion compared to a typical SDW-MWF and the SDW-MWF incorporating the conditional speech presence probability (SPP).
Kim Ngo, Marc Moonen, Søren Holdt Jensen, Jan Wouters
ICASSP3
2011 Sinusoidal Approach for the Single-Channel Speech Separation and Recognition Challenge
abstract
Most of the single-channel speech separation (SCSS) systems use the short-time Fourier transform as their parametric features. Recent studies have shown that employing sinusoidal features for the SCSS application results in a high perceived speech quality. In this paper, we make a systematic study on automatic speech recognition results for a SCSS system that uses sinusoidal features composed of amplitude and frequency. We compare the speech recognition results with those already reported by other participants in the single-channel speech separation and recognition challenge. Our results show that a newly proposed system achieves an overall recognition accuracy of 52.3%, ranges at the median over all other participants in the challenge. Index Terms: sinusoidal modeling, single-channel speech separation and recognition challenge.
Pejman Mowlaee, Rahim Saeidi, Zheng-Hua Tan, Mads Græsbøll Christensen, Tomi Kinnunen, Pasi Fränti, Søren Holdt Jensen
INTERSPEECH7
2011 Output SNR analysis of integrated active noise control and noise reduction in hearing aids under a single speech source scenario
Romain Serizel, Marc Moonen, Jan Wouters, Søren Holdt Jensen
Signal Process.4
2011 An iterative subspace-based multi-pitch estimation algorithm
Johan Xi Zhang, Mads Græsbøll Christensen, Søren Holdt Jensen, Marc Moonen
Signal Process.3
2011 New Results on Perceptual Distortion Minimization and Nonlinear Least-Squares Frequency Estimation
abstract
In a paper in this journal, a framework was presented wherein a number of practical methods for finding the perceptually most important sinusoids in audio signals could be related using a particular perceptually motivated distortion measure, and it was argued that for Gaussian noise and a large number of samples, these methods should attain the Cramér-Rao lower bound. In this correspondence, we report some new results on this subject. Specifically, we analyze the finite-sample performance of these methods in experiments, and we conclude that for a high number of samples, they perform close to the Cramér-Rao lower bound. However, for a low number of observations, we demonstrate that special care must be taken in designing the perceptually motivated distortion measure if high-resolution estimates are desired. In particular, the smoothness of the frequency response of the perceptual filter that implements the distortion measure is shown to be important.
Mads Græsbøll Christensen, Søren Holdt Jensen
IEEE Trans. Speech Audio Process.2
2011 New Results on Single-Channel Speech Separation Using Sinusoidal Modeling
abstract
We present new results on single-channel speech separation and suggest a new separation approach to improve the speech quality of separated signals from an observed mixture. The key idea is to derive a mixture estimator based on sinusoidal parameters. The proposed estimator is aimed at finding sinusoidal parameters in the form of codevectors from vector quantization (VQ) codebooks pre-trained for speakers that, when combined, best fit the observed mixed signal. The selected codevectors are then used to reconstruct the recovered signals for the speakers in the mixture. Compared to the log-max mixture estimator used in binary masks and the Wiener filtering approach, it is observed that the proposed method achieves an acceptable perceptual speech quality with less cross-talk at different signal-to-signal ratios. Moreover, the method is independent of pitch estimates and reduces the computational complexity of the separation by replacing the short-time Fourier transform (STFT) feature vectors of high dimensionality with sinusoidal feature vectors. We report separation results for the proposed method and compare them with respect to other benchmark methods. The improvements made by applying the proposed method over other methods are confirmed by employing perceptual evaluation of speech quality (PESQ) as an objective measure and a MUSHRA listening test as a subjective evaluation for both speaker-dependent and gender-dependent scenarios.
Pejman Mowlaee, Mads Græsbøll Christensen, Søren Holdt Jensen
IEEE Trans. Speech Audio Process.3
2011 Bayesian Interpolation and Parameter Estimation in a Dynamic Sinusoidal Model
abstract
In this paper, we propose a method for restoring the missing or corrupted observations of nonstationary sinusoidal signals which are often encountered in music and speech applications. To model nonstationary signals, we use a time-varying sinusoidal model which is obtained by extending the static sinusoidal model into a dynamic sinusoidal model. In this model, the in-phase and quadrature components of the sinusoids are modeled as first-order Gauss-Markov processes. The inference scheme for the model parameters and missing observations is formulated in a Bayesian framework and is based on a Markov chain Monte Carlo method known as Gibbs sampler. We focus on the parameter estimation in the dynamic sinusoidal model since this constitutes the core of model-based interpolation. In the simulations, we first investigate the applicability of the model and then demonstrate the inference scheme by applying it to the restoration of lost audio packets on a packet-based network. The results show that the proposed method is a reasonable inference scheme for estimating unknown signal parameters and interpolating gaps consisting of missing/corrupted signal segments.
Jesper Kjær Nielsen, Mads Græsbøll Christensen, A. Taylan Cemgil, Simon J. Godsill, Søren Holdt Jensen
IEEE Trans. Speech Audio Process.5
2010 Fixed-Lag Smoothing for Low-Delay Predictive Coding with Noise Shaping for Lossy Networks
abstract
We consider linear predictive coding and noise shaping for coding and transmission of auto-regressive (AR) sources over lossy networks. We generalize an existing framework to arbitrary filter orders and propose use of fixed-lag smoothing at the decoder, in order to further reduce the impact of transmission failures. We show that fixed-lag smoothing up to a certain delay can be obtained without additional computational complexity by exploiting the state-space structure. We prove that the proposed smoothing strategy strictly improves performance under quite general conditions. Finally, we provide simulations on AR sources, and channels with correlated losses, and show that substantial improvements are possible.
Thomas Arildsen, Jan Østergaard, Manohar N. Murthi, Søren Vang Andersen, Søren Holdt Jensen
DCC5
2010 Error-correction of binary masks using hidden Markov models
abstract
Binary masking is a simple and efficient method for source separation, and a high increase in intelligibility can be obtained by applying the target binary mask to noisy speech. The target binary mask can only be calculated under ideal conditions and will contain errors when estimated in real-life applications. This paper proposes a method for correcting these errors. The error-correction is based on a hidden Markov model and uses the Viterbi algorithm to calculate the most probable error-free target binary mask from a target binary mask containing errors. The results demonstrate that it is possible to correct errors in the target binary mask and reduce the noise energy. However, speech energy is also reduced by the error-correction, but the impact on speech intelligibility and speech quality are not established or evaluated in the present study.
Jesper Bünsow Boldt, Michael Syskind Pedersen, Ulrik Kjems, Mads Græsbøll Christensen, Søren Holdt Jensen
ICASSP5
2010 Enhancing sparsity in linear prediction of speech by iteratively reweighted 1-norm minimization
abstract
Linear prediction of speech based on 1-norm minimization has already proved to be an interesting alternative to 2-norm minimization. In particular, choosing the 1-norm as a convex relaxation of the 0-norm, the corresponding linear prediction model offers a sparser residual better suited for coding applications. In this paper, we propose a new speech modeling technique based on reweighted 1-norm minimization. The purpose of the reweighted scheme is to overcome the mismatch between 0-norm minimization and 1-norm minimization while keeping the problem solvable with convex estimation tools. Experimental results prove the effectiveness of the reweighted 1-norm minimization, offering better coding properties compared to 1-norm minimization.
Daniele Giacobello, Mads Græsbøll Christensen, Manohar N. Murthi, Søren Holdt Jensen, Marc Moonen
ICASSP4
2010 Estimation of frame independent and enhancement components for speech communication over packet networks
abstract
In this paper, we describe a new approach to cope with packet loss in speech coders. The idea is to split the information present in each speech packet into two components, one to independently decode the given speech frame and one to enhance it by exploiting inter-frame dependencies. The scheme is based on sparse linear prediction and a redefinition of the analysis-by-synthesis process. We present Mean Opinion Scores for the presented coder with different degrees of packet loss and show that it performs similarly to frame dependent coders for low packet loss probability and similarly to frame independent coders for high packet loss probability. We also present ideas on how to make the coder work synergistically with the channel loss estimate.
Daniele Giacobello, Manohar N. Murthi, Mads Græsbøll Christensen, Søren Holdt Jensen, Marc Moonen
ICASSP4
2010 Iterated smoothing for accelerated gradient convex minimization in signal processing
abstract
In this paper, we consider the problem of minimizing a non-smooth convex problem using first-order methods. The number of iterations required to guarantee a certain accuracy for such problems is often excessive and several methods, e.g., restart methods, have been proposed to speed-up the convergence. In the restart method a smoothness parameter is adjusted such that smoother approximations of the original non-smooth problem are solved in a sequence before the original, and the previous estimate is used as the starting point each time. Instead of adjusting the smoothness parameter after each restart, we propose a method where we modify the smoothness parameter in each iteration. We prove convergence and provide simulation examples for two typical signal processing applications, namely total variation denoising and l1-norm minimization. The simulations demonstrate that the proposed method require fewer iterations and show lower complexity compared to the restart method.
Tobias Lindstrøm Jensen, Jan Østergaard, Søren Holdt Jensen
ICASSP3
2010 Improved single-channel speech separation using sinusoidal modeling
abstract
We present a novel single-channel separation approach to improve the separation performance while recovering the signals from a mixture. The key idea in this research is to employ a mixture estimator based on unconstrained modified sinusoidal parameters. Compared to the mixmax (binary mask) and Wiener filter (softmask) approaches, the proposed approach works independently of pitch estimates. Furthermore, it is observed that it can achieve acceptable perceptual speech quality with less cross-talk at different signal-to-signal ratios while bringing down the complexity by replacing STFT with sinusoidal parameters. Improvementsmade by the proposed approach are demonstrated by employing PESQ as our objective measure and MUSHRA listening test as our subjective evaluation.
Pejman Mowlaee, Mads Græsbøll Christensen, Søren Holdt Jensen
ICASSP3
2010 Sinusoidal masks for single channel speech separation
abstract
In this paper we present a new approach for binary and soft masks used in single-channel speech separation. We present a novel approach called the sinusoidal mask (binary mask and Wiener filter) in a sinusoidal space. Theoretical analysis is presented for the proposed method, and we show that the proposed method is able to minimize the target speech distortion while suppressing the crosstalk to a predetermined threshold. It is observed that compared to the STFT-based masks, the proposed sinusoidal masks improve the separation performance in terms of objective measures (SSNR and PESQ) and are mostly preferred by listeners.
Pejman Mowlaee, Mads Græsbøll Christensen, Søren Holdt Jensen
ICASSP3
2010 Joint single-channel speech separation and speaker identification
abstract
In this paper, we propose a closed loop system to improve the performance of single-channel speech separation in a speaker independent scenario. The system is composed of two interconnected blocks: a separation block and a speaker identification block. The improvement is accomplished by incorporating the speaker identities found by the speaker identification block as additional information for the separation block, which converts the speaker-independent separation problem to a speaker-dependent one where the speaker codebooks are known. Simulation results show that the closed loop system enhances the quality of the separated output signals. To assess the improvements, the results are reported in terms of PESQ for both target and masked signals.
Pejman Mowlaee, Rahim Saeidi, Zheng-Hua Tan, Mads Græsbøll Christensen, Pasi Fränti, Søren Holdt Jensen
ICASSP6
2010 Adaptive feedback cancellation in hearing aids using a sinusoidal near-end signal model
abstract
Acoustic feedback is a well-known problem in hearing aids, which is caused by the undesired acoustic coupling between the loudspeaker and the microphone. Acoustic feedback limits the maximum amplification that can be used in the hearing aid without making it unstable. The goal of adaptive feedback cancellation (AFC) is to adaptively model the feedback path and estimate the feedback signal, which is then subtracted from the microphone signal. The main problem in identifying the feedback path model is the correlation between the near-end signal and the loudspeaker signal, which is caused by the closed signal loop. A possible solution to this problem is to use the prediction error method (PEM)-based AFC with a linear prediction (LP) model for the near-end signal. In this paper, a modification to the PEM-based AFC is presented where the LP model is replaced by a sinusoidal near-end signal model. More specifically, it is shown that using frequency estimation techniques to estimate the sinusoidal near-end signal model improves the performance of the PEM-based AFC compared to using a LP model. Simulation results for a hearing aid scenario indicate a significant improvement in terms of misadjustment and maximum stable gain increase.
Kim Ngo, Toon van Waterschoot, Mads Græsbøll Christensen, Marc Moonen, Søren Holdt Jensen, Jan Wouters
ICASSP5
2010 Signal-to-Signal Ratio Independent Speaker Identification for Co-channel Speech Signals
abstract
In this paper, we consider speaker identification for the co-channel scenario in which speech mixture from speakers is recorded by one microphone only. The goal is to identify both of the speakers from their mixed signal. High recognition accuracies have already been reported when an accurately estimated signal-to-signal ratio (SSR) is available. In this paper, we approach the problem without estimating SSR. We show that a simple method based on fusion of adapted Gaussian mixture models and Kullback-Leibler divergence calculated between models, achieves an accuracy of 97% and 93% when the two target speakers enlisted as three and two most probable speakers, respectively.
Rahim Saeidi, Pejman Mowlaee, Tomi Kinnunen, Zheng-Hua Tan, Mads Græsbøll Christensen, Søren Holdt Jensen, Pasi Fränti
ICPR6
2010 Improving monaural speaker identification by double-talk detection
abstract
This paper describes a novel approach to improve monoaural speaker identification where two speakers are present in a single-microphone recording. The goal is to identify both of the underlying speakers in the given mixture. The proposed approach is composed of a double-talk detector (DTD) as a preprocessor and speaker identification back-end. We demonstrate that including the double-talk detector improves the speaker identification accuracy. Experiments on GRID corpus show that including the DTD improves average recognition accuracy from 96.53% to 97.43%.
Rahim Saeidi, Pejman Mowlaee, Tomi Kinnunen, Zheng-Hua Tan, Mads Græsbøll Christensen, Søren Holdt Jensen, Pasi Fränti
INTERSPEECH6
2010 Retrieving Sparse Patterns Using a Compressed Sensing Framework: Applications to Speech Coding Based on Sparse Linear Prediction
abstract
Encouraged by the promising application of compressed sensing in signal compression, we investigate its formulation and application in the context of speech coding based on sparse linear prediction. In particular, a compressed sensing method can be devised to compute a sparse approximation of speech in the residual domain when sparse linear prediction is involved. We compare the method of computing a sparse prediction residual with the optimal technique based on an exhaustive search of the possible nonzero locations and the well known Multi-Pulse Excitation, the first encoding technique to introduce the sparsity concept in speech coding. Experimental results demonstrate the potential of compressed sensing in speech coding techniques, offering high perceptual quality with a very sparse approximated prediction residual.
Daniele Giacobello, Mads Græsbøll Christensen, Manohar N. Murthi, Søren Holdt Jensen, Marc Moonen
IEEE Signal Process. Lett.4
2010 Integrated Active Noise Control and Noise Reduction in Hearing Aids
abstract
This paper presents combined active noise control and noise reduction schemes for hearing aids to tackle secondary path effects and effects of noise leakage through an open fitting. While such leakage contributions and the secondary acoustic path from the loudspeaker to the tympanic membrane are usually not taken into account in standard noise reduction systems, they appear to have a non-negligible impact on the final signal-to-noise ratio. Using a noise-reduction algorithm and an active noise control system in cascade may be efficient as long as the causality margin of the system is large enough. Putting the two functional blocks in parallel and then integrating them is found to lead to a more robust algorithm. A Filtered-x Multichannel Wiener Filter is presented and applied to integrate noise reduction and active noise control. The cascaded scheme and the integrated scheme are compared experimentally with a Multichannel Wiener Filter in a classic noise reduction framework without active noise control, where the integrated scheme is found to provide the best performance.
Romain Serizel, Marc Moonen, Jan Wouters, Søren Holdt Jensen
IEEE Trans. Speech Audio Process.4
2010 A Robust and Computationally Efficient Subspace-Based Fundamental Frequency Estimator
abstract
This paper presents a method for high-resolution fundamental frequency$(F_{0})$estimation based on subspaces decomposed from a frequency-selective data model, by effectively splitting the signal into a number of subbands. The resulting estimator is termed frequency-selective harmonic MUSIC (F-HMUSIC). The subband-based approach is expected to ensure computational savings and robustness. Additionally, a method for automatic subband signal activity detection is proposed, which is based on information-theoretic criterion where no subjective judgment is needed. The F-HMUSIC algorithm exhibits good statistical performance when evaluated with synthetic signals for both white and colored noises, while its evaluation on real-life audio signal shows the algorithm to be competitive with other estimators. Finally, F-HMUSIC is found to be computationally more efficient and robust than other subspace-based$F_{0}$estimators, besides being robust against recorded data with inharmonicities.
Johan Xi Zhang, Mads Græsbøll Christensen, Søren Holdt Jensen, Marc Moonen
IEEE Trans. Speech Audio Process.3
2009 l1 Compression of Image Sequences Using the Structural Similarity Index Measure
abstract
We consider lossy compression of image sequences using l1-compression with overcomplete dictionaries. As a fidelity measure for the reconstruction quality, we incorporate the recently proposed structural similarity index measure, and we show that this leads to problem formulations that are very similar to conventional l1 compression algorithms. In addition, we develop efficient large-scale algorithms used for joint encoding of multiple image frames.
Joachim Dahl, Jan Østergaard, Tobias Lindstrøm Jensen, Søren Holdt Jensen
DCC4
2009 Complex Wavelet Modulation Subbands for Speech Compression
abstract
Low-frequency modulation of sound carry essential information for speech and music. They must be preserved for compression. The complex modulation spectrum has already been used for audio compression and is commonly obtained by spectral analysis of the sole temporal envelopes of the subbands out of a time/frequency analysis (modified discrete cosine transform combined with a modified discrete sine transform). However, amplitudes and tones of speech or music tend to vary slowly over time thus the temporal envelopes are often smooth and mostly of polynomial type. Processing in this domain usually creates undesirable distortions because only the magnitudes are taken into account and the phase data is often neglected. We remedy this problem with the use of a complex wavelet transform as a more appropriate envelope and phase processing tool. Complex wavelets carry both magnitude and phase explicitly with great sparsity and preserve well polynomials. Moreover an analytic Hilbert-like transform is possible with complex wavelets implemented as an orthogonal filter bank.
Jean-Marc Luneau, Jérôme Lebrun, Søren Holdt Jensen
DCC3
2009 An efficient first-order method for l1 compression of images
abstract
We consider the problem of lossy compression of images using sparse representations from overcomplete dictionaries. This problem is in principle easy to solve using standard algorithms for convex programming, but often the large dimensions render such an approach intractable. We present a highly efficient method based on recently developed first-order methods, which enables us to compute sparse approximations of entire images with modest time and memory consumption.
Joachim Dahl, Jan Østergaard, Tobias Lindstrøm Jensen, Søren Holdt Jensen
ICASSP4
2009 Joint estimation of short-term and long-term predictors in speech coders
abstract
In low bit-rate coders, the near-sample and far-sample redundancies of the speech signal are usually removed by a cascade of a short-term and a long-term linear predictor. These two predictors are usually found in a sequential and therefore suboptimal approach. In this paper we propose an analysis model that jointly finds the two predictors by adding a regularization term in the minimization process to impose sparsity constraints on a high order predictor. The result is a linear predictor that can be easily factorized into the short-term and long-term predictors. This estimation method is then incorporated into an algebraic code excited linear prediction scheme and shows to have a better performance than traditional cascade methods and other joint optimization methods, offering lower distortion and higher perceptual speech quality.
Daniele Giacobello, Mads Græsbøll Christensen, Joachim Dahl, Søren Holdt Jensen, Marc Moonen
ICASSP4
2009 Sub-band implementation of the Harmonic MUSIC algorithm
abstract
In this paper, we present a novel method for joint estimation of the order and fundamental frequency of a set of harmonically related sinusoids. This method uses a subband based approach to estimate the involved parameters using subspace techniques, and the resulting algorithm is termed frequency-selective harmonic MUSIC (F-HMUSIC). The performance of F-HMUSIC is evaluated and compared to both harmonic MUSIC (HMUSIC) and Cramer-Rao lower bound (CRLB). Especially, in a low signal-to-noise ratio (SNR) with colored noise scenarios, where F-HMUSIC outperforms HMUSIC. F-HMUSIC is concluded to be more computationally efficient and more robust against colored noise than other subspace based fundamental frequency estimators.
Johan Xi Zhang, Mads Græsbøll Christensen, Joachim Dahl, Søren Holdt Jensen, Marc Moonen
ICASSP4
2009 Robust implementation of the MUSIC algorithm
abstract
The problem of estimating frequencies of sinusoids in noise has been studied intensively by the signal processing community during the last decades. Traditionally high resolution subspace-based techniques suffer from high computational complexity, and generally sensitive to the colored noise. We present here a frequency-domain based subspace parameter estimation algorithm termed frequency-selective MUltiple SIgnal Classification (F-MUSIC) that is based on the signal and noise subspace orthogonality property. The method is computationally efficient in providing estimates in the selected subband compared to the classic MUSIC. The performance of F-MUSIC is evaluated and compared to both MUSIC and Cramer-Rao lower bound (CRLB). In a low signal to noise ratio (SNR) with colored noise scenarios, F-MUSIC outperforms MUSIC.
Johan Xi Zhang, Mads Græsbøll Christensen, Joachim Dahl, Søren Holdt Jensen, Marc Moonen
ICASSP4
2009 A system for detecting miscues in dyslexic read speech
abstract
While miscue detection in general is a well explored research field little attention has so far been paid to miscue detection in dyslexic read speech. This domain differs substantially from the domains that are commonly researched, as for example dyslexic read speech includes frequent regressions and long pauses between words. A system detecting miscues in dyslexic read speech is presented. It includes an ASR component employing a forced-alignment like grammar adjusted for dyslexic input and uses the GOP score and phone duration to accept or reject the read words. Experimental results show that the system detects miscues at a false alarm rate of 5.3% and a miscue detection rate of 40.1%. These results are worse than current state of the art reading tutors perhaps indicating that dyslexic read speech is a challenge to handle.
Morten Højfeldt Rasmussen, Zheng-Hua Tan, Børge Lindberg, Søren Holdt Jensen
INTERSPEECH4
2009 Robust Parametric Audio Coding Using Multiple Description Coding
abstract
We propose a new multiple description spherical quantization with repetitively coded amplitudes (MDSQRA) scheme suited for quantization of sinusoidal parameters. The quantization scheme is constituted by a set of spherical quantizers inspired by the multiple description spherical trellis-coded quantization (MDSTCQ) scheme. In this scheme, we apply repetitive coding on the amplitudes, while multiple description coding are applied on the phases and frequencies. Thereby, MDSQRA becomes directly implementable, as opposed to MDSTCQ, since the phase and frequency quantizers depend on the amplitudes which have dissimilar descriptions in MDSTCQ. Furthermore, we implement MDSQRA into a perceptual matching pursuit based sinusoidal audio coder. Finally, we evaluate MDSQRA through perceptual distortion measurements and MUSHRA listening tests. The tests show that MDSQRA outperforms MDSTCQ with respect to a expected perceptual distortion measure. The same results are obtained through the MUSHRA tests performed on sound clips coded using MDSQRA and MDSTCQ.
Jesper Rindom Jensen, Mads Græsbøll Christensen, Morten Holm Jensens, Søren Holdt Jensen, Torben Larsen
IEEE Signal Process. Lett.4
2009 Quantitative Analysis of a Common Audio Similarity Measure
abstract
For music information retrieval tasks, a nearest neighbor classifier using the Kullback-Leibler divergence between Gaussian mixture models of songs' melfrequency cepstral coefficients is commonly used to match songs by timbre. In this paper, we analyze this distance measure analytically and experimentally by the use of synthesized MIDI files, and we find that it is highly sensitive to different instrument realizations. Despite the lack of theoretical foundation, it handles the multipitch case quite well when all pitches originate from the same instrument, but it has some weaknesses when different instruments play simultaneously. As a proof of concept, we demonstrate that a source separation frontend can improve performance. Furthermore, we have evaluated the robustness to changes in key, sample rate, and bitrate.
Jesper Højvang Jensen, Mads Græsbøll Christensen, Daniel P. W. Ellis, Søren Holdt Jensen
IEEE Trans. Speech Audio Process.4
2008 A tempo-insensitive distance measure for cover song identification based on chroma features
abstract
We present a distance measure between audio files designed to identify cover songs, which are new renditions of previously recorded songs. For each song we compute the chromagram, remove phase information and apply exponentially distributed bands in order to obtain a feature matrix that compactly describes a song and is insensitive to changes in instrumentation, tempo and time shifts. As distance between two songs, we use the Frobenius norm of the difference between their feature matrices normalized to unit norm. When computing the distance, we take possible transpositions into account. In a test collection of 80 songs with two versions of each, 38% of the covers were identified. The system was also evaluated on an independent, international evaluation where it despite having much lower complexity performed on par with the winner of last year.
Jesper Højvang Jensen, Mads Græsbøll Christensen, Daniel P. W. Ellis, Søren Holdt Jensen
ICASSP4
2008 Multiple description quantization of sinusoidal parameters
abstract
A new scheme for sinusoidal audio coding named multiple description spherical trellis-coded quantization is proposed and analytic expressions for the point densities and expected distortion of the quantizers are derived based on a high-resolution assumption. The proposed quantizers are of variable dimension, i.e., sinusoids can be quantized jointly for each audio segment whereby a lower distortion is achieved. The quantizers are designed to minimize a perceptual distortion measure subject to an entropy constraint for a given packet-loss probability. In experiments, the performance of the quantizers is compared to the corresponding single description spherical quantizer and associated bounds are found to increase robustness towards packet-losses.
Morten Holm Larsen, Mads Græsbøll Christensen, Søren Holdt Jensen
ICASSP3
2008 Sparse linear predictors for speech processing
abstract
This paper presents two new classes of linear prediction schemes. The first one is based on the concept of creating a sparse residual rather than a minimum variance one, which will allow a more efficient quantization; we will show that this works well in presence of voiced speech, where the excitation can be represented by an impulse train, and creates a sparser residual in the case of unvoiced speech. The second class aims at finding sparse prediction coefficients; interesting results can be seen applying it to the joint estimation of long-term and short-term predictors. The proposed estimators are all solutions to convex optimization problems, which can be solved efficiently and reliably using, e.g., interior-point methods. Index Terms: linear prediction, all-pole modeling, convex optimization 1.
Daniele Giacobello, Mads Græsbøll Christensen, Joachim Dahl, Søren Holdt Jensen, Marc Moonen
INTERSPEECH4
2008 Frequency-domain parameter estimations for binary masked signals
abstract
We present an approach for the extraction of parameters of a damped complex exponential model from a spectrogram modified by a binary mask.The parameters are estimated by a frequency domain based methods using subspace techniques, where the core algorithm is F-ESPRIT.The sub-band defined by the binary mask provides a reduced number of DFT-samples for the parameter extractions, which results in a computational efficient scheme with high parameter estimation accuracy.The proposed synthesis system has synthesis performance comparable to the so-called LSEE-MSTFT.The estimated parameters can be used in many applications such as audio/speech coding, pitch estimation and pitch scale modification.
Johan Xi Zhang, Mads Græsbøll Christensen, Joachim Dahl, Søren Holdt Jensen, Marc Moonen
INTERSPEECH4
2008 Multi-pitch estimation
Mads Græsbøll Christensen, Petre Stoica, Andreas Jakobsson, Søren Holdt Jensen
Signal Process.4
2008 On Optimal Filter Designs for Fundamental Frequency Estimation
abstract
Recently, we proposed using Capon's minimum variance principle to find the fundamental frequency of a periodic waveform. The resulting estimator is formed such that it maximizes the output power of a bank of filters. We present an alternative optimal single filter design and then proceed to quantify the similarities and differences between the estimators using asymptotic analysis and Monte Carlo simulations. Our analysis shows that the single filter can be expressed in terms of the optimal filterbank and that the methods are asymptotically equivalent but generally different for finite length signals.
Mads Græsbøll Christensen, Jesper Højvang Jensen, Andreas Jakobsson, Søren Holdt Jensen
IEEE Signal Process. Lett.4
2008 Variable Dimension Trellis-Coded Quantization of Sinusoidal Parameters
abstract
In this letter, we propose joint quantization of the parameters of a set of sinusoids based on the theory of trellis-coded quantization. A particular advantage of this approach is that it allows for joint quantization of a variable number of sinusoids, which is particularly relevant in variable rate parametric audio coding. Under high-resolution assumptions and based on a perceptually relevant distortion measure, we derive analytical expressions for the optimal design subject to an entropy constraint. Numerical experiments show a significant performance gain compared to optimal spherical quantization at the cost of a slight increase in computational complexity.
Morten Holm Larsen, Mads Græsbøll Christensen, Søren Holdt Jensen
IEEE Signal Process. Lett.3
2007 The Multi-Pitch Estimation Problem: some New Solutions
abstract
In this paper, we formulate the multi-pitch estimation problem and propose a number of methods to estimate the set of fundamental frequencies. The methods, which are based on nonlinear least-squares, multiple signal classification (MUSIC) and the Capon principles, have in common the fact that the multiple fundamental frequencies are estimated by means of a one-dimensional search. The statistical properties of the methods are evaluated via Monte Carlo simulations.
Mads Græsbøll Christensen, Petre Stoica, Andreas Jakobsson, Søren Holdt Jensen
ICASSP (3)4
2007 A Convex Programming Approach to Anisotropic Smoothing
abstract
We develop a computationally efficient algorithm for the problem of anisotropic smoothing in images, i.e., smoothing an image along its directional fields. Whereas similar problems are often posed and solved as PDE problems in the literature, we formulate the problem as a convex programming problem. We present efficient implementations of interior-point algorithms for solving large instances of this problem.
Joachim Dahl, Søren Holdt Jensen, Per Christian Hansen
ICIP (2)2
2007 Joint High-Resolution Fundamental Frequency and Order Estimation
abstract
In this paper, we present a novel method for joint estimation of the fundamental frequency and order of a set of harmonically related sinusoids based on the multiple signal classification (MUSIC) estimation criterion. The presented method, termed HMUSIC, is shown to have an efficient implementation using fast Fourier transforms (FFTs). Furthermore, refined estimates can be obtained using a gradient-based method. Illustrative examples of the application of the algorithm to real-life speech and audio signals are given, and the statistical performance of the estimator is evaluated using synthetic signals, demonstrating its good statistical properties.
Mads Græsbøll Christensen, Andreas Jakobsson, Søren Holdt Jensen
IEEE Trans. Speech Audio Process.3
2006 Computationally Efficient Amplitude Modulated Sinusoidal Audio Coding Using Frequency-Domain Linear Prediction
abstract
A method for amplitude modulated sinusoidal audio coding is presented that has low complexity and low delay. This is based on a sub-band processing system, where, in each subband, the signal is modeled as an amplitude modulated sum of sinusoids. The envelopes are estimated using frequency-domain linear prediction and the prediction coefficients are quantized. As a proof of concept, we evaluate different configurations in a subjective listening test, and this shows that the proposed method offers significant improvements in sinusoidal coding. Furthermore, the properties of the frequency-domain linear prediction-based envelope estimator are analyzed
Mads Græsbøll Christensen, Søren Holdt Jensen
ICASSP (5)2
2006 Packet Loss Concealment with Natural Variations using HMM
abstract
Packet loss concealment (PLC) at a receiver has a substantial effect on the speech quality in voice over IP. Most conventional PLC systems have largely relied upon variations of signal repetition and overlap-add interpolation which can produce speech signals that do not follow the larger overall statistical trends. in this paper, we demonstrate how hidden Markov models can be utilized to effect PLC based on statistical signal processing. In particular, we show how HMM-based PLC yields conditional density functions that can be utilized by various statistical estimation methods that produce signal parameter estimates that produce more natural variation than conventional PLC methods, thereby providing much better speech quality
Manohar N. Murthi, Christoffer Rødbro, Søren Vang Andersen, Søren Holdt Jensen
ICASSP (1)4
2006 Amplitude modulated sinusoidal signal decomposition for audio coding
abstract
In this letter, we present a decomposition for sinusoidal coding of audio, based on an amplitude modulation of sinusoids via a linear combination of arbitrary basis vectors. The proposed method, which incorporates a perceptual distortion measure, is based on a relaxation of a nonlinear least-squares minimization. Rate-distortion curves and listening tests show that, compared to a constant-amplitude sinusoidal coder, the proposed decomposition offers perceptually significant improvements in critical transient signals
Mads Græsbøll Christensen, Andreas Jakobsson, Søren Vang Andersen, Søren Holdt Jensen
IEEE Signal Process. Lett.4
2006 On perceptual distortion minimization and nonlinear least-squares frequency estimation
abstract
In this paper, we present a framework for perceptual error minimization and sinusoidal frequency estimation based on a new perceptual distortion measure, and we state its optimal solution. Using this framework, we relate a number of well-known practical methods for perceptual sinusoidal parameter estimation such as the prefiltering method, the weighted matching pursuit, and the perceptual matching pursuit. In particular, we derive and compare the sinusoidal estimation criteria used in these methods. We show that for the sinusoidal estimation problem, the prefiltering method and the weighted matching pursuit are equivalent to the perceptual matching pursuit under certain conditions.
Mads Græsbøll Christensen, Søren Holdt Jensen
IEEE Trans. Speech Audio Process.2
2006 Hidden Markov model-based packet loss concealment for voice over IP
abstract
As voice over IP proliferates, packet loss concealment (PLC) at the receiver has emerged as an important factor in determining voice quality of service. Through the use of heuristic variations of signal and parameter repetition and overlap-add interpolation to handle packet loss, conventional PLC systems largely ignore the dynamics of the statistical evolution of the speech signal, possibly leading to perceptually annoying artifacts. To address this problem, we propose the use of hidden Markov models for PLC. With a hidden Markov model (HMM) tracking the evolution of speech signal parameters, we demonstrate how PLC is performed within a statistical signal processing framework. Moreover, we show how the HMM is used to index a specially designed PLC module for the particular signal context, leading to signal-contingent PLC. Simulation examples, objective tests, and subjective listening tests are provided showing the ability of an HMM-based PLC built with a sinusoidal analysis/synthesis model to provide better loss concealment than a conventional PLC based on the same sinusoidal model for all types of speech signals, including onsets and signal transitions
Christoffer Rødbro, Manohar N. Murthi, Søren Vang Andersen, Søren Holdt Jensen
IEEE Trans. Speech Audio Process.4
2005 Linear AM decomposition for sinusoidal audio coding
abstract
We present a novel decomposition for sinusoidal audio coding using amplitude modulation of sinusoids via a linear combination of arbitrary basis vectors. The proposed method, which incorporates a perceptual distortion measure, is based on a relaxation of a non-linear least squares minimization. It offers benefits in the modeling of transients in audio signals. We compare the decomposition to constant-amplitude sinusoidal coding using rate-distortion curves and listening tests. Both indicate that, at the same bit-rate, perceptually significant improvements can be achieved using the proposed decomposition.
Mads Græsbøll Christensen, Andreas Jakobsson, Søren Vang Andersen, Søren Holdt Jensen
ICASSP (3)4
2005 Open loop rate-distortion optimized audio coding
abstract
The paper addresses complexity reduced rate-distortion optimized audio coding under rate constraint. A technique where distortion minimizing coding templates, chosen from a set of templates, are jointly selected for a set of segments. This optimization requires knowledge of rate-distortion pairs for all segments, and for each coding template, which is often costly to obtain. The proposed framework exchanges true rate-distortion pairs with predicted ones, thereby allowing for complexity reduction. The prediction is based on a property vector extracted for each segment, from which distortion predictions, using Gaussian mixture models, are performed. Here, we evaluate the proposed framework in a sinusoidal coding context. The results show that the proposed framework can increase the distortion performance, compared to a fixed sinusoidal coding scheme.
Fredrik Nordén, Mads Græsbøll Christensen, Søren Holdt Jensen
ICASSP (3)3
2004 Multiband amplitude modulated sinusoidal audio modeling
abstract
In this paper, we investigate the importance of taking frequency-dependent temporal phenomena into account in audio coding. We do this in the context of sinusoidal modeling of audio signals by applying amplitude modulation to the sinusoidal components. Traditionally, audio coders use a fixed time-segmentation for all frequencies despite the fact that it is well-known that the time-frequency resolution of the human auditory system is not constant. The well-known window switching is an example of this. We compare multiband amplitude modulated sinusoidal models to a singleband model using different audio excerpts. Based on both comparative listening tests and a psychoacoustical distortion measure it is concluded that an improvement is generally gained using multiband amplitude modulation, although specific single sources are well-modeled using a singleband model.
Mads Græsbøll Christensen, Steven van de Par, Søren Holdt Jensen, Søren Vang Andersen
ICASSP (4)3
2004 A perceptual subspace approach for modeling of speech and audio signals with damped sinusoids
abstract
The problem of modeling a signal segment as a sum of exponentially damped sinusoidal components arises in many different application areas, including speech and audio processing. Often, model parameters are estimated using subspace based techniques which arrange the input signal in a structured matrix and exploit the so-called shift-invariance property related to certain vector spaces of the input matrix. A problem with this class of estimation algorithms, when used for speech and audio processing, is that the perceptual importance of the sinusoidal components is not taken into account. In this work we propose a solution to this problem. In particular, we show how to combine well-known subspace based estimation techniques with a recently developed perceptual distortion measure, in order to obtain an algorithm for extracting perceptually relevant model components. In analysis-synthesis experiments with wideband audio signals, objective and subjective evaluations show that the proposed algorithm improves perceived signal quality considerable over traditional subspace based analysis methods.
Jesper Jensen 0001, Richard Heusdens, Søren Holdt Jensen
IEEE Trans. Speech Audio Process.3
2003 A perceptual subspace method for sinusoidal speech and audio modeling
abstract
The problem of modeling a signal segment as a sum of exponentially damped sinusoidal components is of interest in a wide range of fields, including speech and audio processing. Often, model parameters are estimated using subspace based techniques that exploit the so-called shift-invariance property. A drawback of these estimation techniques in relation to speech and audio processing is that the perceptual relevance of the model components is not taken into account. In this paper we show how to combine well-known subspace based estimation techniques with a recently developed perceptual distortion measure, to obtain an algorithm for extracting perceptually relevant model components. In analysis-synthesis experiments with wideband audio signals, objective and subjective evaluations show that the proposed algorithm improves perceived signal quality considerably over traditional subspace based analysis methods.
Jesper Jensen 0001, Richard Heusdens, Søren Holdt Jensen
ICASSP (5)3
2003 Compressed domain packet loss concealment of sinusoidally coded speech
abstract
We consider the problem of packet loss concealment for voice over IP (VoIP). The speech signal is compressed at the transmitter using a sinusoidal coding scheme working at 8 kbit/s. At the receiver, packet loss concealment is carried out working directly on the quantized sinusoidal parameters, based on time-scaling of the packets surrounding the missing ones. Subjective listening tests show promising results indicating the potential of sinusoidal speech coding for VoIP.
Christoffer Rødbro, Mads Græsbøll Christensen, Søren Vang Andersen, Søren Holdt Jensen
ICASSP (1)4
2003 Optimization of signature sequences and receiver filters for downlink DS-CDMA systems
abstract
Joint optimization of both signature sequences and linear receiver filters for a downlink DS-CDMA system with multipath propagation channels is considered. The signature sequences and the receiver filters are optimized in order to minimize the sum of the mean-squared-errors at the output of the receivers while maintaining a fixed total transmitting power. The signature sequences and receiver filters are derived using a filtering approach to accommodate for the multipath propagation effects.
Joachim Dahl, Bernard H. Fleury, Søren Holdt Jensen, Jakob Stoustrup
ICC3
2003 Amplitude Modulated Sinusoidal Models for Audio Modeling and Coding
Mads Græsbøll Christensen, Søren Vang Andersen, Søren Holdt Jensen
KES3
2002 Prediction based multi-antenna single-receiver mobile terminal
abstract
We study a multi-antenna single-receiver mobile terminal for interference reduction in a TDMA system. The antenna system consists of two antenna elements where the phase and amplitude of one antenna signal is altered prior to RF combining. The system therefore requires only a single receiver chain. We propose a new basis expansion method for predicting the combiner weight and we show that the proposed algorithm has improved performance compared to a recently proposed method. The evaluation of the algorithm by simulation is based on measurements at 1.8 GHz by a two-antenna mobile terminal carried by different users.
Morten Jeppesen, Søren Holdt Jensen, Gert Frølund Pedersen
PIMRC2
2001 Recursively updated eigenfilterbank for speech enhancement
abstract
A novel signal subspace method for speech enhancement is proposed. The algorithm is derived from the filterbank interpretation of the truncated (quotient) singular value decomposition (T(Q)SVD) algorithm. We derive a recursive version of this algorithm which results in a recursively updated eigenfilterbank. The proposed method benefits from a low system delay and a low amount of musical noise in the enhanced speech signal.
Morten Jeppesen, Christoffer Rødbro, Søren Holdt Jensen
ICASSP3
2000 Harmonic exponential modeling of transitional speech segments
abstract
An extended sinusoidal model for speech signal processing is proposed. In this model, the frequencies of the sinusoidal components are constrained to be harmonically related, but their amplitudes are allowed to vary exponentially with time. Algorithms for model parameter estimation are derived. The proposed model shows considerable objective and subjective improvements in speech segments with transitions, especially in speech onsets, compared to the constant-frequency, constant-amplitude harmonic model. Potential application areas include speech synthesis and multimode speech coding.
Jesper Jensen 0001, Søren Holdt Jensen, Egon Hansen
ICASSP2
1999 Exponential sinusoidal modeling of transitional speech segments
abstract
A generalized sinusoidal model for speech signal processing is studied. The main feature of the model is that the amplitude of each sinusoidal component is allowed to vary exponentially with time. We propose to use the model in transitional speech segments such as speech onsets and voiced/unvoiced transitions. Computer simulations with natural speech signals indicate substantial better modeling performance in both transitional and voiced regions compared with the traditional constant-amplitude sinusoidal model.
Jesper Jensen 0001, Søren Holdt Jensen, Egon Hansen
ICASSP2
1995 Reduction of broad-band noise in speech by truncated QSVD
abstract
We consider an algorithm for reduction of broadband noise in speech based on signal subspaces. The algorithm is formulated by means of the quotient singular value decomposition (QSVD). With this formulation, a prewhitening operation becomes an integral part of the algorithm. We demonstrate that this is essential in connection with updating issues in real-time recursive applications. We also illustrate by examples that we are able to achieve a satisfactory quality of the reconstructed signal.
Søren Holdt Jensen, Per Christian Hansen, Steffen Duus Hansen, John Aasted Sørensen
IEEE Trans. Speech Audio Process.1