EDBT 2026 Demo / reviewers in the wild / expert
Mads Græsbøll Christensen
dblp:69/5399
· DBLP profile ↗
171ranked-venue papers
19as first author
46since 2021 · last 2025
0000-0003-3586-7969ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 131 · 13 first-author · 35 since 2021Artificial intelligence and machine learning · 52 · 6 first-author · 12 since 2021Security and privacy · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Study of the Scale Invariant Signal to Distortion Ratio in Speech Separation with Noisy References*abstractThis paper examines the implications of using the Scale-Invariant Signal-to-Distortion Ratio (SI-SDR) as both evaluation and training objective in supervised speech separation, when the training references contain noise, as is the case with the de facto benchmark WSJ0-2Mix. A derivation of the SI-SDR with noisy references reveals that noise limits the achievable SI-SDR, or leads to undesired noise in the separated outputs. To address this, a method is proposed to enhance references and augment the mixtures with WHAM!, aiming to train models that avoid learning noisy references. Two models trained on these enhanced datasets are evaluated with the non-intrusive NISQA.v2 metric. Results show reduced noise in separated speech but suggest that processing references may introduce artefacts, limiting overall quality gains. Negative correlation is found between SI-SDR and perceived noisiness across models on the WSJ0-2Mix and Libri2Mix test sets, underlining the conclusion from the derivation. Simon Dahl Jepsen, Mads Græsbøll Christensen, Jesper Rindom Jensen |
ASRU | 2 |
| 2025 | Robust Fixed-Filter Sound Zone Control with Audio-Based Position TrackingabstractPerformance of sound zone control (SZC) systems deployed in practical scenarios are highly sensitive to the location of the listener(s) and can degrade significantly when listener(s) are moving. This paper presents a robust SZC system that adapts to dynamic changes such as moving listeners and varying zone locations using a dictionary-based approach. The proposed system continuously monitors the environment and updates the fixed control filters by tracking the listener position using audio signals only. To test the effectiveness of the proposed SZC method, simulation studies are carried out using practically measured impulse responses. These studies show that SZC, when incorporated with the proposed audio-only position tracking scheme, achieves optimal performance when all listener positions are available in the dictionary. Moreover, even when not all listener positions are included in the dictionary, the method still provides good performance improvement compared to a traditional fixed filter SZC scheme. Sankha Subhra Bhattacharjee, Andreas Jonas Fuglsig, Flemming Christensen, Jesper Rindom Jensen, Mads Græsbøll Christensen |
ICASSP | 5 |
| 2025 | Sound Zone Control Robust To Sound Speed ChangeabstractSound zone control (SZC) implemented using static optimal filters is significantly affected by various perturbations in the acoustic environment, an important one being the fluctuation in the speed of sound, which is in turn influenced by changes in temperature and humidity (TH). This issue arises because control algorithms typically use pre-recorded, static impulse responses (IRs) to design the optimal control filters. The IRs, however, may change with time due to TH changes, which renders the derived control filters to become non-optimal. To address this challenge, we propose a straightforward model called sinc interpolation-compression/expansion-resampling (SICER), which adjusts the IRs to account for both sound speed reduction and increase. Using the proposed technique, IRs measured at a certain TH can be corrected for any TH change and control filters can be re-derived without the need of re-measuring the new IRs (which is impractical when SZC is deployed). We integrate the proposed SICER IR correction method with the recently introduced variable span trade-off (VAST) framework for SZC, and propose a SICER-corrected VAST method that is resilient to sound speed variations. Simulation studies show that the proposed SICER-corrected VAST approach significantly improves acoustic contrast and reduces signal distortion in the presence of sound speed changes. Sankha Subhra Bhattacharjee, Jesper Rindom Jensen, Mads Græsbøll Christensen |
ICASSP | 3 |
| 2025 | Robust Exponential Hyperbolic Tangent Geman-McClure Based Identification of Nonlinear SystemsabstractThis manuscript presents a new technique to improve the performance of adaptive filters in handling the non-Gaussian or impulsive noise environment. The conduct of the adaptive filter decays in the presence of impulsive noise or outliers. To improve the efficiency of the filtering technique, this work presents a robust exponential hyperbolic tangent Geman McClure function for nonlinear system identification. The algorithm utilizes the saturation properties of the hyperbolic tangent function to improve the performance under impulsive noise. The simulation results clearly illustrate the potency of the suggested technique. Neetu Chikyal, Vasundhara, Chayan Bhar, Asutosh Kar, Mads Græsbøll Christensen |
ICASSP | 5 |
| 2025 | Next-Generation ANC: Integrating Dynamic Fixed-Filter Strategies With Extended Kalman Filtering for Enhanced Noise SuppressionabstractThe hybrid selective fixed-filter active noise control with filtered reference normalized least mean square (SFANC-FxNLMS) method struggles in dynamic noise environments due to its reliance on static filters, which limits effectiveness when noise characteristics change rapidly. The generative fixed-filter active noise control with Kalman filtering (GFANC-Kalman) approach offers improved adaptability by dynamically adjusting the filtering process but may still underperform in complex noise scenarios. The dynamic fixed-filter active noise control with extended Kalman filter (DFANC-EKF) method overcomes these limitations by integrating an extended Kalman filter with a 2D convolutional neural network for advanced feature extraction. This integration enables the system to better capture and adapt to intricate noise patterns, significantly enhancing noise reduction. Numerical simulations using real-world noise data validate the DFANC-EKF approach's superior performance across various challenging scenarios. Fareedha, Vasundhara, Asutosh Kar, Mads Græsbøll Christensen |
ICASSP | 4 |
| 2025 | Advances in Microphone Array Processing and Multichannel Speech EnhancementabstractThis paper reviews pioneering works in microphone array processing and multichannel speech enhancement, highlighting historical achievements, technological evolution, commercialization aspects, and key challenges. It provides valuable insights into the progression and future direction of these areas. The paper examines foundational developments in microphone array design and optimization, showcasing innovations that improved sound acquisition and enhanced speech intelligibility in noisy and reverberant environments. It then introduces recent advancements and cutting-edge research in the field, particularly the integration of deep learning techniques such as all-neural beamformers. The paper also explores critical applications, discussing their evolution and current state-of-the-art technologies that significantly impact user experience. Finally, the paper outlines future research directions, identifying challenges and potential solutions that could drive further innovation in these fields. By providing a comprehensive overview and forward-looking perspective, this paper aims to inspire ongoing research and contribute to the sustained growth and development of microphone arrays and multichannel speech enhancement. Gongping Huang, Jesper Rindom Jensen, Jingdong Chen, Jacob Benesty, Mads Græsbøll Christensen, Akihiko Sugiyama, Gary W. Elko, Tomas Gänsler |
ICASSP | 5 |
| 2025 | A Modified Gain Normalized Step Size Adaptive Algorithm for Improved Online Secondary Path Modelling in Active Noise ControlabstractThe noise cancellation performance of an active control system decreases when there are temporal variations in the primary and secondary paths. An active noise control (ANC) framework has been introduced in this work, which incorporates four adaptive filters and two decorrelation filters for online secondary path modelling. A novel adaptive algorithm for an active noise control filter has been developed with the combination of modified gain filtered-x recursive least square and normalised step size filtered-x least mean square. The aim is to improve the reduction of mean noise and decrease residual noise while maintaining consistent convergence rate. To update the decorrelation filters in the framework, an adaptive variable step size modified decorrelation normalised least mean square algorithm has been used. These filters are designed to maximize the efficiency of secondary path modelling. Compared to its counterparts, the simulation results illustrate the enhancements of the proposed framework without a substantial increase in overall computational complexity. Asutosh Kar, Pradeep K. Shill, Somanath Pradhan, Vasundhara, Mads Græsbøll Christensen |
ICASSP | 6 |
| 2025 | Fractional-Order Hyperbolic Tangent Based Adaptive Algorithm for Feedback Control in Hearing AidsabstractA new method is suggested to improve the effectiveness of adaptive filters in dealing with unexpected disturbances at the error sensor. This method focuses specifically on the difficult scenario of α-stable noise in feedback cancellation for hearing aids. α-stable noise, which is distinguished by its strong tails and abrupt behavior, poses considerable difficulties for conventional adaptive algorithms. In order to tackle this issue, we propose the implementation of a fractional order hyperbolic tangent (FOHT) algorithm. Fractional order systems, which utilize non-integer order derivatives to represent intricate dynamics, provide improved adaptability and precision, especially in settings where non-Gaussian noise, such as α-stable distributions, is prevalent. The algorithm utilizes the distinct characteristics of fractional calculus to enhance the resilience and adaptability of the system, thereby reducing the influence of α-stable noise. The results of extensive simulations indicate that the FOHT algorithm outperforms existing techniques in terms of steady-state convergence and robustness. Vanitha Devi R, Vasundhara, Asutosh Kar, Mads Græsbøll Christensen |
ICASSP | 4 |
| 2025 | Efficient time-domain speech separation using short encoded sequence network
Debang Liu, Mads Græsbøll Christensen, Baoze Ma |
Speech Commun. | 3 |
| 2025 | Speech Conv-Mamba: Selective Structured State Space Model With Temporal Dilated Convolution for Efficient Speech SeparationabstractAs a selective state space model, Mamba exhibits outstanding performance and efficiency in sequence modeling tasks. Therefore, in this paper, we use Mamba as the fundamental network component to construct a novel speech separation model, Speech Conv-Mamba. Specifically, this model embeds Mamba within a U-shaped convolutional network to build the encoder and decoder network for high-dimensional representation and waveform reconstruction of speech signals. Additionally, we stack multiple temporal dilated convolutions and Mamba to create the separation network for separation task. Our comparative experiments on the GRID2Mix and Libri2Mix datasets demonstrate that the proposed model Speech Conv-Mamba, which achieves 98% and 89% of SepFormer's separation accuracy on two datasets using only 9% (2.4 M) of its model size, provides much less computational complexity and training cost. Debang Liu, Ying Wei 0010, Mads Græsbøll Christensen |
IEEE Signal Process. Lett. | 5 |
| 2025 | Provable Privacy Advantages of Decentralized Federated Learning via Distributed OptimizationabstractFederated learning (FL) emerged as a paradigm designed to improve data privacy by enabling data to reside at its source, thus embedding privacy as a core consideration in FL architectures, whether centralized or decentralized. Contrasting with recent findings by Pasquini et al., which suggest that decentralized FL does not empirically offer any additional privacy or security benefits over centralized models, our study provides compelling evidence to the contrary. We demonstrate that decentralized FL, when deploying distributed optimization, provides enhanced privacy protection - both theoretically and empirically - compared to centralized approaches. The challenge of quantifying privacy loss through iterative processes has traditionally constrained the theoretical exploration of FL protocols. We overcome this by conducting a pioneering in-depth information-theoretical privacy analysis for both frameworks. Our analysis, considering both eavesdropping and passive adversary models, successfully establishes bounds on privacy leakage. In particular, we show information theoretically that the privacy loss in decentralized FL is upper bounded by the loss in centralized FL. Compared to the centralized case where local gradients of individual participants are directly revealed, a key distinction of optimization-based decentralized FL is that the relevant information includes differences of local gradients over successive iterations and the aggregated sum of different nodes’ gradients over the network. This information complicates the adversary’s attempt to infer private data. To bridge our theoretical insights with practical applications, we present detailed case studies involving logistic regression and deep neural networks. These examples demonstrate that while privacy leakage remains comparable in simpler models, complex models like deep neural networks exhibit lower privacy risks under decentralized FL. Extensive numerical tests further validate that decentralized FL is more resistant to privacy attacks, aligning with our theoretical findings. Wenrui Yu, Qiongxiu Li, Milan Lopuhaä-Zwakenberg, Mads Græsbøll Christensen, Richard Heusdens |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Broadband Personal Sound Zone Control in the Presence of NonlinearitiesabstractExisting literature on sound zone control generally consider the signal model to be linear. However, this is seldom true in practice owing to nonlinear distortions arising from the loudspeakers, especially in consumer applications. In this paper, we propose a new signal model for personal sound zone control that takes into consideration any nonlinear behaviour that may arise from the loudspeakers. Following the proposed signal model, the optimization problem is formulated such that it inherently ensures the reduction of nonlinear distortion effects in both the bright and dark zones. In addition, a broadband nonlinear acoustic contrast control - pressure matching approach is proposed for the new signal model. Simulation results on practical data show that our proposed approach can provide improvement in the acoustic contrast and/or the signal distortion performance compared to the traditional linear solution, in the presence of nonlinear distortions. Moreover, important observations are made for the study of nonlinear effects on sound zone control. Sankha Subhra Bhattacharjee, Srikanth Burra, Jesper Rindom Jensen, Liming Shi, Guoli Ping, Jingkai Weng, Mads Græsbøll Christensen |
ICASSP | 7 |
| 2024 | Conjugate Gradient Based Adaptive Algorithm for Nonlinear AECabstractRecently, to mitigate the loudspeaker-based distortion in the acoustic system, the functional link adaptive filter – based nonlinear acoustic echo cancellation (NAEC) algorithm has been proposed. However, the usage of sine and cosine functions in nonlinear modeling coupled with the steepest descent based weight adaption limits the echo cancellation performance when the underlying distortion has faster time varying amplitude levels. Hence, in this paper, we propose a conjugate gradient (CG)-based algorithm referred to as nonlinear improved sparse conjugate algorithm. It employs the sine and cosine terms but, with time-varying coefficients to enhance distortion modeling and also an improved CG method in improving the echo cancellation performance of NAEC. The simulation results demonstrate the effectiveness of the proposed algorithm compared to the existing functional link-based NAEC. Srikanth Burra, Asutosh Kar, Mads Græsbøll Christensen |
ICASSP | 3 |
| 2024 | Improving Speech Attenuation in Headphones using Harmonic Model Decomposition and Multiple-Frequency ANCabstractIn environments such as open offices, call centres, etc., speech is often the main disturbing source of ambient noise, reducing concentration and productivity. Active noise control (ANC) systems have difficulties in dealing with speech due to its non-stationary nature and constraints in the ANC system, which require the optimal filters to be non-causal. The non-causality is due to the delay incurred by, e.g., digital processing or acoustic propagation paths. To deal with this, we propose a new feedforward ANC system for headphone applications, HMD-ANC, which improves voiced speech attenuation. Notably, in HMD-ANC, each speech harmonic and its quadrature version, obtained by the harmonic model decomposition, are predicted by a two-weight adaptive filter to overcome the delay. The results show that HMD-ANC outperforms conventional adaptive feedforward ANC for delays starting from 6 samples (0.125 ms) at a sampling frequency of 48 kHz. Moreover, HMD-ANC extends the attenuation bandwidth, e.g., up to 2.5 kHz, while the conventional ANC is limited to 1 kHz. Yurii Iotov, Sidsel Marie Nørholm, Peter John McCutcheon, Mads Græsbøll Christensen |
ICASSP | 4 |
| 2024 | Multi-layer encoder-decoder time-domain single channel speech separation
Debang Liu, Mads Græsbøll Christensen, Ying Wei 0010 |
Pattern Recognit. Lett. | 3 |
| 2024 | Approximating the zero-norm penalized sparse signal recovery using a hierarchical Bayesian framework
Zonglong Bai, Liming Shi, Mads Græsbøll Christensen |
Signal Process. | 4 |
| 2024 | Nonlinear acoustic echo cancellation using low-complexity low-rank recursive least-squares algorithms
Vinal Patel, Sankha Subhra Bhattacharjee, Jesper Rindom Jensen, Mads Græsbøll Christensen, Jacob Benesty |
Signal Process. | 4 |
| 2024 | Laguerre Kernel Adaptive Filter With Arctangent Criterion for Nonlinear System IdentificationabstractKernel adaptive filters (KAF) have emerged as a prominent method for nonlinear system identification (NSI). However, the KAF becomes computationally intensive as the input signal grows. In complex systems, traditional KAF with a unit time-delay structure may struggle with insufficient control capability. Moreover, KAFs adapted to second-order statistics can be susceptible to non-Gaussian noise. In this letter, we introduce the Laguerre kernel adaptive filter (LKAF) for NSI, using a block-oriented nonlinear model. The LKAF leverages the Laguerre series to approximate the linear block, benefiting from infinite impulse response (IIR) characteristics and a simple feedforward structure. To address non-Gaussian noise, the LKAF employs an arctangent (AT) criterion. This integration leads to the development of the Laguerre kernel arctangent least mean square (L-KATLMS) algorithm and its variations, which utilize random Fourier approximation. Simulation results demonstrate the superiority of our proposed algorithms for NSI. Yingying Zhu 0006, Haiquan Zhao 0001, Mads Græsbøll Christensen |
IEEE Signal Process. Lett. | 3 |
| 2024 | Audio-Visual Fusion With Temporal Convolutional Attention Network for Speech SeparationabstractCurrently, audio-visual speech separation methods utilize the speaker's audio and visual correlation information to help separate the speech of the target speaker. However, these methods commonly use the approach of feature concatenation with linear mapping to obtain the fused audio-visual features, which prompts us to conduct a deeper exploration for audio-visual fusion. Therefore, in this paper, according to the speaker's mouth landmark movements during speech, we propose a novel time-domain single-channel audio-visual speech separation method: audio-visual fusion with temporal convolution attention network for speech separation model (AVTCA). In this method, we design temporal convolution attention network (TCANet) based on the attention mechanism to model the contextual relationships between audio and visual sequences, and use TCANet as the basic unit to construct sequence learning and fusion network. In the whole deep separation framework, we first use cross attention to focus on the cross-correlation information of the audio and visual sequences, and then we use the TCANet to fuse the audio-visual feature sequences with temporal dependencies and cross-correlations. Afterwards, the fused audio-visual features sequences will be used as input to the separation network to predict mask and separate the source of each speaker. Finally, this paper conducts comparative experiments on Vox2, GRID, LRS2 and TCD-TIMIT datasets, indicating that AVTCA outperforms other state-of-the-art (SOTA) separation methods. Furthermore, it exhibits greater efficiency in computational performance and model size. Debang Liu, Mads Græsbøll Christensen, Zeliang An |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2024 | A Two-Stage Deep Representation Learning-Based Speech Enhancement Method Using Variational Autoencoder and Adversarial TrainingabstractThis article focuses on leveraging deep representation learning (DRL) for speech enhancement (SE). In general, the performance of the deep neural network (DNN) is heavily dependent on the learning of data representation. However, the DRL's importance is often ignored in many DNN-based SE algorithms. To obtain a higher quality enhanced speech, we propose a two-stage DRL-based SE method through adversarial training. In the first stage, we disentangle different latent variables because disentangled representations can help DNN generate a better enhanced speech. Specifically, we use the$\beta$-variational autoencoder (VAE) algorithm to obtain the speech and noise posterior estimations and related representations from the observed signal. However, since the posteriors and representations are intractable and we can only apply a conditional assumption to estimate them, it is difficult to ensure that these estimations are always pretty accurate, which may potentially degrade the final accuracy of the signal estimation. To further improve the quality of enhanced speech, in the second stage, we introduce adversarial training to reduce the effect of the inaccurate posterior towards signal reconstruction and improve the signal estimation accuracy, making our algorithm more robust for the potentially inaccurate posterior estimations. As a result, better SE performance can be achieved. The experimental results indicate that the proposed strategy can help similar DNN-based SE algorithms achieve higher short-time objective intelligibility (STOI), perceptual evaluation of speech quality (PESQ), and scale-invariant signal-to-distortion ratio (SI-SDR) scores. Moreover, the proposed algorithm can also outperform recent competitive SE algorithms. Yang Xiang 0008, Jesper Lisby Højvang, Morten Højfeldt Rasmussen, Mads Græsbøll Christensen |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Study And Design Of Robust Personal Sound Zones With Vast Using Low Rank RirsabstractThe performance of sound zone control algorithms are known to degrade significantly with changes in acoustic conditions including perturbations of control microphones' positions. In this work, we study the feasibility and effectiveness of using low rank approximations of RIRs to calculate sound zone control filters, to improve the robustness of sound zone control algorithms to perturbations in the bright zone (BZ) microphones. For algorithm design, we consider the framework of variable span linear filter (VSLF) which allows a wide range of user selectivity between acoustic contrast (AC) and signal distortion (SD) trade off, including acoustic contrast control (ACC) and pressure matching (PM) methods as special cases. Detailed simulation study shows that above a certain rank of the variable span trade-off (VAST) filter, the proposed approach using low rank RIRs to derive the control filters provides higher AC compared to using full rank RIRs, when there are perturbations in BZ microphone positions. Sankha Subhra Bhattacharjee, Liming Shi, Guoli Ping, Xiaoxiang Shen, Mads Græsbøll Christensen |
ICASSP | 5 |
| 2023 | Sparse Bayesian Learning Based Three-Dimensional Imaging for Antenna Array RadarabstractIn recent years, the development of compressed sensing and sparse representation provide us with a broader perspective of three-dimensional (3-D) imaging. In this work, we propose a 3-D imaging method based on a sparse Bayesian learning(SBL) framework for antenna array radar. It solves the problem of long-term accumulation and complicated motion compensation problem that occurs with interferometric inverse synthetic aperture radar (InISAR). Using the framework, the proposed method can automatically learn optimal hyper-parameters from the data at a low computational cost. Experimental results show that the proposed method has advantages in terms of 3-D imaging accuracy and computational efficiency compared to existing methods. Yuhan Li 0002, Jesper Rindom Jensen, Maozhong Fu, Zhenmiao Deng, Mads Græsbøll Christensen |
ICASSP | 5 |
| 2023 | Frequency Bin-Wise Single Channel Speech Presence Probability Estimation Using Multiple DNNSabstractIn this work, we propose a frequency bin-wise method to estimate the single-channel speech presence probability (SPP) with multiple deep neural networks (DNNs) in the short-time Fourier transform domain. Since all frequency bins are typically considered simultaneously as input features for conventional DNN-based SPP estimators, high model complexity is inevitable. To reduce the model complexity and the requirements on the training data, we take a single frequency bin and some of its neighboring frequency bins into account to train separate gate recurrent units. In addition, the noisy speech and the a posteriori probability SPP representation are used to train our model. The experiments were performed on the Deep Noise Suppression challenge dataset. The experimental results show that the speech detection accuracy can be improved when we employ the frequency bin-wise model. Finally, we also demonstrate that our proposed method outperforms most of the state-of-the-art SPP estimation methods in terms of speech detection accuracy and model complexity. Shuai Tao, Himavanth Reddy, Jesper Rindom Jensen, Mads Græsbøll Christensen |
ICASSP | 4 |
| 2023 | Audio-Visual Fusion using Multiscale Temporal Convolutional Attention for Time-Domain Speech SeparationabstractAudio-only speech separation methods cannot fully exploit audio-visual correlation information of speaker, which limits separation performance. Additionally, audio-visual separation methods usually adopt traditional idea of feature splicing and linear mapping to fuse audio-visual features, this approach requires us to think more about fusion process. Therefore, in this paper, combining with the changes of speaker mouth landmarks, we propose a time-domain audio-visual temporal convolution attention speech separation method (AVTA). In AVTA, we design a multiscale temporal convolutional attention (MTCA) to better focus on contextual dependencies of time sequences. We then use sequence learning and fusion network composed of MTCA to build a separation model for speech separation task. On different datasets, AVTA achieves competitive performance, and compared to baseline methods, AVTA is better balanced in training cost, computational complexity and separation performance. Debang Liu, Mads Græsbøll Christensen, Ying Wei 0010, Zeliang An |
INTERSPEECH | 3 |
| 2023 | Space alternating variational estimation based sparse Bayesian learning for complex-value sparse signal recovery using adaptive Laplace priorsabstractAbstract Due to its self‐regularising nature and its ability to quantify uncertainty, the Bayesian approach has achieved excellent recovery performance across a wide range of sparse signal recovery applications. However, most existing methods are based on the real‐value signal model, with the complex‐value signal model rarely considered. Motivated by the adaptive least absolute shrinkage and selection operator (LASSO) and the sparse Bayesian learning framework, a hierarchical model with adaptive Laplace priors is proposed in this paper for recovery of complex sparse signals. Moreover, the space alternating approach is integrated into the algorithm to reduce the computational complexity of the proposed method. In experiments, the proposed algorithm is studied for complex Gaussian random dictionaries and different types of complex signals. These experiments show that the proposed algorithm offers better recovery performance for different types of complex signals than state‐of‐the‐art methods. Zonglong Bai, Liming Shi, Jinwei Sun, Mads Græsbøll Christensen |
IET Signal Process. | 4 |
| 2023 | Widely linear complex-valued hyperbolic secant adaptive filtering algorithm and its performance analysis
Lei Li 0033, Yi-Fei Pu, Sankha Subhra Bhattacharjee, Mads Græsbøll Christensen |
Signal Process. | 4 |
| 2023 | Recursive least-squares algorithm based on a third-order tensor decomposition for low-rank system identification
Constantin Paleologu, Jacob Benesty, Cristian Lucian Stanciu, Jesper Rindom Jensen, Mads Græsbøll Christensen, Silviu Ciochina |
Signal Process. | 5 |
| 2023 | An adaptive autoregressive pre-whitener for speech and acoustic signals based on parametric NMFabstractA common assumption in many speech and acoustic processing methods is that the noise is white and Gaussian (WGN). Although making this assumption results in simple and computationally attractive methods, the assumption is often too simple and crude in many applications. In this paper, we introduce a general purpose and online pre-whitener which can be used as a pre-processor with methods based on the WGN assumption, improving their reliability and performance in applications with colored noise. The pre-whitener is a time-varying filter whose coefficients are found using a parametric non-negative matrix factorization (NMF), based on autoregressive (AR) mixture modeling of both the noise component and the signal component constituting the noisy signal. Compared to other types of pre-whiteners, we show that the proposed pre-whitener has the best performance, especially in applications with non-stationary noise. We also perform a large number of experiments to quantify the benefits of using a pre-whitener as a pre-processor for methods based on the WGN-assumption. The applications of interest were pitch estimation and time-of-arrival (TOA) estimation, where the WGN assumption is very popular. Alfredo Esquivel Jaramillo, Jesper Kjær Nielsen, Mads Græsbøll Christensen |
Speech Commun. | 3 |
| 2023 | Generalized Soft-Root-Sign Based Robust Sparsity-Aware Adaptive FiltersabstractRobust adaptive filters utilizing hyperbolic cosine and correntropy functions have been successfully employed in non-Gaussian noisy environments. However, these filters suffer from high steady-state misalignment due to significant weight update in the presences of outliers. In addition, several practical systems exhibit sparse characteristics, which is not taken into account by these filters. In this paper, a generalized soft-root-sign (GSRS) function is proposed and the corresponding GSRS adaptive filter is designed. The proposed GSRS provides negligible weight update in the occurrence of large outliers and thereby results in lower steady-state misalignment. To further improve modelling performance for sparse systems and to achieve robustness, sparsity-aware GSRS algorithms are also developed in this paper. The bound on learning rate and the computational complexity of proposed algorithm is also investigated. Simulation studies confirmed the improved convergence characteristics achieved by the proposed algorithms over existing algorithms. Vinal Patel, Sankha Subhra Bhattacharjee, Mads Græsbøll Christensen |
IEEE Signal Process. Lett. | 3 |
| 2023 | An Analysis of Traditional Noise Power Spectral Density Estimators Based on the Gaussian Stochastic Volatility ModelabstractMany single- and multi-channel speech enhancement techniques, old and new, rely in one way or another on estimates of the noise power spectral density (PSD). For example, the classical Wiener filter requires that either the speech or noise PSD be estimated. Typically, the noise PSD is estimated, as it is often easier to model and estimate than the speech. As a result, much attention has been paid to this important problem over the past couple of decades, with important scientific milestones being the minimum statistics (MS), the minima controlled recursive averaging (IMCRA), and the minimum mean squared (MMSE) estimators. Despite leading to major progress, these estimators are rather ad hoc, making them difficult to tune and improve in a systematic manner. In this article, we analyse some of the common heuristics employed in such noise PSD estimators to put them on firmer mathematical ground. More specifically, we use the Gaussian stochastic volatility model and show that the MMSE noise PSD estimator can be interpreted as a special case thereof. Moreover, we analyze the related problem of speech presence probability (SPP) estimation and show that the SPP estimation performed in the MMSE noise PSD estimator can be interpreted as an SNR estimator in the context of the Gaussian stochastic volatility model. Jesper Kjær Nielsen, Mads Græsbøll Christensen, Jesper Bünsow Boldt |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | CGMM-Based Sound Zone Generation Using Robust Pressure Matching With ATF Perturbation ConstraintsabstractPersonal sound zone (PSZ) refers to the technique that uses an array of loudspeakers and digital signal processing tools to achieve spatial soundfield control. To generate the target sound zones, this technique generally requires to know the acoustic transfer functions (ATFs) between the loudspeakers and the spots where soundfields are to be controlled. In practical applications, however, the true ATFs are never accessible and they have to be measured or estimated. Due to many sophisticated reasons, the measured ATFs generally deviate from the true ones, which may lead to significant degradation in performance of sound zone reproduction. In this work, a robust pressure matching (RPM) algorithm is presented for sound zone generation. It exploits a complex Gaussian mixture model (CGMM) to model the ATFs and their perturbations. The CGMM parameters are estimated using the expectation-maximization (EM) algorithm. To improve the robustness of the pressure matching method, an uncertainty constraint is applied to the ATF estimates and the pressure matching problem is then formulated as one of biconvex optimization. The coordinate descent algorithm is subsequently used to solve the optimization problem, thereby obtaining the optimal control filter. In comparison with the existing pressure matching methods without considering the effect of ATF perturbations, the presented algorithm is able to achieve lower normalized signal distortion energy and higher signal to interference ratio. Numerical simulations justify the effectiveness of the presented algorithm as well as its advantages over the traditional methods. Junqing Zhang, Liming Shi, Mads Græsbøll Christensen, Wen Zhang 0002, Lijun Zhang 0004, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | A New Virtual Tracking Sub-Algorithm Based Hybrid Active Control System for Narrowband Noise With Impulsive InterferenceabstractMechanical noise is usually a mixture of narrowband and impulsive noise which needs complex active noise control (ANC) algorithms to improve the de-noising performance. But the ANC algorithm with a high computation load will reduce the real-time performance of an ANC system, thus decreasing the attenuation performance and even leading to divergence. To alleviate this contradiction in narrowband ANC systems, a new virtual filtered-x L0 norm discrete Fourier cancellation (FxL0DFC) based hybrid FxNLMS(filtered-x normalized least mean square)-FxDFC framework is proposed to decrease the total computing load and keep good attenuation performance. For a fast-changing noise, the FxNLMS algorithm is employed. The new virtual FxL0DFC algorithm serves to prepare parameters for steady-state, and when this happens, the FxDFC algorithm with the parameters provided by FxL0DFC is applied. Compared to using the FxNLMS algorithm to attenuate narrowband periodical noise, the FxDFC algorithm has nearly the same tracking performance while having a low computational load. As a result, the FxL0DFC-based FxNLMS-FxDFC algorithm leads to a reduction of the total computational load. Moreover, the proposed method performs excellently in terms of tracking in simulations and experiments on actual data, particularly in environments with rapid power changes and impulsive noise. Wenzhao Zhu, Lei Luo 0009, Jinwei Sun, Mads Græsbøll Christensen |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | Sparse Modeling of The Early Part of Noisy Room Impulse Responses with Sparse Bayesian LearningabstractA model of a room impulse response (RIR) is useful for a wide range of applications. Typically, the early part of a RIR is sparse, and its sparse structure allows for accurate and simple modeling of the RIR. The existing ℓp(0 < p ≤ 1)-norm-based methods suffer from the sensitivity to the user-selected regularization parameters or a high computational burden. In this work, we propose to reconstruct the sparse model for the early part of RIRs with sparse Bayesian learning (SBL). Under the framework of SBL, the proposed method can adaptively learn the optimal hyper-parameters from data at a low computational cost. Experiment results show that the proposed method has advantages in terms of noise robustness, reconstruction sparsity, and computational efficiency compared to the existing methods. Maozhong Fu, Jesper Rindom Jensen, Yuhan Li 0002, Mads Græsbøll Christensen |
ICASSP | 4 |
| 2022 | Computationally Efficient Fixed-Filter ANC for Speech Based on Long-Term Prediction for Headphone ApplicationsabstractIn some situations, such as open office spaces, speech can play the role of an unwanted and disturbing source of noise, and ANC headphones or earbuds might help to solve this problem. However, ANC in modern headphones is often based on a pre-calculated fixed-filter for practical reasons, like stability and cost. Moreover, in some cases the optimal filter is non-causal, which cannot be realized with such a filter, and ANC attenuation performance will be significantly decreased. In this paper we propose to solve the causality problem in feedforward fixed-filter ANC systems by integrating a long-term linear prediction filter to predict the incoming disturbance, here speech, by the same amount of samples ahead in time, as the non-causal delay. The proposed ANC system outperforms conventional adaptive feedforward ANC systems in terms of computational complexity, showing comparable or better results on voiced speech attenuation at non-causal delays from 4 to 18 samples (0.5 to 2.25 ms) at a sampling frequency of 8 kHz. Yurii Iotov, Sidsel Marie Nørholm, Valiantsin Belyi, Mads Dyrholm, Mads Græsbøll Christensen |
ICASSP | 5 |
| 2022 | Privacy-Preserving Distributed Expectation Maximization for Gaussian Mixture Model Using Subspace PerturbationabstractPrivacy has become a major concern in machine learning. In fact, the federated learning is motivated by the privacy concern as it does not allow to transmit the private data but only intermediate updates. However, federated learning does not always guarantee privacy-preservation as the intermediate updates may also reveal sensitive information. In this paper, we give an explicit information-theoretical analysis of a federated expectation maximization algorithm for Gaussian mixture model and prove that the intermediate updates can cause severe privacy leakage. To address the privacy issue, we propose a fully decentralized privacy-preserving solution, which is able to securely compute the updates in each maximization step. Additionally, we consider two different types of security attacks: the honest-but-curious and eavesdropping adversary models. Numerical validation shows that the proposed approach has superior performance compared to the existing approach in terms of both the accuracy and privacy level. Qiongxiu Li, Jaron Skovsted Gundersen, Katrine Tjell, Rafael Wisniewski, Mads Græsbøll Christensen |
ICASSP | 5 |
| 2022 | Generation of Personal Sound Fields in Reverberant Environments Using Interframe CorrelationabstractPersonal sound field control techniques aim to produce sound fields for different sound contents in different places of an acoustic space without interference. The limitations of the state-of-the-art methods for sound field control include high latency and computational complexity, especially in the cases when the reverberation time is long and number of loudspeakers is large. In this paper, we propose a personal sound field control approach that exploits interframe correlation. Considering the past frames, the proposed method can accommodate long reverberation time with a low latency. To find the optimal parameters for the physical meaningful constraints, the subspace decomposition and Newton’s method are applied. Furthermore, a sound field distortion oriented subspace construction method is proposed to reduce the subspace dimension. Compared with traditional methods, simulation results show that the proposed algorithm is able to obtain a good trade-off between acoustic contrast and reproduction error with a low latency for measured room impulse responses. Liming Shi, Guoli Ping, Xiaoxiang Shen, Mads Græsbøll Christensen |
ICASSP | 4 |
| 2022 | A Bayesian Permutation Training Deep Representation Learning Method for Speech Enhancement with Variational AutoencoderabstractRecently, variational autoencoder (VAE), a deep representation learning (DRL) model, has been used to perform speech enhancement (SE). However, to the best of our knowledge, current VAE-based SE methods only apply VAE to model speech signal, while noise is modeled using the traditional non-negative matrix factorization (NMF) model. One of the most important reasons for using NMF is that these VAE-based methods cannot disentangle the speech and noise latent variables from the observed signal. Based on Bayesian theory, this paper derives a novel variational lower bound for VAE, which ensures that VAE can be trained in supervision, and can disentangle speech and noise latent variables from the observed signal. This means that the proposed method can apply the VAE to model both speech and noise signals, which is totally different from the previous VAE-based SE works. More specifically, the proposed DRL method can learn to impose speech and noise signal priors to different sets of latent variables for SE. The experimental results show that the proposed method can not only disentangle speech and noise latent variables from the observed signal, but also obtain a higher scale-invariant signal-to-distortion ratio and speech quality score than the similar deep neural network-based (DNN) SE method. Yang Xiang 0008, Jesper Lisby Højvang, Morten Højfeldt Rasmussen, Mads Græsbøll Christensen |
ICASSP | 4 |
| 2022 | Robust Pressure Matching with ATF Perturbation Constraints for Sound Field ControlabstractSound field control systems deployed in room acoustic environments require knowing the acoustic channel impulse responses between the loudspeakers and matching microphones, which are challenging to estimate accurately due to perturbations caused by such factors as temperature changes and sensors’ position mismatches. To deal with this issue, a robust pressure matching algorithm is developed in this work where a perturbation term of the acoustic transfer function (ATF) is modeled as a Gaussian process, based on which an uncertainty constraint is applied to limit the impact of perturbation on pressure matching. This constrained problem is formulated as one of biconvex optimization, and a coordinate descent algorithm is adopted to estimate the optimal control filter. Simulations are performed and results show that the proposed method is able to achieve more accurate control as compared to the standard pressure matching algorithm in the presence of ATF perturbations. Junqing Zhang, Liming Shi, Mads Græsbøll Christensen, Wen Zhang 0002, Lijun Zhang 0004, Jingdong Chen |
ICASSP | 3 |
| 2022 | Communication efficient privacy-preserving distributed optimization using adaptive differential quantizationabstractPrivacy issues and communication cost are both major concerns in distributed optimization in networks. There is often a trade-off between them because the encryption methods used for privacy-preservation often require expensive communication overhead. To address these issues, we, in this paper, propose a quantization-based approach to achieve both communication efficient and privacy-preserving solutions in the context of distributed optimization. By deploying an adaptive differential quantization scheme, we allow each node in the network to achieve its optimum solution with a low communication cost while keeping its private data unrevealed. Additionally, the proposed approach is general and can be applied in various distributed optimization methods, such as the primal-dual method of multipliers (PDMM) and the alternating direction method of multipliers (ADMM). We consider two widely used adversary models, passive and eavesdropping, and investigate the properties of the proposed approach using different applications and demonstrate its superior performance compared to existing privacy-preserving approaches in terms of both accuracy and communication cost. Qiongxiu Li, Richard Heusdens, Mads Græsbøll Christensen |
Signal Process. | 3 |
| 2021 | A Novel NMF-HMM Speech Enhancement Algorithm Based on Poisson Mixture ModelabstractIn this paper, we propose a novel non-negative matrix factorization (NMF) and hidden Markov model (NMF-HMM) based speech enhancement algorithm, which employs a Poisson mixture model (PMM). Compared to the previously proposed NMF-HMM method, the new algorithm, termed PMM-NMF-HMM, uses the Poisson mixture distribution for the state conditional likelihood function for a HMM rather than the single Poisson distribution. This means that there are the more basis matrices that can be used to model the speech and noise signals, so more signal information can be captured by the resulting model. The proposed method is supervised and thus includes a training and an enhancement stage. It is shown that, in the training stage, the proposed method can be implemented efficiently using multiplicative update (MU) for the model parameters, much like the NMF-HMM algorithm. In the speech enhancement stage, which can be performed online, a novel PMM-NMF-HMM minimum mean-square error (MMSE) estimator is developed. The experimental results indicate that the PMM-NMF-HMM method can obtain higher short-time objective intelligibility (STOI) and perceptual evaluation of speech quality (PESQ) score than NMF-HMM. Additionally, the method also outperforms other state-of-the-art NMF- based supervised speech enhancement algorithms. Yang Xiang 0008, Liming Shi, Jesper Lisby Højvang, Morten Højfeldt Rasmussen, Mads Græsbøll Christensen |
ICASSP | 5 |
| 2021 | Speech Decomposition Based on a Hybrid Speech Model and Optimal SegmentationabstractIn a hybrid speech model, both voiced and unvoiced components can coexist in a segment. Often, the voiced speech is regarded as the deterministic component, and the unvoiced speech and additive noise are the stochastic components. Typically, the speech signal is considered stationary within fixed segments of 20-40 ms, but the degree of stationarity varies over time. For decomposing noisy speech into its voiced and unvoiced components, a fixed segmentation may be too crude, and we here propose to adapt the segment length according to the signal local characteristics. The segmentation relies on parameter estimates of a hybrid speech model and the maximum a posteriori (MAP) and log-likelihood criteria as rules for model selection among the possible segment lengths, for voiced and unvoiced speech, respectively. Given the optimal segmentation markers and the estimated statistics, both components are estimated using linear filtering. A codebook-based approach differentiates between unvoiced speech and noise. A better extraction of the components is possible by taking into account the adaptive segmentation, compared to a fixed one. Also, a lower distortion for voiced speech and higher segSNR for both components is possible, as compared to other decomposition methods. Alfredo Esquivel Jaramillo, Jesper Kjær Nielsen, Mads Græsbøll Christensen |
Interspeech | 3 |
| 2021 | Fast algorithms for fundamental frequency estimation in autoregressive noise
Barry G. Quinn, Jesper Kjær Nielsen, Mads Græsbøll Christensen |
Signal Process. | 3 |
| 2021 | Automatic quality control and enhancement for voice-based remote Parkinson's disease detection
Amir Hossein Poorjam, Mathew Shaji Kavalekalam, Liming Shi, Yordan P. Raykov, Jesper Rindom Jensen, Max A. Little, Mads Græsbøll Christensen |
Speech Commun. | 7 |
| 2021 | Fast Generation of Sound Zones Using Variable Span Trade-Off Filters in the DFT-DomainabstractThe creation of sound zones with frequency-domain variable span trade-off filters (VAST) is investigated herein. Both narrowband and broadband discrete Fourier transform (DFT)-domain VAST approaches are proposed, and we discuss their relationship to the existing time-domain VAST approach. The core idea in VAST is to apply a generalized eigenvalue decomposition to the spatial statistics to control the trade-off between acoustic contrast and signal distortion. Moreover, a method for determining the optimal Lagrange multiplier that controls this trade-off is also considered in terms of physical, meaningful parameters. Through analysis and experiments, a performance comparison using measured room impulse responses is conducted not only between the two proposed methods but also between the two proposed methods and the existing time-domain approach. The results confirm that the broadband approach is able to transfer the acoustic contrast from one frequency bin to another, which is not the case for the narrowband approach. Furthermore, the results also show that the proposed DFT-domain VAST approach can be considered to be a special case of the time-domain VAST approach. Taewoong Lee, Liming Shi, Jesper Kjær Nielsen, Mads Græsbøll Christensen |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Generation of Personal Sound Zones With Physical Meaningful Constraints and Conjugate Gradient MethodabstractPersonal sound zones provide users to experience independent listening and quiet areas in the same acoustic environment using multiple loudspeakers. The generalized eigenvalue decomposition (GEVD) has been proposed for sound zones generation, allowing user to control the trade-off between acoustic contrast and signal distortion by adjusting some parameters. Unfortunately, these parameters are not physically meaningful, and the user has to tune them for different source materials and acoustic environments. Moreover, performing a high dimensional GEVD is computational complex. In this article, we first propose various strategies to control the reproduced sound zones as precisely and accurately as possible by reformulating the problem using physically meaningful constraints using regularization approach. Then, a hybrid approach of combining the conjugate gradient method and GEVD is proposed to reduce the computational complexity and signal distortion when the subspace dimension is small. The proposed methods show precise control over the reproduced sound zone via extensive numerical simulations in reverberant environments for different physically meaningful constraints. Liming Shi, Taewoong Lee, Lijun Zhang 0004, Jesper Kjær Nielsen, Mads Græsbøll Christensen |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2021 | Privacy-Preserving Distributed Processing: Metrics, Bounds and AlgorithmsabstractPrivacy-preserving distributed processing has recently attracted considerable attention. It aims to design solutions for conducting signal processing tasks over networks in a decentralized fashion without violating privacy. Many existing algorithms can be adopted to solve this problem such as differential privacy, secure multiparty computation, and the recently proposed distributed optimization based subspace perturbation algorithms. However, since each of them is derived from a different context and has different metrics and assumptions, it is hard to choose or design an appropriate algorithm in the context of distributed processing. In order to address this problem, we first propose general mutual information based information-theoretical metrics that are able to compare and relate these existing algorithms in terms of two key aspects: output utility and individual privacy. We consider two widely-used adversary models, the passive and eavesdropping adversary. Moreover, we derive a lower bound on individual privacy which helps to understand the nature of the problem and provides insights on which algorithm is preferred given different conditions. To validate the above claims, we investigate a concrete example and compare a number of state-of-the-art approaches in terms of the concerned aspects using not only theoretical analysis but also numerical validation. Finally, we discuss and provide principles for designing appropriate algorithms for different applications. Qiongxiu Li, Jaron Skovsted Gundersen, Richard Heusdens, Mads Græsbøll Christensen |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2020 | Autoregressive Parameter Estimation with Dnn-Based Pre-ProcessingabstractIn this paper, a method for estimating the autoregressive parameters from a signal segment is proposed. The method is based on a deep neural network (DNN) in combination with the classical Levinson-Durbin recursion (LDR). The DNN acts as a pre-processor for the LDR and can be trained on different metrics commonly encountered in speech processing using a generalized analysis-by-synthesis (GABS) structure where the LDR acts as the encoder. Unlike end-to-end data-driven approaches, this structure ensures that the DNN is easy to train and initialize since the DNN only has to learn a simple mapping. The results confirm this and show that the proposed method produces an AR-spectrum that efficiently represents the speech spectrum in terms of the Itakura-Saito divergence, Kullback-Leibler divergence, log-spectral distortion, and speech distortion. Zihao Cui, Changchun Bao, Jesper Kjær Nielsen, Mads Græsbøll Christensen |
ICASSP | 4 |
| 2020 | Robust Fundamental Frequency Estimation in Coloured NoiseabstractMost parametric fundamental frequency estimators make the implicit assumption that any corrupting noise is additive, white Gaus-sian. Under this assumption, the maximum likelihood (ML) and the least squares estimators are the same, and statistically efficient. However, in the coloured noise case, the estimators differ, and the spectral shape of the corrupting noise should be taken into account. To allow for this, we here propose two schemes that refine the noise statistics and parameter estimates in an iterative manner, one of them based on an approximate ML solution and the other one based on removing the periodic signal obtained from a linearly constrained minimum variance (LCMV) filter. Evaluations on real speech data indicate that the iteration steps improve the estimation accuracy, therefore offering improvement over traditional non-parametric fundamental frequency methods in most of the evaluated scenarios. Alfredo Esquivel Jaramillo, Andreas Jakobsson, Jesper Kjær Nielsen, Mads Græsbøll Christensen |
ICASSP | 4 |
| 2020 | Convex Optimisation-Based Privacy-Preserving Distributed Average Consensus in Wireless Sensor NetworksabstractIn many applications of wireless sensor networks, it is important that the privacy of the nodes of the network be protected. Therefore, privacy-preserving algorithms have received quite some attention recently. In this paper, we propose a novel convex optimization-based solution to the problem of privacy-preserving distributed average consensus. The proposed method is based on the primal-dual method of multipliers (PDMM), and we show that the introduced dual variables of the PDMM will only converge in a certain subspace determined by the graph topology and will not converge in the orthogonal complement. These properties are exploited to protect the private data from being revealed to others. More specifically, the proposed algorithm is proven to be secure for both passive and eavesdropping adversary models. Finally, the convergence properties and accuracy of the proposed approach are demonstrated by simulations which show that the method is superior to the state-of-the-art. Qiongxiu Li, Richard Heusdens, Mads Græsbøll Christensen |
ICASSP | 3 |
| 2020 | A Fast Reduced-Rank Sound Zone Control Algorithm Using The Conjugate Gradient MethodabstractSound zone control enables different users to enjoy different audio contents in the same acoustic environment. Generalized eigenvalue decomposition (GEVD)-based methods allow us to control the tradeoff between the acoustic contrast (AC) and signal distortion (SD). However, such methods have a high computational complexity. In this paper, we propose a fast reduced-rank sound zone control algorithm using the conjugate gradient (CG) method. Instead of using the eigenvectors as the basis for the solution space, the search directions in the CG method are used to reduce the computational complexity. Then, a low dimensional EVD is applied to obtain the sub-optimal control filter coefficients. The dark zone power can be adjusted by a parameter, which implicitly controls the trade-off between the AC and SD. Compared with GEVD-based methods, experimental results show that the proposed algorithm has a degradation of performance (4-5 dB) in terms of AC or SD but a high improvement on computational efficiency. Liming Shi, Taewoong Lee, Lijun Zhang 0004, Jesper Kjær Nielsen, Mads Græsbøll Christensen |
ICASSP | 5 |
| 2020 | An NMF-HMM Speech Enhancement Method Based on Kullback-Leibler DivergenceabstractIn this paper, we present a novel supervised Non-negative Matrix Factorization (NMF) speech enhancement method, which is based on Hidden Markov Model (HMM) and Kullback- Leibler (KL) divergence (NMF-HMM). Our algorithm applies theHMMto capture the timing information, so the temporal dynamics of speech signal can be considered by comparing with the traditional NMF-based speech enhancement method. More specifically, the sum of Poisson, leading to the KL divergence measure, is used as the observation model for each state of HMM. This ensures that the parameter update rule of the proposed algorithm is identical to the multiplicative update rule, which is quick and efficient. In the training stage, this update rule is applied to train the NMF-HMM model. In the online enhancement stage, a novel minimum mean-square error (MMSE) estimator that combines the NMF-HMM is proposed to conduct speech enhancement. The performance of the proposed algorithm is evaluated by perceptual evaluation of speech quality (PESQ) and short-timeobjective intelligibility (STOI). The experimental results indicate that the STOI score of proposed strategy is able to outperform 7% than current state-of-the-art NMF-based speech enhancement methods. Yang Xiang 0008, Liming Shi, Jesper Lisby Højvang, Morten Højfeldt Rasmussen, Mads Græsbøll Christensen |
INTERSPEECH | 5 |
| 2020 | Harmonic beamformers for speech enhancement and dereverberation in the time domain
Jesper Rindom Jensen, Sam Karimian-Azari, Mads Græsbøll Christensen, Jacob Benesty |
Speech Commun. | 3 |
| 2020 | Signal-Adaptive and Perceptually Optimized Sound Zones With Variable Span Trade-Off FiltersabstractCreating sound zones has been an active research field since the idea was first proposed. So far, most sound zone control methods rely on either an optimization of physical metrics such as acoustic contrast and signal distortion or a mode decomposition of the desired sound field. By using these types of methods, approximately 15 dB of acoustic contrast between the reproduced sound field in the target zone and its leakage to other zone(s) has been reported in practical set-ups, but this is typically not high enough to satisfy the people inside the zones. In this article, we propose a sound zone control method shaping the leakage errors so that they are as inaudible as possible for a given acoustic contrast. The shaping of the leakage errors is performed by taking the time-varying input signal characteristics and the human auditory system into account when the loudspeaker control filters are calculated. We show how this shaping can be performed using variable span trade-off filters, and we show theoretically how these filters can be used for trading signal distortion in the target zone for acoustic contrast. The proposed method is evaluated based on physical metrics such as acoustic contrast and perceptual metrics such as STOI. The computational complexity and processing time of the proposed method for different system set-ups are also investigated. Lastly, the results of a MUSHRA listening test are reported. The test results show that the proposed method provides more than 20% perceptual improvement compared to existing sound zone control methods. Taewoong Lee, Jesper Kjær Nielsen, Mads Græsbøll Christensen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Estimation of Guitar String, Fret and Plucking Position Using Parametric Pitch EstimationabstractIn this paper a fast yet effective method is proposed for analyzing guitar performances. Specifically, the activated string and fret as well as the location of the plucking event along the guitar string are extracted from guitar signal recordings. The method is based on a parametric pitch estimator and is derived from a physically meaningful model that includes inharmonicity. A maximum a posteriori classifier is proposed, which requires training data captured from only one fret per string. The classifier is tested on recordings of electric and acoustic guitar and performs well: the average absolute error of string and fret classification is 1.5%, while the error rate varies depending on the fret used for training. The plucking position estimator is the minimizer of the log spectral distance between the amplitudes of the observed signal and the plucking model and it is evaluated in proof-of-concept experiments with sudden changes of string, fret and plucking positions, which can be estimated accurately. Unlike the state of the art, the proposed method works on very short segments, which makes it suitable for high-tempo and real-time applications. Jacob Møller Hjerrild, Mads Græsbøll Christensen |
ICASSP | 2 |
| 2019 | A Study on How Pre-whitening Influences Fundamental Frequency EstimationabstractThis paper deals with the influence of pre-whitening for the task of fundamental frequency estimation in noisy conditions. Parametric fundamental frequency estimators commonly assume that the noise is white and Gaussian and, therefore, they are only statistically efficient under those conditions. The noise is coloured in many practical applications and this will often result in problems of misidentifying an integer divisor or multiple of the true fundamental frequency (i.e., octave errors). The purpose of this paper is to see if pre-whitening can reduce this problem, based on noise statistics obtained from existing noise PSD estimation algorithms. For this purpose, different noise types and prediction orders of LPC pre-whitening are considered. The results show that pre-whitening improves significantly the estimation accuracy of an NLS pitch estimator when the noise is fairly stationary. For nonstationary noise, the improvements are modest at best, but we hypothesize that this is due to the noise PSD estimation performance rather than the LPC pre-whitening principle. Alfredo Esquivel Jaramillo, Jesper Kjær Nielsen, Mads Græsbøll Christensen |
ICASSP | 3 |
| 2019 | Hearing Aid-controlled Beamformer for Binaural Speech Enhancement Using a Model-based ApproachabstractThe understanding of speech from a particular speaker in the presence of other interfering speakers can be severely degraded for a hearing impaired person. Beamforming techniques have been proven to be effective to improve the speech understanding in such scenarios. However, the number of microphones in a hearing aid (HA) is limited due to the space and power constraints present in the HA. In this paper, we propose to use an external device e.g., a microphone array, that can communicate with the HA to overcome this limitation. We propose a method to control this external device based on the look direction of the HA user. We show, by means of simulations, the robustness of the proposed method at very low SNRs in a reverberant scenario. Moreover, we have also conducted experiments that show the benefit of using this framework for binaural and monaural enhancement. Mathew Shaji Kavalekalam, Jesper Kjær Nielsen, Mads Græsbøll Christensen, Jesper Bünsow Boldt |
ICASSP | 3 |
| 2019 | Towards Perceptually Optimized Sound Zones: A Proof-of-concept StudyabstractThe creation of sound zones has been an active research topic for approximately two decades. Many sound zone control methods have been proposed, and the best approaches result in a target to interferer ratio (TIR) of about 15 dB in a practical set-up. Unfortunately, this is far from a TIR of about 25 dB which is currently believed necessary to make sound zones commercially viable. However, state-of-the-art sound zone control methods take neither the input signal characteristics nor human auditory perception into account. In this paper, we show how a recently proposed sound zone control framework called VAST can be extended into perceptual VAST (P-VAST) which takes input signal characteristics and human auditory perception into account. We also make a proof-of-concept simulation and an AB preference test which both show that P-VAST outperforms traditional sound zone control methods in terms of perceptually meaningful metrics such as STOI and PESQ in a fairly simple set-up. Taewoong Lee, Jesper Kjær Nielsen, Mads Græsbøll Christensen |
ICASSP | 3 |
| 2019 | Quality Control of Voice Recordings in Remote Parkinson's Disease Monitoring Using the Infinite Hidden Markov ModelabstractThe performance of voice-based systems for remote monitoring of Parkinson's disease is highly dependent on the degree of adherence of the recordings to the test protocols, which probe for specific symptoms. Identifying segments of the signal that adhere to the protocol assumptions is typically performed manually by experts. This process is costly, time consuming, and often infeasible for large-scale data sets. In this paper, we propose a method to automatically identify the segments of signals that violate the test protocol with a high accuracy. In our approach, the signal is first split into variable duration segments by fitting an infinite hidden Markov model (iHMM) to the frames of the signals in the mel-frequency cepstral domain. The complexity of the iHMM is capable of growing jointly with the data allowing us to infer a potentially large (asymptotically infinite) number of different phenomena segmented into different hidden states. Then, we identify the segments that adhere to the test protocol by applying a multinomial naive Bayes classifier to the state indicators of segments. The experimental results show that even by using a small amount of training data, we can achieve around 96% accuracy in identifying short-term protocol violations with a 0.2 s resolution. Amir Hossein Poorjam, Yordan P. Raykov, Reham Badawy, Jesper Rindom Jensen, Mads Græsbøll Christensen, Max A. Little |
ICASSP | 5 |
| 2019 | Harmonic Beamformers for Non-Intrusive Speech Intelligibility PredictionabstractIn recent years, research into objective speech intelligibil- ity measures has gained increased interest as a tool to opti- mize speech enhancement algorithms. While most intelligi- bility measures are intrusive, i.e., they require a clean refer- ence signal, this is rarely available in real-time applications. This paper proposes two non-intrusive intelligibility measures, which allow using the intrusive short-time objective intelligibil- ity (STOI) measure without requiring access to the clean signal. Instead, a reference signal is obtained from the degraded sig- nal using either a fixed or an adaptive harmonic spatial filter. This reference signal is then used as input to STOI. The exper- imental results show a high correlation between both proposed non-intrusive speech intelligibility measures and the original in- trusively computed STOI scores. Charlotte Sørensen, Jesper Bünsow Boldt, Mads Græsbøll Christensen |
INTERSPEECH | 3 |
| 2019 | Validation of the Non-Intrusive Codebook-Based Short Time Objective Intelligibility Metric for Processed SpeechabstractIn recent years, objective measures of speech intelligibility have gained increasing interest. However, most speech intelligibil- ity metrics require a clean reference signal, which is often not available in real-life applications. In a recent publication, we proposed a method, the Non-Intrusive Codebook-based Short- Time Objective Intelligibility (NIC-STOI) metric, which allows using an intrusive method without requiring access to the clean signal. The statistics of the reference signal is estimated as a combination of predefined codebooks that best fit the degraded signal by modeling the speech and noisy spectra. In this pa- per, we perform additional validation of the NIC-STOI in more diverse noise condition as well as for speech processed non- linearly with binary masks, where it is shown to outperform existing non-intrusive metrics. Charlotte Sørensen, Jesper Bünsow Boldt, Mads Græsbøll Christensen |
INTERSPEECH | 3 |
| 2019 | Estimation of Fundamental Frequencies in Stereophonic Music MixturesabstractIn this paper, a method for multi-pitch estimation of stereophonic mixtures of harmonic signals, e.g., instrument recordings, is presented. The proposed method is based on a signal model that includes the panning parameters of the sources in a stereophonic mixture, such as those applied artificially in a recording studio. If the sources in a mixture have different panning parameters, this diversity can be used to simplify the pitch estimation problem. The mixing parameters of the sources might be shared, resulting in a multi-pitch estimation problem, which is solved using an approach based on an expectation-maximization algorithm for Gaussian sources, where the fundamental frequencies and model orders are estimated jointly. The fundamental frequencies may be related, resulting in overlapping harmonics, complicating the estimation of the parameters. A codebook of harmonic amplitude vectors is trained on recordings of instruments playing single notes, and used when estimating the amplitudes of the mixture components. The proposed method is evaluated using stereophonic mixtures of instrument recordings and is compared to state-of-the-art transcription and multi-pitch estimation methods. Experiments show an increase in performance when knowledge about the panning parameters is taken into account. The proposed method provides a full parameterization of the components of the observed signal. Possible applications include instrument tuning, audio editing tools, modification of harmonic mixture components, and audio effects. Martin Weiss Hansen, Jesper Rindom Jensen, Mads Græsbøll Christensen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Model-Based Speech Enhancement for Intelligibility Improvement in Binaural Hearing AidsabstractSpeech intelligibility is often severely degraded among hearing impaired individuals in situations such as the cocktail party scenario. The performance of the current hearing aid technology has been observed to be limited in these scenarios. In this paper, we propose a binaural speech enhancement framework that takes into consideration the speech production model. The enhancement framework proposed here is based on the Kalman filter that allows us to take the speech production dynamics into account during the enhancement process. The usage of a Kalman filter requires the estimation of clean speech and noise short term predictor (STP) parameters, and the clean speech pitch parameters. In this work, a binaural codebook-based method is proposed for estimating the STP parameters, and a directional pitch estimator based on the harmonic model and maximum likelihood principle is used to estimate the pitch parameters. The proposed method for estimating the STP and pitch parameters jointly uses the information from left and right ears, leading to a more robust estimation of the filter parameters. Objective measures such as PESQ and STOI have been used to evaluate the enhancement framework in different acoustic scenarios representative of the cocktail party scenario. We have also conducted subjective listening tests on a set of nine normal hearing subjects, to evaluate the performance in terms of intelligibility and quality improvement. The listening tests show that the proposed algorithm, even with access to only a single channel noisy observation, significantly improves the overall speech quality, and the speech intelligibility by up to 15 %. Mathew Shaji Kavalekalam, Jesper Kjær Nielsen, Jesper Bünsow Boldt, Mads Græsbøll Christensen |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | Robust Bayesian Pitch Tracking Based on the Harmonic ModelabstractFundamental frequency is one of the most important characteristics of speech and audio signals. Harmonic model-based fundamental frequency estimators offer a higher estimation accuracy and robustness against noise than the widely used autocorrelation-based methods. However, the traditional harmonic model-based estimators do not take the temporal smoothness of the fundamental frequency, the model order, and the voicing into account as they process each data segment independently. In this paper, a fully Bayesian fundamental frequency tracking algorithm based on the harmonic model and a first-order Markov process model is proposed. Smoothness priors are imposed on the fundamental frequencies, model orders, and voicing using first-order Markov process models. Using these Markov models, fundamental frequency estimation and voicing detection errors can be reduced. Using the harmonic model, the proposed fundamental frequency tracker has an improved robustness to noise. An analytical form of the likelihood function, which can be computed efficiently, is derived. Compared to the state-of-the-art neural network and nonparametric approaches, the proposed fundamental frequency tracking algorithm has superior performance in almost all investigated scenarios, especially in noisy conditions. For example, under 0 dB white Gaussian noise, the proposed algorithm reduces the mean absolute errors and gross errors by 15% and 20% on the Keele pitch database and 36% and 26% on sustained /a/ sounds from a database of Parkinson's disease voices. A MATLAB version of the proposed algorithm is made freely available for reproduction of the results.11An implementation of the proposed algorithm using MATLAB may be found in https://tinyurl.com/yxn4a543. Liming Shi, Jesper Kjær Nielsen, Jesper Rindom Jensen, Max A. Little, Mads Græsbøll Christensen |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2018 | Estimation of Source Panning Parameters and Segmentation of Stereophonic MixturesabstractIn this paper, we propose a method for finding the number of sources and their parameters from stereophonic mixtures. The method is based on clustering of narrowband interaural level and time differences for an unknown number of sources and uses an optimal segmentation on which the clustering is based. The parameter distribution, for both individual segments and across segments that comprise the entire signal, is modelled as a Gaussian mixture. For each segment parameters are estimated using a minimum description length algorithm for mixtures based on the expectation-maximization algorithm. The generalized variance and degree of membership of the Gaussian components across segments is used as a basis for the proposed selection of clusters amongst candidates. Simulations on synthetic and real audio shows promising results for source parameter estimation and number of sources estimated across segments. The optimal segmentation shows an improvement for parameter estimation success rate, compared to the uniform segmentation. Jacob Møller Hjerrild, Mads Græsbøll Christensen |
ICASSP | 2 |
| 2018 | A Study of Noise PSD Estimators for Single Channel Speech EnhancementabstractThe estimation of the noise power spectral density (PSD) forms a critical component of several existing single channel speech enhancement systems. In this paper, we evaluate one new and some of the existing and commonly used noise PSD estimation algorithms in terms of the spectral estimation accuracy and the enhancement performance for different commonly encountered background noises, which are stationary and non-stationary in nature. The evaluated algorithms include the Minimum Statistics, MMSE, IMCRA methods and a new model-based method. Mathew Shaji Kavalekalam, Jesper Kjær Nielsen, Mads Græsbøll Christensen, Jesper Bünsow Boldt |
ICASSP | 3 |
| 2018 | A Unified Approach to Generating Sound Zones Using Variable Span Linear FiltersabstractSound zones are typically created using Acoustic Contrast Control (ACC), Pressure Matching (PM), or variations of the two. ACC maximizes the acoustic potential energy contrast between a listening zone and a quiet zone. Although the contrast is maximized, the phase is not controlled. To control both the amplitude and the phase, PM instead minimizes the difference between the reproduced sound field and the desired sound field in all zones. On the surface, ACC and PM seem to control sound fields differently, but we here demonstrate they are actually extreme special cases of a much more general framework. The framework is inspired by the variable span linear filtering framework for speech enhancement. Using this framework, we demonstrate that 1) ACC gives the best contrast, but the highest signal distortion in the bright zone, and 2) PM gives the smallest signal distortion in the bright zone, but the worst contrast. Aside from showing this mathematically, we also demonstrate this via a small toy example. Taewoong Lee, Jesper Kjær Nielsen, Jesper Rindom Jensen, Mads Græsbøll Christensen |
ICASSP | 4 |
| 2018 | Model-Based Noise PSD Estimation from Speech in Non-Stationary NoiseabstractMost speech enhancement algorithms need an estimate of the noise power spectral density (PSD) to work. In this paper, we introduce a model-based framework for doing noise PSD estimation. The proposed framework allows us to include prior spectral information about the speech and noise sources, can be configured to have zero tracking delay, and does not depend on estimated speech presence probabilities. This is in contrast to other noise PSD estimators which often have a too large tracking delay to give good results in non- stationary situations and offer no consistent way of including prior information about the speech or the noise type. The results show that the proposed method outperforms state-of-the-art noise PSD estima- tors in terms of tracking speed and estimation accuracy. Jesper Kjær Nielsen, Mathew Shaji Kavalekalam, Mads Græsbøll Christensen, Jesper Bünsow Boldt |
ICASSP | 3 |
| 2018 | A Parametric Approach for Classification of Distortions in Pathological VoicesabstractIn biomedical acoustics, distortion in voice signals, commonly present during acquisition and transmission, adversely affects acoustic features extracted from pathological voice. Information on the type of distortion can help in compensating for its effects. This paper proposes a new approach to detecting four major types of commonly encountered distortion in remote analysis of pathological voice, namely background noise, reverberation, clipping and coding. In this approach, by applying factor analysis to Gaussian mixture model mean supervectors, distortions in variable-duration recordings are modeled by fixed-length, low-dimensional channel vectors. Then, linear discriminant analysis (LDA) is used to remove the remaining nuisance effects in the channel vectors. Finally, two different classifiers, namely support vector machines and probabilistic LDA classify the different types of distortion. Experimental results obtained using Parkinson's voices, as an example of pathological voice, show 11.4% relative improvement in performance over systems which directly use acoustic features for distortion classification. Amir Hossein Poorjam, Max A. Little, Jesper Rindom Jensen, Mads Græsbøll Christensen |
ICASSP | 4 |
| 2018 | A Supervised Approach to Global Signal-to-Noise Ratio Estimation for Whispered and Pathological VoicesabstractThe presence of background noise in signals adversely affects the performance of many speech-based algorithms. Accurate estimation of signal-to-noise-ratio (SNR), as a measure of noise level in a signal, can help in compensating for noise effects. Most existing SNR estimation methods have been developed for normal speech and might not provide accurate estimation for special speech types such as whispered or disordered voices, particularly, when they are corrupted by non-stationary noises. In this paper, we first investigate the impact of stationary and non-stationary noise on the behavior of mel-frequency cepstral coefficients (MFCCs) extracted from normal, whispered and pathological voices. We demonstrate that, regardless of the speech type, the mean and the covariance of MFCCs are predictably modified by additive noise and the amount of change is related to the noise level. Then, we propose a new supervised method for SNR estimation which is based on a regression model trained on MFCCs of the noisy signals. Experimental results show that the proposed approach provides accurate estimation and consistent performance for various speech types under different noise conditions. Amir Hossein Poorjam, Max A. Little, Jesper Rindom Jensen, Mads Græsbøll Christensen |
ICASSP | 4 |
| 2018 | Multipitch Estimation Using Block Sparse Bayesian Learning and Intra-Block ClusteringabstractPitch estimation is an important task in speech and audio analysis. In this paper, we present a multi-pitch estimation algorithm based on block sparse Bayesian learning and intra-block clustering for speech analysis. A statistical hierarchical model is formulated based on a pitch dictionary with a fixed maximum number of harmonics for all the candidate pitches. Block sparse Bayesian learning is proposed for estimating the complex amplitudes. To deal with the problem of unknown harmonic orders and subharmonic errors, intra-block clustering structured sparsity prior is also introduced. The statistical update formulas are obtained by the variational Bayesian inference. Compared with the conventional group LASSO-type algorithms for multi-pitch estimation, experimental results indicate robustness against noise and improved estimation accuracy of the proposed method. Liming Shi, Jesper Rindom Jensen, Jesper Kjær Nielsen, Mads Græsbøll Christensen |
ICASSP | 4 |
| 2018 | Non-intrusive codebook-based intelligibility prediction
Charlotte Sørensen, Mathew Shaji Kavalekalam, Angeliki Xenaki, Jesper Bünsow Boldt, Mads Græsbøll Christensen |
Speech Commun. | 5 |
| 2017 | Estimation of multiple pitches in stereophonic mixtures using a codebook-based approachabstractIn this paper, a method for multi-pitch estimation of stereophonic mixtures of multiple harmonic signals is presented. The method is based on a signal model which takes the amplitude and delay panning parameters of the sources in a stereophonic mixture into account. Furthermore, the method is based on the extended invariance principle (EXIP), and a codebook of realistic amplitude vectors. For each fundamental frequency candidate in each of the sources, the amplitude estimates are mapped to entries in the codebook, and the pitch and model order are estimated jointly. The performance of the proposed method is evaluated using mixtures of real signals. Experiments show an increase in performance when knowledge about the panning parameters is utilized together with the codebook of magnitude amplitudes when compared to a state-of-the-art transcription method. Martin Weiss Hansen, Jesper Rindom Jensen, Mads Græsbøll Christensen |
ICASSP | 3 |
| 2017 | Harmonic minimum mean squared error filters for multichannel speech enhancementabstractMany state-of-the-art multichannel speech enhancement methods rely on second-order statistics of the desired speech signal, the noise signal, or both. Estimation of those are difficult in practice, resulting in a practical performance that is typically much lower than their potential theoretical performance. We propose two multichannel enhancement techniques that instead rely on a model for voiced speech. That is, the proposed methods are driven by the signals' fundamental frequencies, which may be accurately estimated even in noisy scenarios. The first method is designed independently of the microphone array geometry and source position, whereas these are utilized in the second approach. Thereby, we can investigate when to exploit such information in the case of localization errors and violations of the spatial assumptions. Numerical results show that the proposed method is able to outperform competing methods in terms of both output SNRs and PESQ scores. Jesper Rindom Jensen, Mads Græsbøll Christensen, Andreas Jakobsson |
ICASSP | 2 |
| 2017 | Model based binaural enhancement of voiced and unvoiced speechabstractThis paper deals with the enhancement of speech in presence of non-stationary babble noise. A binaural speech enhancement framework is proposed which takes into account both the voiced and unvoiced speech production model. The usage of this model in enhancement requires the Short term predictor (STP) parameters and the pitch information to be estimated. This paper uses a codebook based approach for estimating the STP parameters and a parametric binaural method is proposed for estimating the pitch parameters. Improvements in objective score are shown when using the voiced-unvoiced speech model in comparison to the conventional unvoiced speech model. Mathew Shaji Kavalekalam, Mads Græsbøll Christensen, Jesper Bünsow Boldt |
ICASSP | 2 |
| 2017 | Fast harmonic chirp summationabstractThe harmonic chirp signal model has only very recently been introduced for modelling approximately periodic signals with a time-varying fundamental frequency. A number of estimators for the parameters of this model have already been proposed, but they are either inaccurate, non-robust to noise, or very computationally intensive. In this paper, we propose a fast algorithm for the harmonic chirp summation method which has been demonstrated in the literature to be accurate and robust to noise. The proposed algorithm is orders of magnitudes faster than previous algorithms which is also demonstrated via timing studies. Jesper Kjær Nielsen, Tobias Lindstrøm Jensen, Jesper Rindom Jensen, Mads Græsbøll Christensen, Søren Holdt Jensen |
ICASSP | 4 |
| 2017 | Least 1-norm pole-zero modeling with sparse deconvolution for speech analysisabstractIn this paper, we present a speech analysis method based on sparse pole-zero modeling of speech. Instead of using the all-pole model to approximate the speech production filter, a pole-zero model is used for the combined effect of the vocal tract; radiation at the lips and the glottal pulse shape. Moreover, to consider the spiky excitation form of the pulse train during voiced speech, the modeling parameters and sparse residuals are estimated in an iterative fashion using a least 1-norm pole-zero with sparse deconvolution algorithm. Compared with the conventional two-stage least squares pole-zero, linear prediction and sparse linear prediction methods, experimental results show that the proposed speech analysis method has lower spectral distortion, higher reconstruction SNR and sparser residuals. Liming Shi, Jesper Rindom Jensen, Mads Græsbøll Christensen |
ICASSP | 3 |
| 2017 | Pitch-based non-intrusive objective intelligibility predictionabstractAutomatic adjustment of the hearing aid according to the intelligibility for the user in the environment could be beneficial. While most intelligibility metrics require a clean speech reference, i.e. intrusive methods, this is rarely available in real-life. This paper proposes a non-intrusive intelligibility metric in which a reconstruction of the clean speech is used in the established intrusive short-time objective intelligibility (STOI) metric. The reconstruction of the clean speech is based on pitch-features of the desired source using a spatio-temporal harmonic model. This model takes advantage of both the spatial and spectral separation of the desired source and interferers to reconstruct the clean signal. The simulations show a high correlation between the proposed pitch-based STOI (PB-STOI) and the original intrusive STOI and hence is promising for online processing of intelligibility. Charlotte Sørensen, Angeliki Xenaki, Jesper Bünsow Boldt, Mads Græsbøll Christensen |
ICASSP | 4 |
| 2017 | Distributed max-SINR speech enhancement with ad hoc microphone arraysabstractIn recent years, signal processing with ad hoc microphone arrays has attracted a lot of attention. Speech enhancement in noisy, interfered, and reverberant environments is one of the problems targeted by ad hoc microphone arrays. Most of the proposed solutions require knowledge of fingerprints, such as acoustic transfer functions, which may not be known as accurately as required in practical situations. In this paper, a distributed signal subspace filtering method is proposed which is not restricted to a special graph topology. Here, the maximum signal to interference-plus-noise ratio (max-SINR) criterion is used with the primal-dual method of multipliers for distributed filtering. The paper investigates the convergence of the algorithm in both synchronous and asynchronous schemes, and also discusses some practical pros and cons. The applicability of the proposed method is demonstrated by means of simulation results. Vincent Mohammad Tavakoli, Jesper Rindom Jensen, Richard Heusdens, Jacob Benesty, Mads Græsbøll Christensen |
ICASSP | 5 |
| 2017 | Dominant Distortion Classification for Pre-Processing of Vowels in Remote Biomedical Voice AnalysisabstractAdvances in speech signal analysis facilitate the development of techniques for remote biomedical voice assessment. However, the performance of these techniques is affected by noise and distortion in signals. In this paper, we focus on the vowel /a/ as the most widely-used voice signal for pathological voice assessments and investigate the impact of four major types of distortion that are commonly present during recording or transmission in voice analysis, namely: background noise, reverberation, clipping and compression, on Mel-frequency cepstral coefficients (MFCCs) - the most widely-used features in biomedical voice analysis. Then, we propose a new distortion classification approach to detect the most dominant distortion in such voice signals. The proposed method involves MFCCs as frame-level features and a support vector machine as classifier to detect the presence and type of distortion in frames of a given voice signal. Experimental results obtained from the healthy and Parkinson's voices show the effectiveness of the proposed approach in distortion detection and classification. Amir Hossein Poorjam, Jesper Rindom Jensen, Max A. Little, Mads Græsbøll Christensen |
INTERSPEECH | 4 |
| 2017 | Fast fundamental frequency estimation: Making a statistically efficient estimator computationally efficient
Jesper Kjær Nielsen, Tobias Lindstrøm Jensen, Jesper Rindom Jensen, Mads Græsbøll Christensen, Søren Holdt Jensen |
Signal Process. | 4 |
| 2016 | Experimental study of generalized subspace filters for the cocktail party situationabstractThis paper investigates the potential performance of generalized subspace filters for speech enhancement in cocktail party situations with very poor signal/noise ratio, e.g. down to -15 dB. Performance metrics output signal/noise ratio, signal/distortion ratio, speech quality rating and speech intelligibility rating are mapped as functions of two algorithm parameters, revealing clear trade-off options between noise, distortion and subjective performances and a recommended choice of trade-off. Given sufficiently good noise statistics, SNR improvements around 20 dB as well as PESQ quality and STOI intelligibility rating improvements exceeding 1.0 and 0.2 points respectively are found. This shows the potential of the method. Knud B. Christensen, Mads Græsbøll Christensen, Jesper Bünsow Boldt, Fredrik Gran |
ICASSP | 2 |
| 2016 | Variable span filters for speech enhancementabstractIn this work, we consider enhancement of multichannel speech recordings. Linear filtering and subspace approaches have been considered previously for solving the problem. The current linear filtering methods, although many variants exist, have limited control of noise reduction and speech distortion. Subspace approaches, on the other hand, can potentially yield better control by filtering in the eigen-domain, but traditionally these approaches have not been optimized explicitly for traditional noise reduction and signal distortion measures. Herein, we combine these approaches by deriving optimal filters using a joint diagonalization as a basis. This gives excellent control over the performance, as we can optimize for noise reduction or signal distortion performance. Results from real data experiments show that the proposed variable span filters can achieve better performance than existing filters. In terms of output SNR, the gain was more than 8 dB, and more than 0.1 in mean opinion score in the conducted experiments. Jesper Rindom Jensen, Jacob Benesty, Mads Græsbøll Christensen |
ICASSP | 3 |
| 2016 | DOA estimation of audio sources in reverberant environmentsabstractReverberation is well-known to have a detrimental impact on many localization methods for audio sources. We address this problem by imposing a model for the early reflections as well as a model for the audio source itself. Using these models, we propose two iterative localization methods that estimate the direction-of-arrival (DOA) of both the direct path of the audio source and the early reflections. In these methods, the contribution of the early reflections is essentially subtracted from the signal observations before localization of the direct path component, which may reduce the estimation bias. Our simulation results show that we can estimate the DOA of the desired signal more accurately with this procedure compared to state-of-the-art estimator in both synthetic and real data experiments with reverberation. Jesper Rindom Jensen, Jesper Kjær Nielsen, Richard Heusdens, Mads Græsbøll Christensen |
ICASSP | 4 |
| 2016 | Kalman filter for speech enhancement in cocktail party scenarios using a codebook-based approachabstractEnhancement of speech in non-stationary background noise is a challenging task, and conventional single channel speech enhancement algorithms have not been able to improve the speech intelligibility in such scenarios. The work proposed in this paper investigates a single channel Kalman filter based speech enhancement algorithm, whose parameters are estimated using a codebook based approach. The results indicate that the enhancement algorithm is able to improve the speech intelligibility and quality according to objective measures. Moreover, we investigate the effects of utilizing a speaker specific trained codebook over a generic speech codebook in relation to the performance of the speech enhancement system. Mathew Shaji Kavalekalam, Mads Græsbøll Christensen, Fredrik Gran, Jesper Bünsow Boldt |
ICASSP | 2 |
| 2016 | Fast and statistically efficient fundamental frequency estimationabstractFundamental frequency estimation is a very important task in many applications involving periodic signals. For computational reasons, fast autocorrelation-based estimation methods are often used despite parametric estimation methods having superior estimation accuracy. However, these parametric methods are much more costly to run. In this paper, we propose an algorithm which significantly reduces the computational cost of an accurate maximum likelihood-based estimator for real-valued data. The computational cost is reduced by exploiting the matrix structure of the problem and by using a recursive solver. Via benchmarks, we demonstrate that the computation time is reduced by approximately two orders of magnitude. The proposed fast algorithm is available for download online. Jesper Kjær Nielsen, Tobias Lindstrøm Jensen, Jesper Rindom Jensen, Mads Græsbøll Christensen, Søren Holdt Jensen |
ICASSP | 4 |
| 2016 | A partitioned approach to signal separation with microphone ad hoc arraysabstractIn this paper, a blind algorithm is proposed for speech enhancement in multi-speaker scenarios, in which interference rejection is the main objective. Here, the ad hoc array is broken into microphone duples which are used to partition the array into local sub-arrays. The core algorithm takes advantage of differences in signal structure in each duple. A geometric mean filter is then used to merge the output signals obtained with different duples, and to form a global broadband maximum signal-to-interference ratio (SIR) enhancement apparatus. The resulting filter outputs are enhanced acoustic signals in terms of SIR, as shown with experiments. Vincent Mohammad Tavakoli, Jesper Rindom Jensen, Jacob Benesty, Mads Græsbøll Christensen |
ICASSP | 4 |
| 2016 | Fast algorithms for high-order sparse linear prediction with applications to speech processing
Tobias Lindstrøm Jensen, Daniele Giacobello, Toon van Waterschoot, Mads Græsbøll Christensen |
Speech Commun. | 4 |
| 2016 | Noise Reduction with Optimal Variable Span Linear FiltersabstractIn this paper, the problem of noise reduction is addressed as a linear filtering problem in a novel way by using concepts from subspace-based enhancement methods, resulting in variable span linear filters. This is done by forming the filter coefficients as linear combinations of a number of eigenvectors stemming from a joint diagonalization of the covariance matrices of the signal of interest and the noise. The resulting filters are flexible in that it is possible to trade off distortion of the desired signal for improved noise reduction. This tradeoff is controlled by the number of eigenvectors included in forming the filter. Using these concepts, a number of different filter designs are considered, like minimum distortion, Wiener, maximum SNR, and tradeoff filters. Interestingly, all these can be expressed as special cases of variable span filters. We also derive expressions for the speech distortion and noise reduction of the various filter designs. Moreover, we consider an alternative approach, wherein the filter is designed for extracting an estimate of the noise signal, which can then be extracted from the observed signals, which is referred to as the indirect approach. Simulations demonstrate the advantages and properties of the variable span filter designs, and their potential performance gain compared to widely used speech enhancement methods. Jesper Rindom Jensen, Jacob Benesty, Mads Græsbøll Christensen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Computationally Efficient and Noise Robust DOA and Pitch EstimationabstractMany natural signals, such as voiced speech and some musical instruments, are approximately periodic over short intervals. These signals are often described in mathematics by the sum of sinusoids (harmonics) with frequencies that are proportional to the fundamental frequency, or pitch. In sensor (microphone) array signal processing, the periodic signals are estimated from spatio-temporal samples regarding to the direction of arrival (DOA) of the signal of interest. In this paper, we consider the problem of pitch and DOA estimation of quasi-periodic audio signals. In real-life scenarios, recorded signals are often contaminated by different types of noise, which challenges the assumption of white Gaussian noise in most state-of-the-art methods. We establish filtering methods based on noise statistics to apply to nonparametric spectral and spatial parameter estimates of the harmonics. We design minimum variance solutions with distortionless constraints to estimate the pitch from the frequency estimates, and to estimate the DOA from multichannel phase estimates of the harmonics. Applying this filtering method as the sum of weighted frequency and DOA estimates of the harmonics, we also design a joint DOA and pitch estimator. In white Gaussian noise, we derive even more computationally efficient solutions which are designed using the narrowband power spectrum of the harmonics. Numerical results reveal the performance of the estimators in colored noise compared with the Cramér-Rao lower bound. Experiments on real-life signals indicate the applicability of the methods in practical low local signal-to-noise ratios. Sam Karimian-Azari, Jesper Rindom Jensen, Mads Græsbøll Christensen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Enhancement and Noise Statistics Estimation for Non-Stationary Voiced SpeechabstractIn this paper, single channel speech enhancement in the time domain is considered. We address the problem of modelling non-stationary speech by describing the voiced speech parts by a harmonic linear chirp model instead of using the traditional harmonic model. This means that the speech signal is not assumed stationary, instead the fundamental frequency can vary linearly within each frame. The linearly constrained minimum variance (LCMV) filter and the amplitude and phase estimation (APES) filter are derived in this framework and compared to the harmonic versions of the same filters. It is shown through simulations on synthetic and speech signals, that the chirp versions of the filters perform better than their harmonic counterparts in terms of output signal-to-noise ratio (SNR) and signal reduction factor. For synthetic signals, the output SNR for the harmonic chirp APES based filter is increased 3 dB compared to the harmonic APES based filter at an input SNR of 10 dB, and at the same time the signal reduction factor is decreased. For speech signals, the increase is 1.5 dB along with a decrease in the signal reduction factor of 0.7. As an implicit part of the APES filter, a noise covariance matrix estimate is obtained. We suggest using this estimate in combination with other filters such as the Wiener filter. The performance of the Wiener filter and LCMV filter are compared using the APES noise covariance matrix estimate and a power spectral density (PSD) based noise covariance matrix estimate. It is shown that the APES covariance matrix works well in combination with the Wiener filter, and the PSD based covariance matrix works well in combination with the LCMV filter. Sidsel Marie Nørholm, Jesper Rindom Jensen, Mads Græsbøll Christensen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Instantaneous Fundamental Frequency Estimation With Optimal Segmentation for Nonstationary Voiced SpeechabstractIn speech processing, the speech is often considered stationary within segments of 20-30 ms even though it is well known not to be true. In this paper, we take the nonstationarity of voiced speech into account by using a linear chirp model to describe the speech signal. We propose a maximum likelihood estimator of the fundamental frequency and chirp rate of this model, and show that it reaches the Cramer-Rao lower bound. Since the speech varies over time, a fixed segment length is not optimal, and we propose making a segmentation of the signal based on the maximum a posteriori criterion. Using this segmentation method, the segments are on average longer for the chirp model compared to the traditional harmonic model. For the signal under test, the average segment length is 24.4 and 17.1 ms for the chirp model and traditional harmonic model, respectively. This suggests a better fit of the chirp model than the harmonic model to the speech signal. The methods are based on an assumption of white Gaussian noise, and, therefore, two prewhitening filters are also proposed. Sidsel Marie Nørholm, Jesper Rindom Jensen, Mads Græsbøll Christensen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | A Framework for Speech Enhancement With Ad Hoc Microphone ArraysabstractSpeech enhancement is vital for improved listening practices. Ad hoc microphone arrays are promising assets for this purpose. Most well-established enhancement techniques with conventional arrays can be adapted into ad hoc scenarios. Despite recent efforts to introduce various ad hoc speech enhancement apparatus, a common framework for integration of conventional methods into this new scheme is still missing. This paper establishes such an abstraction based on inter and intra subarray speech coherencies. Along with measures for signal quality at the input of subarrays, a measure of coherency is proposed both for subarray selection in local enhancement approaches, and also for selecting a proper global reference when more than one subarray are used. Proposed methods within this framework are evaluated with regard to quantitative and qualitative measures, including array gains, the speech distortion ratio, the PESQ measure, and the STOI intelligibility measure. Major findings in this work are the observed changes in the superiority of different methods for certain conditions. When perceptual quality or intelligibility of the speech are the ultimate goals, there are turning points where the MVDR and the LCMV are superior to Wiener-based methods. Also, for certain scenarios, local approaches may be preferred to global ones. Vincent Mohammad Tavakoli, Jesper Rindom Jensen, Mads Græsbøll Christensen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | Preference learning with evolutionary Multivariate Adaptive Regression Spline modelabstractThis paper introduces a novel approach for pairwise preference learning through combining an evolutionary method with Multivariate Adaptive Regression Spline (MARS). Collecting users' feedback through pairwise preferences is recommended over other ranking approaches as this method is more appealing for human decision making. Learning models from pairwise preference data is however an NP-hard problem. Therefore, constructing models that can effectively learn such data is a challenging task. Models are usually constructed with accuracy being the most important factor. Another vitally important aspect that is usually given less attention is expressiveness, i.e. how easy it is to explain the relationship between the model input and output. Most machine learning techniques are focused either on performance or on expressiveness. This paper employ MARS models which have the advantage of being a powerful method for function approximation as well as being relatively easy to interpret. MARS models are evolved based on their efficiency in learning pairwise data. The method is tested on two datasets that collectively provide pairwise preference data of five cognitive states expressed by users. The method is analysed in terms of the performance, expressiveness and complexity and showed promising results in all aspects. Mohamed Abou-Zleikha, Noor Shaker, Mads Græsbøll Christensen |
CEC | 3 |
| 2015 | Pitch and TDOA-based localization of acoustic sources with distributed arraysabstractIn this paper, a method for acoustic source localization using distributed microphone arrays based on time-differences of arrival (TDOAs) is presented. The TDOAs are used to estimate the location of an acoustic source using a recently proposed method, based on a 4D parameter space defined by the 3D location of the source, and the TDOAs. The performance of the proposed method for acoustic source localization is compared to the performance of a method based on generalized cross-correlation with phase transform (GCC-PHAT) using synthetic and speech signals with varying source position. Results show a decrease in the error of the estimated position when the proposed method is used. Martin Weiss Hansen, Jesper Rindom Jensen, Mads Græsbøll Christensen |
ICASSP | 3 |
| 2015 | A joint audio-visual approach to audio localizationabstractLocalization of audio sources is an important research problem, e.g., to facilitate noise reduction. In the recent years, the problem has been tackled using distributed microphone arrays (DMA). A common approach is to apply direction-of-arrival (DOA) estimation on each array (denoted as nodes), and then map the DOA estimates to a location. In practice, however, the individual nodes contain few microphones, limiting the DOA estimation accuracy and, thereby, also the localization performance. We investigate a new approach, where range estimates are also obtained and utilized from each node, e.g., using time-of-flight cameras. Moreover, we propose an optimal method for weighting such DOA and range information for audio localization. Our experiments on both synthetic and real data show that there is a clear, potential advantage of using the joint audio-visual localization framework. Jesper Rindom Jensen, Mads Græsbøll Christensen |
ICASSP | 2 |
| 2015 | On frequency domain models for TDOA estimationabstractTime-difference-of-arrival (TDOA) estimation is an important problem in many microphone signal processing applications. Traditionally, this problem is solved by using a cross-correlation method, but in this paper we show that the cross-correlation method is actually a restricted special case of a much more general method. In this connection, we establish the conditions under which the crosscorrelation method is a statistically efficient estimator. One of the conditions is that the source signal is periodic with a known fundamental frequency of 2π/N radians per sample, where N is the number of data points, and a known number of harmonics. The more general method only relies on that the source signal is periodic and is, therefore, able to outperform the cross-correlation method in terms of estimation accuracy on both synthetic data and artificially delayed speech data. The simulation code is available online. Jesper Rindom Jensen, Jesper Kjær Nielsen, Mads Græsbøll Christensen, Søren Holdt Jensen |
ICASSP | 3 |
| 2015 | Pitch estimation and tracking with harmonic emphasis on the acoustic spectrumabstractIn this paper, we use unconstrained frequency estimates (UFEs) from a noisy harmonic signal and propose two methods to estimate and track the pitch over time. We assume that the UFEs are multivariate-normally-distributed random variables, and derive a maximum likelihood (ML) pitch estimator by maximizing the likelihood of the UFEs over short time-intervals. As the main contribution of this paper, we propose two state-space representations to model the pitch continuity, and, accordingly, we propose two Bayesian methods, namely a hidden Markov model and a Kalman filter. These methods are designed to optimally use the correlations in the consecutive pitch values, where the past pitch estimates are used to recursively update the prior distribution for the pitch variable. We perform experiments using synthetic data as well as a noisy speech recording, and show that the Bayesian methods provide more accurate estimates than the corresponding ML methods. Sam Karimian-Azari, Nasser Mohammadiha, Jesper Rindom Jensen, Mads Græsbøll Christensen |
ICASSP | 4 |
| 2015 | Pseudo-coherence-based MVDR beamformer for speech enhancement with ad hoc microphone arraysabstractSpeech enhancement with distributed arrays has been met with various methods. On the one hand, data independent methods require information about the position of sensors, so they are not suitable for dynamic geometries. On the other hand, Wiener-based methods cannot assure a distortionless output. This paper proposes minimum variance distortionless response filtering based on multichannel pseudo-coherence for speech enhancement with ad hoc microphone arrays. This method requires neither position information nor control of the trade-off used in the distortion weighted methods. Furthermore, certain performance criteria are derived in terms of the pseudo-coherence vector, and the method is compared with the multichannel Wiener filter. Evaluation shows the suitability of the proposed method in terms of noise reduction with minimum distortion in ad hoc scenarios. Vincent Mohammad Tavakoli, Jesper Rindom Jensen, Mads Græsbøll Christensen, Jacob Benesty |
ICASSP | 3 |
| 2015 | Enhancement of non-stationary speech using harmonic chirp filtersabstractIn this paper, the issue of single channel speech enhancement of non-stationary voiced speech is addressed. The non-stationarity of speech is well known, but state of the art speech enhancement methods assume stationarity within frames of 20–30 ms. We derive optimal distortionless filters that take the non-stationarity nature of voiced speech into account via linear constraints. This is facilitated by imposing a harmonic chirp model on the speech signal. As an implicit part of the filter design, the noise statistics are also estimated based on the observed signal and parameters of the harmonic chirp model. Simulations on real speech show that the chirp based filters perform better than their harmonic counterparts. Further, it is seen that the gain of using the chirp model increases when the estimated chirp parameter is big corresponding to periods in the signal where the instantaneous fundamental frequency changes fast. Sidsel Marie Nørholm, Jesper Rindom Jensen, Mads Græsbøll Christensen |
INTERSPEECH | 3 |
| 2015 | Least squares estimate of the initial phases in STFT based speech enhancementabstractIn this paper, we consider single-channel speech enhancement in the short time Fourier transform (STFT) domain. We suggest to improve an STFT phase estimate by estimating the initial phases. The method is based on the harmonic model and a model for the phase evolution over time. The initial phases are estimated by setting up a least squares problem between the noisy phase and the model for phase evolution. Simulations on synthetic and speech signals show a decreased error on the phase when an estimate of the initial phase is included compared to using the noisy phase as an initialisation. The error on the phase is decreased at input SNRs from -10 to 10 dB. Reconstructing the signal using the clean amplitude, the mean squared error is decreased and the PESQ score is increased. Sidsel Marie Nørholm, Martin Krawczyk-Becker, Timo Gerkmann, Steven van de Par, Jesper Rindom Jensen, Mads Græsbøll Christensen |
INTERSPEECH | 6 |
| 2015 | Multi-pitch estimation exploiting block sparsity
Stefan Ingi Adalbjornsson, Andreas Jakobsson, Mads Græsbøll Christensen |
Signal Process. | 3 |
| 2015 | Joint Spatio-Temporal Filtering Methods for DOA and Fundamental Frequency EstimationabstractIn this paper, spatio-temporal filtering methods are proposed for estimating the direction-of-arrival (DOA) and fundamental frequency of periodic signals, like those produced by the speech production system and many musical instruments using microphone arrays. This topic has quite recently received some attention in the community and is quite promising for several applications. The proposed methods are based on optimal, adaptive filters that leave the desired signal, having a certain DOA and fundamental frequency, undistorted and suppress everything else. The filtering methods simultaneously operate in space and time, whereby it is possible resolve cases that are otherwise problematic for pitch estimators or DOA estimators based on beamforming. Several special cases and improvements are considered, including a method for estimating the covariance matrix based on the recently proposed iterative adaptive approach (IAA). Experiments demonstrate the improved performance of the proposed methods under adverse conditions compared to the state of the art using both synthetic signals and real signals, as well as illustrate the properties of the methods and the filters. Jesper Rindom Jensen, Mads Græsbøll Christensen, Jacob Benesty, Søren Holdt Jensen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Fundamental frequency and model order estimation using spatial filteringabstractIn signal processing applications of harmonic-structured signals, estimates of the fundamental frequency and number of harmonics are often necessary. In real scenarios, a desired signal is contaminated by different levels of noise and interferers, which complicate the estimation of the signal parameters. In this paper, we present an estimation procedure for harmonic-structured signals in situations with strong interference using spatial filtering, or beamforming. We jointly estimate the fundamental frequency and the constrained model order through the output of the beamformers. Besides that, we extend this procedure to account for inharmonicity using unconstrained model order estimation. The simulations show that beamforming improves the performance of the joint estimates of fundamental frequency and the number of harmonics in low signal to interference (SIR) levels, and an experiment on a trumpet signal show the applicability on real signals. Sam Karimian-Azari, Jesper Rindom Jensen, Mads Græsbøll Christensen |
ICASSP | 3 |
| 2014 | Joint sparsity and frequency estimation for spectral compressive sensingabstractParameter estimation from compressively sensed signals has recently received some attention. We here also consider this problem in the context of frequency sparse signals which are encountered in many application. Existing methods perform the estimation using finite dictionaries or incorporate various interpolation techniques to estimate the continuous frequency parameters. In this paper, we show that solving the problem in a probabilistic framework instead produces an asymptotically efficient estimator which outperforms existing methods in terms of estimation accuracy while still having a low computational complexity. Moreover, the proposed algorithm is also able to make inference about the sparsity level of the measured signal. The simulation code is available online. Jesper Kjær Nielsen, Mads Græsbøll Christensen, Søren Holdt Jensen |
ICASSP | 2 |
| 2014 | Model selection and comparison for independents sinusoidsabstractIn the signal processing literature, many methods have been proposed for estimating the number of sinusoidal basis functions from a noisy data set. The most popular method is the asymptotic MAP criterion, which is sometimes also referred to as the BIC. In this paper, we extend and improve this method by considering the problem in a full Bayesian framework instead of the approximate formulation, on which the asymptotic MAP criterion is based. This leads to a new model selection and comparison method, the lp-BIC, whose computational complexity is of the same order as the asymptotic MAP criterion. Through simulations, we demonstrate that the lp-BIC outperforms the asymptotic MAP criterion and other state of the art methods in terms of model selection, de-noising and prediction performance. The simulation code is available online. Jesper Kjær Nielsen, Mads Græsbøll Christensen, Søren Holdt Jensen |
ICASSP | 2 |
| 2014 | Noise reduction in the time domain using joint diagonalizationabstractA new filter design based on joint diagonalization of the clean speech and noise covariance matrices is proposed. First, an estimate of the noise is found by filtering the observed signal. The filter for this is generated by a weighted sum of the eigenvectors from the joint diagonalization. Second, an estimate of the desired signal is found by subtraction of the noise estimate from the observed signal. The filter can be designed to obtain a desired trade-off between noise reduction and signal distortion, depending on the number of eigenvectors included in the filter design. This is explored through simulations using a speech signal corrupted by car noise, and the results confirm that the output signal-to-noise ratio and speech distortion index both increase when more eigenvectors are included in the filter design. Sidsel Marie Nørholm, Jacob Benesty, Jesper Rindom Jensen, Mads Græsbøll Christensen |
ICASSP | 4 |
| 2014 | Stable 1-Norm Error Minimization Based Linear Predictors for Speech ModelingabstractIn linear prediction of speech, the 1-norm error minimization criterion has been shown to provide a valid alternative to the 2-norm minimization criterion. However, unlike 2-norm minimization, 1-norm minimization does not guarantee the stability of the corresponding all-pole filter and can generate saturations when this is used to synthesize speech. In this paper, we introduce two new methods to obtain intrinsically stable predictors with the 1-norm minimization. The first method is based on constraining the roots of the predictor to lie within the unit circle by reducing the numerical range of the shift operator associated with the particular prediction problem considered. The second method uses the alternative Cauchy bound to impose a convex constraint on the predictor in the 1-norm error minimization. These methods are compared with two existing methods: the Burg method, based on the 1-norm minimization of the forward and backward prediction error, and the iteratively reweighted 2-norm minimization known to converge to the 1-norm minimization with an appropriate selection of weights. The evaluation gives proof of the effectiveness of the new methods, performing as well as unconstrained 1-norm based linear prediction for modeling and coding of speech. Daniele Giacobello, Mads Græsbøll Christensen, Tobias Lindstrøm Jensen, Manohar N. Murthi, Søren Holdt Jensen, Marc Moonen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2013 | Estimating multiple pitches using block sparsityabstractWe study the problem of estimating the fundamental frequencies of a signal containing multiple harmonically related sinusoidal signals using a novel block sparsity representation of the signal model. An efficient algorithm for solving the resulting optimization is devised exploiting an alternating directions method of multipliers (ADMM) formulation of the problem. The superiority of the proposed method, as compared to earlier methods, is demonstrated using both simulated and measured audio signals. Stefan Ingi Adalbjornsson, Andreas Jakobsson, Mads Græsbøll Christensen |
ICASSP | 3 |
| 2013 | An exact subspace method for fundamental frequency estimationabstractIn this paper, an exact subspace method for fundamental frequency estimation is presented. The method is based on the principles of the MUSIC algorithm, wherein the orthogonality between the signal and and noise subspace is exploited. Unlike the original MUSIC algorithm, the new method uses an exact measure of the angles between the subspaces. This makes a difference, for example, when the fundamental frequency is low, for real signals, or when the number of samples is low. In Monte Carlo simulations, the performance of the new method is compared to a number of state-of-the-art methods and is demonstrated to lead to improvements in certain, critical cases. Moreover, it is demonstrated on a speech signal that the method can be applied to speech signals and is robust towards noise. Mads Græsbøll Christensen |
ICASSP | 1 |
| 2013 | Multichannel signal enhancement using non-causal, time-domain filtersabstractIn the vast amount of time-domain filtering methods for speech enhancement, the filters are designed to be causal. Recently, however, it was shown that the noise reduction and signal distortion capabilities of such single-channel filters can be improved by allowing the filters to be non-causal. While non-causal filters require knowledge of the future, they can be implemented in practice by introducing a short delay. In this paper, we generalize the idea of exploiting non-causality in optimal filter designs to the multichannel scenario. More specifically, a set of optimal, non-causal, multichannel filters for enhancement based on an orthogonal decomposition is proposed. The evaluation shows that there is a potential gain in noise reduction and signal distortion by introducing non-causality. Moreover, experiments on real-life speech show that we can improve the perceptual quality. Jesper Rindom Jensen, Mads Græsbøll Christensen, Jacob Benesty |
ICASSP | 2 |
| 2013 | Statistically efficient methods for pitch and DOA estimationabstractTraditionally, direction-of-arrival (DOA) and pitch estimation of multichannel, periodic sources have been considered as two separate problems. Separate estimation may render the task of resolving sources with similar DOA or pitch impossible, and it may decrease the estimation accuracy. Therefore, it was recently considered to estimate the DOA and pitch jointly. In this paper, we propose two novel methods for DOA and pitch estimation. They both yield maximum-likelihood estimates in white Gaussian noise scenarios, where the SNR may be different across channels, as opposed to state-of-the-art methods. The first method is a joint estimator, whereas the latter use a cascaded approach, but with a much lower computational complexity. The simulation results confirm that the proposed methods outperform state-of-the-art methods in terms of estimation accuracy in both synthetic and real-life signal scenarios. Jesper Rindom Jensen, Mads Græsbøll Christensen, Søren Holdt Jensen |
ICASSP | 2 |
| 2013 | Real-time implementations of sparse linear prediction for speech processingabstractEmploying sparsity criteria in linear prediction of speech has been proven successful for several analysis and coding purposes. However, sparse linear prediction comes at the expenses of a much higher computational burden and numerical sensitivity compared to the traditional minimum variance approach. This makes sparse linear prediction difficult to deploy in real-time systems. In this paper, we present a step towards real-time implementation of the sparse linear prediction problem using hand-tailored interior-point methods. Using compiled implementations the sparse linear prediction problems corresponding to a frame size of 20ms can be solved on a standard PC in approximately 2ms and orders faster than with general purpose software. Tobias Lindstrøm Jensen, Daniele Giacobello, Mads Græsbøll Christensen, Søren Holdt Jensen, Marc Moonen |
ICASSP | 3 |
| 2013 | Bayesian model comparison and the BIC for regression modelsabstractIn the signal processing literature, many methods have been proposed for solving the important model comparison and selection problem. However, most of these methods only find the most likely model or only work well under particular circumstances such as a large number of data points or a high signal-to-noise ratio (SNR). One of the most successful classes of methods is the Bayesian information criteria (BIC) and in this paper, we extend some of the recent work on the BIC. In particular, we develop methods in a full Bayesian framework which work well across a large/small number of data points and high/low SNR for either real- or complex-valued data originating from a regression model. Aside from selecting the most probable model, these rules can also be used for model averaging as they assign a probability to each candidate model. Through simulations on a polynomial trend model, we demonstrate that the proposed rules outperform other rules in terms of detecting the true model order, de-noising the noisy signal, and making predictions of unobserved data points. The simulation code is available online. Jesper Kjær Nielsen, Mads Græsbøll Christensen, Søren Holdt Jensen |
ICASSP | 2 |
| 2013 | Joint DOA and fundamental frequency estimation based on relaxed iterative adaptive approach and optimal filteringabstractIn this work, the problem of joint direction-of-arrival and fundamental frequency estimation for multi-channel harmonic sinusoidal signals is addressed. Different from the conventional optimal filtering method, we estimate the covariance matrix with the 2-D iterative adaptive approach, which is based on a single snapshot. In addition, to improve the estimation accuracy for the off-grid sources, a relaxation technique is utilized. Then, joint estimation is conducted on this covariance matrix estimate with the optimal filtering method. As a result, the relaxed iterative adaptive approach - optimal filtering method is devised. Statistical evaluation with synthetic signals shows the accurate performance of the proposed method compared with the Cramér-Rao lower bound. Zhenhua Zhou, Mads Græsbøll Christensen, Jesper Rindom Jensen, Hing-Cheung So |
ICASSP | 2 |
| 2013 | Improved prediction error filters for adaptive feedback cancellation in hearing aids
Kim Ngo, Toon van Waterschoot, Mads Græsbøll Christensen, Marc Moonen, Søren Holdt Jensen |
Signal Process. | 3 |
| 2013 | Accurate Estimation of Low Fundamental Frequencies From Real-Valued MeasurementsabstractIn this paper, the difficult problem of estimating low fundamental frequencies from real-valued measurements is addressed. The methods commonly employed do not take the phenomena encountered in this scenario into account and thus fail to deliver accurate estimates. The reason for this is that they employ asymptotic approximations that are violated when the harmonics are not well-separated in frequency, something that happens when the observed signal is real-valued and the fundamental frequency is low. To mitigate this, we analyze the problem and present some exact fundamental frequency estimators that are aimed at solving this problem. These estimators are based on the principles of nonlinear least-squares, harmonic fitting, optimal filtering, subspace orthogonality, and shift-invariance, and they all reduce to already published methods for a high number of observations. In experiments, the methods are compared and the increased accuracy obtained by avoiding asymptotic approximations is demonstrated. Mads Græsbøll Christensen |
IEEE Trans. Speech Audio Process. | 1 |
| 2013 | A Class of Optimal Rectangular Filtering Matrices for Single-Channel Signal Enhancement in the Time DomainabstractIn this paper, we introduce a new class of optimal rectangular filtering matrices for single-channel speech enhancement. The new class of filters exploits the fact that the dimension of the signal subspace is lower than that of the full space. By doing this, extra degrees of freedom in the filters, that are otherwise reserved for preserving the signal subspace, can be used for achieving an improved output signal-to-noise ratio (SNR). Moreover, the filters allow for explicit control of the tradeoff between noise reduction and speech distortion via the chosen rank of the signal subspace. An interesting aspect is that the framework in which the filters are derived unifies the ideas of optimal filtering and subspace methods. A number of different optimal filter designs are derived in this framework, and the properties and performance of these are studied using both synthetic, periodic signals and real signals. The results show a number of interesting things. Firstly, they show how speech distortion can be traded for noise reduction and vice versa in a seamless manner. Moreover, the introduced filter designs are capable of achieving both the upper and lower bounds for the output SNR via the choice of a single parameter. Jesper Rindom Jensen, Jacob Benesty, Mads Græsbøll Christensen, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2013 | Nonlinear Least Squares Methods for Joint DOA and Pitch EstimationabstractIn this paper, we consider the problem of joint direction-of-arrival (DOA) and fundamental frequency estimation. Joint estimation enables robust estimation of these parameters in multi-source scenarios where separate estimators may fail. First, we derive the exact and asymptotic Cramér-Rao bounds for the joint estimation problem. Then, we propose a nonlinear least squares (NLS) and an approximate NLS (aNLS) estimator for joint DOA and fundamental frequency estimation. The proposed estimators are maximum likelihood estimators when: 1) the noise is white Gaussian, 2) the environment is anechoic, and 3) the source of interest is in the far-field. Otherwise, the methods still approximately yield maximum likelihood estimates. Simulations on synthetic data show that the proposed methods have similar or better performance than state-of-the-art methods for DOA and fundamental frequency estimation. Moreover, simulations on real-life data indicate that the NLS and aNLS methods are applicable even when reverberation is present and the noise is not white Gaussian. Jesper Rindom Jensen, Mads Græsbøll Christensen, Søren Holdt Jensen |
IEEE Trans. Speech Audio Process. | 2 |
| 2013 | Default Bayesian Estimation of the Fundamental FrequencyabstractJoint fundamental frequency and model order estimation is an important problem in several applications. In this paper, a default estimation algorithm based on a minimum of prior information is presented. The algorithm is developed in a Bayesian framework, and it can be applied to both real- and complex-valued discrete-time signals which may have missing samples or may have been sampled at a non-uniform sampling frequency. The observation model and prior distributions corresponding to the prior information are derived in a consistent fashion using maximum entropy and invariance arguments. Moreover, several approximations of the posterior distributions on the fundamental frequency and the model order are derived, and one of the state-of-the-art joint fundamental frequency and model order estimators is demonstrated to be a special case of one of these approximations. The performance of the approximations are evaluated in a small-scale simulation study on both synthetic and real world signals. The simulations indicate that the proposed algorithm yields more accurate results than previous algorithms. The simulation code is available online. Jesper Kjær Nielsen, Mads Græsbøll Christensen, Søren Holdt Jensen |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | A method for low-delay pitch tracking and smoothingabstractIn this paper, a new method for pitch tracking is presented. The method is comprised of two steps. In the first step, accurate pitch estimates are obtained on a sample-by-sample basis by updates of the signal statistics with an exponential forgetting factor and subsequent numerical optimization. In the second step, a Kalman filter is used to smooth the estimates and separate the pitch into a slowly varying component and a rapidly varying component. The former represents the mean pitch while the latter represents vibrato, slides and other fast changes. The method is intended for use in applications that require fast and sample-by-sample estimates, like tuners for musical instruments, transcription tasks requiring details like vibrato, and real-time tracking of voiced speech. Mads Græsbøll Christensen |
ICASSP | 1 |
| 2012 | Multi-channel maximum likelihood pitch estimationabstractIn this paper, a method for multi-channel pitch estimation is proposed. The method is a maximum likelihood estimator and is based on a parametric model where the signals in the various channels share the same fundamental frequency but can have different amplitudes, phases, and noise characteristics. This essentially means that the model allows for different conditions in the various channels, like different signal-to-noise ratios, microphone characteristics and reverberation. Moreover, the method does not assume that a certain array structure is used but rather relies on a more general model and is hence suited for a large class of problems. Simulations with real signals shows that the method outperforms a state-of-the-art multi-channel method in terms of gross error rate. Mads Græsbøll Christensen |
ICASSP | 1 |
| 2012 | Subjective and objective quality assessment of single-channel speech separation algorithmsabstractPrevious studies on performance evaluation of single-channel speech separation (SCSS) algorithms mostly focused on automatic speech recognition (ASR) accuracy as their performance measure. Assessing the separated signals by different metrics other than this has the benefit that the results are expected to carry on to other applications beyond ASR. In this paper, in addition to conventional speech quality metrics (PESQ and SNRloss), we also evaluate the separation systems output using different source separation metrics: blind source separation evaluation (BSS EVAL) and perceptual evaluation methods for audio source separation (PEASS) measures. In our experiments, we apply these measures on the separated signals obtained by two well-known systems in the SCSS challenge to assess the objective and subjective quality of their output signals. Comparing subjective and objective measurements shows that PESQ and PEASS quality metrics predict well the subjective quality of separated signals obtained by the separation systems. From the results it is observed that the short-time objective intelligibility (STOI) measure predict the speech intelligibility results. Pejman Mowlaee, Rahim Saeidi, Mads Græsbøll Christensen, Rainer Martin 0001 |
ICASSP | 3 |
| 2012 | On compressed sensing and the estimation of continuous parameters from noisy observationsabstractCompressed sensing (CS) has in recent years become a very popular way of sampling sparse signals. This sparsity is measured with respect to some known dictionary consisting of a finite number of atoms. Most models for real world signals, however, are parametrised by continuous parameters corresponding to a dictionary with an infinite number of atoms. Examples of such parameters are the temporal and spatial frequency. In this paper, we analyse how CS affects the estimation performance of any unbiased estimator when we assume such infinite dictionaries. We base our analysis on the Cramer-Rao lower bound (CRLB) which is frequently used for benchmarking the estimation accuracy of unbiased estimators. For the popular sensing matrices such as the Gaussian sensing matrix, our analysis shows that compressed sensing on average degrades the estimation accuracy by at least the down-sample factor. Jesper Kjær Nielsen, Mads Græsbøll Christensen, Søren Holdt Jensen |
ICASSP | 2 |
| 2012 | An approximate Bayesian fundamental frequency estimatorabstractJoint fundamental frequency and model order estimation is an important problem in several applications such as speech and music processing. In this paper, we develop an approximate estimation algorithm of these quantities using Bayesian inference. The inference about the fundamental frequency and the model order is based on a probability model which corresponds to a minimum of prior information. From this probability model, we give the exact posterior distributions on the fundamental frequency and the model order, and we also present analytical approximations of these distributions which lower the computational load of the algorithm. By use of simulations on both a synthetic signal and a speech signal, the algorithm is demonstrated to be more accurate than a state-of-the-art maximum likelihood-based method. Jesper Kjær Nielsen, Mads Græsbøll Christensen, Søren Holdt Jensen |
ICASSP | 2 |
| 2012 | New Entries to the SPL EDICS for Audio and Acoustic Signal ProcessingabstractThis letter describes some of the new entries to the Signal Processing Letters (SPL) Editor's Information Classification Scheme (EDICS) for the topic Audio and Acoustic Signal Processing. Mads Græsbøll Christensen, Rudolf Rabenstein |
IEEE Signal Process. Lett. | 1 |
| 2012 | Sparse Linear Prediction and Its Applications to Speech ProcessingabstractThe aim of this paper is to provide an overview of Sparse Linear Prediction, a set of speech processing tools created by introducing sparsity constraints into the linear prediction framework. These tools have shown to be effective in several issues related to modeling and coding of speech signals. For speech analysis, we provide predictors that are accurate in modeling the speech production process and overcome problems related to traditional linear prediction. In particular, the predictors obtained offer a more effective decoupling of the vocal tract transfer function and its underlying excitation, making it a very efficient method for the analysis of voiced speech. For speech coding, we provide predictors that shape the residual according to the characteristics of the sparse encoding techniques resulting in more straightforward coding strategies. Furthermore, encouraged by the promising application of compressed sensing in signal compression, we investigate its formulation and application to sparse linear predictive coding. The proposed estimators are all solutions to convex optimization problems, which can be solved efficiently and reliably using, e.g., interior-point methods. Extensive experimental results are provided to support the effectiveness of the proposed methods, showing the improvements over traditional linear prediction in both speech analysis and coding. Daniele Giacobello, Mads Græsbøll Christensen, Manohar N. Murthi, Søren Holdt Jensen, Marc Moonen |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Non-Causal Time-Domain Filters for Single-Channel Noise ReductionabstractIn many existing time-domain filtering methods for noise reduction in, e.g., speech processing, the filters are causal. Such causal filters can be implemented directly in practice. However, it is possible to improve the performance of such noise reduction filtering methods in terms of both noise suppression and signal distortion by allowing the filters to be non-causal. Non-causal time-domain filters require knowledge of the future, and are therefore not directly implementable. If the observed signal is processed in blocks, however, the non-causal filters are implementable. In this paper, we propose such non-causal time-domain filters for noise reduction in speech applications. We also propose some performance measures that enable us to evaluate the performance of non-causal filters. Moreover, it is shown how some of the filters can be updated recursively. Using the recursive expressions, it is also shown that the output SNRs of the filters always increase as we increase the length of the filter when the desired signal is stationary. From both the theoretical and practical evaluations of the filters, it is clearly shown that the performance of time-domain filtering methods for noise reduction can be improved by introducing non-causality. Jesper Rindom Jensen, Jacob Benesty, Mads Græsbøll Christensen, Søren Holdt Jensen |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | Enhancement of Single-Channel Periodic Signals in the Time-DomainabstractMost state-of-the-art filtering methods for speech enhancement require an estimate of the noise statistics, but the noise statistics are difficult to estimate in practice when speech is present. Thus, nonstationary noise will have a detrimental impact on the performance of most speech enhancement filters. The impact of such noise can be reduced by using the signal statistics rather than the noise statistics in the filter design. For example, this is possible by assuming a harmonic model for the desired signal; while this model fits well for voiced speech, it will not be appropriate for unvoiced speech. That is, signal-dependent methods based on the signal statistics will introduce undesired distortion for some parts of speech compared to signal-independent methods based on the noise statistics. Since both the signal-independent and signal-dependent approaches to speech enhancement have advantages, it is relevant to combine them to reduce the impact of their individual disadvantages. In this paper, we give theoretical insights into the relationship between these different approaches, and these reveal a close relationship between the two approaches. This justifies joint use of such filtering methods which can be beneficial from a practical point of view. Our experimental results confirm that both signal-independent and signal-dependent approaches have advantages and that they are closely-related. Moreover, as a part of our experiments, we illustrate the practical usefulness of combining signal-independent and signal-dependent enhancement methods by applying such methods jointly on real-life speech. Jesper Rindom Jensen, Jacob Benesty, Mads Græsbøll Christensen, Søren Holdt Jensen |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | A Joint Approach for Single-Channel Speaker Identification and Speech SeparationabstractIn this paper, we present a novel system for joint speaker identification and speech separation. For speaker identification a single-channel speaker identification algorithm is proposed which provides an estimate of signal-to-signal ratio (SSR) as a by-product. For speech separation, we propose a sinusoidal model-based algorithm. The speech separation algorithm consists of a double-talk/single-talk detector followed by a minimum mean square error estimator of sinusoidal parameters for finding optimal codevectors from pre-trained speaker codebooks. In evaluating the proposed system, we start from a situation where we have prior information of codebook indices, speaker identities and SSR-level, and then, by relaxing these assumptions one by one, we demonstrate the efficiency of the proposed fully blind system. In contrast to previous studies that mostly focus on automatic speech recognition (ASR) accuracy, here, we report the objective and subjective results as well. The results show that the proposed system performs as well as the best of the state-of-the-art in terms of perceived quality while its performance in terms of speaker identification and automatic speech recognition results are generally lower. It outperforms the state-of-the-art in terms of intelligibility showing that the ASR results are not conclusive. The proposed method achieves on average, 52.3% ASR accuracy, 41.2 points in MUSHRA and 85.9% in speech intelligibility. Pejman Mowlaee, Rahim Saeidi, Mads Græsbøll Christensen, Zheng-Hua Tan, Tomi Kinnunen, Pasi Fränti, Søren Holdt Jensen |
IEEE Trans. Speech Audio Process. | 3 |
| 2011 | A new metric for VQ-based speech enhancement and separationabstractSpeech enhancement and separation algorithms frequently employ two-stage processing schemes, where the signal is first mapped to an intermediate low-dimensional parametric description. Then, these parameters are mapped to vectors in codebooks trained on individual noise-free sources using a vector quantizer. To obtain accurate parameters, one must employ an estimator that takes the signal characteristics into account. An open question is, however, how to derive metrics for use in the vector quantization process. In this paper, we present and derive a new metric aimed at exactly this, and we exemplify and demonstrate its use in sinusoidal modeling. The metric takes into account that parameters may have different uncertainties and dependencies associated with them and thus leads to more accurate estimates, as is demonstrated in experiments. Moreover, we incorporate the metric in a recently proposed speech separation algorithm and compare its performance to state-of-the-art methods. Mads Græsbøll Christensen, Pejman Mowlaee |
ICASSP | 1 |
| 2011 | A single snapshot optimal filtering method for fundamental frequency estimationabstractRecently, optimal linearly constrained minimum variance (LCMV) filtering methods have been applied for fundamental frequency estimation. Like many other fundamental frequency estimators, these methods utilize the inverse covariance matrix. Therefore, the covariance matrix needs to be invertible which is typically ensured by using the sample covariance matrix involving data partitioning. The partitioning adversely affects the spectral resolution. We propose a novel optimal filtering method which utilizes the LCMV principle in conjunction with the iterative adaptive approach (IAA). The IAA enables us to estimate the covariance matrix from a single snapshot, i.e., without data partitioning. The experimental results show, that the performance of the proposed method is comparable or better than that of other competing methods in terms of spectral resolution. Jesper Rindom Jensen, Mads Græsbøll Christensen, Søren Holdt Jensen |
ICASSP | 2 |
| 2011 | Sinusoidal Approach for the Single-Channel Speech Separation and Recognition ChallengeabstractMost of the single-channel speech separation (SCSS) systems use the short-time Fourier transform as their parametric features. Recent studies have shown that employing sinusoidal features for the SCSS application results in a high perceived speech quality. In this paper, we make a systematic study on automatic speech recognition results for a SCSS system that uses sinusoidal features composed of amplitude and frequency. We compare the speech recognition results with those already reported by other participants in the single-channel speech separation and recognition challenge. Our results show that a newly proposed system achieves an overall recognition accuracy of 52.3%, ranges at the median over all other participants in the challenge. Index Terms: sinusoidal modeling, single-channel speech separation and recognition challenge. Pejman Mowlaee, Rahim Saeidi, Zheng-Hua Tan, Mads Græsbøll Christensen, Tomi Kinnunen, Pasi Fränti, Søren Holdt Jensen |
INTERSPEECH | 4 |
| 2011 | An iterative subspace-based multi-pitch estimation algorithm
Johan Xi Zhang, Mads Græsbøll Christensen, Søren Holdt Jensen, Marc Moonen |
Signal Process. | 2 |
| 2011 | New Results on Perceptual Distortion Minimization and Nonlinear Least-Squares Frequency EstimationabstractIn a paper in this journal, a framework was presented wherein a number of practical methods for finding the perceptually most important sinusoids in audio signals could be related using a particular perceptually motivated distortion measure, and it was argued that for Gaussian noise and a large number of samples, these methods should attain the Cramér-Rao lower bound. In this correspondence, we report some new results on this subject. Specifically, we analyze the finite-sample performance of these methods in experiments, and we conclude that for a high number of samples, they perform close to the Cramér-Rao lower bound. However, for a low number of observations, we demonstrate that special care must be taken in designing the perceptually motivated distortion measure if high-resolution estimates are desired. In particular, the smoothness of the frequency response of the perceptual filter that implements the distortion measure is shown to be important. Mads Græsbøll Christensen, Søren Holdt Jensen |
IEEE Trans. Speech Audio Process. | 1 |
| 2011 | New Results on Single-Channel Speech Separation Using Sinusoidal ModelingabstractWe present new results on single-channel speech separation and suggest a new separation approach to improve the speech quality of separated signals from an observed mixture. The key idea is to derive a mixture estimator based on sinusoidal parameters. The proposed estimator is aimed at finding sinusoidal parameters in the form of codevectors from vector quantization (VQ) codebooks pre-trained for speakers that, when combined, best fit the observed mixed signal. The selected codevectors are then used to reconstruct the recovered signals for the speakers in the mixture. Compared to the log-max mixture estimator used in binary masks and the Wiener filtering approach, it is observed that the proposed method achieves an acceptable perceptual speech quality with less cross-talk at different signal-to-signal ratios. Moreover, the method is independent of pitch estimates and reduces the computational complexity of the separation by replacing the short-time Fourier transform (STFT) feature vectors of high dimensionality with sinusoidal feature vectors. We report separation results for the proposed method and compare them with respect to other benchmark methods. The improvements made by applying the proposed method over other methods are confirmed by employing perceptual evaluation of speech quality (PESQ) as an objective measure and a MUSHRA listening test as a subjective evaluation for both speaker-dependent and gender-dependent scenarios. Pejman Mowlaee, Mads Græsbøll Christensen, Søren Holdt Jensen |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Bayesian Interpolation and Parameter Estimation in a Dynamic Sinusoidal ModelabstractIn this paper, we propose a method for restoring the missing or corrupted observations of nonstationary sinusoidal signals which are often encountered in music and speech applications. To model nonstationary signals, we use a time-varying sinusoidal model which is obtained by extending the static sinusoidal model into a dynamic sinusoidal model. In this model, the in-phase and quadrature components of the sinusoids are modeled as first-order Gauss-Markov processes. The inference scheme for the model parameters and missing observations is formulated in a Bayesian framework and is based on a Markov chain Monte Carlo method known as Gibbs sampler. We focus on the parameter estimation in the dynamic sinusoidal model since this constitutes the core of model-based interpolation. In the simulations, we first investigate the applicability of the model and then demonstrate the inference scheme by applying it to the restoration of lost audio packets on a packet-based network. The results show that the proposed method is a reasonable inference scheme for estimating unknown signal parameters and interpolating gaps consisting of missing/corrupted signal segments. Jesper Kjær Nielsen, Mads Græsbøll Christensen, A. Taylan Cemgil, Simon J. Godsill, Søren Holdt Jensen |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Error-correction of binary masks using hidden Markov modelsabstractBinary masking is a simple and efficient method for source separation, and a high increase in intelligibility can be obtained by applying the target binary mask to noisy speech. The target binary mask can only be calculated under ideal conditions and will contain errors when estimated in real-life applications. This paper proposes a method for correcting these errors. The error-correction is based on a hidden Markov model and uses the Viterbi algorithm to calculate the most probable error-free target binary mask from a target binary mask containing errors. The results demonstrate that it is possible to correct errors in the target binary mask and reduce the noise energy. However, speech energy is also reduced by the error-correction, but the impact on speech intelligibility and speech quality are not established or evaluated in the present study. Jesper Bünsow Boldt, Michael Syskind Pedersen, Ulrik Kjems, Mads Græsbøll Christensen, Søren Holdt Jensen |
ICASSP | 4 |
| 2010 | Enhancing sparsity in linear prediction of speech by iteratively reweighted 1-norm minimizationabstractLinear prediction of speech based on 1-norm minimization has already proved to be an interesting alternative to 2-norm minimization. In particular, choosing the 1-norm as a convex relaxation of the 0-norm, the corresponding linear prediction model offers a sparser residual better suited for coding applications. In this paper, we propose a new speech modeling technique based on reweighted 1-norm minimization. The purpose of the reweighted scheme is to overcome the mismatch between 0-norm minimization and 1-norm minimization while keeping the problem solvable with convex estimation tools. Experimental results prove the effectiveness of the reweighted 1-norm minimization, offering better coding properties compared to 1-norm minimization. Daniele Giacobello, Mads Græsbøll Christensen, Manohar N. Murthi, Søren Holdt Jensen, Marc Moonen |
ICASSP | 2 |
| 2010 | Estimation of frame independent and enhancement components for speech communication over packet networksabstractIn this paper, we describe a new approach to cope with packet loss in speech coders. The idea is to split the information present in each speech packet into two components, one to independently decode the given speech frame and one to enhance it by exploiting inter-frame dependencies. The scheme is based on sparse linear prediction and a redefinition of the analysis-by-synthesis process. We present Mean Opinion Scores for the presented coder with different degrees of packet loss and show that it performs similarly to frame dependent coders for low packet loss probability and similarly to frame independent coders for high packet loss probability. We also present ideas on how to make the coder work synergistically with the channel loss estimate. Daniele Giacobello, Manohar N. Murthi, Mads Græsbøll Christensen, Søren Holdt Jensen, Marc Moonen |
ICASSP | 3 |
| 2010 | Improved single-channel speech separation using sinusoidal modelingabstractWe present a novel single-channel separation approach to improve the separation performance while recovering the signals from a mixture. The key idea in this research is to employ a mixture estimator based on unconstrained modified sinusoidal parameters. Compared to the mixmax (binary mask) and Wiener filter (softmask) approaches, the proposed approach works independently of pitch estimates. Furthermore, it is observed that it can achieve acceptable perceptual speech quality with less cross-talk at different signal-to-signal ratios while bringing down the complexity by replacing STFT with sinusoidal parameters. Improvementsmade by the proposed approach are demonstrated by employing PESQ as our objective measure and MUSHRA listening test as our subjective evaluation. Pejman Mowlaee, Mads Græsbøll Christensen, Søren Holdt Jensen |
ICASSP | 2 |
| 2010 | Sinusoidal masks for single channel speech separationabstractIn this paper we present a new approach for binary and soft masks used in single-channel speech separation. We present a novel approach called the sinusoidal mask (binary mask and Wiener filter) in a sinusoidal space. Theoretical analysis is presented for the proposed method, and we show that the proposed method is able to minimize the target speech distortion while suppressing the crosstalk to a predetermined threshold. It is observed that compared to the STFT-based masks, the proposed sinusoidal masks improve the separation performance in terms of objective measures (SSNR and PESQ) and are mostly preferred by listeners. Pejman Mowlaee, Mads Græsbøll Christensen, Søren Holdt Jensen |
ICASSP | 2 |
| 2010 | Joint single-channel speech separation and speaker identificationabstractIn this paper, we propose a closed loop system to improve the performance of single-channel speech separation in a speaker independent scenario. The system is composed of two interconnected blocks: a separation block and a speaker identification block. The improvement is accomplished by incorporating the speaker identities found by the speaker identification block as additional information for the separation block, which converts the speaker-independent separation problem to a speaker-dependent one where the speaker codebooks are known. Simulation results show that the closed loop system enhances the quality of the separated output signals. To assess the improvements, the results are reported in terms of PESQ for both target and masked signals. Pejman Mowlaee, Rahim Saeidi, Zheng-Hua Tan, Mads Græsbøll Christensen, Pasi Fränti, Søren Holdt Jensen |
ICASSP | 4 |
| 2010 | Adaptive feedback cancellation in hearing aids using a sinusoidal near-end signal modelabstractAcoustic feedback is a well-known problem in hearing aids, which is caused by the undesired acoustic coupling between the loudspeaker and the microphone. Acoustic feedback limits the maximum amplification that can be used in the hearing aid without making it unstable. The goal of adaptive feedback cancellation (AFC) is to adaptively model the feedback path and estimate the feedback signal, which is then subtracted from the microphone signal. The main problem in identifying the feedback path model is the correlation between the near-end signal and the loudspeaker signal, which is caused by the closed signal loop. A possible solution to this problem is to use the prediction error method (PEM)-based AFC with a linear prediction (LP) model for the near-end signal. In this paper, a modification to the PEM-based AFC is presented where the LP model is replaced by a sinusoidal near-end signal model. More specifically, it is shown that using frequency estimation techniques to estimate the sinusoidal near-end signal model improves the performance of the PEM-based AFC compared to using a LP model. Simulation results for a hearing aid scenario indicate a significant improvement in terms of misadjustment and maximum stable gain increase. Kim Ngo, Toon van Waterschoot, Mads Græsbøll Christensen, Marc Moonen, Søren Holdt Jensen, Jan Wouters |
ICASSP | 3 |
| 2010 | Signal-to-Signal Ratio Independent Speaker Identification for Co-channel Speech SignalsabstractIn this paper, we consider speaker identification for the co-channel scenario in which speech mixture from speakers is recorded by one microphone only. The goal is to identify both of the speakers from their mixed signal. High recognition accuracies have already been reported when an accurately estimated signal-to-signal ratio (SSR) is available. In this paper, we approach the problem without estimating SSR. We show that a simple method based on fusion of adapted Gaussian mixture models and Kullback-Leibler divergence calculated between models, achieves an accuracy of 97% and 93% when the two target speakers enlisted as three and two most probable speakers, respectively. Rahim Saeidi, Pejman Mowlaee, Tomi Kinnunen, Zheng-Hua Tan, Mads Græsbøll Christensen, Søren Holdt Jensen, Pasi Fränti |
ICPR | 5 |
| 2010 | Improving monaural speaker identification by double-talk detectionabstractThis paper describes a novel approach to improve monoaural speaker identification where two speakers are present in a single-microphone recording. The goal is to identify both of the underlying speakers in the given mixture. The proposed approach is composed of a double-talk detector (DTD) as a preprocessor and speaker identification back-end. We demonstrate that including the double-talk detector improves the speaker identification accuracy. Experiments on GRID corpus show that including the DTD improves average recognition accuracy from 96.53% to 97.43%. Rahim Saeidi, Pejman Mowlaee, Tomi Kinnunen, Zheng-Hua Tan, Mads Græsbøll Christensen, Søren Holdt Jensen, Pasi Fränti |
INTERSPEECH | 5 |
| 2010 | Retrieving Sparse Patterns Using a Compressed Sensing Framework: Applications to Speech Coding Based on Sparse Linear PredictionabstractEncouraged by the promising application of compressed sensing in signal compression, we investigate its formulation and application in the context of speech coding based on sparse linear prediction. In particular, a compressed sensing method can be devised to compute a sparse approximation of speech in the residual domain when sparse linear prediction is involved. We compare the method of computing a sparse prediction residual with the optimal technique based on an exhaustive search of the possible nonzero locations and the well known Multi-Pulse Excitation, the first encoding technique to introduce the sparsity concept in speech coding. Experimental results demonstrate the potential of compressed sensing in speech coding techniques, offering high perceptual quality with a very sparse approximated prediction residual. Daniele Giacobello, Mads Græsbøll Christensen, Manohar N. Murthi, Søren Holdt Jensen, Marc Moonen |
IEEE Signal Process. Lett. | 2 |
| 2010 | A Robust and Computationally Efficient Subspace-Based Fundamental Frequency EstimatorabstractThis paper presents a method for high-resolution fundamental frequency$(F_{0})$estimation based on subspaces decomposed from a frequency-selective data model, by effectively splitting the signal into a number of subbands. The resulting estimator is termed frequency-selective harmonic MUSIC (F-HMUSIC). The subband-based approach is expected to ensure computational savings and robustness. Additionally, a method for automatic subband signal activity detection is proposed, which is based on information-theoretic criterion where no subjective judgment is needed. The F-HMUSIC algorithm exhibits good statistical performance when evaluated with synthetic signals for both white and colored noises, while its evaluation on real-life audio signal shows the algorithm to be competitive with other estimators. Finally, F-HMUSIC is found to be computationally more efficient and robust than other subspace-based$F_{0}$estimators, besides being robust against recorded data with inharmonicities. Johan Xi Zhang, Mads Græsbøll Christensen, Søren Holdt Jensen, Marc Moonen |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Joint estimation of short-term and long-term predictors in speech codersabstractIn low bit-rate coders, the near-sample and far-sample redundancies of the speech signal are usually removed by a cascade of a short-term and a long-term linear predictor. These two predictors are usually found in a sequential and therefore suboptimal approach. In this paper we propose an analysis model that jointly finds the two predictors by adding a regularization term in the minimization process to impose sparsity constraints on a high order predictor. The result is a linear predictor that can be easily factorized into the short-term and long-term predictors. This estimation method is then incorporated into an algebraic code excited linear prediction scheme and shows to have a better performance than traditional cascade methods and other joint optimization methods, offering lower distortion and higher perceptual speech quality. Daniele Giacobello, Mads Græsbøll Christensen, Joachim Dahl, Søren Holdt Jensen, Marc Moonen |
ICASSP | 2 |
| 2009 | Sub-band implementation of the Harmonic MUSIC algorithmabstractIn this paper, we present a novel method for joint estimation of the order and fundamental frequency of a set of harmonically related sinusoids. This method uses a subband based approach to estimate the involved parameters using subspace techniques, and the resulting algorithm is termed frequency-selective harmonic MUSIC (F-HMUSIC). The performance of F-HMUSIC is evaluated and compared to both harmonic MUSIC (HMUSIC) and Cramer-Rao lower bound (CRLB). Especially, in a low signal-to-noise ratio (SNR) with colored noise scenarios, where F-HMUSIC outperforms HMUSIC. F-HMUSIC is concluded to be more computationally efficient and more robust against colored noise than other subspace based fundamental frequency estimators. Johan Xi Zhang, Mads Græsbøll Christensen, Joachim Dahl, Søren Holdt Jensen, Marc Moonen |
ICASSP | 2 |
| 2009 | Robust implementation of the MUSIC algorithmabstractThe problem of estimating frequencies of sinusoids in noise has been studied intensively by the signal processing community during the last decades. Traditionally high resolution subspace-based techniques suffer from high computational complexity, and generally sensitive to the colored noise. We present here a frequency-domain based subspace parameter estimation algorithm termed frequency-selective MUltiple SIgnal Classification (F-MUSIC) that is based on the signal and noise subspace orthogonality property. The method is computationally efficient in providing estimates in the selected subband compared to the classic MUSIC. The performance of F-MUSIC is evaluated and compared to both MUSIC and Cramer-Rao lower bound (CRLB). In a low signal to noise ratio (SNR) with colored noise scenarios, F-MUSIC outperforms MUSIC. Johan Xi Zhang, Mads Græsbøll Christensen, Joachim Dahl, Søren Holdt Jensen, Marc Moonen |
ICASSP | 2 |
| 2009 | Robust Parametric Audio Coding Using Multiple Description CodingabstractWe propose a new multiple description spherical quantization with repetitively coded amplitudes (MDSQRA) scheme suited for quantization of sinusoidal parameters. The quantization scheme is constituted by a set of spherical quantizers inspired by the multiple description spherical trellis-coded quantization (MDSTCQ) scheme. In this scheme, we apply repetitive coding on the amplitudes, while multiple description coding are applied on the phases and frequencies. Thereby, MDSQRA becomes directly implementable, as opposed to MDSTCQ, since the phase and frequency quantizers depend on the amplitudes which have dissimilar descriptions in MDSTCQ. Furthermore, we implement MDSQRA into a perceptual matching pursuit based sinusoidal audio coder. Finally, we evaluate MDSQRA through perceptual distortion measurements and MUSHRA listening tests. The tests show that MDSQRA outperforms MDSTCQ with respect to a expected perceptual distortion measure. The same results are obtained through the MUSHRA tests performed on sound clips coded using MDSQRA and MDSTCQ. Jesper Rindom Jensen, Mads Græsbøll Christensen, Morten Holm Jensens, Søren Holdt Jensen, Torben Larsen |
IEEE Signal Process. Lett. | 2 |
| 2009 | Quantitative Analysis of a Common Audio Similarity MeasureabstractFor music information retrieval tasks, a nearest neighbor classifier using the Kullback-Leibler divergence between Gaussian mixture models of songs' melfrequency cepstral coefficients is commonly used to match songs by timbre. In this paper, we analyze this distance measure analytically and experimentally by the use of synthesized MIDI files, and we find that it is highly sensitive to different instrument realizations. Despite the lack of theoretical foundation, it handles the multipitch case quite well when all pitches originate from the same instrument, but it has some weaknesses when different instruments play simultaneously. As a proof of concept, we demonstrate that a source separation frontend can improve performance. Furthermore, we have evaluated the robustness to changes in key, sample rate, and bitrate. Jesper Højvang Jensen, Mads Græsbøll Christensen, Daniel P. W. Ellis, Søren Holdt Jensen |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | Robust subspace-based fundamental frequency estimationabstractThe problem of fundamental frequency estimation is considered in the context of signals where the frequencies of the harmonics are not exact integer multiples of a fundamental frequency. This frequently occurs in audio signals produced by, for example, stiff-stringed musical instruments, and is sometimes referred to as inharmonicity. We derive a novel robust method based on the subspace orthogonality property of MUSIC and show how it may be used for analyzing audio signals. The proposed method is both more general and less complex than a straight-forward implementation of a parametric model of the inharmonicity derived from a physical instrument model. Additionally, it leads to more accurate estimates of the individual frequencies than the method based on the parametric inharmonicity model and a reduced bias of the fundamental frequency compared to the perfectly harmonic model. Mads Græsbøll Christensen, Pedro Vera-Candeas, Samuel Dilshan Somasundaram, Andreas Jakobsson |
ICASSP | 1 |
| 2008 | A tempo-insensitive distance measure for cover song identification based on chroma featuresabstractWe present a distance measure between audio files designed to identify cover songs, which are new renditions of previously recorded songs. For each song we compute the chromagram, remove phase information and apply exponentially distributed bands in order to obtain a feature matrix that compactly describes a song and is insensitive to changes in instrumentation, tempo and time shifts. As distance between two songs, we use the Frobenius norm of the difference between their feature matrices normalized to unit norm. When computing the distance, we take possible transpositions into account. In a test collection of 80 songs with two versions of each, 38% of the covers were identified. The system was also evaluated on an independent, international evaluation where it despite having much lower complexity performed on par with the winner of last year. Jesper Højvang Jensen, Mads Græsbøll Christensen, Daniel P. W. Ellis, Søren Holdt Jensen |
ICASSP | 2 |
| 2008 | Multiple description quantization of sinusoidal parametersabstractA new scheme for sinusoidal audio coding named multiple description spherical trellis-coded quantization is proposed and analytic expressions for the point densities and expected distortion of the quantizers are derived based on a high-resolution assumption. The proposed quantizers are of variable dimension, i.e., sinusoids can be quantized jointly for each audio segment whereby a lower distortion is achieved. The quantizers are designed to minimize a perceptual distortion measure subject to an entropy constraint for a given packet-loss probability. In experiments, the performance of the quantizers is compared to the corresponding single description spherical quantizer and associated bounds are found to increase robustness towards packet-losses. Morten Holm Larsen, Mads Græsbøll Christensen, Søren Holdt Jensen |
ICASSP | 2 |
| 2008 | Sparse linear predictors for speech processingabstractThis paper presents two new classes of linear prediction schemes. The first one is based on the concept of creating a sparse residual rather than a minimum variance one, which will allow a more efficient quantization; we will show that this works well in presence of voiced speech, where the excitation can be represented by an impulse train, and creates a sparser residual in the case of unvoiced speech. The second class aims at finding sparse prediction coefficients; interesting results can be seen applying it to the joint estimation of long-term and short-term predictors. The proposed estimators are all solutions to convex optimization problems, which can be solved efficiently and reliably using, e.g., interior-point methods. Index Terms: linear prediction, all-pole modeling, convex optimization 1. Daniele Giacobello, Mads Græsbøll Christensen, Joachim Dahl, Søren Holdt Jensen, Marc Moonen |
INTERSPEECH | 2 |
| 2008 | Frequency-domain parameter estimations for binary masked signalsabstractWe present an approach for the extraction of parameters of a damped complex exponential model from a spectrogram modified by a binary mask.The parameters are estimated by a frequency domain based methods using subspace techniques, where the core algorithm is F-ESPRIT.The sub-band defined by the binary mask provides a reduced number of DFT-samples for the parameter extractions, which results in a computational efficient scheme with high parameter estimation accuracy.The proposed synthesis system has synthesis performance comparable to the so-called LSEE-MSTFT.The estimated parameters can be used in many applications such as audio/speech coding, pitch estimation and pitch scale modification. Johan Xi Zhang, Mads Græsbøll Christensen, Joachim Dahl, Søren Holdt Jensen, Marc Moonen |
INTERSPEECH | 2 |
| 2008 | Multi-pitch estimation
Mads Græsbøll Christensen, Petre Stoica, Andreas Jakobsson, Søren Holdt Jensen |
Signal Process. | 1 |
| 2008 | On Optimal Filter Designs for Fundamental Frequency EstimationabstractRecently, we proposed using Capon's minimum variance principle to find the fundamental frequency of a periodic waveform. The resulting estimator is formed such that it maximizes the output power of a bank of filters. We present an alternative optimal single filter design and then proceed to quantify the similarities and differences between the estimators using asymptotic analysis and Monte Carlo simulations. Our analysis shows that the single filter can be expressed in terms of the optimal filterbank and that the methods are asymptotically equivalent but generally different for finite length signals. Mads Græsbøll Christensen, Jesper Højvang Jensen, Andreas Jakobsson, Søren Holdt Jensen |
IEEE Signal Process. Lett. | 1 |
| 2008 | Variable Dimension Trellis-Coded Quantization of Sinusoidal ParametersabstractIn this letter, we propose joint quantization of the parameters of a set of sinusoids based on the theory of trellis-coded quantization. A particular advantage of this approach is that it allows for joint quantization of a variable number of sinusoids, which is particularly relevant in variable rate parametric audio coding. Under high-resolution assumptions and based on a perceptually relevant distortion measure, we derive analytical expressions for the optimal design subject to an entropy constraint. Numerical experiments show a significant performance gain compared to optimal spherical quantization at the cost of a slight increase in computational complexity. Morten Holm Larsen, Mads Græsbøll Christensen, Søren Holdt Jensen |
IEEE Signal Process. Lett. | 2 |
| 2007 | The Multi-Pitch Estimation Problem: some New SolutionsabstractIn this paper, we formulate the multi-pitch estimation problem and propose a number of methods to estimate the set of fundamental frequencies. The methods, which are based on nonlinear least-squares, multiple signal classification (MUSIC) and the Capon principles, have in common the fact that the multiple fundamental frequencies are estimated by means of a one-dimensional search. The statistical properties of the methods are evaluated via Monte Carlo simulations. Mads Græsbøll Christensen, Petre Stoica, Andreas Jakobsson, Søren Holdt Jensen |
ICASSP (3) | 1 |
| 2007 | Joint High-Resolution Fundamental Frequency and Order EstimationabstractIn this paper, we present a novel method for joint estimation of the fundamental frequency and order of a set of harmonically related sinusoids based on the multiple signal classification (MUSIC) estimation criterion. The presented method, termed HMUSIC, is shown to have an efficient implementation using fast Fourier transforms (FFTs). Furthermore, refined estimates can be obtained using a gradient-based method. Illustrative examples of the application of the algorithm to real-life speech and audio signals are given, and the statistical performance of the estimator is evaluated using synthetic signals, demonstrating its good statistical properties. Mads Græsbøll Christensen, Andreas Jakobsson, Søren Holdt Jensen |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | Computationally Efficient Amplitude Modulated Sinusoidal Audio Coding Using Frequency-Domain Linear PredictionabstractA method for amplitude modulated sinusoidal audio coding is presented that has low complexity and low delay. This is based on a sub-band processing system, where, in each subband, the signal is modeled as an amplitude modulated sum of sinusoids. The envelopes are estimated using frequency-domain linear prediction and the prediction coefficients are quantized. As a proof of concept, we evaluate different configurations in a subjective listening test, and this shows that the proposed method offers significant improvements in sinusoidal coding. Furthermore, the properties of the frequency-domain linear prediction-based envelope estimator are analyzed Mads Græsbøll Christensen, Søren Holdt Jensen |
ICASSP (5) | 1 |
| 2006 | Amplitude modulated sinusoidal signal decomposition for audio codingabstractIn this letter, we present a decomposition for sinusoidal coding of audio, based on an amplitude modulation of sinusoids via a linear combination of arbitrary basis vectors. The proposed method, which incorporates a perceptual distortion measure, is based on a relaxation of a nonlinear least-squares minimization. Rate-distortion curves and listening tests show that, compared to a constant-amplitude sinusoidal coder, the proposed decomposition offers perceptually significant improvements in critical transient signals Mads Græsbøll Christensen, Andreas Jakobsson, Søren Vang Andersen, Søren Holdt Jensen |
IEEE Signal Process. Lett. | 1 |
| 2006 | On perceptual distortion minimization and nonlinear least-squares frequency estimationabstractIn this paper, we present a framework for perceptual error minimization and sinusoidal frequency estimation based on a new perceptual distortion measure, and we state its optimal solution. Using this framework, we relate a number of well-known practical methods for perceptual sinusoidal parameter estimation such as the prefiltering method, the weighted matching pursuit, and the perceptual matching pursuit. In particular, we derive and compare the sinusoidal estimation criteria used in these methods. We show that for the sinusoidal estimation problem, the prefiltering method and the weighted matching pursuit are equivalent to the perceptual matching pursuit under certain conditions. Mads Græsbøll Christensen, Søren Holdt Jensen |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | Efficient parametric coding of transientsabstractIn this paper, methods for improved parametric coding of transients are presented. We propose a signal model for coding of transients consisting of a sum of sinusoids each being amplitude-modulated by a different gamma envelope. These envelopes are characterized by an onset time, an attack and a decay parameter. An efficient method for estimating these parameters is presented. Further, methods are proposed that combine this transient model with a constant-amplitude sinusoidal model in order to achieve efficient coding of both stationary and transient signal parts. By rate-distortion optimization using a perceptual distortion measure, we combine variable rate bit allocation and segmentation in an optimal way. Formal, as well as informal, listening tests show that significant improvements can be achieved with the proposed model as compared to a state-of-the-art sinusoidal coder by the combination of optimal segmentation and amplitude modulated sinusoidal audio coding Mads Græsbøll Christensen, Steven van de Par |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | Linear AM decomposition for sinusoidal audio codingabstractWe present a novel decomposition for sinusoidal audio coding using amplitude modulation of sinusoids via a linear combination of arbitrary basis vectors. The proposed method, which incorporates a perceptual distortion measure, is based on a relaxation of a non-linear least squares minimization. It offers benefits in the modeling of transients in audio signals. We compare the decomposition to constant-amplitude sinusoidal coding using rate-distortion curves and listening tests. Both indicate that, at the same bit-rate, perceptually significant improvements can be achieved using the proposed decomposition. Mads Græsbøll Christensen, Andreas Jakobsson, Søren Vang Andersen, Søren Holdt Jensen |
ICASSP (3) | 1 |
| 2005 | Open loop rate-distortion optimized audio codingabstractThe paper addresses complexity reduced rate-distortion optimized audio coding under rate constraint. A technique where distortion minimizing coding templates, chosen from a set of templates, are jointly selected for a set of segments. This optimization requires knowledge of rate-distortion pairs for all segments, and for each coding template, which is often costly to obtain. The proposed framework exchanges true rate-distortion pairs with predicted ones, thereby allowing for complexity reduction. The prediction is based on a property vector extracted for each segment, from which distortion predictions, using Gaussian mixture models, are performed. Here, we evaluate the proposed framework in a sinusoidal coding context. The results show that the proposed framework can increase the distortion performance, compared to a fixed sinusoidal coding scheme. Fredrik Nordén, Mads Græsbøll Christensen, Søren Holdt Jensen |
ICASSP (3) | 2 |
| 2004 | Multiband amplitude modulated sinusoidal audio modelingabstractIn this paper, we investigate the importance of taking frequency-dependent temporal phenomena into account in audio coding. We do this in the context of sinusoidal modeling of audio signals by applying amplitude modulation to the sinusoidal components. Traditionally, audio coders use a fixed time-segmentation for all frequencies despite the fact that it is well-known that the time-frequency resolution of the human auditory system is not constant. The well-known window switching is an example of this. We compare multiband amplitude modulated sinusoidal models to a singleband model using different audio excerpts. Based on both comparative listening tests and a psychoacoustical distortion measure it is concluded that an improvement is generally gained using multiband amplitude modulation, although specific single sources are well-modeled using a singleband model. Mads Græsbøll Christensen, Steven van de Par, Søren Holdt Jensen, Søren Vang Andersen |
ICASSP (4) | 1 |
| 2003 | Compressed domain packet loss concealment of sinusoidally coded speechabstractWe consider the problem of packet loss concealment for voice over IP (VoIP). The speech signal is compressed at the transmitter using a sinusoidal coding scheme working at 8 kbit/s. At the receiver, packet loss concealment is carried out working directly on the quantized sinusoidal parameters, based on time-scaling of the packets surrounding the missing ones. Subjective listening tests show promising results indicating the potential of sinusoidal speech coding for VoIP. Christoffer Rødbro, Mads Græsbøll Christensen, Søren Vang Andersen, Søren Holdt Jensen |
ICASSP (1) | 2 |
| 2003 | Amplitude Modulated Sinusoidal Models for Audio Modeling and Coding
Mads Græsbøll Christensen, Søren Vang Andersen, Søren Holdt Jensen |
KES | 1 |