EDBT 2026 Demo / reviewers in the wild / expert
Richard Heusdens
dblp:91/6470
· DBLP profile ↗
122ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0001-5998-1550ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 82 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 34 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4Security and privacy · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3Systems, architecture and hardware · 2Theory of computation · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Optimal pilot design for OTFS in linear time-varying channelsabstractThis paper investigates the positioning of the pilot symbols, as well as the power distribution between the pilot and the communication symbols for the orthogonal time frequency space (OTFS) modulation scheme. We analyze the pilot placements that minimize the mean squared error (MSE) in estimating the channel taps. This allows us to identify two new pilot allocations for OTFS that save approximately 50% of the pilot overhead compared to existing allocations. In addition, we optimize the average channel capacity by adjusting the power distribution. We show that this leads to a significant increase in average capacity. The results provide valuable guidance for designing the OTFS parameters to achieve maximum capacity. Numerical simulations are performed to validate the findings. Ids Van der Werf, Richard Heusdens, Richard C. Hendriks, Geert Leus |
Signal Process. | 2 |
| 2025 | Re-Evaluating Privacy in Centralized and Decentralized Learning: An Information-Theoretical and Empirical StudyabstractDecentralized Federated Learning (DFL) has garnered attention for its robustness and scalability compared to Centralized Federated Learning (CFL). While DFL is commonly believed to offer privacy advantages due to the decentralized control of sensitive data, recent work by Pasquini et, al. challenges this view, demonstrating that DFL does not inherently improve privacy against empirical attacks under certain assumptions. For investigating fully this issue, a formal theoretical framework is required. Our study offers a novel perspective by conducting a rigorous information-theoretical analysis of privacy leakage in FL using mutual information. We further investigate the effectiveness of privacy-enhancing techniques like Secure Aggregation (SA) in both CFL and DFL. Our simulations and real-world experiments show that DFL generally offers stronger privacy preservation than CFL in practical scenarios where a fully trusted server is not available. We address discrepancies in previous research by highlighting limitations in their assumptions about graph topology and privacy attacks, which inadequately capture information leakage in FL. Changlong Ji, Richard Heusdens, Stéphane Maag, Qiongxiu Li |
ICASSP | 2 |
| 2025 | Privacy-Preserving Distributed Maximum Consensus Without Accuracy LossabstractIn distributed networks, calculating the maximum element is a fundamental task in data analysis, known as the distributed maximum consensus problem. However, the sensitive nature of the data involved makes privacy protection essential. Despite its importance, privacy in distributed maximum consensus has received limited attention in the literature. Traditional privacy-preserving methods typically add noise to updates, degrading the accuracy of the final result. To overcome these limitations, we propose a novel distributed optimization-based approach that preserves privacy without sacrificing accuracy. Our method introduces virtual nodes to form an augmented graph and leverages a carefully designed initialization process to ensure the privacy of honest participants, even when all their neighboring nodes are dishonest. Through a comprehensive information-theoretical analysis, we derive a sufficient condition to protect private data against both passive and eavesdropping adversaries. Extensive experiments validate the effectiveness of our approach, demonstrating that it not only preserves perfect privacy but also maintains accuracy, outperforming existing noise-based methods that typically suffer from accuracy loss. Wenrui Yu, Richard Heusdens, Jun Pang 0001, Qiongxiu Li |
ICASSP | 2 |
| 2025 | Provable Privacy Advantages of Decentralized Federated Learning via Distributed OptimizationabstractFederated learning (FL) emerged as a paradigm designed to improve data privacy by enabling data to reside at its source, thus embedding privacy as a core consideration in FL architectures, whether centralized or decentralized. Contrasting with recent findings by Pasquini et al., which suggest that decentralized FL does not empirically offer any additional privacy or security benefits over centralized models, our study provides compelling evidence to the contrary. We demonstrate that decentralized FL, when deploying distributed optimization, provides enhanced privacy protection - both theoretically and empirically - compared to centralized approaches. The challenge of quantifying privacy loss through iterative processes has traditionally constrained the theoretical exploration of FL protocols. We overcome this by conducting a pioneering in-depth information-theoretical privacy analysis for both frameworks. Our analysis, considering both eavesdropping and passive adversary models, successfully establishes bounds on privacy leakage. In particular, we show information theoretically that the privacy loss in decentralized FL is upper bounded by the loss in centralized FL. Compared to the centralized case where local gradients of individual participants are directly revealed, a key distinction of optimization-based decentralized FL is that the relevant information includes differences of local gradients over successive iterations and the aggregated sum of different nodes’ gradients over the network. This information complicates the adversary’s attempt to infer private data. To bridge our theoretical insights with practical applications, we present detailed case studies involving logistic regression and deep neural networks. These examples demonstrate that while privacy leakage remains comparable in simpler models, complex models like deep neural networks exhibit lower privacy risks under decentralized FL. Extensive numerical tests further validate that decentralized FL is more resistant to privacy attacks, aligning with our theoretical findings. Wenrui Yu, Qiongxiu Li, Milan Lopuhaä-Zwakenberg, Mads Græsbøll Christensen, Richard Heusdens |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Privacy-Preserving Distributed Optimisation using Stochastic PDMMabstractPrivacy-preserving distributed processing has received considerable attention recently. The main purpose of these algorithms is to solve certain signal processing tasks over a network in a decentralised fashion without revealing private/secret data to the outside world. Because of the iterative nature of these distributed algorithms, computationally complex approaches such as (homomorphic) encryption are undesired. Recently, an information theoretic method called subspace perturbation has been introduced for synchronous update schemes. The main idea is to exploit a certain structure in the update equations for noise insertion such that the private data is protected without compromising the algorithm’s accuracy. This structure, however, is absent in asynchronous update schemes. In this paper we will investigate such asynchronous schemes and derive a lower bound on the noise variance after random initialisation of the algorithm. This bound shows that the privacy level of asynchronous schemes is always better than or at least equal to that of synchronous schemes. Computer simulations are conducted to consolidate our theoretical results. Sebastian O. Jordan, Qiongxiu Li, Richard Heusdens |
ICASSP | 3 |
| 2024 | Topology-Dependent Privacy Bound for Decentralized Federated LearningabstractDecentralized Federated Learning (FL) has attracted significant attention due to its enhanced robustness and scalability compared to its centralized counterpart. It pivots on peer-to-peer communication rather than depending on a central server for model aggregation. While prior research has delved into various factors of decentralized FL such as aggregation methods and privacy-preserving techniques, one crucial aspect affecting privacy is relatively unexplored: the underlying graph topology. In this paper, we fill the gap by deriving a stringent privacy bound for decentralized FL under the condition that the accuracy is not compromised, highlighting the pivotal role of graph topology. Specifically, we demonstrate that the minimum privacy loss at each model aggregation step is dependent on the size of what we term as ’honest components’, the maximally connected subgraphs once all untrustworthy participants are excluded from the networks, which is closely tied to network robustness. Our analysis suggests that attack-resilient networks will provide a superior privacy guarantee. We further validate this by studying both Poisson and power law networks, showing that the latter, being less robust against attacks, indeed reveals more privacy. In addition to a theoretical analysis, we consolidate our findings by examining two distinct privacy attacks: membership inference and gradient inversion. Qiongxiu Li, Wenrui Yu, Changlong Ji, Richard Heusdens |
ICASSP | 4 |
| 2024 | On the equivalence of OSDM and OTFSabstractIn this paper, we show the mathematical equivalence of two popular modulation schemes: OSDM and OTFS. The former is mainly used in underwater acoustic communications, while the latter scheme is a promising modulation technique in radio-frequency communications. Although literature suggests a link between the two modulation schemes by connecting them to related modulation schemes like V-OFDM and A-OFDM, to the best of the authors’ knowledge, a direct mathematical comparison between the schemes has not been presented yet. The main purpose of this paper is therefore to show the mathematical equivalence of the two schemes. In addition, by combining the knowledge of acoustic and radio-frequency communications, we give insight in the performance of OSDM/OTFS in terms of intersymbol interference (ISI) and intercarrier interference (ICI) by analyzing its signal structure. Ids Van der Werf, Henry Dol, Koen Blom, Richard Heusdens, Richard C. Hendriks, Geert Leus |
Signal Process. | 4 |
| 2024 | Binaural Beamforming Taking Into Account Spatial Release From MaskingabstractHearing impairment is a prevalent problem with daily challenges like impaired speech intelligibility and sound localisation. One of the shortcomings of spatial filtering in hearing aids is that speech intelligibility is often not optimised directly, meaning that different auditory processes contributing to intelligibility are often not considered. One example is the perceptual phenomenon known as spatial release from masking (SRM). This paper develops a signal model that explicitly considers SRM in the beamforming design, achieved by transforming the binaural intelligibility prediction model (BSIM) into a signal processing framework. The resulting extended signal model is used to analyse the performance of reference beamformers and design a novel beamformer that more closely considers how the auditory system perceives binaural sound. It can be shown that the binaural minimum variance distortionless response (BMVDR) beamformer is also an optimal solution for the extended, perceived model, suggesting that SRM does not play a significant role in intelligibility enhancement after optimal beamforming. However, the optimal beamformer is no longer unique in the extended signal model. The additional secondary degrees of freedom can be used to preserve binaural cues of interfering sources while still achieving the same perceived performance of the BMVDR beamformer, though with a possible high sensitivity to intelligibility model mismatch errors. Johannes W. de Vries, Steven van de Par, Geert Leus, Richard Heusdens, Richard C. Hendriks |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2024 | Adaptive Differentially Quantized Subspace Perturbation (ADQSP): A Unified Framework for Privacy-Preserving Distributed Average ConsensusabstractPrivacy-preserving distributed average consensus has received significant attention recently due to its wide applicability. Based on the achieved performances, existing approaches can be broadly classified into perfect accuracy-prioritized approaches such as secure multiparty computation (SMPC), and worst-case privacy-prioritized approaches such as differential privacy (DP). Methods of the first class achieve perfect output accuracy but reveal some private information, while methods from the second class provide privacy against the strongest adversary at the cost of a loss of accuracy. In this paper, we propose a general approach named adaptive differentially quantized subspace perturbation (ADQSP) which combines quantization schemes with so-called subspace perturbation. Although not relying on cryptographic primitives, the proposed approach enjoys the benefits of both accuracy-prioritized and privacy-prioritized methods and is able to unify them. More specifically, we show that by varying a single quantization parameter the proposed method can vary between SMPC-type performances and DP-type performances. Our results show the potential of exploiting traditional distributed signal processing tools for providing cryptographic guarantees. In addition to a comprehensive theoretical analysis, numerical validations are conducted to substantiate our results. Qiongxiu Li, Jaron Skovsted Gundersen, Milan Lopuhaä-Zwakenberg, Richard Heusdens |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2023 | Convergence of Stochastic PDMMabstractIn recent years, the large increase in connected devices and the data that are collected by these devices have caused a heightened interest in distributed processing. Many practical distributed networks are of heterogeneous nature, because different devices in the network can have different specifications. Because of this, it is highly desirable that algorithms operating within these networks can operate asynchronously, since in that case there is no need for clock synchronisation between the nodes, and the algorithm is not slowed down by the slowest device in the network. In this paper, we focus on the primal-dual method of multipliers (PDMM), which is a promising distributed optimisation algorithm that is suitable for distributed optimisation in heterogeneous networks. Most theoretical work that can be found in existing literature focuses on synchronous versions of PDMM. In this work, we prove the convergence of stochastic PDMM, which is a general framework that can model variations such as asynchronous PDMM and PDMM with transmission losses. Sebastian O. Jordan, Thomas Sherson, Richard Heusdens |
ICASSP | 3 |
| 2023 | Sensor Selection for Angle of Arrival Estimation Based on the Two-Target Cramér-Rao BoundabstractSensor selection is a useful method to help reduce data throughput, as well as computational, power, and hardware requirements, while still maintaining acceptable performance. Although minimizing the Cramér-Rao bound has been adopted previously for sparse sensing, it did not consider multiple targets and unknown source models. In this work, we propose to tackle the sensor selection problem for angle of arrival estimation using the worst-case Cramér-Rao bound of two uncorrelated sources. To do so, we cast the problem as a convex semi-definite program and retrieve the binary selection by randomized rounding. Through numerical examples related to a linear array, we illustrate the proposed method and show that it leads to the natural selection of elements at the edges plus the center of the linear array. This contrasts with the typical solutions obtained from minimizing the single-target Cramér-Rao bound. Costas A. Kokke, Mario Coutino, Laura Anitori, Richard Heusdens, Geert Leus |
ICASSP | 4 |
| 2022 | Communication efficient privacy-preserving distributed optimization using adaptive differential quantizationabstractPrivacy issues and communication cost are both major concerns in distributed optimization in networks. There is often a trade-off between them because the encryption methods used for privacy-preservation often require expensive communication overhead. To address these issues, we, in this paper, propose a quantization-based approach to achieve both communication efficient and privacy-preserving solutions in the context of distributed optimization. By deploying an adaptive differential quantization scheme, we allow each node in the network to achieve its optimum solution with a low communication cost while keeping its private data unrevealed. Additionally, the proposed approach is general and can be applied in various distributed optimization methods, such as the primal-dual method of multipliers (PDMM) and the alternating direction method of multipliers (ADMM). We consider two widely used adversary models, passive and eavesdropping, and investigate the properties of the proposed approach using different applications and demonstrate its superior performance compared to existing privacy-preserving approaches in terms of both accuracy and communication cost. Qiongxiu Li, Richard Heusdens, Mads Græsbøll Christensen |
Signal Process. | 2 |
| 2021 | Acoustic Reflectors Localization from Stereo Recordings Using Neural NetworksabstractAcoustic room geometry estimation is often performed in ad hoc settings, i.e., using multiple microphones and sources distributed around the room, or assuming control over the excitation signals. We propose a fully convolutional network (FCN) that localizes reflective surfaces under the relaxed assumptions that (i) a compact array of only two microphones is available, (ii) emitter and receivers are not synchronized, and (iii) both the excitation signals and the impulse responses of the enclosures are unknown. Our FCN is trained in a supervised fashion to predict the likelihood of reflective surfaces at specific distances and directions-of-arrival (DOA). When a single reflective surface is present, up to 80% of real and virtual sources are detected, while this figure approaches 50% in rectangular rooms. Experiments on real-world recordings report similar accuracy as with artificially reverberated speech signals, validating the generalization capabilities of the framework. Giovanni Bologni, Richard Heusdens, Jorge Martínez 0002 |
ICASSP | 2 |
| 2021 | Localization Based on Enhanced Low Frequency Interaural Level DifferenceabstractThe processing of low-frequency interaural time differences is found to be problematic among hearing-impaired people. The current generation of beamformers does not consider this deficiency. In an attempt to tackle this issue, we propose to replace the inaudible interaural time differences in the low-frequency region with the interaural level differences. In addition, a beamformer is introduced and analyzed, which enhances the low-frequency interaural level differences of the sound sources using a near-field transformation. The proposed beamforming problem is relaxed to a convex problem using semi-definite relaxation. The instrumental analysis suggests that the low-frequency interaural level differences are enhanced without hindering the provided intelligibility. A psychoacoustic localization test is done using a listening experiment, which suggests that the replacement of time differences into level differences improves the localization performance of normal-hearing listeners for an anechoic scene but not for a reverberant scene. Metin Calis, Steven van de Par, Richard Heusdens, Richard C. Hendriks |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Privacy-Preserving Distributed Processing: Metrics, Bounds and AlgorithmsabstractPrivacy-preserving distributed processing has recently attracted considerable attention. It aims to design solutions for conducting signal processing tasks over networks in a decentralized fashion without violating privacy. Many existing algorithms can be adopted to solve this problem such as differential privacy, secure multiparty computation, and the recently proposed distributed optimization based subspace perturbation algorithms. However, since each of them is derived from a different context and has different metrics and assumptions, it is hard to choose or design an appropriate algorithm in the context of distributed processing. In order to address this problem, we first propose general mutual information based information-theoretical metrics that are able to compare and relate these existing algorithms in terms of two key aspects: output utility and individual privacy. We consider two widely-used adversary models, the passive and eavesdropping adversary. Moreover, we derive a lower bound on individual privacy which helps to understand the nature of the problem and provides insights on which algorithm is preferred given different conditions. To validate the above claims, we investigate a concrete example and compare a number of state-of-the-art approaches in terms of the concerned aspects using not only theoretical analysis but also numerical validation. Finally, we discuss and provide principles for designing appropriate algorithms for different applications. Qiongxiu Li, Jaron Skovsted Gundersen, Richard Heusdens, Mads Græsbøll Christensen |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2020 | Convex Optimisation-Based Privacy-Preserving Distributed Average Consensus in Wireless Sensor NetworksabstractIn many applications of wireless sensor networks, it is important that the privacy of the nodes of the network be protected. Therefore, privacy-preserving algorithms have received quite some attention recently. In this paper, we propose a novel convex optimization-based solution to the problem of privacy-preserving distributed average consensus. The proposed method is based on the primal-dual method of multipliers (PDMM), and we show that the introduced dual variables of the PDMM will only converge in a certain subspace determined by the graph topology and will not converge in the orthogonal complement. These properties are exploited to protect the private data from being revealed to others. More specifically, the proposed algorithm is proven to be secure for both passive and eavesdropping adversary models. Finally, the convergence properties and accuracy of the proposed approach are demonstrated by simulations which show that the method is superior to the state-of-the-art. Qiongxiu Li, Richard Heusdens, Mads Græsbøll Christensen |
ICASSP | 2 |
| 2020 | Spatially Correct Rate-Constrained Noise Reduction for Binaural Hearing Aids in Wireless Acoustic Sensor NetworksabstractCompared to monaural hearing aids (HAs), binaural hearing aid systems, in which there is a communication link between the two devices, have improved noise reduction capabilities and the ability to preserve binaural spatial information. However, the limited HA battery lifetime puts constraints on the amount of information that can be shared between the two devices. In other words, the rate of transmission between the devices is an important constraint that needs to be considered, while preserving the spatial information. In this article, a linearly constrained noise reduction problem is proposed, which jointly finds the optimal rate allocation and the optimal estimation (beamforming) weights across all sensors and frequencies, while preserving the binaural spatial cues of point sources. The proposed method considers a rate constraint together with linear constraints to preserve the binaural spatial cues of point sources. Minimizing the mean square error on the estimated target speech at the left and the right side beamformers, the optimal weights are found to be rate-constrained linearly constrained minimum variance (LCMV) filters, and the optimal rates are found to be the solutions to a set of reverse water filling problems. The performance of the proposed method is evaluated using the averaged binaural signal-to-noise ratio (SNR), the interaural level difference (ILD) error and the interaural time difference (ITD) error. The results show that the proposed method outperforms spatially correct noise reduction approaches that use naive/random rate allocation strategies. Jamal Amini, Richard C. Hendriks, Richard Heusdens, Meng Guo 0001, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Rate-Constrained Noise Reduction in Wireless Acoustic Sensor NetworksabstractWireless acoustic sensor networks (WASNs) can be used for centralized multi-microphone noise reduction, where the processing is done in a fusion center (FC). To perform the noise reduction, the data needs to be transmitted to the FC. Considering the limited battery life of the devices in a WASN, the total data rate at which the FC can communicate with the different network devices should be constrained. In this article, we propose a rate-constrained multi-microphone noise reduction algorithm, which jointly finds the best rate allocation and estimation weights for the microphones across all frequencies. The optimal linear estimators are found to be the quantized Wiener filters, and the rates are the solutions to a filter-dependent reverse water-filling problem. The performance of the proposed framework is evaluated using simulations in terms of mean square error and predicted speech intelligibility. The results show that the proposed method is very close in performance to that of the existing optimal method based on discrete optimization. However, the proposed approach can do this at a much lower complexity, while the existing optimal reference method needs a non-tractable exhaustive search to find the best rate allocation across microphones. Jamal Amini, Richard C. Hendriks, Richard Heusdens, Meng Guo 0001, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | A Novel Binaural Beamforming Scheme with Low Complexity Minimizing Binaural-cue DistortionsabstractWhile the majority of binaural beamformers aim to minimize the output noise power while (approximately) preserving the binaural cues of the sources using constraints, we propose in this paper to minimize the binaural-cue distortions of the sources in the acoustic scene, such that the output noise power is below a predefined threshold. This new problem formulation is a convex QCQP problem, which leads to an efficient trade-off between noise reduction, binaural-cue preservation and complexity. In particular, the proposed beamformer provides a better trade-off between noise reduction and binaural-cue preservation (in terms of interaural level and phase differences) compared to the well-known binaural minimum variance distortionless response-η beamformer. Andreas I. Koutrouvelis, Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001, Meng Guo 0001 |
ICASSP | 3 |
| 2019 | Distributed Rate-Constrained LCMV BeamformingabstractIn this letter, we propose a decentralized framework for rate-distributed linearly constrained minimum variance (LCMV) beamforming in wireless acoustic sensor networks. To save the energy usage within the network, we propose to minimize the transmission cost and put a constraint on the noise reduction performance. Subsequently, we decentralize the obtained LCMV filter structure by exploiting an imposed block diagonal form of the noise correlation matrix. As a result, the beamformer weights are calculated in a decentralized fashion and each node can determine its quantization rate locally. Finally, numerical results validate the proposed method. Jie Zhang 0042, Andreas I. Koutrouvelis, Richard Heusdens, Richard C. Hendriks |
IEEE Signal Process. Lett. | 3 |
| 2019 | Asymmetric Coding for Rate-Constrained Noise Reduction in Binaural Hearing AidsabstractBinaural hearing aids (HAs) can potentially perform advanced noise reduction algorithms, leading to an improvement over monaural/bilateral HAs. Due to the limited transmission capacities between the HAs and given knowledge of the complete joint noisy signal statistics, the optimal rate-constrained beamforming strategy is known from the literature. However, as these joint statistics are unknown in practice, sub-optimal strategies have been presented. In this paper, we present a unified framework to study the performance of these existing optimal and sub-optimal rate-constrained beamforming methods for binaural HAs. Moreover, we propose to use an asymmetric sequential coding scheme to estimate the joint statistics between the microphones in the two HAs. We show that under certain assumptions, this leads to sub-optimal performance in one HA but allows to obtain the truly optimal performance in the second HA. Based on the mean square error distortion measure, we evaluate the performance improvement between monaural beamforming (no communication) and the proposed scheme, as well as the optimal and the existing sub-optimal strategies in terms of the information bit-rate. The results show that the proposed method outperforms existing practical approaches in most scenarios, especially at middle rates and high rates, without having the prior knowledge of the joint statistics. Jamal Amini, Richard C. Hendriks, Richard Heusdens, Meng Guo 0001, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | A Convex Approximation of the Relaxed Binaural Beamforming Optimization ProblemabstractThe recently proposed relaxed binaural beamforming (RBB) optimization problem provides a flexible tradeoff between noise suppression and binaural-cue preservation of the sound sources in the acoustic scene. It minimizes the output noise power, under the constraints, which guarantee that the target remains unchanged after processing and the binaural-cue distortions of the acoustic sources will be less than a user-defined threshold. However, the RBB problem is a computationally demanding non convex optimization problem. The only existing suboptimal method which approximately solves the RBB is a successive convex optimization (SCO) method which, typically, requires to solve multiple convex optimization problems per frequency bin, in order to converge. Convergence is achieved when all constraints of the RBB optimization problem are satisfied. In this paper, we propose a semidefinite convex relaxation (SDCR) of the RBB optimization problem. The proposed suboptimal SDCR method solves a single convex optimization problem per frequency bin, resulting in a much lower computational complexity than the SCO method. Unlike the SCO method, the SDCR method does not guarantee user-controlled upper-bounded binaural-cue distortions. To tackle this problem, we also propose a suboptimal hybrid method that combines the SDCR and SCO methods. Instrumental measures combined with a listening test show that the SDCR and hybrid methods achieve significantly lower computational complexity than the SCO method, and in most cases better tradeoff between predicted intelligibility and binaural-cue preservation than the SCO method. Andreas I. Koutrouvelis, Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Robust Joint Estimation of Multimicrophone Signal Model ParametersabstractOne of the biggest challenges in multimicrophone applications is the estimation of the parameters of the signal model, such as the power spectral densities (PSDs) of the sources, the early (relative) acoustic transfer functions of the sources with respect to the microphones, the PSD of late reverberation, and the PSDs of microphone-self noise. Typically, existing methods estimate subsets of the aforementioned parameters and assume some of the other parameters to be known a priori. This may result in inconsistencies and inaccurately estimated parameters and potential performance degradation in the applications using these estimated parameters. So far, there is no method to jointly estimate all the aforementioned parameters. In this paper, we propose a robust method for jointly estimating all the aforementioned parameters using confirmatory factor analysis. The estimation accuracy of the signal-model parameters thus obtained outperforms existing methods in most cases. We experimentally show significant performance gains in several multimicrophone applications over state-of-the-art methods. Andreas I. Koutrouvelis, Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Relative Acoustic Transfer Function Estimation in Wireless Acoustic Sensor NetworksabstractIn this paper, we present an algorithm to estimate the relative acoustic transfer function (RTF) of a target source in wireless acoustic sensor networks (WASNs). Two well-known methods to estimate the RTF are the covariance subtraction (CS) method and the covariance whitening (CW) approach, the latter based on the generalized eigenvalue decomposition. Both methods depend on the use of the noisy correlation matrix, which, in practice, has to be estimated using limited and (in WASNs) quantized data. The bit rate and the fact that we use limited data records therefore directly affect the accuracy of the estimated RTFs. Therefore, we first theoretically analyze the estimation performance of the two approaches in terms of bit rate. Second, we propose a rate-distribution method by minimizing the power usage and constraining the expected estimation error for both RTF estimators. The optimal rate distributions are found by using convex optimization techniques. The model-based methods, however, are impractical due to the dependence on the true RTFs. We therefore further develop two greedy rate-distribution methods for both approaches. Finally, numerical simulations on synthetic data and real audio recordings show the superiority of the proposed approaches in power usage compared to uniform rate allocation. We find that in order to satisfy the same RTF estimation accuracy, the rate-distributed CW methods consume much less transmission energy than the CS-based methods. Jie Zhang 0042, Richard Heusdens, Richard C. Hendriks |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Quantisation Effects in Distributed OptimisationabstractIn this paper the effects of quantisation on distributed convex optimisation algorithms are explored via the lens of monotone operator theory. Specifically, by representing transmission quantisation via an additive noise model, we demonstrate how quantisation can be viewed as an instance of an inexact Krasnosel' skiľ-Mann scheme. In the case of two distributed solvers, the Alternating Direction Method of Multipliers and the Primal Dual Method of Multipliers, we further demonstrate how an adaptive quantisation scheme can be constructed to reduce transmission costs between nodes. Finally for the Gaussian channel capacity maximisation problem, we demonstrate convergence even in the presence of one-bit uniform quantisation based on the aforementioned adaptive quantisation scheme. Joseph A. G. Jonkman, Thomas Sherson, Richard Heusdens |
ICASSP | 3 |
| 2018 | Distributed Tdoa-Based Indoor Source LocalisationabstractIndoor localisation is an important research topic with several possible applications. For example, knowing a user's location can be used as navigation aid in hospitals and malls, or for better targeted marketing. In this paper we consider the case where the environment of interest is equipped with several receivers (with known location) from which time-difference-of-arrival (TDOA) measurements are obtained and used to localise the source. We will present a distributed algorithm for localising the source. More specifically, we experimentally show that the distributed algorithm, which only uses time-of-arrival (TOA) measurements obtained from neighbouring receivers to calculate the TDOAs, performs as well as a centralised solution that has access to all TOA measurements in the network. In addition, we propose a method for discarding erroneous TOA measurements which considerably improves the performance in noisy and reverberant environments. Wangyang Yu 0002, Nikolay D. Gaubitch, Richard Heusdens |
ICASSP | 3 |
| 2018 | A Low-Cost Robust Distributed Linearly Constrained Beamformer for Wireless Acoustic Sensor Networks With Arbitrary TopologyabstractWe propose a new robust distributed linearly constrained beamformer that utilizes a set of linear equality constraints to reduce the cross power spectral density matrix to a block-diagonal form. The proposed beamformer has a convenient objective function for use in arbitrary distributed network topologies while having identical performance to a centralized implementation. Moreover, the new optimization problem is robust to relative acoustic transfer function (RATF) estimation errors and to target activity detection (TAD) errors. Two variants of the proposed beamformer are presented and evaluated in the context of multimicrophone speech enhancement in a wireless acoustic sensor network, and are compared with other state-of-the-art distributed beamformers in terms of communication costs and robustness to RATF estimation errors and TAD errors. Andreas I. Koutrouvelis, Thomas Sherson, Richard Heusdens, Richard C. Hendriks |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Microphone Subset Selection for MVDR Beamformer Based Noise ReductionabstractIn large-scale wireless acoustic sensor networks (WASNs), many of the sensors will only have a marginal contribution to a certain estimation task. Involving all sensors increases the energy budget unnecessarily and decreases the lifetime of the WASN. Using microphone subset selection, also termed as sensor selection, the most informative sensors can be chosen from a set of candidate sensors to achieve a prescribed inference performance. In this paper, we consider microphone subset selection for minimum variance distortionless response (MVDR) beamformer based noise reduction. The best subset of sensors is determined by minimizing the transmission cost while constraining the output noise power (or signal-to-noise ratio). Assuming the statistical information on correlation matrices of the sensor measurements is available, the sensor selection problem for this model-driven scheme is first solved by utilizing convex optimization techniques. In addition, to avoid estimating the statistics related to all the candidate sensors beforehand, we also propose a data-driven approach to select the best subset using a greedy strategy. The performance of the greedy algorithm converges to that of the model-driven method, while it displays advantages in dynamic scenarios as well as on computational complexity. Compared to a sparse MVDR or radius-based beamformer, experiments show that the proposed methods can guarantee the desired performance with significantly less transmission costs. Jie Zhang 0042, Sundeep Prabhakar Chepuri, Richard C. Hendriks, Richard Heusdens |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2018 | Rate-Distributed Spatial Filtering Based Noise Reduction in Wireless Acoustic Sensor NetworksabstractIn wireless acoustic sensor networks (WASNs), sensors typically have a limited energy budget as they are often battery-driven. Energy efficiency is, therefore, essential for the design of algorithms in WASNs. One way to reduce energy costs is to select only the sensors that are most informative, a problem known as sensor selection . In this way, only sensors that significantly contribute to the task at hand will be involved. In this paper, we consider a more general approach, which is based on rate-distributed spatial filtering. Depending on the distance over which a transmission takes place, the bit rate directly influences the energy consumption. We try to minimize the battery usage due to transmission, while constraining the noise reduction performance. This results in an efficient rate allocation strategy, which depends on the underlying signal statistics, as well as the distance from sensors to a fusion center (FC). Through the utilization of a linearly constrained minimum variance beamformer, the problem is derived as a semidefinite program. Furthermore, we show that rate allocation is more general than sensor selection, and sensor selection can be seen as a special case of the presented rate-allocation solution, e.g., the best microphone subset can be determined by thresholding the rates. Finally, numerical simulations for estimating several target sources in a WASN demonstrate that the proposed method outperforms the sensor-selection-based approaches in terms of energy usage, and we find that the sensors close to the FC and point sources are allocated with higher rates. Jie Zhang 0042, Richard Heusdens, Richard C. Hendriks |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Greedy alternative for room geometry estimation from acoustic echoes: A subspace-based methodabstractIn this paper, we present a greedy subspace method for the acoustic echoes labeling problem, which occurs in applications such as source localization and room geometry estimation. The orthogonal projection into the null space of the microphones position matrix is used to filter and sort all possible combinations of echoes. A greedy strategy, based on the rank constraint of Euclidean distance matrices (EDMs), is used on the sorted subset of echo combinations to extract the feasible combinations. Numerical simulations using room impulse responses (RIRs) from shoe-box shaped rooms show that the method provides improvements in terms of computational complexity and the number of required measurements with respect to a recently published graph-based method. Mario Coutino, Martin Bo Møller, Jesper Kjær Nielsen, Richard Heusdens |
ICASSP | 4 |
| 2017 | Quantisation effects in PDMM: A first study for synchronous distributed averagingabstractLarge-scale networks of computing units, often characterised by the absence of central control, have become commonplace in many applications. To facilitate data processing in these large-scale networks, distributed signal processing is required. The iterative behaviour of distributed processing algorithms combined with energy, computational power, and bandwidth limitations imposed by such networks, place tight constraints on the transmission capacities of the individual nodes. In this paper we investigate the effects of subtractive dithered uniform quantisation in PDMM for the synchronous distributed averaging problem. This is done by deriving expressions for the mean squared error (MSE) that include quantisation noise. Also, the required data rate for quantised PDMM is considered. It was found that for practical applications quantisation in PDMM can be applied with a fixed-rate quantiser, such that significant data rate reduction can be achieved, without compromising the rate of convergence. Daan H. M. Schellekens, Thomas Sherson, Richard Heusdens |
ICASSP | 3 |
| 2017 | Distributed max-SINR speech enhancement with ad hoc microphone arraysabstractIn recent years, signal processing with ad hoc microphone arrays has attracted a lot of attention. Speech enhancement in noisy, interfered, and reverberant environments is one of the problems targeted by ad hoc microphone arrays. Most of the proposed solutions require knowledge of fingerprints, such as acoustic transfer functions, which may not be known as accurately as required in practical situations. In this paper, a distributed signal subspace filtering method is proposed which is not restricted to a special graph topology. Here, the maximum signal to interference-plus-noise ratio (max-SINR) criterion is used with the primal-dual method of multipliers for distributed filtering. The paper investigates the convergence of the algorithm in both synchronous and asynchronous schemes, and also discusses some practical pros and cons. The applicability of the proposed method is demonstrated by means of simulation results. Vincent Mohammad Tavakoli, Jesper Rindom Jensen, Richard Heusdens, Jacob Benesty, Mads Græsbøll Christensen |
ICASSP | 3 |
| 2017 | Relaxed Binaural LCMV BeamformingabstractIn this paper, we propose a new binaural beamforming technique, which can be seen as a relaxation of the linearly constrained minimum variance (LCMV) framework. The proposed method can achieve simultaneous noise reduction and exact binaural cue preservation of the target source, similar to the binaural minimum variance distortionless response (BMVDR) method. However, unlike BMVDR, the proposed method is also able to preserve the binaural cues of multiple interferers to a certain predefined accuracy. Specifically, it is able to control the trade-off between noise reduction and binaural cue preservation of the interferers by using a separate trade-off parameter per-interferer. Moreover, we provide a robust way of selecting these trade-off parameters in such a way that the preservation accuracy for the binaural cues of the interferers is always better than the corresponding ones of the BMVDR. The relaxation of the constraints in the proposed method achieves approximate binaural cue preservation of more interferers than other previously presented LCMV-based binaural beamforming methods that use strict equality constraints. Andreas I. Koutrouvelis, Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Concurrent localization of sound sources and dual-microphone sub-arrays using TOFs
Mojtaba Farmani, Richard Heusdens, Michael Syskind Pedersen, Zheng-Hua Tan, Jesper Jensen 0001 |
FUSION | 2 |
| 2016 | Room geometry estimation from acoustic echoes using graph-based echo labelingabstractA computer being able to estimate the geometry of a room could benefit applications such as auralization, robot navigation, virtual reality and teleconferencing. When estimating the geometry of a room using multiple microphones, the main challenge is to identify which reflections, or echoes, originate from the same wall and can, therefore, be modeled by a virtual source outside the room using the mirror image source model. In this paper we present a new and efficient method to disambiguate the echoes using a graph theoretical approach where echo combinations are modeled as nodes in a graph and the problem is stated as a maximum independent set problem. Once the echoes are correctly labelled, we know the locations of the virtual sources from which we can infer the room geometry. Experiments for shoe-box shaped rooms show that we can reliably estimate the room geometry within seconds on contemporary hardware and achieve centimeter precision on finding the vertices of the room. Ingmar Jager, Richard Heusdens, Nikolay D. Gaubitch |
ICASSP | 2 |
| 2016 | DOA estimation of audio sources in reverberant environmentsabstractReverberation is well-known to have a detrimental impact on many localization methods for audio sources. We address this problem by imposing a model for the early reflections as well as a model for the audio source itself. Using these models, we propose two iterative localization methods that estimate the direction-of-arrival (DOA) of both the direct path of the audio source and the early reflections. In these methods, the contribution of the early reflections is essentially subtracted from the signal observations before localization of the direct path component, which may reduce the estimation bias. Our simulation results show that we can estimate the DOA of the desired signal more accurately with this procedure compared to state-of-the-art estimator in both synthetic and real data experiments with reverberation. Jesper Rindom Jensen, Jesper Kjær Nielsen, Richard Heusdens, Mads Græsbøll Christensen |
ICASSP | 3 |
| 2016 | Improved multi-microphone noise reduction preserving binaural cuesabstractWe propose a new multi-microphone noise reduction technique for binaural cue preservation of the desired source and the interferers. This method is based on the linearly constrained minimum variance (LCMV) framework, where the constraints are used for the binaural cue preservation of the desired source and of multiple interferers. In this framework there is a trade-off between noise reduction and binaural cue preservation. The more constraints the LCMV uses for preserving binaural cues, the less degrees of freedom can be used for noise suppression. The recently presented binaural LCMV (BLCMV) method and the optimal BLCMV (OBLCMV) method require two constraints per interferer and introduce an additional interference rejection parameter. This unnecessarily reduces the degrees of freedom, available for noise reduction, and negatively influences the trade-off between noise reduction and binaural cue preservation. With the proposed method, binaural cue preservation is obtained using just a single constraint per interferer without the need of an interference rejection parameter. The proposed method can simultaneously achieve noise reduction and perfect binaural cue preservation of more than twice as many interferers as the BLCMV, while the OBLCMV can preserve the binaural cues of only one interferer. Andreas I. Koutrouvelis, Richard C. Hendriks, Jesper Jensen 0001, Richard Heusdens |
ICASSP | 4 |
| 2016 | A distributed algorithm for robust LCMV beamformingabstractIn this paper we propose a distributed reformulation of the linearly constrained minimum variance (LCMV) beamformer for use in acoustic wireless sensor networks. The proposed distributed minimum variance (DMV) algorithm, for which we demonstrate implementations for both cyclic and acyclic networks, allows the optimal beamformer output to be computed at each node without the need for sharing raw data within the network. By exploiting the low rank structure of estimated covariance matrices in time-varying noise fields, the algorithm can also provide a reduction in the total amount of data transmitted during computation when compared to centralised solutions. This is particularly true when multiple microphones are used per node. We also compare the performance of DMV with state of the art distributed beamformers and demonstrate that it achieves greater improvements in SNR in dynamic noise fields with similar transmission costs. Thomas Sherson, W. Bastiaan Kleijn, Richard Heusdens |
ICASSP | 3 |
| 2016 | On simplifying the primal-dual method of multipliersabstractRecently, the primal-dual method of multipliers (PDMM) has been proposed to solve a convex optimization problem defined over a general graph. In this paper, we consider simplifying PDMM for a subclass of the convex optimization problems. This subclass includes the consensus problem as a special form. By using algebra, we show that the update expressions of PDMM can be simplified significantly. We then evaluate PDMM for training a support vector machine (SVM). The experimental results indicate that PDMM converges considerably faster than the alternating direction method of multipliers (ADMM). Guoqiang Zhang 0003, Richard Heusdens |
ICASSP | 2 |
| 2016 | A Fast Method for High-Resolution Voiced/Unvoiced Detection and Glottal Closure/Opening Instant Estimation of SpeechabstractWe propose a fast speech analysis method which simultaneously performs high-resolution voiced/unvoiced detection (VUD) and accurate estimation of glottal closure and glottal opening instants (GCIs and GOIs, respectively). The proposed algorithm exploits the structure of the glottal flow derivative in order to estimate GCIs and GOIs only in voiced speech using simple time-domain criteria. We compare our method with well-known GCI/GOI methods, namely, the dynamic programming projected phase-slope algorithm (DYPSA), the yet another GCI/GOI algorithm (YAGA) and the speech event detection using the residual excitation and a mean-based signal (SEDREAMS). Furthermore, we examine the performance of the aforementioned methods when combined with state-of-the-art VUD algorithms, namely, the robust algorithm for pitch tracking (RAPT) and the summation of residual harmonics (SRH). Experiments conducted on the APLAWD and SAM databases show that the proposed algorithm outperforms the state-of-the-art combinations of VUD and GCI/GOI algorithms with respect to almost all evaluation criteria for clean speech. Experiments on speech contaminated with several noise types (white Gaussian, babble, and car-interior) are also presented and discussed. The proposed algorithm outperforms the state-of-the-art combinations in most evaluation criteria for signal-to-noise ratio greater than 10 dB. Andreas I. Koutrouvelis, George P. Kafentzis, Nikolay D. Gaubitch, Richard Heusdens |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2015 | Bi-alternating direction method of multipliers over graphsabstractIn this paper, we extend the bi-alternating direction method of multipliers (BiADMM) designed on a graph of two nodes to a graph of multiple nodes. In particular, we optimize a sum of convex functions defined over a general graph, where every edge carries a linear equality constraint. In designing the new algorithm, an augmented primal-dual Lagrangian function is carefully constructed which naturally captures the associated graph topology. We show that under both the synchronous and asynchronous updating schemes, the extended BiADMM has the convergence rate of O(1/K) (where K denotes the iteration index) for general closed, proper and convex functions. As an example, we apply the new algorithm for distributed averaging. Experimental results show that the new algorithm remarkably outperforms the state-of-the-art methods. Guoqiang Zhang 0003, Richard Heusdens |
ICASSP | 2 |
| 2014 | Calibration of distributed sound acquisition systems using TOA measurements from a moving acoustic sourceabstractWe present a method for calibrating a distributed microphone array using time-of-arrival (TOA) measurements. The calibration encompasses localization and gain equalization of the microphones, which are both important in applications such as beamforming. The availability of accurate TOA measurements between the microphones and a set of spatially distributed acoustic events is pivotal to the calibration task. We propose to use a moving acoustic source emitting a calibration signal at known intervals. We then show that the TOAs and the observed signals can be used to estimate the gain differences between microphones in addition to the more established microphone localization. Finally, we provide experimental results with simulated and real measured data to demonstrate that our approach facilitates accurate TOA measurements and hence, accurate localization and gain equalization, even in reverberant and noisy conditions. Nikolay D. Gaubitch, W. Bastiaan Kleijn, Richard Heusdens |
ICASSP | 3 |
| 2014 | Time-delay estimation for TOA-based localization of multiple sensorsabstractIn many applications using multiple sensors, knowledge of the relative positions of the sensors is required. The locations of the sensors can be obtained from measured time-of-arrivals (TOAs) of events generated by sources. Although several TOA-based localization techniques exist, practical TOA measurements are incomplete because they include an unknown internal delay; the time taken from the signal reaching the sensor to that it is registered as received by the capturing device. In order to localize the sensors properly, these internal delays need to be estimated accurately. In this paper we propose a method for estimating the internal delays by using a data fitting technique based on structured total least squares. Under reasonable assumptions we show that the algorithm is guaranteed to converge to the optimal solution and ultimately achieves a quadratic rate of convergence. Experimental results show that the execution time is less than 1% of the execution time of existing methods while attaining an even higher accuracy. Richard Heusdens, Nikolay D. Gaubitch |
ICASSP | 1 |
| 2014 | On the convergence rate of the bi-alternating direction method of multipliersabstractIn this paper, we analyze the convergence rate of the bi-alternating direction method of multipliers (BiADMM). Differently from ADMM that optimizes an augmented Lagrangian function, Bi-ADMM optimizes an augmented primal-dual Lagrangian function. The new function involves both the objective functions and their conjugates, thus incorporating more information of the objective functions than the augmented Lagrangian used in ADMM. We show that BiADMM has a convergence rate of O(K-1) (K denotes the number of iterations) for general convex functions. We consider the lasso problem as an example application. Our experimental results show that BiADMM outperforms not only ADMM, but fast-ADMM as well. Guoqiang Zhang 0003, Richard Heusdens, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2014 | Convergence of Min-Sum-Min Message-Passing for Quadratic Optimization
Guoqiang Zhang 0003, Richard Heusdens |
ECML/PKDD (3) | 2 |
| 2014 | Speech energy redistribution for intelligibility improvement in noise based on a perceptual distortion measure
Cees H. Taal, Richard C. Hendriks, Richard Heusdens |
Comput. Speech Lang. | 3 |
| 2013 | Auto-localization in ad-hoc microphone arraysabstractWe present a method for automatic microphone localization in adhoc microphone arrays. The localization is based on time-of-arrival (TOA) measurements obtained from spatially distributed acoustic events. In practice, measured TOAs are an incomplete representation of the true TOAs due to unknown onset times of the acoustic events and internal delays in the capturing devices and make the localization problem insoluble if not addressed appropriately. The main contribution of the proposed method is an algorithm that identifies and corrects for such internal delays and acoustic event onset times in the measured TOAs. Experimental results using both simulated and real-world data demonstrate the performance of the method and highlight the significance of correct estimation of the internal delays and onset times. Nikolay D. Gaubitch, W. Bastiaan Kleijn, Richard Heusdens |
ICASSP | 3 |
| 2013 | Bi-alternating direction method of multipliersabstractThe alternating-direction method of multipliers (ADMM) has been widely applied in the field of distributed optimization and statistic learning. ADMM iteratively approaches the saddle point of an augmented Lagrangian function by performing three updates per-iteration. In this paper, we propose a bi-alternating direction method of multipliers (BiADMM) that iteratively minimizes an augmented bi-conjugate function. As a result, the convergence of BiADMM is naturally established. Unlike ADMM that always involves three updates per iteration, BiADMM opens up an avenue to perform either two or three updates per iteration, depending on the functional construction. As an application, we consider applying BiADMM for the lasso problem. Experimental results demonstrate the effectiveness of our new method. Guoqiang Zhang 0003, Richard Heusdens |
ICASSP | 2 |
| 2013 | Simplified alternating-direction message passing for dual MAP LP-relaxationabstractThe approximate MAP inference over (factor) graphic models is of great importance in many applications. Due to its simplicity, linear-programming (LP) relaxation has become one of the most popular approaches to approximate MAP. In this paper, we propose a new message passing algorithm for the MAP LP-relaxation problem by using the alternating-direction method of multipliers (ADMM). At each iteration, the new algorithm performs two layers of optimization sequentially, that is node-oriented optimization and factor-oriented optimization. On the other hand, the recently proposed augmented dual LP (ADLP) algorithm, also based on the ADMM, has to perform three layers of optimization. We refer to our new algorithm as the simplified ADLP (SiADLP) algorithm. The design of the SiADLP algorithm stems from a new formulation for the dual LP problem. Experimental results show that the SiADLP algorithm outperforms the ADLP method. Guoqiang Zhang 0003, Richard Heusdens |
ICASSP | 2 |
| 2013 | Proximal alternating-direction message-passing for MAP LP relaxationabstractLinear programming (LP) relaxation for MAP inference over (factor) graphic models is one of the fundamental problems in machine learning. In this paper, we propose a new message-passing algorithm for the MAP LP-relaxation by using the proximal alternating-direction method of multipliers (PADMM). At each iteration, the new algorithm performs two layers of optimization, that is node-oriented optimization and factor-oriented optimization. On the other hand, the recently proposed augmented primal LP (APLP) algorithm, based on the ADMM, has to perform three layers of optimization. Our algorithm simplifies the APLP algorithm by removing one layer of optimization, thus reducing the computational complexities and further accelerating the convergence rate. We refer to our new algorithm as the proximal alternating-direction (PAD) algorithm. Experimental results confirm that the PAD algorithm indeed converges faster than the APLP method. Guoqiang Zhang 0003, Richard Heusdens |
ICASSP | 2 |
| 2013 | Objective Estimation of Speech Quality for Communication SystemsabstractThis paper provides an overview of instrumental models for predicting the quality of speech signals. On the basis of perceptual and cognitive characteristics of the human auditory system which are relevant for quality judgment, approaches are presented which aim at predicting overall quality, intelligibility, or other quality dimensions from measurable parameters or signal characteristics. The approaches are discussed with respect to their underlying principles, showing that perception modeling can significantly improve prediction accuracy. Application examples are presented which make use of these algorithms for offline or online prediction, adaptation, or intelligibility improvement. Sebastian Möller 0001, Richard Heusdens |
Proc. IEEE | 2 |
| 2013 | A generalized Fourier domain: Signal processing framework and applications
Jorge Martínez 0002, Richard Heusdens, Richard C. Hendriks |
Signal Process. | 2 |
| 2012 | A spatio-temporal generalized fourier domain framework to acoustic modeling in enclosed spacesabstractIn this paper, we present a spatio-temporal framework for multichannel acoustic modeling in enclosed spaces. Reverberation occurs when the sound field is enclosed between reflective boundaries (e.g. walls). We model the reverberated sound field by proper sampling of the (generalized) Fourier representation of the free-field sound field. We show that the spatial aliasing introduced by spectral sampling represents all the (damped) reflections. From the samples of the generalized spectrum, we compute the spatio-temporal sound field in the enclosed space with very low-complexity, of O(N log N) per measuring position, with N proportional to the reverberation time. Jorge Martínez 0002, Richard Heusdens, Richard C. Hendriks |
ICASSP | 2 |
| 2012 | A speech preprocessing strategy for intelligibility improvement in noise based on a perceptual distortion measureabstractA speech pre-processing algorithm is presented to improve the speech intelligibility in noise for the near-end listener. The algorithm improves the intelligibility by optimally redistributing the speech energy over time and frequency for a perceptual distortion measure, which is based on a spectro-temporal auditory model. In contrast to spectral-only models, short-time information is taken into account. As a consequence, the algorithm is more sensitive to transient regions, which will therefore receive more amplification compared to stationary vowels. It is known from literature that changing the vowel-transient energy ratio is beneficial for improving speech-intelligibility in noise. Objective intelligibility prediction results show that the proposed method has higher speech intelligibility in noise compared to two other reference methods, without modifying the global speech energy. Cees H. Taal, Richard C. Hendriks, Richard Heusdens |
ICASSP | 3 |
| 2012 | Linear coordinate-descent message-passing for quadratic optimizationabstractIn this paper we propose a new message-passing algorithm for quadratic optimization. The design of the new algorithm is based on linear coordinate-descent between neighboring nodes. The updating messages are in a form of linear functions as compared to the min-sum algorithm of which the messages are in a form of quadratic functions. Therefore, the linear coordinate-descent (LiCD) algorithm has simpler updating rules than the min-sum algorithm. It is shown that when the quadratic matrix is walk-summable, the LiCD algorithm converges. As an application, the LiCD algorithm is utilized in solving general linear systems. The performance of the LiCD algorithm is found empirically to be comparable to that of the min-sum algorithm, but at lower complexity in terms of computation and storage. Guoqiang Zhang 0003, Richard Heusdens |
ICASSP | 2 |
| 2012 | Generalized linear coordinate-descent message-passing for convex optimizationabstractIn this paper we propose a generalized linear coordinate-descent (GLiCD) algorithm for a class of unconstrained convex optimization problems. The considered objective function can be decomposed into edge-functions and node-functions of a graphical model. The messages of the GLiCD algorithm are in a form of linear functions, as compared to the min-sum algorithm of which the form of messages depends on the objective function. Thus, the implementation of the GLiCD algorithm is much simpler than that of the min-sum algorithm. A theorem is stated according to which the algorithm converges to the optimal solution if the objective function satisfies a diagonal-dominant condition. As an application, the GLiCD algorithm is exploited in solving the averaging problem in sensor networks, where the performance is compared to that of the min-sum algorithm. Guoqiang Zhang 0003, Richard Heusdens |
ICASSP | 2 |
| 2012 | Convergence of generalized linear coordinate-descent message-passing for quadratic optimizationabstractWe study the generalized linear coordinate-descent (GLiCD) algorithm for the quadratic optimization problem. As an extension of the linear coordinate-descent (LiCD) algorithm, the GLiCD algorithm incorporates feedback from last iteration in generating new messages. We show that if the amount of feedback signal from last iteration is above a threshold and the GLiCD algorithm converges, it computes the optimal solution. Based on the result, we further show that if the feedback signal is large enough, the GLiCD algorithm is guaranteed to converge. Guoqiang Zhang 0003, Richard Heusdens |
ISIT | 2 |
| 2012 | Linear Coordinate-Descent Message Passing for Quadratic OptimizationabstractIn this letter, we propose a new message-passing algorithm for quadratic optimization. The design of the new algorithm is based on linear coordinate descent between neighboring nodes. The updating messages are in a form of linear functions as compared to the min-sum algorithm of which the messages are in a form of quadratic functions. As a result, the linear coordinate-descent (LiCD) algorithm transmits only one parameter per message as opposed to the min-sum algorithm, which transmits two parameters per message. We show that when the quadratic matrix is walk-summable, the LiCD algorithm converges. By taking the LiCD algorithm as a subroutine, we also fix the convergence issue for a general quadratic matrix. The LiCD algorithm works in either a synchronous or asynchronous message-passing manner. Experimental results show that for a general graph with multiple cycles, the LiCD algorithm has comparable convergence speed to the min-sum algorithm, thereby reducing the number of parameters to be transmitted and the computational complexity. Guoqiang Zhang 0003, Richard Heusdens |
Neural Comput. | 2 |
| 2012 | A Low-Complexity Spectro-Temporal Distortion Measure for Audio Processing ApplicationsabstractPerceptual models exploiting auditory masking are frequently used in audio and speech processing applications like coding and watermarking. In most cases, these models only take into account spectral masking in short-time frames. As a consequence, undesired audible artifacts in the temporal domain may be introduced (e.g., pre-echoes). In this article we present a new low-complexity spectro-temporal distortion measure. The model facilitates the computation of analytic expressions for masking thresholds, while advanced spectro-temporal models typically need computationally demanding adaptive procedures to find an estimate of these masking thresholds. We show that the proposed method gives similar masking predictions as an advanced spectro-temporal model with only a fraction of its computational power. The proposed method is also compared with a spectral-only model by means of a listening test. From this test it can be concluded that for non-stationary frames the spectral model underestimates the audibility of introduced errors and therefore overestimates the masking curve. As a consequence, the system of interest incorrectly assumes that errors are masked in a particular frame, which leads to audible artifacts. This is not the case with the proposed method which correctly detects the errors made in the temporal structure of the signal. Cees H. Taal, Richard C. Hendriks, Richard Heusdens |
IEEE Trans. Speech Audio Process. | 3 |
| 2011 | A Statistical Room Impulse Response Model with Frequency Dependent Reverberation Time for Single-Microphone Late Reverberation SuppressionabstractSingle-channel late reverberation suppression algorithms need estimates of the late reverberance spectral variance (LRSV) in order to suppress the late reverberance. Often the LRSV estimators are derived from a statistical room impulse response (RIR) model. Usually the late reverberation is modeled as a white Gaussian noise sequence with exponentially decaying variance. The whiteness assumption means that the same decay constant is assumed for all frequencies. Since there is generally more absorption of sound energy with increasing frequency, there is a need for RIR models that take this into account. We propose a new statistical time-varying RIR model that consists of a sum of decaying cosine functions with random phases, with a frequency dependent decay constant. We show that the resulting LRSV estimators have the same form as existing ones, but with an inherent frequency dependency of the decay constant. Experiments with real measured RIRs, however, indicate that, for the purpose of reverberation suppression, using a frequency independent decay constant is often sufficiently good. A common assumption in the derivation of LRSV estimatorsisthat thedirect signal and earlyreflections areuncorrelated with the late reverberation. We verify this assumption experimentally on measured RIRs and conclude that it is accurate. Index Terms: speech enhancement, echo suppression, statistical room impulse response models, reverberation time. Jan S. Erkelens, Richard Heusdens |
INTERSPEECH | 2 |
| 2011 | A Generalized Poisson Summation Formula and its Application to Fast Linear ConvolutionabstractIn this letter, a generalized Fourier transform is introduced and its corresponding generalized Poisson summation formula is derived. For discrete, Fourier based, signal processing, this formula shows that a special form of control on the periodic repetitions that occur due to sampling in the reciprocal domain is possible. The present paper is focused on the derivation and analysis of a weighted circular convolution theorem. We use this specific result to compute linear convolutions in the generalized Fourier domain, without the need of zero-padding. This results in faster, more resource-efficient computations. Other techniques that achieve this have been introduced in the past using different approaches. The newly proposed theory however, constitutes a unifying framework to the methods previously published. Jorge Martínez 0002, Richard Heusdens, Richard C. Hendriks |
IEEE Signal Process. Lett. | 2 |
| 2011 | An Algorithm for Intelligibility Prediction of Time-Frequency Weighted Noisy SpeechabstractIn the development process of noise-reduction algorithms, an objective machine-driven intelligibility measure which shows high correlation with speech intelligibility is of great interest. Besides reducing time and costs compared to real listening experiments, an objective intelligibility measure could also help provide answers on how to improve the intelligibility of noisy unprocessed speech. In this paper, a short-time objective intelligibility measure (STOI) is presented, which shows high correlation with the intelligibility of noisy and time-frequency weighted noisy speech (e.g., resulting from noise reduction) of three different listening experiments. In general, STOI showed better correlation with speech intelligibility compared to five other reference objective intelligibility models. In contrast to other conventional intelligibility models which tend to rely on global statistics across entire sentences, STOI is based on shorter time segments (386 ms). Experiments indeed show that it is beneficial to take segment lengths of this order into account. In addition, a free Matlab implementation is provided. Cees H. Taal, Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
IEEE Trans. Speech Audio Process. | 3 |
| 2010 | Noise and late-reverberation suppression in time-varying acoustical environmentsabstractWe consider single-channel blind late-reverberation suppression in noisy and time-varying acoustical environments. Existing estimators for the late reverberant spectral variance (LRSV) are derived assuming the room impulse responses (RIRs) to be time-invariant realizations of a stochastic process. In this paper, we go one step further and analyze time-varying RIRs. We show theoretically that existing LRSV estimators may still be used even when the individual RIR filter taps vary rapidly with time, provided that the reverberation time T60and direct-to-reverberation ratio (DRR) remain nearly constant during an interval of the order of a few frames. We show that these parameters can be taken frequency-independent in DFT-based enhancement algorithms. We estimate them blindly. Experiments with time-varying RIRs validate the analysis and show the importance of accurate estimation of the reverberant spectral variance. Experiments with additive non-stationary noise show the influence of T60and DRR estimation. Jan S. Erkelens, Richard Heusdens |
ICASSP | 2 |
| 2010 | On linear versus non-linear magnitude-DFT estimators and the influence of super-Gaussian speech priorsabstractAlthough the linear mean-squared error (MSE) complex-DFT estimator, i.e., the Wiener filter, is well-known, its magnitude-DFT (MDFT) counterpart has never been considered in the context of speech enhancement. Therefore, certain theoretical questions regarding MDFT estimators remained unanswered. For example, it is unknown to which extend the performance of existing MSE MDFT estimators depends on the chosen speech prior, or on the non-linearity of the estimators. In this paper we present linear MSE MDFT estimators for speech enhancement. In contrast to the linear complex-DFT estimator, the presented linear MSE MDFT estimators do depend on the assumed distribution of the speech DFT coefficients. Based on objective and subjective experiments, it can be concluded that the chosen speech prior, i.e., Gaussian versus super-Gaussian has a significant effect on the performance of MDFT estimators, while the linearity as compared to non-linearity has only a minor influence. Richard C. Hendriks, Richard Heusdens |
ICASSP | 2 |
| 2010 | MMSE based noise PSD tracking with low complexityabstractMost speech enhancement algorithms heavily depend on the noise power spectral density (PSD). Because this quantity is unknown in practice, estimation from the noisy data is necessary. We present a low complexity method for noise PSD estimation. The algorithm is based on a minimum mean-squared error estimator of the noise magnitude-squared DFT coefficients. Compared to minimum statistics based noise tracking, segmental SNR and PESQ are improved for non-stationary noise sources with 1 dB and 0.25 MOS points, respectively. Compared to recently published algorithms, similar good noise tracking performance is obtained, but at a computational complexity that is in the order of a factor 40 lower. Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
ICASSP | 2 |
| 2010 | A short-time objective intelligibility measure for time-frequency weighted noisy speechabstractExisting objective speech-intelligibility measures are suitable for several types of degradation, however, it turns out that they are less appropriate for methods where noisy speech is processed by a time-frequency (TF) weighting, e.g., noise reduction and speech separation. In this paper, we present an objective intelligibility measure, which shows high correlation (rho=0.95) with the intelligibility of both noisy, and TF-weighted noisy speech. The proposed method shows significantly better performance than three other, more sophisticated, objective measures. Furthermore, it is based on an intermediate intelligibility measure for short-time (approximately 400 ms) TF-regions, and uses a simple DFT-based TF-decomposition. In addition, a free Matlab implementation is provided. Cees H. Taal, Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
ICASSP | 3 |
| 2010 | On Low-Complexity Simulation of Multichannel Room Impulse ResponsesabstractIn this letter, we present a method for low-complexity simulation of multichannel room impulse responses (RIRs). Low-complexity RIR methods will become inevitable in next generation communication systems having massive amounts of microphones/loudspeakers. For a room with rigid boundaries, we show that proper sampling of the free-field plenacoustic spectrum results in the solution of the wave equation at any position in the room. We show that the spatial aliasing introduced by spectral sampling represents the wall reflections. These wall reflections are usually modelled, at least in low-complexity simulation algorithms, by the creation of virtual free-field sources outside the room, an image source model commonly referred to as the image method. The image method requiresO(N3) operations per receiver position, whereas the newly proposed method requires onlyO(NlogN) operations per receiver position. Jorge Martínez 0002, Richard Heusdens |
IEEE Signal Process. Lett. | 2 |
| 2010 | Correlation-Based and Model-Based Blind Single-Channel Late-Reverberation Suppression in Noisy Time-Varying Acoustical EnvironmentsabstractThis paper considers suppression of late reverberation and additive noise in single-channel speech recordings. The reverberation introduces long-term correlation in the observed signal. In the first part of this work, we show how this correlation can be used to estimate the late reverberant spectral variance (LRSV) without having to assume a specific model for the room impulse responses (RIRs) while no explicit estimates of RIR model parameters are needed. That makes thiscorrelation-basedapproach more robust against RIR modeling errors. However, the correlation-based method can follow only slow time variations in the RIRs. Existingmodel-basedmethods use statistical models for the RIRs, that depend on one or more parameters that have to be estimated blindly. The common statistical models lead to simple expressions for the LRSV that depend on past values of the spectral variance of the reverberant, noise-free, signal. All existing model-based LRSV estimators in the literature are derived assuming the RIRs to betime-invariant realizationsof a stochastic process. In the second part of this paper, we go one step further and analyzetime-varyingRIRs. We show that in this case the reverberance tends to become decorrelated. We discuss the relations between different RIR models and their corresponding LRSV estimators. We show theoretically that similar simple estimators exist as in the time-invariant case, provided that the reverberation timeT60and direct-to-reverberation ratio (DRR) of the RIRs remain nearly constant during an interval of the order of a few frames. We show that the reverberation time can be taken frequency-bin independent in DFT-based enhancement algorithms. Experiments with time-varying RIRs validate the analysis. Experiments with additive nonstationary noise and time-invariant RIRs show the influence of blind estimation of the reverberation time and the DRR. Jan S. Erkelens, Richard Heusdens |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | n -Channel Asymmetric Entropy-Constrained Multiple-Description Lattice Vector QuantizationabstractThis paper is about the design and analysis of an index-assignment (IA)-based multiple-description coding scheme for the n-channel asymmetric case. We use entropy constrained lattice vector quantization and restrict attention to simple reconstruction functions, which are given by the inverse IA function when all descriptions are received or otherwise by a weighted average of the received descriptions. We consider smooth sources with finite differential entropy rate and MSE fidelity criterion. As in previous designs, our construction is based on nested lattices which are combined through a single IA function. The results are exact under high-resolution conditions and asymptotically as the nesting ratios of the lattices approach infinity. For any n, the design is asymptotically optimal within the class of IA-based schemes. Moreover, in the case of two descriptions and finite lattice vector dimensions greater than one, the performance is strictly better than that of existing designs. In the case of three descriptions, we show that in the limit of large lattice vector dimensions, points on the inner bound of Pradhan can be achieved. Furthermore, for three descriptions and finite lattice vector dimensions, we show that the IA-based approach yields, in the symmetric case, a smaller rate loss than the recently proposed source-splitting approach. Jan Østergaard, Richard Heusdens, Jesper Jensen 0001 |
IEEE Trans. Inf. Theory | 2 |
| 2009 | Single-microphone late-reverberation suppression in noisy speech by exploiting long-term correlation in the DFT domainabstractWe consider blind late-reverberation suppression in speech signals measured with a single microphone in noisy environments. We exploit that reverberant speech shows correlation over longer time spans than clean speech by predicting the contribution of reverberant energy to the current observed spectrum from the enhanced spectra of previous frames. The prediction parameters are recursively updated with estimates of the correlation coefficients between the current reverberant spectrum and enhanced previous spectra. The contributions of late reverberation and noise are suppressed by a standard noise reduction algorithm. The algorithm is shown to decrease the long-term correlation. It achieves significant improvements in segmental speech-to-interference ratio and Bark spectral distortion for typical reverberation times and noise levels, while almost no distortions are introduced in clean speech. Jan S. Erkelens, Richard Heusdens |
ICASSP | 2 |
| 2009 | Fast noise PSD estimation with low complexityabstractAlthough noise PSD estimation is a crucial part of noise reduction algorithms, most noise PSD estimators have problems in tracking non-stationary noise sources. Recently, a noise PSD estimator based on DFT-subspace decompositions was proposed, which improves estimation of the PSD of such noise sources. However, as this approach is based on eigenvalue decompositions per DFT bin, it might be too computationally demanding for low-complexity applications like hearing aids. In this paper we present a method with similar noise tracking performance as the DFT-subspace approach, but with low computational costs. This method is based on computation of high resolution perodiograms, and can estimate the noise PSD when both speech and noise are present in a frequency bin. When combined with a complete noise reduction system, the proposed method can lead to an improvement for non-stationary noise sources of more than 1 dB segmental SNR and 0.3 on a PESQ scale, compared to standard noise tracking methods such as minimum statistics and the quantile based approach, while computational complexity is in the same order of magnitude. Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001, Ulrik Kjems |
ICASSP | 2 |
| 2009 | A low-complexity spectro-temporal based perceptual modelabstractThe use of psychoacoustical masking models for audio coding applications has been wide spread over the past decades. In such applications, it is typically assumed that the original input signal serves as a masker for the distortions that are introduced by the lossy coding method that is used. Up to now, these masking models are mostly based on spectral masking. In this paper, we propose a new perceptual model for audio and speech processing algorithms based on spectro-temporal masking. A sophisticated perceptual model is simplified, such that the eventual distortion measure can be written as a frequency-weighted l2-norm. This yields the same computational complexity as conventional spectral-based methods, but with the preservation of the temporal fine structure of the clean signal. It is shown that the new model can successfully avoid pre-echoes and can correctly predict masking curves for various maskers. Cees H. Taal, Richard Heusdens |
ICASSP | 2 |
| 2009 | Log-spectral magnitude MMSE estimators under super-Gaussian densitiesabstractDespite the fact that histograms of speech DFT coefficients are super-Gaussian, not much attention has been paid to develop estimators under these super-Gaussian distributions in combi-nation with perceptual meaningful distortion measures. In this paper we present log-spectral magnitude MMSE estimators un-der super-Gaussian densities, resulting in an estimator that is perceptually more meaningful and in line with measured his-tograms of speech DFT coefficients. Compared to state-of-the-art reference methods, the presented estimator leads to an im-provement of the segmental SNR in the order of 0.5 dB up to 1 dB. Moreover, listening tests show that the proposed estima-tor leads to significant improvement for the presented estimator over state-of-the-art methods. Index Terms: speech enhancement, log-spectral magnitude MMSE, super-Gaussian Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
INTERSPEECH | 2 |
| 2009 | An evaluation of objective quality measures for speech intelligibility predictionabstractIn this research various objective quality measures are evalu-ated in order to predict the intelligibility for a wide range of non-linearly processed speech signals and speech degraded by additive noise. The obtained results are compared with the pre-diction results of a more advanced perceptual-based model pro-posed by Dau et al. and an objective intelligibility measure, namely the coherence speech intelligibility index (cSII). These tests are performed in order to gain more knowledge between the link of speech-quality and speech-intelligibility and may help us to exploit the extensive research done into the field of speech-quality for speech-intelligibility. It is shown that cSII does not necessarily show better performance compared to con-ventional objective (speech)-quality measures. In general, the DAU-model is the only method with reasonable results for all processing conditions. Index Terms: Speech intelligibility prediction, speech quality, objective Measure. Cees H. Taal, Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001, Ulrik Kjems |
INTERSPEECH | 3 |
| 2009 | On Optimal Multichannel Mean-Squared Error Estimators for Speech EnhancementabstractIn this letter we present discrete Fourier transform (DFT) domain minimum mean-squared error (MMSE) estimators for multichannel noise reduction. The estimators are derived assuming that the clean speech magnitude DFT coefficients are generalized-Gamma distributed. We show that for Gaussian distributed noise DFT coefficients, the optimal filtering approach consists of a concatenation of a minimum variance distortionless response (MVDR) beamformer followed by well-known single-channel MMSE estimators. The multichannel Wiener filter follows as a special case of the presented MSE estimators and is in general suboptimal. For non-Gaussian distributed noise DFT coefficients the resulting spatial filter is in general nonlinear with respect to the noisy microphone signals and cannot be decomposed into an MVDR beamformer and a post-filter. Richard C. Hendriks, Richard Heusdens, Ulrik Kjems, Jesper Jensen 0001 |
IEEE Signal Process. Lett. | 2 |
| 2008 | Fast noise tracking based on recursive smoothing of MMSE noise power estimatesabstractWe consider estimation of the noise spectral variance from speech signals contaminated by highly nonstationary noise sources. In each time frame, for each frequency bin, the noise variance estimate is updated recursively with the minimum mean-square error (MMSE) estimate of the current noise power. For the estimation of the noise power, a spectral gain function is used, which is found by an iterative data-driven training method. The proposed noise tracking method can accurately track fast changes in noise level (up to about 10 dB/s). When compared to the minimum statistics method for various noise sources in a speech enhancement system, improvements in segmental signal-to-noise ratio of more than 1 dB are obtained. Jan S. Erkelens, Richard Heusdens |
ICASSP | 2 |
| 2008 | Comparison of complex-DFT estimators with and without the independence assumption of real and imaginary partsabstractMMSE estimators for DFT-domain based single-microphone speech enhancement can broadly be classified in those that estimate the complex-DFT coefficients and those that estimate the DFT magnitudes. Existing complex-DFT MMSE estimators have generally been derived under assumptions that are in conflict with measured histograms and that are inconsistent with the assumptions made to derive DFT magnitude estimators. Recently it has been shown that these inconsistencies can be eliminated, i.e., no independency has to be assumed between real and imaginary parts of DFT coefficients if the phase of DFT coefficients is assumed uniformly distributed. In this paper we discuss the assumptions that underlie the different complex-DFT estimators and show that the uniform phase assumption matches actual speech data. Furthermore, we show experimentally that the estimators without the independence assumption lead to a lower mean-square error. Richard C. Hendriks, Jan S. Erkelens, Richard Heusdens |
ICASSP | 3 |
| 2008 | On the Estimation of Complex Speech DFT Coefficients Without Assuming Independent Real and Imaginary PartsabstractThis letter considers the estimation of speech signals contaminated by additive noise in the discrete Fourier transform (DFT) domain. Existing complex-DFT estimators assume independency of the real and imaginary parts of the speech DFT coefficients, although this is not in line with measurements. In this letter, we derive some general results on these estimators, under more realistic assumptions. Assuming that speech and noise are independent, speech DFT coefficients have uniform phase, and that noise DFT coefficients have a Gaussian density, we show theoretically that the spectral gain function for speech DFT estimation is real and upper-bounded by the corresponding gain function for spectral magnitude estimation. We also show that the minimum mean-square error (MMSE) estimator of the speech phase equals the noisy phase. No assumptions are made about the distribution of the speech spectral magnitudes. Recently, speech spectral amplitude estimators have been derived under a generalized-Gamma amplitude distribution. As an example, we will derive the corresponding complex-DFT estimators, without making the independence assumption. Jan S. Erkelens, Richard C. Hendriks, Richard Heusdens |
IEEE Signal Process. Lett. | 3 |
| 2008 | Tracking of Nonstationary Noise Based on Data-Driven Recursive Noise Power EstimationabstractThis paper considers estimation of the noise spectral variance from speech signals contaminated by highly nonstationary noise sources. The method can accurately track fast changes in noise power level (up to about 10 dB/s). In each time frame, for each frequency bin, the noise variance estimate is updated recursively with the minimum mean-square error (mmse) estimate of the current noise power. A time- and frequency-dependent smoothing parameter is used, which is varied according to an estimate of speech presence probability. In this way, the amount of speech power leaking into the noise estimates is kept low. For the estimation of the noise power, a spectral gain function is used, which is found by an iterative data-driven training method. The proposed noise tracking method is tested on various stationary and nonstationary noise sources, for a wide range of signal-to-noise ratios, and compared with two state-of-the-art methods. When used in a speech enhancement system, improvements in segmental signal-to-noise ratio of more than 1 dB can be obtained for the most nonstationary noise sources at high noise levels. Jan S. Erkelens, Richard Heusdens |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | Noise Tracking Using DFT Domain Subspace DecompositionsabstractAll discrete Fourier transform (DFT) domain-based speech enhancement gain functions rely on knowledge of the noise power spectral density (PSD). Since the noise PSD is unknown in advance, estimation from the noisy speech signal is necessary. An overestimation of the noise PSD will lead to a loss in speech quality, while an underestimation will lead to an unnecessary high level of residual noise. We present a novel approach for noise tracking, which updates the noise PSD for each DFT coefficient in the presence of both speech and noise. This method is based on the eigenvalue decomposition of correlation matrices that are constructed from time series of noisy DFT coefficients. The presented method is very well capable of tracking gradually changing noise types. In comparison to state-of-the-art noise tracking algorithms the proposed method reduces the estimation error between the estimated and the true noise PSD. In combination with an enhancement system the proposed method improves the segmental SNR with several decibels for gradually changing noise types. Listening experiments show that the proposed system is preferred over the state-of-the-art noise tracking algorithm. Richard C. Hendriks, Jesper Jensen 0001, Richard Heusdens |
IEEE Trans. Speech Audio Process. | 3 |
| 2007 | Analysis and Synthesis of Pseudo-Periodic Job Arrivals in Grids: A Matching Pursuit ApproachabstractPseudo-periodicity is one of the basic job arrival patterns on data-intensive clusters and Grids. In this paper, a signal decomposition methodology called matching pursuit is applied for analysis and synthesis of pseudo-periodic job arrival processes. The matching pursuit decomposition is well localized both in time and frequency, and it is naturally suited for analyzing non-stationary as well as stationary signals. The stationarity of the processes can be quantitatively measured by permutation entropy, with which the relationship between stationarity and modeling complexity is excellently explained. Quantitative methods based on the power spectrum are also provided to measure the degree of periodicity present in the data. Matching pursuit is further shown to be able to extract patterns from signals, which is an attractive feature from a modeling perspective. Real world workload data from production clusters and Grids are used to empirically evaluate the proposed measures and methodologies. Hui Li 0025, Richard Heusdens, Michael Muskulus, Lex Wolters |
CCGRID | 2 |
| 2007 | DFT domain subspace based noise tracking for speech enhancementabstractMost DFT domain based speech enhancement methods are de-pendent on an estimate of the noise power spectral density (PSD). For non-stationary noise sources it is desirable to es-timate the noise PSD also in spectral regions where speech is present. In this paper a new method for noise tracking is pre-sented, based on eigenvalue decompositions of correlation ma-trices that are constructed from time series of noisy DFT coef-ficients. The presented method can estimate the noise PSD at time-frequency points where both speech and noise are present. In comparison to state-of-the-art noise tracking algorithms the proposed algorithm reduces the estimation error between the estimated and the true noise PSD and improves segmental SNR when combined with an enhancement system with several dB. Index Terms: Speech enhancement, noise tracking, DFT do-main subspace decompositions. Richard C. Hendriks, Jesper Jensen 0001, Richard Heusdens |
INTERSPEECH | 3 |
| 2007 | A data-driven approach to optimizing spectral speech enhancement methods for various error criteria
Jan S. Erkelens, Jesper Jensen 0001, Richard Heusdens |
Speech Commun. | 3 |
| 2007 | Minimum Mean-Square Error Estimation of Discrete Fourier Coefficients With Generalized Gamma PriorsabstractThis paper considers techniques for single-channel speech enhancement based on the discrete Fourier transform (DFT). Specifically, we derive minimum mean-square error (MMSE) estimators of speech DFT coefficient magnitudes as well as of complex-valued DFT coefficients based on two classes of generalized gamma distributions, under an additive Gaussian noise assumption. The resulting generalized DFT magnitude estimator has as a special case the existing scheme based on a Rayleigh speech prior, while the complex DFT estimators generalize existing schemes based on Gaussian, Laplacian, and Gamma speech priors. Extensive simulation experiments with speech signals degraded by various additive noise sources verify that significant improvements are possible with the more recent estimators based on super-Gaussian priors. The increase in perceptual evaluation of speech quality (PESQ) over the noisy signals is about 0.5 points for street noise and about 1 point for white noise, nearly independent of input signal-to-noise ratio (SNR). The assumptions made for deriving the complex DFT estimators are less accurate than those for the magnitude estimators, leading to a higher maximum achievable speech quality with the magnitude estimators. Jan S. Erkelens, Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
IEEE Trans. Speech Audio Process. | 3 |
| 2007 | An MMSE Estimator for Speech Enhancement Under a Combined Stochastic-Deterministic Speech ModelabstractAlthough many discrete Fourier transform (DFT) domain-based speech enhancement methods rely on stochastic models to derive clean speech estimators, like the Gaussian and Laplace distribution, certain speech sounds clearly show a more deterministic character. In this paper, we study the use of a deterministic model in combination with the well-known stochastic models for speech enhancement. We derive a minimum mean-square error (MMSE) estimator under a combined stochastic-deterministic speech model with speech presence uncertainty and show that for different distributions of the DFT coefficients the combined stochastic-deterministic speech model leads to improved performance of approximately 0.8 dB segmental signal-to-noise ratio (SNR) over the use of a stochastic model alone. Evaluation with perceptual evaluation of speech quality (PESQ) shows performance improvements of approximately 0.15 on an MOS scale Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | Improved Subspace-Based Single-Channel Speech Enhancement Using Generalized Super-Gaussian PriorsabstractTraditional single-channel subspace-based schemes for speech enhancement rely mostly on linear minimum mean-square error estimators, which are globally optimal only if the Karhunen-Loeacuteve transform (KLT) coefficients of the noise and speech processes are Gaussian distributed. We derive in this paper subspace-based nonlinear estimators assuming that the speech KLT coefficients are distributed according to a generalized super-Gaussian distribution which has as special cases the Laplacian and the two-sided Gamma distribution. As with the traditional linear estimators, the derived estimators are functions of the a priori signal-to-noise ratio (SNR) in the subspaces spanned by the KLT transform vectors. We propose a scheme for estimating these a priori SNRs, which is in fact a generalization of the "decision-directed" approach which is well-known from short-time Fourier transform (STFT)-based enhancement schemes. We show that the proposed a priori SNR estimation scheme leads to a significant reduction of the residual noise level, a conclusion which is confirmed in extensive objective speech quality evaluations as well as subjective tests. We also show that the derived estimators based on the super-Gaussian KLT coefficient distribution lead to improvements for different noise sources and levels as compared to when a Gaussian assumption is imposed Jesper Jensen 0001, Richard Heusdens |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | High-Resolution Spherical Quantization of Sinusoidal ParametersabstractSinusoidal coding is an often employed technique in low bit-rate audio coding. Therefore, methods for efficient quantization of sinusoidal parameters are of great importance. In this paper, we use high-resolution assumptions to derive analytical expressions for the optimal entropy-constrained unrestricted spherical quantizers for the amplitude, phase, and frequency parameters of the sinusoidal model. This is done both for the case of a single sinusoid, and for the more practically relevant case of multiple sinusoids distributed across multiple segments. To account for psychoacoustical effects of the auditory system, a perceptual distortion measure is used. The optimal quantizers minimize a high-resolution approximation of the expected perceptual distortion, while the corresponding quantization indices satisfy an entropy constraint. The quantizers turn out to be flexible and of low complexity, in the sense that they can be determined easily for varying bit rate requirements, without any sort of retraining or iterative procedures. In an objective comparison it is shown that for the squared error distortion measure, the rate-distortion performance of the proposed method is very close to that of the theoretically optimal entropy-constrained vector quantization. Furthermore, for the perceptual distortion measure, the proposed scheme is shown to objectively outperform an existing sinusoidal quantization scheme, where frequency quantization is done independently. Finally, a subjective listening test, in which the proposed scheme is compared to an existing state-of-the-art sinusoidal quantization scheme with fixed quantizers for all input signals, indicates that the proposed scheme leads to an average bit rate reduction of 20%, at the same subjective quality level as the existing scheme Pim Korten, Jesper Jensen 0001, Richard Heusdens |
IEEE Trans. Speech Audio Process. | 3 |
| 2006 | Noise Power Spectrum Estimation for Speech Enhancement Using an Autoregressive Model for Speech Power Spectrum DynamicsabstractIn this paper we propose a method for estimating the non-stationary noise power spectral density (PSD) given a noisy speech signal. The method is based on an autoregressive (AR) model of the speech PSD dynamics combined with a Kalman filtering based noise PSD estimation technique. Objective and subjective performance evaluations show that the speech enhancement scheme utilizing the proposed noise PSD estimation technique achieves significant improvements over a system using a stationary noise estimate as well as compared to a system that uses a noise tracker developed in our previous work Ivo Batina, Jesper Jensen 0001, Richard Heusdens |
ICASSP (3) | 3 |
| 2006 | Speech Enhancement Under a Combined Stochastic-Deterministic ModelabstractMost DFT domain based enhancement methods rely on stochastic models to derive clean speech estimators. In this paper we investigate the use of a deterministic speech model and present an MMSE estimator under a combined stochastic-deterministic speech model. Experimental results show an increase in segmental SNR of 1.18 dB, compared to the use of a stochastic model alone. Furthermore, PESQ evaluations lead to an increase of 0.3 on the MOS scale. Listening tests show a preference for the proposed MMSE estimator under combined stochastic-deterministic speech model Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
ICASSP (1) | 2 |
| 2006 | High Resolution Spherical Quantization of Sinusoids with Harmonically Related FrequenciesabstractSinusoidal coding is an essential tool in low-rate audio coding, and developing an efficient quantization scheme for the sinusoidal parameters is therefore crucial. In this work we derive optimal entropy constrained amplitude, phase and frequency quantizers for sinusoids whose frequencies are harmonically related, with respect to the l2distortion measure. This scheme exploits the harmonic structure of many speech and audio signals in the sense that besides amplitudes and phases, only fundamental frequencies need to be quantized, resulting in a significant decrease in the number of bits assigned to frequency parameters. The asymptotically optimal quantizers minimize a high-resolution approximation of the expected l2distortion while the corresponding quantization indices satisfy an entropy constraint. The quantizers turn out to be flexible and of low complexity, in the sense that they can be determined easily for varying bit rate requirements, without any sort of retraining or iterative procedures. In an objective rate-distortion comparison, the proposed scheme is shown to outperform two variants of a recently proposed scheme, in which all frequency parameters are quantized separately, either directly or differentially Pim Korten, Jesper Jensen 0001, Richard Heusdens |
ICASSP (5) | 3 |
| 2006 | RD Optimal Temporal Noise Shaping for Transform Audio CodingabstractIn this article we investigate rate-distortion optimal temporal noise shaping for transform audio coding. Temporal noise shaping, or TNS, is a technique for reshaping the quantization noise in the time domain through open-loop linear predictive coding of frequency domain coefficients. Traditionally, a selection mechanism based on prediction gain is employed to determine whether it is advantageous to apply TNS or not. Although this method is effective for reducing coding artifacts in transient and speech signals, critical adjustment of the prediction gain threshold is necessary to avoid excessive bit rate demands. We propose the use of TNS in a rate-distortion optimization framework. Within this framework a jointly optimal selection of the prediction filter order and the quantizer for coding the coefficients can be made, such that the perceptual distortion is minimized for a given target rate. Experimental results for an MDCT-based audio coding system are presented and it is shown that TNS within an RD optimization framework outperforms the existing TNS method Omar Niamut, Richard Heusdens |
ICASSP (5) | 2 |
| 2006 | Perceptual Audio Coding Using N-Channel Lattice Vector QuantizationabstractWe consider the problem of reliable distribution of audio over packet-switched networks. We make use of multiple-description coding combined with transform coding in order to obtain robustness towards packet losses. Previous approaches to this problem were restricted to the case of only two descriptions. In this work we use n-channel multiple-description lattice vector quantizers (MD-LVQs), which allow for the possibility of using more than two descriptions. For a given packet-loss probability we find the number of descriptions and the bit allocation between transform coefficients which minimizes a perceptual distortion measure subject to an entropy constraint. The optimal quantizers are presented in closed form, thus avoiding any iterative quantizer design procedures. The theoretical results are verified with numerical computer simulations using audio signals and it is shown that in environments with excessive packet losses it is advantageous to use more than two descriptions. We verify in subjective listening tests that using more than two descriptions lead to signals of perceptually higher quality Jan Østergaard, Omar Niamut, Jesper Jensen 0001, Richard Heusdens |
ICASSP (5) | 4 |
| 2006 | MMSE estimation of complex-valued discrete Fourier coefficients with generalized gamma priorsabstractWe consider DFT based techniques for single-channel speech enhancement. Specifically, we derive minimum mean-square error estimators of clean speech DFT coefficients based on generalized gamma prior probability density functions. Our estimators contain as special cases the well-known Wiener estimator and the more recently derived estimators based on Laplacian and twosided gamma priors. Simulation experiments with speech signals degraded by various additive noise sources verifythat theestimator based on the two-sided gamma prior is close to optimal amongst all the estimators considered in this paper. Jesper Jensen 0001, Richard C. Hendriks, Jan S. Erkelens, Richard Heusdens |
INTERSPEECH | 4 |
| 2006 | Source-Channel Erasure Codes with Lattice Codebooks for Multiple Description CodingabstractIt was recently shown that a subset of the rate distortion region of the symmetric K-channel multiple description coding problem can be achieved by use of (K, k) source-channel erasure codes (SCEC). The construction of the previously proposed SCEC made use of source coding with side information and relied upon random codebooks. In this paper we propose to replace the random codebooks of (K, k) SCEC by structured (lattice) codebooks. We then show that, in certain cases and under high-resolution assumptions, this improves the achievable rate distortion region over the traditional (K, k) SCEC Jan Østergaard, Richard Heusdens, Jesper Jensen 0001 |
ISIT | 2 |
| 2006 | Adaptive Time Segmentation for Improved Speech EnhancementabstractSingle-channel enhancement algorithms are widely used to overcome the degradation of noisy speech signals. Speech enhancement gain functions are typically computed from two quantities, namely, an estimate of the noise power spectrum and of the noisy speech power spectrum. The variance of these power spectral estimates degrades the quality of the enhanced signal and smoothing techniques are, therefore, often used to decrease the variance. In this paper, we present a method to determine the noisy speech power spectrum based on an adaptive time segmentation. More specifically, the proposed algorithm determines for each noisy frame which of the surrounding frames should contribute to the corresponding noisy power spectral estimate. Further, we demonstrate the potential of our adaptive segmentation in both maximum likelihood and decision direction-based speech enhancement methods by making a better estimate of the a priori signal-to-noise ratio (SNR) xi. Objective and subjective experiments show that an adaptive time segmentation leads to significant performance improvements in comparison to the conventionally used fixed segmentations, particularly in transitional regions, where we observe local SNR improvements in the order of 5 dB Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Rate-distortion optimal time-segmentation and redundancy selection for VoIPabstractIn this paper, novel techniques for packet loss robust speech coding are proposed. By exploiting knowledge of the receiving end packet loss concealment algorithm, an existing rate-distortion optimal time-segmentation algorithm is extended to taking packet losses into account. To increase robustness in highly nonstationary signals, the technique is complemented by a redundancy selection scheme. A jointly optimal approach ensures that the complementarity between time-segmentation and redundancies is fully exploited. The performance of the methods is investigated through Monte Carlo simulations under various conditions, such as rate, packet loss probability, and algorithmic delay. Finally, subjective listening tests demonstrate perceptual improvements as compared to conventional adaptive time-segmentation not taking packet losses into account. Christoffer Rødbro, Jesper Jensen 0001, Richard Heusdens |
IEEE Trans. Speech Audio Process. | 3 |
| 2006 | n-channel entropy-constrained multiple-description lattice vector quantizationabstractIn this paper, we derive analytical expressions for the central and side quantizers which, under high-resolution assumptions, minimize the expected distortion of a symmetric multiple-description lattice vector quantization (MD-LVQ) system subject to entropy constraints on the side descriptions for given packet-loss probabilities. We consider a special case of the general n-channel symmetric multiple-description problem where only a single parameter controls the redundancy tradeoffs between the central and the side distortions. Previous work on two-channel MD-LVQ showed that the distortions of the side quantizers can be expressed through the normalized second moment of a sphere. We show here that this is also the case for three-channel MD-LVQ. Furthermore, we conjecture that this is true for the general n-channel MD-LVQ. For given source, target rate, and packet-loss probabilities we find the optimal number of descriptions and construct the MD-LVQ system that minimizes the expected distortion. We verify theoretical expressions by numerical simulations and show in a practical setup that significant performance improvements can be achieved over state-of-the-art two-channel MD-LVQ by using three-channel MD-LVQ. Jan Østergaard, Jesper Jensen 0001, Richard Heusdens |
IEEE Trans. Inf. Theory | 3 |
| 2005 | n-Channel Symmetric Multiple-Description Lattice Vector QuantizationabstractWe derive analytical expressions for the central and side quantizers in an n-channel symmetric multiple-description lattice vector quantizer which, under high-resolution assumptions, minimize the expected distortion subject to entropy constraints on the side descriptions for given packet-loss probabilities. The performance of the central quantizer is lattice dependent whereas the performance of the side quantizers is lattice independent. In fact the normalized second moments of the side quantizers are given by that of an L-dimensional sphere. Furthermore, our analytical results reveal a simple way to determine the optimum number of descriptions. We verify theoretical results with numerical experiments and show that with a packet-loss probability of 5%, a gain of 9.1 dB in MSE over state-of-the-art two-description systems can be achieved when quantizing a two-dimensional unit-variance Gaussian source using a total bit budget of 15 bits/dimension and using three descriptions. With 20% packet loss, a similar experiment reveals an MSE reduction of 10.6 dB when using four descriptions. Jan Østergaard, Jesper Jensen 0001, Richard Heusdens |
DCC | 3 |
| 2005 | Adaptive Time Segmentation of Noisy Speech for Improved Speech EnhancementabstractEnhancement algorithms are widely used to overcome the degradation of noisy speech signals. Most enhancement algorithms require an estimate of the noise and noisy speech power spectra in order to compute the gain function used for the noise suppression. The variance of these power spectral estimates degrades the quality of the enhanced signal and smoothing techniques are therefore often used to decrease the variance. We present a method to determine the noisy speech power spectrum based on an adaptive time segmentation. More specifically, the proposed algorithm determines for each noisy frame which of the surrounding frames should contribute to the corresponding noisy power spectral estimate. Objective and subjective experiments show that an adaptive time segmentation leads to significant performance improvements, particularly in transitional speech regions. Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
ICASSP (1) | 2 |
| 2005 | Jointly optimal time segmentation, component selection and quantization for sinusoidal coding of audio and speechabstractWe propose a rate-distortion optimal algorithm for sinusoidal modeling of audio and speech. The algorithm determines, for a pre-specified target bit-rate, the optimal (variable-length) time segmentation, the optimal distribution of sinusoidal components over the segments and the optimal (scalar) quantizers for quantizing the sinusoid parameters. The optimization is done by jointly optimizing the segment lengths, number of sinusoids and quantizers using high-resolution quantization theory and dynamic programming techniques, which makes it possible to solve the algorithm in polynomial time. A particular advantage of the proposed method is that, given a target bit-rate, it solves the problem of finding the optimal balance between total number of sinusoids and number of bits per sinusoid. Richard Heusdens, Jesper Jensen 0001 |
ICASSP (3) | 1 |
| 2005 | High resolution spherical quantization of sinusoidal parameters using a perceptual distortion measureabstractSinusoidal modelling is a key technology in low rate audio coding, and methods for efficient quantization of sinusoidal parameters are therefore of high importance. We derive analytical formulas for the optimal entropy constrained unrestricted spherical quantizers for amplitude, phase and frequency, using a perceptual distortion measure. This is done both for a single sinusoid, and for multiple sinusoids distributed over multiple segments. The quantizers minimize a high-resolution approximation of the expected distortion, while the corresponding quantization indices satisfy an entropy constraint. The quantizers turn out to be flexible and of low complexity, in the sense that they can be determined easily for varying bit rate requirements, without any sort of retraining or iterative procedures. In objective and subjective comparison tests, the proposed method is shown to outperform an existing state-of-the-art sinusoidal quantization scheme, where quantization of frequency parameters is done independently. Pim Korten, Jesper Jensen 0001, Richard Heusdens |
ICASSP (3) | 3 |
| 2005 | Improved decision directed approach for speech enhancement using an adaptive time segmentationabstractShort-time Fourier transform (STFT) methods are often used to overcome the degradation of speech signals affected by noise. STFT-gain functions are usually expressed as a function of the a priori SNR, say ξ, and good techniques to estimate ξ are of vital importance for the quality of enhanced speech. Often, ξ is estimated using the so-called decision directed approach (DD). However, the DD approach builds on a number of approximations, where certain expected values of signal related quantities are approximated by instantaneous estimates. In this paper we present a method to improve these approximations by combining the DD approach with an adaptive time segmentation. Objective and subjective experiments show that the proposed method leads to significant improvements compared to the conventional DD approach. Furthermore, simulation experiments confirm a decreased amount of non-stationary residual noise. Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
INTERSPEECH | 2 |
| 2005 | n-channel asymmetric multiple-description lattice vector quantizationabstractWe present analytical expressions for optimal entropy-constrained multiple-description lattice vector quantizers which, under high-resolutions assumptions, minimize the expected distortion for given packet-loss probabilities. We consider the asymmetric case where packet-loss probabilities and side entropies are allowed to be unequal and find optimal quantizers for any number of descriptions in any dimension. We show that the normalized second moments of the side-quantizers are given by that of an L-dimensional sphere independent of the choice of lattices. Furthermore, we show that the optimal bit-distribution among the descriptions is not unique. In fact, within certain limits, bits can be arbitrarily distributed Jan Østergaard, Richard Heusdens, Jesper Jensen 0001 |
ISIT | 2 |
| 2005 | Optimal time segmentation for overlap-add systems with variable amount of window overlapabstractIn this letter, we propose a new best basis search algorithm for computing the optimal time segmentation of a signal, given a predefined cost measure. The new algorithm solves a problem that arises when the individual signal segments are windowed and overlap-add is applied between adjacent signal segments. When windows having a variable tail shape are employed, the minimization of a cost measure is faced with dependencies between segmental costs due to varying window overlap. A dynamic programming-based algorithm is presented that takes into account these dependencies. It computes both the optimal split positions and the optimal amount of window overlap at these split positions in polynomial time. The proposed algorithm gives an upper bound to the achievable performance of existing algorithms. Experimental results for a modified discrete cosine transform-based processing system are presented, both for entropy and rate-distortion cost measures. These results show a performance gain over existing schemes at the cost of an increased computational complexity. Omar Niamut, Richard Heusdens |
IEEE Signal Process. Lett. | 2 |
| 2004 | Perceptual linear predictive noise modelling for sinusoid-plus-noise audio codingabstractSinusoidal coding of an audio subject to a bit-rate constraint, in general, results in a noise-like residual signal. This residual signal is of high perceptual importance; reconstruction of audio using the sinusoidal representation only typically results in an artificial sounding reconstruction. We present a new method, called perceptual linear predictive coding (PLPC), where the residual is encoded by applying LPC in the perceptual domain. This method minimizes a perceptual modelling error and therefore represents only residual components that are of perceptual relevance, while automatically discarding components masked by the sinusoidally coded part. Subjective listening tests show that PLPC performs significantly better than ordinary LPC as a sinusoidal residual coding technique. Furthermore, PLPC combined with a flexible segmentation and model order allocation algorithm leads to a significant gain in terms of R/D performance for fragments with fast changing characteristics. Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
ICASSP (4) | 2 |
| 2004 | Entropy constrained multiple description lattice vector quantizationabstractRecently, lattice vector quantizers (LVQ) capable of performing close to known information theoretic bounds were introduced in the area of multiple description coding (MDC). We derive analytical expressions for the central and side quantizers which minimize the expected distortion of an LVQ subject to entropy constraints on the side descriptions for given packet loss probabilities. We show that for certain packet loss probabilities, an optimal LVQ for single descriptions might not be optimal for multiple descriptions. Specifically, we show that the Z/sup 2/ lattice performs better than the A/sub 2/ lattice in some cases. Moreover, our results suggest a practical way of determining which lattice quantizers are optimal for given packet loss probabilities. Jan Østergaard, Jesper Jensen 0001, Richard Heusdens |
ICASSP (4) | 3 |
| 2004 | Adaptive time-segmentation for speech coding with limited delayabstractWe investigate the trade-off between delay and signal quality in adaptive time-segmentation for speech coding. A variable rate sinusoidal coder with adaptive segmentation and bit allocations is proposed and implemented with specifiable look-ahead. Objective and subjective results indicate that adaptive time-segmentation is advantageous even with low delay (30 ms), and that quality only increases with the delay until approximately 100 ms. Christoffer Rødbro, Jesper Jensen 0001, Richard Heusdens |
ICASSP (1) | 3 |
| 2004 | A perceptual subspace approach for modeling of speech and audio signals with damped sinusoidsabstractThe problem of modeling a signal segment as a sum of exponentially damped sinusoidal components arises in many different application areas, including speech and audio processing. Often, model parameters are estimated using subspace based techniques which arrange the input signal in a structured matrix and exploit the so-called shift-invariance property related to certain vector spaces of the input matrix. A problem with this class of estimation algorithms, when used for speech and audio processing, is that the perceptual importance of the sinusoidal components is not taken into account. In this work we propose a solution to this problem. In particular, we show how to combine well-known subspace based estimation techniques with a recently developed perceptual distortion measure, in order to obtain an algorithm for extracting perceptually relevant model components. In analysis-synthesis experiments with wideband audio signals, objective and subjective evaluations show that the proposed algorithm improves perceived signal quality considerable over traditional subspace based analysis methods. Jesper Jensen 0001, Richard Heusdens, Søren Holdt Jensen |
IEEE Trans. Speech Audio Process. | 2 |
| 2003 | A perceptual subspace method for sinusoidal speech and audio modelingabstractThe problem of modeling a signal segment as a sum of exponentially damped sinusoidal components is of interest in a wide range of fields, including speech and audio processing. Often, model parameters are estimated using subspace based techniques that exploit the so-called shift-invariance property. A drawback of these estimation techniques in relation to speech and audio processing is that the perceptual relevance of the model components is not taken into account. In this paper we show how to combine well-known subspace based estimation techniques with a recently developed perceptual distortion measure, to obtain an algorithm for extracting perceptually relevant model components. In analysis-synthesis experiments with wideband audio signals, objective and subjective evaluations show that the proposed algorithm improves perceived signal quality considerably over traditional subspace based analysis methods. Jesper Jensen 0001, Richard Heusdens, Søren Holdt Jensen |
ICASSP (5) | 2 |
| 2003 | Flexible frequency decompositions for cosine-modulated filter banksabstractWe investigate the use of nonuniform cosine-modulated filter banks for audio coding. A rate-distortion framework is employed, similar to the work in Herley et al. (1994), to select the filter bank structure from a large library of possible frequency decompositions. A new flexible frequency decomposition algorithm is proposed that jointly optimizes the filter bank structure and the bit allocation over the subband channels. Experimental results for both synthetic and real audio signals are provided. The new algorithm shows significant improvements in comparison with fixed uniform frequency decompositions, but special care has to be taken to reduce the size of the decomposition overhead. Omar Niamut, Richard Heusdens |
ICASSP (5) | 2 |
| 2003 | Schemes for optimal frequency-differential encoding of sinusoidal model parameters
Jesper Jensen 0001, Richard Heusdens |
Signal Process. | 2 |
| 2003 | Subband merging in cosine-modulated filter banksabstractRecently, a new method for constructing nonuniform modulated lapped transforms (MLTs) was introduced, by combining subband filters of a uniform MLT. The design, however, was restricted to combining two or four subband filters only, and no systematic design procedure was given. In this letter, we propose an extension to the above-mentioned method that allows arbitrary numbers of subbands to be combined in a systematic way. We investigate the general case of combining filters in arbitrary cosine-modulated filter banks, and give conditions on how to combine the constituent filters such that the resulting nonuniform filter banks have suitable frequency responses. Omar Niamut, Richard Heusdens |
IEEE Signal Process. Lett. | 2 |
| 2002 | Rate-distortion optimal sinusoidal modeling of audio and speech using psychoacoustical matching pursuitsabstractIn this paper, we propose a rate-distortion optimal algorithm for sinusoidal modeling of audio and speech. The algorithm uses a variable-length analysis window where the total number of sinusoids needed to model the source signal is optimally distributed over the segments. To account for human auditory perception, we use a new perceptually relevant distortion measure which is combined with the psychoacoustical matching pursuit algorithm to select the desired sinusoidal components. We discuss the encoding of the segmentation information and show how to reduce this overhead by restricting the minimum and maximum segment size of the constituent segments. Although this restricts the number of possible partitionings of the input signal, we still have a high accuracy in time at which new segments can start. By doing so, we can decrease the segmentation overhead by 50%, almost without loss of coding efficiency and without introducing pre-echoes. Richard Heusdens, Steven van de Par |
ICASSP | 1 |
| 2002 | Optimal frequency-differential encoding of sinusoidal model parametersabstractSinusoidal coding has proven to be efficient for low bit-rate audio coding. In this paper we consider schemes for frequency-differential (FD) encoding of the sinusoidal model parameters. For a given signal frame, the parameters of a sinusoidal component may be encoded either differentially relative to other components in the same frame, or directly, i.e., without differential encoding. Using basic tools from graph theory, two algorithms are derived for finding bit rate optimal combinations of direct and differential encoding of the sinusoidal parameters. In simulation experiments with audio signals, the algorithms showed bit-rate reductions of up to 27% relative to direct encoding. Furthermore, when compared to a commonly used FD encoding scheme, the proposed algorithms achieved bit rate reductions of up to 7%. Jesper Jensen 0001, Richard Heusdens |
ICASSP | 2 |
| 2002 | A new psychoacoustical masking model for audio coding applicationsabstractThe use of psychoacoustical masking models for audio coding applications has been wide spread over the past decades. In such applications, it is typically assumed that the original input signal serves as a masker for the distortions that are introduced by the lossy coding method that is used. Such masking models are based on the peripheral bandpass filtering properties of the auditory system and basically evaluate the distortion-to-masker ratio within each auditory filter. Up to now these models have been based on the assumption that the masking of distortions is governed by the auditory filter for which the ratio between distortion and masker is largest. This assumption, however, is not in line with some new findings within the field of psychoacoustics. A more accurate assumption would be that the human auditory system is able to integrate distortions that are present within a range of auditory filters. In this contribution a new model is presented which is in line with new psychoacoustical studies and which is suitable for application within an audio codec. Although this model can be used to derive a masking curve, the model also gives a measure for the detectability of distortions provided that distortions are not too large. Steven van de Par, Armin Kohlrausch, Ghassan Charestan, Richard Heusdens |
ICASSP | 4 |
| 2002 | Sinusoidal modeling using psychoacoustic-adaptive matching pursuitsabstractWe propose a segment-based matching-pursuit algorithm where the psychoacoustical properties of the human auditory system are taken into account. Rather than scaling the dictionary elements according to auditory perception, we define a psychoacoustic-adaptive norm on the signal space that can be used for assigning the dictionary elements to the individual segments in a rate-distortion optimal way. The new algorithm is asymptotically equal to signal-to-mask-ratio-based algorithms in the limit of infinite-analysis window length. However, the new algorithm provides a significantly improved selection of the dictionary elements for finite window length. Richard Heusdens, Renat Vafin, W. Bastiaan Kleijn |
IEEE Signal Process. Lett. | 1 |
| 2001 | Sinusoidal modeling of audio and speech using psychoacoustic-adaptive matching pursuitsabstractWe propose a segment-based matching pursuit algorithm where the psychoacoustical properties of the human auditory system are taken into account. Rather than scaling the dictionary elements according to auditory perception, we define a psychoacoustic-adaptive norm on the signal space which can be used for assigning the dictionary elements to the individual segments in a rate-distortion optimal manner. The new algorithm is asymptotically equal to signal-to-mask ratio based algorithms in the limit of infinite analysis window length. However, the new algorithm provides a significantly improved selection of the dictionary elements for finite window length. Richard Heusdens, Renat Vafin, W. Bastiaan Kleijn |
ICASSP | 1 |
| 2001 | Modifying transients for efficient coding of audioabstractWe propose a method for efficient representation of transients in audio signals. We estimate the transient component of an original audio signal and modify the locations of the transients in such a way that the transients can occur only at locations defined by a relatively coarse time grid. This procedure allows an efficient representation of transients with damped sinusoids. We also verify that the introduced modifications do not result in a perceptual difference between the original and the modified audio signals. Renat Vafin, Richard Heusdens, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2000 | Design of orthogonal and biorthogonal lapped transforms satisfying perception related constraintsabstractWe propose a new efficient method for the design of orthogonal and biorthogonal lapped transforms for image coding applications. It is shown how perception related constraints such as decay and smoothness of the filters' impulse responses can be incorporated in the optimization procedure. A decomposition of lapped transforms (orthogonal and biorthogonal) with 50% overlap leads to an efficient recursive optimization procedure, which is robust with respect to initial solutions. The importance of this decomposition lies in the fact that it allows to decouple the design of the even-symmetric and the odd-symmetric filters and hence drastically reduces the number of variables to be optimized. It furthermore reveals all the variables predetermined by perception related and coding-efficiency related constraints imposed on the filters. We present design and coding examples demonstrating the perceptual performance and the rate distortion performance of the resulting transforms. Helmut Bölcskei, Richard Heusdens, Hendrik Theunis, Augustus J. E. M. Janssen |
IEEE Trans. Image Process. | 2 |
| 1998 | Robust exponential modeling of audio signalsabstractWe present a numerically robust method for modeling audio signals which is based on a exponential data model. This model is a generalization of the classical sinusoidal model in the sense that it allows the amplitude of the sinusoids to evolve exponentially. We show that, using this model, so called attacks can be represented very efficiently and we propose an algorithm for finding the exponentials in a robust way. Moreover, we show that by using a proper segmentation of the input data into variable length segments the signal-to-noise ratio can be drastically improved as compared to a fixed-length analysis. Joost Nieuwenhuijse, Richard Heusdens, Ed F. Deprettere |
ICASSP | 2 |
| 1996 | Design of lapped orthogonal transformsabstractWe investigate the design of lapped orthogonal transforms for data compression of images. We present some properties and new results of paraunitary filter banks. We concentrate on the case where the filter length L=2K, where K is the number of channels. The aim is to design perceptually relevant filters, i.e., linear-phase filters that smoothly decay to zero at the boundaries. Richard Heusdens |
IEEE Trans. Image Process. | 1 |
| 1993 | Subband filtering: Cordic modulation and systolic quadrature mirror filter treeabstractThe decomposition (analysis) of a finite-energy signal into a relatively small number of mutually independent signals which allows reconstruction (synthesis) of the original signal is called subband filtering. Subbands can be processed in parallel or recursively. In the latter case, one obtains a so-called quadrature mirror filter tree. The former case leads to cosine-modulated filter banks. The authors present a Cordic based cosine-modulated bank and a systolic algorithm for quadrature mirror tree filtering.> Ed F. Deprettere, Richard Heusdens, Hendrik Theunis |
ASAP | 2 |