VLDB 2026 Research / reviewers in the wild / expert
Richard C. Hendriks
dblp:39/1030 · also Richard Christian Hendriks
· DBLP profile ↗
75ranked-venue papers
20as first author
12since 2021 · last 2026
0000-0001-8297-0251ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 43 · 14 first-author · 6 since 2021Artificial intelligence and machine learning · 37 · 9 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Optimal pilot design for OTFS in linear time-varying channelsabstractThis paper investigates the positioning of the pilot symbols, as well as the power distribution between the pilot and the communication symbols for the orthogonal time frequency space (OTFS) modulation scheme. We analyze the pilot placements that minimize the mean squared error (MSE) in estimating the channel taps. This allows us to identify two new pilot allocations for OTFS that save approximately 50% of the pilot overhead compared to existing allocations. In addition, we optimize the average channel capacity by adjusting the power distribution. We show that this leads to a significant increase in average capacity. The results provide valuable guidance for designing the OTFS parameters to achieve maximum capacity. Numerical simulations are performed to validate the findings. Ids Van der Werf, Richard Heusdens, Richard C. Hendriks, Geert Leus |
Signal Process. | 3 |
| 2024 | Tensor Decomposition-Based Data Fusion for Biomarker Extraction from Multiple EEG ExperimentsabstractThe pursuit of sensitive and dependable biomarkers capable of capturing the neural processes associated with cognition is a prominent area of interest. Event-related potentials (ERPs) hold significant promise for assessing cognitive dysfunction in various neurological disorders. However, existing data analysis techniques often underutilize the available data and may benefit from potential enhancements. In this paper, we investigate biomarker extraction methods based on two ERP experiments. First, we derive average ERPs from the electroencephalography (EEG) recorded during each experiment and store them in third-order tensors with subjects, channels and time samples along the three modes. Then, we extract biomarkers from these datasets via tensor decompositions. We compare single tensor decompositions and joint tensor decompositions that fuse the data from the individual tensors. In a simulated ERP experiment we compare the benefits and limitations of different tensor-based data fusion methods. Finally, we investigate their performance on a real dataset obtained from schizophrenia patients. K. R. Stunnenberg, Richard C. Hendriks, J. L. Vroegop, M. L. Adank, Borbála Hunyadi |
ICASSP | 2 |
| 2024 | On the equivalence of OSDM and OTFSabstractIn this paper, we show the mathematical equivalence of two popular modulation schemes: OSDM and OTFS. The former is mainly used in underwater acoustic communications, while the latter scheme is a promising modulation technique in radio-frequency communications. Although literature suggests a link between the two modulation schemes by connecting them to related modulation schemes like V-OFDM and A-OFDM, to the best of the authors’ knowledge, a direct mathematical comparison between the schemes has not been presented yet. The main purpose of this paper is therefore to show the mathematical equivalence of the two schemes. In addition, by combining the knowledge of acoustic and radio-frequency communications, we give insight in the performance of OSDM/OTFS in terms of intersymbol interference (ISI) and intercarrier interference (ICI) by analyzing its signal structure. Ids Van der Werf, Henry Dol, Koen Blom, Richard Heusdens, Richard C. Hendriks, Geert Leus |
Signal Process. | 5 |
| 2024 | Block-Based Perceptually Adaptive Sound Zones With Reproduction Error ConstraintsabstractSound zone algorithms control the inputs to a loudspeaker array such that spatially distinct zones, each with separate audio content, are created. This work proposes a sound zone approach which includes a model of human auditory perception in the optimization problem designing the loudspeaker control filters. The control filters are therefore optimized directly for human experience, rather than by proxy through sound pressure, as is done in typical approaches. The proposed optimization problem features a perceptually weighted constraint on the bright zone reproduction error, which allows the user of the algorithm to specify the desired bright zone quality. The proposed method achieves 2 to 4 dB of additional acoustic contrast and is expected to yield less distracting dark-zone interference for the same perceived quality when compared to a traditional approach. Niels de Koeijer, Martin Bo Møller, Jorge Martínez 0002, Pablo Martínez-Nuevo, Richard C. Hendriks |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2024 | Binaural Beamforming Taking Into Account Spatial Release From MaskingabstractHearing impairment is a prevalent problem with daily challenges like impaired speech intelligibility and sound localisation. One of the shortcomings of spatial filtering in hearing aids is that speech intelligibility is often not optimised directly, meaning that different auditory processes contributing to intelligibility are often not considered. One example is the perceptual phenomenon known as spatial release from masking (SRM). This paper develops a signal model that explicitly considers SRM in the beamforming design, achieved by transforming the binaural intelligibility prediction model (BSIM) into a signal processing framework. The resulting extended signal model is used to analyse the performance of reference beamformers and design a novel beamformer that more closely considers how the auditory system perceives binaural sound. It can be shown that the binaural minimum variance distortionless response (BMVDR) beamformer is also an optimal solution for the extended, perceived model, suggesting that SRM does not play a significant role in intelligibility enhancement after optimal beamforming. However, the optimal beamformer is no longer unique in the extended signal model. The additional secondary degrees of freedom can be used to preserve binaural cues of interfering sources while still achieving the same perceived performance of the BMVDR beamformer, though with a possible high sensitivity to intelligibility model mismatch errors. Johannes W. de Vries, Steven van de Par, Geert Leus, Richard Heusdens, Richard C. Hendriks |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | Noise PSD Insensitive RTF Estimation in a Reverberant and Noisy EnvironmentabstractSpatial filtering techniques typically rely on estimates of the target relative transfer function (RTF). However, the target speech signal is typically corrupted by late reverberation and ambient noise, which complicates RTF estimation. Existing methods subtract the noise covariance matrix to obtain the target plus late reverberation covariance matrix, from where the RTF is estimated. However, the noise covariance matrix is typically unknown. More specifically, the noise power spectral density (PSD) is typically unknown, while the spatial coherence matrix can be assumed known as it might remain time-invariant for a longer time. Using the spatial coherence matrices we simplify the signal model such that the off-diagonal elements are not affected by the PSDs of the late reverberation and the ambient noise. Then we use these elements to estimate the target covariance matrix, from where the RTF can be obtained. Hence, the resulting estimate of the RTF is insensitive to the noise PSD. Experiments demonstrate the estimation performance of our proposed method. Changheng Li, Richard C. Hendriks |
ICASSP | 2 |
| 2023 | Estimation of Cardiac Fibre Direction Based on Activation MapsabstractEstimating tissue conductivity parameters from electrograms (EGMs) could be an important tool for diagnosing and treating heart rhythm disorders such as atrial fibrillation (AF). One of these parameters is the fibre direction, often assumed to be known in conductivity estimation methods. In this paper, a novel method to estimate the fibre direction from EGMs is presented. This method is based on local conduction slowness vectors of a propagating activation wave. These conduction slowness vectors follow an elliptical pattern that depends on the underlying conductivity parameters. The fibre direction and conductivity anisotropy ratio can therefore be estimated by fitting an ellipse to the conduction slowness vectors. Applying the presented method on simulated data shows that it can estimate the fibre direction more accurately than existing methods, and that its performance depends mostly on the range of wavefront directions present in the measurement area. The main advantage of the presented method is that it still functions relatively well in the presence of conduction blocks, as long as the surrounding tissue is approximately homogeneous. Johannes W. de Vries, Natasja M. S. de Groot, Richard C. Hendriks |
ICASSP | 4 |
| 2023 | Alternating Least-Squares-Based Microphone Array Parameter Estimation for a Single-Source Reverberant and Noisy Acoustic ScenarioabstractAcoustic-scene-related parameters such as relative transfer functions (RTFs) and power spectral densities (PSDs) of the target source, late reverberation and ambient noise are essential for microphone array signal processing and are challenging to estimate. Existing methods typically only estimate a subset of the parameters by assuming the other parameters are known. This can lead to unmatched scenarios and reduced estimation performance on the parameters of interest. Moreover, many methods process time frames independently, despite they share common information such as the same RTF. In this work, we consider a noisy scenario by modelling the noise component as a spatially homogeneous sound field with a time-invariant spatial coherence matrix and time-varying PSD. We first modify an existing alternating least squares (ALS) method to obtain more accurate estimates using a single time frame. Then, we extend the method to use multiple time frames that share the same RTF. Furthermore, we propose more robust constraints on the PSDs to avoid large estimation errors. We compare our proposed methods to several reference methods, among which the state-of-the-art simultaneously confirmatory factor analysis (SCFA) method, a recently developed joint maximum likelihood estimation (JMLE) method and an existing ALS-based method. The experimental results in terms of estimation accuracy, noise reduction performance, predicted speech quality, and predicted speech intelligibility demonstrate that our proposed ALS-based methods achieve similar performance compared to the state-of-the-art SCFA method. Both the proposed ALS-based methods and the SCFA method outperform the existing ALS-based method in all scenarios and outperform the JMLE method particularly in low SNR scenarios. Moreover, in terms of computational complexity, our proposed methods are the least complex of all reference methods. This is confirmed by the measured processing time, which is significantly lower than for SCFA. Changheng Li, Richard C. Hendriks |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Joint Maximum Likelihood Estimation of Microphone Array Parameters for a Reverberant Single Source ScenarioabstractEstimation of the acoustic-scene related parameters such as relative transfer functions (RTFs) from source to microphones, source power spectral densities (PSDs) and PSDs of the late reverberation is essential and also challenging. Existing maximum likelihood estimators typically consider only subsets of these parameters and use each time frame separately. In this paper we explicitly focus on the single source scenario and first propose a joint maximum likelihood estimator (MLE) to estimate all parameters jointly using a single time frame. Since the RTFs are typically invariant for a number of consecutive time frames we also propose a joint maximum likelihood estimator (MLE) using multiple time frames which has similar estimation performance compared to a recently proposed reference algorithm called simultaneously confirmatory factor analysis (SCFA), but at a much lower complexity. Moreover, we present experimental results which demonstrate that the estimation accuracy, together with the performance of noise reduction, speech quality and speech intelligibility, of our proposed joint MLE outperform those of existing MLE based approaches that use only a single time frame. Changheng Li, Jorge Martínez 0002, Richard C. Hendriks |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Low Complex Accurate Multi-Source RTF EstimationabstractMany multi-microphone algorithms depend on knowing the relative acoustic transfer functions (RTFs) of the individual sound sources in the acoustic scene. However, accurate joint RTF estimation for multiple sources is a challenging problem. Existing methods to jointly estimate the RTF for multiple sources have either no satisfying performance, or, suffer from a very large computational complexity. In this paper, we propose a method for robust estimation of the individual RTFs in a multi-source acoustic scenario. The presented algorithm is based on linear algebraic concepts and therefore of lower computational complexity compared to a recently presented state-of-the-art algorithm, while having a similar performance. Experimental results are presented to demonstrate the RTF estimation performance as well as the noise reduction performance when combining the estimated RTFs with a beamformer. Changheng Li, Jorge Martínez 0002, Richard C. Hendriks |
ICASSP | 3 |
| 2021 | Localization Based on Enhanced Low Frequency Interaural Level DifferenceabstractThe processing of low-frequency interaural time differences is found to be problematic among hearing-impaired people. The current generation of beamformers does not consider this deficiency. In an attempt to tackle this issue, we propose to replace the inaudible interaural time differences in the low-frequency region with the interaural level differences. In addition, a beamformer is introduced and analyzed, which enhances the low-frequency interaural level differences of the sound sources using a near-field transformation. The proposed beamforming problem is relaxed to a convex problem using semi-definite relaxation. The instrumental analysis suggests that the low-frequency interaural level differences are enhanced without hindering the provided intelligibility. A psychoacoustic localization test is done using a listening experiment, which suggests that the replacement of time differences into level differences improves the localization performance of normal-hearing listeners for an anechoic scene but not for a reverberant scene. Metin Calis, Steven van de Par, Richard Heusdens, Richard C. Hendriks |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | A Study on Reference Microphone Selection for Multi-Microphone Speech EnhancementabstractMulti-microphone speech enhancement methods typically require a reference position with respect to which the target signal is estimated. Often, this reference position is arbitrarily chosen as one of the reference microphones. However, it has been shown that the choice of the reference microphone can have a significant impact on the final noise reduction performance. In this paper, we therefore theoretically analyze the impact of selecting a reference on the noise reduction performance with near-end noise being taken into account. Following the generalized eigenvalue decomposition (GEVD) based optimal variable span filtering framework, we find that for any linear beamformer, the output signal-to-noise ratio (SNR) taking both the near-end and far-end noise into account is reference dependent. Only when the near-end noise is neglected, the output SNR of rank-1 beamformers does not depend on the reference position. However, in general for rank-r beamformers with r > 1 (e.g., the multichannel Wiener filter) the performance does depend on the reference position. Based on these, we propose an optimal algorithm for microphone reference selection that maximizes the output SNR. In addition, we propose a lower-complexity algorithm that is still optimal for rank-1 beamformers, but sub-optimal for the general r > 1 rank beamformers. Experiments using a simulated microphone array validate the effectiveness of both proposed methods and show that in terms of quality, several dB can be gained by selecting the proper reference microphone. Jie Zhang 0042, Li-Rong Dai 0001, Richard C. Hendriks |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | Spatially Correct Rate-Constrained Noise Reduction for Binaural Hearing Aids in Wireless Acoustic Sensor NetworksabstractCompared to monaural hearing aids (HAs), binaural hearing aid systems, in which there is a communication link between the two devices, have improved noise reduction capabilities and the ability to preserve binaural spatial information. However, the limited HA battery lifetime puts constraints on the amount of information that can be shared between the two devices. In other words, the rate of transmission between the devices is an important constraint that needs to be considered, while preserving the spatial information. In this article, a linearly constrained noise reduction problem is proposed, which jointly finds the optimal rate allocation and the optimal estimation (beamforming) weights across all sensors and frequencies, while preserving the binaural spatial cues of point sources. The proposed method considers a rate constraint together with linear constraints to preserve the binaural spatial cues of point sources. Minimizing the mean square error on the estimated target speech at the left and the right side beamformers, the optimal weights are found to be rate-constrained linearly constrained minimum variance (LCMV) filters, and the optimal rates are found to be the solutions to a set of reverse water filling problems. The performance of the proposed method is evaluated using the averaged binaural signal-to-noise ratio (SNR), the interaural level difference (ILD) error and the interaural time difference (ITD) error. The results show that the proposed method outperforms spatially correct noise reduction approaches that use naive/random rate allocation strategies. Jamal Amini, Richard C. Hendriks, Richard Heusdens, Meng Guo 0001, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Rate-Constrained Noise Reduction in Wireless Acoustic Sensor NetworksabstractWireless acoustic sensor networks (WASNs) can be used for centralized multi-microphone noise reduction, where the processing is done in a fusion center (FC). To perform the noise reduction, the data needs to be transmitted to the FC. Considering the limited battery life of the devices in a WASN, the total data rate at which the FC can communicate with the different network devices should be constrained. In this article, we propose a rate-constrained multi-microphone noise reduction algorithm, which jointly finds the best rate allocation and estimation weights for the microphones across all frequencies. The optimal linear estimators are found to be the quantized Wiener filters, and the rates are the solutions to a filter-dependent reverse water-filling problem. The performance of the proposed framework is evaluated using simulations in terms of mean square error and predicted speech intelligibility. The results show that the proposed method is very close in performance to that of the existing optimal method based on discrete optimization. However, the proposed approach can do this at a much lower complexity, while the existing optimal reference method needs a non-tractable exhaustive search to find the best rate allocation across microphones. Jamal Amini, Richard C. Hendriks, Richard Heusdens, Meng Guo 0001, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Binaural Beamforming Based on Automatic Interferer SelectionabstractBinaural cues are important for sound localization. In addition, spatially separated sound sources are more intelligible than when they are co-located. Binaural cue preservation in multi-microphone hearing assistive devices is therefore important for the user's listening experience and safety. A number of linearly-constrained-minimum-variance (LCMV) based methods exist for this purpose. These are all limited in the number of sources for which they can preserve the binaural cues. We propose a method of automatically selecting the most important interfering sources using convex optimization. The proposed method is compared, using simulation experiments, to existing methods in terms of noise suppression and localization errors. It improves the performance of the joint binaural LCMV beam-former, by giving it more degrees of freedom for noise reduction and allows a larger number of (virtual) sources present in the scene. Costas A. Kokke, Richard C. Hendriks, Andreas I. Koutrouvelis |
ICASSP | 2 |
| 2019 | A Novel Binaural Beamforming Scheme with Low Complexity Minimizing Binaural-cue DistortionsabstractWhile the majority of binaural beamformers aim to minimize the output noise power while (approximately) preserving the binaural cues of the sources using constraints, we propose in this paper to minimize the binaural-cue distortions of the sources in the acoustic scene, such that the output noise power is below a predefined threshold. This new problem formulation is a convex QCQP problem, which leads to an efficient trade-off between noise reduction, binaural-cue preservation and complexity. In particular, the proposed beamformer provides a better trade-off between noise reduction and binaural-cue preservation (in terms of interaural level and phase differences) compared to the well-known binaural minimum variance distortionless response-η beamformer. Andreas I. Koutrouvelis, Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001, Meng Guo 0001 |
ICASSP | 2 |
| 2019 | Distributed Rate-Constrained LCMV BeamformingabstractIn this letter, we propose a decentralized framework for rate-distributed linearly constrained minimum variance (LCMV) beamforming in wireless acoustic sensor networks. To save the energy usage within the network, we propose to minimize the transmission cost and put a constraint on the noise reduction performance. Subsequently, we decentralize the obtained LCMV filter structure by exploiting an imposed block diagonal form of the noise correlation matrix. As a result, the beamformer weights are calculated in a decentralized fashion and each node can determine its quantization rate locally. Finally, numerical results validate the proposed method. Jie Zhang 0042, Andreas I. Koutrouvelis, Richard Heusdens, Richard C. Hendriks |
IEEE Signal Process. Lett. | 4 |
| 2019 | Asymmetric Coding for Rate-Constrained Noise Reduction in Binaural Hearing AidsabstractBinaural hearing aids (HAs) can potentially perform advanced noise reduction algorithms, leading to an improvement over monaural/bilateral HAs. Due to the limited transmission capacities between the HAs and given knowledge of the complete joint noisy signal statistics, the optimal rate-constrained beamforming strategy is known from the literature. However, as these joint statistics are unknown in practice, sub-optimal strategies have been presented. In this paper, we present a unified framework to study the performance of these existing optimal and sub-optimal rate-constrained beamforming methods for binaural HAs. Moreover, we propose to use an asymmetric sequential coding scheme to estimate the joint statistics between the microphones in the two HAs. We show that under certain assumptions, this leads to sub-optimal performance in one HA but allows to obtain the truly optimal performance in the second HA. Based on the mean square error distortion measure, we evaluate the performance improvement between monaural beamforming (no communication) and the proposed scheme, as well as the optimal and the existing sub-optimal strategies in terms of the information bit-rate. The results show that the proposed method outperforms existing practical approaches in most scenarios, especially at middle rates and high rates, without having the prior knowledge of the joint statistics. Jamal Amini, Richard C. Hendriks, Richard Heusdens, Meng Guo 0001, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | A Convex Approximation of the Relaxed Binaural Beamforming Optimization ProblemabstractThe recently proposed relaxed binaural beamforming (RBB) optimization problem provides a flexible tradeoff between noise suppression and binaural-cue preservation of the sound sources in the acoustic scene. It minimizes the output noise power, under the constraints, which guarantee that the target remains unchanged after processing and the binaural-cue distortions of the acoustic sources will be less than a user-defined threshold. However, the RBB problem is a computationally demanding non convex optimization problem. The only existing suboptimal method which approximately solves the RBB is a successive convex optimization (SCO) method which, typically, requires to solve multiple convex optimization problems per frequency bin, in order to converge. Convergence is achieved when all constraints of the RBB optimization problem are satisfied. In this paper, we propose a semidefinite convex relaxation (SDCR) of the RBB optimization problem. The proposed suboptimal SDCR method solves a single convex optimization problem per frequency bin, resulting in a much lower computational complexity than the SCO method. Unlike the SCO method, the SDCR method does not guarantee user-controlled upper-bounded binaural-cue distortions. To tackle this problem, we also propose a suboptimal hybrid method that combines the SDCR and SCO methods. Instrumental measures combined with a listening test show that the SDCR and hybrid methods achieve significantly lower computational complexity than the SCO method, and in most cases better tradeoff between predicted intelligibility and binaural-cue preservation than the SCO method. Andreas I. Koutrouvelis, Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Robust Joint Estimation of Multimicrophone Signal Model ParametersabstractOne of the biggest challenges in multimicrophone applications is the estimation of the parameters of the signal model, such as the power spectral densities (PSDs) of the sources, the early (relative) acoustic transfer functions of the sources with respect to the microphones, the PSD of late reverberation, and the PSDs of microphone-self noise. Typically, existing methods estimate subsets of the aforementioned parameters and assume some of the other parameters to be known a priori. This may result in inconsistencies and inaccurately estimated parameters and potential performance degradation in the applications using these estimated parameters. So far, there is no method to jointly estimate all the aforementioned parameters. In this paper, we propose a robust method for jointly estimating all the aforementioned parameters using confirmatory factor analysis. The estimation accuracy of the signal-model parameters thus obtained outperforms existing methods in most cases. We experimentally show significant performance gains in several multimicrophone applications over state-of-the-art methods. Andreas I. Koutrouvelis, Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Relative Acoustic Transfer Function Estimation in Wireless Acoustic Sensor NetworksabstractIn this paper, we present an algorithm to estimate the relative acoustic transfer function (RTF) of a target source in wireless acoustic sensor networks (WASNs). Two well-known methods to estimate the RTF are the covariance subtraction (CS) method and the covariance whitening (CW) approach, the latter based on the generalized eigenvalue decomposition. Both methods depend on the use of the noisy correlation matrix, which, in practice, has to be estimated using limited and (in WASNs) quantized data. The bit rate and the fact that we use limited data records therefore directly affect the accuracy of the estimated RTFs. Therefore, we first theoretically analyze the estimation performance of the two approaches in terms of bit rate. Second, we propose a rate-distribution method by minimizing the power usage and constraining the expected estimation error for both RTF estimators. The optimal rate distributions are found by using convex optimization techniques. The model-based methods, however, are impractical due to the dependence on the true RTFs. We therefore further develop two greedy rate-distribution methods for both approaches. Finally, numerical simulations on synthetic data and real audio recordings show the superiority of the proposed approaches in power usage compared to uniform rate allocation. We find that in order to satisfy the same RTF estimation accuracy, the rate-distributed CW methods consume much less transmission energy than the CS-based methods. Jie Zhang 0042, Richard Heusdens, Richard C. Hendriks |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | An Instrumental Intelligibility Metric Based on Information TheoryabstractWe propose a monaural intrusive instrumental intelligibility metric called speech intelligibility in bits (SIIB). SIIB is an estimate of the amount of information shared between a talker and a listener in bits per second. Unlike existing information theoretic intelligibility metrics, SIIB accounts for talker variability and statistical dependencies between time-frequency units. Our evaluation shows that relative to state-of-the-art intelligibility metrics, SIIB is highly correlated with the intelligibility of speech that has been degraded by noise and processed by speech enhancement algorithms. Steven Van Kuyk, W. Bastiaan Kleijn, Richard C. Hendriks |
IEEE Signal Process. Lett. | 3 |
| 2018 | A Low-Cost Robust Distributed Linearly Constrained Beamformer for Wireless Acoustic Sensor Networks With Arbitrary TopologyabstractWe propose a new robust distributed linearly constrained beamformer that utilizes a set of linear equality constraints to reduce the cross power spectral density matrix to a block-diagonal form. The proposed beamformer has a convenient objective function for use in arbitrary distributed network topologies while having identical performance to a centralized implementation. Moreover, the new optimization problem is robust to relative acoustic transfer function (RATF) estimation errors and to target activity detection (TAD) errors. Two variants of the proposed beamformer are presented and evaluated in the context of multimicrophone speech enhancement in a wireless acoustic sensor network, and are compared with other state-of-the-art distributed beamformers in terms of communication costs and robustness to RATF estimation errors and TAD errors. Andreas I. Koutrouvelis, Thomas Sherson, Richard Heusdens, Richard C. Hendriks |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2018 | An Evaluation of Intrusive Instrumental Intelligibility MetricsabstractInstrumental intelligibility metrics are commonly used as an alternative to listening tests. This paper evaluates 12 monaural intrusive intelligibility metrics: SII, HEGP, CSII, HASPI, NCM, QSTI, STOI, ESTOI, MIKNN, SIMI, SIIB, and sEPSMcorr. In addition, this paper investigates the ability of intelligibility metrics to generalize to new types of distortions and analyzes why the top performing metrics have high performance. The intelligibility data were obtained from 11 listening tests described in the literature. The stimuli included Dutch, Danish, and English speech that was distorted by additive noise, reverberation, competing talkers, preprocessing enhancement, and postprocessing enhancement. SIIB and HASPI had the highest performance achieving a correlation with listening test scores on average of p = 0.92 and p = 0.89, respectively. The high performance of SIIB may, in part, be the result of SIIBs developers having access to all the intelligibility data considered in the evaluation. The results show that intelligibility metrics tend to perform poorly on datasets that were not used during their development. By modifying the original implementations of SIIB and STOI, the advantage of reducing statistical dependencies between input features is demonstrated. Additionally, this paper presents a new version of SIIB called SIIBGauss, which has similar performance to SIIB and HASPI, but takes less time to compute by two orders of magnitude. Steven Van Kuyk, W. Bastiaan Kleijn, Richard C. Hendriks |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Microphone Subset Selection for MVDR Beamformer Based Noise ReductionabstractIn large-scale wireless acoustic sensor networks (WASNs), many of the sensors will only have a marginal contribution to a certain estimation task. Involving all sensors increases the energy budget unnecessarily and decreases the lifetime of the WASN. Using microphone subset selection, also termed as sensor selection, the most informative sensors can be chosen from a set of candidate sensors to achieve a prescribed inference performance. In this paper, we consider microphone subset selection for minimum variance distortionless response (MVDR) beamformer based noise reduction. The best subset of sensors is determined by minimizing the transmission cost while constraining the output noise power (or signal-to-noise ratio). Assuming the statistical information on correlation matrices of the sensor measurements is available, the sensor selection problem for this model-driven scheme is first solved by utilizing convex optimization techniques. In addition, to avoid estimating the statistics related to all the candidate sensors beforehand, we also propose a data-driven approach to select the best subset using a greedy strategy. The performance of the greedy algorithm converges to that of the model-driven method, while it displays advantages in dynamic scenarios as well as on computational complexity. Compared to a sparse MVDR or radius-based beamformer, experiments show that the proposed methods can guarantee the desired performance with significantly less transmission costs. Jie Zhang 0042, Sundeep Prabhakar Chepuri, Richard C. Hendriks, Richard Heusdens |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Rate-Distributed Spatial Filtering Based Noise Reduction in Wireless Acoustic Sensor NetworksabstractIn wireless acoustic sensor networks (WASNs), sensors typically have a limited energy budget as they are often battery-driven. Energy efficiency is, therefore, essential for the design of algorithms in WASNs. One way to reduce energy costs is to select only the sensors that are most informative, a problem known as sensor selection . In this way, only sensors that significantly contribute to the task at hand will be involved. In this paper, we consider a more general approach, which is based on rate-distributed spatial filtering. Depending on the distance over which a transmission takes place, the bit rate directly influences the energy consumption. We try to minimize the battery usage due to transmission, while constraining the noise reduction performance. This results in an efficient rate allocation strategy, which depends on the underlying signal statistics, as well as the distance from sensors to a fusion center (FC). Through the utilization of a linearly constrained minimum variance beamformer, the problem is derived as a semidefinite program. Furthermore, we show that rate allocation is more general than sensor selection, and sensor selection can be seen as a special case of the presented rate-allocation solution, e.g., the best microphone subset can be determined by thresholding the rates. Finally, numerical simulations for estimating several target sources in a WASN demonstrate that the proposed method outperforms the sensor-selection-based approaches in terms of energy usage, and we find that the sensors close to the FC and point sources are allocated with higher rates. Jie Zhang 0042, Richard Heusdens, Richard C. Hendriks |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | On the information rate of speech communicationabstractThe key to the success of speech-based technology is an understanding of human speech communication. While significant advances have been made, a unified theory of speech communication that is both comprehensive and quantitative is yet to emerge. In this paper we approach speech communication from an information theoretical perspective. Without relying on prior knowledge of speech production, language, or auditory processing, we develop a new methodology for measuring the information rate of speech. Instead we rely on having recordings of multiple talkers saying the same utterance. In general, our results are consistent with a linguistic understanding of speech communication. Steven Van Kuyk, W. Bastiaan Kleijn, Richard C. Hendriks |
ICASSP | 3 |
| 2017 | Intelligibility Enhancement Based on Mutual InformationabstractSpeech intelligibility enhancement is considered for multiple-microphone acquisition and single loudspeaker rendering. This is based on the mutual information measured between the message spoken at far-end environment and the message perceived by a listener at near-end. We prove that the joint optimal processing can be decomposed into far-end and near-end processing. The former is a minimum variance distortionless response beamformer that reduces the noise in the talker environment and the latter is a post-filter that redistributes the power over the frequency bands. Disjoint processing is optimal provided that the post-filtering operation is aware of the residual noise from the beamforming operation. Our results show that both processing steps are necessary for the effective conveyance of a message and, importantly, that the second step must be aware of the remaining noise from the beamforming operation in the first step. In addition, we study the use of the mutual information applied on the perceptually more relevant powers per critical band. Seyran Khademi, Richard C. Hendriks, W. Bastiaan Kleijn |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Relaxed Binaural LCMV BeamformingabstractIn this paper, we propose a new binaural beamforming technique, which can be seen as a relaxation of the linearly constrained minimum variance (LCMV) framework. The proposed method can achieve simultaneous noise reduction and exact binaural cue preservation of the target source, similar to the binaural minimum variance distortionless response (BMVDR) method. However, unlike BMVDR, the proposed method is also able to preserve the binaural cues of multiple interferers to a certain predefined accuracy. Specifically, it is able to control the trade-off between noise reduction and binaural cue preservation of the interferers by using a separate trade-off parameter per-interferer. Moreover, we provide a robust way of selecting these trade-off parameters in such a way that the preservation accuracy for the binaural cues of the interferers is always better than the corresponding ones of the BMVDR. The relaxation of the constraints in the proposed method achieves approximate binaural cue preservation of more interferers than other previously presented LCMV-based binaural beamforming methods that use strict equality constraints. Andreas I. Koutrouvelis, Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Jointly optimal near-end and far-end multi-microphone speech intelligibility enhancement based on mutual informationabstractThe processing required for the global maximization of the intelligibility of speech acquired by multiple microphones and rendered by a single loudspeaker, is considered in this paper. The intelligibility is quantized, based on the mutual information rate between the message spoken by the talker and the message as interpreted by the listener. We prove that then, in each of a set of narrow-band channels, the processing can be decomposed into a minimum variance distortionless response (MVDR) beamforming operation that reduces the noise in the talker environment, followed by a gain operation that, given the far-end noise and beamforming operation, accounts for the noise at the listener end. Our experiments confirm that both processing steps are necessary for the effective conveyance ofa message and, importantly, that the second step must be aware of the first step. Seyran Khademi, Richard C. Hendriks, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2016 | Improved multi-microphone noise reduction preserving binaural cuesabstractWe propose a new multi-microphone noise reduction technique for binaural cue preservation of the desired source and the interferers. This method is based on the linearly constrained minimum variance (LCMV) framework, where the constraints are used for the binaural cue preservation of the desired source and of multiple interferers. In this framework there is a trade-off between noise reduction and binaural cue preservation. The more constraints the LCMV uses for preserving binaural cues, the less degrees of freedom can be used for noise suppression. The recently presented binaural LCMV (BLCMV) method and the optimal BLCMV (OBLCMV) method require two constraints per interferer and introduce an additional interference rejection parameter. This unnecessarily reduces the degrees of freedom, available for noise reduction, and negatively influences the trade-off between noise reduction and binaural cue preservation. With the proposed method, binaural cue preservation is obtained using just a single constraint per interferer without the need of an interference rejection parameter. The proposed method can simultaneously achieve noise reduction and perfect binaural cue preservation of more than twice as many interferers as the BLCMV, while the OBLCMV can preserve the binaural cues of only one interferer. Andreas I. Koutrouvelis, Richard C. Hendriks, Jesper Jensen 0001, Richard Heusdens |
ICASSP | 2 |
| 2015 | Speech reinforcement in noisy reverberant conditions under an approximation of the short-time SIIabstractWhile most contributions on speech reinforcement only consider the presence of environmental noise, late reverberation can also severely degrade the intelligibility of speech. In this paper we address the problem of speech reinforcement in noisy and reverberant environments. We use a short-time version of a recently presented approximation of the speech intelligibility index, which we optimize locally. The resulting time-frequency dependent amplification depends on both the noise and late reverberation power spectral density. The latter is estimated using the Polack model and assumes that prior knowledge of the room geometry is available. Speech intelligibility improvements of around 20% are observed. Richard C. Hendriks, Joao B. Crespo, Jesper Jensen 0001, Cees H. Taal |
ICASSP | 1 |
| 2015 | On clock synchronization for multi-microphone speech processing in wireless acoustic sensor networksabstractIn wireless acoustic sensor networks (WASNs), clock synchronization is crucial for multi-microphone signal processing, since clock differences between capturing devices will cause signal drift. This in turn severely degrades the performance of multi-microphone signal processing. After a theoretical analysis of the effect of clock synchronization, we evaluate the use of three different clock synchronization algorithms in the context of multi-microphone noise reduction. Our experimental study shows that the achieved precision of clock synchronization enables sufficient accuracy of clock synchronization for the MVDR beamformer in ideal scenarios. However, in practical scenarios with measurement noise on the parameters of interest, time-stamp based clock synchronization algorithms get degraded, while signal based algorithms are still accurate enough for the MVDR beamformer, albeit at a much higher transmission cost. Yuan Zeng 0001, Richard C. Hendriks, Nikolay D. Gaubitch |
ICASSP | 2 |
| 2015 | Distributed estimation of the inverse of the correlation matrix for privacy preserving beamforming
Yuan Zeng 0001, Richard C. Hendriks |
Signal Process. | 2 |
| 2015 | A Simple Model of Speech Communication and its Application to Intelligibility EnhancementabstractWe introduce a model of communication that includes noise inherent in the message production process as well as noise inherent in the message interpretation process. The production and interpretation noise processes have a fixed signal-to-noise ratio. The resulting system is a simple but effective model of human communication. The model naturally leads to a method to enhance the intelligibility of speech rendered in a noisy environment. State-of-the-art experimental results confirm the practical value of the model. W. Bastiaan Kleijn, Richard C. Hendriks |
IEEE Signal Process. Lett. | 2 |
| 2015 | Optimal Near-End Speech Intelligibility Improvement Incorporating Additive Noise and Late Reverberation Under an Approximation of the Short-Time SIIabstractThe presence of environmental additive noise in the vicinity of the user typically degrades the speech intelligibility of speech processing applications. This intelligibility loss can be compensated by properly preprocessing the speech signal prior to play-out, often referred to as near-end speech enhancement. Although the majority of such algorithms focus primarily on the presence of additive noise, reverberation can also severely degrade intelligibility. In this paper we investigate how late reverberation and additive noise can be jointly taken into account in the near-end speech enhancement process. For this effort we use a recently presented approximation of the speech intelligibility index under a power constraint, which we optimize for speech degraded by both additive noise and late reverberation. The algorithm results in time-frequency dependent amplification factors that depend on both the additive noise power spectral density as well as the late reverberation energy. These amplification factors redistribute speech energy across frequency and perform a dynamic range compression. Experimental results using both instrumental intelligibility measures as well as intelligibility listening tests show that the proposed approach improves speech intelligibility over state-of-the-art reference methods when speech signals are degraded simultaneously by additive noise and reverberation. Speech intelligibility improvements in the order of 20% are observed. Richard C. Hendriks, Joao B. Crespo, Jesper Jensen 0001, Cees H. Taal |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2014 | Speech reinforcement in noisy reverberant environments using a perceptual distortion measureabstractIn this paper, a time-frequency weighting is proposed for speech reinforcement (near-end listening enhancement) in a noisy and reverberant environment, which optimizes a perceptual distortion measure locally for each time-frequency bin. The algorithm acts as a dynamic range compressor, smearing out the energy of the clean speech along time. Simulations predict an intelligibility increase with respect to the unprocessed condition and two reference methods, for moderate smoothing windows, as measured by the optimized distortion measure and two objective intelligibility measures. Joao B. Crespo, Richard C. Hendriks |
ICASSP | 2 |
| 2014 | Speech energy redistribution for intelligibility improvement in noise based on a perceptual distortion measure
Cees H. Taal, Richard C. Hendriks, Richard Heusdens |
Comput. Speech Lang. | 2 |
| 2014 | Multizone Speech ReinforcementabstractIn this article, we address speech reinforcement (near-end listening enhancement) for a scenario where there are several playback zones. In such a framework, signals from one zone can leak into other zones (crosstalk), causing intelligibility and/or quality degradation. An optimization framework is built by exploring a signal model where effects of noise, reverberation and zone crosstalk are taken into account simultaneously. Through the symbolic usage of a general smooth distortion measure, necessary optimality conditions are derived in terms of distortion measure gradients and the signal model. Subsequently, as an illustrative example of the framework, the conditions are applied for the mean-square error (MSE) expected distortion under a hybrid stochastic-deterministic model for the corruptions. A crosstalk cancellation algorithm follows, which depends on diffuse reverberation and across zone direct path components. Simulations validate the optimality of the algorithm and show a clear benefit in multizone processing, as opposed to the iterated application of a single-zone algorithm. Also, comparisons with least-squares crosstalk cancellers in literature show the profit of using a hybrid model. Joao B. Crespo, Richard C. Hendriks |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Distributed Delay and Sum Beamformer for Speech Enhancement via Randomized GossipabstractIn this paper, we investigate the use of randomized gossip for distributed speech enhancement and present a distributed delay and sum beamformer (DDSB). In a randomly connected wireless acoustic sensor network, the DDSB estimates the desired signal at each node by communicating only with its neighbors. We first provide the asynchronous DDSB (ADDSB) where each pair of neighboring nodes updates its data asynchronously. Then, we introduce an improved general distributed synchronous averaging (IGDSA) algorithm, which can be used in any connected network, and combine that with the DDSB algorithm where multiple node pairs can update their estimates simultaneously. For convergence analysis, we first provide bounds for the worst case averaging time of the ADDSB for the best and worst connected networks, and then we compare the convergence rate of the ADDSB with the original synchronous DDSB (OSDDSB) and the improved synchronous DDSB (ISDDSB) in regular networks. This convergence rate comparison is extended to randomly connected non-regular networks using simulations. The simulation results show that the DDSB using the different updating schemes converges to the optimal estimates of the centralized beamformer and that the proposed IGDSA algorithm converges much faster than the original synchronous communication scheme, in particular for non-regular networks. Moreover, comparisons are performed with several existing distributed speech enhancement methods from literature, assuming that the steering vector is given. In the simulated scenario, the proposed method leads to a slight performance improvement at the expense of a higher communication cost. The presented method is not constrained to a certain network topology (e.g., tree connected or fully connected), while this is the case for many of the reference methods. Yuan Zeng 0001, Richard C. Hendriks |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2013 | Privacy-preserving distributed speech enhancement forwireless sensor networks by processing in the encrypted domainabstractTo improve speech communication in noisy and reverberant environments, an increased interest is shown to develop algorithms that make efficiently use of acoustic wireless sensor networks (WSNs). The processors and sensors forming these WSNs can be owned by multiple users. Sending private data across such a WSN can lead to severe privacy and security issues and may limit its acceptance. Using the advantages of WSNs, while guaranteeing people's privacy, requires therefore to share processors and data in a privacy preserving manner. In this paper we raise attention to the problem of privacy and security for distributed speech enhancement and propose the new paradigm of privacy preserving distributed beamforming. Using cryptographic techniques, particularly homomorphic encryption, we demonstrate how distributed beamforming techniques can be computed in a privacy preserving manner in the encrypted domain. Richard C. Hendriks, Zekeriya Erkin, Timo Gerkmann |
ICASSP | 1 |
| 2013 | A generalized Fourier domain: Signal processing framework and applications
Jorge Martínez 0002, Richard Heusdens, Richard C. Hendriks |
Signal Process. | 3 |
| 2012 | Improved mmse-based noise PSD tracking using temporal cepstrum smoothingabstractRecently, it has been shown that MMSE-based noise power estimation [1] results in an improved noise tracking performance with respect to minimum statistics-based approaches. The MMSE-based approach employs two estimates of the speech power to estimate the unbiased noise power. In this work, we improve the MMSE-based noise power estimator by employing a more advanced estimator of the speech power based on temporal cepstrum smoothing (TCS). TCS can exploit knowledge about the speech spectral structure. As a result, only one speech power estimate is needed for MMSE-based noise power estimation. Moreover, the presented estimator results in an improved noise tracking performance, especially in babble noise, where SNR improvements of 1dB over the original MMSE-based approach can be observed. Timo Gerkmann, Richard C. Hendriks |
ICASSP | 2 |
| 2012 | A spatio-temporal generalized fourier domain framework to acoustic modeling in enclosed spacesabstractIn this paper, we present a spatio-temporal framework for multichannel acoustic modeling in enclosed spaces. Reverberation occurs when the sound field is enclosed between reflective boundaries (e.g. walls). We model the reverberated sound field by proper sampling of the (generalized) Fourier representation of the free-field sound field. We show that the spatial aliasing introduced by spectral sampling represents all the (damped) reflections. From the samples of the generalized spectrum, we compute the spatio-temporal sound field in the enclosed space with very low-complexity, of O(N log N) per measuring position, with N proportional to the reverberation time. Jorge Martínez 0002, Richard Heusdens, Richard C. Hendriks |
ICASSP | 3 |
| 2012 | A speech preprocessing strategy for intelligibility improvement in noise based on a perceptual distortion measureabstractA speech pre-processing algorithm is presented to improve the speech intelligibility in noise for the near-end listener. The algorithm improves the intelligibility by optimally redistributing the speech energy over time and frequency for a perceptual distortion measure, which is based on a spectro-temporal auditory model. In contrast to spectral-only models, short-time information is taken into account. As a consequence, the algorithm is more sensitive to transient regions, which will therefore receive more amplification compared to stationary vowels. It is known from literature that changing the vowel-transient energy ratio is beneficial for improving speech-intelligibility in noise. Objective intelligibility prediction results show that the proposed method has higher speech intelligibility in noise compared to two other reference methods, without modifying the global speech energy. Cees H. Taal, Richard C. Hendriks, Richard Heusdens |
ICASSP | 2 |
| 2012 | On mutual information as a measure of speech intelligibilityabstractSpeech intelligibility prediction of noisy and processed noisy speech is important in a number of application domains such as hearing instruments and forensics. Most available objective intelligibility measures employ either a signal-to-noise ratio (SNR)-based or correlation-based comparison between frequency bands of the clean and the processed speech. In this paper, we approach the speech intelligibility prediction from the angle of information theory and show that an information theoretic concept provides a unified viewpoint on both the SNR and the correlation based approaches. Two objective intelligibility measures are introduced based on estimated mutual information between the clean speech and the processed speech in the time and the frequency subband domain. Our proposed measures show high correlation with subjective intelligibility measure (i.e. word correct scores) and comparative results with the short-term objective intelligibility measure (STOI). Jalal Taghia, Rainer Martin 0001, Richard C. Hendriks |
ICASSP | 3 |
| 2012 | Distributed delay and sum beamformer for speech enhancement in wireless sensor networks via randomized gossipabstractIn this paper, we describe a distributed delay and sum beamformer (DDSB) for speech enhancement based on a randomized gossip algorithm. The proposed algorithm operates in a randomly connected wireless sensor network. Without any network topology constraint, the DDSB estimates the desired signal at each node by only exchanging information with its neighbors. Since the DDSB performs only local signal processing, it is robust and scalable for large sensor networks and dynamic environments. We show that the DDSB converges to the optimal estimate of the centralized beamformer. Furthermore, we provide a bound for the worst-case averaging time of the DDSB for the worst connected network. The simulation results validate the theoretical results of the algorithm. Yuan Zeng 0001, Richard C. Hendriks |
ICASSP | 2 |
| 2012 | Unbiased MMSE-Based Noise Power Estimation With Low Complexity and Low Tracking DelayabstractRecently, it has been proposed to estimate the noise power spectral density by means of minimum mean-square error (MMSE) optimal estimation. We show that the resulting estimator can be interpreted as a voice activity detector (VAD)-based noise power estimator, where the noise power is updated only when speech absence is signaled, compensated with a required bias compensation. We show that the bias compensation is unnecessary when we replace the VAD by a soft speech presence probability (SPP) with fixed priors. Choosing fixed priors also has the benefit of decoupling the noise power estimator from subsequent steps in a speech enhancement framework, such as the estimation of the speech power and the estimation of the clean speech. We show that the proposed speech presence probability (SPP) approach maintains the quick noise tracking performance of the bias compensated minimum mean-square error (MMSE)-based approach while exhibiting less overestimation of the spectral noise power and an even lower computational complexity. Timo Gerkmann, Richard C. Hendriks |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Noise Correlation Matrix Estimation for Multi-Microphone Speech EnhancementabstractFor multi-channel noise reduction algorithms like the minimum variance distortionless response (MVDR) beamformer, or the multi-channel Wiener filter, an estimate of the noise correlation matrix is needed. For its estimation, it is often proposed in the literature to use a voice activity detector (VAD). However, using a VAD the estimated matrix can only be updated in speech absence. As a result, during speech presence the noise correlation matrix estimate does not follow changing noise fields with an appropriate accuracy. This effect is further increased, as in nonstationary noise voice activity detection is a rather difficult task, and false-alarms are likely to occur. In this paper, we present and analyze an algorithm that estimates the noise correlation matrix without using a VAD. This algorithm is based on measuring the correlation of the noisy input and a noise reference which can be obtained, e.g., by steering a null towards the target source. When applied in combination with an MVDR beamformer, it is shown that the proposed noise correlation matrix estimate results in a more accurate beamformer response, a larger signal-to-noise ratio improvement and a larger instrumentally predicted speech intelligibility when compared to competing algorithms such as the generalized sidelobe canceler, a VAD-based MVDR beamformer, and an MVDR based on the noisy correlation matrix. Richard C. Hendriks, Timo Gerkmann |
IEEE Trans. Speech Audio Process. | 1 |
| 2012 | Spectral Magnitude Minimum Mean-Square Error Estimation Using Binary and Continuous Gain FunctionsabstractRecently, binary mask techniques have been proposed as a tool for retrieving a target speech signal from a noisy observation. A binary gain function is applied to time-frequency tiles of the noisy observation in order to suppress noise dominated and retain target dominated time-frequency regions. When implemented using discrete Fourier transform (DFT) techniques, the binary mask techniques can be seen as a special case of the broader class of DFT-based speech enhancement algorithms, for which the applied gain function is not constrained to be binary. In this context, we develop and compare binary mask techniques to state-of-the-art continuous gain techniques. We derive spectral magnitude minimum mean-square error binary gain estimators; the binary gain estimators turn out to be simple functions of the continuous gain estimators. We show that the optimal binary estimators are closely related to a range of existing, heuristically developed, binary gain estimators. The derived binary gain estimators perform better than existing binary gain estimators in simulation experiments with speech signals contaminated by several different noise sources as measured by speech quality and intelligibility measures. However, even the best binary mask method is significantly outperformed by state-of-the-art continuous gain estimators. The instrumental intelligibility results are confirmed in an intelligibility listening test. Jesper Jensen 0001, Richard C. Hendriks |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | A Low-Complexity Spectro-Temporal Distortion Measure for Audio Processing ApplicationsabstractPerceptual models exploiting auditory masking are frequently used in audio and speech processing applications like coding and watermarking. In most cases, these models only take into account spectral masking in short-time frames. As a consequence, undesired audible artifacts in the temporal domain may be introduced (e.g., pre-echoes). In this article we present a new low-complexity spectro-temporal distortion measure. The model facilitates the computation of analytic expressions for masking thresholds, while advanced spectro-temporal models typically need computationally demanding adaptive procedures to find an estimate of these masking thresholds. We show that the proposed method gives similar masking predictions as an advanced spectro-temporal model with only a fraction of its computational power. The proposed method is also compared with a spectral-only model by means of a listening test. From this test it can be concluded that for non-stationary frames the spectral model underestimates the audibility of introduced errors and therefore overestimates the masking curve. As a consequence, the system of interest incorrectly assumes that errors are masked in a particular frame, which leads to audible artifacts. This is not the case with the proposed method which correctly detects the errors made in the temporal structure of the signal. Cees H. Taal, Richard C. Hendriks, Richard Heusdens |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Estimation of the noise correlation matrixabstractTo harvest the potential of multi-channel noise reduction methods, it is crucial to have an accurate estimate of the noise correlation matrix. Existing algorithms either assume speech absence and exploit a voice activity detector (VAD), or make use of additional assumptions like a diffuse noise field. Therefore, these algorithms are limited with respect to their tracking speed and the type of noise fields for which they can estimate the correlation matrix. In this paper we present a new method for noise correlation matrix estimation that makes no assumptions about the type of noise field, nor uses a VAD. The presented method exploits the existence of accurate single-channel noise PSD estimators, as well as the avail ability of one noise reference per microphone pair. For spatially and temporally non-stationary noise fields, the proposed method leads to improved performance compared to widely used state-of-the-art reference methods in terms of both segmental SNR and beamformer response error. Richard C. Hendriks, Timo Gerkmann |
ICASSP | 1 |
| 2011 | Spectral magnitude minimum mean-square error binary masks for DFT based speech enhancementabstractOriginally, ideal binary mask (idbm) techniques have been used as a tool for studying aspects of the auditory system. More recently, idbm techniques have been adapted to the practical problem of retrieving a target speech signal from a noisy observation. In this practical setting, the biliary mask techniques show similarities with existing DFT based speech enhancement techniques. In this context, we derive single-channel, binary mask estimators which minimize the spectral magnitude mean-square error. We show in simulation experiments with natural speech and noise signals that the proposed estimators perform significantly better than existing binary mask estimators. However, even the best of the proposed estimators is clearly out performed by non-binary estimators, both in terms of speech quality and intelligibility. Jesper Jensen 0001, Richard C. Hendriks |
ICASSP | 2 |
| 2011 | A Generalized Poisson Summation Formula and its Application to Fast Linear ConvolutionabstractIn this letter, a generalized Fourier transform is introduced and its corresponding generalized Poisson summation formula is derived. For discrete, Fourier based, signal processing, this formula shows that a special form of control on the periodic repetitions that occur due to sampling in the reciprocal domain is possible. The present paper is focused on the derivation and analysis of a weighted circular convolution theorem. We use this specific result to compute linear convolutions in the generalized Fourier domain, without the need of zero-padding. This results in faster, more resource-efficient computations. Other techniques that achieve this have been introduced in the past using different approaches. The newly proposed theory however, constitutes a unifying framework to the methods previously published. Jorge Martínez 0002, Richard Heusdens, Richard C. Hendriks |
IEEE Signal Process. Lett. | 3 |
| 2011 | An Algorithm for Intelligibility Prediction of Time-Frequency Weighted Noisy SpeechabstractIn the development process of noise-reduction algorithms, an objective machine-driven intelligibility measure which shows high correlation with speech intelligibility is of great interest. Besides reducing time and costs compared to real listening experiments, an objective intelligibility measure could also help provide answers on how to improve the intelligibility of noisy unprocessed speech. In this paper, a short-time objective intelligibility measure (STOI) is presented, which shows high correlation with the intelligibility of noisy and time-frequency weighted noisy speech (e.g., resulting from noise reduction) of three different listening experiments. In general, STOI showed better correlation with speech intelligibility compared to five other reference objective intelligibility models. In contrast to other conventional intelligibility models which tend to rely on global statistics across entire sentences, STOI is based on shorter time segments (386 ms). Experiments indeed show that it is beneficial to take segment lengths of this order into account. In addition, a free Matlab implementation is provided. Cees H. Taal, Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | On linear versus non-linear magnitude-DFT estimators and the influence of super-Gaussian speech priorsabstractAlthough the linear mean-squared error (MSE) complex-DFT estimator, i.e., the Wiener filter, is well-known, its magnitude-DFT (MDFT) counterpart has never been considered in the context of speech enhancement. Therefore, certain theoretical questions regarding MDFT estimators remained unanswered. For example, it is unknown to which extend the performance of existing MSE MDFT estimators depends on the chosen speech prior, or on the non-linearity of the estimators. In this paper we present linear MSE MDFT estimators for speech enhancement. In contrast to the linear complex-DFT estimator, the presented linear MSE MDFT estimators do depend on the assumed distribution of the speech DFT coefficients. Based on objective and subjective experiments, it can be concluded that the chosen speech prior, i.e., Gaussian versus super-Gaussian has a significant effect on the performance of MDFT estimators, while the linearity as compared to non-linearity has only a minor influence. Richard C. Hendriks, Richard Heusdens |
ICASSP | 1 |
| 2010 | MMSE based noise PSD tracking with low complexityabstractMost speech enhancement algorithms heavily depend on the noise power spectral density (PSD). Because this quantity is unknown in practice, estimation from the noisy data is necessary. We present a low complexity method for noise PSD estimation. The algorithm is based on a minimum mean-squared error estimator of the noise magnitude-squared DFT coefficients. Compared to minimum statistics based noise tracking, segmental SNR and PESQ are improved for non-stationary noise sources with 1 dB and 0.25 MOS points, respectively. Compared to recently published algorithms, similar good noise tracking performance is obtained, but at a computational complexity that is in the order of a factor 40 lower. Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
ICASSP | 1 |
| 2010 | A short-time objective intelligibility measure for time-frequency weighted noisy speechabstractExisting objective speech-intelligibility measures are suitable for several types of degradation, however, it turns out that they are less appropriate for methods where noisy speech is processed by a time-frequency (TF) weighting, e.g., noise reduction and speech separation. In this paper, we present an objective intelligibility measure, which shows high correlation (rho=0.95) with the intelligibility of both noisy, and TF-weighted noisy speech. The proposed method shows significantly better performance than three other, more sophisticated, objective measures. Furthermore, it is based on an intermediate intelligibility measure for short-time (approximately 400 ms) TF-regions, and uses a simple DFT-based TF-decomposition. In addition, a free Matlab implementation is provided. Cees H. Taal, Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
ICASSP | 2 |
| 2009 | Fast noise PSD estimation with low complexityabstractAlthough noise PSD estimation is a crucial part of noise reduction algorithms, most noise PSD estimators have problems in tracking non-stationary noise sources. Recently, a noise PSD estimator based on DFT-subspace decompositions was proposed, which improves estimation of the PSD of such noise sources. However, as this approach is based on eigenvalue decompositions per DFT bin, it might be too computationally demanding for low-complexity applications like hearing aids. In this paper we present a method with similar noise tracking performance as the DFT-subspace approach, but with low computational costs. This method is based on computation of high resolution perodiograms, and can estimate the noise PSD when both speech and noise are present in a frequency bin. When combined with a complete noise reduction system, the proposed method can lead to an improvement for non-stationary noise sources of more than 1 dB segmental SNR and 0.3 on a PESQ scale, compared to standard noise tracking methods such as minimum statistics and the quantile based approach, while computational complexity is in the same order of magnitude. Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001, Ulrik Kjems |
ICASSP | 1 |
| 2009 | Log-spectral magnitude MMSE estimators under super-Gaussian densitiesabstractDespite the fact that histograms of speech DFT coefficients are super-Gaussian, not much attention has been paid to develop estimators under these super-Gaussian distributions in combi-nation with perceptual meaningful distortion measures. In this paper we present log-spectral magnitude MMSE estimators un-der super-Gaussian densities, resulting in an estimator that is perceptually more meaningful and in line with measured his-tograms of speech DFT coefficients. Compared to state-of-the-art reference methods, the presented estimator leads to an im-provement of the segmental SNR in the order of 0.5 dB up to 1 dB. Moreover, listening tests show that the proposed estima-tor leads to significant improvement for the presented estimator over state-of-the-art methods. Index Terms: speech enhancement, log-spectral magnitude MMSE, super-Gaussian Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
INTERSPEECH | 1 |
| 2009 | An evaluation of objective quality measures for speech intelligibility predictionabstractIn this research various objective quality measures are evalu-ated in order to predict the intelligibility for a wide range of non-linearly processed speech signals and speech degraded by additive noise. The obtained results are compared with the pre-diction results of a more advanced perceptual-based model pro-posed by Dau et al. and an objective intelligibility measure, namely the coherence speech intelligibility index (cSII). These tests are performed in order to gain more knowledge between the link of speech-quality and speech-intelligibility and may help us to exploit the extensive research done into the field of speech-quality for speech-intelligibility. It is shown that cSII does not necessarily show better performance compared to con-ventional objective (speech)-quality measures. In general, the DAU-model is the only method with reasonable results for all processing conditions. Index Terms: Speech intelligibility prediction, speech quality, objective Measure. Cees H. Taal, Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001, Ulrik Kjems |
INTERSPEECH | 2 |
| 2009 | On Optimal Multichannel Mean-Squared Error Estimators for Speech EnhancementabstractIn this letter we present discrete Fourier transform (DFT) domain minimum mean-squared error (MMSE) estimators for multichannel noise reduction. The estimators are derived assuming that the clean speech magnitude DFT coefficients are generalized-Gamma distributed. We show that for Gaussian distributed noise DFT coefficients, the optimal filtering approach consists of a concatenation of a minimum variance distortionless response (MVDR) beamformer followed by well-known single-channel MMSE estimators. The multichannel Wiener filter follows as a special case of the presented MSE estimators and is in general suboptimal. For non-Gaussian distributed noise DFT coefficients the resulting spatial filter is in general nonlinear with respect to the noisy microphone signals and cannot be decomposed into an MVDR beamformer and a post-filter. Richard C. Hendriks, Richard Heusdens, Ulrik Kjems, Jesper Jensen 0001 |
IEEE Signal Process. Lett. | 1 |
| 2008 | Comparison of complex-DFT estimators with and without the independence assumption of real and imaginary partsabstractMMSE estimators for DFT-domain based single-microphone speech enhancement can broadly be classified in those that estimate the complex-DFT coefficients and those that estimate the DFT magnitudes. Existing complex-DFT MMSE estimators have generally been derived under assumptions that are in conflict with measured histograms and that are inconsistent with the assumptions made to derive DFT magnitude estimators. Recently it has been shown that these inconsistencies can be eliminated, i.e., no independency has to be assumed between real and imaginary parts of DFT coefficients if the phase of DFT coefficients is assumed uniformly distributed. In this paper we discuss the assumptions that underlie the different complex-DFT estimators and show that the uniform phase assumption matches actual speech data. Furthermore, we show experimentally that the estimators without the independence assumption lead to a lower mean-square error. Richard C. Hendriks, Jan S. Erkelens, Richard Heusdens |
ICASSP | 1 |
| 2008 | On the Estimation of Complex Speech DFT Coefficients Without Assuming Independent Real and Imaginary PartsabstractThis letter considers the estimation of speech signals contaminated by additive noise in the discrete Fourier transform (DFT) domain. Existing complex-DFT estimators assume independency of the real and imaginary parts of the speech DFT coefficients, although this is not in line with measurements. In this letter, we derive some general results on these estimators, under more realistic assumptions. Assuming that speech and noise are independent, speech DFT coefficients have uniform phase, and that noise DFT coefficients have a Gaussian density, we show theoretically that the spectral gain function for speech DFT estimation is real and upper-bounded by the corresponding gain function for spectral magnitude estimation. We also show that the minimum mean-square error (MMSE) estimator of the speech phase equals the noisy phase. No assumptions are made about the distribution of the speech spectral magnitudes. Recently, speech spectral amplitude estimators have been derived under a generalized-Gamma amplitude distribution. As an example, we will derive the corresponding complex-DFT estimators, without making the independence assumption. Jan S. Erkelens, Richard C. Hendriks, Richard Heusdens |
IEEE Signal Process. Lett. | 2 |
| 2008 | Noise Tracking Using DFT Domain Subspace DecompositionsabstractAll discrete Fourier transform (DFT) domain-based speech enhancement gain functions rely on knowledge of the noise power spectral density (PSD). Since the noise PSD is unknown in advance, estimation from the noisy speech signal is necessary. An overestimation of the noise PSD will lead to a loss in speech quality, while an underestimation will lead to an unnecessary high level of residual noise. We present a novel approach for noise tracking, which updates the noise PSD for each DFT coefficient in the presence of both speech and noise. This method is based on the eigenvalue decomposition of correlation matrices that are constructed from time series of noisy DFT coefficients. The presented method is very well capable of tracking gradually changing noise types. In comparison to state-of-the-art noise tracking algorithms the proposed method reduces the estimation error between the estimated and the true noise PSD. In combination with an enhancement system the proposed method improves the segmental SNR with several decibels for gradually changing noise types. Listening experiments show that the proposed system is preferred over the state-of-the-art noise tracking algorithm. Richard C. Hendriks, Jesper Jensen 0001, Richard Heusdens |
IEEE Trans. Speech Audio Process. | 1 |
| 2007 | DFT domain subspace based noise tracking for speech enhancementabstractMost DFT domain based speech enhancement methods are de-pendent on an estimate of the noise power spectral density (PSD). For non-stationary noise sources it is desirable to es-timate the noise PSD also in spectral regions where speech is present. In this paper a new method for noise tracking is pre-sented, based on eigenvalue decompositions of correlation ma-trices that are constructed from time series of noisy DFT coef-ficients. The presented method can estimate the noise PSD at time-frequency points where both speech and noise are present. In comparison to state-of-the-art noise tracking algorithms the proposed algorithm reduces the estimation error between the estimated and the true noise PSD and improves segmental SNR when combined with an enhancement system with several dB. Index Terms: Speech enhancement, noise tracking, DFT do-main subspace decompositions. Richard C. Hendriks, Jesper Jensen 0001, Richard Heusdens |
INTERSPEECH | 1 |
| 2007 | Minimum Mean-Square Error Estimation of Discrete Fourier Coefficients With Generalized Gamma PriorsabstractThis paper considers techniques for single-channel speech enhancement based on the discrete Fourier transform (DFT). Specifically, we derive minimum mean-square error (MMSE) estimators of speech DFT coefficient magnitudes as well as of complex-valued DFT coefficients based on two classes of generalized gamma distributions, under an additive Gaussian noise assumption. The resulting generalized DFT magnitude estimator has as a special case the existing scheme based on a Rayleigh speech prior, while the complex DFT estimators generalize existing schemes based on Gaussian, Laplacian, and Gamma speech priors. Extensive simulation experiments with speech signals degraded by various additive noise sources verify that significant improvements are possible with the more recent estimators based on super-Gaussian priors. The increase in perceptual evaluation of speech quality (PESQ) over the noisy signals is about 0.5 points for street noise and about 1 point for white noise, nearly independent of input signal-to-noise ratio (SNR). The assumptions made for deriving the complex DFT estimators are less accurate than those for the magnitude estimators, leading to a higher maximum achievable speech quality with the magnitude estimators. Jan S. Erkelens, Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | An MMSE Estimator for Speech Enhancement Under a Combined Stochastic-Deterministic Speech ModelabstractAlthough many discrete Fourier transform (DFT) domain-based speech enhancement methods rely on stochastic models to derive clean speech estimators, like the Gaussian and Laplace distribution, certain speech sounds clearly show a more deterministic character. In this paper, we study the use of a deterministic model in combination with the well-known stochastic models for speech enhancement. We derive a minimum mean-square error (MMSE) estimator under a combined stochastic-deterministic speech model with speech presence uncertainty and show that for different distributions of the DFT coefficients the combined stochastic-deterministic speech model leads to improved performance of approximately 0.8 dB segmental signal-to-noise ratio (SNR) over the use of a stochastic model alone. Evaluation with perceptual evaluation of speech quality (PESQ) shows performance improvements of approximately 0.15 on an MOS scale Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
IEEE Trans. Speech Audio Process. | 1 |
| 2007 | MAP Estimators for Speech Enhancement Under Normal and Rayleigh Inverse Gaussian DistributionsabstractThis paper presents a new class of estimators for speech enhancement in the discrete Fourier transform (DFT) domain, where we consider a multidimensional normal inverse Gaussian (MNIG) distribution for the speech DFT coefficients. The MNIG distribution can model a wide range of processes, from heavy-tailed to less heavy-tailed processes. Under the MNIG distribution complex DFT and amplitude estimators are derived. In contrast to other estimators, the suppression characteristics of the MNIG-based estimators can be adapted online to the underlying distribution of the speech DFT coefficients. Compared to noise suppression algorithms based on preselected super-Gaussian distributions, the MNIG-based complex DFT and amplitude estimators lead to a performance improvement in terms of segmental signal-to-noise ratio (SNR) in the order of 0.3 to 0.6 dB and 0.2 to 0.6 dB, respectively Richard C. Hendriks, Rainer Martin 0001 |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | Speech Enhancement Under a Combined Stochastic-Deterministic ModelabstractMost DFT domain based enhancement methods rely on stochastic models to derive clean speech estimators. In this paper we investigate the use of a deterministic speech model and present an MMSE estimator under a combined stochastic-deterministic speech model. Experimental results show an increase in segmental SNR of 1.18 dB, compared to the use of a stochastic model alone. Furthermore, PESQ evaluations lead to an increase of 0.3 on the MOS scale. Listening tests show a preference for the proposed MMSE estimator under combined stochastic-deterministic speech model Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
ICASSP (1) | 1 |
| 2006 | MMSE estimation of complex-valued discrete Fourier coefficients with generalized gamma priorsabstractWe consider DFT based techniques for single-channel speech enhancement. Specifically, we derive minimum mean-square error estimators of clean speech DFT coefficients based on generalized gamma prior probability density functions. Our estimators contain as special cases the well-known Wiener estimator and the more recently derived estimators based on Laplacian and twosided gamma priors. Simulation experiments with speech signals degraded by various additive noise sources verifythat theestimator based on the two-sided gamma prior is close to optimal amongst all the estimators considered in this paper. Jesper Jensen 0001, Richard C. Hendriks, Jan S. Erkelens, Richard Heusdens |
INTERSPEECH | 2 |
| 2006 | Adaptive Time Segmentation for Improved Speech EnhancementabstractSingle-channel enhancement algorithms are widely used to overcome the degradation of noisy speech signals. Speech enhancement gain functions are typically computed from two quantities, namely, an estimate of the noise power spectrum and of the noisy speech power spectrum. The variance of these power spectral estimates degrades the quality of the enhanced signal and smoothing techniques are, therefore, often used to decrease the variance. In this paper, we present a method to determine the noisy speech power spectrum based on an adaptive time segmentation. More specifically, the proposed algorithm determines for each noisy frame which of the surrounding frames should contribute to the corresponding noisy power spectral estimate. Further, we demonstrate the potential of our adaptive segmentation in both maximum likelihood and decision direction-based speech enhancement methods by making a better estimate of the a priori signal-to-noise ratio (SNR) xi. Objective and subjective experiments show that an adaptive time segmentation leads to significant performance improvements in comparison to the conventionally used fixed segmentations, particularly in transitional regions, where we observe local SNR improvements in the order of 5 dB Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | Adaptive Time Segmentation of Noisy Speech for Improved Speech EnhancementabstractEnhancement algorithms are widely used to overcome the degradation of noisy speech signals. Most enhancement algorithms require an estimate of the noise and noisy speech power spectra in order to compute the gain function used for the noise suppression. The variance of these power spectral estimates degrades the quality of the enhanced signal and smoothing techniques are therefore often used to decrease the variance. We present a method to determine the noisy speech power spectrum based on an adaptive time segmentation. More specifically, the proposed algorithm determines for each noisy frame which of the surrounding frames should contribute to the corresponding noisy power spectral estimate. Objective and subjective experiments show that an adaptive time segmentation leads to significant performance improvements, particularly in transitional speech regions. Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
ICASSP (1) | 1 |
| 2005 | Improved decision directed approach for speech enhancement using an adaptive time segmentationabstractShort-time Fourier transform (STFT) methods are often used to overcome the degradation of speech signals affected by noise. STFT-gain functions are usually expressed as a function of the a priori SNR, say ξ, and good techniques to estimate ξ are of vital importance for the quality of enhanced speech. Often, ξ is estimated using the so-called decision directed approach (DD). However, the DD approach builds on a number of approximations, where certain expected values of signal related quantities are approximated by instantaneous estimates. In this paper we present a method to improve these approximations by combining the DD approach with an adaptive time segmentation. Objective and subjective experiments show that the proposed method leads to significant improvements compared to the conventional DD approach. Furthermore, simulation experiments confirm a decreased amount of non-stationary residual noise. Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
INTERSPEECH | 1 |
| 2004 | Perceptual linear predictive noise modelling for sinusoid-plus-noise audio codingabstractSinusoidal coding of an audio subject to a bit-rate constraint, in general, results in a noise-like residual signal. This residual signal is of high perceptual importance; reconstruction of audio using the sinusoidal representation only typically results in an artificial sounding reconstruction. We present a new method, called perceptual linear predictive coding (PLPC), where the residual is encoded by applying LPC in the perceptual domain. This method minimizes a perceptual modelling error and therefore represents only residual components that are of perceptual relevance, while automatically discarding components masked by the sinusoidally coded part. Subjective listening tests show that PLPC performs significantly better than ordinary LPC as a sinusoidal residual coding technique. Furthermore, PLPC combined with a flexible segmentation and model order allocation algorithm leads to a significant gain in terms of R/D performance for fragments with fast changing characteristics. Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001 |
ICASSP (4) | 1 |