Toon van Waterschoot

dblp:04/4841 · DBLP profile ↗
← Back
58ranked-venue papers
8as first author
17since 2021 · last 2025
0000-0002-6323-7350ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 31 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 26 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 A Zero-Shot Physics-Informed Dictionary Learning Approach for Sound Field Reconstruction
abstract
Sound field reconstruction aims to estimate pressure fields in areas lacking direct measurements. Existing techniques often rely on strong assumptions or face challenges related to data availability or the explicit modeling of physical properties. To bridge these gaps, this study introduces a zero-shot, physics-informed dictionary learning approach to perform sound field reconstruction. Our method relies only on a few sparse measurements to learn a dictionary, without the need for additional training data. Moreover, by enforcing the Helmholtz equation during the optimization process, the proposed approach ensures that the reconstructed sound field is represented as a linear combination of a few physically meaningful atoms. Evaluations on real-world data show that our approach achieves comparable performance to state-of-the-art dictionary learning techniques, with the advantage of requiring only a few observations of the sound field and no training on a dataset.
Stefano Damiano, Federico Miotello, Mirco Pezzoli, Alberto Bernardini, Fabio Antonacci, Augusto Sarti, Toon van Waterschoot
ICASSP7
2025 A Comparative Analysis of Generalised Echo and Interference Cancelling and Extended Multichannel Wiener Filtering for Combined Noise Reduction and Acoustic Echo Cancellation
abstract
Two algorithms for combined acoustic echo cancellation (AEC) and noise reduction (NR) are analysed, namely the generalised echo and interference canceller (GEIC) and the extended multichannel Wiener filter (MWFext). Previously, these algorithms have been examined for linear echo paths, and assuming access to voice activity detectors (VADs) that separately detect desired speech and echo activity. However, algorithms implementing VADs may introduce detection errors. Therefore, in this paper, the previous analyses are extended by 1) modelling general nonlinear echo paths by means of the generalised Bussgang decomposition, and 2) modelling VAD error effects in each specific algorithm, thereby also allowing to model specific VAD assumptions. It is found and verified with simulations that, generally, the MWFextachieves a higher NR performance, while the GEIC achieves a more robust AEC performance.
Arnout Roebben, Toon van Waterschoot, Marc Moonen
ICASSP2
2025 A state-of-the-art review on acoustic preservation of historical worship spaces through auralization
abstract
Historical Worship Spaces (HWS) are significant architectural landmarks which hold both cultural and spiritual value. The acoustic properties of these spaces play a crucial role in historical and contemporary religious liturgies, rituals, and ceremonies, as well as in the performance of sacred music. However, the original acoustic characteristics of these spaces are often at risk due to repurposing, renovations, natural disasters, or deterioration over time. This paper presents a comprehensive review of the current state of research on the acquisition, analysis, and synthesis of acoustics, with a focus on HWS. An example case study of the Nassau chapel in Brussels, Belgium, is presented to demonstrate the application of these techniques for the preservation and auralization of historical worship space acoustics. The paper concludes with a discussion of the challenges and opportunities in the field, and outlines future research directions.
Hannes Rosseel, Toon van Waterschoot
Signal Process.2
2024 Can Synthetic Data Boost the Training of Deep Acoustic Vehicle Counting Networks?
abstract
In the design of traffic monitoring solutions for optimizing the urban mobility infrastructure, acoustic vehicle counting models have received attention due to their cost effectiveness and energy efficiency. Although deep learning has proven effective for visual traffic monitoring, its use has not been thoroughly investigated in the audio domain, likely due to real-world data scarcity. In this work, we propose a novel approach to acoustic vehicle counting by developing: i) a traffic noise simulation framework to synthesize realistic vehicle pass-by events; ii) a strategy to mix synthetic and real data to train a deep-learning model for traffic counting. The proposed system is capable of simultaneously counting cars and commercial vehicles driving on a two-lane road, and identifying their direction of travel under moderate traffic density conditions. With only 24 hours of labeled real-world traffic noise, we are able to improve counting accuracy on real-world data from 63% to 88% for cars and from 86% to 94% for commercial vehicles.
Stefano Damiano, Luca Bondi, Shabnam Ghaffarzadegan, Andre Guntoro, Toon van Waterschoot
ICASSP5
2024 Microphone Subset Selection for the Weighted Prediction Error Algorithm Using a Group Sparsity Penalty
abstract
Reverberation can severely degrade the quality of speech signals recorded using microphones in an enclosure. In acoustic sensor networks with spatially distributed microphones, a similar dereverberation performance may be achieved using only a subset of all available microphones. Using the popular convex relaxation method, in this paper we propose to perform microphone subset selection for the weighted prediction error (WPE) multi-channel dereverberation algorithm by introducing a group sparsity penalty on the prediction filter coefficients. The resulting problem is shown to be solved efficiently using the accelerated proximal gradient algorithm. Experimental evaluation using measured impulse responses shows that the performance of the proposed method is close to the optimal performance obtained by exhaustive search, both for frequency-dependent as well as frequency-independent microphone subset selection. Furthermore, the performance using only a few microphones for frequency-independent microphone subset selection is only marginally worse than using all available microphones.
Anselm Lohmann, Toon van Waterschoot, Jörg Bitzer, Simon Doclo
ICASSP2
2024 Scalable-Complexity Steered Response Power Based on Low-Rank and Sparse Interpolation
abstract
The steered response power (SRP) is a popular approach to compute a map of the acoustic scene, typically used for acoustic source localization. The SRP map is obtained as the frequency-weighted output power of a beamformer steered towards a grid of candidate locations. Due to the exhaustive search over a fine grid at all frequency bins, conventional frequency domain-based SRP (conv. FD-SRP) results in a high computational complexity. Time domain-based SRP (conv. TD-SRP) implementations reduce computational complexity at the cost of accuracy using the inverse fast Fourier transform (iFFT). In this paper, to enable a more favourable complexity-performance trade-off as compared to conv. FD-SRP and conv. TD-SRP, we consider the problem of constructing a fine SRP map over the entire search space at scalable computational cost. We propose two approaches to this problem. Expressing the conv. FD-SRP map as a matrix transform of frequency-domain GCCs, we decompose the SRP matrix into a sampling matrix and an interpolation matrix. While sampling can be implemented by the iFFT, we propose to use optimal low-rank or sparse approximations of the interpolation matrix for complexity reduction. The proposed approaches, refered to as sampling + low-rank interpolation-based SRP (SLRI-SRP) and sampling + sparse interpolation-based SRP (SSPI-SRP), are evaluated in various localization scenarios with speech as source signals and compared to the state-of-the-art. The results indicate that SSPI-SRP performs better if large array apertures are used, while SLRI-SRP performs better at small array apertures or a large number of microphones. In comparison to conv. FD-SRP, two to three orders of magnitude of complexity reduction can achieved, often times enabling a more favourable complexity-performance trade-off as compared to conv. TD-SRP. A MATLAB implementation is available online.
Thomas Dietzen, Enzo De Sena, Toon van Waterschoot
IEEE ACM Trans. Audio Speech Lang. Process.3
2024 Multi-Source Direction-of-Arrival Estimation Using Steered Response Power and Group-Sparse Optimization
abstract
In this paper, a method is proposed for estimating the direction of arrival (DOA) of multiple broadband sound sources. This is achieved by solving a group-sparse optimization problem, which models an observed broadband steered response power (SRP) map as a linear function of power spectral densities (PSDs), associated to a set of candidate DOAs. The estimation of the source DOAs is then accomplished by identifying peaks in the resulting spatial power density, i.e., the estimated direction-specific PSDs integrated over frequency. The proposed method is motivated by its potential to reveal more distinct peaks in the estimated spatial power density than those directly observed in the SRP map, which can be beneficial to the robustness in DOA estimation performance when multiple sources need to be distinguished under varying acoustic conditions. An implementation of the proposed method using the alternating direction method of multipliers (ADMM) is presented, and the DOA estimation performance is evaluated with both simulated and experimental data. Results show that, especially in reverberant scenarios, the proposed method presents an advantage in locating closely spaced sources when compared to SRP-PHAT, the group-sparse iterative covariance-based estimation (GSPICE) method, and the wideband MUSIC method with geometric averaging. Furthermore, it is observed that for a compact microphone array, the proposed method overall maintained its performance even when using SRP maps with lower grid resolutions than the sampling requirements of the broadband SRP function. Finally, results obtained with experimental data showed the applicability of the proposed method in a practical meeting room environment.
Elisa Tengan, Thomas Dietzen, Filip Elvander, Toon van Waterschoot
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Real-Time Acoustic Perception for Automotive Applications
abstract
In recent years the automotive industry has been strongly promoting the development of smart cars, equipped with multi-modal sensors to gather information about the surroundings, in order to aid human drivers or make autonomous decisions. While the focus has mostly been on visual sensors, also acoustic events are crucial to detect situations that require a change in the driving behavior, such as a car honking, or the sirens of approaching emergency vehicles. In this paper, we summarize the results achieved so far in the Marie Sklodowska-Curie Actions (MSCA) Eruopean Industrial Doctorates (EID) project “Intelligent Ultra Low-Power Signal Processing for Automotive (I-SPOT)”. On the algorithmic side, the I-SPOT Project aims to enable detecting, localizing and tracking environmental audio signals by jointly developing microphone array processing and deep learning techniques that specifically target automotive applications. Data generation software has been developed to cover the I-SPOT target scenarios and research challenges. This tool is currently being used to develop low-complexity deep learning techniques for emergency sound detection. On the hardware side, the goal impels workflows for hardware-algorithm co-design to ease the generation of architectures that are sufficiently flexible towards algorithmic evolutions without giving up on efficiency, as well as enable rapid feedback of hardware implications of algorithmic decision. This is pursued though a hierarchical workflow that breaks the hardware-algorithm design space into reasonable subsets, which has been tested for operator-level optimizations on state-of-the-art robust sound source localization for edge devices. Further, several open challenges towards an end-to-end system are clarified for the next stage of I-SPOT.
Jun Yin 0001, Stefano Damiano, Marian Verhelst, Toon van Waterschoot, Andre Guntoro
DATE4
2023 Distributed Adaptive Norm Estimation for Blind System Identification in Wireless Sensor Networks
abstract
Distributed signal-processing algorithms in (wireless) sensor networks often aim to decentralize processing tasks to reduce communication cost and computational complexity or avoid reliance on a single device (i.e., fusion center) for processing. In this contribution, we extend a distributed adaptive algorithm for blind system identification that relies on the estimation of a stacked network-wide consensus vector at each node, the computation of which requires either broadcasting or relaying of node-specific values (i.e., local vector norms) to all other nodes. The extended algorithm employs a distributed-averaging-based scheme to estimate the network-wide consensus norm value by only using the local vector norm provided by neighboring sensor nodes. We introduce an adaptive mixing factor between instantaneous and recursive estimates of these norms for adaptivity in a time-varying system. Simulation results show that the extension provides estimation results close to the optimal fully-connected-network or broadcasting case while reducing inter-node transmission significantly.
Matthias Blochberger, Filip Elvander, Randall Ali, Jan Østergaard, Jesper Jensen 0001, Marc Moonen, Toon van Waterschoot
ICASSP7
2023 Fast Low-Latency Convolution by Low-Rank Tensor Approximation
abstract
In this paper we consider fast time-domain convolution, exploiting low-rank properties of an impulse response (IR). This reduces the computational complexity, speeding up the convolution, without introducing latency. Previous work has considered a truncated singular value decomposition (SVD) of a two-dimensional matricization, or reshaping, of the IR. We here build upon this idea, by providing an algorithm for convolution with a three-dimensional tensorization of the IR. We provide simulations using real-life acoustic room impulse responses (RIRs) of various lengths, convolving them with music, as well as speech signals. The proposed algorithm is shown to outperform the comparable existing algorithm in terms of signal quality degradation, for all considered scenarios, without increasing the computational complexity, or the memory usage.
Martin Jälmby, Filip Elvander, Toon van Waterschoot
ICASSP3
2023 Dereverberation in Acoustic Sensor Networks Using weighted Prediction Error with Microphone-Dependent Prediction Delays
abstract
In the last decades several multi-microphone speech dereverberation algorithms have been proposed, among which the weighted prediction error (WPE) algorithm. In the WPE algorithm, a prediction delay is required to reduce the correlation between the prediction signals and the direct component in the reference microphone signal. In compact arrays with closely-spaced microphones, the prediction delay is often chosen microphone-independent. In acoustic sensor networks with spatially distributed microphones, large time-differences-of-arrival (TDOAs) of the speech source between the reference microphone and other microphones typically occur. Hence, when using a microphone-independent prediction delay1the reference and prediction signals may still be significantly correlated, leading to distortion in the dereverberated output signal. In order to decorrelate the signals, in this paper we propose to apply TDOA compensation with respect to the reference microphone, resulting in microphone-dependent prediction delays for the WPE algorithm. We consider both optimal TDOA compensation using crossband filtering in the short-time Fourier transform domain as well as band-to-band and integer delay approximations. Simulation results for different reverberation times using oracle as well as estimated TDOAs clearly show the benefit of using microphone-dependent prediction delays.
Anselm Lohmann, Toon van Waterschoot, Jörg Bitzer, Simon Doclo
ICASSP2
2023 Simultaneous Acoustic Echo Sorting and 3-D Room Geometry Inference
abstract
Room geometry estimation from multiple acoustic room impulse responses (RIRs) relies on being able to correctly identify first-order echoes corresponding to the same reflector. This can be done by exploiting the properties of various mathematical concepts but often results in needing to solve large combinatorial problems owing to the increased number of measurements these tools require for efficacy. A low-complexity iterative method is proposed for determining the boundaries in convex polygonal rooms using common tangent planes to sets of ellipsoids. Prior knowledge of the order of echoes is unnecessary and only three RIRs are required, which is one less than the minimum number typically needed to find a unique reflector in 3-D, allowing for simpler common tangent plane estimation and reducing possible echo combinations. Candidate partial rooms are built by checking whether expected second-order echoes appear in the RIRs, adding a reflector at each iterate until a room is found. The proposed method is validated by means of computer simulations.
Kathleen MacWilliam, Filip Elvander, Toon van Waterschoot
ICASSP3
2023 Centralized Cascade Multi-Channel Noise Reduction and Acoustic Feedback Cancellation in a Wireless Acoustic Sensor And Actuator Network
abstract
This paper presents a centralized cascade multi-channel noise reduction (NR) and acoustic feedback cancellation (AFC) algorithm for speech applications in a wireless acoustic sensor and actuator network (WASAN). The algorithm consists of a multi-channel Wiener filter (MWF) based NR stage, where M microphone and L loud-speaker signals in the network are used to estimate, for each node, the speech component in its reference microphone and loudspeaker signal. For each node then a prediction error method based AFC stage is applied using these estimates to estimate the desired signal. Closed-loop simulations show that the proposed centralized algorithm outperforms a node working in isolation, in terms of the signal-to-interference-plus-noise ratio improvement (∆SINR) and short-time objective intelligibility (STOI) metric. In particular, the performance is improved when for instance one of the nodes has a lower input signal-to-noise ratio (iSNR) than the other nodes. In addition, it is also shown that a single-channel local adaptive filter per node is sufficient in the AFC stage, instead of a K-channel adaptive filter. Based on this and on the availability of a distributed algorithm for the NR stage, the presented algorithm is viewed as a stepping stone towards a fully distributed NR and AFC algorithm.
Santiago Ruiz, Toon van Waterschoot, Marc Moonen
ICASSP2
2023 Low-Rank Room Impulse Response Estimation
abstract
In this paper we consider low-rank estimation of room impulse responses (RIRs). Inspired by a physics-driven room-acoustical model, we propose an estimator of RIRs that promotes a low-rank structure for a matricization, or reshaping, of the estimated RIR. This low-rank prior acts as a regularizer for the inverse problem of estimating an RIR from input-output observations, preventing overfitting and improving estimation accuracy. As directly enforcing a low rank of the estimate results is an NP-hard problem, we consider two different relaxations, one using the nuclear norm, and one using the recently introduced concept of quadratic envelopes. Both relaxations allow for implementing the proposed estimator using a first-order algorithm with convergence guarantees. When evaluated on both synthetic and recorded RIRs, it is shown that under noisy output conditions, or when the spectral excitation of the input signal is poor, the proposed estimator outperforms comparable existing methods. The performance of the two low-rank relaxations methods is similar, but the quadratic envelope has the benefit of superior robustness to the choice of regularization hyperparameter in the case when the signal-to-noise ratio is unknown. The performance of the proposed method is compared to that of ordinary least squares, Tikhonov least squares, as well as the Cramér-Rao lower bound (CRLB).
Martin Jälmby, Filip Elvander, Toon van Waterschoot
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Cascade Multi-Channel Noise Reduction and Acoustic Feedback Cancellation
abstract
Acoustic feedback and noise are common problems that corrupt microphone signals and affect the performance of speech and audio signal processing applications and devices. In this paper, a cascade noise reduction (NR) and acoustic feedback cancellation (AFC) algorithm is presented for speech applications where a multi-channel Wiener filter (MWF) based NR is applied first followed by a single-channel prediction-error method (PEM) based adaptive feedback cancellation stage. It is shown that by using a rank-2 estimate of the speech correlation matrix in the NR stage it is possible to obtain a good feedback path estimate for the reference microphone in the AFC stage. Closed-loop simulations with M microphones and 1 loudspeaker are presented using both an M-channel rank-1 and an (M + 1)-channel rank-2 MWF and it is shown that for the considered input signal-to-noise ratios the proposed algorithm increases the added stable gain (ASG) of the system.
Santiago Ruiz, Toon van Waterschoot, Marc Moonen
ICASSP2
2022 Distributed Combined Acoustic Echo Cancellation and Noise Reduction in Wireless Acoustic Sensor and Actuator Networks
abstract
The paper presents distributed algorithms for combined acoustic echo cancellation (AEC) and noise reduction (NR) in a wireless acoustic sensor and actuator network (WASAN) where each node may have multiple microphones and multiple loudspeakers, and where the desired signal is a speech signal. A centralized integrated AEC and NR algorithm, i.e., multichannel Wiener filter (MWF), is used as starting point where echo signals are viewed as background noise signals and loudspeaker signals are used as additional input signals to the algorithm. By including prior knowledge (PK), namely that the loudspeaker signals do not contain any desired signal component, an alternative centralized cascade algorithm (PK-MWF) is obtained with an AEC stage first followed by an MWF-based NR stage which has a lower computational complexity. Distributed algorithms can then be obtained from the MWF and PK-MWF algorithm, i.e., the generalized eigenvalue decomposition (GEVD)-based distributed adaptive node-specific signal estimation (DANSE) and PK-GEVD-DANSE algorithm, respectively. In the former, each node performs a reduced dimensional integrated AEC and NR algorithm and broadcasts only 1 fused signal (instead of all its signals) to the other nodes. In the PK-GEVD-DANSE algorithm, each node performs a reduced dimensional cascade AEC and NR algorithm and broadcasts only 2 fused signals (instead of all its signals) to the other nodes. The distributed algorithms achieve the same performance, upon convergence, as the corresponding centralized integrated (MWF) and centralized cascade (PK-MWF) algorithm. It is observed, however, that the communication cost in the PK-GEVD-DANSE algorithm can also be reduced, where each node then broadcasts only 1 fused signal (instead of 2 signals) to the other nodes. The resulting algorithm, referred to as the pruned PK-GEVD-DANSE (pPK-GEVD-DANSE) algorithm, then effectively combines the lowest possible communication cost (as low as in the GEVD-DANSE algorithm) with a lowest possible computational complexity in each node (further reduced from the PK-GEVD-DANSE computational complexity), within the class of algorithms considered in this paper.
Santiago Ruiz, Toon van Waterschoot, Marc Moonen
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 Evaluation of a Novel Speech-in-Noise Test for Hearing Screening: Classification Performance and Transducers' Characteristics
abstract
One of the current gaps in teleaudiology is the lack of methods for adult hearing screening viable for use in individuals of unknown language and in varying environments. We have developed a novel automated speech-in-noise test that uses stimuli viable for use in non-native listeners. The test reliability has been demonstrated in laboratory settings and in uncontrolled environmental noise settings in previous studies. The aim of this study was: (i) to evaluate the ability of the test to identify hearing loss using multivariate logistic regression classifiers in a population of 148 unscreened adults and (ii) to evaluate the ear-level sound pressure levels generated by different earphones and headphones as a function of the test volume. The multivariate classifiers had sensitivity equal to 0.79 and specificity equal to 0.79 using both the full set of features extracted from the test as well as a subset of three features (speech recognition threshold, age, and number of correct responses). The analysis of the ear-level sound pressure levels showed substantial variability across transducer types and models, with earphones levels being up to 22 dB lower than those of headphones. Overall, these results suggest that the proposed approach might be viable for hearing screening in varying environments if an option to self-adjust the test volume is included and if headphones are used. Future research is needed to assess the viability of the test for screening at a distance, for example by addressing the influence of user interface, device, and settings, on a large sample of subjects with varying hearing loss.
Marco Zanet, Edoardo Maria Polo, Marta Lenatti, Toon van Waterschoot, Maurizio Mongelli, Riccardo Barbieri, Alessia Paglialonga
IEEE J. Biomed. Health Informatics4
2020 Integrated Sidelobe Cancellation and Linear Prediction Kalman Filter for Joint Multi-Microphone Speech Dereverberation, Interfering Speech Cancellation, and Noise Reduction
abstract
In multi-microphone speech enhancement, reverberation as well as additive noise and/or interfering speech are commonly suppressed by deconvolution and spatial filtering, e.g., using multi-channel linear prediction (MCLP) on the one hand and beamforming, e.g., a generalized sidelobe canceler (GSC), on the other hand. In this article, we consider several reverberant speech components, whereof some are to be dereverberated and others to be canceled, as well as a diffuse (e.g., babble) noise component to be suppressed. In order to perform both deconvolution and spatial filtering, we integrate MCLP and the GSC into a novel architecture referred to as integrated sidelobe cancellation and linear prediction (ISCLP), where the sidelobe-cancellation (SC) filter and the linear prediction (LP) filter operate in parallel, but on different microphone signal frames. Within ISCLP, we estimate both filters jointly by means of a single Kalman filter. We further propose a spectral Wiener gain post-processor, which is shown to relate to the Kalman filter's posterior state estimate. The presented ISCLP Kalman filter is benchmarked against two state-of-the-art approaches, namely first a pair of alternating Kalman filters respectively performing dereverberation and noise reduction, and second an MCLP+GSC Kalman filter cascade. While the ISCLP Kalman filter is roughly $M^2$ times less expensive than both reference algorithms, where $M$ denotes the number of microphones, it is shown to perform at least similarly as compared to the former, and to outperform the latter. A MATLAB implementation is available.
Thomas Dietzen, Simon Doclo, Marc Moonen, Toon van Waterschoot
IEEE ACM Trans. Audio Speech Lang. Process.4
2020 Square Root-Based Multi-Source Early PSD Estimation and Recursive RETF Update in Reverberant Environments by Means of the Orthogonal Procrustes Problem
abstract
Multi-channel short-time Fourier transform (STFT) domain-based processing of reverberant microphone signals commonly relies on power-spectral-density (PSD) estimates of early source images, where early refers to reflections contained within the same STFT frame. State-of-the-art approaches to multi-source early PSD estimation, given an estimate of the associated relative early transfer functions (RETFs), conventionally minimize the approximation error defined with respect to the early correlation matrix, requiring non-negative inequality constraints on the PSDs. Instead, we here propose to factorize the early correlation matrix and minimize the approximation error defined with respect to the early-correlation-matrix square root. The proposed minimization problem-constituting a generalization of the so-called orthogonal Procrustes problem-seeks a unitary matrix and the square roots of the early PSDs up to an arbitrary complex argument, whereby non-negative inequality constraints become redundant. A solution is obtained iteratively, requiring one singular value decomposition (SVD) per iteration. The estimated unitary matrix and early PSD square roots further allow to recursively update the RETF estimate, which is not inherently possible in the conventional approach. An estimate of the said early-correlation-matrix square root itself is obtained by means of the generalized eigenvalue decomposition (GEVD), where we further propose to restore non-stationarities by desmoothing the generalized eigenvalues in order to compensate for inevitable recursive averaging. Simulation results indicate fast convergence of the proposed multi-source early PSD estimation approach in only one iteration if initialized appropriately, and better performance as compared to the conventional approach. A MATLAB implementation is available.
Thomas Dietzen, Simon Doclo, Marc Moonen, Toon van Waterschoot
IEEE ACM Trans. Audio Speech Lang. Process.4
2020 Localization Uncertainty in Time-Amplitude Stereophonic Reproduction
abstract
This article studies the effects of inter-channel time and level differences in stereophonic reproduction on perceived localization uncertainty, which is defined as how difficult it is for a listener to tell where a sound source is located. Towards this end, a computational model of localization uncertainty is proposed first. The model calculates inter-aural time and level difference cues, and compares them to those associated to free-field point-like sources. The comparison is carried out using a particular distance functional that replicates the increased uncertainty observed experimentally with inconsistent inter-aural time and level difference cues. The model is validated by formal listening tests, achieving a Pearson correlation of 0.99. The model is then used to predict localization uncertainty for stereophonic setups and a listener in central and off-central positions. Results show that amplitude methods achieve a slightly lower localization uncertainty for a listener positioned exactly in the center of the sweet spot. As soon as the listener moves away from that position, the situation reverses, with time-amplitude methods achieving a lower localization uncertainty.
Enzo De Sena, Zoran Cvetkovic, Hüseyin Hacihabiboglu, Marc Moonen, Toon van Waterschoot
IEEE ACM Trans. Audio Speech Lang. Process.5
2019 Methods of Extending a Generalized Sidelobe Canceller With External Microphones
abstract
While substantial noise reduction and speech enhancement can be achieved with multiple microphones organized in an array, in some cases, such as when the microphone spacings are quite close, it can also be quite limited. This degradation can, however, be resolved by the introduction of one or more external microphones (XMs) into the same physical space as the local microphone array (LMA). In this paper, three methods of extending an LMA-based generalized sidelobe canceller (GSC-LMA) with multiple XMs are proposed in such a manner that the relative transfer function pertaining to the LMA is treated as a priori knowledge. Two of these methods involve a procedure for completing an extended blocking matrix, whereas the third uses the speech estimate from the GSC-LMA directly with an orthogonalized version of the XM signals to obtain an improved speech estimate via a rank-1 generalized eigenvalue decomposition. All three methods were evaluated with recorded data from an office room and it was found that the third method could offer the most improvement. It was also shown that in using this method, the speech estimate from the GSC-LMA was not compromised and would be available to the listener if so desired, along with the improved speech estimate that uses both the LMA and XMs.
Randall Ali, Giuliano Bernardi, Toon van Waterschoot, Marc Moonen
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 Integration of a Priori and Estimated Constraints Into an MVDR Beamformer for Speech Enhancement
abstract
Conventionally, the single constraint of the minimum variance distortionless response (MVDR) beamformer for speech enhancement has been defined using one of two approaches. Either it is based on a priori assumptions such as microphone characteristics, position, speech source location, and room acoustics, or on a relative transfer function (RTF) vector estimate using a data dependent method. Each approach has its respective merits and drawbacks and a decision usually has to be made between one of the approaches. In this paper, an alternative approach of using an integrated MVDR beamformer is investigated, where both the hard constraints from the two conventional approaches are softened to yield two tuning parameters. It will be shown that this integrated MVDR beamformer can be expressed as a convex combination of the conventional MVDR beamformers, a linearly constrained minimum variance (LCMV) beamformer, and an all-zero vector, with real, positive-valued coefficients. By analysing how the tuning parameters affect these coefficients, two tuning rules for a practical implementation of the integrated MVDR are subsequently proposed. An evaluation with simulated and recorded data demonstrates that the integrated MVDR beamformer can be beneficial as opposed to relying on either of the conventional MVDR beamformers.
Randall Ali, Toon van Waterschoot, Marc Moonen
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 Joint Acoustic Localization and Dereverberation Through Plane Wave Decomposition and Sparse Regularization
abstract
Acoustic source localization and dereverberation are formulated jointly as an inverse problem. The inverse problem consists of the approximation of the sound field measured by a set of microphones. The recorded sound pressure is matched with that of a particular acoustic model based on a collection of plane waves arriving from different directions at the microphone positions. In order to achieve meaningful results, spatial and spatio-spectral sparsity can be promoted in the weight signals controlling the plane waves. The large-scale optimization problem resulting from the inverse problem formulation is solved using a first order optimization algorithm combined with a weighted overlap-add procedure. It is shown that once the weight signals capable of effectively approximating the sound field are obtained, they can be readily used to localize a moving sound source in terms of direction of arrival (DOA) and to perform dereverberation in a highly reverberant environment. Results from simulation experiments and from real measurements show that the proposed algorithm is robust against both localized and diffuse noise exhibiting a noise reduction in the dereverberated signals.
Niccolò Antonello, Enzo De Sena, Marc Moonen, Patrick A. Naylor, Toon van Waterschoot
IEEE ACM Trans. Audio Speech Lang. Process.5
2019 Comparative Analysis of Generalized Sidelobe Cancellation and Multi-Channel Linear Prediction for Speech Dereverberation and Noise Reduction
abstract
For blind speech dereverberation, two frameworks are commonly used: on the one hand, the multi-channel linear prediction (MCLP) framework, and on the other hand, data-dependent beamforming, e.g., the generalized sidelobe canceler (GSC) framework. The MCLP framework is designed to perform deconvolution and hence has gained increased prominence in blind speech dereverberation. The GSC framework is commonly used for noise reduction, but may be applied for dereverberation as well. In previous work, we have shown that for the noiseless case, MCLP and the GSC yield in theory mathematically equivalent results in terms of dereverberation. In this paper, we assume additional coherent as well as incoherent-noise components and formally analyze and compare both frameworks in terms of dereverberation and noise reduction performance. Both the theoretical analysis and time domain simulation results demonstrate that unlike the GSC, MCLP expectably shows limited performance in terms of noise reduction, while both perform equally well in terms of dereverberation, provided that the GSC blocking matrix achieves complete blocking of the early reverberant-speech component and sufficiently many microphones are available. In case of incomplete blocking, however, the GSC performs inferior to MCLP in terms of dereverberation, as shown in short-time Fourier transform domain simulations.
Thomas Dietzen, Ann Spriet, Wouter Tirry, Simon Doclo, Marc Moonen, Toon van Waterschoot
IEEE ACM Trans. Audio Speech Lang. Process.6
2018 Generalised Sidelobe Canceller for Noise Reduction in Hearing Devices Using an External Microphone
abstract
The use of an external microphone in conjunction with an existing local microphone array can be particularly beneficial for noise reduction tasks that are critical for hearing devices, such as cochlear implants and hearing aids. Recent work has already demonstrated how an external microphone signal can be effectively incorporated into the common noise reduction technique of using a Minimum Variance Distortionless Response (MVDR) beamformer. In this paper, we provide a further extension, whereby an external microphone signal can be incorporated into an existing framework of a Generalised Sidelobe Canceller (GSC) that has been designed for a local microphone signal array. It will be shown that the resulting GSC with an external microphone results in an easily implementable addition to the existing GSC framework for a local microphone array, and can exhibit an improved noise reduction performance.
Randall Ali, Toon van Waterschoot, Marc Moonen
ICASSP2
2018 Joint Source Localization and Dereverberation by Sound Field Interpolation Using Sparse Regularization
abstract
In this paper, source localization and dereverberation are formulated jointly as an inverse problem. The inverse problem consists in the interpolation of the sound field measured by a set of microphones by matching the recorded sound pressure with that of a particular acoustic model. This model is based on a collection of equivalent sources creating either spherical or plane waves. In order to achieve meaningful results, spatial, spatio-temporal and spatio-spectral sparsity can be promoted in the signals originating from the equivalent sources. The inverse problem consists of a large-scale optimization problem that is solved using a first order matrix-free optimization algorithm. It is shown that once the equivalent source signals capable of effectively interpolating the sound field are obtained, they can be readily used to localize a speech sound source in terms of Direction of Arrival (DOA) and to perform dereverberation in a highly reverberant environment.
Niccolò Antonello, Enzo De Sena, Marc Moonen, Patrick A. Naylor, Toon van Waterschoot
ICASSP5
2018 Subjective and Objective Sound-Quality Evaluation of Adaptive Feedback Cancellation Algorithms
abstract
Objective measures are widely used for the perceptual sound-quality evaluation of audio signal processing algorithms. Nevertheless, the use of subjective-evaluation measures remains relevant, in particular when application-specific objective measures are lacking. In this paper, we present a perceptual sound-quality evaluation of different algorithms for adaptive feedback cancellation (AFC), with both speech and music signals. Three algorithms are compared: the block normalized least mean squares algorithm, the prediction-error method (PEM) based frequency-domain adaptive filter, and the PEM-based frequency-domain Kalman filter (PEM-FDKF). The subjective evaluation results for the tested algorithms suggest that there is a large difference in statistical significance, and a corresponding large effect size, between the PEM-FDKF and the other algorithms, when using speech signals. A smaller statistical significance, and a lower effect size, is reported when using music signals. The subjective evaluation results are then compared with the results obtained with several objective measures. The correlation between subjective and objective scores shows that objective measures can be effectively used to predict the sound-quality degradation caused by acoustic feedback and AFC artifacts.
Giuliano Bernardi, Toon van Waterschoot, Jan Wouters, Marc Moonen
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Adaptive Speech Dereverberation Using Constrained Sparse Multichannel Linear Prediction
abstract
In this letter, we present an adaptive speech dereverberation method based on constrained sparse multichannel linear prediction (MCLP), minimizing the mixed ℓ2,pnorm of the desired component. In order to prevent overestimation of the undesired reverberant component, possibly leading to severe distortions of the output, we propose to use a statistical model for late reverberation to limit the power of the MCLP-based estimate. The resulting constrained optimization problem is solved by using the alternating direction method of multipliers, resulting in two variants of the dereverberation algorithm. Simulation results show that the proposed constraint increases the robustness with respect to parameter selection and improves the usability for dynamic scenarios in comparison to the unconstrained method.
Ante Jukic, Toon van Waterschoot, Simon Doclo
IEEE Signal Process. Lett.2
2017 Room Impulse Response Interpolation Using a Sparse Spatio-Temporal Representation of the Sound Field
abstract
Room Impulse Responses (RIRs) are typically measured using a set of microphones and a loudspeaker. When RIRs spanning a large volume are needed, many microphone measurements must be used to spatially sample the sound field. In order to reduce the number of microphone measurements, RIRs can be spatially interpolated. In the present study, RIR interpolation is formulated as an inverse problem. This inverse problem relies on a particular acoustic model capable of representing the measurements. Two different acoustic models are compared: the plane wave decomposition model and a novel time-domain model, which consists of a collection of equivalent sources creating spherical waves. These acoustic models can both approximate any reverberant sound field created by a far-field sound source. In order to produce an accurate RIR interpolation, sparsity regularization is employed when solving the inverse problem. In particular, by combining different acoustic models with different sparsity promoting regularizations, spatial sparsity, spatio-spectral sparsity, and spatio-temporal sparsity are compared. The inverse problem is solved using a matrix-free large-scale optimization algorithm. Simulations show that the best RIR interpolation is obtained when combining the novel time-domain acoustic model with the spatio-temporal sparsity regularization, outperforming the results of the plane wave decomposition model even when far fewer microphone measurements are available.
Niccolò Antonello, Enzo De Sena, Marc Moonen, Patrick A. Naylor, Toon van Waterschoot
IEEE ACM Trans. Audio Speech Lang. Process.5
2017 Adaptive Feedback Cancellation Using a Partitioned-Block Frequency-Domain Kalman Filter Approach With PEM-Based Signal Prewhitening
abstract
Adaptive filtering based feedback cancellation is a widespread approach to acoustic feedback control. However, traditional adaptive filtering algorithms have to be modified in order to work satisfactorily in a closed-loop scenario. In particular, the undesired signal correlation between the loudspeaker signal and the source signal in a closed-loop scenario is one of the major problems to address when using adaptive filters for feedback cancellation. Slow convergence speed and limited tracking capabilities are other important limitations to be considered. Additionally, computationally expensive algorithms as well as long delays should be avoided, for instance, in hearing aid applications, because of power constraints, important to extend battery life, and real-time implementations requirements, respectively. We present an algorithm combining good decorrelation properties, by means of the prediction-error method based signal prewhitening, fast convergence, good tracking behavior, and low computational complexity by means of the frequency-domain Kalman filter, and low delay by means of a partitioned-block implementation.
Giuliano Bernardi, Toon van Waterschoot, Jan Wouters, Marc Moonen
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 A Scalable Algorithm for Physically Motivated and Sparse Approximation of Room Impulse Responses With Orthonormal Basis Functions
abstract
Parametric modeling of room acoustics aims at representing room transfer functions by means of digital filters and finds application in many acoustic signal enhancement algorithms. In previous work by other authors, the use of orthonormal basis functions (OBFs) for modeling room acoustics has been proposed. Some advantages of OBF models over all-zero and pole-zero models have been illustrated, mainly focusing on the fact that OBF models typically require less model parameters to provide the same model accuracy. In this paper, it is shown that the orthogonality of the OBF model brings several additional advantages, which can be exploited if a suitable algorithm for identifying the OBF model parameters is applied. Specifically, the orthogonality of OBF models does not only lead to improved model efficiency (as pointed out in previous work), but also leads to improved model scalability and model stability. Its appealing scalability property derives from a previously unexplored interpretation of the OBF model as an approximation to a solution of the inhomogeneous acoustic wave equation. Following this interpretation, a novel identification algorithm is proposed that takes advantage of the OBF model orthogonality to deliver efficient, scalable, and stable OBF model estimates, which is not necessarily the case for nonlinear estimation techniques that are normally applied.
Giacomo Vairetti, Enzo De Sena, Michael Catrysse, Søren Holdt Jensen, Marc Moonen, Toon van Waterschoot
IEEE ACM Trans. Audio Speech Lang. Process.6
2016 Multichannel identification of room acoustic systems with adaptive filters based on orthonormal basis functions
abstract
Many acoustic signal enhancement applications require adaptive filters with a long impulse response, but with a small number of filter parameters. Fixed-poles infinite impulse response (IIR) adaptive filters based on orthonormal basis functions (OBFs) present advantages over finite impulse response filters and other IIR filters, assuring stability and fast global convergence in the adaptation of the filter parameters. A scalable algorithm is introduced for the estimation of the poles of an adaptive OBF filter from multichannel input-output data. The set of poles, common to all the acoustic channels considered, is estimated in parallel to the adaptation of the linear filter parameters. It will be shown that the result of the identification with common poles is quite robust to variations in the room transfer function, suggesting the possibility that poles may be kept fixed after estimation.
Giacomo Vairetti, Søren Holdt Jensen, Enzo De Sena, Marc Moonen, Michael Catrysse, Toon van Waterschoot
ICASSP6
2016 Fast algorithms for high-order sparse linear prediction with applications to speech processing
Tobias Lindstrøm Jensen, Daniele Giacobello, Toon van Waterschoot, Mads Græsbøll Christensen
Speech Commun.3
2016 A Single-Channel Non-Intrusive C50 Estimator Correlated With Speech Recognition Performance
abstract
Several intrusive measures of reverberation can be computed from measured and simulated room impulse responses, over the full frequency band or for each individual mel-frequency subband. It is initially shown that full-band clarity index C50is the most correlated measure on average with reverberant speech recognition performance. This corroborates previous findings but now for the dataset to be used in this study. We extend the previous findings to show that C50also exhibits the highest mutual information on average. Motivated by these extended findings, a nonintrusive room acoustic (NIRA) estimation method is proposed to estimate C50from only the reverberant speech signal. The NIRA method is a data-driven approach based on computing a number of features from the speech signal and it employs these features to train a model used to perform the estimation. The choice of features and learning techniques are explored in this work using an evaluation set which comprises approximately 100 000 different reverberant signals (around 93 h of speech) including reverberation from measured and simulated room impulse responses. The feature importance of each feature with respect to the estimation of the target C50is analysed following two different approaches. In both cases, the newly chosen set of features shows high importance for the target. The best C50estimator provides a root-mean-square deviation around 3 dB on average for all reverberant test environments.
Pablo Peso Parada, Dushyant Sharma, Jose Lainez, Daniel Barreda, Toon van Waterschoot, Patrick A. Naylor
IEEE ACM Trans. Audio Speech Lang. Process.5
2015 Multi-channel linear prediction-based speech dereverberation with low-rank power spectrogram approximation
abstract
In many acoustic conditions the recorded speech signals may be severely affected by reverberation, leading to a reduced speech quality and intelligibility. In this paper we focus on a blind speech dereverberation method based on multi-channel linear prediction (MCLP) in the short-time Fourier transform domain, which is typically performed in each frequency bin independently without taking into account the spectral structure of the speech signal. Since it is widely accepted that a speech spectrogram can be well approximated with a low-rank matrix, e.g., using a spectral dictionary, in this paper we propose to incorporate a low-rank matrix approximation of the speech spectrogram into the MCLP-based speech dereverberation. The low-rank approximation is obtained using nonnegative matrix factorization with Itakura-Saito divergence. Experimental results for several measured acoustic systems show that incorporating a low-rank approximation improves the dereverberation performance in terms of instrumental speech quality measures.
Ante Jukic, Nasser Mohammadiha, Toon van Waterschoot, Timo Gerkmann, Simon Doclo
ICASSP3
2015 A multi-channel speech enhancement framework for robust NMF-based speech recognition for speech-impaired users
abstract
In this paper a multi-channel speech enhancement framework for distant speech acquisition in noisy and reverberant environments for Non-negative Matrix Factorization (NMF)-based Automatic Speech Recognition (ASR) is proposed. The system is evaluated for its use in an assistive vocal interface for physically impaired and speech-impaired users. The framework utilises the Spatially Pre-processed Speech Distortion Weighted Multi-channel Wiener Filter (SP-SDW-MWF) in combination with a postfilter to reduce noise and reverberation. Additionally, the estimation uncertainty of the speech enhancement framework is propagated through the Mel-Frequency Cepstrum Coefficients (MFCC) feature extraction to allow for feature compensation in a later stage. Results indicate that a) using a trade-off parameter between noise reduction and speech distortion has a positive effect on the recognition performance with respect to the well-known GSC and MWF and b) the addition of a postfilter and the feature compensation increases performance with respect to several baselines for a non-pathological and pathological speaker.
Gert Dekkers, Toon van Waterschoot, Bart Vanrumste, Bert Van Den Broeck, Jort F. Gemmeke, Hugo Van hamme, Peter Karsmakers
INTERSPEECH2
2015 Special issue on wireless acoustic sensor networks and ad hoc microphone arrays
Alexander Bertrand, Simon Doclo, Sharon Gannot, Nobutaka Ono, Toon van Waterschoot
Signal Process.5
2015 Multi-Channel Linear Prediction-Based Speech Dereverberation With Sparse Priors
abstract
The quality of speech signals recorded in an enclosure can be severely degraded by room reverberation. In this paper, we focus on a class of blind batch methods for speech dereverberation in a noiseless scenario with a single source, which are based on multi-channel linear prediction in the short-time Fourier transform domain. Dereverberation is performed by maximum-likelihood estimation of the model parameters that are subsequently used to recover the desired speech signal. Contrary to the conventional method, we propose to model the desired speech signal using a general sparse prior that can be represented in a convex form as a maximization over scaled complex Gaussian distributions. The proposed model can be interpreted as a generalization of the commonly used time-varying Gaussian model. Furthermore, we reformulate both the conventional and the proposed method as an optimization problem with an lp-norm cost function, emphasizing the role of sparsity in the considered speech dereverberation methods. Experimental evaluation in different acoustic scenarios show that the proposed approach results in an improved performance compared to the conventional approach in terms of instrumental measures for speech quality.
Ante Jukic, Toon van Waterschoot, Timo Gerkmann, Simon Doclo
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 On the Modeling of Rectangular Geometries in Room Acoustic Simulations
abstract
This paper is concerned with an acoustical phenomenon called sweeping echo, which manifests itself in a room impulse response as a distinctive, continuous pitch increase. In this paper, it is shown that sweeping echoes are present (although to greatly varying degrees) in all perfectly rectangular rooms. The theoretical analysis is based on the rigid-wall image solution of the wave equation. Sweeping echoes are found to be caused by the orderly time-alignment of high-order reflections arriving from directions close to the three axial directions. While sweeping echoes have been previously observed in real rooms with a geometry very similar to the rectangular model (e.g., a squash court), they are not perceived in commonly encountered rooms. Room acoustic simulators such as the image method (IM) and finite difference time-domain (FDTD) correctly predict the presence of this phenomenon, which means that rectangular geometries should be used with caution when the objective is to model commonly encountered rooms. Small out-of-square asymmetries in the room geometry are shown to reduce the phenomenon significantly. Randomization of the image sources' position is shown to remove sweeping echoes without the need to model an asymmetrical geometry explicitly. Finally, the performance of three speech and audio processing algorithms is shown to be sensitive to strong sweeping echoes, thus highlighting the need to avoid their occurrence.
Enzo De Sena, Niccolò Antonello, Marc Moonen, Toon van Waterschoot
IEEE ACM Trans. Audio Speech Lang. Process.4
2014 Wiener variable step size and gradient spectral variance smoothing for double-talk-robust acoustic echo cancellation and acoustic feedback cancellation
Jose Manuel Gil-Cacho, Toon van Waterschoot, Marc Moonen, Søren Holdt Jensen
Signal Process.2
2014 Embedded-Optimization-Based Loudspeaker Precompensation Using a Hammerstein Loudspeaker Model
abstract
This paper presents an embedded-optimization-based loudspeaker precompensation algorithm using a Hammerstein loudspeaker model, i.e. a cascade of a memoryless nonlinearity and a linear finite impulse response filter. The loudspeaker precompensation consists in a per-frame signal optimization. In order to minimize the perceptible distortion incurred in the loudspeaker, a psychoacoustically motivated optimization criterion is proposed. The resulting per-frame signal optimization problems are solved efficiently using first-order optimization methods. Depending on the invertibility and the smoothness of the memoryless nonlinearity, different first-order optimization methods are proposed and their convergence properties are analyzed. Objective evaluation experiments using synthetic loudspeaker models and real loudspeakers show that the proposed loudspeaker precompensation algorithm provides a significant audio quality improvement, especially so at high playback levels.
Bruno Defraene, Toon van Waterschoot, Moritz Diehl, Marc Moonen
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 A frequency-domain adaptive filter (FDAF) prediction error method (PEM) framework for double-talk-robust acoustic echo cancellation
abstract
In this paper, we propose a new framework to tackle the double-talk (DT) problem in acoustic echo cancellation (AEC). It is based on a frequency-domain adaptive filter (FDAF) implementation of the so-called prediction error method adaptive filtering using row operations (PEM-AFROW) leading to the FDAF-PEM-AFROW algorithm. We show that FDAF-PEM-AFROW is by construction related to the best linear unbiased estimate (BLUE) of the echo path. We depart from this framework to show an improvement in performance with respect to other adaptive filters minimizing the BLUE criterion, namely the PEM-AFROW and the FDAF-NLMS with near-end signal normalization. One of the contributions is to propose the instantaneous pseudo-correlation (IPC) measure between the near-end signal and the loudspeaker signal. The IPC measure serves as an indication of the effect of a DT situation occurring during adaptation. We motivate the choice of FDAF-PEM-AFROW over PEM-AFROW and FDAF-NLMS with near-end signal normalization, based on performance, computational complexity and related IPC measure values. Moreover, we use the FDAF-PEM-AFROW framework to improve several state-of-the-art variable step-size (VSS) and variable regularization (VR) algorithms. The FDAF-PEM-AFROW versions significantly outperform the original versions in every simulation. In terms of computational complexity, the FDAF-PEM-AFROW versions are themselves about two orders of magnitude cheaper than the original versions.
Jose Manuel Gil-Cacho, Toon van Waterschoot, Marc Moonen, Søren Holdt Jensen
IEEE ACM Trans. Audio Speech Lang. Process.2
2013 Improved prediction error filters for adaptive feedback cancellation in hearing aids
Kim Ngo, Toon van Waterschoot, Mads Græsbøll Christensen, Marc Moonen, Søren Holdt Jensen
Signal Process.2
2013 Declipping of Audio Signals Using Perceptual Compressed Sensing
abstract
The restoration of clipped audio signals, commonly known as declipping, is important to achieve an improved level of audio quality in many audio applications. In this paper, a novel declipping algorithm is presented, jointly based on the theory of compressed sensing (CS) and on well-established properties of human auditory perception. Declipping is formulated as a sparse signal recovery problem using the CS framework. By additionally exploiting knowledge of human auditory perception, a novel perceptual compressed sensing (PCS) framework is devised. A PCS-based declipping algorithm is proposed which uses ℓ1-norm type reconstruction. Comparative objective and subjective evaluation experiments reveal a significant audio quality increase for the proposed PCS-based declipping algorithm compared to CS-based declipping algorithms.
Bruno Defraene, Naim Mansour, Steven De Hertogh, Toon van Waterschoot, Moritz Diehl, Marc Moonen
IEEE ACM Trans. Audio Speech Lang. Process.4
2013 Nonlinear Acoustic Echo Cancellation Based on a Sliding-Window Leaky Kernel Affine Projection Algorithm
abstract
Acoustic echo cancellation (AEC) is used in speech communication systems where the existence of echoes degrades the speech intelligibility. Standard approaches to AEC rely on the assumption that the echo path to be identified can be modeled by a linear filter. However, some elements introduce nonlinear distortion and must be modeled as nonlinear systems. Several nonlinear models have been used with more or less success. The kernel affine projection algorithm (KAPA) has been successfully applied to many areas in signal processing but not yet to nonlinear AEC (NLAEC). The contribution of this paper is three-fold: (1) to apply KAPA to the NLAEC problem, (2) to develop a sliding-window leaky KAPA (SWL-KAPA) that is well suited for NLAEC applications, and (3) to propose a kernel function, consisting of a weighted sum of a linear and a Gaussian kernel. In our experiment set-up, the proposed SWL-KAPA for NLAEC consistently outperforms the linear APA, resulting in up to 12 dB of improvement in ERLE at a computational cost that is only 4.6 times higher. Moreover, it is shown that the SWL-KAPA outperforms, by 4-6 dB, a Volterra-based NLAEC, which itself has a much higher 413 times computational cost than the linear APA.
Jose Manuel Gil-Cacho, Marco Signoretto, Toon van Waterschoot, Marc Moonen, Søren Holdt Jensen
IEEE Trans. Speech Audio Process.3
2012 A psychoacoustically motivated speech distortion weighted multi-channel wiener filter for noise reduction
abstract
The aim of this paper is to improve the performance of existing speech distortion weighted multi-channel Wiener filter (SDW-MWFμ) based noise reduction (NR) algorithms. It is well known that for the SDW-MWFμthe improved NR performance comes at the cost of higher speech distortion when a fixed speech distortion weighting factor is used. In this paper we propose two psychoacoustically motivated weighting factor selection strategies, devised to exploit masking properties of the human ear. Experimental results based on PESQ scores, SNR improvement, and signal distortion confirm that both proposed psychoacoustically motivated weighting factor selection strategies do improve the NR performance compared to using a fixed weighting factor. In some of the analyzed scenarios, the fixed weighting factor approach is even seen to degrade the PESQ scores, while the psychoacoustically motivated approaches are seen to significantly improve the PESQ scores in all of the analyzed scenarios.
Bruno Defraene, Kim Ngo, Toon van Waterschoot, Moritz Diehl, Marc Moonen
ICASSP3
2012 Nonlinear acoustic echo cancellation based on a parallel-cascade kernel affine projection algorithm
abstract
In acoustic echo cancellation (AEC) applications, oftentimes an acoustic path from a loudspeaker to a microphone is estimated by means of a linear adaptive filter. However, loudspeakers introduce nonlinear distortions which may strongly degrade the adaptive filter performance, thus nonlinear filters have to be considered. This paper proposes two adaptive algorithms namely the parallel and cascade sliding-window kernel based affine projection algorithm (PSW-KAPA and CSW-KAPA) to solve the problem of nonlinear AEC (NLAEC) while keeping the computational complexity low. They are based on a leaky KAPA which employs the theory and algorithms of kernel methods. The basic concept is to perform adaptive filtering in a linear space that is nonlinearly related to the original input space. A kernel specifically designed for acoustic applications is proposed, which consists in a weighted sum of the linear and the Gaussian kernels. The motivation is basically to separate the problem into linear and nonlinear subproblems. The weights in the kernel also impose different forgetting mechanisms in the sliding window which in turn translates to a more flexible regularization. Simulation results show that PSW-KAPA and CSW-KAPA consistently outperform the linear NLMS, and generalize well both in high and low linear to nonlinear ratio (LNLR).
Jose Manuel Gil-Cacho, Toon van Waterschoot, Marc Moonen, Søren Holdt Jensen
ICASSP2
2012 Distributed estimation of static fields in wireless sensor networks using the finite element method
abstract
This paper deals with the distributed implementation of a recently proposed algorithm for the estimation of static fields. The algorithm combines wireless sensor network (WSN) field measurements with a physical field model in the form of a partial differential equation (PDE), such that the field can be estimated at locations different from the WSN sensor node locations. By discretizing the PDE using the finite element method (FEM), the physical field model reduces to a highly sparse linear system of equations. It is shown how this FEM-induced sparsity pattern can be exploited in the design of a distributed implementation such as to minimize the communication effort and data storage required in the WSN. Simulation results illustrate that a significant improvement in field estimation accuracy can be obtained, compared to the case when only WSN measurements (without a physical model) are used.
Toon van Waterschoot, Geert Leus
ICASSP1
2012 Real-Time Perception-Based Clipping of Audio Signals Using Convex Optimization
abstract
Clipping is an essential signal processing operation in many real-time audio applications, yet the use of existing clipping techniques generally has a detrimental effect on the perceived audio signal quality. In this paper, we present a novel multidisciplinary approach to clipping which aims to explicitly minimize the perceptible clipping-induced distortion by embedding a convex optimization criterion and a psychoacoustic model into a frame-based algorithm. The core of this perception-based clipping algorithm consists in solving a convex optimization problem for each time frame in a fast and reliable way. To this end, three different structure-exploiting optimization methods are derived in the common mathematical framework of convex optimization, and corresponding theoretical complexity bounds are provided. From comparative audio quality evaluation experiments, it is concluded that the perception-based clipping algorithm results in significantly higher objective audio quality scores than existing clipping techniques. Moreover, the algorithm is shown to be capable to adhere to real-time deadlines without making a sacrifice in terms of audio quality.
Bruno Defraene, Toon van Waterschoot, Hans Joachim Ferreau, Moritz Diehl, Marc Moonen
IEEE Trans. Speech Audio Process.2
2011 A fast projected gradient optimization method for real-time perception-based clipping of audio signals
abstract
Clipping is a necessary signal processing operation in many real time audio applications, yet it often reduces the sound quality of the signal. The recently proposed perception-based clipping algorithm has been shown to significantly outperform other clipping techniques in terms of objective sound quality scores. However, the real-time solution of the optimization problems that form the core of this algorithm, poses a challenge. In this paper, a fast gradient projection optimization method is proposed and incorporated into the perception based clipping algorithm. The optimization method will be shown to have an extremely low computational complexity per iteration, allowing the perception-based clipping algorithm to be applied in real-time for a broad range of clipping factors.
Bruno Defraene, Toon van Waterschoot, Moritz Diehl, Marc Moonen
ICASSP2
2011 Fifty Years of Acoustic Feedback Control: State of the Art and Future Challenges
abstract
The acoustic feedback problem has intrigued researchers over the past five decades, and a multitude of solutions has been proposed. In this survey paper, we aim to provide an overview of the state of the art in acoustic feedback control, to report results of a comparative evaluation with a selection of existing methods, and to cast a glance at the challenges for future research.
Toon van Waterschoot, Marc Moonen
Proc. IEEE1
2010 Adaptive feedback cancellation in hearing aids using a sinusoidal near-end signal model
abstract
Acoustic feedback is a well-known problem in hearing aids, which is caused by the undesired acoustic coupling between the loudspeaker and the microphone. Acoustic feedback limits the maximum amplification that can be used in the hearing aid without making it unstable. The goal of adaptive feedback cancellation (AFC) is to adaptively model the feedback path and estimate the feedback signal, which is then subtracted from the microphone signal. The main problem in identifying the feedback path model is the correlation between the near-end signal and the loudspeaker signal, which is caused by the closed signal loop. A possible solution to this problem is to use the prediction error method (PEM)-based AFC with a linear prediction (LP) model for the near-end signal. In this paper, a modification to the PEM-based AFC is presented where the LP model is replaced by a sinusoidal near-end signal model. More specifically, it is shown that using frequency estimation techniques to estimate the sinusoidal near-end signal model improves the performance of the PEM-based AFC compared to using a LP model. Simulation results for a hearing aid scenario indicate a significant improvement in terms of misadjustment and maximum stable gain increase.
Kim Ngo, Toon van Waterschoot, Mads Græsbøll Christensen, Marc Moonen, Søren Holdt Jensen, Jan Wouters
ICASSP2
2010 Analytical Expressions for the Power Spectral Density of CP-OFDM and ZP-OFDM Signals
abstract
In this letter, analytical expressions are derived for the power spectral density (PSD) of orthogonal frequency division multiplex (OFDM) signals employing a cyclic prefix (CP-OFDM) or zero padding (ZP-OFDM) time guard interval. Under the relatively weak assumptions that (i) the data are independent and identically distributed on all OFDM subcarriers and (ii) the OFDM pulse shape is sufficiently localized in time, simple closed-form PSD expressions can be obtained. These expressions are then compared to existing OFDM PSD expressions and validated by inspecting the power spectra of some standardized OFDM signals.
Toon van Waterschoot, Vincent Le Nir, Jonathan Duplicy, Marc Moonen
IEEE Signal Process. Lett.1
2009 Adaptive feedback cancellation for audio applications
Toon van Waterschoot, Marc Moonen
Signal Process.1
2008 Adaptive feedback cancellation for audio signals using a warped all-pole near-end signal model
abstract
Sound amplification systems having a closed signal loop often suffer from acoustic feedback, which limits the achievable amount of amplification and severely affects sound quality. A promising solution to the feedback problem consists in predicting the feedback signal using an adaptive filter, however, a bias is then introduced due to signal correlation. In speech applications, a prediction-error-method- based approach to adaptive feedback cancellation has proven to be capable of providing sufficient decorrelation without sacrificing speech quality. This approach, which is based on estimating an all-pole near-end signal model, appears to be unappropriate for musical audio signals because of their large degree of tonality. We propose a novel prediction-error-method-based adaptive feedback cancellation algorithm that features a frequency-warped all-pole near-end signal model, which is better suited for tonal audio signals. Simulation results show a doubling of the convergence speed, with only a relatively small increase in computational complexity.
Toon van Waterschoot, Marc Moonen
ICASSP1
2008 Optimally regularized adaptive filtering algorithms for room acoustic signal enhancement
Toon van Waterschoot, Geert Rombouts, Marc Moonen
Signal Process.1
2007 Linear prediction of audio signals
Toon van Waterschoot, Marc Moonen
INTERSPEECH1
2007 A Pole-Zero Placement Technique for Designing Second-Order IIR Parametric Equalizer Filters
abstract
A new procedure is presented for designing second-order parametric equalizer filters. In contrast to the traditional approach, in which the design is based on a bilinear transform of an analog filter, the presented procedure allows for designing the filter directly in the digital domain. A rather intuitive technique known as pole-zero placement, is treated here in a quantitative way. It is shown that by making some meaningful approximations, a set of relatively simple design equations can be obtained. Design examples of both notch and resonance filters are included to illustrate the performance of the proposed method and to compare with state-of-the-art solutions.
Toon van Waterschoot, Marc Moonen
IEEE Trans. Speech Audio Process.1