Oliver Thiergart

dblp:70/9877 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
6since 2021 · last 2025
0000-0001-5681-882XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 2 first-author
YearPublicationVenuePosition
2025 A first-order DirAC-based parametric Ambisonic coder for immersive communications
abstract
Directional Audio Coding (DirAC) is a proven method for parametrically representing a 3D audio scene in B-format and is capable of reproducing it on arbitrary loudspeaker layouts. Although such a method seems well suited for low bitrate Ambisonic transmission, little work has been done on the feasibility of building a real system upon it. In this paper, we present a DirAC-based coding for Higher-Order Ambisonics (HOA), developed as part of a standardisation effort to extend the 3GPP EVS codec to immersive communications. Starting from the first-order DirAC model, we show how to reduce algorithmic delay, the bitrate required for the parameters and complexity by bringing the full synthesis in the spherical harmonic domain. The evaluation of the proposed technique for coding 3rdorder Ambisonics at bitrates from 32 to 128 kbps shows the relevance of the parametric approach compared with existing solutions.
Guillaume Fuchs, Florin Ghido, Dominik Weckbecker, Oliver Thiergart
ICASSP4
2025 Low-Complexity Neural Speech Dereverberation With Adaptive Target Control
abstract
Existing neural network-based speech dereverberation approaches use a fixed-length early reflection part of the reverberant signal as the target for estimation, irrespective of the severity of reverberation. Such an approach often leads to distortions in the enhanced signals in highly reverberant scenarios. In practice, while some listeners prefer minimal speech distortions, others have a higher tolerance for distortions and prefer a clean signal. To address these points, we propose a novel target definition and a low-complexity neural network for user-controlled single-channel dereverberation. Our target definition is parameterized by the relative amount of reduction in the late reverberation energy. Further, the same parameter is passed as a control input to the dereverberation network for adaptability during inference. Objective and subjective evaluation shows the feasibility of the proposed dereverberation approach.
Nagashree K. S. Rao, Srikanth Raj Chetupalli, Shrishti Saha Shetu, Emanuël A. P. Habets, Oliver Thiergart
ICASSP5
2024 Binaural Rendering of Heterogeneous Sound Sources with Extent
abstract
In spatial audio rendering applications, it is often desired to render sound sources with a certain spatial extent in a realistic way. While existing methods mainly consider rendering of homogeneously extended sound sources (i.e., with constant radiation characteristics over the extent), rendering of heterogeneously extended sound sources (i.e., with position-dependent radiation characteristics) has barely been discussed in the literature. In this paper, we propose an approach for binaural rendering of heterogeneously extended sound sources. Input to the algorithm is a two-channel signal, which provides information about the position-dependent radiation characteristics of the sound source. Based on a model for an extended sound source with position-dependent energy and spectral content, the target covariance matrix of the binaural output signal is determined. Using a previously proposed optimal mixing approach, a binaural output signal with the desired properties is obtained, ensuring that the spatial characteristics encoded in the two-channel input signal are preserved. The proposed approach is evaluated both objectively and subjectively by comparing it to two homogeneous extent-rendering baselines as well as to simple point source reproduction.
Carlotta Anemüller, Oliver Thiergart, Emanuël A. P. Habets
ICASSP2
2024 Sector-Based Interference Cancellation for Robust Keyword Spotting Applications Using an Informed MPDR Beamformer
abstract
A low-complexity, sector-based interference cancellation approach is proposed for voice-controlled devices, e.g., smart speakers. We propose an informed minimum power distortionless response beamformer that provides an optimal trade-off between noise reduction, dereverberation, and interference cancellation, with a minimal amount of target speaker distortions. Low complexity is achieved by using information on the target speaker in both the linear constraint and the beamformer’s minimization term, which allows sharing of the most complex beamformer computations across the different sectors. The results show that the proposed approach significantly improves keyword spotter performance compared to other approaches, such as the delay-and-sum beamformer, linearly constrained minimum variance beamformer, and traditional minimum power distortionless response beamformer.
Guendalina Milano, Oliver Thiergart, Emanuël A. P. Habets
ICASSP2
2024 Ultra Low Complexity Deep Learning Based Noise Suppression
abstract
This paper introduces an innovative method for reducing the computational complexity of deep neural networks in real-time speech enhancement on resource-constrained devices. The proposed approach utilizes a two-stage processing framework, employing channelwise feature reorientation to reduce the computational load of convolutional operations. By combining this with a modified power law compression technique for enhanced perceptual quality, this approach achieves noise suppression performance comparable to state-of-the-art methods with significantly less computational requirements. Notably, our algorithm exhibits 3 to 4 times less computational complexity and memory usage than prior state-of-the-art approaches.
Shrishti Saha Shetu, Soumitro Chakrabarty, Oliver Thiergart, Edwin Mabande
ICASSP3
2022 A Data-Driven Approach to Audio Decorrelation
abstract
The degree of correlation between two audio signals entering the ears is known to have a significant impact on the spatial perception of a sound image. Audio signal decorrelation is therefore a widely used tool in various applications within the field of spatial audio processing. This paper explores for the first time the use of a data-driven approach for audio decorrelation. We propose a convolutional neural network architecture that is trained with the help of a state-of-the-art reference decorrelator. The proposed approach is evaluated using music and applause signals by means of objective evaluations as well as through a listening test. The proposed approach can serve as a proof of concept to address common limitations of existing decorrelation techniques in future work, which include introduction of temporal smearing and coloration artifacts and the production of a limited number of mutually uncorrelated output signals.
Carlotta Anemüller, Oliver Thiergart, Emanuël A. P. Habets
IEEE Signal Process. Lett.2
2019 Combining Linear Spatial Filtering and Non-linear Parametric Processing for High-quality Spatial Sound Capturing
abstract
Flexible spatial sound capturing and reproduction can be achieved with multiple microphones by using linear spatial filtering or non-linear parametric processing. The non-linear approaches usually provide a superior spatial resolution compared to the linear approaches but can result in artifacts due to violations of the sound field model. In this paper, we combine both approaches to achieve a high robustness against model violations and a high spatial resolution. We assume linear spatial filters that approximate the spatial responses of the desired output format and compensate remaining deviations with an optimal post filter. The post filter is computed such that the proposed approach behaves like a linear system when the spatial filters achieve the desired spatial response, and scales towards a non-linear system otherwise. Experimental results show that the proposed approach can significantly reduce distortions of existing parametric processing schemes especially when a sufficiently high number of microphones is available.
Oliver Thiergart, Guendalina Milano, Emanuël A. P. Habets
ICASSP1
2018 Evaluation and Comparison of Late Reverberation Power Spectral Density Estimators
abstract
Reduction of late reverberation can be achieved using spatio-spectral filters, such as the multichannel Wiener filter. To compute this filter, an estimate of the late reverberation power spectral density (PSD) is required. In recent years, a multitude of late reverberation PSD estimators have been proposed. In this paper, these estimators are categorized into several classes, their relations and differences are discussed, and a comprehensive experimental comparison is provided. To compare their performance, simulations in controlled as well as practical scenarios are conducted. It is shown that a common weakness of spatial coherence-based estimators is their performance in high direct-to-diffuse ratio conditions. To mitigate this problem, a correction method is proposed and evaluated. It is shown that the proposed correction method can decrease the speech distortion without significantly affecting the reverberation reduction.
Sebastian Braun, Adam Kuklasinski, Ofer Schwartz, Oliver Thiergart, Emanuël A. P. Habets, Sharon Gannot, Simon Doclo, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2016 A low complexity weighted least squares narrowband DOA estimator for arbitrary array geometries
abstract
An increasing number of spatial filtering approaches requires narrowband direction-of-arrival (DOA) estimates. State-of-the-art (SOA) estimators such as root-MUSIC and ESPRIT are computationally complex and can be used only with specific array geometries. In this work, a low complexity DOA estimator is proposed that can be applied to arbitrary array geometries. The DOA is estimated by minimizing the weighted error between the observed and expected inter-microphone phase differences. The complexity of the proposed DOA estimator is significantly lower compared to that of the SOA estimators while providing a similar performance.
Oliver Thiergart, Weilong Huang, Emanuël A. P. Habets
ICASSP1
2015 A Bayesian approach to spatial filtering and diffuse power estimation for joint dereverberation and noise reduction
abstract
A spatial filter, with L linear constraints that are based on instantaneous narrowband direction-of-arrival (DOA) estimates, was recently proposed to obtain a desired spatial response for at most L sound sources. In noisy and reverberant environments, it becomes difficult to get reliable instantaneous DOA estimates and hence obtain the desired spatial response. In this work, we develop a Bayesian approach to spatial filtering that is more robust to DOA estimation errors. The resulting filter is a weighted sum of spatial filters pointed at a discrete set of DOAs, with the relative contribution of each filter determined by the posterior distribution of the discrete DOAs given the microphone signals. In addition, the proposed spatial filter is able to reduce both reverberation and noise. In this work, the required diffuse sound power is estimated using the posterior distribution of the discrete set of DOAs. Simulation results demonstrate the ability of the proposed filter to achieve strong suppression of the undesired signal components with small amount of signal distortion, in noisy and reverberant conditions.
Soumitro Chakrabarty, Oliver Thiergart, Emanuël A. P. Habets
ICASSP2
2014 Automatic spatial gain control for an informed spatial filter
abstract
When capturing speech in a multi-talker telecommunication scenario, it is desirable to keep the enhanced signal at an equal loudness level for each speaker. Single-channel automatic gain control systems are not able to adjust the level of different talkers when they are simultaneously active. In this work, an automatic spatial gain control (ASGC) algorithm is proposed that adjusts the directional response of an existing informed spatial filter such that the direct sound of multiple sources can be kept at a constant desired loudness level at the output. The spatial filter additionally reduces diffuse sound and ambient noise. It is shown that the proposed AGSC works well within the tested scenario, and is able to adjust the levels of different speakers even during double talk scenarios.
Sebastian Braun, Oliver Thiergart, Emanuël A. P. Habets
ICASSP2
2014 Power-based signal-to-diffuse ratio estimation using noisy directional microphones
abstract
The signal-to-diffuse ratio (SDR), which describes the power ratio between the direct and diffuse component of a sound field, is an important parameter in many applications. This paper proposes a power-based SDR estimator which considers the auto power spectral densities obtained by noisy directional microphones. Compared to recently proposed estimators that exploit the spatial coherence between two microphones, the power-based estimator is more robust at lower frequencies given that the microphone directivities are known with sufficiently high accuracy. The proposed estimator can incorporate more than two microphones and can therefore provide accurate SDR estimates independently of the direction-of-arrival of the direct sound. We further propose a method to determine the optimal microphone orientations for a given set of directional microphones. Simulations show the practical applicability.
Oliver Thiergart, Tobias Ascherl, Emanuël A. P. Habets
ICASSP1
2014 Extracting Reverberant Sound Using a Linearly Constrained Minimum Variance Spatial Filter
abstract
Microphone arrays are typically used to extract the direct sound of sound sources while suppressing noise and reverberation. Applications such as immersive spatial sound reproduction commonly also require an estimate of the reverberant sound. A linearly constrained minimum variance filter, of which one of the constraints is related to the spatial coherence of the assumed reverberant sound field, is proposed to obtain an estimate of the sound pressure of the reverberant field. The proposed spatial filter provides an almost omnidirectional directivity pattern with spatial nulls for the directions-of-arrival of the direct sound. The filter is computationally efficient and outperforms existing methods.
Oliver Thiergart, Emanuël A. P. Habets
IEEE Signal Process. Lett.1
2014 An informed parametric spatial filter based on instantaneous direction-of-arrival estimates
abstract
Extracting desired source signals in noisy and reverberant environments is required in many hands-free communication systems. In practical situations, where the position and number of active sources may be unknown and time-varying, conventional implementations of spatial filters do not provide sufficiently good performance. Recently, informed spatial filters have been introduced that incorporate almost instantaneous parametric information on the sound field, thereby enabling adaptation to new acoustic conditions and moving sources. In this contribution, we propose a spatial filter which generalizes the recently proposed informed linearly constrained minimum variance filter and informed minimum mean square error filter. The proposed filter uses multiple direction-of-arrival estimates and second-order statistics of the noise and diffuse sound. To determine those statistics, an optimal diffuse power estimator is proposed that outperforms state-of-the-art estimators. Extensive performance evaluation demonstrates the effectiveness of the proposed filter in dynamic acoustic conditions. For this purpose, we have considered a challenging scenario which consists of quickly moving sound sources during double-talk. The performance of the proposed spatial filter was evaluated in terms of objective measures including segmental signal-to-reverberation ratio and log spectral distance, and by means of a listening test confirming the objective results.
Oliver Thiergart, Maja Taseska, Emanuël A. P. Habets
IEEE ACM Trans. Audio Speech Lang. Process.1
2013 An informed LCMV filter based on multiple instantaneous direction-of-arrival estimates
abstract
Extracting sound sources in noisy and reverberant conditions remains a challenging task that is commonly found in modern communication systems. In this work, we consider the problem of obtaining a desired spatial response for at most L simultaneously active sound sources. The proposed spatial filter is obtained by minimizing the diffuse plus self-noise power at the output of the filter subject to L linear constraints. In contrast to earlier works, the L constraints are based on instantaneous narrowband direction-of-arrival estimates. In addition, a novel estimator for the diffuse-to-noise ratio is developed that exhibits a sufficiently high temporal and spectral resolution to achieve both dereverberation and noise reduction. The presented results demonstrate that an optimal tradeoff between maximum white noise gain and maximum directivity is achieved.
Oliver Thiergart, Emanuël A. P. Habets
ICASSP1
2013 Geometry-Based Spatial Sound Acquisition Using Distributed Microphone Arrays
abstract
Traditional spatial sound acquisition aims at capturing a sound field with multiple microphones such that at the reproduction side a listener can perceive the sound image as it was at the recording location. Standard techniques for spatial sound acquisition usually use spaced omnidirectional microphones or coincident directional microphones. Alternatively, microphone arrays and spatial filters can be used to capture the sound field. From a geometric point of view, the perspective of the sound field is fixed when using such techniques. In this paper, a geometry-based spatial sound acquisition technique is proposed to compute virtual microphone signals that manifest a different perspective of the sound field. The proposed technique uses a parametric sound field model that is formulated in the time-frequency domain. It is assumed that each time-frequency instant of a microphone signal can be decomposed into one direct and one diffuse sound component. It is further assumed that the direct component is the response of a single isotropic point-like source (IPLS) of which the position is estimated for each time-frequency instant using distributed microphone arrays. Given the sound components and the position of the IPLS, it is possible to synthesize a signal that corresponds to a virtual microphone at an arbitrary position and with an arbitrary pick-up pattern.
Oliver Thiergart, Giovanni Del Galdo, Maja Taseska, Emanuël A. P. Habets
IEEE ACM Trans. Audio Speech Lang. Process.1
2012 Signal-to-reverberant ratio estimation based on the complex spatial coherence between omnidirectional microphones
abstract
The signal-to-reverberant ratio (SRR) is an important parameter in several applications such as speech enhancement, dereverberation, and parametric spatial audio coding. In this contribution, an SRR estimator is derived from the direction-of-arrival dependent complex spatial coherence function computed via two omnidirectional microphones. It is shown that by employing a computationally inexpensive DOA estimator, the proposed SRR estimator outperforms existing approaches.
Oliver Thiergart, Giovanni Del Galdo, Emanuël A. P. Habets
ICASSP1
2011 Resolving spatial sampling effects in parametric directional filtering
abstract
Directional Audio Coding (DirAC) represents an efficient scheme to analyze and reproduce spatial sound; the coded stream consists of a single-channel audio signal and few parameters. Estimation of these parameters ideally relies on figure-of-eight microphones and one omnidirectional microphone. The figure-of-eight directivity can be efficiently approximated by first-order differential microphone arrays. However, due to spatial sampling, there are considerable deviations from the required directivity at high frequencies. These deviations lead to incorrect spatial parameter estimates; especially, the instantaneous direction-of-arrival (DOA) becomes biased and - beyond a certain aliasing frequency - ambiguous. Parameters like the DOA drive subsequent processing units, e. g., for spatial filtering. In this paper we propose a novel strategy to compensate bias and ambiguity not directly w. r. t. the DOA but the output parameters of a spatial filtering processing unit. Simulation results confirm the benefits of the novel method, which does not involve additional computational load at run-time.
Markus Kallinger, Michael Buerger, Oliver Thiergart, Fabian Kuech, Dirk Mahne
ICASSP3