EDBT 2026 Demo / reviewers in the wild / expert
Emanuël A. P. Habets
dblp:42/16 · also Emanuël Anco Peter Habets
· DBLP profile ↗
142ranked-venue papers
13as first author
36since 2021 · last 2025
0000-0002-2613-8046ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 102 · 7 first-author · 32 since 2021Artificial intelligence and machine learning · 48 · 6 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 8 · 6 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AdaBit-TasNet: Speech Separation with Inference Adaptable PrecisionabstractDeploying advanced neural network-based speech separation (SS) models on resource-constrained devices is challenging due to their high computational and memory demands. Conventional network compression techniques, such as pruning and quantization, can alleviate these demands without significantly compromising performance. However, they lack the flexibility to select the compression factor at run-time to suit varying operating conditions, such as changing computational and energy budgets in battery-powered devices. In this paper, we introduce AdaBit-TasNet, an adaptable-precision network (APN) for SS that enables flexible bit-width selection during inference. Experimental evaluation on the Libri2Mix dataset demonstrates that AdaBit-TasNet achieves comparable performance to that of individually trained fixed-precision networks at several bitwidths. Mohamed Elminshawi, Srikanth Raj Chetupalli, Emanuël A. P. Habets |
ASRU | 3 |
| 2025 | On the Relation Between Speech Quality and Quantized Latent Representations of Neural CodecsabstractNeural audio signal codecs have attracted significant attention in recent years. In essence, the impressive low bitrate achieved by such encoders is enabled by learning an abstract representation that captures the properties of encoded signals, e.g., speech. In this work, we investigate the relation between the latent representation of the input signal learned by a neural codec and the quality of speech signals. To do so, we introduce Latent-representation-to-Quantization error Ratio (LQR) measures, which quantify the distance from the idealized neural codec’s speech signal model for a given speech signal. We compare the proposed metrics to intrusive measures as well as data-driven supervised methods using two subjective speech quality datasets. This analysis shows that the proposed LQR correlates strongly (up to 0.9 Pearson’s correlation) with the subjective quality of speech. Despite being a non-intrusive metric, this yields a competitive performance with, or even better than, other pre-trained and intrusive measures. These results show that LQR is a promising basis for more sophisticated speech quality measures. Mhd Modar Halimeh, Matteo Torcoli, Philipp Grundhuber, Emanuël A. P. Habets |
ICASSP | 4 |
| 2025 | Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotronabstractIn recent years, several text-to-speech systems have been proposed to synthesize natural speech in zero-shot, few-shot, and low-resource scenarios. However, these methods typically require training with data from many different speakers. The speech quality across the speaker set typically is diverse and imposes an upper limit on the quality achievable for the low-resource speaker. In the current work, we achieve high-quality speech synthesis using as little as five minutes of speech from the desired speaker by augmenting the low-resource speaker data with noise and employing multiple sampling techniques during training. Our method requires only four high-quality, high-resource speakers, which are easy to obtain and use in practice. Our low-complexity method achieves improved speaker similarity compared to the state-of-the-art zero-shot method HierSpeech++ and the recent low-resource method AdapterMix while maintaining comparable naturalness. Our proposed approach can also reduce the data requirements for speech synthesis for new speakers and languages. Kishor Kayyar Lakshminarayana, Frank Zalkow, Christian Dittmar, Nicola Pia, Emanuël A. P. Habets |
ICASSP | 5 |
| 2025 | Low-Complexity Neural Speech Dereverberation With Adaptive Target ControlabstractExisting neural network-based speech dereverberation approaches use a fixed-length early reflection part of the reverberant signal as the target for estimation, irrespective of the severity of reverberation. Such an approach often leads to distortions in the enhanced signals in highly reverberant scenarios. In practice, while some listeners prefer minimal speech distortions, others have a higher tolerance for distortions and prefer a clean signal. To address these points, we propose a novel target definition and a low-complexity neural network for user-controlled single-channel dereverberation. Our target definition is parameterized by the relative amount of reduction in the late reverberation energy. Further, the same parameter is passed as a control input to the dereverberation network for adaptability during inference. Objective and subjective evaluation shows the feasibility of the proposed dereverberation approach. Nagashree K. S. Rao, Srikanth Raj Chetupalli, Shrishti Saha Shetu, Emanuël A. P. Habets, Oliver Thiergart |
ICASSP | 4 |
| 2025 | GAN-Based Speech Enhancement for Low SNR Using Latent Feature ConditioningabstractEnhancing speech quality under adverse SNR conditions remains a significant challenge for discriminative deep neural network (DNN)-based approaches. In this work, we propose DisCoGAN, which is a time-frequency-domain generative adversarial network (GAN) conditioned by the latent features of a discriminative model pre-trained for speech enhancement in low SNR scenarios. Our proposed method achieves superior performance compared to state-of-the-art discriminative methods and also surpasses end-to-end (E2E) trained GAN models. We also investigate the impact of various configurations for conditioning the proposed GAN model with the discriminative model and assess their influence on enhancing speech quality. Shrishti Saha Shetu, Emanuël A. P. Habets, Andreas Brendel |
ICASSP | 2 |
| 2025 | You Are What You Say: Exploiting Linguistic Content for VoicePrivacy Attacksabstract4238 Ünal Ege Gaznepoglu, Anna Leschanowsky, Ahmad Aloradi, Prachi Singh, Daniel Tenbrinck, Emanuël A. P. Habets, Nils Peters |
INTERSPEECH | 6 |
| 2025 | Benchmarking Neural Speech Codec Intelligibility with SIToolabstract5488 Anna Leschanowsky, Kishor Kayyar Lakshminarayana, Anjana Rajasekhar, Lyonel Behringer, Ibrahim Kilinc, Guillaume Fuchs, Emanuël A. P. Habets |
INTERSPEECH | 7 |
| 2025 | Bridging the Training-Inference Gap in TTS: Training Strategies for Robust Generative Postprocessing for Low-Resource Speakersabstract2470 Frank Zalkow, Paolo Sani, Kishor Kayyar Lakshminarayana, Emanuël A. P. Habets, Nicola Pia, Christian Dittmar |
INTERSPEECH | 4 |
| 2024 | Binaural Rendering of Heterogeneous Sound Sources with ExtentabstractIn spatial audio rendering applications, it is often desired to render sound sources with a certain spatial extent in a realistic way. While existing methods mainly consider rendering of homogeneously extended sound sources (i.e., with constant radiation characteristics over the extent), rendering of heterogeneously extended sound sources (i.e., with position-dependent radiation characteristics) has barely been discussed in the literature. In this paper, we propose an approach for binaural rendering of heterogeneously extended sound sources. Input to the algorithm is a two-channel signal, which provides information about the position-dependent radiation characteristics of the sound source. Based on a model for an extended sound source with position-dependent energy and spectral content, the target covariance matrix of the binaural output signal is determined. Using a previously proposed optimal mixing approach, a binaural output signal with the desired properties is obtained, ensuring that the spatial characteristics encoded in the two-channel input signal are preserved. The proposed approach is evaluated both objectively and subjectively by comparing it to two homogeneous extent-rendering baselines as well as to simple point source reproduction. Carlotta Anemüller, Oliver Thiergart, Emanuël A. P. Habets |
ICASSP | 3 |
| 2024 | Sector-Based Interference Cancellation for Robust Keyword Spotting Applications Using an Informed MPDR BeamformerabstractA low-complexity, sector-based interference cancellation approach is proposed for voice-controlled devices, e.g., smart speakers. We propose an informed minimum power distortionless response beamformer that provides an optimal trade-off between noise reduction, dereverberation, and interference cancellation, with a minimal amount of target speaker distortions. Low complexity is achieved by using information on the target speaker in both the linear constraint and the beamformer’s minimization term, which allows sharing of the most complex beamformer computations across the different sectors. The results show that the proposed approach significantly improves keyword spotter performance compared to other approaches, such as the delay-and-sum beamformer, linearly constrained minimum variance beamformer, and traditional minimum power distortionless response beamformer. Guendalina Milano, Oliver Thiergart, Emanuël A. P. Habets |
ICASSP | 3 |
| 2024 | Odaq: Open Dataset of Audio QualityabstractResearch into the prediction and analysis of perceived audio quality is hampered by the scarcity of openly available datasets of audio signals accompanied by corresponding subjective quality scores. To address this problem, we present the Open Dataset of Audio Quality (ODAQ), a new dataset containing the results of a MUSHRA listening test conducted with expert listeners from 2 international laboratories. ODAQ contains 240 audio samples and corresponding quality scores. Each audio sample is rated by 26 listeners. The audio samples are stereo audio signals sampled at 44.1 or 48 kHz and are processed by a total of 6 method classes, each operating at different quality levels. The processing method classes are designed to generate quality degradations possibly encountered during audio coding and source separation, and the quality levels for each method class span the entire quality range. The diversity of the processing methods, the large span of quality levels, the high sampling frequency, and the pool of international listeners make ODAQ particularly suited for further research into subjective and objective audio quality. The dataset is released with permissive licenses, and the software used to conduct the listening test is also made publicly available. Matteo Torcoli, Chih-Wei Wu, Sascha Dick, Phillip A. Williams, Mhd Modar Halimeh, William Wolcott, Emanuël A. P. Habets |
ICASSP | 7 |
| 2024 | Meta Learning Text-to-Speech Synthesis in over 7000 Languagesabstract4958 Florian Lux, Sarina Meyer, Lyonel Behringer, Frank Zalkow, Phat Do, Matt Coler, Emanuël A. P. Habets, Ngoc Thang Vu |
INTERSPEECH | 7 |
| 2024 | Dynamic Slimmable Network for Speech SeparationabstractNeural networks for speech separation generally exhibit high computational costs and large memory footprints. Moreover, typical separation networks have a fixed computational graph that processes all input frames at a uniform computational cost, even though intensive processing may not be necessary for frames containing silence or a single active speaker. Addressing this computational inefficiency is especially crucial when these networks are deployed on resource-constrained devices. In this letter, we propose a dynamic slimmable network for speech separation that mitigates the computational inefficiency of existing networks. We introduce slimmable layers with a gating mechanism that can adapt their computational complexity based on the input characteristics. As an example, we propose to use the slimmable layers in the intra-chunk blocks of a dual-path structure-based network to facilitate adaptation based on the local characteristics of the input signal. Experimental evaluation on simulated two-speaker mixtures from the WSJ0-2mix dataset demonstrates that the proposed method substantially reduces the computational cost while maintaining comparable performance to fully utilized static networks. Mohamed Elminshawi, Srikanth Raj Chetupalli, Emanuël A. P. Habets |
IEEE Signal Process. Lett. | 3 |
| 2023 | Beamformer-Guided Target Speaker ExtractionabstractWe propose a Beamformer-guided Target Speaker Extraction (BG-TSE) method to extract a target speaker’s voice from a multi-channel recording informed by the direction of arrival of the target. The proposed method employs a front-end beamformer steered towards the target speaker to pro-vide an auxiliary signal to a single-channel TSE system. By allowing for time-varying embeddings in the single-channel TSE block, the proposed method fully exploits the correspondence between the front-end beamformer output and the tar-get speech in the microphone signal. Experimental evaluation on simulated multi-channel 2-speaker mixtures, in both anechoic and reverberant conditions, demonstrates the advantage of the proposed method compared to recent single-channel and multi-channel baselines. Mohamed Elminshawi, Srikanth Raj Chetupalli, Emanuël A. P. Habets |
ICASSP | 3 |
| 2023 | Contrastive Representation Learning for Acoustic Parameter EstimationabstractA study is presented in which a contrastive learning approach is used to extract low-dimensional representations of the acoustic environment from single-channel, reverberant speech signals. Convolution of room impulse responses (RIRs) with anechoic source signals is leveraged as a data augmentation technique that offers considerable flexibility in the design of the upstream task. We evaluate the embeddings across three different downstream tasks, which include the regression of acoustic parameters reverberation time RT60and clarity index C50, and the classification into small and large rooms. We demonstrate that the learned representations generalize well to unseen data and perform similarly to a fully-supervised baseline. Philipp Götz, Cagdas Tuna, Andreas Walther 0001, Emanuël A. P. Habets |
ICASSP | 4 |
| 2023 | Better Together: Dialogue Separation and Voice Activity Detection for Audio Personalization in TVabstractIn TV services, dialogue level personalization is key to meeting user preferences and needs. When dialogue and background sounds are not separately available from the production stage, Dialogue Separation (DS) can estimate them to enable personalization. DS was shown to provide clear benefits for the end user. Still, the estimated signals are not perfect, and some leakage can be introduced. This is undesired, especially during passages without dialogue. We propose to combine DS and Voice Activity Detection (VAD), both recently proposed for TV audio. When their combination suggests dialogue inactivity, background components leaking in the dialogue estimate are reassigned to the background estimate. A clear improvement of the audio quality is shown for dialogue-free signals, without performance drops when dialogue is active. A post-processed VAD estimate with improved detection accuracy is also generated. It is concluded that DS and VAD can improve each other and are better used together. Matteo Torcoli, Emanuël A. P. Habets |
ICASSP | 2 |
| 2023 | Multi-Microphone Speaker Separation by Spatial RegionsabstractWe consider the task of region-based source separation of reverberant multi-microphone recordings. We assume pre-defined spatial regions with a single active source per region. The objective is to estimate the signals from the individual spatial regions as captured by a reference microphone while retaining a correspondence between signals and spatial regions. We propose a data-driven approach using a modified version of a state-of-the-art network, where different layers model spatial and spectro-temporal information. The network is trained to enforce a fixed mapping of regions to network outputs. Using speech from LibriMix, we construct a data set specifically designed to contain the region information. Additionally, we train the network with permutation invariant training. We show that both training methods result in a fixed mapping of regions to network outputs, achieve comparable performance, and that the networks exploit spatial information. The proposed network outperforms a baseline network by 1.5 dB in scale-invariant signal-to-distortion ratio. Julian Wechsler, Srikanth Raj Chetupalli, Wolfgang Mack, Emanuël A. P. Habets |
ICASSP | 4 |
| 2023 | Evaluating Speech-Phoneme Alignment and its Impact on Neural Text-To-Speech SynthesisabstractIn recent years, the quality of text-to-speech (TTS) synthesis vastly improved due to deep-learning techniques, with parallel architectures, in particular, providing excellent synthesis quality at fast inference. Training these models usually requires speech recordings, corresponding phoneme-level transcripts, and the temporal alignment of each phoneme to the utterances. Since manually creating such fine-grained alignments requires expert knowledge and is time-consuming, it is common practice to estimate them using automatic speech–phoneme alignment methods. In the literature, either the estimation methods’ accuracy or their impact on the TTS system’s synthesis quality is evaluated. In this study, we perform experiments with five state-of-the-art speech–phoneme aligners and evaluate their output with objective and subjective measures. As our main result, we show that small alignment errors (below 75 ms error) do not decrease the synthesis quality, which implies that the alignment error may not be the crucial factor when choosing an aligner for TTS training. Frank Zalkow, Prachi Govalkar, Meinard Müller, Emanuël A. P. Habets, Christian Dittmar |
ICASSP | 4 |
| 2023 | Predicting Preferred Dialogue-to-Background Loudness Difference in Dialogue-Separated AudioabstractDialogue Enhancement (DE) enables the rebalancing of dialogue and background sounds to fit personal preferences and needs in the context of broadcast audio. When individual audio stems are unavailable from production, Dialogue Separation (DS) can be applied to the final audio mixture to obtain esti-mates of these stems. This work focuses on Preferred Loudness Differences (PLDs) between dialogue and background sounds. While previous studies determined the PLD through a listening test employing original stems from production, stems estimated by DS are used in the present study. In addition, a larger variety of signal classes is considered. PLDs vary substantially across individuals (average interquartile range: 5.7 LU). Despite this variability, PLDs are found to be highly dependent on the signal type under consideration, and it is shown that median PLDs can be predicted using objective intelligibility metrics. Two existing baseline prediction methods - intended for use with original stems - displayed a Mean Absolute Error (MAE) of 7.5 LU and 5 LU, respectively. A modified baseline (MAE: 3.2 LU) and an alternative approach (MAE: 2.5 LU) are proposed. Results support the viability of processing final broadcast mixtures with DS and offering an alternative remixing that accounts for median PLDs. Luca Resti, Martin Strauss 0003, Matteo Torcoli, Emanuël A. P. Habets, Bernd Edler |
QoMEX | 4 |
| 2023 | Saliency of Omnidirectional Videos with Different Audio Presentations: Analyses and DatasetabstractThere is an increased interest in understanding users' behavior when exploring omnidirectional (360°) videos, especially in the presence of spatial audio. Several studies demonstrate the effect of no, mono, or spatial audio on visual saliency. However, no studies investigate the influence of higher-order (i.e., 4t h- order) Ambisonics on subjective exploration in virtual reality settings. In this work, a between-subjects test design is employed to collect users' exploration data of 360° videos in a free-form viewing scenario using the Varjo XR-3 Head Mounted Display, in the presence of no, mono, and 4th-order Ambisonics audio. Saliency information was captured as head-saliency in terms of the center of a viewport at 50 Hz. For each item, subjects were asked to describe the scene with a short free-verbalization task. Moreover, cybersickness was assessed using the simulator sickness questionnaire at the beginning and at the end of the test. The head-saliency results over time show that with the presence of higher-order Ambisonics audio, subjects concentrate more on the directions sound is coming from. No influence of audio scenario on cybersickness scores was observed. From the analysis of the verbal scene descriptions, it was found that users were attentive to the omnidirectional video, but only for the ‘no audio’ scenario provided minute and insignificant details of the scene objects. The audiovisual saliency dataset is made available following the open science approach already used for the audiovisual scene recordings we previously published. The data is sought to enable training of visual and audiovisual saliency prediction models for interactive experiences. Ashutosh Singla, Thomas Robotham, Abhinav Bhattacharya, William Menz, Emanuël A. P. Habets, Alexander Raake |
QoMEX | 5 |
| 2023 | Influence of Multi-Modal Interactive Formats on Subjective Audio Quality and Exploration BehaviorabstractThis study uses a mixed between- and within-subjects test design to evaluate the influence of interactive formats on the quality of binaurally rendered 360° spatial audio content. Focusing on ecological validity using real-world recordings of 60 s duration, three independent groups of subjects () were exposed to three formats: audio only (A), audio with 2D visuals (A2DV), and audio with head-mounted display (AHMD) visuals. Within each interactive format, two sessions were conducted to evaluate degraded audio conditions: bit-rate and Ambisonics order. Our results show a statistically significant effect (p < .05) of format only on spatial audio quality ratings for Ambisonics order. Exploration data analysis shows that format A yields little variability in exploration, while formats A2DV and AHMD yield broader viewing distribution of 360° content. The results imply audio quality factors can be optimized depending on the interactive format. Thomas Robotham, Ashutosh Singla, Alexander Raake, Olli Rummukainen, Emanuël A. P. Habets |
IMX | 5 |
| 2023 | Speaker Counting and Separation From Single-Channel Noisy MixturesabstractWe address the problem of speaker counting and separation from a noisy, single-channel, multi-source, recording. Most of the works in the literature assume mixtures containing two to five speakers. In this work, we consider noisy speech mixtures with one to five speakers and noise-only recordings. We propose a deep neural network (DNN) architecture, that predicts a speaker count of zero for noise-only recordings and predicts the individual clean speaker signals and speaker count for mixtures of one to five speakers. The DNN is composed of transformer layers and processes the recordings using the long-time and short-time sequence modeling approach to masking in a learned time-feature domain. The network uses an encoder-decoder attractor module with long-short term memory units to generate a variable number of outputs. The network is trained with simulated noisy speech mixtures composed of the speech recordings from WSJ0 corpus, and noise recordings from the WHAM! corpus. We show that the network achieves 99% speaker counting accuracy and more than 19 dB improvement in the scale-invariant signal-to-noise ratio for mixtures of up to three speakers. Srikanth Raj Chetupalli, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | AmbiSep: Joint Ambisonic-to-Ambisonic Speech Separation and Noise ReductionabstractBlind separation of the sounds in an Ambisonic sound scene is a challenging problem, especially when the spatial impression of these sounds needs to be preserved. In this work, we consider Ambisonic-to-Ambisonic separation of reverberant speech mixtures, optionally containing noise. A supervised learning approach is adopted utilizing a transformer-based deep neural network, denoted by AmbiSep. AmbiSep takes mutichannel Ambisonic signals as input and estimates separate multichannel Ambisonic signals for each speaker while preserving their spatial images including reverberation. The GPU memory requirement of AmbiSep during training increases with the number of Ambisonic channels. To overcome this issue, we propose different aggregation methods. The model is trained and evaluated for first-order and second-order Ambisonics using simulated speech mixtures. Experimental results show that the model performs well on clean and noisy reverberant speech mixtures, and also generalizes to mixtures generated with measured Ambisonic impulse responses. Adrian Herzog, Srikanth Raj Chetupalli, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Blind Reverberation Time Estimation in Dynamic Acoustic ConditionsabstractThe estimation of reverberation time from real-world signals plays a central role in a wide range of applications. In many scenarios, acoustic conditions change over time which in turn requires the estimate to be updated continuously. Previously proposed methods involving deep neural networks were mostly designed and tested under the assumption of static acoustic conditions. In this work, we show that these approaches can perform poorly in dynamically evolving acoustic environments. Motivated by a recent trend towards data-centric approaches in machine learning, we propose a novel way of generating training data and demonstrate, using an existing deep neural network architecture, the considerable improvement in the ability to follow temporal changes in reverberation time. Philipp Götz, Cagdas Tuna, Andreas Walther 0001, Emanuël A. P. Habets |
ICASSP | 4 |
| 2022 | Speech Separation for an Unknown Number of Speakers Using Transformers With Encoder-Decoder Attractorsabstract5393 Srikanth Raj Chetupalli, Emanuël A. P. Habets |
INTERSPEECH | 2 |
| 2022 | Audiovisual Database with 360° Video and Higher-Order Ambisonics Audio for Perception, Cognition, Behavior, and QoE Evaluation ResearchabstractResearch into multi-modal perception, human cog-nition, behavior, and attention can benefit from high-fidelity content that may recreate real-life-like scenes when rendered on head-mounted displays. Moreover, aspects of audiovisual perception, cognitive processes, and behavior may complement questionnaire-based Quality of Experience (QoE) evaluation of interactive virtual environments. Currently, there is a lack of high-quality open-source audiovisual databases that can be used to evaluate such aspects or systems capable of reproducing high-quality content. With this paper, we provide a publicly available audiovisual database consisting of twelve scenes capturing real-life nature and urban environments with a video resolution of 7680×3840 at 60 frames-per-second and with 4th-order Ambison-ics audio. These 360° video sequences, with an average duration of 60 seconds, represent real-life settings for systematically evaluating various dimensions of uni-/multi-modal perception, cognition, behavior, and QoE. The paper provides details of the scene requirements, recording approach, and scene descriptions. The database provides high-quality reference material with a balanced focus on auditory and visual sensory information. The database will be continuously updated with additional scenes and further metadata such as human ratings and saliency information. Thomas Robotham, Ashutosh Singla, Olli Rummukainen, Alexander Raake, Emanuël A. P. Habets |
QoMEX | 5 |
| 2022 | Dialogue Enhancement and Listening Effort in Broadcast Audio: A Multimodal EvaluationabstractDialogue enhancement (DE) plays a vital role in broadcasting, enabling the personalization of the relative level between foreground speech and background music and effects. DE has been shown to improve the quality of experience, intel-ligibility, and self-reported listening effort (LE). A physiological indicator of LE known from audiology studies is pupil size. The relation between pupil size and LE is typically studied using artificial sentences and background noises not encountered in broadcast content. This work evaluates the effect of DE on LE in a multimodal manner that includes pupil size (tracked by a VR headset) and real-world audio excerpts from TV. Under ideal listening conditions, 28 normal-hearing participants listened to 30 audio excerpts presented in random order and processed by conditions varying the relative level between foreground and background audio. One of these conditions employed a recently proposed source separation system to attenuate the background given the original mixture as the sole input. After listening to each excerpt, subjects were asked to repeat the heard sentence and self-report the LE. Mean pupil dilation and peak pupil dilation were analyzed and compared with the self-report and the word recall rate. The multimodal evaluation shows a consistent trend of decreasing LE along with decreasing background level. DE, also when enabled by source separation, significantly reduces the pupil size as well as the self-reported LE. This highlights the benefit of personalization functionalities at the user's end. Matteo Torcoli, Thomas Robotham, Emanuël A. P. Habets |
QoMEX | 3 |
| 2022 | A Hybrid Acoustic Echo Reduction Approach Using Kalman Filtering and Informed Source Extraction with Improved TrainingabstractState-of-the-art acoustic echo and noise reduction combines adaptive filters with a deep neural network-based postfilter. While the signal-to-distortion ratio is often used for training, it is not well-defined for all echo-reduction scenarios. We propose well-defined loss functions for training and modifications of a recently proposed echo reduction system that is based on informed source extraction. The modifications include using a Kalman filter as a prefilter and a cyclical learning rate scheduler. The proposed modifications improve the performance on the blind test set of the Interspeech 2021 AEC challenge. A comparison to the challenge-winner shows that the proposed system underperforms the winner by 0.1 mean opinion score (MOS) points in double-talk echo reduction. However, it outperforms the winner by 0.3 MOS points in echo-only echo reduction. In all other scenarios, both algorithms perform comparably. Wolfgang Mack, Emanuël A. P. Habets |
SLT | 2 |
| 2022 | Signal-aware direction-of-arrival estimation using attention mechanisms
Wolfgang Mack, Julian Wechsler, Emanuël A. P. Habets |
Comput. Speech Lang. | 3 |
| 2022 | A Data-Driven Approach to Audio DecorrelationabstractThe degree of correlation between two audio signals entering the ears is known to have a significant impact on the spatial perception of a sound image. Audio signal decorrelation is therefore a widely used tool in various applications within the field of spatial audio processing. This paper explores for the first time the use of a data-driven approach for audio decorrelation. We propose a convolutional neural network architecture that is trained with the help of a state-of-the-art reference decorrelator. The proposed approach is evaluated using music and applause signals by means of objective evaluations as well as through a listening test. The proposed approach can serve as a proof of concept to address common limitations of existing decorrelation techniques in future work, which include introduction of temporal smearing and coloration artifacts and the production of a limited number of mutually uncorrelated output signals. Carlotta Anemüller, Oliver Thiergart, Emanuël A. P. Habets |
IEEE Signal Process. Lett. | 3 |
| 2022 | TDOA-Based Robust Sound Source Localization With Sparse Regularization in Wireless Acoustic Sensor NetworksabstractTime difference of arrival (TDOA) measurements, which are contaminated by large values of error, known as outliers, would have a significant impact on the accuracy of sound source localization (SSL) in wireless acoustic sensor networks (WASNs). Few techniques are reported in the literature to tackle SSL in WASNs by taking TDOA outliers into consideration. To mitigate the effect of outliers on the accuracy of SSL, we propose outlier-resistant robust sound source localization (RSSL) algorithms based on sparse regularization using an unsynchronized network of microphone arrays. The TDOA errors are divided into two components: a) energy-bounded inliers and b) outliers. Assuming that outliers are sparse in the measurement set, we formulate the RSSL problem as that of minimizing the number of outliers, mathematically, a$\ell _{0}$(pseudo)-norm optimization problem with non-convex constraints. Five sub-optimal RSSL solvers are derived, among which the first two solvers are applicable to the scenario involving only outliers while the last three solvers concentrate on the scenario incorporating both outliers and inliers. In common, these solvers exploit a convex approximation technique called Concave Convex Procedure to dispose of the non-convex constraints. Differently, the first solver approximates the original$\ell _{0}$(pseudo)-norm cost function with the$\ell _{1}$norm while a concave surrogate function is adopted in the second solver to yield a tighter approximation to the$\ell _{0}$(pseudo)-norm. Apart from the application of these two approximation techniques, the third and fourth solvers relax the non-convex$\ell _{2}$norm constraint with the$\ell _{\infty }$norm. The fifth solver is dedicated to the$\ell _{1}$norm regularization problem with the Lasso formulation, which is equivalent to the M-estimator of Huber’s function solved via the iteratively reweighted least squares paradigm. Experimental results validate the effectiveness and robustness of the proposed algorithms. Xudong Dang, Emanuël A. P. Habets, Hongyan Zhu |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Comparing Direct and Indirect Methods of Audio Quality Evaluation in Virtual Reality Scenes of Varying ComplexityabstractMany quality evaluation methods are used to assess uni-modal audio or video content without considering perceptual, cognitive, and interactive aspects present in virtual reality (VR) settings. Consequently, little is known regarding the repercussions of the employed evaluation method, content, and subject behavior on the quality ratings in VR. This mixed between- and within-subjects study uses four subjective audio quality evaluation methods (viz. multiple-stimulus with and without reference for direct scaling, and rank-order elimination and pairwise comparison for indirect scaling) to investigate the contributing factors present in multi-modal 6-DoF VR on quality ratings of real-time audio rendering. For each between-subjects employed method, two sets of conditions in five VR scenes were evaluated within-subjects. The conditions targeted relevant attributes for binaural audio reproduction using scenes with various amounts of user interactivity. Our results show all referenceless methods produce similar results using both condition sets. However, rank-order elimination proved to be the fastest method, required the least amount of repetitive motion, and yielded the highest discrimination between spatial conditions. Scene complexity was found to be a main effect within results, with behavioral and task load index results implying more complex scenes and interactive aspects of 6-DoF VR can impede quality judgments. Thomas Robotham, Olli Rummukainen, Miriam Kurz, Marie Eckert, Emanuël A. P. Habets |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2021 | Direction Preserving Wind Noise Reduction Of B-Format SignalsabstractNoise reduction in B-format recordings is particularly challenging as it concurrently requires to suppress undesired signals and preserve the spatial properties of the acoustic environment. In particular, wind noise poses an undesirable acoustic condition outdoors. In this work, methods to reduce wind noise while limiting the spatial distortions of the original signal are proposed based on recent works of the present authors. The main contributions are the derivation of the mentioned methods for B-format signals and the usage of the dipole-to-omnidirectional power ratio to control the trade-off between the desired-signal distortion and the noise reduction. The proposed methods are evaluated using B-format wind-noise recordings. Adrian Herzog, Daniele Mirabilii, Emanuël A. P. Habets |
ICASSP | 3 |
| 2021 | Efficient Training Data Generation for Phase-Based DOA EstimationabstractDeep learning (DL) based direction of arrival (DOA) estimation is an active research topic and currently represents the state-of-the-art. Usually, DL-based DOA estimators are trained with recorded data or computationally expensive generated data. Both data types require significant storage and excessive time to, respectively, record or generate. We propose a low complexity online data generation method to train DL models with a phase-based feature input. The data generation method models the phases of the microphone signals in the frequency domain by employing a deterministic model for the direct path and a statistical model for the late reverberation of the room transfer function. By an evaluation using data from measured room impulse responses, we demonstrate that a model trained with the proposed training data generation method performs comparably to models trained with data generated based on the source-image method. Fabian Hübner, Wolfgang Mack, Emanuël A. P. Habets |
ICASSP | 3 |
| 2021 | An Empirical Study of Visual Features for DNN Based Audio-Visual Speech Enhancement in Multi-Talker EnvironmentsabstractAudio-visual speech enhancement (AVSE) methods use both audio and visual features for the task of speech enhancement and the use of visual features has been shown to be particularly effective in multi-speaker scenarios. In the majority of deep neural network (DNN) based AVSE methods, the audio and visual data are first processed separately using different sub-networks, and then the learned features are fused to uti-lize the information from both modalities. There have been various studies on suitable audio input features and network architectures, however, to the best of our knowledge, there is no published study that has investigated which visual features are best suited for this specific task. In this work, we perform an empirical study of the most commonly used visual features for DNN based AVSE, the pre-processing requirements for each of these features, and investigate their influence on the performance. Our study shows that despite the overall better performance of embedding-based features, their computationally intensive pre-processing makes their use difficult in low resource systems. For such systems, optical flow or raw pixels-based features might be better suited. Shrishti Saha Shetu, Soumitro Chakrabarty, Emanuël A. P. Habets |
ICASSP | 3 |
| 2021 | Cognitive Load Estimation Based on Pupillometry in Virtual Reality with Uncontrolled Scene LightingabstractVirtual reality (VR) technology enables and requires new ways of user experience testing in immersive environments. For various aspects of user experience, objective assessment of the cognitive load can be a useful parameter. With eye-tracking becoming a more widespread feature of current VR headsets, pupillometry is an appealing option to unobtrusively measure cognitive load during a VR experience. This paper shows that pupil size measured by an off-the-shelf VR headset with an integrated eye tracker positively correlates with the self-reported cognitive load during a standard n-back task adapted to a VR environment. To overcome the need for steady scene-lighting conditions, we present a method to correct for the light-induced pupil size changes, otherwise masking the cognitive load effects. Our results show that a commercially available VR headset with eye tracking can be used to measure the cognitive load in unpredictable lighting conditions without additional hardware. Marie Eckert, Emanuël A. P. Habets, Olli Rummukainen |
QoMEX | 2 |
| 2020 | Signal-Aware Broadband DOA Estimation Using Attention MechanismsabstractWe refer to direction-of-arrivals (DOAs) estimation of a user-defined subset of directional (desired) sound sources as signal-aware DOA estimation. Source selection, thereby, can be achieved with time-frequency masks to apply attention to TF bins dominated by desired sources. With deep neural networks (DNNs), another option is to train the DNN to estimate the DOAs only of specific classes, like speech, and disregard the DOAs of other classes. Consequently, changing the desired classes requires retraining the DNN. Also, the mask-based approaches are trained for sources known prior to DNN training. To obtain a flexible signal-aware DOA estimator, we propose to use binary mask attention with a DNN for multi-source DOA estimation trained with artificial noise. The desired sources are determined via binary masks, which allows a redefinition by changing the masks. Consequently, the DOA estimator is independent of the desired sources. We experiment with attention in form of oracle and estimated binary masks. Wolfgang Mack, Ullas Bharadwaj, Soumitro Chakrabarty, Emanuël A. P. Habets |
ICASSP | 4 |
| 2020 | Data-Driven Wind Speed Estimation Using Multiple MicrophonesabstractA deep neural network (DNN) based approach for estimating the speed of airflows using closely-spaced microphones is proposed. The spatial characteristics of wind noise measured with a smallaperture array are exploited, i.e., the low-frequency spatial coherence of wind noise signals is used as an input feature. The output is an estimate of the wind speed averaged over a specific time interval. The DNN is trained using synthetic wind noise, which overcomes the time-consuming data collection and allows to isolate wind noise from different acoustic sources. The dataset used for testing comprises wind noise measured outdoors with a circular linear array and a ground truth obtained using an ultrasonic anemometer. The obtained model is applied to generated and measured wind noise. The performance of the proposed method is assessed across a wide range of wind speeds and directions, using different time resolutions. Daniele Mirabilii, Kishor Kayyar Lakshminarayana, Wolfgang Mack, Emanuël A. P. Habets |
ICASSP | 4 |
| 2020 | Low Complexity NLMS for Multiple Loudspeaker Acoustic ECHO Canceller Using Relative Loudspeaker Transfer FunctionsabstractSpeech signals captured by a microphone mounted to a smart soundbar or speaker are inherently contaminated by echos. Modern smart devices are usually characterized by low computational capabilities and low memory resources; in these cases, a low-complexity acoustic echo canceller (AEC) may be preferred even though a tolerable degradation in the cancellation occurs. In principle, devices with multiple loudspeakers need an individual AEC for each loudspeaker because the transfer function (TF) from each loudspeaker to the microphone must be estimated. In this paper, we present an normalized least mean square (NLMS) algorithm for a multi-loudspeaker case using relative loudspeaker transfer functions (RLTFs). In each iteration, the RLTFs between each loudspeaker and the reference loudspeaker are estimated first, and then the primary TF between the reference loudspeaker and the microphone. Assuming loudspeakers that are close to each other, the RLTFs can be estimated using fewer coefficients w.r.t. the primary TF, yielding a reduction of 3:4 in computational complexity and 1:2 in memory usage. The algorithm is evaluated using both simulated and real room impulse responses (RIRs) of two loudspeakers with a reverberation time set to 0.3 s and several distances between the loudspeakers. Ofer Schwartz, Emanuël A. P. Habets, Sharon Gannot |
ICASSP | 2 |
| 2020 | Online Blind Reverberation Time Estimation Using CRNNsabstractS.5061-5065 Shuwen Deng, Wolfgang Mack, Emanuël A. P. Habets |
INTERSPEECH | 3 |
| 2020 | Single-Channel Blind Direct-to-Reverberation Ratio Estimation Using MaskingabstractAcoustic parameters, like the direct-to-reverberation ratio (DRR), can be used in audio processing algorithms to perform, e.g., dereverberation or in audio augmented reality. Often, the DRR is not available and has to be estimated blindly from recorded audio signals. State-of-the-art DRR estimation is achieved by deep neural networks (DNNs), which directly map a feature representation of the acquired signals to the DRR. Motivated by the equality of the signal-to-reverberation ratio and the (channel-based) DRR under certain conditions, we formulate single-channel DRR estimation as an extraction task of two signal components from the recorded audio. The DRR can be obtained by inserting the estimated signals in the definition of the DRR. The extraction is performed using time-frequency masks. The masks are estimated by a DNN trained end-to-end to minimize the mean-squared error between the estimated and the oracle DRR. We conduct experiments with different preprocessing and mask estimation schemes. The proposed method outperforms state-of-the-art single- and multi-channel methods on the ACE challenge data corpus. Wolfgang Mack, Shuwen Deng, Emanuël A. P. Habets |
INTERSPEECH | 3 |
| 2020 | Deep Filtering: Signal Extraction and Reconstruction Using Complex Time-Frequency FiltersabstractSignal extraction from a single-channel mixture with additional undesired signals is most commonly performed using time-frequency (TF) masks. Typically, the mask is estimated with a deep neural network (DNN), and element-wise applied to the complex mixture short-time Fourier transform (STFT) representation to perform the extraction. Ideal mask magnitudes are zero for solely undesired signals in a TF bin and undefined for total destructive interference. Usually, masks have an upper bound to provide well-defined DNN outputs at the cost of limited extraction capabilities. We propose to estimate with a DNN a complex TF filter for each mixture TF bin which maps an STFT area in the respective mixture to the desired TF bin to address destructive interference in mixture TF bins. The DNN is optimized by minimizing the error between the extracted and the ground-truth desired signal allowing to learn the TF filters without having to specify ground-truth TF filters. We compare our approach with complex and real-valued TF masks by separating speech from a variety of different sound and noise classes from the Google AudioSet corpus. We also process the mixture STFT with notch-filters and zero whole time-frames, to simulate packet-loss during transmission, to demonstrate the reconstruction capabilities of our approach. The proposed method outperformed the baselines, especially when notch-filters and time-frame zeroing were applied. Wolfgang Mack, Emanuël A. P. Habets |
IEEE Signal Process. Lett. | 2 |
| 2020 | Direction and Reverberation Preserving Noise Reduction of Ambisonics SignalsabstractAmbisonics encodes the directional information of a sound field w.r.t. a reference position in an efficient and scalable way. However, the sound field might contain undesired noise. Reducing the noise while preserving the directional distribution of all sound field components is a challenging task. Recently, the present authors proposed a direction-preserving noise reduction method for higher-order Ambisonics (HOA) signals, which, in contrast to e.g. binaural beamforming methods, yields a HOA signal at the output. In this work, we investigate the direction-preserving noise reduction method further and compare it against a beamforming-based method and the matrix multi-channel Wiener filter. Different methods to estimate the power spectral densities which are needed for the noise reduction methods are discussed. Moreover, a method to preserve the reverberation of the desired signal is proposed. In the evaluation, the discussed methods are compared for different speech sources in anechoic and reverberant conditions and different noise types. Adrian Herzog, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Spatial Coherence-Aware Multi-Channel Wind Noise ReductionabstractOutdoor recording is particularly challenging in the presence of wind, which induces highly non-stationary noise in the microphone signals. To enhance a desired signal, e.g., speech, a dedicated noise reduction processing is required. The reduction is usually performed by estimating an unknown set of parameters, e.g., the noise and the speech power spectral densities (PSDs). In contrast to the commonly used assumption of uncorrelated wind noise in multi-channel recordings, we assume the spatial correlation of wind noise contributions to be non-zero at low frequencies when closely-spaced microphones are employed. In our earlier work, we assumed that the spatial coherence was known, i.e., given by a model which depends on the free-field speed and direction of the air stream. As these quantities are unknown in practice, we propose in this work a method to recursively estimate the spatial coherence matrix based on the microphone observations. In addition, we prove the equivalence of two recently developed noise PSD estimation methods when uncorrelated wind noise is assumed, and we propose an approximation of both estimators which is independent of the propagation vector of the speech source at sufficiently low frequency and for a small inter-microphone distance. An evaluation in terms of improvements in speech quality, signal-to-noise ratio and intelligibility is carried out using both simulated and measured wind noise samples. Daniele Mirabilii, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Scattering in Feedback Delay NetworksabstractFeedback delay networks (FDNs) are recursive filters, which are widely used for artificial reverberation and decorrelation. One central challenge in the design of FDNs is the generation of sufficient echo density in the impulse response without compromising the computational efficiency. In a previous contribution, we have demonstrated that the echo density of an FDN can be increased by introducing so-called delay feedback matrices where each matrix entry is a scalar gain and a delay. In this contribution, we generalize the feedback matrix to arbitrary lossless filter feedback matrices (FFMs). As a special case, we propose the velvet feedback matrix, which can create dense impulse responses at a minimal computational cost. Further, FFMs can be used to emulate the scattering effects of non-specular reflections. We demonstrate the effectiveness of FFMs in terms of echo density and modal distribution. Sebastian J. Schlecht, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | 3D Room Geometry Inference Using a Linear Loudspeaker Array and a Single MicrophoneabstractSound reproduction systems may highly benefit from detailed knowledge of the acoustic space to enhance the spatial sound experience. This article presents a room geometry inference method based on identification of reflective boundaries using a high-resolution direction-of-arrival map produced via room impulse responses (RIRs) measured with a linear loudspeaker array and a single microphone. Exploiting the sparse nature of the early part of the RIRs, Elastic Net regularization is applied to obtain a 2D polar-coordinate map, on which the direct path and early reflections appear as distinct peaks, described by their propagation distance and direction of arrival. Assuming a separable room geometry with four side-walls perpendicular to the floor and ceiling, and imposing pre-defined geometrical constraints on the walls, the 2D-map is segmented into six regions, each corresponding to a particular wall. The salient peaks within each region are selected as candidates for the first-order wall reflections, and a set of potential room geometries is formed by considering all possible combinations of the associated peaks. The room geometry is then inferred using a cost function evaluated on the higher-order reflections computed via beam tracing. The proposed method is tested with both simulated and measured data. Cagdas Tuna, Antonio Canclini, Federico Borra, Philipp Götz, Fabio Antonacci, Andreas Walther 0001, Augusto Sarti, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 8 |
| 2019 | Direction Preserving Wiener Matrix Filtering for Ambisonic Input-output SystemsabstractWe present a spatial matrix filtering framework for noise reduction in the spherical harmonics (ambisonics) domain (SHD), which outputs an SHD signal vector rather than one signal as commonly provided by beamforming approaches. We discuss two spatial matrix filtering methods: A multi-beamformer method using known propagation vectors of the desired signal components and a method preserving the directional information in an optimal way. Parametric multi-channel Wiener filter solutions for both methods are discussed and a performance evaluation is conducted. It is shown that the direction preserving method preserves the spatial distribution of the desired sounds and residual noise at the cost of less noise reduction and higher signal distortion when compared to the multi-beamformer approach. Moreover, no spatial parameters have to be estimated. Adrian Herzog, Emanuël A. P. Habets |
ICASSP | 2 |
| 2019 | Multi-channel Wind Noise Reduction Using the Corcos ModelabstractOutdoor recordings of speech are often corrupted by wind noise, which is difficult to reduce due to its high non-stationarity. In this work, a multi-channel wind noise reduction method is presented, based on a joint estimation of the speech and wind noise power spectral densities. In contrast to existing methods that assume uncorrelated wind noise, the estimation phase is performed exploiting the spatial characteristics of wind noise measured by closely-spaced microphones. Here the characteristics are approximated by a fluid-dynamics model, termed the Corcos model. An additional contribution is the employment of a frequency dependent trade-off parameter in the reduction phase, which depends on the ratio of the difference-signal power to the sum-signal power of a sub-set of two microphones. In particular, the trade-off parameter of the parametric multi-channel Wiener filter is used to control the trade-off between noise reduction and speech distortion. The proposed method can be used also to reduce spatially uncorrelated wind noise. An evaluation in terms of speech quality and signal-to-noise ratio improvements is carried out to compare the proposed method to an existing multichannel wind noise reduction method, using recorded and simulated wind noise samples. Daniele Mirabilii, Emanuël A. P. Habets |
ICASSP | 2 |
| 2019 | Combining Linear Spatial Filtering and Non-linear Parametric Processing for High-quality Spatial Sound CapturingabstractFlexible spatial sound capturing and reproduction can be achieved with multiple microphones by using linear spatial filtering or non-linear parametric processing. The non-linear approaches usually provide a superior spatial resolution compared to the linear approaches but can result in artifacts due to violations of the sound field model. In this paper, we combine both approaches to achieve a high robustness against model violations and a high spatial resolution. We assume linear spatial filters that approximate the spatial responses of the desired output format and compensate remaining deviations with an optimal post filter. The post filter is computed such that the proposed approach behaves like a linear system when the spatial filters achieve the desired spatial response, and scales towards a non-linear system otherwise. Experimental results show that the proposed approach can significantly reduce distortions of existing parametric processing schemes especially when a sufficiently high number of microphones is available. Oliver Thiergart, Guendalina Milano, Emanuël A. P. Habets |
ICASSP | 3 |
| 2019 | Perceptual Study of Near-Field Binaural Audio Rendering in Six-Degrees-of-Freedom Virtual RealityabstractAuditory localization cues in the near-field are significantly different than in the far-field. The near-field region is within an arm's length of the listener allowing to integrate proprioceptive cues to determine the location of an object in space. This perceptual study compares three non-individualized methods to apply head-related transfer functions (HRTFs) in six-degrees-of-freedom near-field audio rendering, namely, far-field measured HRTFs, multi-distance measured HRTFs, and spherical-model-based HRTFs with near-field extrapolation. To set our findings in context, we provide a real-world hand-held audio source for comparison along with a distance-invariant condition. Two modes of interaction are compared in an audio-visual virtual reality: one allowing the participant to move the audio object dynamically and the other with a stationary audio object but a freely moving listener. Olli S. Hurnmukainen, Sebastian J. Schlecht, Thomas Robotham, Axel Plinge, Emanuël A. P. Habets |
VR | 5 |
| 2019 | Eigenbeam-ESPRIT for DOA-Vector EstimationabstractSeveral techniques exist to estimate the directions of arrival (DOAs) of sound sources captured with a spherical microphone array. The eigenbeam rotational invariance technique (EB-ESPRIT) uses recurrence relations of spherical harmonics to estimate the DOAs. In this letter, we propose a new EB-ESPRIT method which uses three recurrence relations and a joint-diagonalization procedure to estimate the unit-vectors pointing to the source DOAs. We evaluate the angular estimation errors of the proposed method under noisy and reverberant conditions and compare it to existing EB-ESPRIT methods. We find that our proposed method can estimate the source DOAs with higher accuracy compared to the discussed existing EB-ESPRIT methods, and the estimation accuracy does not depend on the source DOAs. Adrian Herzog, Emanuël A. P. Habets |
IEEE Signal Process. Lett. | 2 |
| 2019 | Iterative DFT-Domain Inverse Filter Optimization Using a Weighted Least-Squares CriterionabstractFor many inverse filtering problems, finite impulse response filters are designed according to least-squares criteria, where time-domain and frequency-domain weights are often applied to achieve optimal results for the considered application. While least-squares-optimal filter coefficients are given by an explicit formula, the computation cost to compute its solution is proportional up to the third power of the number of jointly optimized filter coefficients. A joint optimization of all filter coefficients is necessary whenever a time-domain or a frequency-domain weight is introduced. This imposes limits for filter lengths and numbers of channels in many real-world scenarios. In this contribution, an algorithm is presented that yields time-domain filter coefficients optimized to meet such a weighted least-squares criterion, while performing the most expensive computation steps efficiently in the discrete Fourier transform domain. As a consequence, the demands on computational power and memory are kept on a moderate level, even for large-scale problems. A rigorous mathematical derivation is provided that identifies all approximations used in the algorithm. Additionally, an effective regularization method is proposed that does not depend much on the regularization parameters. Furthermore, the proposed approach is experimentally evaluated considering a sound-zones scenario, which is one of many possible application areas. In that way, the applicability of the proposed approach is verified. Martin Schneider 0009, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | CountNet: Estimating the Number of Concurrent Speakers Using Supervised LearningabstractEstimating the maximum number of concurrent speakers from single-channel mixtures is a challenging problem and an essential first step to address various audio-based tasks such as blind source separation, speaker diarization, and audio surveillance. We propose a unifying probabilistic paradigm, where deep neural network architectures are used to infer output posterior distributions. These probabilities are in turn processed to yield discrete point estimates. Designing such architectures often involves two important and complementary aspects that we investigate and discuss. First, we study how recent advances in deep architectures may be exploited for the task of speaker count estimation. In particular, we show that convolutional recurrent neural networks outperform recurrent networks used in a previous study when adequate input features are used. Even for short segments of speech mixtures, we can estimate up to five speakers, with a significantly lower error than other methods. Second, through comprehensive evaluation, we compare the best-performing method to several baselines, as well as the influence of gain variations, different data sets, and reverberation. The output of our proposed method is compared to human performance. Finally, we give insights into the strategy used by our proposed method. Fabian-Robert Stöter, Soumitro Chakrabarty, Bernd Edler, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | Low-Complexity Multi-Microphone Acoustic Echo Control in the Short-Time Fourier Transform DomainabstractMany modern communication and smart devices are equipped with several microphones, in addition to one or more loudspeakers. Each microphone not only acquires sounds produced in the near-end room, i.e., desired near-end speech, background noise, and other interferences, but also a far-end signal that is reproduced by the loudspeaker(s). This particular type of acoustic coupling, commonly denoted as acoustic echo, can be reduced in a distortionless manner by employing multi-microphone acoustic echo cancellation (MM-AEC) techniques. However, under noisy conditions, the performance of AEC is limited by the echo-to-noise ratio, and additional echo reduction is needed. Further, to ensure high-quality end-to-end communication in noisy environments, background noise has to be reduced as well. To achieve the latter, multi-microphone speech enhancement techniques, such as beamforming (BF), are often used as they are capable of reducing undesired signal components while causing little distortion to the desired near-end speech. In spite of its high computational cost, the most effective solution to reduce acoustic echoes and background noise is to cascade MM-AEC and BF. In this work, a low-complexity multi-microphone echo controller is introduced, which not only combines low-complexity MM-AEC with BF, but also integrates residual echo reduction into the beamformer design. Maria Luis Valero, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Dual-Channel Modulation Energy Metric for Direct-to-Reverberation Ratio EstimationabstractNon-intrusive estimators for acoustic parameters like the direct-to-reverberation ratio (DRR) are useful tools but still perform weakly as shown in the acoustic characterization of environments (ACE) challenge. In this paper, we develop a novel dual-channel metric based on the modulation energy domain for DRR estimation. In contrast to established modulation based single-channel metrics like the speech-to-reverberation modulation energy ratio (SRMR), we exploit the spatial information from two microphones as well as the temporal dynamics in the modulation energy domain. The developed metric shows a strong linear correlation to the DRR, which allows a simple mapping. It is shown that the metric is robust against the microphone array configuration, room characteristics and the speech signal. The proposed metric is compared to a reference method based on the spectral variance of the room transfer functions, and both metrics are evaluated using simulated and measured data. In our experiments, the proposed metric achieved a higher correlation and lower RMSE compared the reference method, and outperforms existing SRMR based DRR estimators. Sebastian Braun, João Felipe Santos, Emanuël A. P. Habets, Tiago H. Falk |
ICASSP | 3 |
| 2018 | A Weighted Least Squares Beam Shaping Technique for Sound Field ControlabstractA weighted least squares beam shaping technique for sound field control using a loudspeaker array is proposed. Given a desired spatial response at prescribed control points, the space-time filter is designed by solving a least squares minimization problem. To reduce the computational effort, we propose to place control points only along an arc of circumference centered at the center of the array and passing through a region of interest. Furthermore, we adopt a weighted least squares approach for the design of the space-time filter, so that control points at directions towards which we admit a looser control of the sound field are less relevant in the filter design. The choice of the weights depends on the specific application and we demonstrate the feasibility of the proposed approach for sound zones scenario with one bright and one dark zone. Antonio Canclini, Dejan Markovic, Martin Schneider 0009, Fabio Antonacci, Emanuël A. P. Habets, Andreas Walther 0001, Augusto Sarti |
ICASSP | 5 |
| 2018 | Classification vs. Regression in Supervised Learning for Single Channel Speaker Count EstimationabstractThe task of estimating the maximum number of concurrent speakers from single channel mixtures is important for various audio-based applications, such as blind source separation, speaker diarisation, audio surveillance or auditory scene classification. Building upon powerful machine learning methodology, we develop a Deep Neural Network (DNN) that estimates a speaker count. While DNNs efficiently map input representations to output targets, it remains unclear how to best handle the network output to infer integer source count estimates, as a discrete count estimate can either be tackled as a regression or a classification problem. In this paper, we investigate this important design decision and also address complementary parameter choices such as the input representation. We evaluate a state-of-the-art DNN audio model based on a Bi-directional Long Short-Term Memory network architecture for speaker count estimations. Through experimental evaluations aimed at identifying the best overall strategy for the task and show results for five seconds speech segments in mixtures of up to ten speakers. Fabian-Robert Stöter, Soumitro Chakrabarty, Bernd Edler, Emanuël A. P. Habets |
ICASSP | 4 |
| 2018 | Single-Channel Dereverberation Using Direct MMSE Optimization and Bidirectional LSTM NetworksabstractDereverberation is useful in hands-free communication and voice controlled devices for distant speech acquisition. Single-channel dereverberation can be achieved by applying a time-frequency (TF) mask to the short-time Fourier transform (STFT) representation of a reverberant signal. Recent approaches have used deep neural networks (DNNs) to estimate such masks. Previously proposed DNN-based mask estimation methods train a DNN to minimize the mean-squared-error (MSE) between the desired and estimated masks. Recent TF mask estimation methods for signal separation directly minimize instead the MSE between the desired and estimated STFT magnitudes. We apply this direct optimization concept to dereverberation. Moreover, as reverberation exceeds the duration of a single STFT frame, we propose to use a bidirectional long short-term memory (LSTM) network which is able to take the relation between multiple STFT frames into account. We evaluated our method for different reverberation times and source-microphone distances using simulated as well as measured room impulse responses of different rooms. An evaluation of the proposed method and a comparison with a state-of-the-art method demonstrate the superiority of our approach and its robustness to different acoustic conditions. Wolfgang Mack, Soumitro Chakrabarty, Fabian-Robert Stöter, Sebastian Braun, Bernd Edler, Emanuël A. P. Habets |
INTERSPEECH | 6 |
| 2018 | DoA Reliability for Distributed Acoustic TrackingabstractDistributed acoustic tracking estimates the trajectories of source positions using an acoustic sensor network. As it is often difficult to estimate the source-sensor range from individual nodes, the source positions have to be inferred from the direction-of-arrival (DoA) estimates. Due to reverberation and noise, the sound field becomes increasingly diffuse with increasing source-sensor distance, leading to a decreased Direction of Arrival (DoA)-estimation accuracy. To distinguish between accurate and uncertain DoA estimates, this letter proposes to incorporate the coherent-to-diffuse ratio as a measure of DoA reliability for single-source tracking. It is shown that the source positions, therefore, can be probabilistically triangulated by exploiting the spatial diversity of all nodes. Christine Evers, Emanuël A. P. Habets, Sharon Gannot, Patrick A. Naylor |
IEEE Signal Process. Lett. | 2 |
| 2018 | 3D Room Geometry Inference Based on Room Impulse Response StacksabstractRoom geometry inference is concerned with the localization of reflective boundaries in an enclosed space. This paper outlines a method for inferring room geometry based on the positions of loudspeakers and real or image microphones, which are computed using sets of times of arrival (TOAs) obtained from room impulse responses (RIRs). These RIRs describe the acoustic propagation between the loudspeakers in an array and a single microphone. First, peaks corresponding to TOAs in these RIRs are detected and labeled using an automated method. Second, the labeled TOA sets are used to estimate the real and image microphone positions, with knowledge of the loudspeaker array geometry. Third, using all these positions, the positions of reflection points on the available reflectors in the room are determined. The reflection points determine the reflectors' locations and orientations. This approach is largely automated and usable in real-world scenarios. Youssef El Baba, Andreas Walther 0001, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Linear Prediction-Based Online Dereverberation and Noise Reduction Using Alternating Kalman FiltersabstractMultichannel linear prediction-based dereverberation in the short-time Fourier transform (STFT) domain has been shown to be highly effective. Using this framework, the desired dereverberated multichannel signal is obtained by filtering the noise-free reverberant signals using the estimated multichannel autoregressive (MAR) coefficients. To use such methods in the presence of noise, especially in the case of online processing, remains a challenging problem. Existing sequential enhancement structures, which first remove the noise and then estimate the MAR coefficients, suffer from a causality problem as both the optimal noise reduction and dereverberation stages depend on the current output of each other. To address this problem, an algorithm that consists of two alternating Kalman filters to estimate the noise-free reverberant signals and the (MAR) coefficients is proposed. The causality of the estimation procedure is important when dealing with time-variant acoustic scenarios, where the MAR coefficients are time-varying. The proposed method is evaluated using simulated and measured acoustic impulse responses and is compared to a method based on the same signal model. In addition, a method to control the reverberation reduction and noise reduction independently is derived. Sebastian Braun, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Evaluation and Comparison of Late Reverberation Power Spectral Density EstimatorsabstractReduction of late reverberation can be achieved using spatio-spectral filters, such as the multichannel Wiener filter. To compute this filter, an estimate of the late reverberation power spectral density (PSD) is required. In recent years, a multitude of late reverberation PSD estimators have been proposed. In this paper, these estimators are categorized into several classes, their relations and differences are discussed, and a comprehensive experimental comparison is provided. To compare their performance, simulations in controlled as well as practical scenarios are conducted. It is shown that a common weakness of spatial coherence-based estimators is their performance in high direct-to-diffuse ratio conditions. To mitigate this problem, a correction method is proposed and evaluated. It is shown that the proposed correction method can decrease the speech distortion without significantly affecting the reverberation reduction. Sebastian Braun, Adam Kuklasinski, Ofer Schwartz, Oliver Thiergart, Emanuël A. P. Habets, Sharon Gannot, Simon Doclo, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2018 | A Bayesian Approach to Informed Spatial Filtering With Robustness Against DOA Estimation ErrorsabstractA Bayesian approach to spatial filtering is presented, which is robust to uncertain or erroneous direction-of-arrival (DOA) information. The proposed framework aims to capture multiple sound sources at each time-frequency instant with an arbitrary direction-dependent gain, while attenuating diffuse sound and noise. For robustness, the DOA corresponding to each sound source is assumed to be a discrete random variable with a prior defined on a discrete set of candidate DOAs over the whole DOA space. With this assumption, the desired spatial filter is given as a weighted sum of spatial filters corresponding to a specific combination of probable DOA values, where the weights are given by the joint posterior probabilities of the combination of DOA values. Assuming the whole DOA space as the support for each random variable results in redundant computations and contributes to a high computational cost. To alleviate this problem, a narrowband DOA estimate-based posterior probability approximation method is proposed, which isolates regions in the DOA space with high probability of containing the actual source DOAs to compute time-adaptive supports for each random variable. Through experimental analysis, we demonstrate the robustness of the proposed framework against DOA estimation errors. Experimental evaluation with simulated and measured room impulse responses, in terms of objective performance measures, demonstrates the effectiveness of the framework to perform spatial filtering in noisy and reverberant acoustic environments. Soumitro Chakrabarty, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Blind Source Separation of Moving Sources Using Sparsity-Based Source Detection and TrackingabstractSparsity-based blind source separation (BSS) algorithms in the short time-frequency (TF) domain have received a lot of attention due to their versatility and noise reduction capabilities. In most of these algorithms, the estimation of the BSS filters relies on the accurate association of each time-frequency bin to the dominant source at that bin. The TF bin associations are then used to estimate the statistics of the source signals, and BSS is achieved by optimal spatial filters computed using the estimated statistics. The main objective of this paper is to apply such a framework to scenarios with an unknown number of moving sources. While state-of-the-art approaches employ online clustering algorithms to solve the problem for moving sources, we propose an approximate Bayesian tracker and perform the association of each TF bin to the dominant source using the tracker's measurement-to-source association probabilities. Therefore, the choice of the underlying narrowband models and measurements for the tracker as well as the resulting tracking algorithm constitute the main contributions of this paper. The TF bin associations obtained from the tracker are then used to estimate the statistics of the source signals. The performance of the resulting BSS filters is compared to the performance of state-of-the-art sparsity-based and independent vector analysis-based BSS algorithms. Our proposed approach targets scenarios with at least two spatially separated microphone arrays, with known microphone positions and relative orientations. The framework also allows for efficient management of a time-varying number of sources. Maja Taseska, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Time of arrival disambiguation using the linear Radon transformabstractEcho labeling, the challenging task of assigning acoustic reflections to image sources, is equivalent to the highly-important disambiguation task in room geometry inference. A method using the Radon transform, an image processing tool, is proposed to address this challenge. The method relies on acoustic wavefront detection in room impulse response stacks, obtained with a uniform linear array of loudspeakers and one microphone. We show in our experiments that the proposed method can both label and detect echoes. Youssef El Baba, Andreas Walther 0001, Emanuël A. P. Habets |
ICASSP | 3 |
| 2017 | Evaluation of binaural reproduction systems from behavioral patterns in a six-degrees-of-freedom wayfinding taskabstractThis paper proposes a new method for evaluating real-time binaural reproduction systems by means of a wayfinding task in six degrees of freedom. Participants physically walk to sound objects in a virtual reality created by a head-mounted display and binaural audio. We show how the localization accuracy of spatial audio rendering is reflected by objective measures of the participants' behavior. The method allows for comparative evaluation of different rendering systems as well as the subjective assessment of the quality of experience. Olli Rummukainen, Sebastian J. Schlecht, Axel Plinge, Emanuël A. P. Habets |
QoMEX | 4 |
| 2017 | Feedback Delay Networks: Echo Density and Mixing TimeabstractFeedback delay networks (FDNs) are frequently used to generate artificial reverberation. This paper discusses the temporal features of impulse responses produced by FDNs, i.e., the number of echoes per time unit and its evolution over time. This so-called echo density is related to known measures of mixing time and their psychoacoustic correlates such as auditive perception of the room size. It is shown that the echo density of FDNs follows a polynomial function, whereby the polynomial coefficients can be derived from the lengths of the delays for which an explicit method is given. The mixing time of impulse responses can be predicted from the echo density, and conversely, a desired mixing time can be achieved by a derived mean delay length. A Monte Carlo simulation confirms the accuracy of the derived relation of mixing time and delay lengths. Sebastian J. Schlecht, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Multispeaker LCMV Beamformer and Postfilter for Source Separation and Noise ReductionabstractThe problem of source separation and noise reduction using multiple microphones is addressed. The minimum mean square error (MMSE) estimator for the multispeaker case is derived and a novel decomposition of this estimator is presented. The MMSE estimator is decomposed into two stages: first, a multispeaker linearly constrained minimum variance (LCMV) beamformer (BF); and second, a subsequent multispeaker Wiener postfilter. The first stage separates and enhances the signals of the individual speakers by utilizing the spatial characteristics of the speakers [as manifested by the respective acoustic transfer functions (ATFs)] and the noise power spectral density (PSD) matrix, while the second stage exploits the speakers' PSD matrix to reduce the residual noise at the output of the first stage. The output vector of the multispeaker LCMV BF is proven to be the sufficient statistic for estimating the marginal speech signals in both the classic sense and the Bayesian sense. The log spectral amplitude estimator for the multispeaker case is also derived given the multispeaker LCMV BF outputs. The performance evaluation was conducted using measured ATFs and directional noise with various signal-to-noise ratio levels. It is empirically verified that the multispeaker postfilters are beneficial in terms of signal-to-interference plus noise ratio improvement when compared with the single-speaker postfilter. Ofer Schwartz, Sharon Gannot, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | Cramér-Rao Bound Analysis of Reverberation Level Estimators for Dereverberation and Noise ReductionabstractThe reverberation power spectral density (PSD) is often required for dereverberation and noise reduction algorithms. In this work, we compare two maximum likelihood (ML) estimators of the reverberation PSD in a noisy environment. In the first estimator, the direct path is first blocked. Then, the ML criterion for estimating the reverberation PSD is stated according to the probability density function of the blocking matrix (BM) outputs. In the second estimator, the speech component is not blocked. Instead, the ML criterion for estimating the speech and reverberation PSD is stated according to the probability density function of the microphone signals. To compare the expected mean square error (MSE) between the two ML estimators of the reverberation PSD, the Cramér-Rao Bounds (CRBs) for the two ML estimators are derived. We show that the CRB for the joint reverberation and speech PSD estimator is lower than the CRB for estimating the reverberation PSD from the BM outputs. Experimental results show that the MSE of the two estimators indeed obeys the CRB curves. Experimental results of multimicrophone dereverberation and noise reduction algorithm show the benefits of using the ML estimators in comparison with another baseline estimators. Ofer Schwartz, Sharon Gannot, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | Two Model-Based EM Algorithms for Blind Source Separation in Noisy EnvironmentsabstractThe problem of blind separation of speech signals in the presence of noise using multiple microphones is addressed. Blind estimation of the acoustic parameters and the individual source signals is carried out by applying the expectation-maximization (EM) algorithm. Two models for the speech signals are used, namely an unknown deterministic signal model and a complex-Gaussian signal model. For the two alternatives, we define a statistical model and develop EM-based algorithms to jointly estimate the acoustic parameters and the speech signals. The resulting algorithms are then compared from both theoretical and performance perspectives. In both cases, the latent data (differently defined for each alternative) are estimated in the E-step, where in the M-step, the two algorithms estimate the acoustic transfer functions of each source and the noise covariance matrix. The algorithms differ in the way the clean speech signals are used in the EM scheme. When the clean signal is assumed deterministic unknown, only the a posteriori probabilities of the presence of each source are estimated in the E-step, whereas their time-frequency coefficients are the parameters that are estimated in the M-step using the minimum variance distortionless response beamformer. If the clean speech signals are modeled as complex Gaussian signals, their power spectral densities are estimated in the E-step using the multichannel Wiener filter output. The proposed algorithms were tested using reverberant noisy mixtures of two speech sources in different reverberation and noise conditions. Boaz Schwartz, Sharon Gannot, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | Nonstationary Noise PSD Matrix Estimation for Multichannel Blind Speech ExtractionabstractNoise power spectral density (PSD) matrix estimation is one of the most important components of a multichannel blind speech extraction framework, as it largely determines the amount of residual noise at the output of a spatial filter. Optimality of well-known spatial filters, such as the multichannel Wiener filter, is only ensured if the PSD matrices of the noise and the desired speech are accurately estimated. In practical situations, where the noise is nonstationary, temporal averaging over time frames where the desired signal is inactive does not provide sufficiently fast tracking of the noise PSD matrix, resulting in high residual noise at the spatial filter output. Therefore, approaches that estimate the PSD matrices using narrowband signal detection have been proposed. Following the well-known single- and multichannel minima-controlled recursive averaging (MCRA) approaches, in this paper, we focus on narrowband speech presence probability-based noise PSD matrix estimators, which are suitable for blind scenarios where the location and the propagation vector of the desired speech source are unknown. The main contributions of the paper are a maximum likelihood interpretation of the multichannel MCRA, and a coherent-to-diffuse ratio-based a priori speech absence probability (SAP) estimator. The latter is a key parameter that determines the accuracy of the noise PSD matrix estimates in nonstationary scenarios. In this paper, we confirm the importance of the a priori SAP and show that its control is crucial for source extraction in nonstationary environments. Maja Taseska, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Conditional MMSE-based single-channel speech enhancement using inter-frame and inter-band correlationsabstractObtaining an estimate of clean speech for each time-frequency (TF) unit continues to be of importance in single-channel speech enhancement. Recently, it has been proposed to exploit inter-frame and interband correlations in a variety of speech processing applications. To estimate the clean speech, we propose in this contribution a conditional minimum mean squared error (MMSE)-based filter which exploits both inter-frame and inter-band correlations and takes into account the speech presence uncertainty. The speech presence uncertainty is provided by a recently proposed a posteriori speech presence probability (SPP) estimator that can also take into account the inter-frame and inter-band correlations. Simulation results demonstrate that the conditional MMSE-based filter in combination with the previously proposed SPP estimator and a fixed a priori SPP results in less distorted speech compared to the other SPP estimators. Hajar Momeni, Hamid Reza Abutalebi, Emanuël A. P. Habets |
ICASSP | 3 |
| 2016 | Joint maximum likelihood estimation of late reverberant and speech power spectral density in noisy environmentsabstractAn estimate of the power spectral density (PSD) of the late reverberation is often required by dereverberation algorithms. In this work, we derive a novel multichannel maximum likelihood (ML) estimator for the PSD of the reverberation that can be applied in noisy environments. Since the anechoic speech PSD is usually unknown in advance, it is estimated as well. As a closed-form solution for the maximum likelihood estimator is unavailable, a Newton method for maximizing the ML criterion is derived. Experimental results show that the proposed estimator provides an accurate estimate of the PSD, and outperforms competing estimators. Moreover, when used in a multi-microphone dereverberation and noise reduction algorithm, the best performance in terms of the log-spectral distance is achieved when employing the proposed PSD estimator. Ofer Schwartz, Sharon Gannot, Emanuël A. P. Habets |
ICASSP | 3 |
| 2016 | A low complexity weighted least squares narrowband DOA estimator for arbitrary array geometriesabstractAn increasing number of spatial filtering approaches requires narrowband direction-of-arrival (DOA) estimates. State-of-the-art (SOA) estimators such as root-MUSIC and ESPRIT are computationally complex and can be used only with specific array geometries. In this work, a low complexity DOA estimator is proposed that can be applied to arbitrary array geometries. The DOA is estimated by minimizing the weighted error between the observed and expected inter-microphone phase differences. The complexity of the proposed DOA estimator is significantly lower compared to that of the SOA estimators while providing a similar performance. Oliver Thiergart, Weilong Huang, Emanuël A. P. Habets |
ICASSP | 3 |
| 2016 | Insight into a phase modulation technique for signal decorrelation in multi-channel acoustic echo cancellationabstractHigh coherence between the loudspeaker signals in a multichannel communication set-up is known to be detrimental to the performance of a multi-channel acoustic echo cancellation (MC-AEC) system. The MC-AEC performance can be improved by decorrelating the loudspeaker signals prior to their reproduction. The decorrelation process can however degrade the subjective sound quality. A technique which has proven to provide a good trade-off between the MC-AEC performance enhancement and the subjective quality of the decorrelated signals, is applying a time-varying and perceptually motivated phase modulation in the sub-band domain. The aim of this paper is to provide further insight into this technique by analysing the influence on the signals' coherence as well as the MC-AEC performance in terms of its convergence speed. Maria Luis Valero, Emanuël A. P. Habets |
ICASSP | 2 |
| 2016 | Near-field signal acquisition for smartglasses using two acoustic vector-sensors
Dovid Levin, Emanuël A. P. Habets, Sharon Gannot |
Speech Commun. | 2 |
| 2016 | Online Dereverberation for Dynamic Scenarios Using a Kalman Filter With an Autoregressive ModelabstractReverberant signals can be modeled in the short-time Fourier transform domain using a frequency-dependent autoregressive (AR) model. In state-of-the-art, these AR coefficients have been considered stationary, which does not hold in time-varying environments. We propose to model these AR coefficients using a first-order Markov process, whereas the reverberant microphone signal observations are considered deterministic. This leads to a framework where the AR coefficients can be optimally estimated using a Kalman filter per subband. As a consequence, we can dereverberate the observed signals by applying the estimated AR coefficients as an adaptive linear filter per subband. Estimators for the required statistical parameters in the Kalman filter are derived. Due to the adaptive solution, the algorithm is suitable for real-time applications. It is shown that the proposed method outperforms an existing recursive least-squares solution in terms of reverberation reduction, convergence time, and tracking changes in the acoustic scene. Sebastian Braun, Emanuël A. P. Habets |
IEEE Signal Process. Lett. | 2 |
| 2016 | On the Numerical Instability of an LCMV Beamformer for a Uniform Linear ArrayabstractWe analyze the conditions for numerical instability in the solution of a linearly constrained minimum variance (LCMV) beamformer with multiple directional constraints for a uniform linear array. An analytic expression is presented to determine the frequencies (for broadband signals such as speech) where the inverse term in the solution of the LCMV beamformer does not exist. Simulation results and power patterns are provided to further illustrate the problem. In addition, we investigate and discuss possible solutions to the problem. Soumitro Chakrabarty, Emanuël A. P. Habets |
IEEE Signal Process. Lett. | 2 |
| 2016 | An Alternative Complexity Reduction Method for Partitioned-Block Frequency-Domain Adaptive FiltersabstractPartitioned-block frequency-domain adaptive filters provide lower algorithmic complexity than their time-domain counterparts. However, for certain applications, it is desirable to reduce complexity even further. This can be achieved by reducing the complexity of the constraint, or linearization, of the circular correlations. Existing methods are either suboptimal or are not sufficiently flexible in terms of the design parameters. Therefore, an alternative complexity reduction method is proposed, which achieves near-optimal performance and which, in addition, is flexible in terms of the design parameters, as these can be arbitrarily chosen. Maria Luis Valero, Edwin Mabande, Emanuël A. P. Habets |
IEEE Signal Process. Lett. | 3 |
| 2016 | An Expectation-Maximization Algorithm for Multimicrophone Speech Dereverberation and Noise Reduction With Coherence Matrix EstimationabstractIn speech communication systems, the microphone signals are degraded by reverberation and ambient noise. The reverberant speech can be separated into two components, namely, an early speech component that consists of the direct path and some early reflections and a late reverberant component that consists of all late reflections. In this paper, a novel algorithm to simultaneously suppress early reflections, late reverberation, and ambient noise is presented. The expectation-maximization (EM) algorithm is used to estimate the signals and spatial parameters of the early speech component and the late reverberation components. As a result, a spatially filtered version of the early speech component is estimated in the E-step. The power spectral density (PSD) of the anechoic speech, the relative early transfer functions, and the PSD matrix of the late reverberation are estimated in the M-step of the EM algorithm. The algorithm is evaluated using real room impulse response recorded in our acoustic lab with a reverberation time set to 0.36 s and 0.61 s and several signal-to-noise ratio levels. It is shown that significant improvement is obtained and that the proposed algorithm outperforms baseline single-channel and multichannel dereverberation algorithms, as well as a state-of-the-art multichannel dereverberation algorithm. Ofer Schwartz, Sharon Gannot, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Spotforming: Spatial Filtering With Distributed Arrays for Position-Selective Sound AcquisitionabstractHands-free capture of speech often requires extraction of sources from a certain spot of interest (SOI), while reducing interferers and background noise. Although state-of-the-art spatial filters are fully data-dependent and computed using the power spectral density (PSD) matrices of the desired and the undesired signals, the existing solutions to extract sources from a SOI are only partially data-dependent. Estimating the time-varying PSD matrices from the data is a challenging problem, especially in dynamic and quickly time-varying acoustic scenes. Hence, the spot signal statistics are often pre-computed based on a near-field propagation model, resulting in suboptimal filters. In this work, we propose a fully data-dependent spatial filtering framework for extraction of speech signals that originate from a SOI. To achieve position-based spatial selectivity, distributed arrays are used, which offer larger spatial diversity compared to arrays of closely spaced microphones. The PSD matrices of the desired and the undesired signals are updated at each time-frequency bin by using a minimum Bayes risk detector that is based on a probabilistic model of narrowband position estimates. The proposed framework is applicable in challenging multitalk situations, without requiring any prior information, except the geometry, location, and orientation of the arrays. Maja Taseska, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Residual noise control using a parametric multichannel Wiener filterabstractMultichannel noise reduction techniques are commonly used in speech communication applications. In these applications, it is often desired to maintain a residual amount of background noise to avoid perceptually unpleasant artifacts, such as musical tones or time periods of complete silence. Noise reduction can be achieved by the parametric multichannel Wiener filter (PMWF), which provides a trade-off between speech distortion and noise reduction. To additionally control the maximum noise reduction, the PMWF can be decomposed into a spatial filter and a spectral gain, which is limited to a desired minimum value. Such decomposition is however only possible if the desired source power spectral density matrix is rank-one, which in general does not even hold for a single source in reverberant environments. In the proposed approach, we define the desired signal as a sum of the speech signal plus the desired residual noise, and derive an optimum filter in the minimum mean-square error sense. The resulting filter has the advantage that it enables direct control of the maximum noise reduction without the need for a gain limiting step and is furthermore applicable to desired signals of higher rank. We analyze the derived filter thoroughly and show its relation to the standard PMWF that results as a special case. Furthermore, we propose a solution for keeping the residual noise level constant in slowly time-varying noise fields. Sebastian Braun, Konrad Kowalczyk, Emanuël A. P. Habets |
ICASSP | 3 |
| 2015 | A Bayesian approach to spatial filtering and diffuse power estimation for joint dereverberation and noise reductionabstractA spatial filter, with L linear constraints that are based on instantaneous narrowband direction-of-arrival (DOA) estimates, was recently proposed to obtain a desired spatial response for at most L sound sources. In noisy and reverberant environments, it becomes difficult to get reliable instantaneous DOA estimates and hence obtain the desired spatial response. In this work, we develop a Bayesian approach to spatial filtering that is more robust to DOA estimation errors. The resulting filter is a weighted sum of spatial filters pointed at a discrete set of DOAs, with the relative contribution of each filter determined by the posterior distribution of the discrete DOAs given the microphone signals. In addition, the proposed spatial filter is able to reduce both reverberation and noise. In this work, the required diffuse sound power is estimated using the posterior distribution of the discrete set of DOAs. Simulation results demonstrate the ability of the proposed filter to achieve strong suppression of the undesired signal components with small amount of signal distortion, in noisy and reverberant conditions. Soumitro Chakrabarty, Oliver Thiergart, Emanuël A. P. Habets |
ICASSP | 3 |
| 2015 | Nested generalized sidelobe canceller for joint dereverberation and noise reductionabstractSpeech signal is often contaminated by both room reverberation and ambient noise. In this contribution, we propose a nested generalized sidelobe canceller (GSC) beamforming structure, comprising an inner and an outer GSC beamformers (BFs), that decouple the speech dereverberation and the noise reduction operations. The BFs are implemented in the short-time Fourier transform (STFT) domain. Two alternative reverberation models are adopted. In the first, used in the inner GSC, reverberation is assumed to comprise a coherent early component and a late reverberant component. In the second, used in the outer GSC, the influence of the entire acoustic transfer function (ATF) is modeled as a convolution along the frame index in each frequency. Unlike other BF designs for this problem that must be updated in each time-frame, the proposed BF is time-invariant in static scenarios. Experiments with both simulated and recorded environments verify the effectiveness of the proposed structure. Ofer Schwartz, Sharon Gannot, Emanuël A. P. Habets |
ICASSP | 3 |
| 2015 | Minimum Bayes risk signal detection for speech enhancement based on a narrowband DOA modelabstractA desired speech signal in hands-free communication systems is often degraded by background noise and interferers. Data-dependent spatial filters for desired speech extraction depend on the power spectral density (PSD) matrices of the desired and the undesired signals, which are commonly estimated recursively using a signal model-based speech presence probability (SPP). The SPP and the PSD matrix estimates are only accurate, if the statistics of the undesired signals vary more slowly compared to the desired signal. In practical situations with competing talkers, this assumption is violated. To estimate the PSD matrices of highly non-stationary signals, we propose a minimum Bayes risk detector based on a model for the narrowband direction-of-arrival estimates. The performance of the proposed detector and the objective quality of the extracted desired speech are evaluated using simulated and measured data. Maja Taseska, Emanuël A. P. Habets |
ICASSP | 2 |
| 2015 | Direct-ambient decomposition using parametric wiener filtering with spatial cue controlabstractA method for decomposing audio signals into direct signals and ambient signals is described that can be applied to sound post-production and reproduction. The proposed method is based on a parametric multichannel Wiener filter (MWF) that enables a trade-off between the attenuation of the interfering signal and the distortion of the desired signal. We show that the MWF leads to distortions of the spatial cues of the ambient output signal, namely inter-channel correlation and inter-channel level difference. Our proposed solution is to control the trade-off parameter of the parametric MWF to ensure that these spatial distortions are inaudible. Christian Uhle, Emanuël A. P. Habets |
ICASSP | 2 |
| 2015 | A state-space partitioned-block adaptive filter for echo cancellation using inter-band correlations in the Kalman gain computationabstractA partitioned-block-based architecture for a model-based acoustic echo canceller in the frequency domain was recently presented. Partitioned-block-based frequency domain adaptive filters provide a lower algorithmic delay compared to the non-partitioned formulations, which is achieved by partitioning the acoustic echo path and hence shortening the time-frequency transforms. Under these circumstances, avoiding the linearisation of the involved circular convolutions can significantly impair the performance. This paper extends the already proposed diagonalization of the Kalman gain matrix by taking into account the neglected inter-band correlations to improve the performance of the acoustic echo canceller. This comes at the cost of a moderate increase in the algorithmic complexity. Maria Luis Valero, Edwin Mabande, Emanuël A. P. Habets |
ICASSP | 3 |
| 2015 | On the Average Directivity Factor Attainable With a Beamformer Incorporating Null ConstraintsabstractThe directivity factor (DF) of a beamformer describes its spatial selectivity and ability to suppress diffuse noise which arrives from all directions. For a given array constellation, it is possible to select beamforming weights which maximize the DF for a particular look-direction, while enforcing nulls for a set of undesired directions. In general, the resulting DF is dependent upon the specific look- and null directions. Using the same array, one may apply a different set of weights designed for any other feasible set of look- and null directions. In this contribution, we show that when the optimal DF is averaged over all look directions, the result equals the number of sensors minus the number of null constraints. This result holds regardless of the positions and spatial responses of the individual sensors and regardless of the null directions. The result generalizes to more complex wave-propagation domains (e.g., reverberation). Dovid Levin, Emanuël A. P. Habets, Sharon Gannot |
IEEE Signal Process. Lett. | 2 |
| 2015 | Multi-Microphone Speech Dereverberation and Noise Reduction Using Relative Early Transfer FunctionsabstractIn speech communication systems, the microphone signals are degraded by reverberation and ambient noise. The reverberant speech can be separated into two components, namely, an early speech component that includes the direct path and some early reflections, and a late reverberant component that includes all the late reflections. In this paper, a novel algorithm to simultaneously suppress early reflections, late reverberation and ambient noise is presented. A multi-microphone minimum mean square error estimator is used to obtain a spatially filtered version of the early speech component. The estimator constructed as a minimum variance distortionless response (MVDR) beamformer (BF) followed by a postfilter (PF). Three unique design features characterize the proposed method. First, the MVDR BF is implemented in a special structure, named the nonorthogonal generalized sidelobe canceller (NO-GSC). Compared with the more conventional orthogonal GSC structure, the new structure allows for a simpler implementation of the GSC blocks for various MVDR constraints. Second, In contrast to earlier works, RETFs are used in the MVDR criterion rather than either the entire RTFs or only the direct-path of the desired speech signal. An estimator of the RETFs is proposed as well. Third, the late reverberation and noise are processed by both the beamforming stage and the PF stage. Since the relative power of the noise and the late reverberation varies with the frame index, a computationally efficient method for the required matrix inversion is proposed to circumvent the cumbersome mathematical operation. The algorithm was evaluated and compared with two alternative multichannel algorithms and one single-channel algorithm using simulated data and data recorded in a room with a reverberation time of 0.5 s for various source-microphone array distances (1-4 m) and several signal-to-noise levels. The processed signals were tested using two commonly used objective measures, namely perceptual evaluation of speech quality and log-spectral distance. As an additional objective measure, the improvement in word accuracy percentage of an acoustic speech recognition system is also demonstrated. Ofer Schwartz, Sharon Gannot, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | Online Speech Dereverberation Using Kalman Filter and EM AlgorithmabstractSpeech signals recorded in a room are commonly degraded by reverberation. In most cases, both the speech signal and the acoustic system of the room are unknown and time-varying. In this paper, a scenario with a single desired sound source and slowly time-varying and spatially-white noise is considered, and a multi-microphone algorithm that simultaneously estimates the clean speech signal and the time-varying acoustic system is proposed. The recursive expectation-maximization scheme is employed to obtain both the clean speech signal and the acoustic system in an online manner. In the expectation step, the Kalman filter is applied to extract a new sample of the clean signal, and in the maximization step, the system estimate is updated according to the output of the Kalman filter. Experimental results show that the proposed method is able to significantly reduce reverberation and increase the speech quality. Moreover, the tracking ability of the algorithm was validated in practical scenarios using human speakers moving in a natural manner. Boaz Schwartz, Sharon Gannot, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | Down-mixing using coherence suppressionabstractA common problem in audio signal processing is to mix two or more signals into one sum signal. The mixing procedure, known as down-mixing, usually introduces some signal impairments, especially if two signals contain similar but phase shifted signal components. Summing up such signals results in severe comb-filter artifacts. In this paper, we propose a novel down-mix method which prevents comb-filter effects. This is achieved by suppressing the coherent signal parts of one input signal prior to mixing. During the down-mixing, a scaling gain is applied ensuring the preservation of overall signal energy. Furthermore, a phase-align extension to the down-mixer is introduced. The proposed down-mixer is evaluated qualitatively in comparison with two existing down-mix approaches as well as quantitatively by means of a distortion analysis. Alexander Adami, Emanuël A. P. Habets, Jürgen Herre |
ICASSP | 2 |
| 2014 | Automatic spatial gain control for an informed spatial filterabstractWhen capturing speech in a multi-talker telecommunication scenario, it is desirable to keep the enhanced signal at an equal loudness level for each speaker. Single-channel automatic gain control systems are not able to adjust the level of different talkers when they are simultaneously active. In this work, an automatic spatial gain control (ASGC) algorithm is proposed that adjusts the directional response of an existing informed spatial filter such that the direct sound of multiple sources can be kept at a constant desired loudness level at the output. The spatial filter additionally reduces diffuse sound and ambient noise. It is shown that the proposed AGSC works well within the tested scenario, and is able to adjust the levels of different speakers even during double talk scenarios. Sebastian Braun, Oliver Thiergart, Emanuël A. P. Habets |
ICASSP | 3 |
| 2014 | Extended Kalman filter with probabilistic data association for multiple non-concurrent speaker localization in reverberant environmentsabstractAcoustic source localization and tracking (ASLT) in reverberant environments is a challenging task due to the multi-path propagation of acoustic waves. ASLT is often based on the use of a Kalman filter or a particle filter, with time-difference-of-arrival (TDOA) estimates used as measurements. In this work, we aim to track non-concurrent speakers by applying an extended Kalman filter (EKF) with probabilistic data association (PDA) that takes into account multiple measurements simultaneously. By using PDA, the inaccuracy of the measurements caused by room reflections and noise is explicitly taken into account. Unlike in typical approaches where the measurements consist of broadband TDOA estimates, the measurements in the proposed approach consist of multiple narrowband direction-of-arrival (DOA) estimates obtained from distributed microphone arrays. Experimental results demonstrate that incorporating PDA and using properly selected narrowband DOA estimates leads to a better tracking performance, as compared to the standard EKF with a single narrowband or broadband measurement. Soumitro Chakrabarty, Konrad Kowalczyk, Maja Taseska, Emanuël A. P. Habets |
ICASSP | 4 |
| 2014 | The influence of low order reflections on the interaural time differences in crosstalk cancellation systemsabstractReflections can play an important role in human perception of sound. While they can positively contribute to the perceived sound quality, they may also interfere with the reproduction of, for example, crosstalk cancelled binaural sounds through loudspeakers. In this paper, we study the influence of 1st and 2nd order reflections on a crosstalk-cancelled desktop reproduction system, through an analysis of the interaural time differences. The direct and reflected sounds are calculated using an image-source model. Meanwhile, the crosstalk cancellation filters are calculated assuming anechoic conditions, and therefore proper sound reproduction is now questionable. In this scenario, the reflections are found to introduce changes in the interaural phase differences and in the interaural group delay. These changes are analyzed and the possible effect on sound localization is investigated using a subjective localization experiment. The results indicated that for the studied setup, the localization accuracy was, practically, unaffected by the low order reflections. Dimitrios Kosmidis, Yesenia Lacouture-Parodi, Emanuël A. P. Habets |
ICASSP | 3 |
| 2014 | Single-channel speech presence probability estimation using inter-frame and inter-band correlationsabstractThe speech presence probability (SPP) plays an important role in many noise reduction and noise estimation methods. The SPP is commonly computed per time and frequency in the short time Fourier transform (STFT) domain based on the a priori speech absence probability and the a priori and a posteriori signal-to-noise ratios. Due to the STFT as well as the nature of the speech signal, there exists a correlation between subsequent time frames and neighboring frequency bands. In this work, we explicitly take these inter-frame and inter-band correlations into account when computing the SPP. The presented results demonstrate that we can increase the detection accuracy of the SPP estimator by taking a few neighboring time and frequency bins into account. Hajar Momeni, Emanuël A. P. Habets, Hamid Reza Abutalebi |
ICASSP | 2 |
| 2014 | Power-based signal-to-diffuse ratio estimation using noisy directional microphonesabstractThe signal-to-diffuse ratio (SDR), which describes the power ratio between the direct and diffuse component of a sound field, is an important parameter in many applications. This paper proposes a power-based SDR estimator which considers the auto power spectral densities obtained by noisy directional microphones. Compared to recently proposed estimators that exploit the spatial coherence between two microphones, the power-based estimator is more robust at lower frequencies given that the microphone directivities are known with sufficiently high accuracy. The proposed estimator can incorporate more than two microphones and can therefore provide accurate SDR estimates independently of the direction-of-arrival of the direct sound. We further propose a method to determine the optimal microphone orientations for a given set of directional microphones. Simulations show the practical applicability. Oliver Thiergart, Tobias Ascherl, Emanuël A. P. Habets |
ICASSP | 3 |
| 2014 | Signal-based late residual echo spectral variance estimationabstractA commonly used technique to attenuate acoustic echo signals in hands-free devices is acoustic echo cancellation (AEC). In practice, the AEC is unable to provide sufficient echo reduction due to system misalignment, modeling errors and the insufficient length of the estimated acoustic echo path. Hence, AEC is usually used in conjunction with one or more postfilters to suppress the remaining echo. For the estimation of the postfilter the residual echo is required. In the past, several approaches have been proposed which estimate the late residual echo spectral variance using the far-end signal and parameters that are obtained from the estimated echo path. In this work, we propose a signal-based algorithm to estimate the parameters. Thus, among other advantages, the postfilter can also be used in combination with acoustic echo suppression. Maria Luis Valero, Edwin Mabande, Emanuël A. P. Habets |
ICASSP | 3 |
| 2014 | Extracting Reverberant Sound Using a Linearly Constrained Minimum Variance Spatial FilterabstractMicrophone arrays are typically used to extract the direct sound of sound sources while suppressing noise and reverberation. Applications such as immersive spatial sound reproduction commonly also require an estimate of the reverberant sound. A linearly constrained minimum variance filter, of which one of the constraints is related to the spatial coherence of the assumed reverberant sound field, is proposed to obtain an estimate of the sound pressure of the reverberant field. The proposed spatial filter provides an almost omnidirectional directivity pattern with spatial nulls for the directions-of-arrival of the direct sound. The filter is computationally efficient and outperforms existing methods. Oliver Thiergart, Emanuël A. P. Habets |
IEEE Signal Process. Lett. | 2 |
| 2014 | Noise Reduction in the Spherical Harmonic Domain Using a Tradeoff Beamformer and Narrowband DOA EstimatesabstractIn noise reduction, a common approach is to use a microphone array with a beamformer that combines the individual microphone signals to extract a desired speech signal. The beamformer weights usually depend on the statistics of the noise and desired speech signals, which cannot be directly observed and must be estimated. Estimators based on the speech presence probability (SPP) seek to update the statistics estimates only when desired speech is known to be absent or present. However, they do not normally distinguish between desired and undesired speech sources. In this contribution, an algorithm is proposed to distinguish between these two types of sources using additional spatial information, by estimating a desired speech presence probability based on the combination of a multichannel SPP and a direction of arrival (DOA) based probability. The DOA-based probability is computed using DOA estimates for each time-frequency bin. The estimated statistics are then used to compute the weights of a spherical harmonic domain tradeoff beamformer, which achieves a balance between noise reduction and speech distortion. The performance evaluation demonstrates the effectiveness of the proposed approach at suppressing both background noise and spatially coherent noise. A number of audio examples and sample spectrograms are also provided. Daniel P. Jarrett, Maja Taseska, Emanuël A. P. Habets, Patrick A. Naylor |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | Multichannel Noise Reduction in the Karhunen-Loève Expansion DomainabstractThe noise reduction problem is traditionally approached in the time, frequency, or transform domain. Having a signal dependent transform has shown some advantages over the traditional signal independent transform. Recently, the single-channel noise reduction problem in the Karhunen-Loève expansion (KLE) domain has received special attention. In this paper, the noise reduction problem in the KLE domain is studied from a multichannel perspective. We present a new formulation of the problem, in which inter-channel and inter-mode correlations are optimally exploited. We derive different optimal noise reduction filters and present a set of useful performance measures within this framework. The performance of the different filters is then evaluated through experiments in which not only noise but also competing speech sources are present. It is shown that the proposed multichannel formulation is more robust to competing speech sources than the single-channel approach and that a better compromise between noise reduction and speech distortion can be obtained. Yesenia Lacouture-Parodi, Emanuël A. P. Habets, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Robust Multichannel Dereverberation using Relaxed Multichannel Least SquaresabstractA novel approach is proposed for robust multichannel dereverberation in the presence of system identification error (SIEs), based on channel shortening. A mathematical link is derived between the well known multiple-input/output inverse theorem (MINT) algorithm and channel shortening. The relaxed multichannel least squares (RMCLS) algorithm is then proposed as an efficient realization within the channel shortening paradigm and is shown through experimental results to outperform MINT in the presence of SIEs. While the RMCLS is robust to SIEs, the coloration of the output cannot be controlled. Two extensions to RMCLS are proposed to control the level of coloration and the performances of both extensions are evaluated comparatively. It is shown that both substantially maintain the dereverberation performance and robustness to SIEs obtained from RMCLS while effectively controlling the level of coloration introduced. Felicia Lim, Wancheng Zhang, Emanuël A. P. Habets, Patrick A. Naylor |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | Informed Spatial Filtering for Sound Extraction Using Distributed Microphone ArraysabstractHands-free acquisition of speech is required in many human-machine interfaces and communication systems. The signals received by integrated microphones contain a desired speech signal, spatially coherent interfering signals, and background noise. In order to enhance the desired speech signal, state-of-the-art techniques apply data-dependent spatial filters which require the second order statistics (SOS) of the desired signal, the interfering signals and the background noise. As the number of sources and the reverberation time increase, the estimation accuracy of the SOS deteriorates, often resulting in insufficient noise and interference reduction. In this paper, a signal extraction framework with distributed microphone arrays is developed. An expectation maximization (EM)-based algorithm detects the number of coherent speech sources and estimates source clusters using time-frequency (TF) bin-wise position estimates. Subsequently, the second order statistics (SOS) are estimated using bin-wise speech presence probability (SPP) and a source probability for each source. Finally, a desired source is extracted using a minimum variance distortionless response (MVDR) filter, a multichannel Wiener filter (MWF) and a parametric multichannel Wiener filter (PMWF). The same framework can be employed for source separation, where a spatial filter is computed for each source considering the remaining sources as interferers. Evaluation using simulated and measured data demonstrates the effectiveness of the framework in estimating the number of sources, clustering, signal enhancement, and source separation. Maja Taseska, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | An informed parametric spatial filter based on instantaneous direction-of-arrival estimatesabstractExtracting desired source signals in noisy and reverberant environments is required in many hands-free communication systems. In practical situations, where the position and number of active sources may be unknown and time-varying, conventional implementations of spatial filters do not provide sufficiently good performance. Recently, informed spatial filters have been introduced that incorporate almost instantaneous parametric information on the sound field, thereby enabling adaptation to new acoustic conditions and moving sources. In this contribution, we propose a spatial filter which generalizes the recently proposed informed linearly constrained minimum variance filter and informed minimum mean square error filter. The proposed filter uses multiple direction-of-arrival estimates and second-order statistics of the noise and diffuse sound. To determine those statistics, an optimal diffuse power estimator is proposed that outperforms state-of-the-art estimators. Extensive performance evaluation demonstrates the effectiveness of the proposed filter in dynamic acoustic conditions. For this purpose, we have considered a challenging scenario which consists of quickly moving sound sources during double-talk. The performance of the proposed spatial filter was evaluated in terms of objective measures including segmental signal-to-reverberation ratio and log spectral distance, and by means of a listening test confirming the objective results. Oliver Thiergart, Maja Taseska, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2013 | An informed spatial filter for dereverberation in the spherical harmonic domainabstractIn speech communication systems the received microphone signals are commonly degraded by reverberation and ambient noise that can decrease the fidelity and intelligibility of a desired speaker. Reverberation can be modeled as non-stationary diffuse sound which is not directly observable. In this work, we derive a multichannel Wiener filter in the spherical harmonic domain to reduce both reverberation and noise. The filter depends on the direction-of-arrival of the direct sound of the desired speaker and an interference power spectral density matrix for which an estimator is developed. The resulting informed spatial filter incorporates instantaneous information about the diffuseness of the sound field into the design of the filter. In addition, it is shown how the proposed filter relates to the well-known robust minimum variance distortionless response filter that is also used for comparison in the evaluation. Experimental results show that the proposed spatial filter provides a tradeoff between noise reduction and dereverberation depending on the diffuse sound PSD. Sebastian Braun, Daniel P. Jarrett, Johannes Fischer 0002, Emanuël A. P. Habets |
ICASSP | 4 |
| 2013 | Multichannel object-based audio coding with controllable qualityabstractIn this paper a new multichannel object-based audio coding scheme with scalable signal quality is proposed. The novel scheme is based on controlled downmixing and demixing. By means of a dedicated control mechanism, a number of distinct audio objects are mixed into a lower number of channels. The latter is chosen such that the desired quality level is met after demixing. The quality is assessed with two new psychoacoustically motivated metrics. Following the informed source separation approach, the downmix is decomposed via optimum spatial filtering guided by short-time power spectral densities of the audio objects. In an experiment it is shown that the raw data rate of an exemplary 10-track recording can be reduced by at least 30 % using linear pulse-code modulation while maintaining perceptual transparency. Stanislaw Gorlow, Emanuël A. P. Habets, Sylvain Marchand |
ICASSP | 2 |
| 2013 | Spherical harmonic domain noise reduction using an MVDR beamformer and DOA-based second-order statistics estimationabstractMost beamformers used for noise reduction rely on the accurate estimation of the second-order statistics of the noise, and in some cases, of the desired signal. Speech presence probability (SPP) based statistics estimators seek to update the estimates only when speech is absent/present, however, when used with a fixed a priori SPP, they cannot distinguish between a coherent desired source and coherent noise sources. We propose to distinguish between desired and noise sources by estimating the second-order statistics with a direction of arrival dependent a priori SPP, which we then use to compute the weights of a spherical harmonic domain minimum variance distortionless response filter. Daniel P. Jarrett, Emanuël A. P. Habets, Patrick A. Naylor |
ICASSP | 2 |
| 2013 | Application of particle filtering to an interaural time difference based head tracker for crosstalk cancellationabstractCrosstalk cancellation systems (CCSs) suffer from a rather narrow sweet spot and small head rotations introduce significant interaural time difference (ITD) errors at the ears of the listener that destroy the 3D experience. In a previous study, we proposed a CCS with a microphone-based head tracker that uses two microphones placed closed to the ears of the listener. The head orientation was estimated by minimizing the difference between the ITD of desired binaural signals and the ITD of the microphones signals. We present here an extension of the previously proposed system in which the tracking of the ITD error is performed using a particle filtering (PF) approach that allows to take into account the dynamics of the human head. The head dynamics are modelled as a Ornstein-Uhlenbeck process, which shows to be a closer approximation of the natural movements of the head. Experimental results shows that with the proposed PF approach, head rotations can be accurately tracked in the presence of noise and when multiple virtual sound sources are reproduced simultaneously. Yesenia Lacouture-Parodi, Emanuël A. P. Habets |
ICASSP | 2 |
| 2013 | Robust beamforming using sensors with nonidentical directivity patternsabstractThe optimal weights for a beamformer that provide maximum directivity, are often found to be severely lacking in terms of robustness. Although an ideal implementation of the beamformer with these weights provides high directivity, minor perturbations of the weights or of sensor placement cause severe degradation. Therefore, a robustness constraint is often imposed during the beamformer's design stage. The classical method of diagonal loading is commonly used for this purpose. There are known results in this field which pertain to an array consisting of sensors with identical directivity-patterns and orientations. We extend these results to account for sensors with nonidentical directivity patterns, and sensors which share placement errors. We show that in such cases, modification of the classical loading scheme to incorporate nonidentical diagonal elements and off-diagonal elements is beneficial. Dovid Levin, Emanuël A. P. Habets, Sharon Gannot |
ICASSP | 2 |
| 2013 | Blind reverberation time estimation by intrinsic modeling of reverberant speechabstractThe reverberation time (RT) is a very important measure that quantifies the acoustic properties of a room and provides information about the quality and intelligibility of speech recorded in that room. Moreover, information about the RT can be used to improve the performance of automatic speech recognition systems and speech dereverberation algorithms. In a recent study, it has been shown that existing methods for blind estimation of the RT are highly sensitive to additive noise. In this paper, a novel method is proposed to blindly estimate the RT based on the decay rate distribution. Firstly, a data-driven representation of the underlying decay rates of several training rooms is obtained via the eigenvalue decomposition of a specially-tailored kernel. Secondly, the representation is extended to a room under test and used to estimate its decay rate (and hence its RT). The presented results show that the proposed method outperforms a competing method and is significantly more robust to noise. Ronen Talmon, Emanuël A. P. Habets |
ICASSP | 2 |
| 2013 | MMSE-based source extraction using position-based posterior probabilitiesabstractA scenario with multiple talkers and additive background noise is considered, where some talkers are active simultaneously and the activity of the talkers changes with time. We propose an MMSE-based method to blindly extract any talker using bin-wise position estimates obtained from distributed microphone arrays. In order to distinguish between different talkers, the position estimates are clustered using the expectation maximization algorithm. The resulting posterior probabilities allow to estimate the PSD matrices of the talkers and compute an MMSE-optimal linear filter for extracting each talker. We evaluate the performance of the proposed method in terms of noise and interference reduction and distortion of the desired speech signal at the output of a multichannel Wiener filter. Maja Taseska, Emanuël A. P. Habets |
ICASSP | 2 |
| 2013 | An informed LCMV filter based on multiple instantaneous direction-of-arrival estimatesabstractExtracting sound sources in noisy and reverberant conditions remains a challenging task that is commonly found in modern communication systems. In this work, we consider the problem of obtaining a desired spatial response for at most L simultaneously active sound sources. The proposed spatial filter is obtained by minimizing the diffuse plus self-noise power at the output of the filter subject to L linear constraints. In contrast to earlier works, the L constraints are based on instantaneous narrowband direction-of-arrival estimates. In addition, a novel estimator for the diffuse-to-noise ratio is developed that exhibits a sufficiently high temporal and spectral resolution to achieve both dereverberation and noise reduction. The presented results demonstrate that an optimal tradeoff between maximum white noise gain and maximum directivity is achieved. Oliver Thiergart, Emanuël A. P. Habets |
ICASSP | 2 |
| 2013 | Blind System Identification Using Sparse Learning for TDOA Estimation of Room ReflectionsabstractLocalization of early room reflections can be achieved by estimating the time-differences-of-arrival (TDOAs) of reflected waves between elements of a microphone array. For an unknown source, we propose to apply sparse blind system identification (BSI) methods to identify the acoustic impulse responses, from which the TDOAs of temporally sparse reflections are estimated. The proposed time- and frequency-domain adaptive algorithms based on crossrelation formulation are regularized by incorporating an l1-norm sparseness constraint, which is realized using a split Bregman method. These algorithms are shown to outperform standard crossrelation-based BSI techniques when estimating TDOAs of reflections in the presence of background noise. Konrad Kowalczyk, Emanuël A. P. Habets, Walter Kellermann, Patrick A. Naylor |
IEEE Signal Process. Lett. | 2 |
| 2013 | A Generalized Theorem on the Average Array Directivity FactorabstractThe beampattern of an array consisting of$N$elements is determined by the beampatterns of the individual elements, their placement, and the weights assigned to them. For each look direction, it is possible to design weights that maximize the array directivity factor (DF). For the case of an array of omnidirectional elements using optimal weights, it has been shown that the average DF over all look directions equals the number of elements. The validity of this theorem is not dependent on array geometry. We generalize this theorem by means of an alternative proof. The chief contributions of this letter are a) a compact and direct proof, b) generalization to arrays containing directional elements (such as cardioids and dipoles), and c) generalization to arbitrary wave propagation models. A discussion of the theorem's ramifications on array processing is provided. Dovid Levin, Emanuël A. P. Habets, Sharon Gannot |
IEEE Signal Process. Lett. | 2 |
| 2013 | A Two-Stage Beamforming Approach for Noise Reduction and DereverberationabstractIn general, the signal-to-noise ratio as well as the signal-to-reverberation ratio of speech received by a microphone decrease when the distance between the talker and microphone increases. Dereverberation and noise reduction algorithm are essential for many applications such as videoconferencing, hearing aids, and automatic speech recognition to improve the quality and intelligibility of the received desired speech that is corrupted by reverberation and noise. In the last decade, researchers have aimed at estimating the reverberant desired speech signal as received by one of the microphones. Although this approach has let to practical noise reduction algorithms, the spatial diversity of the received desired signal is not exploited to dereverberate the speech signal. In this paper, a two-stage beamforming approach is presented for dereverberation and noise reduction. In the first stage, a signal-independent beamformer is used to generate a reference signal which contains a dereverberated version of the desired speech signal as received at the microphones and residual noise. In the second stage, the filtered microphone signals and the noisy reference signal are used to obtain an estimate of the dereverberated desired speech signal. In this stage, different signal-dependent beamformers can be used depending on the desired operating point in terms of noise reduction and speech distortion. The presented performance evaluation demonstrates the effectiveness of the proposed two-stage approach. Emanuël A. P. Habets, Jacob Benesty |
IEEE Trans. Speech Audio Process. | 1 |
| 2013 | Multi-Microphone Noise Reduction Based on Orthogonal Noise Signal DecompositionsabstractMulti-microphone noise reduction plays an increasing and important role in acoustic communication systems. Existing multichannel noise reduction filters are commonly computed based on a single noise covariance matrix. Recently, an orthogonal noise signal decomposition was proposed that uses a single noise signal as a reference. Using this decomposition, it was possible to reformulate the noise reduction problem and derived a multichannel noise reduction filter that allows a tradeoff between the noise that is coherent and incoherent with respect to the reference signal. In this contribution, we analyze the previously proposed decomposition and propose an orthogonal decomposition that is based on a rank-one projection of all noise signals. The projection is chosen such that the total variance of the coherent noise component is maximized. To further improve the separation between coherent noise and incoherent noise, a rank-Qprojection of the observed noise signals is proposed. The decomposed noise covariance matrix is then used to derive a minimum variance distortionless response beamformer that allows a tradeoff between coherent and incoherent noise reduction, and to form a constraint matrix for a linearly constrained minimum variance beamformer. The results of the performance evaluation demonstrate the advantage of the proposed decompositions over the previously proposed decomposition. Emanuël A. P. Habets, Jacob Benesty |
IEEE Trans. Speech Audio Process. | 1 |
| 2013 | Geometry-Based Spatial Sound Acquisition Using Distributed Microphone ArraysabstractTraditional spatial sound acquisition aims at capturing a sound field with multiple microphones such that at the reproduction side a listener can perceive the sound image as it was at the recording location. Standard techniques for spatial sound acquisition usually use spaced omnidirectional microphones or coincident directional microphones. Alternatively, microphone arrays and spatial filters can be used to capture the sound field. From a geometric point of view, the perspective of the sound field is fixed when using such techniques. In this paper, a geometry-based spatial sound acquisition technique is proposed to compute virtual microphone signals that manifest a different perspective of the sound field. The proposed technique uses a parametric sound field model that is formulated in the time-frequency domain. It is assumed that each time-frequency instant of a microphone signal can be decomposed into one direct and one diffuse sound component. It is further assumed that the direct component is the response of a single isotropic point-like source (IPLS) of which the position is estimated for each time-frequency instant using distributed microphone arrays. Given the sound components and the position of the IPLS, it is possible to synthesize a signal that corresponds to a virtual microphone at an arbitrary position and with an arbitrary pick-up pattern. Oliver Thiergart, Giovanni Del Galdo, Maja Taseska, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2012 | Multi-microphone noise reduction using interchannel and interframe correlationsabstractMulti-microphone noise reduction methods often operate in the time-frequency domain in which a complex gain is applied to each time-frame and subband. These methods can achieve good noise reduction with little speech distortion by exploiting the fact that the desired signal is correlated across the channels. In the context of single-microphone noise reduction, it has been shown recently that the performance in terms of noise reduction and speech distortion can be improved by exploiting the correlation between subsequent time-frames, i.e., by exploiting the interframe correlation. In this paper, we exploit both interchannel and interframe correlations in the context of multi-microphone noise reduction. Now the interframe correlation is taken into account, i.e., a filter is applied in each subband and channel instead of just a gain. The results of our experimental study show that we can improve the fullband signal-to-noise ratios (SNRs) by using interchannel and interframe correlations when dealing with signals, such as speech, that exhibit a sufficiently large interframe correlation. Emanuël A. P. Habets, Jacob Benesty, Jingdong Chen |
ICASSP | 1 |
| 2012 | Multichannel noise reduction wiener filter in the Karhunen-Loève expansion domainabstractThis paper explores the noise reduction problem in the Karhunen-Loève expansion (KLE) domain from a multichannel perspective. Based on formulations proposed for the design of optimal single-channel noise reduction in the KLE domain, we formulate the multichannel noise reduction in the KLE domain. Two different performance measures are presented: the noise reduction and speech distortion. The optimal multichannel Wiener filter is derived and its performance in terms of noise reduction and speech distortion is compared with the performance of the optimal single-channel Wiener filter. Experimental results show that a significant improvement in performance is obtained when using multiple microphone signals. The multichannel Wiener filter results also in better noise reduction in the presence of coherent noise sources. Yesenia Lacouture-Parodi, Emanuël A. P. Habets, Jacob Benesty |
ICASSP | 2 |
| 2012 | On the application of reverberation suppression to robust speech recognitionabstractIn this paper, we study the effect of the design parameters of a single-channel reverberation suppression algorithm on reverberation-robust speech recognition. At the same time, reverberation compensation at the speech recognizer is investigated. The analysis reveals that it is highly beneficial to attenuate only the reverberation tail after approximately 50 ms while coping with the early reflections and residual late-reverberation by training the recognizer on moderately reverberant data. It will be shown that the overall system at its optimum configuration yields a very promising recognition performance even in strongly reverberant environments. Since the reverberation suppression algorithm is evidenced to significantly reduce the dependency on the training data, it allows for a very efficient training of acoustic models that are suitable for a wide range of reverberation conditions. Finally, experiments with an “ideal” reverberation suppression algorithm are carried out to cross-check the inferred guidelines. Roland Maas, Emanuël A. P. Habets, Armin Sehr, Walter Kellermann |
ICASSP | 2 |
| 2012 | Signal-to-reverberant ratio estimation based on the complex spatial coherence between omnidirectional microphonesabstractThe signal-to-reverberant ratio (SRR) is an important parameter in several applications such as speech enhancement, dereverberation, and parametric spatial audio coding. In this contribution, an SRR estimator is derived from the direction-of-arrival dependent complex spatial coherence function computed via two omnidirectional microphones. It is shown that by employing a computationally inexpensive DOA estimator, the proposed SRR estimator outperforms existing approaches. Oliver Thiergart, Giovanni Del Galdo, Emanuël A. P. Habets |
ICASSP | 3 |
| 2012 | An insight into common filtering in noisy SIMO blind system identificationabstractThe effect of additive sensor noise on single-input-multiple-output (SIMO) blind system identification (BSI) algorithms based upon cross-relation (CR) error is investigated. Previous studies have shown that additive noise in the observed signal results in systems comprising the true estimated channels convolved with an erroneous ‘common filter’, and additionally that identification and removal of this filter significantly improves estimation error. However, the source of the common filter remained an open question. This paper explains the common filter through a first-order perturbation analysis of the CR matrix, showing that it be estimated from the perturbation and the eigenvectors of the noiseless CR matrix. The analysis given in this paper provides a new insight into the effect of noise on SIMO BSI algorithms and forms the first step towards an overall noise robust solution. Mark R. P. Thomas, Nikolay D. Gaubitch, Emanuël A. P. Habets, Patrick A. Naylor |
ICASSP | 3 |
| 2012 | On the Noise Reduction Performance of a Spherical Harmonic Domain Tradeoff BeamformerabstractIn this letter, we derive an expression for the expected incoherent noise reduction factor of a spherical harmonic domain (SHD) tradeoff beamformer. The tradeoff beamformer attempts to reduce noise while minimizing speech distortion, and includes the minimum variance distortionless response (MVDR) and multichannel Wiener filters as special cases. For the open spherical microphone array, we find a number of analogies between the expressions for the expected noise reduction factor of the SHD and spatial domain MVDR beamformers. In an anechoic environment we find that the performance of the SHD MVDR beamformer with an open array depends almost entirely on the number of microphones, as in the spatial domain. Daniel P. Jarrett, Emanuël A. P. Habets |
IEEE Signal Process. Lett. | 2 |
| 2012 | Inference of Room Geometry From Acoustic Impulse ResponsesabstractAcoustic scene reconstruction is a process that aims to infer characteristics of the environment from acoustic measurements. We investigate the problem of locating planar reflectors in rooms, such as walls and furniture, from signals obtained using distributed microphones. Specifically, localization of multiple two- dimensional (2-D) reflectors is achieved by estimation of the time of arrival (TOA) of reflected signals by analysis of acoustic impulse responses (AIRs). The estimated TOAs are converted into elliptical constraints about the location of the line reflector, which is then localized by combining multiple constraints. When multiple walls are present in the acoustic scene, an ambiguity problem arises, which we show can be addressed using the Hough transform. Additionally, the Hough transform significantly improves the robustness of the estimation for noisy measurements. The proposed approach is evaluated using simulated rooms under a variety of different controlled conditions where the floor and ceiling are perfectly absorbing. Results using AIRs measured in a real environment are also given. Additionally, results showing the robustness to additive noise in the TOA information are presented, with particular reference to the improvement achieved through the use of the Hough transform. Fabio Antonacci, Jason Filos, Mark R. P. Thomas, Emanuël A. P. Habets, Augusto Sarti, Patrick A. Naylor, Stefano Tubaro |
IEEE Trans. Speech Audio Process. | 4 |
| 2012 | A Perspective on Frequency-Domain Beamformers in Room AcousticsabstractSignals captured by a set of microphones in a speech communication system are mixtures of desired signals and noise. In this paper, a different perspective on frequency-domain beamformers in room acoustics is provided. Specifically, the observed noise signals are divided into coherent and incoherent signal components while no assumptions are being made regarding the number of coherent noise sources and the noise sound field. From this perspective, performance measures are defined and existing beamformers are deduced. In addition, a new and general tradeoff beamformer is proposed that enables a compromise between noise reduction and speech distortion on the one hand, and coherent noise versus incoherent noise reductions on the other hand. The presented performance evaluation shows how existing beamformers and the tradeoff beamformer perform in different scenarios. Emanuël A. P. Habets, Jacob Benesty |
IEEE Trans. Speech Audio Process. | 1 |
| 2012 | A Speech Distortion and Interference Rejection Constraint BeamformerabstractSignals captured by a set of microphones in a speech communication system are mixtures of desired and undesired signals and ambient noise. Existing beamformers can be divided into those that preserve or distort the desired signal. Beamformers that preserve the desired signal are, for example, the linearly constrained minimum variance (LCMV) beamformer that is supposed, ideally, to reject the undesired signal and reduce the ambient noise power, and the minimum variance distortionless response (MVDR) beamformer that reduces the interference-plus-noise power. The multichannel Wiener filter, on the other hand, reduces the interference-plus-noise power without preserving the desired signal. In this paper, a speech distortion and interference rejection constraint (SDIRC) beamformer is derived that minimizes the ambient noise power subject to specific constraints that allow a tradeoff between speech distortion and interference-plus-noise reduction on the one hand, and undesired signal and ambient noise reductions on the other hand. Closed-form expressions for the performance measures of the SDIRC beamformer are derived and the relations to the aforementioned beamformers are derived. The performance evaluation demonstrates the tradeoffs that can be made using the SDIRC beamformer. Emanuël A. P. Habets, Jacob Benesty, Patrick A. Naylor |
IEEE Trans. Speech Audio Process. | 1 |
| 2011 | Simulating room impulse responses for spherical microphone arraysabstractA method is proposed for simulating the sound pressure signals on a spherical microphone array in a reverberant enclosure. The method employs spherical harmonic decomposition and takes into account scattering from a solid sphere. An analysis shows that the error in the decomposition can be made arbitrarily small given a sufficient number of spherical harmonics. Daniel P. Jarrett, Emanuël A. P. Habets, Mark R. P. Thomas, Patrick A. Naylor |
ICASSP | 2 |
| 2011 | Direction-of-arrival estimation using acoustic vector sensors in the presence of noiseabstractA vector-sensor consisting of a monopole sensor collocated with orthogonally oriented dipole sensors can be used for direction-of arrival (DOA) estimation. A method is proposed to estimate the DOA based on the direction of maximum power. Algorithms mentioned in earlier works are shown to be special cases of the proposed method. An iterative algorithm based on the principal of gradient ascent is presented for the solution of the maximum power problem. The proposed maximum-power method is shown to approach the Cramer-Rao lower bound (CRLB) with a suitable choice of parameter. Dovid Levin, Sharon Gannot, Emanuël A. P. Habets |
ICASSP | 3 |
| 2011 | A proportionate adaptive algorithm with variable partitioned block length for acoustic echo cancellationabstractDue to the nature of an acoustic enclosure, the early part (i.e., direct path and early reflections) of the acoustic echo path is often sparse while the late reverberant part of the acoustic path is normally dispersive. In order to account for this structure within the acoustic impulse response when performing acoustic echo cancellation, we propose an adaptive filter that consists of two time-domain partition blocks, with adaptive block partitioning, such that different adaptive algorithms can be used for each block. Specifically, the improved proportionate normalized least-mean-square (IPNLMS) algorithm is used. Simulation results show that the proposed variable length partitioned block IPNLMS (VLPB-IPNLMS) algorithm works well in both sparse and dispersive circumstances and in practical applications involving time-varying systems. Pradeep Loganathan, Emanuël A. P. Habets, Patrick A. Naylor |
ICASSP | 2 |
| 2011 | Multiple-Hypothesis Extended Particle Filter for Acoustic Source Localization in Reverberant EnvironmentsabstractParticle filtering has been shown to be an effective approach to solving the problem of acoustic source localization in reverberant environments. In reverberant environment, the direct- arrival of the single source is accompanied by multiple spurious arrivals. Multiple-hypothesis model associated with these arrivals can be used to alleviate the unreliability often attributed to the acoustic source localization problem. Until recently, this multiple- hypothesis approach was only applied to bootstrap-based particle filter schemes. Recently, the extended Kalman particle filter (EPF) scheme which allows for an improved tracking capability was proposed for the localization problem. The EPF scheme utilizes a global extended Kalman filter (EKF) which strongly depends on prior knowledge of the correct hypotheses. Due to this, the extension of the multiple-hypothesis model for this scheme is not trivial. In this paper, the EPF scheme is adapted to the multiple-hypothesis model to track a single acoustic source in reverberant environments. Our work is supported by an extensive experimental study using both simulated data and data recorded in our acoustic lab. Various algorithms and array constellations were evaluated. The results demonstrate the superiority of the proposed algorithm in both tracking and switching scenarios. It is further shown that splitting the array into several sub-arrays improves the robustness of the estimated source location. A. Levy, Sharon Gannot, Emanuël A. P. Habets |
IEEE Trans. Speech Audio Process. | 3 |
| 2010 | An online quasi-Newton algorithm for blind SIMO identificationabstractIn the last decade various time- and frequency-domain algorithms were derived to blindly identify acoustic systems. One of these algorithms is the multichannel Newton (MCN) algorithm, which is also the basis of the well known normalized multichannel frequency-domain least-mean-square (NMCFLMS) algorithm. A major drawback of the MCN is that it requires the computation and inversion of a Hessian matrix, which involves extensive computation making it unsuitable for online applications. In this paper, we therefore derive and investigate an efficient online multichannel quasi-Newton (MCQN) algorithm that updates the inverse of the Hessian by analyzing successive gradient vectors. The new MCQN is shown to exhibit similar performance to MCN but with much reduced complexity. Emanuël A. P. Habets, Patrick A. Naylor |
ICASSP | 1 |
| 2010 | Performance analysis of IPNLMS for identification of time-varying systemsabstractThe tracking performance of adaptive filters is crucially important in practical applications involving time-varying systems. We present an analysis of the tracking performance for IPNLMS, one of the best known and best performing algorithms originally targeted at sparse system identification. We then validate our analytic results in practical simulations for echo cancellation for sparse and dispersive time-varying unknown echo path systems. These results show the analysis to be highly accurate in all the cases studied. Pradeep Loganathan, Emanuël A. P. Habets, Patrick A. Naylor |
ICASSP | 2 |
| 2010 | A System-Identification-Error-Robust Method for equalization of multichannel acoustic systemsabstractIn hands-free communications, speech received by a microphone is distorted by room reverberation that can reduce the intelligibility of speech. An approach to dereverberation is firstly to estimate the impulse responses of the acoustic channels between the speaker and the microphones and secondly to design a multichannel equalization system based on the estimated impulse responses. Traditional equalization techniques are designed without the consideration of estimation errors that are commonly introduced by the system identification process. In this work, a System-Identification-Error-Robust Equalization Method (SIEREM) for the equalization of multichannel room acoustic systems is presented. Experimental results for dereverberation using SIEREM applied to estimates of single-input multiple-output acoustic systems with known level of estimation errors show that the proposed equalization design significantly outperforms existing methods in the presence of both synthetic and real system identification errors. Wancheng Zhang, Emanuël A. P. Habets, Patrick A. Naylor |
ICASSP | 2 |
| 2010 | New Insights Into the MVDR Beamformer in Room AcousticsabstractThe minimum variance distortionless response (MVDR) beamformer, also known as Capon's beamformer, is widely studied in the area of speech enhancement. The MVDR beamformer can be used for both speech dereverberation and noise reduction. This paper provides new insights into the MVDR beamformer. Specifically, the local and global behavior of the MVDR beamformer is analyzed and novel forms of the MVDR filter are derived and discussed. In earlier works it was observed that there is a tradeoff between the amount of speech dereverberation and noise reduction when the MVDR beamformer is used. Here, the tradeoff between speech dereverberation and noise reduction is analyzed thoroughly. The local and global behavior, as well as the tradeoff, is analyzed for different noise fields such as, for example, a mixture of coherent and non-coherent noise fields, entirely non-coherent noise fields and diffuse noise fields. It is shown that maximum noise reduction is achieved when the MVDR beamformer is used for noise reduction only. The amount of noise reduction that is sacrificed when complete dereverberation is required depends on the direct-to-reverberation ratio of the acoustic impulse response between the source and the reference microphone. The performance evaluation supports the theoretical analysis and demonstrates the tradeoff between speech dereverberation and noise reduction. When desiring both speech dereverberation and noise reduction, the results also demonstrate that the amount of noise reduction that is sacrificed decreases when the number of microphones increases. Emanuël A. P. Habets, Jacob Benesty, Israel Cohen, Sharon Gannot, Jacek Dmochowski |
IEEE Trans. Speech Audio Process. | 1 |
| 2009 | On a tradeoff between dereverberation and noise reduction using the MVDR beamformerabstractThe minimum variance distortionless response (MVDR) beamformer can be used for both speech dereverberation and noise reduction. In this paper we analyse the tradeoff between the amount of speech dereverberation and noise reduction achieved by the MVDR beamformer. We show that the amount of noise reduction that is sacrificed when desiring both speech dereverberation and noise reduction depends on the direct-to-reverberation ratio of the acoustic transfer function between the desired source and a reference microphone. The performance evaluation supports the theoretical analysis and demonstrates the tradeoff between speech dereverberation and noise reduction. Emanuël A. P. Habets, Jacob Benesty, Israel Cohen, Sharon Gannot |
ICASSP | 1 |
| 2009 | Late Reverberant Spectral Variance Estimation Based on a Statistical ModelabstractIn speech communication systems the received microphone signals are degraded by room reverberation and ambient noise that decrease the fidelity and intelligibility of the desired speaker. Reverberant speech can be separated into two components, viz. early speech and late reverberant speech. Recently, various algorithms have been developed to suppress late reverberant speech. One of the main challenges is to develop an estimator for the so-called late reverberant spectral variance (LRSV) which is required by most of these algorithms. In this letter a statistical reverberation model is proposed that takes the energy contribution of the direct-path into account. This model is then used to derive a more general LRSV estimator, which in a particular case reduces to an existing LRSV estimator. Experimental results show that the developed estimator is advantageous in case the source-microphone distance is smaller than the critical distance. Emanuël A. P. Habets, Sharon Gannot, Israel Cohen |
IEEE Signal Process. Lett. | 1 |
| 2008 | Dual-microphone speech dereverberation using GARCH modelingabstractIn this paper, we develop a dual-microphone speech dereverberation algorithm for noisy environments, which is aimed at suppressing late reverberation and background noise. The spectral variance of the late reverberation is obtained with adaptively-estimated direct path compensation. A Markov-switching generalized autoregressive conditional heteroscedasticity (GARCH) model is used to estimate the spectral variance of the desired signal, which includes the direct sound and early reverberation. Experimental results demonstrate the advantage of the proposed algorithm compared to a decision-directed-based algorithm. Ari Abramson, Emanuël A. P. Habets, Sharon Gannot, Israel Cohen |
ICASSP | 2 |
| 2008 | Temporal selective dereverberation of noisy speech using one microphoneabstractReverberant speech can be described as sounding distant with noticeable coloration and echo. These detrimental perceptual effects are caused by early and late reflections, respectively, and reduces the fidelity and intelligibility of speech. It is well-known that the echo density of the reflections increases with time. Therefore, the temporal structure of early and late reflections differs. In this paper, we combine two different dereverberation techniques that were recently developed to suppress early and late reverberation separately. First, late reverberation is suppressed using a spectral processing technique that is based on a statistical reverberation model. Secondly, early reverberation and residual late reverberation are suppressed using a linear prediction (LP) residual processing technique. In addition, an objective measure based on the kurtosis of the LP residual is proposed to measure the coloration caused by early reflections. Experimental results demonstrate the beneficial use of the new single microphone system that reduces echo and coloration with little speech distortion. Emanuël A. P. Habets, Nikolay D. Gaubitch, Patrick A. Naylor |
ICASSP | 1 |
| 2008 | Blind estimation of reverberation time based on the distribution of signal decay ratesabstractThe reverberation time is one of the most prominent acoustic characteristics of an enclosure. Its value can be used to predict speech intelligibility, and is used by speech enhancement techniques to suppress reverberation. The reverberation time is usually obtained by analysing the decay rate of (i) the energy decay curve that is observed when a noise source is switched off, and (ii) the energy decay curve of the room impulse response. Estimating the reverberation time using only the observed reverberant speech signal, i.e., blind estimation, is required for speech evaluation and enhancement techniques. Recently, (semi) blind methods have been developed. Unfortunately, these methods are not very accurate when the source consists of a human speaker, and unnatural speech pauses are required to detect and/or track the decay. In this paper we extract and analyse the decay rate of the energy envelope blindly from the observed reverberation speech signal in the short-time Fourier transform domain. We develop a method to estimate the reverberation time using a property of the distribution of the decay rates. Experimental results using simulated and real reverberant speech signals demonstrate the performance of the new method. Jimi Yung-Chuan Wen, Emanuël A. P. Habets, Patrick A. Naylor |
ICASSP | 2 |
| 2008 | Multimicrophone speech dereverberation using spatiotemporal and spectral processingabstractSpeech signals acquired in a reverberant room with microphones positioned at a distance from the talker are degraded in quality due to reverberation and measurement noise. Therefore, enhancement of reverberant speech is important in hands-free telecommunications applications. The perceptual effects of reverberation can be linked to the room impulse response (RIR) between the talker and the microphone and are characterized by: (i) colouration, due to the strong early reflections and (ii) a distant 'echoey' quality due to the decaying tail of the RIR. Accordingly, we present a two-stage multimicrophone method for speech dereverberation. First, spatiotemporal averaging is performed on the linear prediction residual, which primarily reduces the effects of the early reflections. Secondly, a spectral subtraction method is employed to reduce late reverberation. Simulation results with measured RIRs and additive white Gaussian noise illustrate the performance of this method and show that the combined approach performs better than each of the two stages individually. Nikolay D. Gaubitch, Emanuël A. P. Habets, Patrick A. Naylor |
ISCAS | 2 |
| 2008 | Joint Dereverberation and Residual Echo Suppression of Speech Signals in Noisy EnvironmentsabstractHands-free devices are often used in a noisy and reverberant environment. Therefore, the received microphone signal does not only contain the desired near-end speech signal but also interferences such as room reverberation that is caused by the near-end source, background noise and a far-end echo signal that results from the acoustic coupling between the loudspeaker and the microphone. These interferences degrade the fidelity and intelligibility of near-end speech. In the last two decades, postfilters have been developed that can be used in conjunction with a single microphone acoustic echo canceller to enhance the near-end speech. In previous works, spectral enhancement techniques have been used to suppress residual echo and background noise for single microphone acoustic echo cancellers. However, dereverberation of the near-end speech was not addressed in this context. Recently, practically feasible spectral enhancement techniques to suppress reverberation have emerged. In this paper, we derive a novel spectral variance estimator for the late reverberation of the near-end speech. Residual echo will be present at the output of the acoustic echo canceller when the acoustic echo path cannot be completely modeled by the adaptive filter. A spectral variance estimator for the so-called late residual echo that results from the deficient length of the adaptive filter is derived. Both estimators are based on a statistical reverberation model. The model parameters depend on the reverberation time of the room, which can be obtained using the estimated acoustic echo path. A novel postfilter is developed which suppresses late reverberation of the near-end speech, residual echo and background noise, and maintains a constant residual background noise level. Experimental results demonstrate the beneficial use of the developed system for reducing reverberation, residual echo, and background noise. Emanuël A. P. Habets, Sharon Gannot, Israel Cohen, P. Sommen |
IEEE Trans. Speech Audio Process. | 1 |
| 2007 | Dual-Microphone Speech Dereverberation using a Reference SignalabstractSpeech signals recorded with a distant microphone usually contain reverberation, which degrades the fidelity and intelligibility of speech, and the recognition performance of automatic speech recognition systems. In this paper we propose a speech dereverberation system which uses two microphones. A generalized sidelobe canceller (GSC) type of structure is used to enhance the desired speech signal. The GSC structure is used to create two signals. The first signal is the output of a standard delay and sum beamformer, and the second signal is a reference signal which is constructed such that the direct speech signal is blocked. We propose to utilize the reverberation which is present in the reference signal to enhance the output of the delay and sum beamformer. The power envelope of the reference signal and the power envelope of the output of the delay and sum beamformer are used to estimate the residual reverberation in the output of the delay and sum beamformer. The output of the delay and sum beamformer is then enhanced using a spectral enhancement technique. The proposed method only requires an estimate of the direction of arrival of the desired speech source. Experiments using simulated room impulse responses are presented and show significant reverberation reduction while keeping the speech distortion low. Emanuël A. P. Habets, Sharon Gannot |
ICASSP (4) | 1 |
| 2005 | Multi-channel speech dereverberation based on a statistical model of late reverberationabstractSpeech signals recorded with a distant microphone usually contain reverberation, which degrades the fidelity and intelligibility of speech, and the recognition performance of automatic speech recognition systems. A multi-channel speech dereverberation algorithm is presented which reduces spectral coloration and late reverberation. A spatially averaged amplitude spectrum is used to estimate the instantaneous amplitude spectrum of the clean speech signal, which is then further enhanced using an estimate of the power spectrum of the late reverberant signal. The power spectrum of the late reverberant signal is constructed from multiple microphone signals and a statistical model of late reverberation. The algorithm is tested using synthetic reverberated signals. The performances for different room impulse responses with reverberation times ranging from approximately 150 ms to 350 ms show significant reverberation reduction with little signal distortion. Emanuël A. P. Habets |
ICASSP (4) | 1 |