EDBT 2026 Demo / reviewers in the wild / expert
Walter Kellermann
dblp:83/1301 · also Walter L. Kellermann
· DBLP profile ↗
130ranked-venue papers
6as first author
17since 2021 · last 2024
0000-0002-6501-3174ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 104 · 5 first-author · 14 since 2021Artificial intelligence and machine learning · 31 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Nonlinear acoustic echo cancellation based on pipelined Hermite filters
Mhd Modar Halimeh, Yi-Fei Pu, Lu Lu 0005, Walter Kellermann |
Signal Process. | 5 |
| 2024 | End-to-End Deep Learning-Based Adaptation Control for Linear Acoustic Echo CancellationabstractThe attenuation of acoustic loudspeaker echoes remains to be one of the open challenges to achieve pleasant full-duplex hands free speech communication. In many modern signal enhancement interfaces, this problem is addressed by a linear acoustic echo canceler which subtracts a loudspeaker echo estimate from the recorded microphone signal. To obtain precise echo estimates, the parameters of the echo canceler, i.e., the filter coefficients, need to be estimated quickly and precisely from the observed loudspeaker and microphone signals. For this a sophisticated adaptation control is required to deal with high-power double-talk and rapidly track time-varying acoustic environments which are often faced with portable devices. In this paper, we address this problem by end-to-end deep learning. In particular, we suggest to infer the step-size for a least mean squares frequency-domain adaptive filter update by a Deep Neural Network (DNN). Two different step-size inference approaches are investigated. On the one hand broadband approaches, which use a single DNN to jointly infer step-sizes for all frequency bands, and on the other hand narrowband methods, which exploit individual DNNs per frequency band. The discussion of benefits and disadvantages of both approaches leads to a novel hybrid approach which shows improved echo cancellation while requiring only small DNN architectures. Furthermore, we investigate the effect of different loss functions, signal feature vectors, and DNN output layer architectures on the echo cancellation performance from which we obtain valuable insights into the general design and functionality of DNN-based adaptation control algorithms. Thomas Haubner, Andreas Brendel, Walter Kellermann |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2024 | Erratum to "End-to-End Deep Learning-Based Adaptation Control for Linear Acoustic Echo Cancellation"abstractPresents corrections to the article “End-to-End Deep Learning-Based Adaptation Control for Linear Acoustic Echo Cancellation”. Thomas Haubner, Andreas Brendel, Walter Kellermann |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2024 | On Semi-Blind Source Separation-Based Approaches to Nonlinear Echo Cancellation Based on Bilinear Alternating OptimizationabstractAcoustic echo cancellation (AEC) is a crucial task in full duplex communications. As conventional linear filtering approaches are ineffective to deal with double-talk, various semi-blind source separation (SBSS)-based AEC algorithms are deceived, most of which are formulated and implemented in the frequency domain based on the multiplicative transfer function (MTF) model for computational efficiency. To avoid large latency and in order to deal with loudspeaker nonlinearities, the convolutive transfer function (CTF) model and odd power series expansion are leveraged, which are employed by numerous SBSS-based nonlinear AEC (SBSS-NAEC) algorithms. Conventional SBSS-NAEC methods estimate the series expansion coefficients and the CTF filter simultaneously making the number of free parameters to estimate large. Hence, the corresponding algorithms are computationally expensive and are difficult to optimize. In this work, we propose to decouple the series expansion coefficients and the CTF filters into a bilinear form and present a bilinear alternating optimization framework for estimating the model parameters. An alternating iterative projection (AIP) algorithm and an alternating element-wise iterative source steering (AEISS) algorithm are proposed. As the bilinear representation consists of less parameters compared to the conventional methods, the proposed algorithms not only improve the AEC performance but also reduce the computational complexity, which is validated by comprehensive simulations and experiments. Xianrui Wang, Yichen Yang 0010, Andreas Brendel, Tetsuya Ueda, Shoji Makino, Jacob Benesty, Walter Kellermann, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2023 | Exploiting Spatial Information with the Informed Complex-Valued Spatial Autoencoder for Target Speaker ExtractionabstractIn conventional multichannel audio signal enhancement, spatial and spectral filtering are often performed sequentially. In contrast, it has been shown that for neural spatial filtering a joint approach of spectro-spatial filtering is more beneficial. In this contribution, we investigate the spatial filtering performed by such a time-varying spectro-spatial filter. We extend the recently proposed complex-valued spatial autoencoder (COSPA) for the task of target speaker extraction by leveraging its interpretable structure and purposefully informing the network of the target speaker’s position. We show that the resulting informed COSPA (iCOSPA) effectively and flexibly extracts a target speaker from a mixture of speakers. We also find that the proposed architecture is well capable of learning pronounced spatial selectivity patterns and show that the results depend significantly on the training target and the reference signal when computing various evaluation metrics. Annika Briegleb, Mhd Modar Halimeh, Walter Kellermann |
ICASSP | 3 |
| 2023 | Switching Kronecker Product Linear Filtering for Multispeaker Adaptive Speech DereverberationabstractDereverberation, a process to mitigate or eliminate the reverberation effect, plays an important role in hands-free speech communication and human-machine interfaces. Tremendous efforts have been devoted to this problem and various methods have been developed over the last three decades. Those methods generally assume that there is only a single speaker in the acoustic environment and, consequently, they suffer from significant performance degradation if multiple speakers participate in the conversation. How to deal with reverberation in multiple-speaker scenarios is still a challenging problem, which is studied in this work. We present a switching multichannel linear prediction filtering method, which designs multiple linear filters with each tracking one speaker. When some speaker is active, the corresponding filter and the weighted cross-correlation matrix are updated while the other filters are kept unchanged. To further improve the performance and reduce complexity, we apply the Kronecker product to decompose every linear prediction filter into a Kronecker product of two shorter filters: one is time-invariant and the other is time-varying. The former is estimated with a batch method (using only a few seconds of speech signal when the corresponding speaker starts to talk in the entire conversation) while a recursive least-squares algorithm is derived for identifying the time-varying set of Kronecker filters. Gongping Huang, Jacob Benesty, Israel Cohen, Emil Winebrand, Jingdong Chen, Walter Kellermann |
ICASSP | 6 |
| 2023 | Spatially Informed Independent vector analysis for Source Extraction based on the convolutive Transfer Function ModelabstractSpatial information can help improve source separation performance. Numerous spatially informed source extraction methods based on the independent vector analysis (IVA) have been developed, which can achieve reasonably good performance in non- or weakly reverberant environments. However, the performance of those methods degrades quickly as the reverberation increases. The underlying reason is that those methods are derived based on the multiplicative transfer function model with a rank-1 assumption, which does not hold true if reverberation is strong. To circumvent this issue, this paper proposes to use the convolutive transfer function (CTF) model to improve the source extraction performance and develop a spatially informed IVA algorithm. Simulations demonstrate the efficacy of the developed method even in highly reverberant environments. Xianrui Wang, Andreas Brendel, Gongping Huang, Yichen Yang 0010, Walter Kellermann, Jingdong Chen |
ICASSP | 5 |
| 2022 | Manifold Learning-Supported Estimation of Relative Transfer Functions For Spatial FilteringabstractMany spatial filtering algorithms used for voice capture in, e.g., teleconferencing applications, can benefit from or even rely on knowledge of Relative Transfer Functions (RTFs). Accordingly, many RTF estimators have been proposed which, however, suffer from performance degradation under acoustically adverse conditions or need prior knowledge on the properties of the interfering sources. While state-of-the-art RTF estimators ignore prior knowledge about the acoustic enclosure, audio signal processing algorithms for teleconferencing equipment are often operating in the same or at least a similar acoustic enclosure, e.g., a car or an office, such that training data can be collected. In this contribution, we use such data to train Variational Autoencoders (VAEs) in an unsupervised manner and apply the trained VAEs to enhance imprecise RTF estimates. Furthermore, a hybrid between classic RTF estimation and the trained VAE is investigated. Comprehensive experiments with real-world data confirm the efficacy for the proposed method. Andreas Brendel, Johannes Zeitler, Walter Kellermann |
ICASSP | 3 |
| 2022 | Complex-Valued Spatial Autoencoders for Multichannel Speech EnhancementabstractIn this contribution, we present a novel online approach to multichannel speech enhancement. The proposed method estimates the enhanced signal through a filter-and-sum framework. More specifically, complex-valued masks are estimated by a deep complex-valued neural network, termed the complex-valued spatial autoencoder. The proposed network is capable of manipulating both the phase and the amplitude of the microphone signals and hence, the network is able to exploit both spatial and spectral characteristics of the desired source signal resulting in a physically plausible spatial selectivity and superior speech quality. Mhd Modar Halimeh, Walter Kellermann |
ICASSP | 2 |
| 2022 | End-To-End Deep Learning-Based Adaptation Control for Frequency-Domain Adaptive System IdentificationabstractWe present a novel end-to-end deep learning-based adaptation control algorithm for frequency-domain adaptive system identification. The proposed method exploits a deep neural network to map observed signal features to corresponding step-sizes which control the filter adaptation. The parameters of the network are optimized in an end-to-end fashion by minimizing the average normalized system distance of the adaptive filter. This avoids the need of explicit signal power spectral density estimation as required for model-based adaptation control and further auxiliary mechanisms to deal with model inaccuracies. The proposed algorithm achieves fast convergence and robust steady-state performance for scenarios characterized by high-level, non-white and non-stationary additive noise signals, abrupt environment changes and additional model inaccuracies. Thomas Haubner, Andreas Brendel, Walter Kellermann |
ICASSP | 3 |
| 2021 | Network-Aware Optimal Microphone Channel Selection in Wireless Acoustic Sensor NetworksabstractTo address the vital problem of selecting the most useful microphones in wireless acoustic sensor networks, this paper proposes a novel, general-purpose approach that accounts for both acoustic and network aspects and remains application-agnostic for broad applicability. The inter-channel correlation of single-channel signal features, together with tools from spectral graph theory, is used to assess the usefulness from an acoustic perspective. By only transmitting the features that characterize signal frames as opposed to the full signal waveform, the unique constraints of wireless sensor networks are accommodated. The source-to-sink transmission delay, resulting from embedding a distributed signal processing application into a wireless network, captures the usefulness from a network perspective. The experiments demonstrate the efficacy of the proposed method for an exemplary multichannel signal processing application. Michael Günther 0003, Haitham Afifi, Andreas Brendel, Holger Karl, Walter Kellermann |
ICASSP | 5 |
| 2021 | Accelerating Auxiliary Function-Based Independent Vector AnalysisabstractIndependent Vector Analysis (IVA) is an effective approach for Blind Source Separation (BSS) of convolutive mixtures of audio signals. As a practical realization of an IVA-based BSS algorithm, the so-called AuxIVA update rules based on the Majorize-Minimize (MM) principle have been proposed which allow for fast and computationally efficient optimization of the IVA cost function. For many real-time applications, however, update rules for IVA exhibiting even faster convergence are highly desirable. To this end, we investigate techniques which accelerate the convergence of the AuxIVA update rules without extra computational cost. The efficacy of the proposed methods is verified in experiments representing real-world acoustic scenarios. Andreas Brendel, Walter Kellermann |
ICASSP | 2 |
| 2021 | Combining Adaptive Filtering And Complex-Valued Deep Postfiltering For Acoustic Echo CancellationabstractIn this contribution, we introduce a novel approach to noise-robust acoustic echo cancellation employing a complex-valued Deep Neural Network (DNN) for postfiltering. In a first step, early linear echo components are removed using a double-talk robust adaptive filter. The residual signal is subsequently processed by the proposed post-filter (PF). Due to its complex-valued nature, the PF allows to sup-press unwanted signal components without introducing distortions to the near-end speaker. For training and evaluation, we exclusively use data from the ICASSP 2021 AEC challenge. Exploiting only a moderate amount of training data, we demonstrate the efficacy of the proposed method. Specifically, we show that the PF (i) benefits significantly from a preceding linear adaptive filter and (ii) significantly outperforms a conventional real-valued DNN-based PF. Mhd Modar Halimeh, Thomas Haubner, Annika Briegleb, Alexander Schmidt 0004, Walter Kellermann |
ICASSP | 5 |
| 2021 | Noise-Robust Adaptation Control for Supervised Acoustic System Identification Exploiting a Noise DictionaryabstractWe present a noise-robust adaptation control strategy for block-online supervised acoustic system identification by exploiting a noise dictionary. The proposed algorithm takes advantage of the pronounced spectral structure which characterizes many types of interfering noise signals. We model the noisy observations by a linear Gaussian Discrete Fourier Transform-domain state space model whose parameters are estimated by an online generalized Expectation-Maximization algorithm. Unlike all other state-of-the-art approaches we suggest to model the covariance matrix of the observation probability density function by a dictionary model. We propose to learn the noise dictionary from training data, which can be gathered either offline or online whenever the system is not excited, while we infer the activations continuously. The proposed algorithm represents a novel machine-learning-based approach to noise-robust adaptation control which allows for faster convergence in applications characterized by high-level and non-stationary interfering noise signals and abrupt system changes. Thomas Haubner, Andreas Brendel, Mohamed Elminshawi, Walter Kellermann |
ICASSP | 4 |
| 2021 | Effective Rank-Based Estimation of the Coherent-to-Diffuse Power RatioabstractMany algorithms for speech dereverberation and noise reduction rely on an estimate of the coherent-to-diffuse power ratio (CDR). Such systems typically operate in very diverse acoustic conditions, and CDR estimators relying on very weak model assumptions about the acoustic sound field of the desired speech and interfering noise are hence desirable. A CDR estimator whose design is based on this premise is devised in this contribution. The proposed non-iterative CDR estimator exploits the effective rank of the spatial covariance matrix of the recorded input signals by assuming it to be lower for a coherent sound field than for a diffuse sound field. In addition to this weak assumption, related methods usually require information about, e.g., the array geometry, direction-of-arrival (DOA) of the desired source or rely on a coherence model for the desired signal or background noise, which is not required for the proposed method. Despite the use of little a priori information about the acoustic sound field, the new estimator achieves a significantly higher estimation accuracy for the CDR in comparison to related state-of-the-art approaches which use explicit coherence models. Heinrich W. Löllmann, Andreas Brendel, Walter Kellermann |
ICASSP | 3 |
| 2021 | Misalignment Recognition in Acoustic Sensor Networks Using a Semi-Supervised Source Estimation Method and Markov Random FieldsabstractIn this paper, we consider the problem of acoustic source localization by acoustic sensor networks (ASNs) using a promising, learning-based technique that adapts to the acoustic environment. In particular, we look at the scenario when a node in the ASN is displaced from its position during training. As the mismatch between the ASN used for learning the localization model and the one after a node displacement leads to erroneous position estimates, a displacement has to be detected and the displaced nodes need to be identified. We propose a method that considers the disparity in position estimates made by leave-one-node-out (LONO) sub-networks and uses a Markov random field (MRF) framework to infer the probability of each LONO position estimate being aligned, misaligned or unreliable while accounting for the noise inherent to the estimator. This probabilistic approach is advantageous over naïve detection methods, as it outputs a normalized value that encapsulates conditional information provided by each LONO sub-network on whether the reading is in misalignment with the overall network. Experimental results confirm that the performance of the proposed method is consistent in identifying compromised nodes in various acoustic conditions. Gabriel F. Miller, Andreas Brendel, Walter Kellermann, Sharon Gannot |
ICASSP | 3 |
| 2021 | Robust Dereverberation With Kronecker Product Based Multichannel Linear PredictionabstractReverberation impairs not only the speech quality, but also intelligibility. The weighted-prediction-error (WPE) method, which estimates the late reverberation component based on a multichannel linear predictor, is by far one of the most effective algorithms for dereverberation. Generally, the WPE prediction filter in every short-time-Fourier-transform (STFT) subband has to be long enough to estimate accurately the late reverberation component. As a consequence, WPE is computationally expensive, which makes it difficult to implement into real-time embedded or edge computing devices. Moreover, WPE is sensitive to additive noise and its performance may suffer from dramatic degradation even in environments where the signal-to-noise ratio (SNR) is high. To address these drawbacks, this letter proposes to decompose the multichannel linear prediction filter as a Kronecker product of a temporal (interframe) prediction filter and a spatial filter. An iterative algorithm is then developed to optimize the two filters. In comparison with the original WPE algorithm, the presented method not only exhibits better performance in terms of dereverberation and robustness to additive noise, as there are fewer parameters to estimate for a given number of observation signal samples, but is also computationally more efficient, since the dimensions of the covariance matrices after Kronecker product decomposition are smaller. Wenxing Yang, Gongping Huang, Jingdong Chen, Jacob Benesty, Israel Cohen, Walter Kellermann |
IEEE Signal Process. Lett. | 6 |
| 2020 | Spatially Guided Independent Vector AnalysisabstractWe present a Maximum A Posteriori (MAP) derivation of the Independent Vector Analysis (IVA) algorithm for blind source separation incorporating an additional spatial prior over the demixing matrices. In this way, the outer permutation ambiguity of IVA is avoided and the algorithm can be guided towards a desired solution in adverse acoustic conditions. The resulting MAP optimization problem is solved by deriving majorize-minimize update rules to achieve convergence speed comparable to the well-known auxiliary function IVA algorithm, i.e., the convergence is not impaired by the additional constraint. The proposed algorithm exhibits superior performance at lower computational cost than a state-of-the-art spatially constrained IVA algorithm in a setup defined by real-world Room Impulse Responses (RIRs). Andreas Brendel, Thomas Haubner, Walter Kellermann |
ICASSP | 3 |
| 2020 | Efficient Multichannel Nonlinear Acoustic Echo Cancellation Based on a Cooperative StrategyabstractWhile a common approach to address nonlinear distortions, emitted by multiple loudspeakers and observed by multiple microphones, is to use post-filtering techniques, this paper proposes a cooperative strategy to rather model and then cancel such distortions. In this approach, the overall problem of modeling distortions emitted by a number of loudspeakers is divided into multiple simpler and easier tasks of estimating distortions emitted by subsets of loudspeakers. This approach allows also the exploitation of the physical configuration of the loudspeakers and microphones to select certain microphone signals for estimating the nonlinearity of loudspeakers that contribute the predominant part of the acoustic echo to this microphone signal. The proposed strategy is realized using the elitist resampling particle filter and the Gaussian particle filter. Both variants are evaluated and compared to a linear approach using synthesized and real recordings. Mhd Modar Halimeh, Walter Kellermann |
ICASSP | 2 |
| 2020 | Generalized Coherence-Based Signal EnhancementabstractThis contribution presents a novel approach for coherence-based signal enhancement. An estimator for the coherent-to-diffuse ratio (CDR) is devised, which exploits the concept of generalized magnitude coherence and thus, unlike common state-of-the-art schemes, can simultaneously take advantage of more than two microphones. Moreover, the speech enhancement by CDR-based spectral weighting is not performed as a post-filtering step, but by enhancing the most appropriate microphone signal. This signal is implicitly determined as part of the CDR estimation such that the presented technique does not depend on an estimation of the direction-of-arrival (DOA) or similar side-information about the desired source.The application of the new approach to binaural hearings aids shows that it achieves a consistently better speech enhancement performance than comparable state-of-the-art approaches. Heinrich W. Löllmann, Andreas Brendel, Walter Kellermann |
ICASSP | 3 |
| 2020 | Acoustic Self-Awareness of Autonomous Systems in a World of SoundsabstractAutonomous systems (ASs) operating in real-world environments are exposed to a plurality and diversity of sounds that carry a wealth of information for perception in cognitive dynamic systems. While the importance of the acoustic modality for humans as “ASs” is obvious, it is investigated to what extent current technical ASs operating in scenarios filled with airborne sound exploit their potential for supporting self-awareness. As a first step, the state of the art of relevant generic techniques for acoustic scene analysis (ASA) is reviewed, i.e., source localization and the various facets of signal enhancement, including spatial filtering, source separation, noise suppression, dereverberation, and echo cancellation. Then, a comprehensive overview of current techniques for ego-noise suppression, as a specific additional challenge for ASs, is presented. Not only generic methods for robust source localization and signal extraction but also specific models and estimation methods for ego-noise based on various learning techniques are discussed. Finally, active sensing is considered with its unique potential for ASA and, thus, for supporting self-awareness of ASs. Therefore, recent techniques for binaural listening exploiting head motion, for active localization and exploration, and for active signal enhancement are presented, with humanoïd robots as typical platforms. Underlining the multimodal nature of self-awareness, links to other modalities and nonacoustic reference information are pointed out where appropriate. Alexander Schmidt 0004, Heinrich W. Löllmann, Walter Kellermann |
Proc. IEEE | 3 |
| 2020 | The LOCATA Challenge: Acoustic Source Localization and TrackingabstractThe ability to localize and track acoustic events is a fundamental prerequisite for equipping machines with the ability to be aware of and engage with humans in their surrounding environment. However, in realistic scenarios, audio signals are adversely affected by reverberation, noise, interference, and periods of speech inactivity. In dynamic scenarios, where the sources and microphone platforms may be moving, the signals are additionally affected by variations in the source-sensor geometries. In practice, approaches to sound source localization and tracking are often impeded by missing estimates of active sources, estimation errors, as well as false estimates. The aim of the LOCAlization and TrAcking (LOCATA) Challenge is an open-access framework for the objective evaluation and benchmarking of broad classes of algorithms for sound source localization and tracking. This article provides a review of relevant localization and tracking algorithms and, within the context of the existing literature, a detailed evaluation and dissemination of the LOCATA submissions. The evaluation highlights achievements in the field, open challenges, and identifies potential future directions. Christine Evers, Heinrich W. Löllmann, Heinrich Mellmann, Alexander Schmidt 0004, Hendrik Barfuss, Patrick A. Naylor, Walter Kellermann |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2019 | Localization of an Unknown Number of Speakers in Adverse Acoustic Conditions Using Reliability Information and DiarizationabstractThis paper investigates localization of an arbitrary number of simultaneously active speakers in an acoustic enclosure. We propose an algorithm capable of estimating the number of speakers, using reliability information to obtain robust estimation results in adverse acoustic scenarios and estimating individual probability distributions describing the position of each speaker using convex geometry tools. To this end, we start from an established algorithm for localization of acoustic sources based on the EM algorithm. There, the estimation of the number of sources as well as the handling of reverberation has not been addressed sufficiently. We show improvement in the localization of a higher number of sources and in the robustness in adverse conditions including interference from competing speakers, reverberation and noise. Andreas Brendel, Bracha Laufer-Goldshtein, Sharon Gannot, Ronen Talmon, Walter Kellermann |
ICASSP | 5 |
| 2019 | Neural Networks Sequential Training Using Variational Gaussian Particle FilterabstractIn this paper, we propose a sequential training algorithm for feed-forward neural networks based on particle filtering. The proposed algorithm uses variational learning to tailor a proposal density by minimizing the variational energy. This density is then incorporated into the Gaussian particle filter framework. The proposed algorithm and an extension to it using evolutionary resampling are compared to training a neural network using a random walk-based particle filter, an extended Kalman filter, the use of variational learning only, and the backpropagation algorithm, using a synthetic dataset generated by a time-varying random process and a real dataset, where the proposed approach resulted in a moderately lower training and testing errors and a better convergence behavior, rendering the algorithm attractive for uses such as neural networks pre-training. Mhd Modar Halimeh, Andreas Brendel, Walter Kellermann |
ICASSP | 3 |
| 2019 | Informed Ego-noise Suppression Using Motor Data-driven DictionariesabstractThe suppression of ego-noise for (humanoid) robots is typically addressed by learning-based techniques. In this paper, we propose a novel approach which models significant parts of ego-noise spectrograms based on motor data and does not require a prior training step. Accordingly, the intrinsic harmonic structure of ego noise is taken into account by introducing a nonnegative matrix factorization (NMF) framework with motor data-driven dictionaries. Limited improvement was observed by employing an additional pre-trained small-sized dictionary accounting for the residual ego-noise. The presented approach exhibits comparable suppression performance to an audio only-based approach trained specifically to the scenario, while the number of dictionary elements which require prior learning can be reduced by a factor of two. For ego-noise resulting from previously unseen movements, the proposed method shows consistently superior suppression results while the audio only-based approach degrades drastically. Alexander Schmidt 0004, Walter Kellermann |
ICASSP | 2 |
| 2019 | A Neural Network-Based Nonlinear Acoustic Echo CancellerabstractIn this letter, we introduce a novel approach for nonlinear acoustic echo cancellation. The proposed approach uses the principle of transfer learning to train a neural network that approximates the nonlinear function responsible for the nonlinear distortions and generalizes this network to different acoustic conditions. The topology of the proposed network is inspired by the conventional adaptive filtering approaches for nonlinear acoustic echo cancellation. The network is trained to model the nonlinear distortions using the conventional error backpropagation algorithm. In deployment, and in order to account for any variation or discrepancy between training and deployment conditions, only a subset of the network's parameters is adapted using the significance-aware elitist resampling particle filter. The proposed approach is evaluated and verified using synthesized nonlinear distortions and real nonlinear distortions recorded by a commercial mobile phone. Mhd Modar Halimeh, Christian Huemmer 0001, Walter Kellermann |
IEEE Signal Process. Lett. | 3 |
| 2018 | Learning-Based Acoustic Source-Microphone Distance Estimation Using the Coherent-to-Diffuse Power RatioabstractWe propose a method for estimating the distance between a sound source and a pair of recording microphones. The developed algorithm operates in the short-time Fourier transform domain and is based on estimates of the coherent-to-diffuse power ratio, which provides a measure for the amount of reverberation in each time-frequency bin. For a direct use of these estimates, precise knowledge on the room characteristics is necessary, which is in practice usually not available and hard to obtain. Therefore, we use a learning-based method, which adapts to the characteristics of the room in a training phase and estimates the source-microphone distance in a testing phase. The experiments comprise various setups with simulated and real data. It is shown that the proposed method generalizes well for different microphone positions and works robustly for different source signals, directions of arrival, reverberation times, and signal observation intervals. This leads to a high estimation accuracy at a low computational complexity with a small amount of training data. Andreas Brendel, Walter Kellermann |
ICASSP | 2 |
| 2018 | Nonlinear Acoustic Echo Cancellation Using Elitist Resampling Particle FilterabstractThis paper considers an effective method for nonlinear acoustic echo cancellation (NL-AEC). More specifically, we model the nonlinear echo path by a latent state vector capturing the coefficients of a memoryless processor and a linear finite impulse response filter. To estimate the posterior probability distribution of the state vector, an elitist particle filter based on evolutionary strategies (EPFES) has been proposed, which evaluates realizations of the latent state vector based on long-term fitness measures. This method includes a manually-tuned recursive calculation of the probabilities that the observation has been produced by the state-vector realizations. For avoiding this manual tuning, we introduce a new approach denoted as Elitist Resampling Particle Filtering (ERPF) which can also be shown to combine the advantages of the Sequential Importance Sampling Particle Filter (SIS-PF) and the Sequential Importance Sampling/Resampling Particle Filter (SIR-PF). This new approach allows universal use and leads to superior system identification performance compared to both the original EPFES as well as the SIR-PF, as verified for a simulated scenario and a real smartphone recording. Mhd Modar Halimeh, Christian Huemmer 0001, Walter Kellermann |
ICASSP | 3 |
| 2018 | A Novel Ego-Noise Suppression Algorithm for Acoustic Signal Enhancement in Autonomous SystemsabstractThe use of autonomous systems (ASs), such as humanoid robots, drones or self-driving vehicles, has expanded significantly in recent years. For such systems, acoustic scene analysis can provide useful information about the environment and supports the AS to react appropriately. However, compared to most other application areas, analysis and enhancement of acoustic signals captured by ASs is not only complicated by external sources of signal degradation but also by very specific challenges like internal and self-created ego-noise. This paper first gives an overview of a typical acoustic scenario an AS is exposed to. Then, we consider the specific problem of ego-noise suppression and propose to use motor data to predict the characteristic time-varying harmonic structure of ego-noise. This knowledge is then incorporated into a multichannel dictionary-based algorithm. The resulting two-stage ego-noise reduction scheme is evaluated for ego-noise of a humanoid robot and outperforms a comparable method that uses no motor data but a a larger dictionary. Alexander Schmidt 0004, Heinrich W. Löllmann, Walter Kellermann |
ICASSP | 3 |
| 2018 | Estimating Parameters of Nonlinear Systems Using the Elitist Particle Filter Based on Evolutionary StrategiesabstractIn this paper, we present the elitist particle filter based on evolutionary strategies (EPFES) as an efficient approach to estimate the statistics of a latent state vector capturing the relevant information of a nonlinear system. Similar to classical particle filtering, the EPFES consists of a set of particles and respective weights which represent different realizations of the latent state vector and their likelihood of being the solution of the optimization problem. As main innovation, the EPFES includes an evolutionary elitist-particle selection scheme which combines long-term information with instantaneous sampling from an approximated continuous posterior distribution. In this paper, we propose two advancements of the previously published elitist-particle selection process. Further, the EPFES is shown to be a generalization of the widely-used Gaussian particle filter and thus evaluated with respect to the latter: First, we consider the univariate nonstationary growth model with time-variant latent state variable to evaluate the tracking capabilities of the EPFES for instantaneously calculated particle weights. This is followed by addressing the problem of single-channel nonlinear acoustic echo cancellation as a challenging benchmark task for identifying an unknown system of large search space: the nonlinear acoustic echo path is modeled by a cascade of a parameterized preprocessor (to model the loudspeaker signal distortions) and a linear FIR filter (to model the sound wave propagation and the microphone). By using long-term information, we highlight the efficacy of the well-generalizing EPFES in estimating the preprocessor parameters for a simulated scenario and a real smartphone recording. Finally, we illustrate similarities between the EPFES and evolutionary algorithms to outline future improvements by fusing the achievements of both fields of research. Christian Huemmer 0001, Christian Hofmann 0001, Roland Maas, Walter Kellermann |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2018 | Spatial Noise-Field Control With Online Secondary Path Modeling: A Wave-Domain ApproachabstractDue to strong interchannel interference in multichannel active noise control (ANC), there are fundamental problems associated with the filter adaptation and online secondary path modeling remains a major challenge. This paper proposes a wave-domain adaptation algorithm for multichannel ANC with online secondary path modelling to cancel tonal noise over an extended region of two-dimensional plane in a reverberant room. The design is based on exploiting the diagonal-dominance property of the secondary path in the wave domain. The proposed wave-domain secondary path model is applicable to both concentric and nonconcentric circular loudspeakers and microphone array placement, and is also robust against array positioning errors. Normalized least mean squares-type algorithms are adopted for adaptive feedback control. Computational complexity is analyzed and compared with the conventional time-domain and frequency-domain multichannel ANCs. Through simulation-based verification in comparison with existing methods, the proposed algorithm demonstrates more efficient adaptation with low-level auxiliary noise. Wen Zhang 0002, Christian Hofmann 0001, Michael Buerger, Thushara D. Abhayapala, Walter Kellermann |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2018 | Statistical Analysis of the Multichannel Wiener Filter Using a Bivariate Normal Distribution for Sample Covariance MatricesabstractThis paper studies the statistical performance of the multichannel Wiener filter (MWF) when the weights are computed using estimates of the sample covariance matrices of the noisy and the noise signals. It is well known that the optimal weights of the minimum variance distortionless response beamformer are only determined by the noisy sample covariance matrix or the noise sample covariance matrix, while those of the MWF are determined by both of them. Therefore, the difficulty increases dramatically in statistically analyzing the MWF when compared to analyzing the MVDR, where the main reason is that expressing the general joint probability density function (p.d.f.) of the two sample covariance matrices presented a Hitherto unsolved problem, to the best of our knowledge. For a deeper insight into the statistical performance of the MWF, this paper first introduces a bivariate normal distribution to approximately model the joint p.d.f. of the noisy and the noise sample covariance matrices. Each sample covariance matrix is approximately modeled by a random scalar multiplied by its true covariance matrix. This approximation is designed to preserve both the bias and the mean squared error of the matrix with respect to a natural distance on covariance matrices. The correlation of the bivariate normal distribution, referred to as the sample covariance matrices intrinsic correlation coefficient, captures all second-order dependencies of the noisy and the noise sample covariance matrices. By using the proposed bivariate normal distribution, the performance of the MWF can be predicted from the derived analytical expressions and many interesting results are revealed. As an example, the theoretical analysis demonstrates that the MWF performance may degrade in terms of noise reduction and signal-to-noise-ratio improvement when using more sensors in some noise scenarios. Chengshi Zheng, Antoine Deleforge, Xiaodong Li 0002, Walter Kellermann |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2017 | Robust audio localization with phase unwrappingabstractMost of multichannel sound source Direction Of Arrival (DOA) estimation algorithms suffer from spatial aliasing problems. The phase differences between a pair of microphones are wrapped beyond the spatial aliasing frequency. A common solution is to adjust the distance between the microphones to obtain a suitable aliasing frequency, and take only the frequency band below the aliasing frequency for localization. With correct phase unwrapping, a broader frequency band can be utilized for localization. In this paper, we investigate a method for phase unwrapping solving the spatial aliasing problem for scenarios with a single source and high-level diffuse background noise (around 0dB SNR). The aliasing frequency is estimated from the signal, and is used to unwrap a phase difference vector. Pre- and post-processing steps are applied to increase the robustness. Our experiments with a large number of simulated and real signals demonstrate the robustness of our method in noise. Kainan Chen, Jürgen T. Geiger, Walter Kellermann |
ICASSP | 3 |
| 2017 | Online environmental adaptation of CNN-based acoustic models using spatial diffuseness featuresabstractWe propose a new concept for adapting CNN-based acoustic models using spatial diffuseness features as auxiliary information about the acoustic environment: the spatial diffuseness features are simultaneously employed as acoustic-model input features and to estimate environmental cues for context adaptation, where one convolutional layer is factorized into several sub-layers to represent different acoustic conditions. This context-adaptive CNN-based acoustic model facilitates an online environmental adaptation and is experimentally verified for the real-world recordings provided by the CHiME-3 task. The best performing setup reduces the average word error rate scores achieved by the baseline system (without using spatial diffuseness features) from 19.4% to 15.9% and 12.2% to 10.7% considering two experimental setups with and without front-end signal enhancement, respectively. Christian Huemmer 0001, Marc Delcroix, Atsunori Ogawa, Keisuke Kinoshita, Tomohiro Nakatani, Walter Kellermann |
ICASSP | 6 |
| 2017 | Online secondary path modelling in wave-domain active noise controlabstractThe performance of an ANC system largely depends on the availability of an accurate secondary path model. This is however a major challenge in multichannel ANC where the computational complexity increases significantly with the number of secondary sources and error sensors. This paper proposes wave-domain adaptive processing algorithm for multichannel ANC with online secondary modelling to cancel a tonal noise over a region of space within a reverberant room. The design is based on exploiting a special property of the secondary path model in the wave domain. A feedback control system is implemented, where a single microphone array is placed at the boundary of the control region to measure the residual signals, and a loudspeaker array reproduces secondary sources to generate the anti-noise signals and auxiliary noise for secondary path modelling. Through experimental verification in comparison with existing methods the proposed algorithm demonstrates more efficient adaptation with low-level auxiliary noise. Wen Zhang 0002, Christian Hofmann 0001, Michael Buerger, Thushara D. Abhayapala, Walter Kellermann |
ICASSP | 5 |
| 2017 | Robust coherence-based spectral enhancement for speech recognition in adverse real-world environments
Hendrik Barfuss, Christian Huemmer 0001, Andreas Schwarz, Walter Kellermann |
Comput. Speech Lang. | 4 |
| 2017 | Combined LCMV-TRINICON Beamforming for Separating Multiple Speech Sources in Noisy and Reverberant EnvironmentsabstractThe problem of source separation using an array of microphones in reverberant and noisy conditions is addressed. We consider applying the well-known linearly constrained minimum variance (LCMV) beamformer (BF) for extracting individual speakers. Constraints are defined using relative transfer functions (RTFs) for the sources, which are ratios of acoustic transfer functions (ATFs) between any microphone and a reference microphone. The latter are usually estimated by methods that rely on single-talk time segments where only a single source is active and on reliable knowledge of the source activity. Two novel algorithms for estimation of RTFs using the “Triple N” ICA for convolutive mixtures (TRINICON) framework are proposed, not resorting to the usually unavailable source activity pattern. The first algorithm estimates the RTFs of the sources by applying multiple two-channel geometrically constrained (GC) TRINICON units, where approximate direction of arrival information for the sources is utilized for ensuring convergence to the desired solution. The GC-TRINICON is applied to all microphone pairs using a common reference microphone. In the second algorithm, we propose to estimate RTFs iteratively using GC-TRINICON, where instead of using a fixed reference microphone as before, we suggest to use the output signals of LCMV-BFs from the previous iteration as spatially processed references with improved signal-to-interference-and-noise ratio. For both algorithms, a simple detection of noise-only time segments is required for estimating the covariance matrix of noise and interference. We conduct an experimental study in which the performance of the proposed methods is confirmed and compared to corresponding supervised methods. Shmulik Markovich-Golan, Sharon Gannot, Walter Kellermann |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Higher-order listening room compensation with additive compensation signalsabstractThe performance of sound reproduction systems for spatial audio is impaired by time-variant, reverberant listening environments. To tackle this issue, the Loudspeaker-Enclosure-Microphone System (LEMS) between the loudspeakers and reference microphones in the listening environment can be identified adaptively to allow an LEMS-specific pre-processing of the loudspeaker signals. This contribution introduces a broadband implementation of a narrowband Listening Room Compensation (LRC) method with additive compensation signals, recently proposed by Talagala et al. [1], it extends the concept to higher-order compensation, and compares LRC to Listening Room Equalization (LRE) analytically. Evaluations in an image-source environment confirm the efficacy of higher-order LRC and its suitability as a complexity-reduced alternative to LRE. Christian Hofmann 0001, Michael Günther 0003, Michael Buerger, Walter Kellermann |
ICASSP | 4 |
| 2016 | Source-specific system identificationabstractMany applications in audio communication require the identification of Loudspeaker-Enclosure-Microphone Systems (LEMS) with multiple inputs and outputs. The according computational complexity typically grows at least proportionally along the number of acoustic paths, which is the product of the number of loudspeakers and the number of microphones. Furthermore, the typical, highly correlated loudspeaker signals preclude an exact identification of the LEMS. To this end, a novel system identification scheme employing prior information from an object-based rendering system, e.g., Ambisonics [1,2] or Wave Field Synthesis (WFS) [3,4], is proposed. In this scheme, only a source-specific system from each virtual source to each microphone is identified adaptively and uniquely. This estimate for a source-specific system can then be transformed into a statistically optimal estimate of the LEMS, which could have been found by a computationally expensive direct LEMS estimation as well. The basic concept is extended to time-varying acoustic scenes and simulations of a WFS application confirm the validity of this novel approach for system identification. Christian Hofmann 0001, Walter Kellermann |
ICASSP | 2 |
| 2016 | Generalized wave-domain transforms for listening room equalization with azimuthally irregularly spaced loudspeaker arraysabstractIn reverberant environments, Listening Room Equalization (LRE) by pre-filtering of loudspeaker signals is highly desirable for premium sound reproduction systems with a high number of loudspeakers. In this contribution, the efficient concept of LRE by wave-domain adaptive filtering is extended by deriving generalized loudspeaker signal transforms for loudspeaker arrays with an irregular azimuthal spacing. In particular, a matrix approximation allows to restore the unitarity property of a transform matrix which lost the unitarity when applied to irregular arrays instead of regular ones. Simulations of adaptive LRE systems with irregular loudspeaker arrays confirm the efficacy of the proposed, novel, generalized transforms. Christian Hofmann 0001, Walter Kellermann |
ICASSP | 2 |
| 2016 | A new uncertainty decoding scheme for DNN-HMM hybrid systems with multichannel speech enhancementabstractUncertainty decoding combines a probabilistic feature description with the acoustic model of a speech recognition system. For DNN-HMM hybrid systems, this can be realized by averaging the DNN outputs produced by a finite set of feature samples (drawn from an estimated probability distribution). In this article, we employ this sampling approach in combination with a multi-microphone speech enhancement system. We propose a new strategy for generating feature samples from multichannel signals, based on modeling the spatial coherence estimates between different microphone pairs as realizations of a latent random variable. From each coherence estimate, a spectral enhancement gain is computed and an enhanced feature vector is obtained, thus producing a finite set of feature samples, of which we average the respective DNN outputs. In the experimental part, this new uncertainty decoding strategy is shown to consistently improve the recognition accuracy of a DNN-HMM hybrid system for the 8-channel REVERB Challenge task. Christian Huemmer 0001, Andreas Schwarz, Roland Maas, Hendrik Barfuss, Ramón Fernandez Astudillo, Walter Kellermann |
ICASSP | 6 |
| 2016 | Artificial Neural Network-Based Feature Combination for Spatial Voice Activity Detection
Stefan Meier, Walter Kellermann |
INTERSPEECH | 2 |
| 2016 | Ego-noise reduction using a motor data-guided multichannel dictionaryabstractWe address the problem of ego-noise reduction, i.e., suppressing the noise a robot causes by its own motions. Such noise degrades the recorded microphone signal massively such that the robot's auditory capabilities suffer. To suppress it, it is intuitive to use also motor data, since it provides additional information about the robot's joints and thereby the noise sources. We propose to fuse motor data to a recently proposed multichannel dictionary algorithm for ego-noise reduction. At training, a dictionary is learned that captures spatial and spectral characteristics of ego-noise. At testing, nonlinear classifiers are used to efficiently associate the current robot's motor state to relevant sets of entries in the learned dictionary. By this, computational load is reduced by one third in typical scenarios while achieving at least the same noise reduction performance. Moreover, we propose to train dictionaries on different microphone array geometries and use them for ego-noise reduction while the head to which the microphones are mounted is moving. In such scenarios, the motor guided approach results in significantly better performance values. Alexander Schmidt 0004, Antoine Deleforge, Walter Kellermann |
IROS | 3 |
| 2016 | Analysis of Additional Stable Gain by Frequency Shifting for Acoustic Feedback Suppression using Statistical Room AcousticsabstractCurrently, a quantitative prediction of the performance of the frequency shifting (FS) approach for stabilizing acoustic feedback loops of public address (PA) systems is only possible in case of negligibly small direct sound components in the feedback path. By employing a statistical room acoustics model, this letter derives a novel analytical expression for the additional stable gain (ASG) with the FS approach. This allows a quantitative prediction of the performance of FS in the presence of significant direct sound components for the first time. Simulation results confirm the theoretical analysis. Chengshi Zheng, Christian Hofmann 0001, Xiaodong Li 0002, Walter Kellermann |
IEEE Signal Process. Lett. | 4 |
| 2016 | Multichannel Acoustic Echo Cancellation in the Wave Domain With Increased Robustness to NonuniquenessabstractAcoustic echo cancellation (AEC) is significantly more challenging in the case of multichannel reproduction compared to the single-channel case. This is mainly due to the so-called nonuniqueness problem, which becomes especially severe for massive multichannel reproduction systems like, e.g., wave field synthesis systems or higher-order ambisonics, where the different loudspeaker signals are often strongly correlated. In this paper, a multichannel wave-domain AEC structure is proposed that maintains robustness without altering the loudspeaker signals, as is commonly required to preserve the reproduction quality. The proposed method exploits the dominance of certain couplings in the wave-domain representation of the loudspeaker-enclosure-microphone system. In this representation, preferring a strong weight of the dominant couplings in the identified system improves the system identification performance. A modification of the well-known generalized frequency-domain adaptive filtering algorithm is then presented to implement the proposed approach, including a self-contained description of all necessary components. An experimental evaluation compares the proposed method to competing approaches and shows that the proposed method compares well in terms of convergence behavior while avoiding the performance degradation, which is intrinsic to other approaches. Martin Schneider 0009, Walter Kellermann |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Trinicon-BSS system incorporating robust dual beamformers for noise reductionabstractIn this paper, a method of adaptive noise suppression combining spatially robust fixed beamforming and the TRINICON blind source separation algorithm is presented. A multichannel sensor array is first processed using complementary fixed beamformers into maximum and minimum SINR channels. The channels form the inputs to a single 2×2 second-order statistics TRINICON-BSS system which adaptively compensates for imperfections of the fixed beamformer design relative to the acoustic scenario. It is demonstrated that integrating the TRINICON-BSS algorithm leads to improved SINR performance over the initial imperfect beamformer design, and achieves a performance comparable to a perfect MVDR beamformer. Craig A. Anderson, Stefan Meier, Walter Kellermann, Paul D. Teal, Mark A. Poletti |
ICASSP | 3 |
| 2015 | Phase-optimized K-SVD for signal extraction from underdetermined multichannel sparse mixturesabstractWe propose a novel sparse representation for heavily underdetermined multichannel sound mixtures, i.e., with much more sources than microphones. The proposed approach operates in the complex Fourier domain, thus preserving spatial characteristics carried by phase differences. We derive a generalization of K-SVD which jointly estimates a dictionary capturing both spectral and spatial features, a sparse activation matrix, and all instantaneous source phases from a set of signal examples. This dictionary can be used to extract the learned signal from a new input mixture. The method is applied to the challenging problem of ego-noise reduction for robot audition. We demonstrate its superiority relative to conventional dictionary-based techniques using real-room recordings. Antoine Deleforge, Walter Kellermann |
ICASSP | 2 |
| 2015 | Spatial diffuseness features for DNN-based speech recognition in noisy and reverberant environmentsabstractWe propose a spatial diffuseness feature for deep neural network (DNN)-based automatic speech recognition to improve recognition accuracy in reverberant and noisy environments. The feature is computed in real-time from multiple microphone signals without requiring knowledge or estimation of the direction of arrival, and represents the relative amount of diffuse noise in each time and frequency bin. It is shown that using the diffuseness feature as an additional input to a DNN-based acoustic model leads to a reduced word error rate for the REVERB challenge corpus, both compared to logmelspec features extracted from noisy signals, and features enhanced by spectral subtraction. Andreas Schwarz, Christian Huemmer 0001, Roland Maas, Walter Kellermann |
ICASSP | 4 |
| 2015 | Enhanced robot audition by dynamic acoustic sensing in moving humanoidsabstractAuditory systems of humanoid robots usually acquire the surrounding sound field by means of microphone arrays. These arrays can undergo motion related to the robot's activity. The conventional approach to dealing with this motion is to stop the robot during sound acquisition. This approach avoids changing the positions of the microphones during the acquisition and reduces the robot's ego-noise. However, stopping the robot can interfere with the naturalness of its behaviour. Moreover, the potential performance improvement due to motion of the sound acquiring system can not be attained. This potential is analysed in the current paper. The analysis considers two different types of motion: (i) rotation of the robot's head and (ii) limb gestures. The study presented here combines both theoretical and numerical simulation approaches. The results show that rotation of the head improves the high-frequency performance of the microphone array positioned on the head of the robot. This is complemented by the limb gestures, which improve the low-frequency performance of the array positioned on the torso and limbs of the robot. Vladimir Tourbabin, Hendrik Barfuss, Boaz Rafaely, Walter Kellermann |
ICASSP | 4 |
| 2015 | Uncertainty decoding for DNN-HMM hybrid systems based on numerical sampling
Christian Huemmer 0001, Roland Maas, Andreas Schwarz, Ramón Fernandez Astudillo, Walter Kellermann |
INTERSPEECH | 5 |
| 2015 | The NLMS Algorithm with Time-Variant Optimum Stepsize Derived from a Bayesian Network PerspectiveabstractIn this letter, we derive a new stepsize adaptation for the normalized least mean square algorithm (NLMS) by describing the task of linear acoustic echo cancellation from a Bayesian network perspective. Similar to the well-known Kalman filter equations, we model the acoustic wave propagation from the loudspeaker to the microphone by a latent state vector and define a linear observation equation (to model the relation between the state vector and the observation) as well as a linear process equation (to model the temporal progress of the state vector). Based on additional assumptions on the statistics of the random variables in observation and process equation, we apply the expectation-maximization (EM) algorithm to derive an NLMS-like filter adaptation. By exploiting the conditional independence rules for Bayesian networks, we reveal that the resulting EM-NLMS algorithm has a stepsize update equivalent to the optimal-stepsize calculation proposed by Yamamoto and Kitayama in 1982, which has been adopted in many textbooks. As main difference, the instantaneous stepsize value is estimated in the M step of the EM algorithm (instead of being approximated by artificially extending the acoustic echo path). The EM-NLMS algorithm is experimentally verified for synthesized scenarios with both, white noise and male speech as input signal. Christian Huemmer 0001, Roland Maas, Walter Kellermann |
IEEE Signal Process. Lett. | 3 |
| 2015 | Coherent-to-Diffuse Power Ratio Estimation for DereverberationabstractThe estimation of the time- and frequency-dependent coherent-to-diffuse power ratio (CDR) from the measured spatial coherence between two omnidirectional microphones is investigated. Known CDR estimators are formulated in a common framework, illustrated using a geometric interpretation in the complex plane, and investigated with respect to bias and robustness towards model errors. Several novel unbiased CDR estimators are proposed, and it is shown that knowledge of either the direction of arrival (DOA) of the target source or the coherence of the noise field is sufficient for unbiased CDR estimation. The validity of the model for the application of CDR estimates to dereverberation is investigated using measured and simulated impulse responses. A CDR-based dereverberation system is presented and evaluated using signal-based quality measures as well as automatic speech recognition accuracy. The results show that the proposed unbiased estimators have a practical advantage over existing estimators, and that the proposed DOA-independent estimator can be used for effective blind dereverberation. Andreas Schwarz, Walter Kellermann |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Significance-aware Hammerstein group models for nonlinear acoustic echo cancellationabstractIn this work, a novel approach for nonlinear acoustic echo cancellation is proposed. The main innovative idea of the proposed method is to model only the small region of the echo path around the direct path by a group of parallel Hammerstein models, to estimate a nonlinear preprocessor by correlations between the linear kernels of the Hammerstein submodels, and to describe the remaining echo path by a simple Hammerstein model with the preprocessor determined in the aforementioned way. While the computational complexity of such a system increases only slightly in comparison to a linear echo canceller, experiments with speech recordings from a smartphone in different environments confirm a significantly increased echo cancellation performance. Christian Hofmann 0001, Christian Huemmer 0001, Walter Kellermann |
ICASSP | 3 |
| 2014 | The elitist particle filter based on evolutionary strategies as novel approach for nonlinear acoustic echo cancellationabstractIn this article, we introduce a novel approach for nonlinear acoustic echo cancellation based on a combination of particle filtering and evolutionary strategies. The nonlinear echo path is modeled as a state vector with non-Gaussian probability distribution and the relation to the observed signals and near-end interferences are captured by nonlinear functions. To estimate the probability distribution of the state vector and the model parameters, we apply the numerical sampling method of particle filtering, where each set of particles represents different realizations of the nonlinear echo path. While the classical particle-filter approach is unsuitable for system identification with large search spaces, we introduce a modified particle filter to select elitist particles based on long-term fitness measures and to create new particles based on the approximated probability distribution of the state vector. The validity of the novel approach is experimentally verified with real recordings for a nonlinear echo path stemming from a commercial smartphone. Christian Huemmer 0001, Christian Hofmann 0001, Roland Maas, Andreas Schwarz, Walter Kellermann |
ICASSP | 5 |
| 2014 | Estimation of Acoustic Reflection Coefficients Through Pseudospectrum MatchingabstractEstimating the geometric and reflective properties of the environment is important for a wide range of applications of space-time audio processing, from acoustic scene analysis to room equalization and spatial audio rendering. In this manuscript, we propose a methodology for frequency-subband in-situ estimation of the reflection coefficients of planar surfaces. This is a rather challenging task, as the reflection coefficients depend on the frequency and the angle of incidence and their estimate is highly sensitive to background noise and interfering sources. Our method is based on the assumption that we know the geometry of the reflectors; the position and the radiation pattern of the source; the position and the spatial response of the array. Applying beamforming algorithms on a single set of measured sensor data, we estimate the angular distribution of the acoustic energy (angular pseudospectrum) that impinges on a microphone array. We then apply a two-step iterative estimation technique based on an Expectation-Maximization (EM) algorithm. The first step estimates the scaling factors. The second one infers the reflection coefficients from the scaling factors. Under the assumption of additive white Gaussian noise, we finally determine the reflection coefficients with a Maximum Likelihood (ML) estimation method. The effectiveness and the accuracy of the proposed technique are assessed through experiments based on measured data. Dejan Markovic, Konrad Kowalczyk, Fabio Antonacci, Christian Hofmann 0001, Augusto Sarti, Walter Kellermann |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2014 | Minimum Mutual Information-Based Linearly Constrained Broadband Signal ExtractionabstractIn this contribution, the problem of broadband acoustic signal extraction is treated as a specific source separation problem, where the desired signal components are to be separated from all remaining undesired components. For this, we exploit the generic TRIple-N Independent component analysis for CONvolutive mixtures (TRINICON) framework. The TRINICON optimization criterion is complemented with linear constraints leading to the Linearly Constrained Minimum Mutual Information (LCMMI) criterion for desired signal extraction. A general linearly constrained update rule for iterative filter optimization is derived, which can efficiently be realized in a novel Minimum Mutual Information (MMI)-Generalized Sidelobe Canceler (GSC). The general treatment of the signal extraction problem using an MMI criterion provides several advantages: Firstly, new insights into the signal extraction problem can be derived by establishing links to both the original GSC and the Multichannel Wiener Filter (MWF). Secondly, by exploiting fundamental properties characteristic for speech and audio signals, complicated and often unreliable Voice Activity Detection (VAD)-based control mechanisms become unnecessary. Thirdly, the overall realization requires only prior information of the desired source position. An evaluation of the MMI-GSC for the double-talk situation with two concurrently active speech sources under reverberant and noisy conditions demonstrates the effectiveness of this novel approach. Klaus Reindl, Stefan Meier, Hendrik Barfuss, Walter Kellermann |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2013 | An uncertainty decoding approach to noise- and reverberation-robust speech recognitionabstractThe generic REMOS (REverberation MOdeling for robust Speech recognition) concept is extended in this contribution to cope with additional noise components. REMOS originally embeds an explicit reverberation model into a hiddenMarkov model (HMM) leading to a relaxed conditional independence assumption for the observed feature vectors. During recognition, a nonlinear optimization problem is to be solved in order to adapt the HMMs' output probability density functions to the current reverberation conditions. The extension for additional noise components necessitates a modified numerical solver for the nonlinear optimization problem. We propose an approximation scheme based on continuous piecewise linear regression. Connected-digit recognition experiments demonstrate the potential of REMOS in reverberant and noisy environments. They furthermore reveal that the benefit of an explicit reverberation model, overcoming the conditional independence assumption, increases with increasing signal-to-noise-ratios. Roland Maas, Akshaya Thippur, Armin Sehr, Walter Kellermann |
ICASSP | 4 |
| 2013 | Wave-domain loudspeaker signal decorrelation for system identification in multichannel audio reproduction scenariosabstractFor applications like acoustic echo cancellation (AEC) or listening room equalization (LRE), a loudspeaker-enclosure-microphone system (LEMS) must be identified. When using a large number of reproduction channels, as, e. g., for wave field synthesis (WFS) or Higher-Order Ambisonics (HOA), the strong correlation of the loud-speaker signals will hamper a unique identification. A state-of-the-art remedy against this so-called nonuniqueness problem is a decorrelation of the loudspeaker signals, which facilitates a unique identification. However, most of the known approaches are not suitable for acoustic wave field reproduction schemes, as they would distort the reproduced wave field in an uncontrolled manner or degrade the audio quality. In this contribution, we propose a wave-domain time-varying filtering of the loudspeaker signals, so that the reproduced wave field is rotated within a perceptually acceptable range, while preserving its shape. Martin Schneider 0009, Christian Huemmer 0001, Walter Kellermann |
ICASSP | 3 |
| 2013 | Conditional emission densities for combining speech enhancement and recognition systems
Armin Sehr, Takuya Yoshioka, Marc Delcroix, Keisuke Kinoshita, Tomohiro Nakatani, Roland Maas, Walter Kellermann |
INTERSPEECH | 7 |
| 2013 | A stereophonic acoustic signal extraction scheme for noisy and reverberant environments
Klaus Reindl, Yuanhang Zheng, Andreas Schwarz, Stefan Meier, Roland Maas, Armin Sehr, Walter Kellermann |
Comput. Speech Lang. | 7 |
| 2013 | Identification of Active Sources in Single-Channel Convolutive Mixtures Using Known Source ModelsabstractWe address the problem of identifying the constituent sources in a single-sensor mixture signal consisting of contributions from multiple simultaneously active sources. We propose a generic framework for mixture signal analysis based on a latent variable approach. The basic idea of the approach is to detect known sources represented as stochastic models, in a single-channel mixture signal without performing signal separation. A given mixture signal is modeled as a convex combination of known source models and the weights of the models are estimated using the mixture signal. We show experimentally that these weights indicate the presence/absence of the respective sources. The performance of the proposed approach is illustrated through mixture speech data in a reverberant enclosure. For the task of identifying the constituent speakers using data from a single microphone, the proposed approach is able to identify the dominant source with up to 8 simultaneously active background sources in a room with RT60= 250 ms, using models obtained from clean speech data for a Source to Interference Ratio (SIR) greater than 2 dB. Sundar Harshavardhan, Thippur V. Sreenivas, Walter Kellermann |
IEEE Signal Process. Lett. | 3 |
| 2013 | Blind System Identification Using Sparse Learning for TDOA Estimation of Room ReflectionsabstractLocalization of early room reflections can be achieved by estimating the time-differences-of-arrival (TDOAs) of reflected waves between elements of a microphone array. For an unknown source, we propose to apply sparse blind system identification (BSI) methods to identify the acoustic impulse responses, from which the TDOAs of temporally sparse reflections are estimated. The proposed time- and frequency-domain adaptive algorithms based on crossrelation formulation are regularized by incorporating an l1-norm sparseness constraint, which is realized using a split Bregman method. These algorithms are shown to outperform standard crossrelation-based BSI techniques when estimating TDOAs of reflections in the presence of background noise. Konrad Kowalczyk, Emanuël A. P. Habets, Walter Kellermann, Patrick A. Naylor |
IEEE Signal Process. Lett. | 3 |
| 2012 | On the application of reverberation suppression to robust speech recognitionabstractIn this paper, we study the effect of the design parameters of a single-channel reverberation suppression algorithm on reverberation-robust speech recognition. At the same time, reverberation compensation at the speech recognizer is investigated. The analysis reveals that it is highly beneficial to attenuate only the reverberation tail after approximately 50 ms while coping with the early reflections and residual late-reverberation by training the recognizer on moderately reverberant data. It will be shown that the overall system at its optimum configuration yields a very promising recognition performance even in strongly reverberant environments. Since the reverberation suppression algorithm is evidenced to significantly reduce the dependency on the training data, it allows for a very efficient training of acoustic models that are suitable for a wide range of reverberation conditions. Finally, experiments with an “ideal” reverberation suppression algorithm are carried out to cross-check the inferred guidelines. Roland Maas, Emanuël A. P. Habets, Armin Sehr, Walter Kellermann |
ICASSP | 4 |
| 2012 | Design of robust polynomial beamformers for symmetric arraysabstractPolynomial broadband beamforming designs enable an easy, smooth, and dynamic steering of the main beam. A number of design methods based on constrained optimization have been proposed recently which allow for the control of the robustness of these designs. Of course, the addition of robustness constraints reduces the number of degrees of freedom of the design. In this paper, we present a method to enhance the spatial selectivity of the robust polynomial beamformer design by exploiting the structure of symmetric arrays while still satisfying the robustness constraints. The effectiveness of this method is shown in design examples for symmetric linear and circular arrays. Edwin Mabande, Michael Buerger, Walter Kellermann |
ICASSP | 3 |
| 2012 | Adaptive listening room equalization using a scalable filtering structure in thewave domainabstractMassive multichannel reproduction systems like wave field synthesis (WFS) are potentially well suited to be complemented by listening room equalization (LRE). However, their typically large number of reproduction channels makes this task challenging for both computational and algorithmic reasons. Wave-domain adaptive filtering (WDAF) was proposed earlier and is especially well-suited to adaptive filtering tasks in the context of WFS. In this paper, we propose to generalize the model originally used for WDAF to allow an adaptive LRE for a broader range of reproduction scenarios, while maintaining the advantages of the original approach. The proposed approach is evaluated for filtering structures of varying complexity along with considering the robustness to varying listener positions. Martin Schneider 0009, Walter Kellermann |
ICASSP | 2 |
| 2012 | A two-channel reverberation suppression scheme based on blind signal separation and wiener filteringabstractIn this paper, we apply a blind signal extraction scheme for two microphones to the problem of dereverberation. The system consists of a blocking matrix that cancels the target signal as well as reverberated components up to a certain time lag, thus obtaining a reference not only for noise and interference, but also for late reverberation, which can then be suppressed with a Wiener filter, while leaving early reverberation components largely intact. The performance is assessed in terms of recognition rate of an automatic speech recognizer trained on clean speech, using sentences from the GRID corpus convolved with measured room impulse responses. We show that the system, although primarily developed for noise and interference suppression in low SNR conditions, can significantly suppress reverberation and thereby improve recognition results. Andreas Schwarz, Klaus Reindl, Walter Kellermann |
ICASSP | 3 |
| 2011 | Synthesis of ICA-based methods for localization of multiple broadband sound sourcesabstractIn this paper, minimization of the statistical dependence is exploited for acoustic source localization purposes. Originally developed for the separation of signal mixtures, we show that Independent Component Analysis (ICA) can also be successfully applied to localize multiple simultaneously active sound sources, with possibly less sensors than sources. First, the recently proposed Averaged Directivity Pattern (ADP) and State Coherence Transform (SCT) methods are reviewed. Similarities and differences between both approaches are underlined and analyzed, leading to a new method merging elements from both concepts, which we call the Modified ADP (MADP). Since the investigated methods do not suffer from the permutation ambiguity, they can be applied in combination with any narrowband or broadband ICA algorithm, without the need to solve the still challenging permutation issue. Experimental results are presented for speech sources in a reverberant environment. Anthony Lombard, Yuanhang Zheng, Walter Kellermann |
ICASSP | 3 |
| 2011 | On 2D localization of reflectors using robust beamforming techniquesabstractThis paper presents a method for the localization of reflectors in an acoustic environment, using robust beamforming techniques and a cylindrical microphone array, for which an intuitive and highly efficient three-step procedure is proposed. First, the directions of ar rival (DOAs) corresponding to the sound source and reflectors are estimated by a robust Minimum Variance Distortionless Response (MVDR) beamformer. Next, signals originating from the estimated DOAs are extracted by a robust superdirective beamformer, from which time differences of arrival (TDOAs) of major reflections are estimated by crosscorrelation analysis. Finally, by using additional information about the position of the direct sound source relative to the array, the positions of the reflective boundaries of the room can be inferred. Experiments based on real measurements carried out in a moderately reverberant room show the effectiveness of this method. Edwin Mabande, Haohai Sun, Konrad Kowalczyk, Walter Kellermann |
ICASSP | 4 |
| 2011 | Frame-wise HMM adaptation using state-dependent reverberation estimatesabstractA novel frame-wise model adaptation approach for reverberation robust distant-talking speech recognition is proposed. It adjusts the means of static cepstral features to capture the statistics of reverber ant feature vector sequences obtained from distant-talking speech recordings. The means of the HMMs are adapted during decoding using a state-dependent estimate of the late reverberation determined by joint use of a feature-domain reverberation model and optimum partial state sequences. Since the parameters of the HMMs and the reverberation model can be estimated completely independently, the approach is very flexible with respect to changing acoustic environments. Due to the frame-wise model adaptation, some of the HMM limitations are relieved, and recognition results surpassing that of matched reverberant training are obtained at the cost of a moderately increased decoding complexity. Armin Sehr, Roland Maas, Walter Kellermann |
ICASSP | 3 |
| 2011 | Joint DOA and TDOA estimation for 3D localization of reflective surfaces using eigenbeam MVDR and spherical microphone arraysabstractMethods of 3D direction of arrival (DOA) estimation, coherent source detection and reflective surface localization are studied, based on recordings by a spherical microphone array. First, the spherical harmonics domain minimum variance distortionless response (EB-MVDR) beamformer is employed for the localization of broadband coherent sources, which is characterized by simpler frequency focusing matrices than the corresponding element-space implementation, and by a higher resolution than conventional spherical array beamformers. After the DOA estimation step, the source signals are extracted by EB-MVDRs. Then, by computing the crosscorrelation functions between the extracted signals, the coherent sources are detected and their time differences of arrival (TDOA) are estimated. Given the positions of the array and the reference source, and the estimated DOA and TDOA of the coherent sources, the positions of the major reflectors can be inferred. Experimental results in a real room validate the proposed method. Haohai Sun, Edwin Mabande, Konrad Kowalczyk, Walter Kellermann |
ICASSP | 4 |
| 2011 | Robust localization of multiple sources in reverberant environments using EB-ESPRIT with spherical microphone arraysabstractSpherical microphone array eigenbeam (EB)-ESPRIT gives an elegant closed-form solution for 3D broadband source localization based on the spherical harmonics (eigenbeam) framework. However, in practical implementations, there are still several issues not being rigorously studied, e.g. how to avoid the ill-conditioning of an EB-ESPRIT matrix, solve the ambiguity problem, handle a large number of sources, and localize coherent broadband sources, etc. In this work, we propose to use the condition number of the EB-ESPRIT matrix as a robustness measure, and to use Wigner-D weighting to avoid the ill-conditioning issue and improve the robustness. In addition, power spectrum testing, frequency smoothing, and manifold vector extension techniques are employed to address the ambiguity, coherent source localization, and large source number problems, respectively. Experimental results based on measurements taken with a real spherical microphone array in a room environment show the effectiveness of the proposed methods. Haohai Sun, Heinz Teutsch, Edwin Mabande, Walter Kellermann |
ICASSP | 4 |
| 2011 | Adaptive Combination of Volterra Kernels and Its Application to Nonlinear Acoustic Echo CancellationabstractThe combination of filters concept is a simple and flexible method to circumvent various compromises hampering the operation of adaptive linear filters. Recently, applications which require the identification of not only linear, but also nonlinear systems are widely studied. In this paper, we propose a combination of adaptive Volterra filters as the most versatile nonlinear models with memory. Moreover, we develop a novel approach that shows a similar behavior but significantly reduces the computational load by combining Volterra kernels rather than complete Volterra filters. Following an outline of the basic principles, the second part of the paper focuses on the application to nonlinear acoustic echo cancellation scenarios. As the ratio of the linear to nonlinear echo signal power is, in general, a priori unknown and time-variant, the performance of nonlinear echo cancellers may be inferior to a linear echo canceller if the nonlinear distortion is very low. Therefore, a modified version of the combination of kernels is developed obtaining a robust behavior regardless of the level of nonlinear distortion. Experiments with noise and speech signals demonstrate the desired behavior and the robustness of both the combination of Volterra filters and the combination of kernels approaches in different application scenarios. Luis Antonio Azpicueta-Ruiz, Marcus Zeller, Aníbal R. Figueiras-Vidal, Jerónimo Arenas-García, Walter Kellermann |
IEEE Trans. Speech Audio Process. | 5 |
| 2011 | TDOA Estimation for Multiple Sound Sources in Noisy and Reverberant Environments Using Broadband Independent Component AnalysisabstractIn this paper, we show that minimization of the statistical dependence using broadband independent component analysis (ICA) can be successfully exploited for acoustic source localization. As the ICA signal model inherently accounts for the presence of several sources and multiple sound propagation paths, the ICA criterion offers a theoretically more rigorous framework than conventional techniques based on an idealized single-path and single-source signal model. This leads to algorithms which outperform other localization methods, especially in the presence of multiple simultaneously active sound sources and under adverse conditions, notably in reverberant environments. Three methods are investigated to extract the time difference of arrival (TDOA) information contained in the filters of a two-channel broadband ICA scheme. While for the first, the blind system identification (BSI) approach, the number of sources should be restricted to the number of sensors, the other methods, the averaged directivity pattern (ADP) and composite mapped filter (CMF) approaches can be used even when the number of sources exceeds the number of sensors. To allow fast tracking of moving sources, the ICA algorithm operates in block-wise batch mode, with a proportionate weighting of the natural gradient to speed up the convergence of the algorithm. The TDOA estimation accuracy of the proposed schemes is assessed in highly noisy and reverberant environments for two, three, and four stationary noise sources with speech-weighted spectral envelopes as well as for moving real speech sources. Anthony Lombard, Yuanhang Zheng, Herbert Buchner, Walter Kellermann |
IEEE Trans. Speech Audio Process. | 4 |
| 2010 | Model-based dereverberation in the logmelspec domain for robust distant-talking speech recognitionabstractThe REMOS (REverberation MOdeling for Speech recognition) concept for reverberation-robust distant-talking speech recognition, introduced in for melspectral features, is extended in this contribution to logarithmic melspectral (logmelspec) features. Based on a combined acoustic model consisting of a hidden Markov model network and a reverberation model, REMOS determines clean-speech and reverberation estimates during recognition by an inner optimization operation. A reformulation of this inner optimization problem for logmelspec features, allowing an efficient solution by nonlinear optimization algorithms, is derived in this paper so that an efficient implementation of REMOS for logmelspec features becomes possible. Connected digit recognition experiments show that the proposed REMOS implementation significantly outperforms reverberantly-trained HMMs in highly reverberant environments. Armin Sehr, Roland Maas, Walter Kellermann |
ICASSP | 3 |
| 2010 | Disambiguation in multidimensional tracking of multiple acoustic sources using a Gaussian likelihood criterionabstractThe TDOA-based acoustic source localization approaches suffer from a pairing problem in the case of multiple dimensions and multiple sources. To resolve this problem, usually three or more microphone pairs are required. In this paper, we propose a new solution based on a Gaussian likelihood function to pair TDOAs, which allows us to localize multiple moving acoustic sources in several dimensions using a minimum number of microphone pairs. Moreover, the proposed method is combined with a particle filter for tracking multiple moving sources. Experimental results demonstrate the effectiveness of the proposed scheme. Pengxiao Teng, Anthony Lombard, Walter Kellermann |
ICASSP | 3 |
| 2010 | Efficient adaptive DFT-domain Volterra filters using an automatically controlled number of quadratic kernel diagonalsabstractThis paper presents a method for estimating the optimum number of second-order kernel diagonals of an adaptive Volterra filter in system identification tasks. To this end, a recently proposed time-domain mechanism is carried over to the very efficient partitioned-block DFT-domain Volterra filtering technique. The size of the nonlinear memory is controlled by monitoring the performance of an adaptive combination scheme with two differently-sized quadratic kernels. Subsequently, an efficient version is derived, requiring only minor additional computations as compared to a single Volterra filter. The effectiveness of the outlined estimation procedure is demonstrated by various simulations with real nonlinear systems and both noise and speech inputs in an acoustic echo cancellation scenario. Marcus Zeller, Luis Antonio Azpicueta-Ruiz, Jerónimo Arenas-García, Walter Kellermann |
ICASSP | 4 |
| 2010 | A novel approach for matched reverberant training of HMMs using data pairs
Armin Sehr, Christian Hofmann 0001, Roland Maas, Walter Kellermann |
INTERSPEECH | 4 |
| 2010 | Introduction to the Special Issue on Processing Reverberant Speech: Methodologies and ApplicationsabstractThe 17 papers in this special issue focus on the methodologies and applications of processing reverberant speech. The issue highlights some major aspects of the recent progress in the field. Tomohiro Nakatani, Walter Kellermann, Patrick A. Naylor, Masato Miyoshi, Biing-Hwang Juang |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Reverberation Model-Based Decoding in the Logmelspec Domain for Robust Distant-Talking Speech RecognitionabstractThe REMOS (REverberation MOdeling for Speech recognition) concept for reverberation-robust distant-talking speech recognition, introduced in “Distant-talking continuous speech recognition based on a novel reverberation model in the feature domain” (A. Sehr , in Proc. Interspeech, 2006, pp. 769-772) for melspectral features, is extended to logarithmic melspectral (logmelspec) features in this contribution. Thus, the favorable properties of REMOS, including its high flexibility with respect to changing reverberation conditions, become available in the more competitive logmelspec domain. Based on a combined acoustic model consisting of a hidden Markov model (HMM) network and a reverberation model (RM), REMOS determines clean-speech and reverberation estimates during recognition. Therefore, in each iteration of a modified Viterbi algorithm, an inner optimization operation maximizes the joint density of the current HMM output and the RM output subject to the constraint that their combination is equal to the current reverberant observation. Since the combination operation in the logmelspec domain is nonlinear, numerical methods appear necessary for solving the constrained inner optimization problem. A novel reformulation of the constraint, which allows for an efficient solution by nonlinear optimization algorithms, is derived in this paper so that a practicable implementation of REMOS for logmelspec features becomes possible. An in-depth analysis of this REMOS implementation investigates the statistical properties of its reverberation estimates and thus derives possibilities for further improving the performance of REMOS. Connected digit recognition experiments show that the proposed REMOS version in the logmelspec domain significantly outperforms the melspec version. While the proposed RMs with parameters estimated by straightforward training for a given room are robust to a mismatch of the speaker-microphone distance, their performance significantly decreases if they are used in a room with substantially different conditions. However, by training multi-style RMs with data from several rooms, good performance can be achieved across different rooms. Armin Sehr, Roland Maas, Walter Kellermann |
IEEE Trans. Speech Audio Process. | 3 |
| 2009 | Novel schemes for nonlinear acoustic echo cancellation based on filter combinationsabstractNonlinear acoustic echo cancellers (NLAEC) are becoming increasingly important in hands-free applications. However, in some situations, an NLAEC is inferior to a linear AEC, especially when the channel generates a negligible (or no) nonlinear echo. In general, the ratio of the linear to nonlinear echo signal power is unknown a priori, and will vary over time, thus making it difficult to know if an NLAEC would improve or degrade the cancellation. In this paper, we present two novel solutions to this problem based on the adaptive combination of linear and nonlinear echo cancellers. Both solutions perform efficiently regardless of the level of nonlinear echo. The benefits and robustness of both schemes are illustrated by experiments using Laplacian colored noise and speech input signals. Luis Antonio Azpicueta-Ruiz, Marcus Zeller, Jerónimo Arenas-García, Walter Kellermann |
ICASSP | 4 |
| 2009 | Multidimensional localization of multiple sound sources using averaged directivity patterns of Blind Source Separation systemsabstractIn this paper, we propose a versatile acoustic source localization framework exploiting the self-steering capability of Blind Source Separation (BSS) algorithms. We provide a way to produce an acoustical map of the scene by computing the averaged directivity pattern of BSS demixing systems. Since BSS explicitly accounts for multiple sources in its signal propagation model, several simultaneously active sound sources can be located using this method. Moreover, the framework is suitable to any microphone array geometry, which allows application for multiple dimensions, in the near field as well as in the far field. Experiments demonstrate the efficiency of the proposed scheme in a reverberant environment for the localization of speech sources. Anthony Lombard, Tobias Rosenkranz, Herbert Buchner, Walter Kellermann |
ICASSP | 4 |
| 2009 | Design of robust superdirective beamformers as a convex optimization problemabstractBroadband data-independent beamforming designs aiming at constant beamwidth often lead to superdirective beamformers for low frequencies, if the sensor spacing is small relative to the wavelengths. Superdirective beamformers are extremely sensitive to spatially white noise and to small errors in the array characteristics. These errors are nearly uncorrelated from sensor to sensor and affect the beamformer in a manner similar to spatially white noise. Hence the White Noise Gain (WNG) is a commonly used measure for the robustness of beamformer designs. In this paper, we present a method which incorporates a constraint for the WNG into a least-squares beamformer design and still leads to a convex optimization problem that can be solved directly, e.g. by sequential quadratic programming. The effectiveness of this method is demonstrated by design examples. Edwin Mabande, Adrian Schad, Walter Kellermann |
ICASSP | 3 |
| 2009 | Strategies for modeling reverberant speech in the feature domainabstractThe length of the room impulse response characterizing the acoustic path between speaker and microphone is significantly larger than the length of the analysis window used for feature extraction in automatic speech recognition (ASR) systems. Therefore, reverberation caused by multi-path propagation of sound waves from the speaker to distant-talking microphones has a dispersive effect on speech feature sequences. This dispersive effect causes a mismatch between the input speech and the acoustic models of the recognizer, usually trained on clean speech, and leads to a significant reduction of recognition performance. In this contribution, different strategies for obtaining acoustic models capturing the dispersive effect of reverberation are investigated in terms of modeling accuracy, flexibility with respect to changing reverberation conditions, effort for obtaining the reverberation representation and decoding complexity. Armin Sehr, Walter Kellermann |
ICASSP | 2 |
| 2009 | Online estimation of the optimum quadratic kernel size of second-order Volterra filters using a convex combination schemeabstractThis paper presents a method for estimating the optimum memory size for identification of an unknown second-order Volterra kernel. As these structures may imply considerable computational demands, it is highly desirable to design adaptive realizations with a minimum number of coefficients. Therefore, we propose a combination scheme comprising two Volterra filters with time-variant sizes of the actually used quadratic kernels. By following some simple rules, the number of diagonals in the quadratic kernels is increased or decreased in order to find the optimum memory configuration in parallel to the coefficient adaptation. Thus, the arbitrary choice of the nonlinear system size is overcome by a dynamically growing/shrinking system. Experimental results for various signals and nonlinear scenarios demonstrate the effectiveness of the proposed method. Marcus Zeller, Luis Antonio Azpicueta-Ruiz, Walter Kellermann |
ICASSP | 3 |
| 2008 | Detection and localization of multiple wideband acoustic sources based on wavefield decomposition using spherical aperturesabstractThis paper discusses novel methods for detecting and localizing multiple wideband acoustic sources using spherical apertures. In contrast to traditional methods the techniques presented here are not based on processing the output of individual microphones directly. Instead, the microphone signals are used to decompose the wave- field into its spherical harmonics which are subsequently used as a basis for novel frequency-independent multiple-source localization and detection methods. Heinz Teutsch, Walter Kellermann |
ICASSP | 2 |
| 2008 | WOZ Acoustic Data Collection for Interactive TV
Alessio Brutti, Luca Cristoforetti, Walter Kellermann, Lutz Marquardt, Maurizio Omologo |
LREC | 3 |
| 2007 | Multi-Channel Source Separation Preserving Spatial InformationabstractIn this paper we propose two novel methods for preserving the spatial information in source separation algorithms. Our approach is applicable to any source separation algorithm and is based on an additional supervised adaptive filtering with the reference signals generated by the source separation system. If a special constrained optimization scheme is applied to derive the source separation algorithm then the novel approach can be simplified. The quality of the spatial representation and the separation performance of both methods and two state-of-the-art approaches from the literature have been evaluated by a MUSHRA listening test according to the relevant ITU recommendation showing that the novel methods clearly outperform the state-of-the-art approaches. Robert Aichner, Herbert Buchner, Meray Zourub, Walter Kellermann |
ICASSP (1) | 4 |
| 2007 | Acoustic Echo Cancellation for Surround Sound using Perceptually Motivated Convergence EnhancementabstractAcoustic echo cancellation (AEC) has become an essential and well-known enabling technology for hands-free communication and human-machine interfaces. AEC for two or more reproduction channels aims at identifying the echo paths between the microphone and each audio reproduction source in order to cancel the associated echo contribution. A number of preprocessing methods have been proposed to decorrelate stereo audio signals in order to enable an unambiguous identification of each echo path and to thus ensure robustness to changing sound source locations. While several of these methods provide enough decorrelation to achieve proper AEC convergence in the stereo case, considerations of subjective sound quality have frequently not been addressed adequately. This paper compares the performance of several methods in terms of both convergence speed and aspects of sound perception, and proposes a novel signal decorrelation approach with attractive properties. The superior performance of the proposed method is demonstrated for 5.1 surround sound reproduction. Jürgen Herre, Herbert Buchner, Walter Kellermann |
ICASSP (1) | 3 |
| 2007 | Nonlinear Residual Echo Suppression using a Power Filter Model of the Acoustic Echo PathabstractLoudspeakers and amplifiers of mobile communication receivers may cause significant nonlinear distortion in the acoustic echo path, resulting in a limitation of the performance of linear echo cancelers. In this contribution, we present a nonlinear acoustic echo suppressor in order to increase the attenuation of the nonlinearly distorted residual echo. The proposed approach is based on a power filter model of the acoustic echo path which is applied to the estimation of the power spectral density of the nonlinear residual echo. These time-variant estimates are used to appropriately adjust the frequency-dependent gain values of the echo suppressor. The performance of the proposed approach is evaluated in realistic experimental set-ups. Fabian Kuech, Walter Kellermann |
ICASSP (1) | 2 |
| 2007 | A New Concept for Feature-Domain Dereverberation for Robust Distant-Talking ASRabstractThe feature-domain dereverberation capabilities of a novel approach for automatic speech recognition in reverberant environments are investigated in this paper. By combining a network of clean speech HMMs and a reverberation model, the most likely combination of the HMM output and the reverberation model output is found during decoding time by an extended version of the Viterbi algorithm. We show in this paper that the most likely HMM output represents a good estimate of the clean speech feature sequence and can be used as input to subsequent speech recognizers. Armin Sehr, Walter Kellermann |
ICASSP (4) | 2 |
| 2007 | Iterated Coefficient Updates of Partitioned Block Frequency Domain Second-Order Volterra Filters for Nonlinear AECabstractThis paper presents the benefits of iterated coefficient updates for the adaptation of partitioned block frequency domain second-order Volterra filters when applied to nonlinear acoustic echo cancellation. In order to increase the convergence speed of an NLMS algorithm with separate kernel normalization, each input frame is used for several coefficient updates. This procedure effectively accelerates the convergence of the employed adaptive Volterra filters and is shown to be superior to processing with increased data overlap. The advantages of this novel approach are illustrated by experimental results for noise and speech input and guidelines for determining suitable numbers of iterations for the filter kernels are given. Marcus Zeller, Walter Kellermann |
ICASSP (3) | 2 |
| 2007 | Improved Wideband Blind Adaptive System Identification Using Decorrelation Filters for the Localization of Multiple SpeakersabstractThis paper addresses the TDOA extraction problem for localizing multiple sources in noisy and reverberant environments with emphasis on speech excitation. TDOAs are estimated by performing blind adaptive MIMO system identification using a gradient-based BSS variant of the TRINICON framework. We present a novel method to improve the TDOA estimation for signals with lowpass-like spectral characteristics such as speech, for which the standard approach achieves only imperfect system identification. To this end, we propose to combine the BSS algorithm with decorrelation filters, thereby achieving a greatly improved wideband identification of the acoustical system. The approach is verified in a number of scenarios, where it provides more accurate TDOA estimates for the speaker localization at a negligible additional computational cost. Anthony Lombard, Herbert Buchner, Walter Kellermann |
ISCAS | 3 |
| 2007 | Multichannel Bin-Wise Robust Frequency-Domain Adaptive Filtering and Its Application to Adaptive BeamformingabstractLeast-squares error (LSE) or mean-squared error (MSE) optimization criteria lead to adaptive filters that are highly sensitive to impulsive noise. The sensitivity to noise bursts increases with the convergence speed of the adaptation algorithm and limits the performance of signal processing algorithms, especially when fast convergence is required, as for example, in adaptive beamforming for speech and audio signal acquisition or acoustic echo cancellation. In these applications, noise bursts are frequently due to undetected double-talk. In this paper, we present impulsive noise robust multichannel frequency-domain adaptive filters (MC-FDAFs) based on outlier-robust M-estimation using a Newton algorithm and a discrete Newton algorithm, which are especially designed for frequency bin-wise adaptation control. Bin-wise adaptation and control in the frequency-domain enables the application of the outlier-robust MC-FDAFs to a generalized sidelobe canceler (GSC) using an adaptive blocking matrix for speech and audio signal acquisition. It is shown that the improved robustness leads to faster convergence and to higher interference suppression relative to nonrobust adaptation algorithms, especially during periods of strong interference Wolfgang Herbordt, Herbert Buchner, Satoshi Nakamura 0001, Walter Kellermann |
IEEE Trans. Speech Audio Process. | 4 |
| 2006 | Post-Processing for Convolutive Blind Source SeparationabstractConvolutive blind source separation (BSS) aims at separating point sources from mixtures picked up by several sensors. In real-world environments moving speakers, background noise and long reverberation are encountered which often degrade the performance of BSS algorithms. In such cases, the application of a post-filter can improve the output signal quality by suppression of residual cross-talk and of background noise. In this paper we propose a novel technique to estimate the necessary power spectral densities of the cross-talk components and present a robust system which allows to further suppress both, the remaining interference from point sources and the background noise. Experimental results show the benefit of this post-processing method in realistic environments Robert Aichner, Meray Zourub, Herbert Buchner, Walter Kellermann |
ICASSP (5) | 4 |
| 2006 | Separating Convolutive Mixtures with TriniconabstractBlind source separation (BSS) algorithms are often categorized as either narrowband or broadband algorithms depending on whether their respective cost functions aim at individual DFT bins or the entire broadband signal. In this contribution, we present comparable general natural gradient-based formulations of both concepts based on the TRINICON framework. As a distinctive feature, narrowband algorithms imply an internal permutation and scaling problem. We show that the common DOA estimation-based methods for aligning the permutations effectively rely on geometric a-priori knowledge, and we explain why they need to be complemented by additional repair mechanisms for robust BSS. The latter can already be viewed as approximations of the generic TRINICON broadband algorithm. As a conclusion, we propose to always use a generic broadband algorithm as a starting point for the design of new BSS algorithms Walter Kellermann, Herbert Buchner, Robert Aichner |
ICASSP (5) | 1 |
| 2006 | Distant-talking continuous speech recognition based on a novel reverberation model in the feature domain
Armin Sehr, Marcus Zeller, Walter Kellermann |
INTERSPEECH | 3 |
| 2006 | A real-time blind source separation scheme and its application to reverberant and noisy acoustic environments
Robert Aichner, Herbert Buchner, Walter Kellermann |
Signal Process. | 4 |
| 2006 | Orthogonalized power filters for nonlinear acoustic echo cancellation
Fabian Kuech, Walter Kellermann |
Signal Process. | 2 |
| 2006 | Multichannel parametric speech enhancementabstractWe present a parametric model-based multichannel approach for speech enhancement. By employing an autoregressive model for the speech signal and using a trained codebook of speech linear predictive coefficients, minimum mean square error estimation of the speech signal is performed. By explicitly accounting for steering errors in the signal model, robust estimates are obtained. Experiments show that the proposed method results in significant performance gains. Sriram Srinivasan 0003, Robert Aichner, W. Bastiaan Kleijn, Walter Kellermann |
IEEE Signal Process. Lett. | 4 |
| 2006 | Robust extended multidelay filter and double-talk detector for acoustic echo cancellationabstractWe propose an integrated acoustic echo cancellation solution based on a novel class of efficient and robust adaptive algorithms in the frequency domain, the extended multidelay filter (EMDF). The approach is tailored to very long adaptive filters and highly auto-correlated input signals as they arise in wideband full-duplex audio applications. The EMDF algorithm allows an attractive tradeoff between the well-known multidelay filter and the recursive least-squares algorithm. It exhibits fast convergence, superior tracking capabilities of the signal statistics, and very low delay. The low computational complexity of the conventional frequency-domain adaptive algorithms can be maintained thanks to efficient fast realizations. We also show how this approach can be combined efficiently with a suitable double-talk detector (DTD). We consider a corresponding extension of a recently proposed DTD based on a normalized cross-correlation vector whose performance was shown to be superior compared to other DTDs based on the cross-correlation coefficient. Since the resulting DTD also has an EMDF structure it is easy to implement, and the fast realization also carries over to the DTD scheme. Moreover, as the robustness issue during double talk is particularly crucial for fast-converging algorithms, we apply the concept of robust statistics into our extended frequency-domain approach. Due to the robust generalization of the cost function leading to a so-called M-estimator, the algorithms become inherently less sensitive to outliers, i.e., short bursts that may be caused by inevitable detection failures of a DTD. The proposed structure is also well suited for an efficient generalization to the multichannel case Herbert Buchner, Jacob Benesty, Tomas Gänsler, Walter Kellermann |
IEEE Trans. Speech Audio Process. | 4 |
| 2005 | On the causality problem in time-domain blind source separation and deconvolution algorithmsabstractUsing a recently presented generic framework for multichannel blind signal processing for convolutive mixtures, we investigate the problem of incorporating acausal delays which are necessary with certain geometric constellations. Starting from a generic update equation which is applicable to blind source separation (BSS), multichannel blind deconvolution (MCBD), and multichannel blind partial deconvolution (MCBPD) for dereverberation of speech signals, two formulations of the natural gradient are derived. It is shown that one expression is applicable to mere causal filters whereas the other also allows an implementation of noncausal filters. Moreover, proper initialization methods for both cases are given. For the implementation of these algorithms, cross-relation estimation techniques, known from linear prediction, are discussed. Based on these results, relationships between traditional MCBD algorithms can be established. Experimental results of different acoustic scenarios show the applicability of the presented algorithms. Robert Aichner, Herbert Buchner, Walter Kellermann |
ICASSP (5) | 3 |
| 2005 | Simultaneous localization of multiple sound sources using blind adaptive MIMO filteringabstractBlind adaptive filtering for time delay of arrival (TDOA) estimation is a very powerful method for acoustic source localization in reverberant environments with broadband signals like speech. Based on a recently presented generic framework for blind signal processing for convolutive mixtures, called TRINICON, we present a TDOA estimation method for simultaneous multidimensional localization of multiple sources. Moreover, an interesting link to the known single-input multiple-output (SIMO)-based adaptive eigenvalue decomposition (AED) method is shown. We evaluate the novel multiple-input multiple-output (MIMO)-based approach and compare it with the known SIMO-based method in a reverberant acoustic environment using reference data of the positions obtained from infrared sensors. The results show that the new approach is very robust against reverberation and background noise. Herbert Buchner, Robert Aichner, Jochen Stenglein, Heinz Teutsch, Walter Kellermann |
ICASSP (3) | 5 |
| 2005 | Joint optimization of LCMV beamforming and acoustic echo cancellation for automatic speech recognitionabstractFor full-duplex hands-free acoustic human/machine interfaces, a combination of acoustic echo cancellation and speech enhancement is often required in order to suppress acoustic echoes, local interference and noise. In order to exploit positive synergies between acoustic echo cancellation and speech enhancement optimally, we previously presented a combined least-squares (LS) optimization criterion for the integration of acoustic echo cancellation and adaptive linearly-constrained minimum variance (LCMV) beamforming (Herbordt, W. et al., Proc. EURASIP European Sig. Process. Conf., 2004). By means of speech recognition experiments, we now illustrate the efficiency of the proposed solution in situations with high levels of background noise and with time-varying echo paths and frequent double-talk. Wolfgang Herbordt, Satoshi Nakamura 0001, Walter Kellermann |
ICASSP (3) | 3 |
| 2005 | Nonlinear acoustic echo cancellation using adaptive orthogonalized power filtersabstractIn acoustic echo cancellation as, e.g., for mobile communication receivers, loudspeakers and their amplifiers cause significant nonlinear distortion in the echo path, resulting in a degradation of the performance of linear echo cancelers. In order to cope with this type of nonlinear echo path, we discuss an orthogonalized version of power filter that can be considered as a parallelized realization of the cascade of a memoryless polynomial followed by a linear filter. As, in the echo cancellation context, the statistics of the speech input are non-stationary and not known in advance, the orthogonalization follows the signal statistics. The performance of the resulting novel nonlinear structure is evaluated by experiments using real hardware. Fabian Kuech, Andreas Mitnacht, Walter Kellermann |
ICASSP (3) | 3 |
| 2005 | EB-ESPRIT: 2D localization of multiple wideband acoustic sources using eigen-beamsabstractThis paper is concerned with the problem of localizing multiple wideband acoustic sources. In contrast to existing techniques, this method takes the physics of wave propagation into account. 2D wave fields are decomposed using cylindrical harmonics as basis functions by a circular array mounted into a rigid cylindrical baffle. The obtained wave field representation is then used to serve as a basis for high-resolution subspace beamforming methods, most notably ESPRIT. It is shown that acoustic source localization based on wave field decomposition has the potential to unambiguously localize multiple simultaneously active wideband sources in the array's full 360 degrees field-of-view. Heinz Teutsch, Walter Kellermann |
ICASSP (3) | 2 |
| 2005 | Precise Visibility Determination of Displays in Camera ImagesabstractIn many novel application scenarios such as smart rooms or sensing rooms visual sensors (such as cameras) need to know which visual actuators (such as displays) are visible to them. Often only parts of a display are visible from a camera. Therefore, a novel algorithm for precise visibility determination is presented. The algorithm makes the assumption that the displays are active, i.e., they can be controlled by the application. Under these conditions the algorithm determines precisely where which parts of a display are imaged by a camera Eva Hörster, Rainer Lienhart, Walter Kellermann, Jean-Yves Bouguet |
ICME | 3 |
| 2005 | Generalized multichannel frequency-domain adaptive filtering: efficient realization and application to hands-free speech communication
Herbert Buchner, Jacob Benesty, Walter Kellermann |
Signal Process. | 3 |
| 2005 | A generalization of blind source separation algorithms for convolutive mixtures based on second-order statisticsabstractWe present a general broadband approach to blind source separation (BSS) for convolutive mixtures based on second-order statistics. This avoids several known limitations of the conventional narrowband approximation, such as the internal permutation problem. In contrast to traditional narrowband approaches, the new framework simultaneously exploits the nonwhiteness property and nonstationarity property of the source signals. Using a novel matrix formulation, we rigorously derive the corresponding time-domain and frequency-domain broadband algorithms by generalizing a known cost-function which inherently allows joint optimization for several time-lags of the correlations. Based on the broadband approach time-domain, constraints are obtained which provide a deeper understanding of the internal permutation problem in traditional narrowband frequency-domain BSS. For both the time-domain and the frequency-domain versions, we discuss links to well-known, and also, to novel algorithms that constitute special cases. Moreover, using the so-called generalized coherence, links between the time-domain and the frequency-domain algorithms can be established, showing that our cost function leads to an update equation with an inherent normalization ensuring a robust adaptation behavior. The concept is applicable to offline, online, and block-online algorithms by introducing a general weighting function allowing for tracking of time-varying real acoustic environments. Herbert Buchner, Robert Aichner, Walter Kellermann |
IEEE Trans. Speech Audio Process. | 3 |
| 2004 | TRINICON: a versatile framework for multichannel blind signal processingabstractIn this paper we present a framework for multichannel blind signal processing for convolutive mixtures, such as blind source separation (BSS) and multichannel blind deconvolution (MCBD). It is based on the use of multivariate pdf and a compact matrix notation which considerably simplifies the representation and handling of the algorithms. By introducing these techniques into an information theoretic cost function, we can exploit the three fundamental signal properties nonwhiteness, nongaussianity, and nonstationarity. This results in a versatile tool that we call TRINICON (Triple-N ICA for convolutive mixtures). Both, links to popular algorithms and several novel algorithms follow from the general approach. In particular, we introduce a new concept of multichannel blind partial deconvolution (MCBPD) for speech which prevents a complete whitening of the output signals, i.e., the vocal tract is excluded from the equalization. This is especially interesting for automatic speech recognition applications. Moreover, we show results for BSS using multivariate spherically invariant random processes (SIRP) to efficiently model speech, and show how the approach carries over to MCBPD. These concepts are also suitable for an efficient implementation in the frequency domain by using a rigorous broadband derivation avoiding the internal permutation problem and circularity effects. Herbert Buchner, Robert Aichner, Walter Kellermann |
ICASSP (3) | 3 |
| 2004 | Wave-domain adaptive filtering: acoustic echo cancellation for full-duplex systems based on wave-field synthesisabstractFor high-quality multimedia communication systems such as teleconferencing or virtual reality applications, multichannel sound reproduction is highly desirable. While progress has been made in stereo and multichannel acoustic echo cancellation (MC AEC) in recent years, the corresponding sound reproduction systems still imply a restrained listening area ('sweet spot'). A volume solution for spatial sound in a large listening area is offered by wave field synthesis (WFS) or by ambisonics, where arrays of loudspeakers generate a prespecified sound field. However, before this new technique can be utilized for full-duplex systems, an efficient solution to the MC AEC problem has to be found. This paper presents a novel approach that extends the current state of the art of MC AEC and transform-domain adaptive filtering by reconciling the flexibility of adaptive filtering and the underlying physics of acoustic waves in a systematic and efficient way. To achieve this, the new framework of wave-domain adaptive filtering (WDAF) exploits the spatial information provided by densely sampled contours for both recording and reproduction. Experimental results with a 32-channel AEC verify the concept for both simulated and actually measured room acoustics. Herbert Buchner, Sascha Spors, Walter Kellermann |
ICASSP (4) | 3 |
| 2004 | A novel multidelay adaptive algorithm for Volterra filters in diagonal coordinate representation [nonlinear acoustic echo cancellation example]abstractIn this contribution, we present a novel algorithm for the efficient computation of the output of Volterra filters in diagonal coordinate representation (DCR), while allowing for different memory length for each kernel. This is achieved by extending partitioned block filtering methods for fast convolution in the discrete Fourier transform (DFT) domain to Volterra filters. It is shown that the DCR is particularly favorable for modeling the cascade of a Volterra filter followed by a linear filter, as required for applications such as nonlinear acoustic echo cancellation. To obtain a corresponding adaptive structure of the proposed approach, we introduce a generalization of a known DFT-domain adaptive algorithm for linear systems to Volterra filters. Fabian Kuech, Walter Kellermann |
ICASSP (2) | 2 |
| 2004 | Introduction to the Special Issue on Multichannel Signal Processing for Audio and Acoustics ApplicationsabstractHE IEEE Signal Processing Society has its roots in an area where acoustics, speech, and signal processing converge, as was reflected in the former name of the society when it was founded in 1974. The interface between acoustics, speech, and signal processing is still an area of great interest to the society, with many fundamental problems still unsolved. Research is driven by applications where acoustic signals have to be captured, transmitted, and/or reproduced in an acoustic environment that includes echoes, noise, and reverberation Considering human/machine interfaces as a major area of applications, it is obvious that signal processing becomes more challenging as the distance between humans and the machines increases, as the signal bandwidth increases, and as the acoustic environment becomes more complex and hostile. Increasingly sophisticated algorithms have been developed since the mid-1970s and along with the availability of greatly increased and affordable computational power, multichannel signal processing algorithms naturally evolved for exploiting the spatial dimension of acoustic signals. The importance and popularity of this field was well reflected by the large number of submissions to this special issue. The volume of high-quality papers could not be fitted into the page budget allotted to us. Thus, we regrettably had to decide to publish some of them in a second section of this special issue as part of a regular issue of the TRANSACTIONS in early 2005. For sound reproduction, where we want to provide a pair of desired signals at the listeners’ ear drums, seamless human/machine interfaces based on multichannel techniques have been implemented since the invention of stereo systems. However, providing the true spatial sound experience in large listening spaces became possible only with new multichannel signal processing techniques, such as wavefield synthesis. Still, major challenges remain, especially phase-true equalization of listening room acoustics and the cancellation of local noise sources and interferers. On the other hand, acquisition of audio and speech signals has been a research topic since the invention of the microphone and still today presents major challenges for the signal processing community. Structurally the simplest problem, the acoustic feedback from loudspeakers into microphones is addressed by acoustic echo cancellation: From the single-channel case which has been investigated since the 1970s, research has moved on to stereo and multichannel reproduction, recently culminating in a new wave-domain adaptive filtering concept which has been presented for the first time at ICASSP 2004. For removing unwanted interference and noise from desired signals, multichannel techniques utilize spatial diversity to discriminate between desired and undesired components, either by exploiting different spatial coherence properties or by beamforming, which directs a beam of increased sensitivity towards the desired source. For traditional beamforming, source localization is necessary if the location of the source is not known a priori. Reflecting current research emphasis in these areas, the paper by Cohen addresses filtering of the Walter Kellermann, Man Mohan Sondhi, Diemer de Vries |
IEEE Trans. Speech Audio Process. | 1 |
| 2003 | An extended multidelay filter: fast low-delay algorithms for very high-order adaptive systemsabstractWe propose a novel class of efficient adaptive algorithms in the frequency domain that is tailored to very long adaptive filters and highly autocorrelated input signals as they arise, e.g., in high-quality full-duplex audio applications. The approach exhibits good tracking capabilities of the signal statistics and very low delay. Moreover, it is shown that the low order of computational complexity of the conventional frequency-domain adaptive algorithms can be maintained thanks to efficient realizations. The algorithm allows a tradeoff between the well-known multidelay filter (MDF) and the recursive least-squares (RLS) algorithm. It is also well suited for an efficient generalization to the multichannel case. Herbert Buchner, Walter Kellermann, Jacob Benesty |
ICASSP (5) | 2 |
| 2003 | Full-duplex multichannel communication: real-time implementations in a general frameworkabstractIn this combustion, we embed full-duplex multichannel communication interfaces for tele-presence systems into a general framework. On the reproduction side, we consider a wide range of multichannel acoustic rendering techniques including traditional stereophony, '5.1' systems, and wave field synthesis using loud speaker arrays for sound immersion. On the recording side, microphone arrays are discussed for capturing clean desired signals with spatial information. Based on this general framework, real-time implementations of such full-duplex multichannel communication systems are then described. We combine wave field synthesis with multichannel acoustic echo cancellation and adaptive beamforming and discuss a real-time implementation on standard desktop and laptop PCs. Wolfgang Herbordt, Herbert Buchner, Walter Kellermann, Rudolf Rabenstein, Sascha Spors, Heinz Teutsch |
ICME | 3 |
| 2002 | Improved Kalman gain computation for multichannel frequency-domain adaptive filtering and application to acoustic echo cancellationabstractFor multichannel adaptive filtering applications like acoustic echo cancellation, sophisticated, but efficient algorithms are necessary that take both auto- and cross-correlations of the excitation signals into account. Based on a recently published frequency-domain framework, we examine in this contribution the computation of the so-called frequency-domain Kalman gain, which is crucial in the multichannel case to meet the above requirements for correlated input signals. We present an extended scheme with increased computational efficiency for a higher number of input channels. Robust convergence behaviour during perturbation of the desired signal is ensured by a new dynamical regularization. Application of the new algorithm to multichannel acoustic echo cancellation with real-world signals achieves improved performance. Herbert Buchner, Walter Kellermann |
ICASSP | 2 |
| 2002 | Analysis of blocking matrices for generalized sidelobe cancellers for non-stationary broadband signalsabstractThis contribution relates the robust generalized sidelobe canceller (RGSC) with an adaptive blocking matrix (ABM) after Hoshuyama et al. for non-stationary broadband signals with linearly constrained minimum variance (LCMV) beamformers. While alternative approaches only exploit spatial wave field characteristics, it is shown that the ABM introduces dependency on the desired signal power spectral density (PSD) into the array look-direction constraints. The ABM yields better interference suppression than for stationary frequency independent look-direction constraints, since the ABM does not suppress frequencies without desired signal components. Wolfgang Herbordt, Walter Kellermann |
ICASSP | 2 |
| 2002 | Nonlinear line echo cancellation using a simplified second order Volterra filterabstractThis paper deals with nonlinearities in line echo cancellation. If a high attenuation of the echo signal is required, the nonlinear behaviour of the actual echo path has to be taken into account. As the echo path inherently has memory, the nonlinear approaches have to be capable to model nonlinear systems with memory. Adaptive Volterra filters represent a common method to deal with such systems. Unfortunately, they suffer from a high computational complexity. Therefore, we propose an adaptive structure representing a simplified realization of a special second order Volterra filter (SVF) and provide an efficient step-size control for speech input. For the given application this structure achieves the same performance as second order Volterra filters at remarkably reduced computational cost. Fabian Kuech, Walter Kellermann |
ICASSP | 2 |
| 2002 | Full-duplex communication systems using loudspeaker arrays and microphone arraysabstractFor high-quality multimedia communication systems, such as teleconferencing, or tele-teaching (especially of music), multichannel sound reproduction is highly desirable. While current approaches still rely on a restrained listening area, the sweet spot, a volume solution for a large listening space is offered by the wave field synthesis (WFS) method, where arrays of loudspeakers generate a prespecified sound field. On the recording side of the two-way systems, the use of microphone arrays is an effective approach to cope with undesired signal components in the receiving room. However, before full-duplex communication can be deployed, efficient approaches to the acoustic echo cancellation (AEC) problem in this challenging scenario have to be found. We investigate different options for system integration, after a brief discussion of the current state of the art. We then present a first real-time solution on a regular PC platform, based on an efficient AEC for MIMO (multi-input and multi-output) systems in the frequency-domain. Herbert Buchner, Sascha Spors, Walter Kellermann, Rudolf Rabenstein |
ICME (1) | 3 |
| 2002 | A real-time acoustic human-machine front-end for multimedia applications integrating robust adaptive beamforming and stereophonic acoustic echo cancellation
Wolfgang Herbordt, J. Ying, Herbert Buchner, Walter Kellermann |
INTERSPEECH | 4 |
| 2001 | Limits for generalized sidelobe cancellers with embedded acoustic echo cancellationabstractThis paper analyzes positive synergies and theoretical limits of the combination of acoustic echo cancellation (AEC) and generalized sidelobe cancellers (GSC) for the removal of echoes of loudspeaker signals and local interferers. While the proposed system only requires one AEC for an arbitrary number of microphones, the array gain is limited by the number of sensor channels when all interferers arrive from different directions-of-arrival (DOAs). The paper also shows that the degrees of freedom of the adaptive sidelobe cancelling path are not sufficient when local interferers and acoustic echoes have common DOAs. Wolfgang Herbordt, Walter Kellermann |
ICASSP | 2 |
| 2001 | Computationally efficient frequency-domain combination of acoustic echo cancellation and robust adaptive beamformingabstractFor hands-free acoustical human/machine interfaces, e. g. for automatic speech recognition or teleconferencing systems, microphone arrays using robust Generalized Sidelobe Cancellers (GSCs) in conjunction with acoustic echo cancellation (AEC) can be efficiently applied for optimum communication. This contribution devises a new structure for combining AEC and GSC. It reduces the computational complexity by more than a factor of ten relative to a time-domain arrangement, increases convergence speed, and preserves positive synergies. 1. Wolfgang Herbordt, Herbert Buchner, Walter Kellermann |
INTERSPEECH | 3 |
| 2001 | An acoustic human-machine interface with multi-channel sound reproductionabstractFor hands-free man-machine audio interfaces with multichannel sound reproduction and automatic speech recognition (ASR), sometimes both an acoustic echo canceller (AEC) and a beamforming (BF) microphone array are necessary for sufficient recognition rates. In the context of multimedia systems, multi-channel sound reproduction (e.g. stereo or 5.1 channel-surround systems) typically requires multi-channel AEC (M-C AEC). With M-C AEC being known to be a very demanding problem in signal processing, no practically relevant simulation results have yet been presented for more than two channels. We examine a previously proposed frequency-domain adaptive filtering scheme with some extensions for the case of more than two loudspeaker channels. The results using this approach show a remarkable performance at a relatively moderate computational complexity. Moreover, we show how to efficiently combine this M-C AEC structure with microphone arrays. Herbert Buchner, Walter Kellermann |
MMSP | 2 |
| 2001 | Efficient frequency-domain realization of robust generalized, sidelobe cancellersabstractThis paper deals with hands-free acoustical front-ends for human/machine interaction in competing talker situations using efficient frequency-domain implementations of a robust generalized sidelobe canceller (GSC). A new scheme is devised which leads to a factor of three of computational savings. The GSC tracking behavior is not impaired as the block lengths are kept sufficiently small. Wolfgang Herbordt, Walter Kellermann |
MMSP | 2 |
| 2000 | Nonlinear acoustic echo cancellation with fast converging memoryless preprocessorabstractLow-cost audio components in hands-free telephone applications call for nonlinear adaptive echo cancellation (AEC). It has been demonstrated that a cascade of a polynomial and an FIR filter can cancel such nonlinear echoes (Stenger and Rabenstein 1998). Another technique employs a hard-clipping curve with LMS adapted saturation parameter (Nollet and Jones 1997). For both cascaded systems we derive an LMS-type adaptation using a common framework, and propose stepsize normalizations for both approaches. To achieve sufficiently fast convergence for practical use, an RLS-type adaptation for the polynomial is derived and experimentally verified. Both techniques are compared using real hardware and speech signals, and show robust convergence behaviour and an echo reduction gain of up to 10 dB compared to a linear AEC. Alexander Stenger, Walter Kellermann |
ICASSP | 2 |
| 2000 | Special section on current topics in adaptive filtering for hands-free acoustic communication and beyond
Walter Kellermann |
Signal Process. | 1 |
| 2000 | Adaptation of a memoryless preprocessor for nonlinear acoustic echo cancelling
Alexander Stenger, Walter Kellermann |
Signal Process. | 2 |
| 1997 | Strategies for combining acoustic echo cancellation and adaptive beamforming microphone arraysabstractNew concepts for efficient combination of acoustic echo cancellation (AEC) and adaptive beamforming microphone arrays (ABMA) are presented. By decomposing common beamforming methods into a time-invariant part, which the AEC can integrate, and a separate time-variant part, the number of echo cancellers is minimized without rendering the system identification problem more difficult. Methods for controlling the interaction of ABMA and AEC are outlined and implementations for typical microphone array applications are discussed. Walter Kellermann |
ICASSP | 1 |
| 1996 | Editorial
André Gilloire, Eberhard Hänsler, Walter Kellermann, J. Svean |
Speech Commun. | 3 |
| 1991 | A self-steering digital microphone arrayabstractA self-steering microphone array for teleconferencing in which the digitally implemented steering algorithm consists of two parts is presented. The first part, the beamforming, is based on known concepts. The second part, a novel voting algorithm, integrates elements of pattern classification and exploits temporal characteristics of speech signals. It accounts for perceptual criteria and the acoustic environment. A real-time implementation is outlined, and results are discussed. The results confirm that the proposed concept deals successfully with teleconferencing environments and that it yields substantially better performance than earlier concepts based on analog hardware.> Walter Kellermann |
ICASSP | 1 |
| 1988 | Analysis and design of multirate systems for cancellation of acoustical echoesabstractAn analysis of the multirate concept for the echo cancellation problem is presented that results in the consideration of two classes of solutions. For the first class the solution is derived explicitly, whereas for the second class a general approach and a subclass are discussed. Regarding the application to the cancelling of acoustical echoes the design problem for the frequency-subband concept-which is shown to be a solution of the first class-is examined in more detail and results are presented verifying the efficiency of this method.> Walter Kellermann |
ICASSP | 1 |