EDBT 2026 Demo / reviewers in the wild / expert
Andreas Brendel
dblp:172/1185
· DBLP profile ↗
18ranked-venue papers
5as first author
13since 2021 · last 2025
0000-0002-6051-6346ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GAN-Based Speech Enhancement for Low SNR Using Latent Feature ConditioningabstractEnhancing speech quality under adverse SNR conditions remains a significant challenge for discriminative deep neural network (DNN)-based approaches. In this work, we propose DisCoGAN, which is a time-frequency-domain generative adversarial network (GAN) conditioned by the latent features of a discriminative model pre-trained for speech enhancement in low SNR scenarios. Our proposed method achieves superior performance compared to state-of-the-art discriminative methods and also surpasses end-to-end (E2E) trained GAN models. We also investigate the impact of various configurations for conditioning the proposed GAN model with the discriminative model and assess their influence on enhancing speech quality. Shrishti Saha Shetu, Emanuël A. P. Habets, Andreas Brendel |
ICASSP | 3 |
| 2024 | On Improving Error Resilience of Neural End-to-End Speech Codersabstract1755 Kishan Gupta, Nicola Pia, Srikanth Korse, Andreas Brendel, Guillaume Fuchs, Markus Multrus |
INTERSPEECH | 4 |
| 2024 | End-to-End Deep Learning-Based Adaptation Control for Linear Acoustic Echo CancellationabstractThe attenuation of acoustic loudspeaker echoes remains to be one of the open challenges to achieve pleasant full-duplex hands free speech communication. In many modern signal enhancement interfaces, this problem is addressed by a linear acoustic echo canceler which subtracts a loudspeaker echo estimate from the recorded microphone signal. To obtain precise echo estimates, the parameters of the echo canceler, i.e., the filter coefficients, need to be estimated quickly and precisely from the observed loudspeaker and microphone signals. For this a sophisticated adaptation control is required to deal with high-power double-talk and rapidly track time-varying acoustic environments which are often faced with portable devices. In this paper, we address this problem by end-to-end deep learning. In particular, we suggest to infer the step-size for a least mean squares frequency-domain adaptive filter update by a Deep Neural Network (DNN). Two different step-size inference approaches are investigated. On the one hand broadband approaches, which use a single DNN to jointly infer step-sizes for all frequency bands, and on the other hand narrowband methods, which exploit individual DNNs per frequency band. The discussion of benefits and disadvantages of both approaches leads to a novel hybrid approach which shows improved echo cancellation while requiring only small DNN architectures. Furthermore, we investigate the effect of different loss functions, signal feature vectors, and DNN output layer architectures on the echo cancellation performance from which we obtain valuable insights into the general design and functionality of DNN-based adaptation control algorithms. Thomas Haubner, Andreas Brendel, Walter Kellermann |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2024 | Erratum to "End-to-End Deep Learning-Based Adaptation Control for Linear Acoustic Echo Cancellation"abstractPresents corrections to the article “End-to-End Deep Learning-Based Adaptation Control for Linear Acoustic Echo Cancellation”. Thomas Haubner, Andreas Brendel, Walter Kellermann |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2024 | On Semi-Blind Source Separation-Based Approaches to Nonlinear Echo Cancellation Based on Bilinear Alternating OptimizationabstractAcoustic echo cancellation (AEC) is a crucial task in full duplex communications. As conventional linear filtering approaches are ineffective to deal with double-talk, various semi-blind source separation (SBSS)-based AEC algorithms are deceived, most of which are formulated and implemented in the frequency domain based on the multiplicative transfer function (MTF) model for computational efficiency. To avoid large latency and in order to deal with loudspeaker nonlinearities, the convolutive transfer function (CTF) model and odd power series expansion are leveraged, which are employed by numerous SBSS-based nonlinear AEC (SBSS-NAEC) algorithms. Conventional SBSS-NAEC methods estimate the series expansion coefficients and the CTF filter simultaneously making the number of free parameters to estimate large. Hence, the corresponding algorithms are computationally expensive and are difficult to optimize. In this work, we propose to decouple the series expansion coefficients and the CTF filters into a bilinear form and present a bilinear alternating optimization framework for estimating the model parameters. An alternating iterative projection (AIP) algorithm and an alternating element-wise iterative source steering (AEISS) algorithm are proposed. As the bilinear representation consists of less parameters compared to the conventional methods, the proposed algorithms not only improve the AEC performance but also reduce the computational complexity, which is validated by comprehensive simulations and experiments. Xianrui Wang, Yichen Yang 0010, Andreas Brendel, Tetsuya Ueda, Shoji Makino, Jacob Benesty, Walter Kellermann, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Spatially Informed Independent vector analysis for Source Extraction based on the convolutive Transfer Function ModelabstractSpatial information can help improve source separation performance. Numerous spatially informed source extraction methods based on the independent vector analysis (IVA) have been developed, which can achieve reasonably good performance in non- or weakly reverberant environments. However, the performance of those methods degrades quickly as the reverberation increases. The underlying reason is that those methods are derived based on the multiplicative transfer function model with a rank-1 assumption, which does not hold true if reverberation is strong. To circumvent this issue, this paper proposes to use the convolutive transfer function (CTF) model to improve the source extraction performance and develop a spatially informed IVA algorithm. Simulations demonstrate the efficacy of the developed method even in highly reverberant environments. Xianrui Wang, Andreas Brendel, Gongping Huang, Yichen Yang 0010, Walter Kellermann, Jingdong Chen |
ICASSP | 2 |
| 2022 | Manifold Learning-Supported Estimation of Relative Transfer Functions For Spatial FilteringabstractMany spatial filtering algorithms used for voice capture in, e.g., teleconferencing applications, can benefit from or even rely on knowledge of Relative Transfer Functions (RTFs). Accordingly, many RTF estimators have been proposed which, however, suffer from performance degradation under acoustically adverse conditions or need prior knowledge on the properties of the interfering sources. While state-of-the-art RTF estimators ignore prior knowledge about the acoustic enclosure, audio signal processing algorithms for teleconferencing equipment are often operating in the same or at least a similar acoustic enclosure, e.g., a car or an office, such that training data can be collected. In this contribution, we use such data to train Variational Autoencoders (VAEs) in an unsupervised manner and apply the trained VAEs to enhance imprecise RTF estimates. Furthermore, a hybrid between classic RTF estimation and the trained VAE is investigated. Comprehensive experiments with real-world data confirm the efficacy for the proposed method. Andreas Brendel, Johannes Zeitler, Walter Kellermann |
ICASSP | 1 |
| 2022 | End-To-End Deep Learning-Based Adaptation Control for Frequency-Domain Adaptive System IdentificationabstractWe present a novel end-to-end deep learning-based adaptation control algorithm for frequency-domain adaptive system identification. The proposed method exploits a deep neural network to map observed signal features to corresponding step-sizes which control the filter adaptation. The parameters of the network are optimized in an end-to-end fashion by minimizing the average normalized system distance of the adaptive filter. This avoids the need of explicit signal power spectral density estimation as required for model-based adaptation control and further auxiliary mechanisms to deal with model inaccuracies. The proposed algorithm achieves fast convergence and robust steady-state performance for scenarios characterized by high-level, non-white and non-stationary additive noise signals, abrupt environment changes and additional model inaccuracies. Thomas Haubner, Andreas Brendel, Walter Kellermann |
ICASSP | 2 |
| 2021 | Network-Aware Optimal Microphone Channel Selection in Wireless Acoustic Sensor NetworksabstractTo address the vital problem of selecting the most useful microphones in wireless acoustic sensor networks, this paper proposes a novel, general-purpose approach that accounts for both acoustic and network aspects and remains application-agnostic for broad applicability. The inter-channel correlation of single-channel signal features, together with tools from spectral graph theory, is used to assess the usefulness from an acoustic perspective. By only transmitting the features that characterize signal frames as opposed to the full signal waveform, the unique constraints of wireless sensor networks are accommodated. The source-to-sink transmission delay, resulting from embedding a distributed signal processing application into a wireless network, captures the usefulness from a network perspective. The experiments demonstrate the efficacy of the proposed method for an exemplary multichannel signal processing application. Michael Günther 0003, Haitham Afifi, Andreas Brendel, Holger Karl, Walter Kellermann |
ICASSP | 3 |
| 2021 | Accelerating Auxiliary Function-Based Independent Vector AnalysisabstractIndependent Vector Analysis (IVA) is an effective approach for Blind Source Separation (BSS) of convolutive mixtures of audio signals. As a practical realization of an IVA-based BSS algorithm, the so-called AuxIVA update rules based on the Majorize-Minimize (MM) principle have been proposed which allow for fast and computationally efficient optimization of the IVA cost function. For many real-time applications, however, update rules for IVA exhibiting even faster convergence are highly desirable. To this end, we investigate techniques which accelerate the convergence of the AuxIVA update rules without extra computational cost. The efficacy of the proposed methods is verified in experiments representing real-world acoustic scenarios. Andreas Brendel, Walter Kellermann |
ICASSP | 1 |
| 2021 | Noise-Robust Adaptation Control for Supervised Acoustic System Identification Exploiting a Noise DictionaryabstractWe present a noise-robust adaptation control strategy for block-online supervised acoustic system identification by exploiting a noise dictionary. The proposed algorithm takes advantage of the pronounced spectral structure which characterizes many types of interfering noise signals. We model the noisy observations by a linear Gaussian Discrete Fourier Transform-domain state space model whose parameters are estimated by an online generalized Expectation-Maximization algorithm. Unlike all other state-of-the-art approaches we suggest to model the covariance matrix of the observation probability density function by a dictionary model. We propose to learn the noise dictionary from training data, which can be gathered either offline or online whenever the system is not excited, while we infer the activations continuously. The proposed algorithm represents a novel machine-learning-based approach to noise-robust adaptation control which allows for faster convergence in applications characterized by high-level and non-stationary interfering noise signals and abrupt system changes. Thomas Haubner, Andreas Brendel, Mohamed Elminshawi, Walter Kellermann |
ICASSP | 2 |
| 2021 | Effective Rank-Based Estimation of the Coherent-to-Diffuse Power RatioabstractMany algorithms for speech dereverberation and noise reduction rely on an estimate of the coherent-to-diffuse power ratio (CDR). Such systems typically operate in very diverse acoustic conditions, and CDR estimators relying on very weak model assumptions about the acoustic sound field of the desired speech and interfering noise are hence desirable. A CDR estimator whose design is based on this premise is devised in this contribution. The proposed non-iterative CDR estimator exploits the effective rank of the spatial covariance matrix of the recorded input signals by assuming it to be lower for a coherent sound field than for a diffuse sound field. In addition to this weak assumption, related methods usually require information about, e.g., the array geometry, direction-of-arrival (DOA) of the desired source or rely on a coherence model for the desired signal or background noise, which is not required for the proposed method. Despite the use of little a priori information about the acoustic sound field, the new estimator achieves a significantly higher estimation accuracy for the CDR in comparison to related state-of-the-art approaches which use explicit coherence models. Heinrich W. Löllmann, Andreas Brendel, Walter Kellermann |
ICASSP | 2 |
| 2021 | Misalignment Recognition in Acoustic Sensor Networks Using a Semi-Supervised Source Estimation Method and Markov Random FieldsabstractIn this paper, we consider the problem of acoustic source localization by acoustic sensor networks (ASNs) using a promising, learning-based technique that adapts to the acoustic environment. In particular, we look at the scenario when a node in the ASN is displaced from its position during training. As the mismatch between the ASN used for learning the localization model and the one after a node displacement leads to erroneous position estimates, a displacement has to be detected and the displaced nodes need to be identified. We propose a method that considers the disparity in position estimates made by leave-one-node-out (LONO) sub-networks and uses a Markov random field (MRF) framework to infer the probability of each LONO position estimate being aligned, misaligned or unreliable while accounting for the noise inherent to the estimator. This probabilistic approach is advantageous over naïve detection methods, as it outputs a normalized value that encapsulates conditional information provided by each LONO sub-network on whether the reading is in misalignment with the overall network. Experimental results confirm that the performance of the proposed method is consistent in identifying compromised nodes in various acoustic conditions. Gabriel F. Miller, Andreas Brendel, Walter Kellermann, Sharon Gannot |
ICASSP | 2 |
| 2020 | Spatially Guided Independent Vector AnalysisabstractWe present a Maximum A Posteriori (MAP) derivation of the Independent Vector Analysis (IVA) algorithm for blind source separation incorporating an additional spatial prior over the demixing matrices. In this way, the outer permutation ambiguity of IVA is avoided and the algorithm can be guided towards a desired solution in adverse acoustic conditions. The resulting MAP optimization problem is solved by deriving majorize-minimize update rules to achieve convergence speed comparable to the well-known auxiliary function IVA algorithm, i.e., the convergence is not impaired by the additional constraint. The proposed algorithm exhibits superior performance at lower computational cost than a state-of-the-art spatially constrained IVA algorithm in a setup defined by real-world Room Impulse Responses (RIRs). Andreas Brendel, Thomas Haubner, Walter Kellermann |
ICASSP | 1 |
| 2020 | Generalized Coherence-Based Signal EnhancementabstractThis contribution presents a novel approach for coherence-based signal enhancement. An estimator for the coherent-to-diffuse ratio (CDR) is devised, which exploits the concept of generalized magnitude coherence and thus, unlike common state-of-the-art schemes, can simultaneously take advantage of more than two microphones. Moreover, the speech enhancement by CDR-based spectral weighting is not performed as a post-filtering step, but by enhancing the most appropriate microphone signal. This signal is implicitly determined as part of the CDR estimation such that the presented technique does not depend on an estimation of the direction-of-arrival (DOA) or similar side-information about the desired source.The application of the new approach to binaural hearings aids shows that it achieves a consistently better speech enhancement performance than comparable state-of-the-art approaches. Heinrich W. Löllmann, Andreas Brendel, Walter Kellermann |
ICASSP | 2 |
| 2019 | Localization of an Unknown Number of Speakers in Adverse Acoustic Conditions Using Reliability Information and DiarizationabstractThis paper investigates localization of an arbitrary number of simultaneously active speakers in an acoustic enclosure. We propose an algorithm capable of estimating the number of speakers, using reliability information to obtain robust estimation results in adverse acoustic scenarios and estimating individual probability distributions describing the position of each speaker using convex geometry tools. To this end, we start from an established algorithm for localization of acoustic sources based on the EM algorithm. There, the estimation of the number of sources as well as the handling of reverberation has not been addressed sufficiently. We show improvement in the localization of a higher number of sources and in the robustness in adverse conditions including interference from competing speakers, reverberation and noise. Andreas Brendel, Bracha Laufer-Goldshtein, Sharon Gannot, Ronen Talmon, Walter Kellermann |
ICASSP | 1 |
| 2019 | Neural Networks Sequential Training Using Variational Gaussian Particle FilterabstractIn this paper, we propose a sequential training algorithm for feed-forward neural networks based on particle filtering. The proposed algorithm uses variational learning to tailor a proposal density by minimizing the variational energy. This density is then incorporated into the Gaussian particle filter framework. The proposed algorithm and an extension to it using evolutionary resampling are compared to training a neural network using a random walk-based particle filter, an extended Kalman filter, the use of variational learning only, and the backpropagation algorithm, using a synthetic dataset generated by a time-varying random process and a real dataset, where the proposed approach resulted in a moderately lower training and testing errors and a better convergence behavior, rendering the algorithm attractive for uses such as neural networks pre-training. Mhd Modar Halimeh, Andreas Brendel, Walter Kellermann |
ICASSP | 2 |
| 2018 | Learning-Based Acoustic Source-Microphone Distance Estimation Using the Coherent-to-Diffuse Power RatioabstractWe propose a method for estimating the distance between a sound source and a pair of recording microphones. The developed algorithm operates in the short-time Fourier transform domain and is based on estimates of the coherent-to-diffuse power ratio, which provides a measure for the amount of reverberation in each time-frequency bin. For a direct use of these estimates, precise knowledge on the room characteristics is necessary, which is in practice usually not available and hard to obtain. Therefore, we use a learning-based method, which adapts to the characteristics of the room in a training phase and estimates the source-microphone distance in a testing phase. The experiments comprise various setups with simulated and real data. It is shown that the proposed method generalizes well for different microphone positions and works robustly for different source signals, directions of arrival, reverberation times, and signal observation intervals. This leads to a high estimation accuracy at a low computational complexity with a small amount of training data. Andreas Brendel, Walter Kellermann |
ICASSP | 1 |