Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Gerald Enzner

dblp:83/4272 · DBLP profile ↗
← Back
43ranked-venue papers
9as first author
7since 2021 · last 2023
0000-0002-3096-1342ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 32 · 7 first-author · 4 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 3 since 2021Computer networks · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
8 papers
Audio and music processing · 100%
Computer networks
2 papers
Physical-layer communications · 100%

Topics — the 23 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing
speech enhancement
1.442023
Binaural-Projection Multichannel Wiener Filter for Cue-Preserving Binaural Speech Enhancement · IEEE ACM Trans. Audio Speech Lang. Process. 2023
Binaural Speaker Localization Integrated Into an Adaptive Beamformer for Hearing Aids · IEEE ACM Trans. Audio Speech Lang. Process. 2018
Bayesian MMSE Filtering of Noisy Speech by SNR Marginalization With Global PSD Priors · IEEE ACM Trans. Audio Speech Lang. Process. 2018
Audio and music processing › acoustic signal processing
acoustic system identification
0.712023
Hybrid-Frequency-Resolution Adaptive Kalman Filter for Online Identification of Long Acoustic Responses With Low Input-Output Latency · IEEE ACM Trans. Audio Speech Lang. Process. 2023
Audio and music processing › spatial audio
binaural cue preservation
0.712023
Binaural-Projection Multichannel Wiener Filter for Cue-Preserving Binaural Speech Enhancement · IEEE ACM Trans. Audio Speech Lang. Process. 2023
Audio and music processing › speech enhancement
binaural speech enhancement
0.712023
Binaural-Projection Multichannel Wiener Filter for Cue-Preserving Binaural Speech Enhancement · IEEE ACM Trans. Audio Speech Lang. Process. 2023
Audio and music processing › speech enhancement
multichannel wiener filter
0.712023
Binaural-Projection Multichannel Wiener Filter for Cue-Preserving Binaural Speech Enhancement · IEEE ACM Trans. Audio Speech Lang. Process. 2023
Audio and music processing › acoustic signal processing › acoustic sensing
acoustic sensor network
0.512021
Double-Cross-Correlation Processing for Blind Sampling-Rate and Time-Offset Estimation · IEEE ACM Trans. Audio Speech Lang. Process. 2021
Audio and music processing › acoustic signal processing
sampling rate offset estimation
0.512021
Double-Cross-Correlation Processing for Blind Sampling-Rate and Time-Offset Estimation · IEEE ACM Trans. Audio Speech Lang. Process. 2021
Audio and music processing
spatial audio
0.412020
Direct Spatial-Fourier Regression of HRIRs from Multi-Elevation Continuous-Azimuth Recordings · IEEE ACM Trans. Audio Speech Lang. Process. 2020
Audio and music processing
acoustic echo cancellation
0.322023
Hybrid-Frequency-Resolution Adaptive Kalman Filter for Online Identification of Long Acoustic Responses With Low Input-Output Latency · IEEE ACM Trans. Audio Speech Lang. Process. 2023
State-Space Frequency-Domain Adaptive Filtering for Nonlinear Acoustic Echo Cancellation · IEEE Trans. Speech Audio Process. 2012
Audio and music processing › speech enhancement
binaural beamforming
0.312018
Binaural Speaker Localization Integrated Into an Adaptive Beamformer for Hearing Aids · IEEE ACM Trans. Audio Speech Lang. Process. 2018
Audio and music processing
sound source localization
0.312018
Binaural Speaker Localization Integrated Into an Adaptive Beamformer for Hearing Aids · IEEE ACM Trans. Audio Speech Lang. Process. 2018
Physical-layer communications › signal processing for communications
adaptive filtering
0.312017
Statistical Convergence Analysis for Optimal Control of DFT-Domain Adaptive Echo Canceler · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Physical-layer communications › interference cancellation
echo cancellation
0.312017
Statistical Convergence Analysis for Optimal Control of DFT-Domain Adaptive Echo Canceler · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Physical-layer communications
signal processing for communications
0.312017
Statistical Convergence Analysis for Optimal Control of DFT-Domain Adaptive Echo Canceler · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Audio and music processing › speech enhancement
dereverberation
0.212014
Variational Bayesian Inference for Multichannel Dereverberation and Noise Reduction · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Audio and music processing › speech enhancement › dereverberation
multichannel dereverberation
0.212014
Variational Bayesian Inference for Multichannel Dereverberation and Noise Reduction · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Audio and music processing › speech enhancement
noise reduction
0.222018
Bayesian MMSE Filtering of Noisy Speech by SNR Marginalization With Global PSD Priors · IEEE ACM Trans. Audio Speech Lang. Process. 2018
Variational Bayesian Inference for Multichannel Dereverberation and Noise Reduction · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Audio and music processing › acoustic echo cancellation
nonlinear acoustic echo cancellation
0.112012
State-Space Frequency-Domain Adaptive Filtering for Nonlinear Acoustic Echo Cancellation · IEEE Trans. Speech Audio Process. 2012
Audio and music processing › spatial audio
virtual auditory space
0.112020
Direct Spatial-Fourier Regression of HRIRs from Multi-Elevation Continuous-Azimuth Recordings · IEEE ACM Trans. Audio Speech Lang. Process. 2020
Physical-layer communications › modulation › continuous phase modulation
coded CPM
0.012002
Coded continuous phase modulation with low-complexity noncoherent reception · IEEE Trans. Commun. 2002
Physical-layer communications › modulation
continuous phase modulation
0.012002
Coded continuous phase modulation with low-complexity noncoherent reception · IEEE Trans. Commun. 2002
Physical-layer communications › signal detection
noncoherent detection
0.012002
Coded continuous phase modulation with low-complexity noncoherent reception · IEEE Trans. Commun. 2002
Physical-layer communications › signal detection
sequence estimation
0.012002
Coded continuous phase modulation with low-complexity noncoherent reception · IEEE Trans. Commun. 2002

Methods — techniques the papers use, named apart from their topics

partitioned-block frequency-domain adaptive filter · 0.7multichannel wiener filter · 0.7hybrid-frequency-resolution kalman filter · 0.7convex MMSE estimation · 0.7generalized cross correlation with phase transform · 0.5double-cross-correlation · 0.5STFT · 0.5least-squares fitting · 0.4continuous-azimuth recording · 0.4bayesian marginalization · 0.3statistical convergence analysis · 0.3mean-square deviation minimization · 0.3frequency-domain kalman filter · 0.3reduced-state viterbi decoding · 0.0noncoherent sequence estimation · 0.0
YearPublicationVenuePosition
2023 Long-Term Synchronization of Wireless Acoustic Sensor Networks with Nonpersistent Acoustic Activity Using Coherence State
abstract
Sample-accurate synchronization of nodes is required to enable the full potential of acoustic sensor networks for cooperative and enhanced signal acquisition. While metrics of spatio-temporal sensor utility are key to successful aggregation of sensor nodes, for instance, to perform sound localization or beamforming, the same is true for waveform-based assessment and compensation of sampling-rate offset (SRO). This paper therefore proposes an acoustic coherence state (ACS) metric to support systems for SRO estimation to integrate estimations of various utility due to nonpersistent acoustic activity and geometrical diversity. Specifically, we consider systems with SRO estimation and compensation in open- and closed-loop structures and outline the architectures for embedding ACS-based sensor utility. It is demonstrated in both cases that the acoustic coherence metric is more appropriate in terms of end-to-end synchronization performance than voice or sound activity detectors.
Aleksej Chinaev, Niklas Knaepper, Gerald Enzner
ICASSP3
2023 Neural Network Models with Integrated Training and Adaptation For Nonlinear Acoustic System Identification
abstract
System identification is instrumental in various tasks of acoustics, including acoustic measurement and acoustic echo cancellation. Adaptive filtering is proven to be successful for the identification of linear parts with variable impulse responses. Neural network frameworks are recently considered for nonlinear modeling, but the desire for variable system identification on individual sequences may be incompatible with the general concept of training across batch data. This paper therefore pursues architectures mixed from trainable and non-trainable adaptive sections. By seamlessly integrating adaptive FIR filters into end-to-end neural network architectures, the desired flexibility in the network is achieved. In-network adaptivity is here implemented as a recurrent computational layer according to the normalized least mean-square (NLMS) algorithm. Altogether six architectures are explored in order to study what part of a system better relies on training or adaptation. The proposed implementations are thus verified on synthetic data with dedicated variations of speech or noise input, linear or nonlinear system components, and fixed or variant acoustic impulse responses.
Svantje Voit, Gerald Enzner
ICASSP2
2023 Hybrid-Frequency-Resolution Adaptive Kalman Filter for Online Identification of Long Acoustic Responses With Low Input-Output Latency
abstract
Online acoustic system identification is one of the most challenging tasks for adaptive filters. Along with the desired accuracy in applications such as acoustic echo cancellation, it bears requirements of accommodating high-order systems (i.e., long acoustic impulse responses) while maintaining low input-output latency. These simultaneous requirements by now have been frequently addressed by multi-delay filters (MDF) a.k.a. partitioned-block frequency-domain adaptive filters (PBFDAF) with sectioned filter representation. In this paper, on the contrary, we consider a monolithic representation of high-order acoustic systems and yet constrain to identification with low input-output latency. For block-frequency-domain processing this approach therefore entails a hybrid frequency resolution with respect to system input and output. Further considering a first-order state-space temporal evolution of the acoustic system, we can rely on the Kalman filter framework for derivation of a hybrid-frequency-resolution adaptive Kalman filter (HyKF) algorithm for acoustic system identification with inherent robustness to double-talk and acoustic noise. Given the broad landscape of existing adaptive filter algorithms, the HyKF here turns out to be an entity of its own. Experimentally, it is demonstrated that the HyKF solution considerably improves rate of convergence and system accuracy over former MDF and former frequency-domain adaptive Kalman filters (FDKF). This development is motivated by recently increased attention for employing adaptive filters in deep learning frameworks for acoustic echo control.
Gerald Enzner, Svantje Voit
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 Binaural-Projection Multichannel Wiener Filter for Cue-Preserving Binaural Speech Enhancement
abstract
Former research in binaural speech enhancement has demonstrated a demand of binaural cue preservation beyond the requirements of noise suppression and speech quality. The binaural state-of-the-art is frequently grouped into the class of spatio-temporal optimum filters with composite cost functions dedicated to a compromise of simultaneous requirements and the class of common-gain spectral filters with exact cue preservation by construction. In this paper, we pursue spatio-temporal filtering by convex MMSE estimation constrained to strict binaural cue preservation. To this end, we rely on a frequency-domain representation of well-known interaural-level (ILD) and interaural time differences (ITD) for setting up a complex-valued constraint. It is then demonstrated that the sought spatial filter effectively falls into the class of common-gain spectral filtering, where the gain consists of a new arrangement of two spectral weightings related to acoustic transfer function (ATF) and power-spectral density (PSD), respectively. Moreover, its equivalence to an unconstrained multiple-input/multiple-output multichannel Wiener filter (MIMO-MWF) with binaural projection onto original noisy spatial cues is shown, hence the naming of the proposed solution as a binaural-projection multichannel Wiener filter (BP-MWF). Experimental results in terms of ILD/ITD spectral histograms and distance metrics confirm that BP-MWF meets the desire of spatial cue preservation. Regarding noise suppression and speech quality, BP-MWF turns out to improve instrumental segSNR, PESQ and STOI metrics over binaural state-of-the-art, such as the partial-noise-estimation forms of MVDR and MWF, and is competitive with the unconstrained MIMO-MWF as an upper bound. The results are finally supported by a formal listening test including various SNR, source directions, and noise types.
Stefan Thaleiser, Gerald Enzner
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 Control Architecture of the Double-Cross-Correlation Processor for Sampling-Rate-Offset Estimation in Acoustic Sensor Networks
abstract
Distributed hardware of acoustic sensor networks bears inconsistency of local sampling frequencies, which is detrimental to signal processing. Fundamentally, sampling rate offset (SRO) nonlinearly relates the discrete-time signals acquired by different sensor nodes. As such, retrieval of SRO from the available signals requires non-linear estimation, like double-cross-correlation processing (DXCP), and frequently results in biased estimation. SRO compensation by asynchronous sampling rate conversion (ASRC) on the signals then leaves an unacceptable residual. As a remedy to this problem, multi-stage procedures have been devised to diminish the SRO residual with multiple iterations of SRO estimation and ASRC over the entire signal. This paper converts the mechanism of offline multi-stage processing into a continuous feedback-control loop comprising a controlled ASRC unit followed by an online implementation of DXCP-based SRO estimation. To support the design of an optimum internal model control unit for this closed-loop system, the paper deploys an analytical dynamical model of the proposed online DXCP. The resulting control architecture then merely applies a single treatment of each signal frame, while efficiently diminishing SRO bias with time. Evaluations with both speech and Gaussian input demonstrate that the high accuracy of multi-stage processing is maintained at the low complexity of single-stage (open-loop) processing.
Aleksej Chinaev, Sven Wienand, Gerald Enzner
ICASSP3
2021 Cue-Preserving MMSE Filter with Bayesian SNR Marginalization for Binaural Speech Enhancement
abstract
Binaural speech enhancement has often suffered from the trade-off between noise reduction and spatial cue preservation. The common-gain filtering of noisy speech under minimum mean-square error (MMSE) turned out as a viable approach, which resembles the format of Wiener-filtering spectral enhancement. Those techniques critically require the estimation of the local time-varying a-priori SNR. In single-channel approaches, it has been recently shown that local a-priori SNR can be marginalized in a Bayesian sense with an MMSE approach. In this paper, we translate the single-channel approach into a binaural Bayesian SNR marginalization, based on a binaural a-priori SNR definition and a related hyperprior. The overall MMSE solution then turns into a posterior expectation of an informed cue-preserving Wiener filter function, the computation of which is governed by binaural a-posteriori SNR and global SNR (i.e., the hyper-prior mean). The resulting MMSE solution is thus easy to implement and performance consistently stands at the top of our evaluation by segmental SNR, PESQ, and STOI computational metrics.
Stefan Thaleiser, Gerald Enzner
ICASSP2
2021 Double-Cross-Correlation Processing for Blind Sampling-Rate and Time-Offset Estimation
abstract
Coherent processing of signals captured by a wireless acoustic sensor network (WASN) requires an estimation of unknown parameters such as the sampling-rate and sampling-time offset (SRO and STO) of asynchronous sensor clocks. Although some sophisticated techniques for blind parameter estimation have become available in this young field of research, further development of new methods is required especially regarding environmental robustness, online estimation, and computational efficiency. As the main contribution of this work, we therefore introduce a novel approach for blind SRO estimation in the spirit of both a recently available time-domain double-cross-correlation processor (DXCP) and the well-known generalized cross-correlation with phase transform (GCC-PhaT). For the proposed approach, called DXCP-PhaT, we specifically introduce secondary cross-quantities to restore ergodicity of the otherwise drifting cross-correlation function or cross-spectrum of asynchronous input signals, which is an important property for estimation. Based upon this theoretical contribution of the paper, we then derive the DXCP-PhaT algorithm in the STFT domain, which brings along a number of improvements over state-of-the-art: advanced accuracy in terms of SRO assessment, environmental robustness to larger microphone distances as of real sensor networks, realtime-applicability in terms of online architecture and reduced complexity, and eventually an extension with online STO estimation (up to the inherent ambiguity of digital STO and acoustic TDOA) reliant on SRO compensation. Those claims are confirmed by comprehensive experiments, both offline and online, on simulated data and real recordings from a Rasberry-Pi-based WASN.
Aleksej Chinaev, Philipp Thuene, Gerald Enzner
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 A Computationally Light Algorithm for Bayesian Speech Enhancement with SNR Marginalization
abstract
While speech enhancement has critically required the estimation of local time-varying SNR, it was recently shown that SNR can be marginalized in a Bayesian sense from the minimum-mean-square-error (MMSE) solution. Precisely, the local SNR is introduced as a stochastic variable and Bayesian integration can be approximately realized under consideration of a hyperprior distribution. In our paper, the proposed approach then takes the multimodal nature of the involved posterior distribution into account for speech inference. Specifically, the extrema of the posterior distribution, which can easily be obtained via differentiation, are combined according to their widths, heights and abscissa. The corresponding solution is not closed form, however, it is found within few iterations. This approach delivers a spectral weighting of noisy speech that simultaneously maximizes instrumental criteria of speech quality, specifically the segmental SNR, STOI score and PESQ.
Stefan Thaleiser, Gerald Enzner
ICASSP2
2020 Direct Spatial-Fourier Regression of HRIRs from Multi-Elevation Continuous-Azimuth Recordings
abstract
Individual head-related impulse responses (HRIRs) have been recognized as a key to creating high-fidelity virtual auditory spaces. Thus, fast and comprehensive acquisition of individual HRIRs has been a subject of continued research. Traditional stop-and-go measurement at discrete angles is time consuming and additionally requires spatial interpolation that has been tackled by mapping discrete HRIR tables to a spatial Fourier format. Uniformly continuous-azimuth recording with a moving apparatus, on the other hand, reduces acquisition time, but is noise-limited due to a very short observation time per angle. In the interest of both fast acquisition and high accuracy, in this paper, we propose direct retrieval of a spatial Fourier format from continuous-azimuth recordings at multiple simultaneous discrete elevations. Specifically, we fit a generative continuous-azimuth model of the recorded signal, based on the spatial Fourier representation, to the continuous recordings of individuals by least-squares. In this approach, the model is meant to entirely capture the spatial variation of the HRIR in azimuth, while the duration of the recording then systematically controls the noise rejection. The proposed time-domain treatment is free of block artifacts, but is numerically demanding. We outline how to take the special structure of the involved covariance matrices into account. Experimental results with simulated data and real recordings demonstrate that the HRIR performance in terms of binaural cues and reproducibility benefits from the proposed algorithm. Our method is hence practical in terms of low measurement time and high performance, while benefiting from increased computational power of current computers.
Christoph Urbanietz, Gerald Enzner
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 A Double-cross-correlation Processor for Blind Sampling Rate Offset Estimation in Acoustic Sensor Networks
abstract
Signal synchronization in wireless acoustic sensor networks requires an accurate estimation of the sampling rate offset (SRO) inevitably present in signals acquired by sensors of ad-hoc networks. Although some sophisticated methods for blind SRO estimation have been recently proposed in this very young field of research, there is still a need for the development of new ideas and concepts especially regarding robust approaches with low computational complexity. We therefore propose a novel time-domain method based on the calculation of a double-cross-correlation function in this contribution. Experimental evaluation of the introduced approach in a challenging acoustic environment and comparison with a state-of-the-art frequency-domain method confirms the high accuracy and low computational load of the proposed technique.
Aleksej Chinaev, Philipp Thuene, Gerald Enzner
ICASSP3
2019 Spatial-fourier Retrieval of Head-related Impulse Responses from Fast Continuous-azimuth Recordings in the Time-domain
abstract
Fast and comprehensive acquisition of head-related impulse responses (HRIRs) continues in the interest of rich application scenarios. Various HRIR resolutions have been presented based on discrete stop-and-go measurement, or comprehensive measurement equipment, or continuous-azimuth acquisition with moving apparatus. Some researchers have further mapped the recorded HRIR table to spatial Fourier format. In the interest of both fast acquisition and high accuracy, in this paper, we propose direct retrieval of the spatial Fourier format from time-domain recordings of the measurement signal. Specifically, we revert to the moving apparatus for HRIR measurement and express a generative model of the recorded signal based on a hypothetical spatial Fourier basis. In this approach, the spatial-Fourier model is meant to entirely capture the spatial variation of the HRIR. By least-squares inference we can then fit that model to the entire recording. Due to our time-domain treatment the model and the processing are entirely clean from block artifacts, while the huge dimension of the problem requires attention in the implementation. It turns out that the proposed algorithm significantly outperforms the previous NLMS-based continuous-azimuth acquisition of HRIR with moving apparatus.
Christoph Urbanietz, Gerald Enzner
ICASSP2
2018 Binaural Rendering of Dynamic Head and Sound Source Orientation Using High-Resolution HRTF and Retarded Time
abstract
This paper is devoted to high-fidelity implementation of HRTF-based binaural rendering with fast head and source rotations in virtual acoustic reality. With an intuitive physical standpoint, we argue that head rotations should be rendered by a convolution model anchored in the sound receive-time. Conversely, the rendering of moving sound sources should be anchored in the sound emission-time, using retarded HRTF. Distant sound sources with bulk delay on the impulse response are handled efficiently by substituting the delay by additional retardation. The proposed rendering engine can utilize fast changing filter coefficients from a high resolution HRTF and satisfies some acoustic features as Doppler shift and dislocation inherently, which have been treated separately in conventional approaches or are neglected at all. To benchmark our rendering engine, we compare these acoustic features of the rendered signal with corresponding physical expectations.
Christoph Urbanietz, Gerald Enzner
ICASSP2
2018 Bayesian MMSE Filtering of Noisy Speech by SNR Marginalization With Global PSD Priors
abstract
MMSE filtering of speech with additive noise and latent speech power-spectral density (PSD) is addressed. This problem is strong in single-channel speech enhancement and restricts the utility of stationary Wiener filters or other statistical estimators based on PSDs. The issue typically manifests itself in residual noise after filtering, despite the availability of the noise PSD. Our paper therefore incorporates the latent speech PSD state via marginalization into the MMSE estimation framework of complex speech spectral amplitudes. The hence involved joint posterior distribution of the complex speech amplitude and speech PSD, conditioned on just the noisy observations, is then resolved in the Bayesian sense into a speech and a speech-PSD posterior. The latter is expressed via the local data likelihood and a hyper-prior of the local speech PSD or a-priori SNR-i.e., a global distribution across the entire speech signal. Marginalization, in this way, turns into expectation over a latent Wiener filter, such that explicit estimation of local a-priori SNR is eliminated. The local input data in the form of the a-posteriori SNR and the global SNR value as a descriptor of the overall speech-in-noise condition turns out sufficient to control our resulting MMSE spectral gain function, and, potentially, can be provided much easier than the latent and time-varying a-priori SNR. An improved balance of residual noise and speech quality in the enhancement of noisy speech is demonstrated by objective experimental evaluation.
Gerald Enzner, Philipp Thuene
IEEE ACM Trans. Audio Speech Lang. Process.1
2018 Binaural Speaker Localization Integrated Into an Adaptive Beamformer for Hearing Aids
abstract
In this paper, we present and compare novel algorithms to localize simultaneous speakers using four microphones distributed on a pair of binaural hearing aids. The framework consists of two groups of localization algorithms, namely, beamforming-based and statistical model based localization algorithms. We first generalize our previously proposed methods based on beamforming techniques to the binaural configuration with 2 × 2 microphones. Next, we contribute two statistical model based methods for binaural localization using the maximum likelihood approach that also takes head-related transfer functions and unknown noise conditions into account. The methods enable the localization of multiple source positions for all azimuth angles and do not require prior training of binaural cues. The proposed localization algorithms are integrated into a generalized side-lobe canceller (GSC) to extract the desired speaker in the presence of competing speakers and background noise and when the head of the listener turns. The GSC components are adapted with the frequency-wise target presence probability and the frame-wise broadband direction-of-arrival (DOA) estimates that track the turns of the listener's head. We evaluate the performance of the localization algorithms individually and also in the context of the adaptive binaural beamformer in various noisy and reverberant conditions. Finally, we introduce a new adaptive beamformer, which combines the GSC with multichannel speech presence probability estimation and achieves superior source separation performance in noisy environment.
Mehdi Zohourian, Gerald Enzner, Rainer Martin 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Realtime binaural speech enhancement demo on raspberry Pi
abstract
We demonstrate the feasibility of the realtime implementation of advanced binaural noise reduction algorithms in a single-chip computer called Raspberry Pi. The implementation of the considered algorithms is realized in Simulink, a graphical programming add-on to the integrated development environment Matlab. Using a complementary support pack- age for Simulink, the Raspberry Pi is connected/hosted. The implemented binaural noise reduction algorithm comprises two stages. First, the noise power spectral density (PSD) is estimated by one of the speech blocking-based noise PSD estimators previously proposed. The adaptively estimated noise PSD in each frame is then employed in a cue-preserving MMSE-based noise-reduction spectral gain function. The objective of this demonstration is to present the capabilities of the proposed algorithms with the application for hearing aids in a challenging noisy environment, i.e., congress babble noise in the show-and-tell area. The proposed solution suppresses the noise without having prior information on noise statistics, target speaker location and voice activity detection (VAD). Moreover, we would like to exhibit the powerful solution entirely executed with low-cost hardware.
Masoumeh Azarpour, Jan Siska, Gerald Enzner
ICASSP3
2017 Robust MMSE filtering for single-microphone speech enhancement
abstract
MMSE filtering of signals contaminated with additive noise is addressed with explicit uncertainty of the second-order target signal statistics. The unfortunate lack of stationarity of speech, and hence the phenomenon of musical noise in speech enhancement, is an ideal problem for the proposed approach. Specifically, we complement the established short-time power-spectral subtraction for speech power estimation with a prior of the momentary speech-power level. The MMSE estimator for Gaussian speech amplitudes is then derived under these circumstances. The potential for the enhancement of noisy speech is briefly demonstrated by SNR and PESQ analysis.
Gerald Enzner, Philipp Thuene
ICASSP1
2017 Frequency-Domain Adaptive Kalman Filter With Fast Recovery of Abrupt Echo-Path Changes
abstract
The frequency-domain Kalman filter (FDKF) was successfully applied in echo-cancelation systems due to its fast initial convergence and robustness to double-talk. However, the reconvergence ability of the FDKF was not comprehensively resolved and analyzed in the literature. This letter presents the steady-state solutions of the predicted system distance and corresponding step size of the FDKF, which is found to be related to the signal-to-noise ratio and the transition parameter. On this basis, it is pointed out that a frequently observed slow reconvergence of the FDKF, when the sudden echo-path change occurs, can be mainly attributed to the over-estimated noise power spectral density. Finally, we propose a shadow filter approach to resolve the algorithm's lack of reconvergence. Computer simulations of acoustic echo cancelation confirm the improved performance of the proposed method.
Feiran Yang 0001, Gerald Enzner, Jun Yang 0004
IEEE Signal Process. Lett.2
2017 Statistical Convergence Analysis for Optimal Control of DFT-Domain Adaptive Echo Canceler
abstract
The frequency-domain adaptive filter (FDAF) is widely used in echo cancellation systems due to its low complexity and fast convergence rate. However, the FDAF algorithm with a fixed step size exhibits a tradeoff among the convergence rate, steady-state misalignment, tracking ability, and robustness to near-end speech interferences. Several variable step-size FDAF algorithms were presented to address this problem. However, the state-of-the-art variable step-size FDAF algorithms did not handle this problem comprehensively. This paper presents a new robust variable step-size control approach to the FDAF algorithm. Based on a statistical analysis of the FDAF algorithm, an optimal step size for each frequency bin is derived by minimizing the mean-square deviation (MSD) between the true weight vector and estimated weight vector at each frame. Calculation of the step size requires the system distance and the observation noise power spectral density (PSD). The system distance is estimated using the deterministic recursive equations of MSD and the noise PSD is computed using the magnitude squared coherence function between the far-end signal and error signal. Moreover, a close link between the proposed FDAF and the frequency-domain Kalman filter is revealed. Specifically, the work presented here can be understood as a means to adaptively monitor and control the underlying acoustic state space of the Kalman filter, including means for fast readaptation of the adaptive filter after abrupt echo path changes. Simulation results demonstrate that the proposed algorithm can achieve fast convergence and low steady-state misalignment. Furthermore, the algorithm is robust to the double-talk interferences, but it does not require an explicit double-talk detector.
Feiran Yang 0001, Gerald Enzner, Jun Yang 0004
IEEE ACM Trans. Audio Speech Lang. Process.2
2016 Evaluation of estimated hammerstein models via normalized projection misalignment of linear and nonlinear subsystems
abstract
In linear system identification, the coexistence of parameter-misadjustment and output-error metrics has turned out very practical and their relation is well understood. In nonlinear system identification, however, such tools for performance evaluation are far less developed and each nonlinear type may need its own treatment. This paper focuses on the Hammerstein model as an instance of nonlinear systems. Irrespective of particular identification algorithms, we generalize the framework of parameter-and output-based performance metrics known from linear systems. An ambiguity in system parameters is resolved via the projection misalignment technique.
Gerald Enzner, Tim C. Kranemann, Philipp Thuene
ICASSP1
2015 Binaural speech enhancement with instantaneous coherence smoothing using the cepstral correlation coefficient
abstract
In this paper we propose a novel approach to cepstral smoothing for reducing musical noise fluctuations in binaural speech enhancement. Similar to other methods, our approach computes a preliminary spectral gain function using the magnitude-squared coherence function and applies an instantaneous weighting to the gain function in the cepstral domain. In this contribution, the weighting function is based on the binaural cepstral correlation coefficient (CCC). We introduce the CCC and briefly discuss its properties. Similar to the cepstrum, the CCC emphasizes the spectral envelope and fundamental frequency information of the target signal, however, in a representation normalized to a range of [-1,1] and less sensitive to spatially uncorrelated noise. Thus, it can be easily and effectively used as a weight in the cepstral domain. The utility of the CCC is confirmed via experiments with different noise types and several instrumental measures.
Rainer Martin 0001, Masoumeh Azarpour, Gerald Enzner
ICASSP3
2015 Multichannel Wiener filtering via multichannel decorrelation
abstract
Extracting a target source signal from multiple noisy observations is an essential task in many applications of signal processing such as digital communications or speech and audio processing. The multi-channel Wiener filter is able to solve this task in a minimum-mean-square-error (MMSE) optimal way by applying a spatial filter succeeded by a spectral postfilter. Its direct implementation, however, is difficult due to requiring the statistics of the unobservable source and noise signals. In this paper, we apply the signal-separation-based technique of multichannel decorrelation and reveal its relation to the Wiener post-filtering component. On this basis, we present a numerically robust and efficient adaptive algorithm to find an estimate of the MMSE-optimal postfilter based on the statistics of the observable signals alone. Experimental evaluation demonstrates the validity of the proposed approach and confirms the convergence of the adaptive algorithm to the MMSE-optimal postfilter solution.
Philipp Thuene, Gerald Enzner
ICASSP2
2014 Binaural noise PSD estimation for binaural speech enhancement
abstract
In this paper we propose a novel binaural algorithm to estimate the power spectral density (PSD) of the background noise at the left and right ear separately. Inspired by equalization-cancelation considered in binaural hearing, the target speech is canceled at both left and right ears by means of the FLMS (fast least-mean square) algorithm. Assuming the ideal equalization, the output error of the blocking filter is a biased estimation of noise PSD. The estimated noise PSD is further corrected by exploiting the estimated left and right interaural transfer functions and the coherence model of the noise field. In addition to noise power estimation assessment, the estimated noise PSD is integrated in a binaural speech enhancement framework in order to evaluate the overall noise reduction performance.
Masoumeh Azarpour, Gerald Enzner, Rainer Martin 0001
ICASSP2
2014 State-space architecture of the partitioned-block-based acoustic echo controller
abstract
Acoustic echo cancellation has traditionally employed basically all variants known from deterministic adaptive filter design, such as least mean-square (LMS), recursive least-squares (RLS), and frequency-domain adaptive filters (FDAF). More recently, a stochastic adaptive filter design based on the concept of acoustic state-space modeling of the echo path has been introduced to accommodate for an ever sought unification of adaptive filtering and adaptation control. The corresponding Kalman filter theory has been formulated for single-channel, multi-channel, and nonlinear echo cancellation problems. This paper closes an important gap by formulating the state-space model and the corresponding adaptive algorithm for the partitioned-block filtering structure which is especially relevant in practice. This structure allows for the use of significantly longer filter lengths in comparison to previous work, and for the flexible design and implementation of acoustic echo cancellers for widely differing acoustic conditions.
Fabian Kuech, Edwin Mabande, Gerald Enzner
ICASSP3
2014 Variational Bayesian Inference for Multichannel Dereverberation and Noise Reduction
abstract
Room reverberation and background noise severely degrade the quality of hands-free speech communication systems. In this work, we address the problem of combined speech dereverberation and noise reduction using a variational Bayesian (VB) inference approach. Our method relies on a multichannel state-space model for the acoustic channels that combines frame-based observation equations in the frequency domain with a first-order Markov model to describe the time-varying nature of the room impulse responses. By modeling the channels and the source signal as latent random variables, we formulate a lower bound on the log-likelihood function of the model parameters given the observed microphone signals and iteratively maximize it using an online expectation-maximization approach. Our derivation yields update equations to jointly estimate the channel and source posterior distributions and the remaining model parameters. An inspection of the resulting VB algorithm for blind equalization and channel identification (VB-BENCH) reveals that the presented framework includes previously proposed methods as special cases. Finally, we evaluate the performance of our approach in terms of speech quality, adaptation times, and speech recognition results to demonstrate its effectiveness for a wide range of reverberation and noise conditions.
Dominic Schmid, Gerald Enzner, Sarmad Malik, Dorothea Kolossa, Rainer Martin 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2013 Adaptive binaural noise reduction based on matched-filter equalization and post-filtering
abstract
In this paper a binaural noise reduction system based on the adaptive matched filter array (MFA) and post-filtering is presented. The binaural MFA filter is formed using the interaural impulse response between left and right microphone signal which is estimated by means of the NLMS algorithm. The residual noise reduction by three post-filter types is then compared and evaluated according to objective and subjective measures. Furthermore, the performance of the algorithm for preserving binaural cues is discussed.
Masoumeh Azarpour, Gerald Enzner, Rainer Martin 0001
ICASSP2
2013 Advanced system options for binaural rendering of Ambisonic format
abstract
Ambisonics uses a well-respected soundfield representation and thus may become a standard in soundfield transmission, storage, and reproduction. Then we will need signal processing options in order to decode and apply the Ambisonic format universally in various applications from loudspeaker to headphone reproduction. More specifically, the use of Ambisonics is envisioned also in mobile applications where headphone reproduction is clearly preferred. In this context, we present two different system options for the advanced implementation of binaural soundfield rendering. Both options utilize the concept of virtual loudspeakers and subsequent binaural rendering via HRTF. In order to achieve spatial realism, one option relies on adaptive soundfield rotation in the Ambisonic domain to mimic a corresponding head rotation of the listener in the given soundfield. The other system option relies on a new continuous-azimuth HRTF format for direct representation of head rotations within the virtual loudspeaker setup. For both system options, we will discuss the pros and cons regarding the implementation. We conclude with a comparison of the resulting localization accuracies in listening tests.
Gerald Enzner, Michael W. Weinert, Stefan Abeling, Jan-Mark Batke, Peter Jax
ICASSP1
2013 Sound source localization with binaural hearing aids using adaptive blind channel identification
abstract
This paper applies blind channel identification (BCI) to estimate direction of arrival (DOA) of sound sources with a pair of binaural hearing aids. It compares the Adaptive Eigenvalue Decomposition Algorithm (AEDA) with the Adaptive Principal Component Algorithm (APCA) for blindly estimating the impulse responses from target to hearing aids and these impulse responses are used to estimate DOA. This paper investigates how both the time difference and the level difference of the impulse responses can be used to estimate DOA, and the performance of both algorithms is evaluated for scenarios with different reverberation times, different SNR, and different source positions. The paper also evaluates the tracking behavior in a dual-talker scenario. The results show that AEDA's DOA performance suffers in the presence of noise and reverberation. APCA is fairly insensitive to noise, but it can only handle moderate levels of reverberation.
Ivo Merks, Gerald Enzner, Tao Zhang 0024
ICASSP2
2013 Improved online identification of acoustic MISO systems based on separated input signal components
abstract
Creating an immersive listening experience and providing the audience with improved spatial realism is the goal of many adaptive audio reproduction techniques such as room equalization or crosstalk cancellation. The majority of these approaches currently relies on acoustic impulse responses (AIRs) that have been measured prior to the actual audio reproduction. In order to maintain a high degree of adaptivity, however, the AIRs need to be estimated online during the reproduction process, which turns out to be a severely ill-conditioned problem due to the high inter-channel correlation of the loudspeaker signals. In this paper, we present a novel approach to MISO system identification with realistically correlated excitation. Based on the idea of separate treatment of correlated and uncorrelated signal components, we propose two extended filter structures for gradient-descent-based adaptive system identification and provide theoretical analysis and experimental validation of their effectiveness.
Philipp Thuene, Gerald Enzner
ICASSP2
2012 Perfect-sweep NLMS for time-variant acoustic system identification
abstract
Fast and robust acoustic system identification is still a research topic of interest, because of the typically time-variant nature of acoustic systems and the natural performance limitation of electroacoustic measurement equipment. In this paper, we propose NLMS-type adaptive identification with perfect-sweep excitation. The perfect-sweep is derived from the more general class of perfect sequences and, thus, it inherits periodicity and especially the desired decorrelation property known from perfect sequences. Moreover, the perfect-sweep shows the desirable characteristics of swept sine signals regarding the immunity against non-linear loudspeaker distortions. On this basis, we first demonstrate the fast tracking ability of the perfect-sweep NLMS algorithm via computer generated simulation of a time-variant acoustic system. Then, the robustness of the perfect-sweep NLMS algorithm against non-linear characteristics of real measurements in a time-invariant case is presented. By finally addressing the measurement of quasi-continuous head-related impulse responses, we face the combined challenge of time-variant and possibly non-linear distorted acoustic system identification in a real application scenario and we can demonstrate the superiority of the perfect-sweep NLMS algorithm.
Christiane Antweiler, Aulis Telle, Peter Vary, Gerald Enzner
ICASSP4
2012 Variational Bayesian inference for nonlinear acoustic echo cancellation using adaptive cascade modeling
abstract
In this contribution, we present a variational Bayesian framework for the acoustic echo cancellation problem in the presence of a memoryless loudspeaker nonlinearity. We pursue a cascade modeling strategy, where first-order Markov models are described over the acoustic echo path and the nonlinear expansion coefficients. An iterative algorithm is then derived that learns the posterior on the echo path and the nonlinear coefficients to fit the evidence distribution. We show that the formulated variational Bayesian state-space frequency-domain adaptive filter is efficiently implementable and performs joint learning of the echo path and the loudspeaker nonlinearity. The algorithm exploits the internal exchange of the reliability information, resulting in effective linear and nonlinear echo cancellation.
Sarmad Malik, Gerald Enzner
ICASSP2
2012 An expectation-maximization algorithm for multichannel adaptive speech dereverberation in the frequency-domain
abstract
This paper presents an online dereverberation algorithm that is derived within the maximum-likelihood expectation-maximization (ML-EM) framework. We formulate an overlap-save observation model for the multichannel blind problem in the DFT-domain. The modeling of acoustic channel impulse responses as random variables with a first-order Markov property facilitates the ensuing algorithm to cope with time-varying conditions. We then show that the ML-EM learning rules for the multichannel state-space model at hand take the form of a recursive posterior estimator for the channels, followed by an equalization stage for recovering the speech signal subject to an expectation with respect to the estimated channel posterior. Our derivation thus results in an iterative ML algorithm for blind equalization and channel identification (ML-BENCH) which comprises two distinct and coupled subsystems. The dereverberation performance of the proposed system is evaluated by considering spectrograms and instrumental quality measures.
Dominic Schmid, Sarmad Malik, Gerald Enzner
ICASSP3
2012 A State-Space Cross-Relation Approach to Adaptive Blind SIMO System Identification
abstract
In this work, we address blind single-input multiple-output (SIMO) system identification in conjunction with dynamical modeling of the underlying system. A multichannel cross-relation observation model in the DFT domain is employed to derive a blind adaptive algorithm that recursively learns the posterior distribution on the unknown SIMO system. The proposed algorithm inherently incorporates the time-varying nature of the channels and a representation of the observation noise. We show that the resulting cross-relation state-space frequency-domain adaptive filter (CR-SSFDAF), owing to its stable and diagonalized structure and near-optimal step-size control, can be efficiently operated in time-varying and noisy conditions.
Sarmad Malik, Dominic Schmid, Gerald Enzner
IEEE Signal Process. Lett.3
2012 State-Space Frequency-Domain Adaptive Filtering for Nonlinear Acoustic Echo Cancellation
abstract
In this paper, we address adaptive acoustic echo cancellation in the presence of an unknown memoryless nonlinearity preceding the echo path. We approach the problem by considering a basis-generic expansion of the memoryless nonlinearity. By absorbing the coefficients of the nonlinear expansion into the unknown echo path, the cascade observation model is transformed into an equivalent multichannel structure, which we further augment with a multichannel first-order Markov model. For the resulting multichannel state-space model, we then derive a recursive Bayesian estimator that takes the form of an adaptive Kalman algorithm in the discrete Fourier transform (DFT) domain. We show that such a recursive estimator can be realized via a stable and structurally efficient multichannel state-space frequency-domain adaptive filter. We demonstrate that our algorithm, which stems from a contained framework, provides effective nonlinear echo cancellation in the presence of continuous double-talk, varying degree of nonlinear distortion, and changes in the echo path.
Sarmad Malik, Gerald Enzner
IEEE Trans. Speech Audio Process.2
2011 Fourier expansion of hammerstein models for nonlinear acoustic system identification
abstract
We consider the task of acoustic system identification, where the input signal undergoes a memoryless nonlinear transformation before convolving with an unknown linear system. We focus on the possibility of modeling the nonlinearity with different basis functions, namely the established power series and the proposed Fourier expansion. In this work the unknown coefficients of generic basis functions are merged with the unknown linear system to obtain an equivalent multichannel structure. We use a multichannel DFT-domain algorithm for learning the underlying coefficients of both types of basis functions. We show that the Fourier modeling achieves faster convergence and better learning of the underlying nonlinearity than the polynomial basis.
Sarmad Malik, Gerald Enzner
ICASSP2
2011 Evaluation of adaptive blind SIMO identification in terms of a normalized filter-projection misalignment
abstract
The blind identification of single-input multiple-output (SIMO) systems suffers in the presence of near-common and exact common zeros between the channels, particularly in conjunction with observation noise. In general, we notice an ambiguity of the identification which cannot be resolved without further a priori information on the channel coefficients. In order to enable an adequate evaluation of blind SIMO identification in such cases, we develop the normalized filter-projection misalignment (NFPM), which represents a multichannel squared-error distance between true and estimated channels, while absorbing a common filter error due to a possible lack of identifiability. Using the NFPM measure, we demonstrate experimentally that the steady-state performance of the blind multichannel least mean-square (MCLMS) algorithm in the presence of missing channel diversity and noise is in line with the results obtained from supervised least mean-square (LMS) system identification.
Dominic Schmid, Gerald Enzner
ICASSP2
2011 Recursive Bayesian Control of Multichannel Acoustic Echo Cancellation
abstract
We present a novel recursive Bayesian method in the DFT-domain to address the multichannel acoustic echo cancellation problem. We model the echo paths between the loudspeakers and the near-end microphone as a multichannel random variable with a first-order Markov property. The incorporation of the near-end observation noise, in conjunction with the multichannel Markov model, leads to a multichannel state-space model. We derive a recursive Bayesian solution to the multichannel state-space model, which turns out to be well suited for input signals that are not only auto-correlated but also cross-correlated. We show that the resulting multichannel state-space frequency-domain adaptive filter (MCSSFDAF) can be efficiently implemented due to the submatrix-diagonality of the state-error covariance. The filter offers optimal tracking and robust adaptation in the presence of near-end noise and echo path variability.
Sarmad Malik, Gerald Enzner
IEEE Signal Process. Lett.2
2010 Online maximum-likelihood learning of time-varying dynamical models in block-frequency-domain
abstract
A linear dynamical model can be used to describe the evolution of an unknown system in noisy conditions. However, in most applications model parameters of a dynamical system are not known a priori, bringing into question the optimality of traditional state-only estimators. In this paper, we consider block-frequency-domain dynamical models and formulate an optimal framework for low-latency joint state and parameter estimation. We show that the resulting variational expectation-maximization algorithm in the block-frequency-domain offers a comprehensive and efficient solution for the joint estimation task.
Sarmad Malik, Gerald Enzner
ICASSP2
2008 Analysis and optimal control of LMS-type adaptive filtering for continuous-azimuth acquisition of head related impulse responses
abstract
Head related impulse responses (HRIRs) are the key to spatial realism in auditory virtual environments (AVEs). However, the measurement of discrete-azimuth HRIRs and their interpolation has been recognized as a tedious and delicate experimental procedure. We therefore suggest an adaptive filtering concept for continuous HRIR acquisition that completely avoids the traditional sampling and interpolation issue. Using an LMS-type adaptive algorithm, the HRIRs - at any azimuth - are extracted from a one-shot binaural recording. During data acquisition, the subject of interest is continuously rotated in the horizontal plane in order to capture the corresponding spatial information. In particular, the paper provides a profound theoretical and experimental analysis of the resulting HRIR inaccuracy in terms of the mean-square error. Furthermore, the optimal step- size parameter of the LMS-type adaptive algorithm is determined for which the minimum HRIR inaccuracy is attained.
Gerald Enzner
ICASSP1
2006 Frequency-domain adaptive Kalman filter for acoustic echo control in hands-free telephones
Gerald Enzner, Peter Vary
Signal Process.1
2003 A soft-partitioned frequency-domain adaptive filter for acoustic echo cancellation
abstract
We present a soft-partitioned frequency-domain adaptive filter (SPFDAF) and compare it in the context of other LMS type acoustic echo cancelers. The SPFDAF uses a nonrectangular window to partition the model taps (impulse response) of an echo canceler. Our soft window can be approximated very efficiently in the DFT domain and still attenuates aliasing components of the DFT by 40 dB. This technique results in a very low computational complexity and maintains an excellent signal quality.
Gerald Enzner, Peter Vary
ICASSP (5)1
2002 Unbiased residual echo power estimation for hands-free telephony
abstract
Residual echo arises in hands-free telephony equipment due to insufficient echo canceler convergence, but can be suppressed using a postfilter. The most important control parameter for postfilter adaptation is therefore the residual echo power spectral density (PSD). In this contribution we present and compare residual echo PSD estimation techniques. We introduce a new partitioned block-adaptive estimator delivering unbiased residual echo PSD estimates in strongly reverberant and noisy acoustic environments.
Gerald Enzner, Rainer Martin 0001, Peter Vary
ICASSP1
2002 Coded continuous phase modulation with low-complexity noncoherent reception
abstract
Coded continuous phase modulation based on a feedback-free modulator with noncoherent detection is discussed. Low-complexity receiver processing is achieved by using only two or three linear filters for demodulation and applying noncoherent sequence estimation with reduced-state Viterbi decoding and simple branch metric calculation. Overall, the proposed noncoherent receiver provides significant advantages over previously presented approaches.
Lutz Lampe, Robert Schober, Gerald Enzner, Johannes B. Huber
IEEE Trans. Commun.3
2001 Noncoherent coded continuous phase modulation
abstract
Continuous phase modulation (CPM) systems for noncoherent coded transmission are proposed and analyzed. Specifically, the application of a feedback-free modulator is regarded. For demodulation, a receiver structure which requires only two or three linear filters is considered. For decoding, noncoherent sequence estimation (NSE) with Viterbi decoding and per-survivor processing is applied. Noteworthy, in our approach the problems of low-complexity filtering and reduced-state decoding can be treated separately. Since we give a recursive formula for the phase reference symbol necessary for NSE metric calculation, computational effort is further decreased. Overall, in terms of complexity, the proposed noncoherent receiver provides significant advantages over approaches in the literature. The high performance of the novel noncoherent CPM system is confirmed by simulation results.
Lutz Lampe, Robert Schober, Gerald Enzner, Johannes B. Huber
ICC3