Wolfgang Mack

dblp:226/2104 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Exploring Task Prioritization in VR Flight Environments: Can Eye-Tracking Uncover Cognitive Control?
abstract
Cognitive control is important for multitasking performance, which is essential for safety and efficiency in human-machine interaction domains such as aviation. Previous research has shown that the stability-flexibility dilemma of cognitive control can be manipulated through task prioritization in low-fidelity flight environments, with eye-tracking metrics effectively classifying this dilemma in a low-fidelity flight environment using a machine learning approach. However, it is unclear whether these results extend to higher-fidelity environments. This study examines the applicability of this approach to eye-tracking metrics assessed in a VR flight simulator. Linear mixed-effects models reveal significant differences in fixation duration, relative number of fixations, coefficient K, stationary entropy, and explore-exploit ratio across flight scenarios with varying task prioritization. Additionally, these metrics were used as input features in a machine learning model to classify the mission scenarios. The findings have important implications for designing adaptive assistance systems aiming at supporting multitasking performance in safety-critical environments.
Sophie-Marie Stasch, Wolfgang Mack
ETRA2
2025 Improving Audio Classification by Transitioning from Zero- to Few-Shot
Wolfgang Mack
INTERSPEECH2
2025 Inference-Adaptive Steering of Neural Networks for Real-Time Area-Based Sound Source Separation
abstract
We propose a novel adaptive steering technique that changes the target area of a spatial-aware multi-microphone sound source separation algorithm during inference without the necessity of retraining the deep neural network (DNN). To achieve this, we first train a DNN aiming to retain speech within a target region, defined by an angular span, while suppressing sound sources stemming from other directions. Afterward, a phase shift is applied to the microphone signals, allowing us to shift the center of the target area during inference at negligible additional cost in computational complexity. Further, we show that the proposed approach performs well in a wide variety of acoustic scenarios, including several speakers inside and outside the target area and additional noise. More precisely, the proposed approach performs on par with DNNs trained explicitly for the steered target area in terms of DNSMOS and SI-SDR.
Martin Strauss 0003, Wolfgang Mack, Maria Luis Valero, Okan Köpüklü
IEEE Signal Process. Lett.2
2023 Multi-Microphone Speaker Separation by Spatial Regions
abstract
We consider the task of region-based source separation of reverberant multi-microphone recordings. We assume pre-defined spatial regions with a single active source per region. The objective is to estimate the signals from the individual spatial regions as captured by a reference microphone while retaining a correspondence between signals and spatial regions. We propose a data-driven approach using a modified version of a state-of-the-art network, where different layers model spatial and spectro-temporal information. The network is trained to enforce a fixed mapping of regions to network outputs. Using speech from LibriMix, we construct a data set specifically designed to contain the region information. Additionally, we train the network with permutation invariant training. We show that both training methods result in a fixed mapping of regions to network outputs, achieve comparable performance, and that the networks exploit spatial information. The proposed network outperforms a baseline network by 1.5 dB in scale-invariant signal-to-distortion ratio.
Julian Wechsler, Srikanth Raj Chetupalli, Wolfgang Mack, Emanuël A. P. Habets
ICASSP3
2022 A Hybrid Acoustic Echo Reduction Approach Using Kalman Filtering and Informed Source Extraction with Improved Training
abstract
State-of-the-art acoustic echo and noise reduction combines adaptive filters with a deep neural network-based postfilter. While the signal-to-distortion ratio is often used for training, it is not well-defined for all echo-reduction scenarios. We propose well-defined loss functions for training and modifications of a recently proposed echo reduction system that is based on informed source extraction. The modifications include using a Kalman filter as a prefilter and a cyclical learning rate scheduler. The proposed modifications improve the performance on the blind test set of the Interspeech 2021 AEC challenge. A comparison to the challenge-winner shows that the proposed system underperforms the winner by 0.1 mean opinion score (MOS) points in double-talk echo reduction. However, it outperforms the winner by 0.3 MOS points in echo-only echo reduction. In all other scenarios, both algorithms perform comparably.
Wolfgang Mack, Emanuël A. P. Habets
SLT1
2022 Signal-aware direction-of-arrival estimation using attention mechanisms
Wolfgang Mack, Julian Wechsler, Emanuël A. P. Habets
Comput. Speech Lang.1
2021 Efficient Training Data Generation for Phase-Based DOA Estimation
abstract
Deep learning (DL) based direction of arrival (DOA) estimation is an active research topic and currently represents the state-of-the-art. Usually, DL-based DOA estimators are trained with recorded data or computationally expensive generated data. Both data types require significant storage and excessive time to, respectively, record or generate. We propose a low complexity online data generation method to train DL models with a phase-based feature input. The data generation method models the phases of the microphone signals in the frequency domain by employing a deterministic model for the direct path and a statistical model for the late reverberation of the room transfer function. By an evaluation using data from measured room impulse responses, we demonstrate that a model trained with the proposed training data generation method performs comparably to models trained with data generated based on the source-image method.
Fabian Hübner, Wolfgang Mack, Emanuël A. P. Habets
ICASSP2
2020 Signal-Aware Broadband DOA Estimation Using Attention Mechanisms
abstract
We refer to direction-of-arrivals (DOAs) estimation of a user-defined subset of directional (desired) sound sources as signal-aware DOA estimation. Source selection, thereby, can be achieved with time-frequency masks to apply attention to TF bins dominated by desired sources. With deep neural networks (DNNs), another option is to train the DNN to estimate the DOAs only of specific classes, like speech, and disregard the DOAs of other classes. Consequently, changing the desired classes requires retraining the DNN. Also, the mask-based approaches are trained for sources known prior to DNN training. To obtain a flexible signal-aware DOA estimator, we propose to use binary mask attention with a DNN for multi-source DOA estimation trained with artificial noise. The desired sources are determined via binary masks, which allows a redefinition by changing the masks. Consequently, the DOA estimator is independent of the desired sources. We experiment with attention in form of oracle and estimated binary masks.
Wolfgang Mack, Ullas Bharadwaj, Soumitro Chakrabarty, Emanuël A. P. Habets
ICASSP1
2020 Data-Driven Wind Speed Estimation Using Multiple Microphones
abstract
A deep neural network (DNN) based approach for estimating the speed of airflows using closely-spaced microphones is proposed. The spatial characteristics of wind noise measured with a smallaperture array are exploited, i.e., the low-frequency spatial coherence of wind noise signals is used as an input feature. The output is an estimate of the wind speed averaged over a specific time interval. The DNN is trained using synthetic wind noise, which overcomes the time-consuming data collection and allows to isolate wind noise from different acoustic sources. The dataset used for testing comprises wind noise measured outdoors with a circular linear array and a ground truth obtained using an ultrasonic anemometer. The obtained model is applied to generated and measured wind noise. The performance of the proposed method is assessed across a wide range of wind speeds and directions, using different time resolutions.
Daniele Mirabilii, Kishor Kayyar Lakshminarayana, Wolfgang Mack, Emanuël A. P. Habets
ICASSP3
2020 Online Blind Reverberation Time Estimation Using CRNNs
abstract
S.5061-5065
Shuwen Deng, Wolfgang Mack, Emanuël A. P. Habets
INTERSPEECH2
2020 Single-Channel Blind Direct-to-Reverberation Ratio Estimation Using Masking
abstract
Acoustic parameters, like the direct-to-reverberation ratio (DRR), can be used in audio processing algorithms to perform, e.g., dereverberation or in audio augmented reality. Often, the DRR is not available and has to be estimated blindly from recorded audio signals. State-of-the-art DRR estimation is achieved by deep neural networks (DNNs), which directly map a feature representation of the acquired signals to the DRR. Motivated by the equality of the signal-to-reverberation ratio and the (channel-based) DRR under certain conditions, we formulate single-channel DRR estimation as an extraction task of two signal components from the recorded audio. The DRR can be obtained by inserting the estimated signals in the definition of the DRR. The extraction is performed using time-frequency masks. The masks are estimated by a DNN trained end-to-end to minimize the mean-squared error between the estimated and the oracle DRR. We conduct experiments with different preprocessing and mask estimation schemes. The proposed method outperforms state-of-the-art single- and multi-channel methods on the ACE challenge data corpus.
Wolfgang Mack, Shuwen Deng, Emanuël A. P. Habets
INTERSPEECH1
2020 Deep Filtering: Signal Extraction and Reconstruction Using Complex Time-Frequency Filters
abstract
Signal extraction from a single-channel mixture with additional undesired signals is most commonly performed using time-frequency (TF) masks. Typically, the mask is estimated with a deep neural network (DNN), and element-wise applied to the complex mixture short-time Fourier transform (STFT) representation to perform the extraction. Ideal mask magnitudes are zero for solely undesired signals in a TF bin and undefined for total destructive interference. Usually, masks have an upper bound to provide well-defined DNN outputs at the cost of limited extraction capabilities. We propose to estimate with a DNN a complex TF filter for each mixture TF bin which maps an STFT area in the respective mixture to the desired TF bin to address destructive interference in mixture TF bins. The DNN is optimized by minimizing the error between the extracted and the ground-truth desired signal allowing to learn the TF filters without having to specify ground-truth TF filters. We compare our approach with complex and real-valued TF masks by separating speech from a variety of different sound and noise classes from the Google AudioSet corpus. We also process the mixture STFT with notch-filters and zero whole time-frames, to simulate packet-loss during transmission, to demonstrate the reconstruction capabilities of our approach. The proposed method outperformed the baselines, especially when notch-filters and time-frame zeroing were applied.
Wolfgang Mack, Emanuël A. P. Habets
IEEE Signal Process. Lett.1
2018 Single-Channel Dereverberation Using Direct MMSE Optimization and Bidirectional LSTM Networks
abstract
Dereverberation is useful in hands-free communication and voice controlled devices for distant speech acquisition. Single-channel dereverberation can be achieved by applying a time-frequency (TF) mask to the short-time Fourier transform (STFT) representation of a reverberant signal. Recent approaches have used deep neural networks (DNNs) to estimate such masks. Previously proposed DNN-based mask estimation methods train a DNN to minimize the mean-squared-error (MSE) between the desired and estimated masks. Recent TF mask estimation methods for signal separation directly minimize instead the MSE between the desired and estimated STFT magnitudes. We apply this direct optimization concept to dereverberation. Moreover, as reverberation exceeds the duration of a single STFT frame, we propose to use a bidirectional long short-term memory (LSTM) network which is able to take the relation between multiple STFT frames into account. We evaluated our method for different reverberation times and source-microphone distances using simulated as well as measured room impulse responses of different rooms. An evaluation of the proposed method and a comparison with a state-of-the-art method demonstrate the superiority of our approach and its robustness to different acoustic conditions.
Wolfgang Mack, Soumitro Chakrabarty, Fabian-Robert Stöter, Sebastian Braun, Bernd Edler, Emanuël A. P. Habets
INTERSPEECH1