VLDB 2026 Research / reviewers in the wild / expert
Hendrik Schröter
dblp:243/6805
· DBLP profile ↗
9ranked-venue papers
8as first author
6since 2021 · last 2025
0000-0003-3147-3014ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 5 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DFingerNet: Noise-Adaptive Speech Enhancement for Hearing AidsabstractThe DeepFilterNet (DFN) architecture was recently proposed as a deep learning model suited for hearing aid devices. Despite its competitive performance on numerous benchmarks, it still follows a ‘one-size-fits-all’ approach, which aims to train a single, monolithic architecture that generalises across different noises and environments. However, its limited size and computation budget can hamper its generalisability. Recent work has shown that in-context adaptation can improve performance by conditioning the denoising process on additional information extracted from background recordings to mitigate this. These recordings can be offloaded outside the hearing aid, thus improving performance while adding minimal computational overhead. We introduce these principles to the DFN model, thus proposing the DFingerNet (DFiN) model, which shows superior performance on various benchmarks inspired by the DNS Challenge. Iosif Tsangko, Andreas Triantafyllopoulos, Hendrik Schröter, Björn W. Schuller |
ICASSP | 4 |
| 2023 | DeepFilterNet: Perceptually Motivated Real-Time Speech Enhancement
Hendrik Schröter, Alberto N. Escalante, Tobias Rosenkranz, Andreas K. Maier |
INTERSPEECH | 1 |
| 2023 | Deep Multi-Frame Filtering for Hearing Aids
Hendrik Schröter, Tobias Rosenkranz, Alberto N. Escalante, Andreas K. Maier |
INTERSPEECH | 1 |
| 2022 | Deepfilternet: A Low Complexity Speech Enhancement Framework for Full-Band Audio Based On Deep FilteringabstractComplex-valued processing has brought deep learning-based speech enhancement and signal extraction to a new level. Typically, the process is based on a time-frequency (TF) mask which is applied to a noisy spectrogram, while complex masks (CM) are usually preferred over real-valued masks due to their ability to modify the phase. Recent work proposed to use a complex filter instead of a point-wise multiplication with a mask. This allows to incorporate information from previous and future time steps exploiting local correlations within each frequency band.In this work, we propose DeepFilterNet, a two stage speech enhancement framework utilizing deep filtering. First, we enhance the spectral envelope using ERB-scaled gains modeling the human frequency perception. The second stage employs deep filtering to enhance the periodic components of speech. Additionally to taking advantage of perceptual properties of speech, we enforce network sparsity via separable convolutions and extensive grouping in linear and recurrent layers to design a low complexity architecture.We further show that our two stage deep filtering approach outperforms complex masks over a variety of frequency resolutions and latencies and demonstrate convincing performance compared to other state-of-the-art models. Hendrik Schröter, Alberto N. Escalante, Tobias Rosenkranz, Andreas K. Maier |
ICASSP | 1 |
| 2022 | Low Latency Speech Enhancement for Hearing Aids Using Deep FilteringabstractNoise reduction is an important feature supporting hearing aid (HA) users in their daily routines and is thus included in most commercially available devices. Latency requirements of HAs require short processing windows resulting in a poor frequency resolution in the whole processing chain including noise reduction. Previous studies have shown that deep neural network (DNN) based algorithms outperform conventional noise reduction algorithms especially for non-stationary noises. This study explores a DNN based noise reduction method using deep filtering targeted for wideband spectrograms given the employed HA filter bank. That is, we predict complex filter coefficients that are linearly applied to the noisy spectrum. We assess different filter sizes over time and frequency axis, and provide evidence for a superior performance over a complex ratio mask. Furthermore, we introduce a frequency response loss that operates on a per-frequency-band basis to fully utilize the deep filtering concept. We objectively demonstrate on-par performance with related state-of-the-art deep learning methods and show in a subjective user study that our method is perceptually preferred to existing HA noise reduction algorithms. Hendrik Schröter, Tobias Rosenkranz, Alberto N. Escalante, Andreas K. Maier |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2021 | LACOPE: Latency-Constrained Pitch Estimation for Speech Enhancement
Hendrik Schröter, Tobias Rosenkranz, Alberto N. Escalante, Andreas K. Maier |
Interspeech | 1 |
| 2020 | CLCNET: Deep Learning-Based Noise Reduction for Hearing aids using Complex Linear CodingabstractNoise reduction is an important part of modern hearing aids and is included in most commercially available devices. Deep learning-based state-of-the-art algorithms, however, either do not consider real-time and frequency resolution constrains or result in poor quality under very noisy conditions.To improve monaural speech enhancement in noisy environments, we propose CLCNet, a framework based on complex valued linear coding. First, we define complex linear coding (CLC) motivated by linear predictive coding (LPC) that is applied in the complex frequency domain. Second, we propose a framework that incorporates complex spectrogram input and coefficient output. Third, we define a parametric normalization for complex valued spectrograms that complies with low-latency and on-line processing.Our CLCNet was evaluated on a mixture of the EUROM database and a real-world noise dataset recorded with hearing aids and compared to traditional real-valued Wiener-Filter gains. Hendrik Schröter, Tobias Rosenkranz, Alberto N. Escalante, Marc Aubreville, Andreas K. Maier |
ICASSP | 1 |
| 2020 | Lightweight Online Noise Reduction on Embedded Devices Using Hierarchical Recurrent Neural NetworksabstractDeep-learning based noise reduction algorithms have proven their success especially for non-stationary noises, which makes it desirable to also use them for embedded devices like hearing aids (HAs). This, however, is currently not possible with state-of-the-art methods. They either require a lot of parameters and computational power and thus are only feasible using modern CPUs. Or they are not suitable for online processing, which requires constraints like low-latency by the filter bank and the algorithm itself. In this work, we propose a mask-based noise reduction approach. Using hierarchical recurrent neural networks, we are able to drastically reduce the number of neurons per layer while including temporal context via hierarchical connections. This allows us to optimize our model towards a minimum number of parameters and floating-point operations (FLOPs), while preserving noise reduction quality compared to previous work. Our smallest network contains only 5k parameters, which makes this algorithm applicable on embedded devices. We evaluate our model on a mixture of EUROM and a real-world noise database and report objective metrics on unseen noise. Hendrik Schröter, Tobias Rosenkranz, Alberto N. Escalante, Pascal Zobel, Andreas K. Maier |
INTERSPEECH | 1 |
| 2019 | Segmentation, Classification, and Visualization of Orca Calls Using Deep LearningabstractAudiovisual media are increasingly used to study the communication and behavior of animal groups, e.g. by placing microphones in the animals habitat resulting in huge datasets with only a small amount of animal interactions. The Orcalab has recorded orca whales since 1973 using stationary underwater hydrophones and made it publicly available on the Orchive. There exist over 15 000 manually extracted orca/noise annotations and about 20 000 h unseen audio data. To analyze the behavior and communication of killer whales we need to interpret the different call types. In this work, we present a two-stage classification approach using the labeled call/noise files and a few labeled call-type files. Results indicate a reliable accuracy of 95.0 % for call segmentation and 87 % for classification of 12 call classes. We further visualize the learned orca call representations in the convolutional neural network (CNN) activations to explain the potential of CNN based recognition for bioaccousitc signals. Hendrik Schröter, Elmar Nöth, Andreas K. Maier, Rachael Cheng, Volker Barth, Christian Bergler |
ICASSP | 1 |