Tobias Rosenkranz

dblp:93/8055 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
7since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2023 Short-term Extrapolation of Speech Signals Using Recursive Neural Networks in the STFT Domain
Maurice Oberhag, Daniel Neudek, Rainer Martin 0001, Tobias Rosenkranz, Henning Puder
INTERSPEECH4
2023 DeepFilterNet: Perceptually Motivated Real-Time Speech Enhancement
Hendrik Schröter, Alberto N. Escalante, Tobias Rosenkranz, Andreas K. Maier
INTERSPEECH3
2023 Deep Multi-Frame Filtering for Hearing Aids
Hendrik Schröter, Tobias Rosenkranz, Alberto N. Escalante, Andreas K. Maier
INTERSPEECH2
2022 Deepfilternet: A Low Complexity Speech Enhancement Framework for Full-Band Audio Based On Deep Filtering
abstract
Complex-valued processing has brought deep learning-based speech enhancement and signal extraction to a new level. Typically, the process is based on a time-frequency (TF) mask which is applied to a noisy spectrogram, while complex masks (CM) are usually preferred over real-valued masks due to their ability to modify the phase. Recent work proposed to use a complex filter instead of a point-wise multiplication with a mask. This allows to incorporate information from previous and future time steps exploiting local correlations within each frequency band.In this work, we propose DeepFilterNet, a two stage speech enhancement framework utilizing deep filtering. First, we enhance the spectral envelope using ERB-scaled gains modeling the human frequency perception. The second stage employs deep filtering to enhance the periodic components of speech. Additionally to taking advantage of perceptual properties of speech, we enforce network sparsity via separable convolutions and extensive grouping in linear and recurrent layers to design a low complexity architecture.We further show that our two stage deep filtering approach outperforms complex masks over a variety of frequency resolutions and latencies and demonstrate convincing performance compared to other state-of-the-art models.
Hendrik Schröter, Alberto N. Escalante, Tobias Rosenkranz, Andreas K. Maier
ICASSP3
2022 First-Order Recursive Smoothing of Short-Time Power Spectra in the Presence of Interference
abstract
In this letter we derive an optimal time-varying smoothing factor for smoothing non-stationary power spectra of a target signal when a noisy observation of this signal is given. The proposed approach is based on a complex Gaussian signal model in the short-time discrete Fourier domain and on mean squared error optimization. We find that the optimal smoothing factor depends on the signal-to-noise ratio as well as on the deviation between the smoothed estimate and the target signal power spectra. We investigate the properties of the proposed adaptive smoothing system and demonstrate its utility in a basic validation experiment.
Jalal Taghia, Daniel Neudek, Tobias Rosenkranz, Henning Puder, Rainer Martin 0001
IEEE Signal Process. Lett.3
2022 Low Latency Speech Enhancement for Hearing Aids Using Deep Filtering
abstract
Noise reduction is an important feature supporting hearing aid (HA) users in their daily routines and is thus included in most commercially available devices. Latency requirements of HAs require short processing windows resulting in a poor frequency resolution in the whole processing chain including noise reduction. Previous studies have shown that deep neural network (DNN) based algorithms outperform conventional noise reduction algorithms especially for non-stationary noises. This study explores a DNN based noise reduction method using deep filtering targeted for wideband spectrograms given the employed HA filter bank. That is, we predict complex filter coefficients that are linearly applied to the noisy spectrum. We assess different filter sizes over time and frequency axis, and provide evidence for a superior performance over a complex ratio mask. Furthermore, we introduce a frequency response loss that operates on a per-frequency-band basis to fully utilize the deep filtering concept. We objectively demonstrate on-par performance with related state-of-the-art deep learning methods and show in a subjective user study that our method is perceptually preferred to existing HA noise reduction algorithms.
Hendrik Schröter, Tobias Rosenkranz, Alberto N. Escalante, Andreas K. Maier
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 LACOPE: Latency-Constrained Pitch Estimation for Speech Enhancement
Hendrik Schröter, Tobias Rosenkranz, Alberto N. Escalante, Andreas K. Maier
Interspeech2
2020 CLCNET: Deep Learning-Based Noise Reduction for Hearing aids using Complex Linear Coding
abstract
Noise reduction is an important part of modern hearing aids and is included in most commercially available devices. Deep learning-based state-of-the-art algorithms, however, either do not consider real-time and frequency resolution constrains or result in poor quality under very noisy conditions.To improve monaural speech enhancement in noisy environments, we propose CLCNet, a framework based on complex valued linear coding. First, we define complex linear coding (CLC) motivated by linear predictive coding (LPC) that is applied in the complex frequency domain. Second, we propose a framework that incorporates complex spectrogram input and coefficient output. Third, we define a parametric normalization for complex valued spectrograms that complies with low-latency and on-line processing.Our CLCNet was evaluated on a mixture of the EUROM database and a real-world noise dataset recorded with hearing aids and compared to traditional real-valued Wiener-Filter gains.
Hendrik Schröter, Tobias Rosenkranz, Alberto N. Escalante, Marc Aubreville, Andreas K. Maier
ICASSP2
2020 Lightweight Online Noise Reduction on Embedded Devices Using Hierarchical Recurrent Neural Networks
abstract
Deep-learning based noise reduction algorithms have proven their success especially for non-stationary noises, which makes it desirable to also use them for embedded devices like hearing aids (HAs). This, however, is currently not possible with state-of-the-art methods. They either require a lot of parameters and computational power and thus are only feasible using modern CPUs. Or they are not suitable for online processing, which requires constraints like low-latency by the filter bank and the algorithm itself. In this work, we propose a mask-based noise reduction approach. Using hierarchical recurrent neural networks, we are able to drastically reduce the number of neurons per layer while including temporal context via hierarchical connections. This allows us to optimize our model towards a minimum number of parameters and floating-point operations (FLOPs), while preserving noise reduction quality compared to previous work. Our smallest network contains only 5k parameters, which makes this algorithm applicable on embedded devices. We evaluate our model on a mixture of EUROM and a real-world noise database and report objective metrics on unseen noise.
Hendrik Schröter, Tobias Rosenkranz, Alberto N. Escalante, Pascal Zobel, Andreas K. Maier
INTERSPEECH2
2012 Integrating recursive minimum tracking and codebook-based noise estimation for improved reduction of non-stationary noise
Tobias Rosenkranz, Henning Puder
Signal Process.1
2012 Erratum to "Integrating recursive minimum tracking and codebook-based noise estimation for improved reduction of non-stationary noise" [Signal Process. 92(3) (2012) 767-779]
Tobias Rosenkranz, Henning Puder
Signal Process.1
2012 Improving Robustness of Codebook-Based Noise Estimation Approaches With Delta Codebooks
abstract
We present a new codebook-based speech enhancement approach which is able to increase robustness of conventional codebook-based approaches against model mismatch and unknown noise types. This is achieved by training only the difference between the actual noise and a robust estimate (e.g., obtained by minimum statistics or recursive minimum tracking) in the cepstral domain instead of the noise itself. The noise codebook is then generated by shifting the so obtained delta-codebook by the cepstral representation of a robust noise estimate. We use the recursive minimum tracking approach as robust estimate. It is thus guaranteed that the robust estimate is also a valid estimate of the codebook-based algorithm. Consequently, the codebook-based algorithm inherits the robustness from the recursive minimum tracking approach. Objective and subjective experiments show that the proposed method yields a consistent quality improvement over the basic codebook-based approach and recursive minimum tracking.
Tobias Rosenkranz, Henning Puder
IEEE Trans. Speech Audio Process.1
2009 Multidimensional localization of multiple sound sources using averaged directivity patterns of Blind Source Separation systems
abstract
In this paper, we propose a versatile acoustic source localization framework exploiting the self-steering capability of Blind Source Separation (BSS) algorithms. We provide a way to produce an acoustical map of the scene by computing the averaged directivity pattern of BSS demixing systems. Since BSS explicitly accounts for multiple sources in its signal propagation model, several simultaneously active sound sources can be located using this method. Moreover, the framework is suitable to any microphone array geometry, which allows application for multiple dimensions, in the near field as well as in the far field. Experiments demonstrate the efficiency of the proposed scheme in a reverberant environment for the localization of speech sources.
Anthony Lombard, Tobias Rosenkranz, Herbert Buchner, Walter Kellermann
ICASSP2