Timo Lohrenz

dblp:199/9553 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
9since 2021 · last 2024
0000-0001-7103-4302ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 10 · 6 first-author · 7 since 2021
YearPublicationVenuePosition
2024 Mixed Children/Adult/Childrenized Fine-Tuning for Children's ASR: How to Reduce Age Mismatch and Speaking Style Mismatch
Thomas Graave, Timo Lohrenz, Tim Fingscheidt
INTERSPEECH3
2024 Interleaved Audio/Audiovisual Transfer Learning for AV-ASR in Low-Resourced Languages
Patrick Blumenberg, Thomas Graave, Timo Lohrenz, Siegfried Kunzmann, Tim Fingscheidt
INTERSPEECH5
2023 Parameter-Efficient Cross-Language Transfer Learning for a Language-Modular Audiovisual Speech Recognition
abstract
In audiovisual speech recognition (AV-ASR), for many languages only few audiovisual data is available. Building upon an English model, in this work, we first apply and analyze various adapters for cross-language transfer learning to build a parameter-efficient and easy-to-extend AV-ASR in multiple languages. Fine-tuning only the bottleneck adapter with 4% of encoder’s parameters and the decoder shows comparable performance to full fine-tuning in French and Spanish AV-ASR. Second, we investigate the effectiveness of various encoder components in cross-language transfer learning. Our proposed modular linguistic transfer learning approach excels the full fine-tuning method for German, French, and Spanish AV-ASR in almost all clean and noisy conditions (8/9). On low-resourced German AV data (13h), our proposed linguistic transfer learning achieves a 4.1% abs. WER reduction on average for clean and noisy speech, while fine-tuning only 50% of the encoder’s parameters. Our code is at GitHub.11https://github.com/ifnspaml/Cross_Language_Transfer_Learning_AVASR.git
Thomas Graave, Timo Lohrenz, Siegfried Kunzmann, Tim Fingscheidt
ASRU4
2023 Relaxed Attention for Transformer Models
abstract
The powerful modeling capabilities of all-attention-based transformer architectures often cause overfitting and-for natural language processing tasks-lead to an implicitly learned internal language model in the autoregressive transformer decoder complicating the integration of external language models. In this paper, we explore relaxed attention, a simple and easy-to-implement smoothing of the attention weights, yielding a two-fold improvement to the general transformer architecture: First, relaxed attention provides regularization when applied to the self-attention layers in the encoder. Second, we show that it naturally supports the integration of an external language model as it suppresses the implicitly learned internal language model by relaxing the cross attention in the decoder. We demonstrate the benefit of relaxed attention across several tasks from different applications with clear improvement in combination with recent benchmark approaches using various transformer model variants and sizes. Specifically, we exceed the former state-of-the-art performance of 26.90% word error rate on the largest public lip-reading LRS3 benchmark with a word error rate of 26.31%, as well as we achieve a top-performing BLEU score of 37.67 on the IWSLT14 (DE → EN) machine translation task without external language models and virtually no additional model parameters.
Timo Lohrenz, Björn Möller, Tim Fingscheidt
IJCNN1
2023 An Efficient and Noise-Robust Audiovisual Encoder for Audiovisual Speech Recognition
Chenwei Liang, Timo Lohrenz, Marvin Sach, Björn Möller, Tim Fingscheidt
INTERSPEECH3
2022 Transformer-Based Lip-Reading with Regularized Dropout and Relaxed Attention
abstract
End-to-end automatic lip-reading usually comprises an encoder-decoder model and an optional external language model. In this work, we introduce two regularization methods to the field of lip-reading: First, we apply the regularized dropout (R-Drop) method to transformer-based lip-reading to improve their training-inference consistency. Second, the relaxed attention technique is applied during training for a better external language model integration. We are the first to show that these two complementary approaches yield particu1arly strong performance if combined in the right manner. In particular, by adding an additional R - Drop loss and smoothing the attention weights in cross multi-head attention during training only, we achieve a new state of the art with a word error rate of 22.2% on Lip Reading Sentences 2 (LRS2). On LRS3, we are 2nd ranked with 25.5% WER using only 1,759 h of training data, while the 1 st rank uses about 90,000 h. Our code is available at GitHub.11https://github.com/ifnspaml/Lipreading-RDrop-RA
Timo Lohrenz, Matthias Dunkelberg, Tim Fingscheidt
SLT2
2021 Relaxed Attention: A Simple Method to Boost Performance of End-to-End Automatic Speech Recognition
abstract
Recently, attention-based encoder-decoder (AED) models have shown high performance for end-to-end automatic speech recognition (ASR) across several tasks. Addressing overconfidence in such models, in this paper we introduce the concept of relaxed attention, which is a simple gradual injection of a uniform distribution to the encoder-decoder attention weights during training that is easily implemented with two lines of code. We investigate the effect of relaxed attention across different AED model architectures and two prominent ASR tasks, Wall Street Journal (WSJ) and Librispeech. We found that transformers trained with relaxed attention outperform the standard baseline models consistently during decoding with external language models. On WSJ, we set a new benchmark for transformer-based end-to-end speech recognition with a word error rate of 3.65%, outperforming state of the art (4.20%) by 13.1% relative, while introducing only a single hyperparameter.
Timo Lohrenz, Patrick Schwarz, Tim Fingscheidt
ASRU1
2021 A New DCASE 2017 Rare Sound Event Detection Benchmark Under Equal Training Data: CRNN With Multi-Width Kernels
abstract
Rare sound event detection (rare SED) deals with obtaining valuable information from data consisting mostly of acoustic background noises. It has meanwhile a long research history and was part of the DCASE 2017 Challenge. State-of-the-art performance is currently reached using a stacked combination of a CNN and an RNN, dubbed CRNN, which was also successfully applied in other domains such as in hybrid automatic speech recognition. In this work, we propose a new CRNN model for rare SED. This new model uses a set of parallel convolutions with multiple kernel widths in the CRNN and is based on an extended feature representation of the log-mel spectrogram. Furthermore, we apply and optimize different evaluation postprocessing methods and analyze the modifications in an ablation study. The proposed model outperforms the so-far top-scoring networks of the DCASE Challenge – using the same training material for all methods – by an error rate of 6.13% absolute and by 4.39% absolute in the F1 score on the test set and under these conditions achieves a new benchmark result on the DCASE 2017 Rare SED data set.
Jan Baumann, Patrick Meyer, Timo Lohrenz, Alexander Roy, Michael Papendieck, Tim Fingscheidt
ICASSP3
2021 Multi-Encoder Learning and Stream Fusion for Transformer-Based End-to-End Automatic Speech Recognition
abstract
Stream fusion, also known as system combination, is a common technique in automatic speech recognition for traditional hybrid hidden Markov model approaches, yet mostly unexplored for modern deep neural network end-to-end model architectures. Here, we investigate various fusion techniques for the all-attention-based encoder-decoder architecture known as the transformer, striving to achieve optimal fusion by investigating different fusion levels in an example single-microphone setting with fusion of standard magnitude and phase features. We introduce a novel multi-encoder learning method that performs a weighted combination of two encoder-decoder multi-head attention outputs only during training. Employing then only the magnitude feature encoder in inference, we are able to show consistent improvement on Wall Street Journal (WSJ) with language model and on Librispeech, without increase in runtime or parameters. Combining two such multi-encoder trained models by a simple late fusion in inference, we achieve state-of-the-art performance for transformer-based models on WSJ with a significant WER reduction of 19% relative compared to the current benchmark approach.
Timo Lohrenz, Tim Fingscheidt
Interspeech1
2020 Beyond the Dcase 2017 Challenge on Rare Sound Event Detection: A Proposal for a More Realistic Training and Test Framework
abstract
There are many ways to evaluate rare sound event detection (SED) approaches, e.g., the DCASE 2017 challenge provides a widely employed framework. This paper proposes a rare SED training and test framework, which is reflecting an SED application in a more realistic way. Our setup gets rid of too much prior knowledge on the test data, and assumes additional unknown acoustic events both in training and test data, which in practice have to be identified as background. Taking this into account during training, the robustness in real-world scenarios can be significantly increased, with an average event-based error rate reduction of an absolute 34 percentage points. Further we show and compare the performance of multi-event (polyphonic) classifiers vs. single-event classifiers while outlining the benefits of multi-event training.
Jan Baumann, Timo Lohrenz, Alexander Roy, Tim Fingscheidt
ICASSP2
2020 BLSTM-Driven Stream Fusion for Automatic Speech Recognition: Novel Methods and a Multi-Size Window Fusion Example
Timo Lohrenz, Tim Fingscheidt
INTERSPEECH1
2019 On Temporal Context Information for Hybrid BLSTM-Based Phoneme Recognition
abstract
The modern approach to include long-term temporal context information into speech recognition systems is the use of recurrent neural networks, e.g., bi-directional long short-term memory (BLSTM) networks. In this paper, we decouple the BLSTM from a preceding CNN-based feature extractor network allowing us to investigate the use of temporal context in both models in a modular fashion. Accordingly, we train the BLSTMs on posteriors, stemming from preceding CNNs which use various amounts of limited context in their input layer, and investigate to what extent the BLSTM is able to effectively make use of its long-term modeling capabilities. We show that it is beneficial to train the BLSTM on posteriors stemming from a temporal context-free acoustic model. Remarkably, the best performing combination of CNN acoustic model and BLSTM afterwards is a large-context CNN (expected), followed by a BLSTM which has been trained on context-free CNN output posteriors (surprising).
Timo Lohrenz, Maximilian Strake, Tim Fingscheidt
ASRU1
2018 A New Timit Benchmark for Context-Independent Phone Recognition Using Turbo Fusion
abstract
In this work, we apply the recently proposed turbo fusion in conjunction with state-of-the-art convolutional neural networks as acoustic models to the standard phone recognition task on the TIMIT database. The turbo fusion operates on posterior streams stemming from standard filterbank features and from group delay (phase) features. By the iterative exchange of posterior information, the phone error rate is decreased down to 16.91% absolute, which is to our knowledge the best reported result on the TIMIT core test set so far using context-independent acoustic models, outperforming the previous respective benchmark by 4.4% relative.
Timo Lohrenz, Wei Li 0174, Tim Fingscheidt
SLT1
2018 Densenet Blstm for Acoustic Modeling in Robust ASR
abstract
In recent years, robust automatic speech recognition (ASR) has greatly taken benefit from the use of neural networks for acoustic modeling, although performance still degrades in severe noise conditions. Based on the previous success of models using convolutional and subsequent bidirectional long short-term memory (BLSTM) layers in the same network, we propose to use a densely connected convolutional network (DenseNet) as the first part of such a model, while the second is a BLSTM network. A particular contribution of our work is that we modify the DenseNet topology to become a kind of feature extractor for the subsequent BLSTM network operating on whole speech utterances. We evaluate our model on the 6-channel task of CHiME-4, and are able to consistently outperform a top-performing baseline based on wide residual networks and BLSTMs providing a 2.4% relative WER reduction on the real test set.
Maximilian Strake, Pascal Behr, Timo Lohrenz, Tim Fingscheidt
SLT3
2017 Turbo fusion of magnitude and phase information for DNN-based phoneme recognition
abstract
In this work we propose the so-called turbo fusion as competitive method for information fusion of Mel-filterbank magnitude and phase feature streams for automatic speech recognition (ASR). Based on the recently introduced turbo ASR paradigm, our contribution is fourfold: First, we introduce DNN-based acoustic modeling into turbo ASR, then we take steps towards LVCSR by omitting the costly state space transform and by investigating the classical TIMIT phoneme recognition task. Finally, replacing the typical stream weighting in fusion methods, we introduce a new dynamic range limitation of the exchanged posteriors between the involved magnitude and phase recognizers, resulting in a smoother information exchange. The proposed turbo fusion outperforms classical benchmarks on the TIMIT dataset both with and without dropout in DNN training, and also is first if compared to several state-of-the-art reference fusion methods.
Timo Lohrenz, Tim Fingscheidt
ASRU1