Tim Fingscheidt

dblp:03/7820 · DBLP profile ↗
← Back
121ranked-venue papers
11as first author
38since 2021 · last 2025
0000-0002-8895-5041ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 83 · 8 first-author · 25 since 2021Artificial intelligence and machine learning · 68 · 6 first-author · 26 since 2021Computer networks · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Less is More: Data Curation Matters in Scaling Speech Enhancement
abstract
The vast majority of modern speech enhancement systems rely on data-driven neural network models. Conventionally, larger datasets are presumed to yield superior model performance, an observation empirically validated across numerous tasks in other domains. However, recent studies reveal diminishing returns when scaling speech enhancement data. We focus on a critical factor: prevalent quality issues in “clean” training labels within large-scale datasets. This work re-examines this phenomenon and demonstrates that, within large-scale training sets, prioritizing high-quality training data is more important than merely expanding the data volume. Experimental findings suggest that models trained on a carefully curated subset of 700 hours can outperform models trained on the 2,500 -hour full dataset. This outcome highlights the crucial role of data curation in scaling speech enhancement systems effectively.
Chenda Li, Wangyou Zhang, Wei Wang 0010, Robin Scheibler, Kohei Saijo, Samuele Cornell, Yihui Fu, Marvin Sach, Zhaoheng Ni, Anurag Kumar 0003, Tim Fingscheidt, Shinji Watanabe 0001, Yanmin Qian
ASRU11
2025 URGENT-PK: Perceptually-Aligned Ranking Model Designed for Speech Enhancement Competition
abstract
The Mean Opinion Score (MOS) is fundamental to speech quality assessment. However, its acquisition requires significant human annotation. Although deep neural network approaches, such as DNSMOS and UTMOS, have been developed to predict MOS to avoid this issue, they often suffer from insufficient training data. Recognizing that the comparison of speech enhancement (SE) systems prioritizes a reliable system comparison over absolute scores, we propose URGENT-PK, a novel ranking approach leveraging pairwise comparisons. URGENT-PK takes homologous enhanced speech pairs as input to predict relative quality rankings. This pairwise paradigm efficiently utilizes limited training data, as all pairwise permutations of multiple systems constitute a training instance. Experiments across multiple open test sets demonstrate URGENT-PK’s superior system-level ranking performance over state-of-the-art baselines, despite its simple network architecture and limited training data.
Chenda Li, Wei Wang 0010, Wangyou Zhang, Samuele Cornell, Marvin Sach, Robin Scheibler, Kohei Saijo, Yihui Fu, Zhaoheng Ni, Anurag Kumar 0003, Tim Fingscheidt, Shinji Watanabe 0001, Yanmin Qian
ASRU12
2025 Efficient Noise-Robust Hybrid Audiovisual Encoder with Joint Distillation and Pruning for Audiovisual Speech Recognition
Pascal Reichert, Thomas Graave, Patrick Blumenberg, Tim Fingscheidt
INTERSPEECH5
2025 Interspeech 2025 URGENT Speech Enhancement Challenge
Kohei Saijo, Wangyou Zhang, Samuele Cornell, Robin Scheibler, Chenda Li, Zhaoheng Ni, Anurag Kumar 0003, Marvin Sach, Yihui Fu, Wei Wang 0010, Tim Fingscheidt, Shinji Watanabe 0001
INTERSPEECH11
2025 Lessons Learned from the URGENT 2024 Speech Enhancement Challenge
Wangyou Zhang, Kohei Saijo, Samuele Cornell, Robin Scheibler, Chenda Li, Zhaoheng Ni, Anurag Kumar 0003, Marvin Sach, Wei Wang 0010, Yihui Fu, Shinji Watanabe 0001, Tim Fingscheidt, Yanmin Qian
INTERSPEECH12
2025 Empirical Spatial Error Bounds for Reliable Semantic Segmentation of Pedestrians and Riders
abstract
The mean intersection over union (mloU) is a standard metric for evaluating semantic segmentation models. While steady improvements in mloU have been achieved on automotive benchmarks like Cityscapes, their impact on reliably detecting vulnerable road users, such as pedestrians and riders, remains unclear. This study empirically analyzes 167 semantic segmentation models w.r.t. the spatial distribution of the false positive rate and false negative rate in the Cityscapes dataset. Our analysis reveals that many segmentation errors occur at object contours, which hardly influence driving decisions and road user safety. Accordingly, we propose to exclude such irrelevant errors. We define spatial error bounds within which models reliably detect pedestrians and riders. Since time-to-collision is strongly related to distance, and a vertical pixel position is roughly related to distance, the vertical position of segmentation errors provides an effective way to evaluate the reliability of semantic segmentation models on an entire dataset. Our evaluation of such empirical spatial error bounds reveals that strong models (w.r.t. mloU) are related to an improved detection of existing pedestrians (false negative rate, FNR). On the other hand, mloU in general is only weakly related to hallucinations of pedestrians and riders (false positive rate, FPR). Some models even exhibit a higher FPR despite having a 11.2% absolute higher mloU.
Timo Bartels, Malte Stelzer, Jan Bickerdt, Volker Schomerus, Jan Piewek, Thorsten Bagdonat, Tim Fingscheidt
IV7
2025 A Generalized Waypoint Loss for End-to-End Autonomous Driving
abstract
Many approaches in autonomous driving generate future waypoints to form trajectories, which are then used to derive driving commands. During imitation learning for end-to-end autonomous driving, these trajectories are typically learned using a straightforward$L_{1}$loss, which compares the model's predictions to those of an expert. In this paper, we propose a separation of longitudinal and lateral components of the$L_{1}$loss that weighs these independently, thereby aligning more with the separate handling of longitudinal and lateral control by PID controllers in the model pipeline. We employ this novel generalized waypoint loss with the TransFuser architecture in the CARLA simulator and show that we can control and improve on certain infraction types, without a performance loss in any other metric. Additionally, we investigate a novel ensemble technique that produces a more cautious ensemble, reducing infractions while maintaining overall performance. For future work, our novel loss formulation enables the definition of a time-variant loss tailored to specific traffic scenarios in the training data.
Malte Stelzer, Timo Bartels, Jan Bickerdt, Volker Schomerus, Jan Piewek, Thorsten Bagdonat, Tim Fingscheidt
IV7
2024 Distributed Semantic Segmentation with Efficient Joint Source and Task Decoding
Danish Nazir, Timo Bartels, Jan Piewek, Thorsten Bagdonat, Tim Fingscheidt
ECCV (88)5
2024 Efficient High-Performance Bark-Scale Neural Network for Residual Echo and Noise Suppression
abstract
In recent years, the introduction of neural networks (NNs) into the field of speech enhancement has brought significant improvements. However, many of the proposed methods are quite demanding in terms of computational complexity and memory footprint. For the application in dedicated communication devices, such as speakerphones, hands-free car systems, or smartphones, efficiency plays a major role along with performance. In this context, we present an efficient, high-performance hybrid joint acoustic echo control and noise suppression system, whereby our main contribution is the post-filter NN, performing both noise and residual echo suppression. The preservation of nearend speech is improved by a Bark-scale auditory filterbank for the NN postfilter. The proposed hybrid method is benchmarked with state-of-the-art methods and its effectiveness is demonstrated on the ICASSP 2023 AEC Challenge blind test set. We demonstrate that it offers high-quality nearend speech preservation during both double-talk and nearend speech conditions. At the same time, it is capable of efficient removal of echo leaks, achieving a comparable performance to already small state-of-the-art models such as the end-to-end DeepVQE-S, while requiring only around 10% of its computational complexity. This makes it easily realtime implementable on a speakerphone device.
Ernst Seidel, Pejman Mowlaee, Tim Fingscheidt
ICASSP3
2024 Employing Real Training Data for Deep Noise Suppression
abstract
Most deep noise suppression (DNS) models are trained with reference-based losses requiring access to clean speech. However, sometimes an additive microphone model is insufficient for real-world applications. Accordingly, ways to use real training data in supervised learning for DNS models promise to reduce a potential training/inference mismatch. Employing real data for DNS training requires either generative approaches or a reference-free loss without access to the corresponding clean speech. In this work, we propose to employ an end-to-end non-intrusive deep neural network (DNN), named PESQ-DNN, to estimate perceptual evaluation of speech quality (PESQ) scores of enhanced real data. It provides a reference-free perceptual loss for employing real data during DNS training, maximizing the PESQ scores. Furthermore, we use an epoch-wise alternating training protocol, updating the DNS model on real data, followed by PESQ-DNN updating on synthetic data. The DNS model trained with the PESQ-DNN employing real data outperforms all reference methods employing only synthetic training data. On synthetic test data, our proposed method excels the Inter-speech 2021 DNS Challenge baseline by a significant 0.32 PESQ points. Both on synthetic and real test data, the proposed method beats the baseline by 0.05 DNSMOS points – although PESQ-DNN optimizes for a different perceptual metric.
Marvin Sach, Jan Pirklbauer, Tim Fingscheidt
ICASSP4
2024 Mixed Children/Adult/Childrenized Fine-Tuning for Children's ASR: How to Reduce Age Mismatch and Speaking Style Mismatch
Thomas Graave, Timo Lohrenz, Tim Fingscheidt
INTERSPEECH4
2024 Interleaved Audio/Audiovisual Transfer Learning for AV-ASR in Low-Resourced Languages
Patrick Blumenberg, Thomas Graave, Timo Lohrenz, Siegfried Kunzmann, Tim Fingscheidt
INTERSPEECH7
2024 URGENT Challenge: Universality, Robustness, and Generalizability For Speech Enhancement
Wangyou Zhang, Robin Scheibler, Kohei Saijo, Samuele Cornell, Chenda Li, Zhaoheng Ni, Jan Pirklbauer, Marvin Sach, Shinji Watanabe 0001, Tim Fingscheidt, Yanmin Qian
INTERSPEECH10
2024 Generalization by Adaptation: Diffusion-Based Domain Extension for Domain-Generalized Semantic Segmentation
abstract
When models, e.g., for semantic segmentation, are applied to images that are vastly different from training data, the performance will drop significantly. Domain adaptation methods try to overcome this issue, but need samples from the target domain. However, this might not always be feasible for various reasons and therefore domain generalization methods are useful as they do not require any target data. We present a new diffusion-based domain extension (DIDEX) method and employ a diffusion model to generate a pseudo-target domain with diverse text prompts. In contrast to existing methods, this allows to control the style and content of the generated images and to introduce a high diversity. In a second step, we train a generalizing model by adapting towards this pseudo-target domain. We outperform previous approaches by a large margin across various datasets and architectures without using any real data. For the generalization from GTA5, we improve state-of-the-art mIoU performance by 3.8 % absolute on average and for SYNTHIA by 11.8 % absolute, marking a big step for the generalization performance on these benchmarks. Code is available at https://github.com/JNiemeijer/DIDEX
Joshua Niemeijer, Manuel Schwonberg, Jan-Aike Termöhlen, Nico M. Schmidt, Tim Fingscheidt
WACV5
2024 Convergence and Performance Analysis of Classical, Hybrid, and Deep Acoustic Echo Control
abstract
Acoustic echo cancellation (AEC) and suppression (AES) are widely researched topics. However, only few papers about hybrid or deep acoustic echo control provide a solid comparative analysis of their methods as it was common with classical signal processing approaches. There can be distinct differences in the behaviour of an AEC/AES model which cannot be fully represented by a single metric or test condition, especially when comparing classical signal processing and machine-learned approaches. These characteristics include convergence behaviour, reliability under varying speech levels or far-end signal types, as well as robustness to adverse conditions such as harsh nonlinearities, room impulse response switches or continuous changes, or delayed echo. A first contribution of this article is to present an extended set oftest conditionsandmetricsthat yields a proper characterization of an AEC/AES model and provides researchers with a useful toolbox to benchmark their systems. Second, we evaluate multiple AEC/AES models, each representing a classical, machine-learned, or hybrid paradigm, in various test conditions. We provide an analysis and new insights into their strengths and weaknesses and identify limitations of common metrics in some cases. Our entire toolbox of evaluation metrics and testing conditions is available on GitHub1.
Ernst Seidel, Pejman Mowlaee, Tim Fingscheidt
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Parameter-Efficient Cross-Language Transfer Learning for a Language-Modular Audiovisual Speech Recognition
abstract
In audiovisual speech recognition (AV-ASR), for many languages only few audiovisual data is available. Building upon an English model, in this work, we first apply and analyze various adapters for cross-language transfer learning to build a parameter-efficient and easy-to-extend AV-ASR in multiple languages. Fine-tuning only the bottleneck adapter with 4% of encoder’s parameters and the decoder shows comparable performance to full fine-tuning in French and Spanish AV-ASR. Second, we investigate the effectiveness of various encoder components in cross-language transfer learning. Our proposed modular linguistic transfer learning approach excels the full fine-tuning method for German, French, and Spanish AV-ASR in almost all clean and noisy conditions (8/9). On low-resourced German AV data (13h), our proposed linguistic transfer learning achieves a 4.1% abs. WER reduction on average for clean and noisy speech, while fine-tuning only 50% of the encoder’s parameters. Our code is at GitHub.11https://github.com/ifnspaml/Cross_Language_Transfer_Learning_AVASR.git
Thomas Graave, Timo Lohrenz, Siegfried Kunzmann, Tim Fingscheidt
ASRU6
2023 Relaxed Attention for Transformer Models
abstract
The powerful modeling capabilities of all-attention-based transformer architectures often cause overfitting and-for natural language processing tasks-lead to an implicitly learned internal language model in the autoregressive transformer decoder complicating the integration of external language models. In this paper, we explore relaxed attention, a simple and easy-to-implement smoothing of the attention weights, yielding a two-fold improvement to the general transformer architecture: First, relaxed attention provides regularization when applied to the self-attention layers in the encoder. Second, we show that it naturally supports the integration of an external language model as it suppresses the implicitly learned internal language model by relaxing the cross attention in the decoder. We demonstrate the benefit of relaxed attention across several tasks from different applications with clear improvement in combination with recent benchmark approaches using various transformer model variants and sizes. Specifically, we exceed the former state-of-the-art performance of 26.90% word error rate on the largest public lip-reading LRS3 benchmark with a word error rate of 26.31%, as well as we achieve a top-performing BLEU score of 37.67 on the IWSLT14 (DE → EN) machine translation task without external language models and virtually no additional model parameters.
Timo Lohrenz, Björn Möller, Tim Fingscheidt
IJCNN4
2023 An Efficient and Noise-Robust Audiovisual Encoder for Audiovisual Speech Recognition
Chenwei Liang, Timo Lohrenz, Marvin Sach, Björn Möller, Tim Fingscheidt
INTERSPEECH6
2023 EffCRN: An Efficient Convolutional Recurrent Network for High-Performance Speech Enhancement
Marvin Sach, Jan Franzen, Bruno Defraene, Kristoff Fluyt, Maximilian Strake, Wouter Tirry, Tim Fingscheidt
INTERSPEECH7
2023 Coded Speech Quality Measurement by a Non-Intrusive PESQ-DNN
abstract
Wideband codecs such as AMR-WB or EVS are widely used in (mobile) speech communication. Evaluation of coded speech quality is often performed subjectively by an absolute category rating (ACR) listening test. However, the ACR test is impractical for online monitoring of speech communication networks. Perceptual evaluation of speech quality (PESQ) is one of the widely used metrics instrumentally predicting the results of an ACR test. However, the PESQ algorithm requires an original reference signal, which is usually unavailable in network monitoring, thus limiting its applicability.NISQAis a new non-intrusive neural-network-based speech quality measure, focusing on super-wideband speech signals. In this work, however, we aim at predicting the well-known PESQ metric using a non-intrusivePESQ-DNNmodel. We illustrate the potential of this model by predicting the PESQ scores of wideband-coded speech obtained from AMR-WB or EVS codecs operating at different bitrates in noisy, tandeming, and error-prone transmission conditions. We compare our methods with the state-of-the-art network topologies ofQualityNet,WaweNet, andDNSMOS—all applied to PESQ prediction—by measuring the mean absolute error (MAE) and the linear correlation coefficient (LCC). The proposedPESQ-DNNoffers the best total MAE and LCC of 0.11 and 0.92, respectively, in conditions without frame loss, and still is best when including frame loss. Note that our model could be similarly used to non-intrusively predict POLQA or other (intrusive) metrics. The proposedPESQ-DNNmodel definition and the code are provided athttps://github.com/ifnspaml/PESQDNN.
Ziyue Zhao 0007, Tim Fingscheidt
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Deep Residual Echo Suppression and Noise Reduction: A Multi-Input FCRN Approach in a Hybrid Speech Enhancement System
abstract
Deep neural network (DNN)-based approaches to acoustic echo cancellation (AEC) and hybrid speech enhancement systems have gained increasing attention recently, introducing significant performance improvements to this research field. Using the fully convolutional recurrent network (FCRN) architecture that is among state of the art topologies for noise reduction, we present a novel deep residual echo suppression and noise reduction with up to four input signals as part of a hybrid speech enhancement system with a linear frequency domain adaptive Kalman filter AEC. In an extensive ablation study, we reveal trade-offs with regard to echo suppression, noise reduction, and near-end speech quality, and provide surprising insights to the choice of the FCRN inputs: In contrast to often seen input combinations for this task, we propose not to use the loudspeaker reference signal, but the enhanced signal after AEC, the microphone signal, and the echo estimate, yielding improvements over previous approaches by more than 0.2PESQ points.
Jan Franzen, Tim Fingscheidt
ICASSP2
2022 Detecting Adversarial Perturbations in Multi-Task Perception
abstract
While deep neural networks (DNNs) achieve impressive performance on environment perception tasks, their sensitivity to adversarial perturbations limits their use in practical applications. In this paper, we (i) propose a novel adversarial perturbation detection scheme based on multi-task perception of complex vision tasks (i.e., depth estimation and semantic segmentation). Specifically, adversarial perturbations are detected by inconsistencies between extracted edges of the input image, the depth output, and the segmentation output. To further improve this technique, we (ii) develop a novel edge consistency loss between all three modalities, thereby improving their initial consistency which in turn supports our detection scheme. We verify our detection scheme's effectiveness by employing various known attacks and image noises. In addition, we (iii) develop a multi-task adversarial attack, aiming at fooling both tasks as well as our detection scheme. Experimental evaluation on the Cityscapes and KITTI datasets shows that under an assumption of a 5% false positive rate up to 100% of images are correctly detected as adversarially perturbed, depending on the strength of the perturbation. Code is available at https://github.com/ifnspaml/AdvAttackDet. A short video at https://youtu.be/KKa6gOyWmH4 provides qualitative results.
Marvin Klingner, Varun Ravi Kumar, Senthil Kumar Yogamani, Andreas Bär, Tim Fingscheidt
IROS5
2022 Amodal Cityscapes: A New Dataset, its Generation, and an Amodal Semantic Segmentation Challenge Baseline
abstract
Amodal perception terms the ability of humans to imagine the entire shapes of occluded objects. This gives humans an advantage to keep track of everything that is going on, especially in crowded situations. Typical perception functions, however, lack amodal perception abilities and are therefore at a disadvantage in situations with occlusions. Complex urban driving scenarios often experience many different types of occlusions and, therefore, amodal perception for automated vehicles is an important task to investigate. In this paper, we consider the task of amodal semantic segmentation and propose a generic way to generate datasets to train amodal semantic segmentation methods. We use this approach to generate an amodal Cityscapes dataset. Moreover, we propose and evaluate a method as baseline on Amodal Cityscapes, showing its applicability for amodal semantic segmentation in automotive environment perception. We provide the means to re-generate this dataset on github1.1https://github.com/ifnspaml/AmodalCityscapes
Jasmin Breitenstein, Tim Fingscheidt
IV2
2022 Transformer-Based Lip-Reading with Regularized Dropout and Relaxed Attention
abstract
End-to-end automatic lip-reading usually comprises an encoder-decoder model and an optional external language model. In this work, we introduce two regularization methods to the field of lip-reading: First, we apply the regularized dropout (R-Drop) method to transformer-based lip-reading to improve their training-inference consistency. Second, the relaxed attention technique is applied during training for a better external language model integration. We are the first to show that these two complementary approaches yield particu1arly strong performance if combined in the right manner. In particular, by adding an additional R - Drop loss and smoothing the attention weights in cross multi-head attention during training only, we achieve a new state of the art with a word error rate of 22.2% on Lip Reading Sentences 2 (LRS2). On LRS3, we are 2nd ranked with 25.5% WER using only 1,759 h of training data, while the 1 st rank uses about 90,000 h. Our code is available at GitHub.11https://github.com/ifnspaml/Lipreading-RDrop-RA
Timo Lohrenz, Matthias Dunkelberg, Tim Fingscheidt
SLT4
2022 Deep Noise Suppression Maximizing Non-Differentiable PESQ Mediated by a Non-Intrusive PESQNet
abstract
Speech enhancement employing deep neural networks (DNNs) for denoising is called deep noise suppression (DNS). The DNS trained with mean squared error (MSE) losses cannot guarantee good perceptual quality. Perceptual evaluation of speech quality (PESQ) is a widely used metric for evaluating speech quality. However, the original PESQ algorithm is non-differentiable, therefore, cannot directly be used as optimization criterion for gradient-based learning. In this work, we propose an end-to-end non-intrusivePESQNetDNN to estimate the PESQ scores of the enhanced speech signal. Thus, by providing a reference-free perceptual loss, it serves as a mediator towards the DNS training, allowing to maximize the PESQ score of the enhanced speech signal. We illustrate the potential of our proposedPESQNet-mediated training on a strong baseline DNS. As further novelty, we propose to train the DNS and thePESQNetalternatingly to keep thePESQNetup-to-date and perform well specifically for the DNS under training. Detailed analysis shows that thePESQNetmediation further increases the DNS performance by about 0.1 PESQ points on synthetic test data and by 0.03 DNSMOS points on real test data, compared to training with the MSE-based loss. Our proposed method outperforms the Interspeech 2021 DNS Challenge baseline by 0.2 PESQ points on synthetic test data and 0.1 DNSMOS points on real test data. Furthermore, it improves on the same DNS trained with an approximated differentiable PESQ loss by about 0.4 PESQ points on synthetic test data and 0.2 DNSMOS points on real test data.
Maximilian Strake, Tim Fingscheidt
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Continual BatchNorm Adaptation (CBNA) for Semantic Segmentation
abstract
Environment perception in autonomous driving vehicles often heavily relies on deep neural networks (DNNs), which are subject to domain shifts, leading to a significantly decreased performance during DNN deployment. Usually, this problem is addressed by unsupervised domain adaptation (UDA) approaches trained either simultaneously on source and target domain datasets or even source-free only on target data in an offline fashion. In this work, we further expand a source-free UDA approach to a continual and therefore online-capable UDA on a single-image basis for semantic segmentation. Accordingly, our method only requires the pre-trained model from the supplier (trained in the source domain) and the current (unlabeled target domain) camera image. Our method Continual BatchNorm Adaptation (CBNA) modifies the source domain statistics in the batch normalization layers, using target domain images in an unsupervised fashion, which yields consistent performance improvements during inference. Thereby, in contrast to existing works, our approach can be applied to improve a DNN continuously on a single-image basis during deployment without access to source data, without algorithmic delay, and nearly without computational overhead. We show the consistent effectiveness of our method across a wide variety of source/target domain settings for semantic segmentation. Code is available athttps://github.com/ifnspaml/CBNA
Marvin Klingner, Mouadh Ayache, Tim Fingscheidt
IEEE Trans. Intell. Transp. Syst.3
2022 SVDistNet: Self-Supervised Near-Field Distance Estimation on Surround View Fisheye Cameras
abstract
A 360° perception of scene geometry is essential for automated driving, notably for parking and urban driving scenarios. Typically, it is achieved using surround-view fisheye cameras, focusing on the near-field area around the vehicle. The majority of current depth estimation approaches focus on employing just a single camera, which cannot be straightforwardly generalized to multiple cameras. The depth estimation model must be tested on a variety of cameras equipped to millions of cars with varying camera geometries. Even within a single car, intrinsics vary due to manufacturing tolerances. Deep learning models are sensitive to these changes, and it is practically infeasible to train and test on each camera variant. As a result, we present novel camera-geometry adaptive multi-scale convolutions which utilize the camera parameters as a conditional input, enabling the model to generalize to previously unseen fisheye cameras. Additionally, we improve the distance estimation by pairwise and patchwise vector-based self-attention encoder networks. We evaluate our approach on the Fisheye WoodScape surround-view dataset, significantly improving over previous approaches. We also show a generalization of our approach across different camera viewing angles and perform extensive experiments to support our contributions. To enable comparison with other approaches, we evaluate the front camera data on the KITTI dataset (pinhole camera images) and achieve state-of-the-art performance among self-supervised monocular methods. An overview video with qualitative results is provided athttps://youtu.be/bmX0UcU9wtA. Baseline code and dataset will be made public.1
Varun Ravi Kumar, Marvin Klingner, Senthil Kumar Yogamani, Markus Bach, Stefan Milz, Tim Fingscheidt, Patrick Mäder
IEEE Trans. Intell. Transp. Syst.6
2021 Relaxed Attention: A Simple Method to Boost Performance of End-to-End Automatic Speech Recognition
abstract
Recently, attention-based encoder-decoder (AED) models have shown high performance for end-to-end automatic speech recognition (ASR) across several tasks. Addressing overconfidence in such models, in this paper we introduce the concept of relaxed attention, which is a simple gradual injection of a uniform distribution to the encoder-decoder attention weights during training that is easily implemented with two lines of code. We investigate the effect of relaxed attention across different AED model architectures and two prominent ASR tasks, Wall Street Journal (WSJ) and Librispeech. We found that transformers trained with relaxed attention outperform the standard baseline models consistently during decoding with external language models. On WSJ, we set a new benchmark for transformer-based end-to-end speech recognition with a word error rate of 3.65%, outperforming state of the art (4.20%) by 13.1% relative, while introducing only a single hyperparameter.
Timo Lohrenz, Patrick Schwarz, Tim Fingscheidt
ASRU4
2021 A New DCASE 2017 Rare Sound Event Detection Benchmark Under Equal Training Data: CRNN With Multi-Width Kernels
abstract
Rare sound event detection (rare SED) deals with obtaining valuable information from data consisting mostly of acoustic background noises. It has meanwhile a long research history and was part of the DCASE 2017 Challenge. State-of-the-art performance is currently reached using a stacked combination of a CNN and an RNN, dubbed CRNN, which was also successfully applied in other domains such as in hybrid automatic speech recognition. In this work, we propose a new CRNN model for rare SED. This new model uses a set of parallel convolutions with multiple kernel widths in the CRNN and is based on an extended feature representation of the log-mel spectrogram. Furthermore, we apply and optimize different evaluation postprocessing methods and analyze the modifications in an ablation study. The proposed model outperforms the so-far top-scoring networks of the DCASE Challenge – using the same training material for all methods – by an error rate of 6.13% absolute and by 4.39% absolute in the F1 score on the test set and under these conditions achieves a new benchmark result on the DCASE 2017 Rare SED data set.
Jan Baumann, Patrick Meyer, Timo Lohrenz, Alexander Roy, Michael Papendieck, Tim Fingscheidt
ICASSP6
2021 AEC in A Netshell: on Target and Topology Choices for FCRN Acoustic Echo Cancellation
abstract
Acoustic echo cancellation (AEC) algorithms have a long-term steady role in signal processing, with approaches improving the performance of applications such as automotive hands-free systems, smart home and loudspeaker devices, or web conference systems. Just recently, very first deep neural network (DNN)-based approaches were proposed with a DNN for joint AEC and residual echo suppression (RES)/noise reduction, showing significant improvements in terms of echo suppression performance. Noise reduction algorithms, on the other hand, have enjoyed already a lot of attention with regard to DNN approaches, with the fully convolutional recurrent network (FCRN) architecture being among state of the art topologies. The recently published impressive echo cancellation performance of joint AEC/RES DNNs, however, so far came along with an undeniable impairment of speech quality. In this work we will heal this issue and significantly improve the near-end speech component quality over existing approaches. Also, we propose for the first time—to the best of our knowledge—a pure DNN AEC in the form of an echo estimator, that is based on a competitive FCRN structure and delivers a quality useful for practical applications.
Jan Franzen, Ernst Seidel, Tim Fingscheidt
ICASSP3
2021 From a Fourier-Domain Perspective on Adversarial Examples to a Wiener Filter Defense for Semantic Segmentation
abstract
Despite recent advancements, deep neural networks are not robust against adversarial perturbations. Many of the proposed adversarial defense approaches use computationally expensive training mechanisms that do not scale to complex real-world tasks such as semantic segmentation, and offer only marginal improvements. In addition, fundamental questions on the nature of adversarial perturbations and their relation to the network architecture are largely understudied. In this work, we study the adversarial problem from a frequency domain perspective. More specifically, we analyze discrete Fourier transform (DFT) spectra of several adversarial images and report two major findings: First, there exists a strong connection between a model architecture and the nature of adversarial perturbations that can be observed and addressed in the frequency domain. Second, the observed frequency patterns are largely image- and attack-type independent, which is important for the practical impact of any defense making use of such patterns. Motivated by these findings, we additionally propose an adversarial defense method based on the well-known Wiener filters that captures and suppresses adversarial frequencies in a data-driven manner. Our proposed method not only generalizes across unseen attacks but also excels five existing state-of-the-art methods across two models in a variety of attack settings.
Nikhil Kapoor, Andreas Bär, Serin Varghese, Jan David Schneider, Fabian Hüger, Peter Schlicht, Tim Fingscheidt
IJCNN7
2021 Multi-Encoder Learning and Stream Fusion for Transformer-Based End-to-End Automatic Speech Recognition
abstract
Stream fusion, also known as system combination, is a common technique in automatic speech recognition for traditional hybrid hidden Markov model approaches, yet mostly unexplored for modern deep neural network end-to-end model architectures. Here, we investigate various fusion techniques for the all-attention-based encoder-decoder architecture known as the transformer, striving to achieve optimal fusion by investigating different fusion levels in an example single-microphone setting with fusion of standard magnitude and phase features. We introduce a novel multi-encoder learning method that performs a weighted combination of two encoder-decoder multi-head attention outputs only during training. Employing then only the magnitude feature encoder in inference, we are able to show consistent improvement on Wall Street Journal (WSJ) with language model and on Librispeech, without increase in runtime or parameters. Combining two such multi-encoder trained models by a simple late fusion in inference, we achieve state-of-the-art performance for transformer-based models on WSJ with a significant WER reduction of 19% relative compared to the current benchmark approach.
Timo Lohrenz, Tim Fingscheidt
Interspeech3
2021 Y2-Net FCRN for Acoustic Echo and Noise Suppression
abstract
In recent years, deep neural networks (DNNs) were studied as an alternative to traditional acoustic echo cancellation (AEC) algorithms. The proposed models achieved remarkable performance for the separate tasks of AEC and residual echo suppression (RES). A promising network topology is a fully convolutional recurrent network (FCRN) structure, which has already proven its performance on both noise suppression and AEC tasks, individually. However, the combination of AEC, postfiltering, and noise suppression to a single network typically leads to a noticeable decline in the quality of the near-end speech component due to the lack of a separate loss for echo estimation. In this paper, we propose a two-stage model (Y$^2$-Net) which consists of two FCRNs, each with two inputs and one output (Y-Net). The first stage (AEC) yields an echo estimate, which - as a novelty for a DNN AEC model - is further used by the second stage to perform RES and noise suppression. While the subjective listening test of the Interspeech 2021 AEC Challenge mostly yielded results close to the baseline, the proposed method scored an average improvement of 0.46 points over the baseline on the blind testset in double-talk on the instrumental metric DECMOS, provided by the challenge organizers.
Ernst Seidel, Jan Franzen, Maximilian Strake, Tim Fingscheidt
Interspeech4
2021 Deep Noise Suppression with Non-Intrusive PESQNet Supervision Enabling the Use of Real Training Data
abstract
Data-driven speech enhancement employing deep neural networks (DNNs) can provide state-of-the-art performance even in the presence of non-stationary noise. During the training process, most of the speech enhancement neural networks are trained in a fully supervised way with losses requiring noisy speech to be synthesized by clean speech and additive noise. However, in a real implementation, only the noisy speech mixture is available, which leads to the question, how such data could be advantageously employed in training. In this work, we propose an end-to-end non-intrusive PESQNet DNN which estimates perceptual evaluation of speech quality (PESQ) scores, allowing a reference-free loss for real data. As a further novelty, we combine the PESQNet loss with denoising and dereverberation loss terms, and train a complex mask-based fully convolutional recurrent neural network (FCRN) in a weakly supervised way, each training cycle employing some synthetic data, some real data, and again synthetic data to keep the PESQNet up-to-date. In a subjective listening test, our proposed framework outperforms the Interspeech 2021 Deep Noise Suppression (DNS) Challenge baseline overall by 0.09 MOS points and in particular by 0.45 background noise MOS points.
Maximilian Strake, Tim Fingscheidt
Interspeech3
2021 An Application-Driven Conceptualization of Corner Cases for Perception in Highly Automated Driving
abstract
Systems and functions that rely on machine learning (ML) are the basis of highly automated driving. An essential task of such ML models is to reliably detect and interpret unusual, new, and potentially dangerous situations. The detection of those situations, which we refer to as corner cases, is highly relevant for successfully developing, applying, and validating automotive perception functions in future vehicles where multiple sensor modalities will be used. A complication for the development of corner case detectors is the lack of consistent definitions, terms, and corner case descriptions, especially when taking into account various automotive sensors. In this work, we provide an application-driven view of corner cases in highly automated driving. To achieve this goal, we first consider existing definitions of the general outlier, novelty, anomaly, and out-of-distribution detection to show relations and differences to corner cases. Moreover, we extend an existing camera-focused systematization of corner cases by adding RADAR (radio detection and ranging) and LiDAR (light detection and ranging) sensors. For this, we describe an exemplary toolchain for data acquisition and processing, highlighting the interfaces of corner case detection. We also define a novel level of corner cases, the method layer corner cases, which appear due to uncertainty inherent in the methodology.
Florian Heidecker, Jasmin Breitenstein, Kevin Rösch, Jonas Löhdefink, Maarten Bieshaar, Christoph Stiller, Tim Fingscheidt, Bernhard Sick
IV7
2021 Improving Convolutional Recurrent Neural Networks for Speech Emotion Recognition
abstract
Deep learning has increased the interest in speech emotion recognition (SER) and has put forth diverse structures and methods to improve performance. In recent years it has turned out that applying SER on a (log-mel) spectrogram and thus, interpreting SER as an image recognition task is a promising method. Following the trend towards using a convolutional neural network (CNN) in combination with a bidirectional long short-term memory (BLSTM) layer, and some subsequent fully connected layers, in this work, we advance the performance of this topology by several contributions: We integrate a multi-kernel width CNN, propose a BLSTM output summarization function, apply an enhanced feature representation, and introduce an effective training method. In order to foster insight into our proposed methods, we separately evaluate the impact of each modification in an ablation study. Based on our modifications, we obtain top results for this type of topology on IEMOCAP with an unweighted average recall of 64.5% on average.
Patrick Meyer, Tim Fingscheidt
SLT3
2021 SynDistNet: Self-Supervised Monocular Fisheye Camera Distance Estimation Synergized with Semantic Segmentation for Autonomous Driving
abstract
State-of-the-art self-supervised learning approaches for monocular depth estimation usually suffer from scale ambiguity. They do not generalize well when applied on distance estimation for complex projection models such as in fisheye and omnidirectional cameras. This paper introduces a novel multi-task learning strategy to improve self-supervised monocular distance estimation on fisheye and pinhole camera images. Our contribution to this work is threefold: Firstly, we introduce a novel distance estimation network architecture using a self-attention based encoder coupled with robust semantic feature guidance to the decoder that can be trained in a one-stage fashion. Secondly, we integrate a generalized robust loss function, which improves performance significantly while removing the need for hyperparameter tuning with the reprojection loss. Finally, we reduce the artifacts caused by dynamic objects violating static world assumptions using a semantic masking strategy. We significantly improve upon the RMSE of previous work on fisheye by 25% reduction in RMSE. As there is little work on fisheye cameras, we evaluated the proposed method on KITTI using a pinhole model. We achieved state-of-the-art performance among self-supervised methods without requiring an external scale estimation.
Varun Ravi Kumar, Marvin Klingner, Senthil Kumar Yogamani, Stefan Milz, Tim Fingscheidt, Patrick Mäder
WACV5
2021 Online Performance Prediction of Perception DNNs by Multi-Task Learning With Depth Estimation
abstract
Online performance prediction (or: observation) of deep neural networks (DNNs) in highly automated driving presents an unsolved task until now, as most DNNs are evaluated offline requiring datasets with ground truth labels. In practice, however, DNN performance depends on the used camera type, lighting and weather conditions, and on various other kinds of domain shift. Also, the input to DNN-based perception systems can be perturbed by adversarial attacks requiring means to detect these input perturbations. In this work we propose a method to mitigate these problems by a multi-task learning approach with monocular depth estimation as a secondary task, which enables us to predict the DNN's performance for various other (primary) tasks by evaluating only the depth estimation task with a physical depth measurement provided, e.g., by a LiDAR sensor. We show the effectiveness of our method for the primary task of semantic segmentation using various training datasets, test datasets, model architectures, and input perturbations. Our method provides an effective way to predict (observe) the performance of DNNs for semantic segmentation even on a single-image basis and is transferable to other primary DNN-based perception tasks in a straightforward manner.
Marvin Klingner, Tim Fingscheidt
IEEE Trans. Intell. Transp. Syst.2
2020 Self-supervised Monocular Depth Estimation: Solving the Dynamic Object Problem by Semantic Guidance
Marvin Klingner, Jan-Aike Termöhlen, Jonas Mikolajczyk, Tim Fingscheidt
ECCV (20)4
2020 Beyond the Dcase 2017 Challenge on Rare Sound Event Detection: A Proposal for a More Realistic Training and Test Framework
abstract
There are many ways to evaluate rare sound event detection (SED) approaches, e.g., the DCASE 2017 challenge provides a widely employed framework. This paper proposes a rare SED training and test framework, which is reflecting an SED application in a more realistic way. Our setup gets rid of too much prior knowledge on the test data, and assumes additional unknown acoustic events both in training and test data, which in practice have to be identified as background. Taking this into account during training, the robustness in real-world scenarios can be significantly increased, with an average event-based error rate reduction of an absolute 34 percentage points. Further we show and compare the performance of multi-event (polyphonic) classifiers vs. single-event classifiers while outlining the benefits of multi-event training.
Jan Baumann, Timo Lohrenz, Alexander Roy, Tim Fingscheidt
ICASSP4
2020 A Multichannel Kalman-Based Wiener Filter Approach for Speaker Interference Reduction in Meetings
abstract
Recording a meeting and obtaining clean speech signals of each speaker is a challenging task. Even with a multichannel recording, in which all speakers are equipped with a close-talk microphone, speech of an active speaker still couples not only into his dedicated microphone, but also into all other microphone channels at a certain level. This is denoted as crosstalk and requires a multichannel speaker interference reduction to enhance the microphone channels for further processing. To solve this issue, we use a Wiener filter which is based on all individual microphone channels. We extend an existing approach by integrating methods from acoustic echo cancellation to improve the estimation of the interferer (noise) components of the filter, which leads to an improved signal-to-interferer ratio by up to 2.1 dB absolute at constant speech component quality.
Patrick Meyer, Samy Elshamy, Tim Fingscheidt
ICASSP3
2020 Fully Convolutional Recurrent Networks for Speech Enhancement
abstract
Convolutional recurrent neural networks (CRNs) using convolutional encoder-decoder (CED) structures have shown promising performance for single-channel speech enhancement. These CRNs handle temporal modeling through integrating long short-term memory (LSTM) layers in between convolutional encoder and decoder. However, in such a CRN, the organization of internal representations in feature maps and the focus on local structure of the convolutional mappings has to be discarded for fully-connected LSTM processing. Furthermore, CRNs can be quite restricted concerning the feature space dimension at the input of the LSTM, which, through its fully-connected nature, requires a large amount of trainable parameters. As first novelty, we propose to replace the fully-connected LSTM by a convolutional LSTM (ConvLSTM) and call the resulting network a fully convolutional recurrent network (FCRN). Secondly, since the ConvLSTM retains the structured organization of its input feature maps, we can show that this helps to internally represent the harmonic structure of speech, allowing us to handle high-dimensional input features using less trainable parameters than an LSTM. The proposed FCRN clearly outperforms CRN reference models with similar amounts of trainable parameters in terms of PESQ, STOI, and segmental ΔSNR.
Maximilian Strake, Bruno Defraene, Kristoff Fluyt, Wouter Tirry, Tim Fingscheidt
ICASSP5
2020 Using Separate Losses for Speech and Noise in Mask-Based Speech Enhancement
abstract
Estimating time-frequency domain masks for speech enhancement using deep learning approaches has recently become a popular field in research. In this paper, we propose a novel components loss (CL) for the training of neural networks for speech enhancement. During the training process, the proposed CL offers separate control over suppression of the noise component and preservation of the speech component. We illustrate the potential of the proposed CL by example of a convolutional neural network (CNN) for mask-based speech enhancement. We show improvement in almost all employed instrumental quality metrics over the baseline losses, which comprises the conventional mean squared error (MSE) loss and also perceptual evaluation of speech quality (PESQ) loss. On average, more than 0.3 dB higher SNR improvement and an at least 0.1 points higher PESQ score on the speech component are obtained. In addition to that, a more naturally sounding residual noise and a consistently best PESQ on the enhanced speech is obtained. All improvements are more distinct at low SNR conditions.
Samy Elshamy, Tim Fingscheidt
ICASSP3
2020 BLSTM-Driven Stream Fusion for Automatic Speech Recognition: Novel Methods and a Multi-Size Window Fusion Example
Timo Lohrenz, Tim Fingscheidt
INTERSPEECH2
2020 INTERSPEECH 2020 Deep Noise Suppression Challenge: A Fully Convolutional Recurrent Network (FCRN) for Joint Dereverberation and Denoising
Maximilian Strake, Bruno Defraene, Kristoff Fluyt, Wouter Tirry, Tim Fingscheidt
INTERSPEECH5
2020 Systematization of Corner Cases for Visual Perception in Automated Driving
abstract
One major task in automated driving is the development of robust and safe visual perception modules. It is of utmost importance that visual perception reacts adequately to so-called corner cases, which range from overexposure of the image sensor to unexpected and potentially dangerous traffic situations. Their detection thus has high significance both as an online system in the intelligent vehicle, but also in the extraction of relevant training and test data for perception modules. In this paper, we provide a systematization of corner cases for visual perception in automated driving, with the categories being structured by detection complexity. Furthermore, we discuss existing metrics and datasets which can be used for the evaluation of corner case detection methods depending on their suitability to provide beneficial information for the various categories.
Jasmin Breitenstein, Jan-Aike Termöhlen, Daniel Lipinski, Tim Fingscheidt
IV4
2020 Focussing Learned Image Compression to Semantic Classes for V2X Applications
abstract
Cooperative perception with many sensors involved greatly improves the performance of perceptual systems in autonomous vehicles. However, the increasing amount of sensor data leads to a bottleneck due to limited capacity of vehicle-to-X (V2X) communication channels. We leverage lossy learned image compression by means of an autoencoder with adversarial loss function to reduce the overall bitrate. Our key contribution is to focus image compression on regions of interest (ROIs) governed by a binary mask. A transmitter-sided semantic segmentation network extracts semantically important classes being the basis for the generation of a ROI. A second key contribution is that the mask is not transmitted as side information, only the quantized bottleneck data is transmitted. To train the network, we use a loss function operating only on the pixels in the ROI. We report peak-signal-to-noise ratio (PSNR) both in the entire image and only in the ROI, evaluating various fusion architectures and fusion operations involving input image and mask. Showing the high generalizability of our approach, we achieve consistent improvements in the ROI in all experiments on the Cityscapes dataset.
Jonas Löhdefink, Andreas Bär, Nico M. Schmidt, Fabian Hüger, Peter Schlicht, Tim Fingscheidt
IV6
2020 Terminology and Analysis of Map Deviations in Urban Domains: Towards Dependability for HD Maps in Automated Vehicles
abstract
A driving function relying on map data to operate is prone to failures that are due to deviations between the real world and the map data. Hence, we transfer the concept of dependability known from systems engineering to high-definition (HD) maps to allow for a system ensuring three major aspects of dependability: reliability, availability, and the safe use of map data within the vehicle. In this paper, we therefore define a coherent terminology in the field, particularly introducing necessary terms for describing and measuring map deviations. To substantiate our terminology, we present the results of a measurement campaign and analyze map deviations in an urban domain HD map after a period of 2 years. The results show that the analyzed map contains relatively few errors (0.07 / km) and no obvious persistent changes over the years but is subject to a comparatively high number of temporary changes (0.74 / km) rendering almost 9% of the map outdated.
Christopher Plachetka, Niels Maier, Jenny Fricke, Jan-Aike Termöhlen, Tim Fingscheidt
IV5
2019 On Temporal Context Information for Hybrid BLSTM-Based Phoneme Recognition
abstract
The modern approach to include long-term temporal context information into speech recognition systems is the use of recurrent neural networks, e.g., bi-directional long short-term memory (BLSTM) networks. In this paper, we decouple the BLSTM from a preceding CNN-based feature extractor network allowing us to investigate the use of temporal context in both models in a modular fashion. Accordingly, we train the BLSTMs on posteriors, stemming from preceding CNNs which use various amounts of limited context in their input layer, and investigate to what extent the BLSTM is able to effectively make use of its long-term modeling capabilities. We show that it is beneficial to train the BLSTM on posteriors stemming from a temporal context-free acoustic model. Remarkably, the best performing combination of CNN acoustic model and BLSTM afterwards is a large-context CNN (expected), followed by a BLSTM which has been trained on context-free CNN output posteriors (surprising).
Timo Lohrenz, Maximilian Strake, Tim Fingscheidt
ASRU3
2019 Learning to Dequantize Speech Signals by Primal-dual Networks: an Approach for Acoustic Sensor Networks
abstract
We introduce a method to improve the quality of simple scalar quantization in the context of acoustic sensor networks by combining ideas from sparse reconstruction, artificial neural networks and weighting filters. We start from the observation that optimization methods based on sparse reconstruction resemble the structure of a neural network. Hence, building upon a successful enhancement method, we unroll the algorithms and use this to build a neural network which we train to obtain enhanced decoding. In addition, the weighting filter from code-excited linear predictive (CELP) speech coding is integrated into the loss function of the neural network, achieving perceptually improved reconstructed speech. Our experiments show that our proposed trained methods allow for better speech reconstruction than the reference optimization methods.
Christoph Brauer, Ziyue Zhao 0007, Dirk A. Lorenz, Tim Fingscheidt
ICASSP4
2019 Improved Measurement Noise Covariance Estimation for N-channel Feedback Cancellation Based on the Frequency Domain Adaptive Kalman Filter
abstract
Acoustic feedback cancellation has gained a major and steady role in the research fields of signal processing over the past decades, since it is inevitable for numerous applications such as hearing aids or in-car communication systems. In this paper, we investigate measurement noise covariance estimation approaches for feedback cancellation based on the frequency domain adaptive Kalman filter (FDAKF). The capabilities of these estimation methods significantly affect the performance of the FDAKF. We summarize and investigate existing approaches from literature and furthermore provide two new proposals that are explicitly motivated for the use in acoustic feedback cancellation. Experimental validation in the context of an in-car communication system shows that our proposals obtain much better speech quality compared to existing approaches and additionally increase the overall feedback suppression.
Jan Franzen, Tim Fingscheidt
ICASSP2
2019 Towards Corner Case Detection for Autonomous Driving
abstract
The progress in autonomous driving is also due to the increased availability of vast amounts of training data for the underlying machine learning approaches. Machine learning systems are generally known to lack robustness, e.g., if the training data did rarely or not at all cover critical situations. The challenging task of corner case detection in video, which is also somehow related to unusual event or anomaly detection, aims at detecting these unusual situations, which could become critical, and to communicate this to the autonomous driving system (online use case). Such a system, however, could be also used in offline mode to screen vast amounts of data and select only the relevant situations for storing and (re)training machine learning algorithms. So far, the approaches for corner case detection have been limited to videos recorded from a fixed camera, mostly for security surveillance. In this paper, we provide a formal definition of a corner case and propose a system framework for both the online and the offline use case that can handle video signals from front cameras of a naturally moving vehicle and can output a corner case score.
Jan-Aike Termöhlen, Andreas Bär, Daniel Lipinski, Tim Fingscheidt
IV4
2019 Towards Tactical Maneuver Detection for Autonomous Driving Based on Vision Only
abstract
The detection of tactical maneuvers performed by other traffic participants is a key component for modern autonomous driving systems. We present a novel vision-based approach to detect tactical maneuvers for vehicles, pedestrians, and bicycles. Our approach uses neural networks with temporal convolutions to incorporate temporal information into the detection. The training and evaluation of the presented architecture is performed on a video dataset obtained in a city and industrial environment spanning 3.5 hours. Our approach detects nine distinct maneuver classes with an average accuracy of 54.21%, with single detection accuracies of up to 88.17%.
Antonia Breuer, Jana Kirschner, Silviu Homoceanu, Tim Fingscheidt
IV4
2019 On Low-Bitrate Image Compression for Distributed Automotive Perception: Higher Peak SNR Does Not Mean Better Semantic Segmentation
abstract
The high amount of sensors required for autonomous driving poses enormous challenges on the capacity of automotive bus systems. There is a need to understand tradeoffs between bitrate and perception performance. In this paper, we compare the image compression standards JPEG, JPEG2000, and WebP to a modern encoder/decoder image compression approach based on generative adversarial networks (GANs). We evaluate both the pure compression performance using typical metrics such as peak signal-to-noise ratio (PSNR), structural similarity (SSIM) and others, but also the performance of a subsequent perception function, namely a semantic segmentation (characterized by the mean intersection over union (mIoU) measure). Not surprisingly, for all investigated compression methods, a higher bitrate means better results in all investigated quality metrics. Interestingly, however, we show that the semantic segmentation mIoU of the GAN autoencoder in the highly relevant low-bitrate regime (at 0.0625 bit/pixel) is better by 3.9 % absolute than JPEG2000, although the latter still is considerably better in terms of PSNR (5.91dB difference). This effect can greatly be enlarged by training the semantic segmentation model with images originating from the decoder, so that the mIoU using the segmentation model trained by GAN reconstructions exceeds the use of the model trained with original images by almost 20 % absolute. We conclude that distributed perception in future autonomous driving will most probably not provide a solution to the automotive bus capacity bottleneck by using standard compression schemes such as JPEG2000, but requires modern coding approaches, with the GAN encoder/decoder method being a promising candidate.
Jonas Löhdefink, Andreas Bär, Nico M. Schmidt, Fabian Hüger, Peter Schlicht, Tim Fingscheidt
IV6
2019 Sinusoidal-Based Lowband Synthesis for Artificial Speech Bandwidth Extension
abstract
Conventional narrowband (NB) telephony suffers from limited acoustic bandwidth at the receiver side, leading to degraded speech quality and intelligibility. In this paper, artificial speech bandwidth extension (ABE) of NB speech toward missing frequencies below about 300 Hz (low-frequency band, LB) is proposed to enhance the speech quality. The LB-ABE in this paper is employed together with a preexisting ABE toward high-frequency components to obtain spectrally balanced speech signals. In an instrumental quality assessment, the spectral distance in the LB was improved by more than 5 dB compared to NB speech. In a subjective listening test, the gap of speech quality between wideband and NB speech was significantly reduced when employing the proposed ABE toward low frequencies. The LB extension was found to further improve the preexisting ABE toward higher frequencies by a significant 0.26 CMOS points.
Johannes Abel, Tim Fingscheidt
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 DNN-Based Cepstral Excitation Manipulation for Speech Enhancement
abstract
This contribution aims at speech model-based speech enhancement by exploiting the source-filter model of human speech production. The proposed method enhances the excitation signal in the cepstral domain by making use of a deep neural network (DNN). We investigate two types of target representations along with the significant effects of their normalization. The new approach exceeds the performance of a formerly introduced classical signal processing-based cepstral excitation manipulation (CEM) method in terms of noise attenuation by about 1.5 dB. We show that this gain also holds true when comparing serial combinations of envelope and excitation enhancement. In the important low-SNR conditions, no significant trade-off for speech component quality or speech intelligibility is induced, while allowing for substantially higher noise attenuation. In total, a traditional purely statistical state-of-the-art speech enhancement system is outperformed by more than 3 dB noise attenuation.
Samy Elshamy, Tim Fingscheidt
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 Convolutional Neural Networks to Enhance Coded Speech
abstract
Enhancing coded speech suffering from far-end acoustic background noise, quantization noise, and potentially transmission errors is a challenging task. In this paper, we propose two postprocessing approaches applying convolutional neural networks either in the time domain or the cepstral domain to enhance the coded speech without any modification of the codecs. The time-domain approach follows an end-to-end fashion, whereas the cepstral domain approach uses analysis-synthesis with cepstral domain features. The proposed postprocessors in both domains are evaluated for various narrowband and wideband speech codecs in a wide range of conditions. The proposed postprocessor improves perceptual evaluation of speech quality by up to 0.25 mean opinion score listening quality objective points for G.711, 0.30 points for G.726, 0.82 points for G.722, and 0.26 points for adaptive multirate wideband codec. In a subjective comparison category rating listening test, the proposed postprocessor on G.711-coded speech exceeds the speech quality of an ITU-T-standardized postfilter by 0.36 CMOS points, and obtains a clear preference of 1.77 CMOS points compared to legacy G.711, even better than uncoded speech with statistical significance. The source code for the cepstral domain approach to enhance G.711-coded speech is made available.11https://github.com/ifnspaml/Enhancement-Coded-Speech.
Ziyue Zhao 0007, Huijun Liu 0001, Tim Fingscheidt
IEEE ACM Trans. Audio Speech Lang. Process.3
2018 A Simple Cepstral Domain DNN Approach to Artificial Speech Bandwidth Extension
abstract
In this work, we present a simple deep neural network (DNN)-based regression approach to artificial speech bandwidth extension (ABE) in the frequency domain for estimating missing speech components in the range 4 ... 7 kHz. The upper band (UB) spectral magnitudes are found by first estimating the UB cepstrum by means of a DNN regression and subsequent conversion to the spectral domain, leading to a more efficient and generalizing model training rather than estimating highly redundant UB magnitudes directly. As second novelty the phase information for the estimated upper band spectral magnitudes is generated by spectrally shifting the NB phase. Apart from framing, this very simple approach does not introduce additional algorithmic delay. A cross-database and cross-language task is defined for training and evaluation of the ABE framework. In a subjective comparison category rating test, the proposed ABE solution significantly outperforms the competing ABE baseline and was found to improve NB speech quality by 0.80 CMOS points, while the computation time is reduced to about 3 % compared to the ABE baseline.
Johannes Abel, Maximilian Strake, Tim Fingscheidt
ICASSP3
2018 An Efficient Residual Echo Suppression for Multi-Channel Acoustic Echo Cancellation Based on the Frequency-Domain Adaptive Kalman Filter
abstract
Emerging use cases, such as keyword spotting while listening to FM radio, or participating in a teleconference utilizing the hands-free system in a vehicle, require the utilization of multi-channel acoustic echo cancellation (AEC). Addressing the typically remaining residual echo, it is common practice to apply a postfilter for residual echo suppression (RES) in a subsequent processing stage. In this paper we propose two RES approaches for the multi-channel frequency-domain adaptive Kalman filter, both being optimal under certain assumptions, extremely efficient and robust by exploiting a tight relation to the Kalman stepsize already available from the AEC.
Jan Franzen, Tim Fingscheidt
ICASSP2
2018 Multichannel Speaker Activity Detection for Meetings
abstract
Multichannel recordings of meetings with a (wireless) headset for each person deliver commonly the best audio quality for subsequent analyses. However, still speech portions of other participants can couple into the microphone channel of the associated target speaker. Due to this crosstalk, a speaker activity detection (SAD) is required in order to identify only the speech portions of the target speaker in the related microphone channel. While most solutions are either complex and need a training process, or achieve insufficient results in multi-talk situations, we propose a low complexity method, which can handle both crosstalk and multi-talk situations. We investigate single- and multi-talk in a wide range of different crosstalk levels, and improved the detection accuracy towards a standardized voice activity detection overall by 12.89 % absolute, whereas a state-of-the-art multichannel SAD was exceeded even by 13.76 % absolute.
Patrick Meyer, Rolf Jongebloed, Tim Fingscheidt
ICASSP3
2018 A Priori SNR Estimation Using Discriminative Non-Negative Matrix Factorization
abstract
A priori signal-to-noise ratio (SNR) contains critical information about the single-channel mixture of a speech and noise signal, and can be used by speech enhancement algorithms. In this paper, we propose a novel a priori SNR estimator using the estimates obtained from discriminative non-negative matrix factorization (DNMF). The idea of our new approach is to utilize the DNMF to perform the preliminary speech components estimation, which can be either directly used to estimate the a priori SNR, or can be combined with the well-known decision-directed (DD) approach by Ephraim and Malah to perform the a priori SNR estimation. We present a speaker-independent but noise-dependent DNMF-based a priori SNR estimator. Speech enhancement simulation results in the presence of non-stationary noise validate our new approach combined with well-known spectral weighting rules, outperforming several NMF-based and non-NMF-based state-of-the-art methods, w.r.t. both SNR improvement and speech perceptual quality.
Samy Elshamy, Tim Fingscheidt
ICASSP3
2018 What Do Classifiers Actually Learn? a Case Study on Emotion Recognition Datasets
Patrick Meyer, Eric Buschermöhle, Tim Fingscheidt
INTERSPEECH3
2018 A New Timit Benchmark for Context-Independent Phone Recognition Using Turbo Fusion
abstract
In this work, we apply the recently proposed turbo fusion in conjunction with state-of-the-art convolutional neural networks as acoustic models to the standard phone recognition task on the TIMIT database. The turbo fusion operates on posterior streams stemming from standard filterbank features and from group delay (phase) features. By the iterative exchange of posterior information, the phone error rate is decreased down to 16.91% absolute, which is to our knowledge the best reported result on the TIMIT core test set so far using context-independent acoustic models, outperforming the previous respective benchmark by 4.4% relative.
Timo Lohrenz, Wei Li 0174, Tim Fingscheidt
SLT3
2018 Densenet Blstm for Acoustic Modeling in Robust ASR
abstract
In recent years, robust automatic speech recognition (ASR) has greatly taken benefit from the use of neural networks for acoustic modeling, although performance still degrades in severe noise conditions. Based on the previous success of models using convolutional and subsequent bidirectional long short-term memory (BLSTM) layers in the same network, we propose to use a densely connected convolutional network (DenseNet) as the first part of such a model, while the second is a BLSTM network. A particular contribution of our work is that we modify the DenseNet topology to become a kind of feature extractor for the subsequent BLSTM network operating on whole speech utterances. We evaluate our model on the 6-channel task of CHiME-4, and are able to consistently outperform a top-performing baseline based on wide residual networks and BLSTMs providing a 2.4% relative WER reduction on the real test set.
Maximilian Strake, Pascal Behr, Timo Lohrenz, Tim Fingscheidt
SLT4
2018 Artificial Speech Bandwidth Extension Using Deep Neural Networks for Wideband Spectral Envelope Estimation
abstract
Estimating a wideband spectral envelope having only narrowband speech at hand is a challenging task. In this paper, we explore ways to do so in the context of an artificial speech bandwidth extension (ABE) framework. Starting from a typical hidden Markov model (HMM)/Gaussian mixture model baseline scheme, we investigate two types of features, topologies, and regularization approaches of deep neural networks (DNNs) to obtain estimates of wideband spectral envelopes with smallest cepstral distance to the original ones. In order to draw realistic conclusions, we employ a database for test, which is acoustically different to the training and validation speech material. Interestingly, it turns out that a DNN regression approach outperforms all other investigated methods, although the HMM has been dropped. Cepstral distance was reduced by 1.18 dB, wideband PESQ was improved by 0.23 MOS points, and a subjective comparison category rating listening test showed a significant preference of the best DNN ABE approach versus narrowband speech of 1.37 CMOS points.
Johannes Abel, Tim Fingscheidt
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 DNN-Supported Speech Enhancement With Cepstral Estimation of Both Excitation and Envelope
abstract
In this paper, we propose and compare various techniques for the estimation of clean spectral envelopes in noisy conditions. The source-filter model of human speech production is employed in combination with a hidden Markov model and/or a deep neural network approach to estimate clean envelope-representing coefficients in the cepstral domain. The cepstral estimators for speech spectral envelope-based noise reduction are both evaluated alone and also in combination with the recently introduced cepstral excitation manipulation (CEM) technique for a priori SNR estimation in a noise reduction framework. Relative to the classical MMSE short time spectral amplitude estimator, we obtain more than 2 dB higher noise attenuation, and relative to our recent CEM technique still 0.5 dB more, in both cases maintaining the quality of the speech component and obtaining considerable SNR improvement.
Samy Elshamy, Nilesh Madhu, Wouter Tirry, Tim Fingscheidt
IEEE ACM Trans. Audio Speech Lang. Process.4
2017 Turbo fusion of magnitude and phase information for DNN-based phoneme recognition
abstract
In this work we propose the so-called turbo fusion as competitive method for information fusion of Mel-filterbank magnitude and phase feature streams for automatic speech recognition (ASR). Based on the recently introduced turbo ASR paradigm, our contribution is fourfold: First, we introduce DNN-based acoustic modeling into turbo ASR, then we take steps towards LVCSR by omitting the costly state space transform and by investigating the classical TIMIT phoneme recognition task. Finally, replacing the typical stream weighting in fusion methods, we introduce a new dynamic range limitation of the exchanged posteriors between the involved magnitude and phase recognizers, resulting in a smoother information exchange. The proposed turbo fusion outperforms classical benchmarks on the TIMIT dataset both with and without dropout in DNN training, and also is first if compared to several state-of-the-art reference fusion methods.
Timo Lohrenz, Tim Fingscheidt
ASRU2
2017 A Delay-Flexible Stereo Acoustic Echo Cancellation for DFT-Based In-Car Communication (ICC) Systems
Jan Franzen, Tim Fingscheidt
INTERSPEECH2
2017 An Instrumental Quality Measure for Artificially Bandwidth-Extended Speech Signals
abstract
Various studies have shown that the instrumental measures wideband PESQ and POLQA are not reliably predicting speech quality for artificial speech bandwidth extension (ABE) test conditions, as this has never been their scope. Based on data from a coordinated subjective listening test with 12 ABE variants developed by 6 different institutions, conducted in 4 languages, we propose in this work a novel instrumental quality measure that is specifically suited for narrowband-to-wideband ABE test conditions. In particular, our contributions are fourfold: First, we propose quality indicators particularly being able to detect ABE-related distortions. Second, we investigate the combination of perceptually and nonperceptually motivated distortion-related statistics. Third, we propose a support-vector-machine-based high-performance MOS predictor for ABE speech quality assessment, finally, we present the training process based on the subjective listening test data. A k-fold cross-validation test on 1) disjoint languages, 2) disjoint speakers, and 3) disjoint ABE solutions proves the superiority of our proposed measure in the ITU-T-recommended categories accuracy, consistency, and linearity compared to both, wideband PESQ and POLQA.
Johannes Abel, Magdalena Kaniewska, Cyril Guillaume, Wouter Tirry, Tim Fingscheidt
IEEE ACM Trans. Audio Speech Lang. Process.5
2017 Instantaneous A Priori SNR Estimation by Cepstral Excitation Manipulation
abstract
As the a priori signal-to-noise ratio (SNR) contains crucial information about a signal's mixture of speech and noise, its estimation is subject to steady research. In this paper, we introduce a novel a priori SNR estimator based on synthesizing an idealized excitation signal in the cepstral domain. Our approach utilizes a source-filter decomposition in combination with a cepstral excitation manipulation in order to recreate an idealized excitation, which is subsequently shaped by an immanent envelope. In contrast to the well-known decision-directed approach by Ephraim and Malah, an instantaneous estimate is obtained, which is less prone to sudden acoustic environmental changes and musical noise. Additionally, the proposed estimator is able to preserve weak harmonic structures resulting in a spectrum that is more full-bodied. We present both a speaker-independent and a speaker-dependent variant of the new a priori SNR estimator, both showing more than 2 dB ΔSNR improvement versus state of the art, without any significant increase in speech distortion.
Samy Elshamy, Nilesh Madhu, Wouter Tirry, Tim Fingscheidt
IEEE ACM Trans. Audio Speech Lang. Process.4
2016 A subjective listening test of six different artificial bandwidth extension approaches in English, Chinese, German, and Korean
abstract
In studies on artificial bandwidth extension (ABE), there is a lack of international coordination in subjective tests between multiple methods and languages. Here we present the design of absolute category rating listening tests evaluating 12 ABE variants of six approaches in multiple languages, namely in American English, Chinese, German, and Korean. Since the number of ABE variants caused a higher-than-recommended length of the listening test, ABE variants were distributed into two separate listening tests per language. The paper focuses on the listening test design, which aimed at merging the subjective scores of both tests and thus allows for a joint analysis of all ABE variants under test at once. A language-dependent analysis, evaluating ABE variants in the context of the underlying coded narrowband speech condition showed statistical significant improvement in English, German, and Korean for some ABE solutions.
Johannes Abel, Magdalena Kaniewska, Cyril Guillaume, Wouter Tirry, Hannu Pulakka, Ville Myllylä, Jari Sjoberg, Paavo Alku, Itai Katsir, David Malah, Israel Cohen, M. A. Tugtekin Turan, Engin Erzin, Thomas Schlien, Peter Vary, Amr H. Nour-Eldin, Peter Kabal, Tim Fingscheidt
ICASSP18
2016 Evaluating instrumental measures of speech quality using Bayesian model selection: Correlations can be misleading!
abstract
Choosing among competing models of collected data is crucial for all sciences. In the last decade there has been an increasing tendency to use Bayesian methods throughout many fields. When assessing the performance of instrumental measures of speech quality, classical measures such as correlation coefficients are still used. While these methods have their merits, they discard information about the data distribution, such as variability. They are useful as absolute measures of fit, but often not suitable for comparing different models. This paper uses Bayesian model selection, which does not suffer from these shortcomings, as it takes all information about the distribution of data into account and yields easily interpretable model probabilities. Two instrumental measures of speech quality are evaluated using data obtained in an absolute category rating (ACR) test. The results are compared and discussed. Bayesian methods prove superior for comparing instrumental measures, especially when the correlation of both measures is either poor or nearly identical. The proposed estimation procedure is highly recommended in selection phases for standardization bodies such as ITU-T, ETSI, 3GPP.
Antonio Kolossa, Johannes Abel, Tim Fingscheidt
ICASSP3
2016 System-compatible robustness improvement for new generation dect decoders by G.722 soft-decision decoding
abstract
The ITU-T Recommendation G.722 about subband adaptive differential pulse code modulation (SB-ADPCM) is the mandatory wideband speech codec in the new generation digital enhanced cordless telephony (NG-DECT). Although in ADPCM the difference signal instead of the original signal is quantized and adaptive prediction is employed, redundancy is yet observed within the quantized samples. In this paper we apply a soft-decision speech decoding technique which exploits this redundancy in terms of a priori knowledge and the channel reliability information to NG-DECT. In that way, we propose a novel scheme in a standard-compliant fashion which improves the robustness of the decoder. The performance of our proposal is evaluated in terms of speech quality and a noticeable improvement over the standard codec and its own packet loss concealment algorithm is observed.
Domingo López-Oller, Sai Han, Ángel M. Gómez, José L. Pérez-Córdoba, Tim Fingscheidt
ICASSP5
2016 Soft linear discriminant analysis (SLDA) for pattern recognition with ambiguous reference labels: Application to social signal processing
abstract
While most pattern recognition approaches are designed and trained with clearly defined reference labels, there are a few new applications working with ambiguous ones. Since the linear discriminant analysis (LDA) is one of the most utilized methods in pattern recognition to reduce the dimensionality of feature vectors, typically increasing the robustness of the features, we propose in this work a modification of the LDA in order to be able to handle ambiguous reference labels in a soft-decision way. In the field of social signal processing (here: emotion recognition) we demonstrate that using a soft accuracy measure evaluating the classifier's confidence output by means of a soft-labeled emotional speech database really provides a degree of similarity to (naturally ambiguous) human votes. The adaptation of our classifier to such soft accuracy measure takes place by a retraining w.r.t. the human vote distribution. Applying this soft accuracy measure to emotion recognition with ambiguous reference labels both retraining the classifier and using the new soft LDA method leads to around 22% relative increase of accuracy.
Patrick Meyer, Tim Fingscheidt
ICASSP2
2016 Turbo Automatic Speech Recognition
abstract
Performance of automatic speech recognition (ASR) systems can significantly be improved by integrating further sources of information such as additional modalities, or acoustic channels, or acoustic models. Given the arising problem of information fusion, striking parallels to problems in digital communications are exhibited, where the discovery of the turbo codes by Berrou et al. was a groundbreaking innovation. In this paper, we show ways how to successfully apply the turbo principle to the domain of ASR and thereby provide solutions to the abovementioned information fusion problem. The contribution of our work is fourfold: First, we review the turbo decoding forward-backward algorithm (FBA), giving detailed insights into turbo ASR, and providing a new interpretation and formulation of the so-called extrinsic information being passed between the recognizers. Second, we present a real-time capable turbo-decoding Viterbi algorithm suitable for practical information fusion and recognition tasks. Then we present simulation results for a multimodal example of information fusion. Finally, we prove the suitability of both our turbo FBA and turbo Viterbi algorithm also for a single-channel multimodel recognition task obtained by using two acoustic feature extraction methods. On a small vocabulary task (challenging, since spelling is included), our proposed turbo ASR approach outperforms even the best reference system on average over all SNR conditions and investigated noise types by a relative word error rate (WER) reduction of 22.4% (audio-visual task) and 18.2% (audio-only task), respectively.
Simon Receveur, Robin Weib, Tim Fingscheidt
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 Acoustic event source localization for surveillance in reverberant environments supported by an event onset detection
abstract
This contribution presents a robust approach to acoustic event source localization for surveillance under reverberant environmental conditions. In particular, we support the classical generalized cross-correlation algorithm with phase transform weighting (GCC-PHAT) and the steered response power (SRP) algorithm by a sound activity detection and an event onset detector. The proposed algorithmic framework including spatial minimum tracking and smoothing for the suppression of artifacts in the spatial likelihood function significantly outperforms a respective reference approach, decreasing both the miss ratio by up to 9% absolute, and the average angular estimation error by up to 4°.
Peter Transfeld, Uwe Martens, Harald Binder, Thomas Schypior, Tim Fingscheidt
ICASSP5
2015 An iterative speech model-based a priori SNR estimator
Samy Elshamy, Nilesh Madhu, Wouter Tirry, Tim Fingscheidt
INTERSPEECH4
2015 An acoustic event detection framework and evaluation metric for surveillance in cars
Peter Transfeld, Simon Receveur, Tim Fingscheidt
INTERSPEECH3
2015 A Priori SNR Estimation Using Air- and Bone-Conduction Microphones
abstract
This paper proposes an a priori signal-to-noise ratio (SNR) estimator using an air-conduction (AC) and a bone-conduction (BC) microphone. Among various ways of combining AC and BC microphones for speech enhancement, it is shown that the total enhancement performance can be maximized if the BC microphone is utilized for estimating the power spectral density (PSD) of the desired speech signal. Considering the fact that a small deviation in the speech PSD estimation process brings severe spectral distortion, this paper focuses on controlling weighting factors while estimating the a priori SNR with the decision-directed approach framework. The time–frequency varying weighting factor that is determined by taking a minimum mean square error criterion improves the capability of eliminating residual noise and minimizing speech distortion. Since the weighting factors are also adjusted by measuring the usefulness of the AC and BC microphones, the proposed approach is suitable for tracking the parameter even if the characteristic of environment changes rapidly. The simulation results confirm the superiority of the proposed algorithm to conventional algorithms in high noise environments.
Ho Seon Shin, Tim Fingscheidt, Hong-Goo Kang
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 On speech quality assessment of artificial bandwidth extension
abstract
During the transition to wideband speech telephony, artificial bandwidth extension (ABE) could help to preserve customer satisfaction by enhancing speech quality in case of narrowband (NB) calls. However, the assessment of speech quality for ABE systems is still an open question. In the literature, instrumental measures are often used to judge the quality of ABE solutions. When subjective listening tests are considered, they most often use a comparison category rating (CCR) scale and, more rarely, an absolute category rating (ACR) scale. This paper investigates the relevance of instrumental and subjective assessment methods for ABE systems. An ACR and a CCR test are organized. Their results are compared and discussed. Discrepancies between these two tests open the discussion for the design of a proper subjective listening test for ABE systems. Some instrumental measures are also evaluated. A poor correlation between these measures and the subjective results is observed.
Patrick Bauer, Cyril Guillaume, Wouter Tirry, Tim Fingscheidt
ICASSP4
2014 Variable-length versus fixed-length coding: On tradeoffs for soft-decision decoding
abstract
Variable-length codes (VLCs) are widely used in media transmission. Compared to fixed-length codes (FLCs), VLCs can represent the same message with a lower bit rate, thus having a better compression performance. But inevitably, VLCs are very sensitive to transmission errors. In this work, based on the trellis representation for VLCs and the BCJR algorithm, we present a variable-length soft-decision decoder utilizing bit-wise channel reliability information and achieving a better error robustness in contrast to hard-decision decoding. Given the application of VLCs in audio coding showing both source correlation and variable block lengths, a strong dependency of performance is observed for both. Therefore, we point out tradeoffs of (soft-decision) decoded FLCs and VLCs depending on quantization bit rate, source correlation, and block length. We find that VLCs over AWGN channels are only recommended for very low source correlation in combination with very short block lengths and soft-decision decoding.
Sai Han, Tim Fingscheidt
ICASSP2
2014 A compact formulation of turbo audio-visual speech recognition
abstract
Since most automatic speech recognition (ASR) systems still suffer from adverse acoustic conditions and insufficient acoustic modeling, recognition robustness can be improved by integrating further information sources such as additional acoustic channels, modalities, or models. Considering the question of information fusion, interesting parallels to problems in digital communications can be observed, where the turbo principle revolutionized reliable communication. In this paper, we provide new perspectives on turbo ASR: First, we introduce a compact formulation of turbo automatic speech recognition; second, we present a shape-based visual feature extraction algorithm without any learning paradigms. Third, we show an application to an audio-visual speech recognition task on a large data set, where our proposed method clearly outperforms the iterative approach introduced by Shivappa et al. as well as a conventional coupled-hidden-Markov-model approach by up to 23.8% relative reduction in word error rate.
Simon Receveur, Patrick Meyer, Tim Fingscheidt
ICASSP3
2014 Document Writer Analysis with Rejection for Historical Arabic Manuscripts
abstract
Determining the individuality of handwriting in ancient manuscripts is an important aspect of the manuscript analysis process. Automatic identification of writers in historical manuscripts can support historians to gain insights into manuscripts with missing metadata such as writer name, period, and origin. In this paper writer classification and retrieval approaches for multi-page documents in the context of historical manuscripts are presented. The main contribution is a learning-based rejection strategy which utilizes writer retrieval and support vector machines for rejecting a decision if no corresponding writer can be found for a query manuscript. Experiments using different feature extraction methods demonstrate the abilities of our proposed methods. A dedicated data set based on a publicly available database of historical Arabic manuscripts was used and the experiments show promising results.
Daniel Fecker, Abedelkadir Asi, Werner Pantke, Volker Märgner, Jihad El-Sana, Tim Fingscheidt
ICFHR6
2014 An Historical Handwritten Arabic Dataset for Segmentation-Free Word Spotting - HADARA80P
abstract
In this paper, we present a new and freely available dataset comprising 80 pages of an historical handwritten Arabic document in conjunction with a detailed ground truth for the development and evaluation of segmentation-free word spotting approaches. Besides information on the underlying manuscript and technical details, we introduce a comprehensive list of tags that each word is labeled with. These tags can be used for research on specific issues such as dealing with text in different colors. For comparison of different word spotters, a fixed set of 25 keywords with different properties is included. Furthermore, some specifics of spotting on Arabic manuscripts are discussed. We exemplarily present a state-of-the-art word spotting algorithm in its original and a new extended implementation and evaluate both approaches on the new dataset. For comparison, they are also tested on the widely used George Washington dataset. It is shown that the extended word spotter outperforms the original version in terms of mean average precision on both datasets.
Werner Pantke, Martin Dennhardt, Daniel Fecker, Volker Märgner, Tim Fingscheidt
ICFHR5
2014 Writer Identification for Historical Arabic Documents
abstract
Identification of writers of handwritten historical documents is an important and challenging task. In this paper we present several feature extraction and classification approaches for the identification of writers in historical Arabic manuscripts. The approaches are able to successfully identify writers of multipage documents. The feature extraction methods rely on different principles, such as contour-, textural- and key point-based and the classification schemes are based on averaging and voting. For all experiments a dedicated data set based on a publicly available database is used. The experiments show promising results and the best performance was achieved using a novel feature extraction based on key point descriptors.
Daniel Fecker, Abedelkadir Asi, Volker Märgner, Jihad El-Sana, Tim Fingscheidt
ICPR5
2013 Impact of hearing impairment on fricative intelligibility for artificially bandwidth-extended telephone speech in noise
abstract
Because of its limited bandwidth, telephone speech is poorly intelligible. Artificial bandwidth extension (ABWE) reconstructs themissing frequencies aiming at, e.g., higher intelligibility. It was recently demonstrated that hearing-impaired persons wearing a hearing aid benefit from ABWE-enhanced telephone speech. However, it is unclear, whether persons without hearing impairment also take profit from ABWE in the same test conditions and if so, to what extent. This paper presents a subjective listening test with normal-hearing subjects based on meaningless German syllables simulating narrowband (NB), ABWE-enhanced and wideband (WB) telephone speech in two noisy listening conditions. The test results reveal a clear impact of hearing impairment on the ABWE capability to improve telephone intelligibility. For a signal-to-noise ratio (SNR) of 0 dB, subjects with and without hearing impairment similarly benefit from ABWE. At 20 dB SNR, hearing-impaired subjects take even more profit in contrast to normal-hearing subjects.
Patrick Bauer, Tim Fingscheidt
ICASSP3
2013 Towards reproducible evaluation of automotive hands-free systems in dynamic conditions
abstract
Reproducible evaluation of dynamic and nonlinear systems is a nontrivial problem. However, automotive speech processing algorithms such as hands-free systems have to be tested under numerous time-variant conditions in a repeatable fashion. The current way of generating time-varying echo paths, as described in ITU-T Recommendation P.1110, relies on a rotating reflecting surface in a car interior, which lacks both flexibility and reproducibility. We propose an automotive loudspeaker-enclosure-microphone (LEM) system-identification approach based on the normalized least mean squares (NLMS) algorithm and a perfect sweep excitation signal. Time-variant simulations of a nonlinear system model show a significant improvement of error signal attenuation by over 7 dB, compared to a white noise excitation, also confirmed by automotive measurements. We present the necessary steps to identify dynamic automotive LEM systems to obtain traces of impulse responses for later reproducible tests of automotive hands-free systems. The method has been proposed to ITU-T standardization in focus group (FG) CARCOM.
Marc-André Jung, Lucca Richter, Tim Fingscheidt
ICASSP3
2013 On the use of explicit redundancy for delayless soft-decision audio decoding
abstract
Wireless transmission systems for high-quality digital audio signals require a low end-to-end delay and strong robustness against channel distortions. In this work we investigate a Bayesian approach to delayless soft-decision decoding of high-quality audio signals jointly exploiting both implicit redundancy within the audio signal and explicit sample-wise redundancy appended by a channel (block) encoder. Because our approach introduces no algorithmic delay, it can be employed in audio transmission systems that are extremely sensitive to latency like, e. g., wireless digital microphones. Experiments carried out with representative audio signals transmitted over AWGN channels show a significant increase in signal quality.
Florian Pflug, Tim Fingscheidt
ICASSP2
2013 On Evaluation of Segmentation-Free Word Spotting Approaches without Hard Decisions
abstract
Word spotting systems are intended to retrieve occurrences of a given keyword in document images without actually recognizing the full document content. As there is a trend towards segmentation-free word spotting methods, we propose a methodology to evaluate these methods by employing measures that take the quality of the retrieved word locations into account without making hard decisions. We derive a desired evaluation behavior with the help of synthetic examples and show discrepancies of existing evaluation methods. New measures following this behavior are introduced and their differences exemplarily described. The proposed evaluation method is applied to a state-of-the-art word spotting approach.
Werner Pantke, Volker Märgner, Tim Fingscheidt
ICDAR3
2013 Speech quality prediction for artificial bandwidth extension algorithms
abstract
During the transition period from narrowband to wideband speech transmission services, Artificial Bandwidth Extension (ABE) algorithms are able to reduce the perceptual degradation of narrowband-transmitted speech signals by extending the audio bandwidth. In this paper, we analyze whether the resulting speech quality can be predicted reliably with instrumental models. Estimations from the new ITU standard POLQA, its predecessor WB-PESQ and the diagnostic DIAL model are compared to subjective listener judgments. This comparison reveals that the instrumental measures are not fully able to cope with ABE-processed speech, particularly in predicting ABE rank orders reliably. Reasons for this finding and corresponding diagnoses are discussed. Index Terms: speech quality, artificial bandwidth extension, instrumental quality prediction, speech transmission, diagnosis
Sebastian Möller 0001, Emilia Kelaidi, Friedemann Köster, Nicolas Côté, Patrick Bauer, Tim Fingscheidt, Thomas Schlien, Hannu Pulakka, Paavo Alku
INTERSPEECH6
2013 Robust Ultra-Low Latency Soft-Decision Decoding of Linear PCM Audio
abstract
Applications such as professional wireless digital microphones require a transmission of practically uncoded high-quality audio with ultra-low latency on the one hand and robustness to error-prone channels on the other hand. The delay restrictions, however, prohibit the utilization of efficient block or convolutional channel codes for error protection. The contribution of this work is fourfold: We revise and summarize concisely a Bayesian framework for soft-decision audio decoding and present three novel approaches to (almost) latency-free robust decoding of uncompressed audio. Bit reliability information from the transmission channel is exploited, as well as short-term and long-term residual redundancy within the audio signal, and optionally some explicit redundancy in terms of a sample-individual block code. In all cases we utilize variants of higher-order linear prediction to compute prediction probabilities in three novel ways: Firstly by employing a serial cascade of multiple predictors, secondly by exploiting explicit redundancy in form of parity bits, and thirdly by utilizing an interpolative forward/backward prediction algorithm. The first two presented approaches work fully delayless, while the third one introduces an ultra-low algorithmic delay of just a few samples. The effectiveness of the proposed algorithms is proven in simulations with BPSK and typical digital microphone FSK modulation schemes on AWGN and bursty fading channels.
Florian Pflug, Tim Fingscheidt
IEEE Trans. Speech Audio Process.2
2012 MMSE speech enhancement under speech presence uncertainty assuming (generalized) gamma speech priors throughout
abstract
Several investigations showed that speech enhancement approaches can be improved by speech presence uncertainty (SPU) estimation. Although there has been a strong focus on the use of correct statistical models for spectral weighting rules for the last few decades, there is just a few publications about SPU estimation based on a speech prior consistent with the spectral weighting rule. This contribution presents a new consistent solution for MMSE speech amplitude (SA) estimation under SPU, being based on the generalized gamma distribution representing a variety of speech priors. Employing the gamma speech model which is a special case of the generalized gamma distribution, the new approach is shown to outperform both the SPU-based MMSE-SA estimator relying on a Gaussian speech prior, and the gamma MMSE-SA estimation without SPU.
Balázs Fodor, Tim Fingscheidt
ICASSP2
2012 Black box measurement of musical tones produced by noise reduction systems
abstract
In the context of noise reduction algorithms, three instrumental measures are of major interest: the speech component quality, the level of noise attenuation, and noise distortion in terms of musical tones. As several proposals are made for the first two, the amount of musical tones is commonly still subjectively evaluated. Recent exploration of the log-kurtosis ratio for instrumentally measuring musical tones has led to white box test methodologies requiring specific information about the particular noise reduction algorithm. In this paper we propose a simple yet robust instrumental musical tones measurement, which is applicable to arbitrary unknown noise reduction systems, i.e., a black box measurement. A subjective listening test has been conducted to verify the proposed instrumental measure. Our measurement methodology has been proposed as part of an ITU-T Recommendation in Study Group 12, FG CarCOM.
Tim Fingscheidt
ICASSP2
2011 Speech enhancement using a joint map estimator with Gaussian mixture model for (non-)stationary noise
abstract
In many applications non-stationary Gaussian or stationary non Gaussian noises can be observed. In this paper we present a maximum a posteriori estimation jointly of spectral amplitude and phase (JMAP). It principally allows for arbitrary speech models (Gaussian, super-Gaussian, ...), while the noise DFT coefficients pdf is modeled as Gaussian mixture (GMM). Such a GMM covers both a non-Gaussian stationary noise process, but also a non-stationary process that changes between Gaussian noise modes of different variance with probability of the GMM weight. Accordingly, we provide results for these two types of noise, showing superiority over the Gaussian noise model JMAP estimator even in case of ideal noise power estimation.
Balázs Fodor, Tim Fingscheidt
ICASSP2
2011 Delayless soft-decision decoding of high-quality audio transmitted over awgn channels
abstract
Short-range wireless audio transmission with high quality on the one hand often encounters error-prone channels, while on the other hand decoding delay plays a critical role in the application. A lot of prior art in audio error concealment is either only intuitively motivated or adds too much delay to the transmission. In this paper we pro pose a framework for transmitted audio error concealment without any algorithmic delay. As a novelty, it strongly follows a Bayesian approach for quantized but uncompressed audio, which can in principle be applied to any type of wireless channel yielding some kind of reliability information. Two methods to compute audio prediction coefficients are presented, one based on the autocorrelation method, the other one on the normalized least-mean-square (NLMS) algorithm. Simulation results for additive white Gaussian noise (AWGN) channels show the significant effect of the proposed approaches.
Florian Pflug, Tim Fingscheidt
ICASSP2
2011 A data-driven post-filter design based on spatially and temporally smoothed a priori SNR
abstract
A microphone array beamformer combined with a post-filter estimation based on spatial smoothing can deliver good noise attenuation preserving the speech component, however, with disturbing musical tones. On the other hand, the temporal smoothing of the decision-directed (DD) a priori signal-to-noise ratio (SNR) estimation for single-channel noise reduction can suppress musical tones well, however, with speech distortion particularly in speech onset. Based on these facts, we derive a new data-driven multi channel a priori SNR estimation based on both spatial and temporal smoothing for the use in a beamformer post-filter. The new a priori SNR estimation is able to find an optimum compromise between noise attenuation, quality of the speech component, and musical tones suppression.
Tim Fingscheidt
ICASSP2
2011 A Data-Driven Approach to A Priori SNR Estimation
abstract
The a priori signal-to-noise ratio (SNR) plays an important role in many speech enhancement algorithms. In this paper, we present a data-driven approach to a priori SNR estimation. It may be used with a wide range of speech enhancement techniques, such as, e.g., the minimum mean square error (MMSE) (log) spectral amplitude estimator, the super Gaussian joint maximum a posteriori (JMAP) estimator, or the Wiener filter. The proposed SNR estimator employs two trained artificial neural networks, one for speech presence, one for speech absence. The classical decision-directed a priori SNR estimator by Ephraim and Malah is broken down into its two additive components, which now represent the two input signals to the neural networks. Both output nodes are combined to represent the new a priori SNR estimate. As an alternative to the neural networks, also simple lookup tables are investigated. Employment of these data-driven nonlinear a priori SNR estimators reduces speech distortion, particularly in speech onset, while retaining a high level of noise attenuation in speech absence.
Suhadi Suhadi, Carsten Last, Tim Fingscheidt
IEEE Trans. Speech Audio Process.3
2011 A Two-Dimensional Channel Model for Digital Data Storage on Microfilm
abstract
Photographic microfilm has become a promising medium for long-term storage of digital data. We present an end-to-end channel model of the whole processing chain including a novel soft-output demodulator for channels with non-Gaussian noise, such as film. The proposed channel model describes the digital microfilm storage channel with good preciseness.
Christoph Voges, Tim Fingscheidt
IEEE Trans. Commun.2
2010 Performance Evaluation of Iterative Channel Codes for Digital Data Storage on Microfilm
abstract
In the past few years microfilm has gained new research interest as a medium for long-term storage of digital data. This became particularly possible by recent advances in laser film recording technology. In contrast to other optical or magnetic storage media, the microfilm digital channel (MDC) still has been subject to characterization in only a few publications. In this paper we investigate iterative channel codes for the MDC, in particular low-density parity-check and turbo convolutional codes. Simulation results show that practically error-free storage can be achieved with code rates even above 0.85 on a monochrome MDC with binary amplitude-shift keying modulation.
Florian Pflug, Christoph Voges, Tim Fingscheidt
GLOBECOM3
2010 WTIMIT: The TIMIT Speech Corpus Transmitted Over The 3G AMR Wideband Mobile Network
Patrick Bauer, David Scheler, Tim Fingscheidt
LREC3
2009 Entropy-based feature analysis for speech recognition
Panji Setiawan, Harald Höge, Tim Fingscheidt
INTERSPEECH3
2008 An HMM-based artificial bandwidth extension evaluated by cross-language training and test
abstract
Artificial bandwidth extension techniques can be employed in mobile terminals to improve the quality of the far-end speaker's signal at the receiver. To accomplish this, usually statistical models are trained requiring wideband speech material from a language that is expected to be used in the conversation. In practice however, the language of a certain phone conversation is not known to the user equipment. Therefore we investigated the performance of an HMM- based multilingually trained artificial bandwidth extension on speech signals of which the language was unseen in training. The cross-language training and test turned out to cause only minor degradations compared to the use of monolingually trained acoustic models of the language used in test. Our findings indicate that artificial bandwidth extension can be efficiently trained with multilingual speech data without significant losses in speech quality.
Patrick Bauer, Tim Fingscheidt
ICASSP2
2008 Towards objective quality assessment of speech enhancement systems in a black box approach
abstract
Quality assessment of speech enhancement systems has to deal with aspects such as distortion of the near-end talker's speech, and with the attenuation and distortion of the noise and the echo in different test cases. We propose first steps into the direction of a new black box objective quality assessment of speech enhancement schemes, based on our previous work on decomposition of the (enhanced) speech signal into its components speech, (residual) noise, and (residual) echo. Having these signals available, to our knowledge, for the first time a black box objective quality assessment of an entire speech enhancement system is proposed allowing for simultaneous measurement of, e.g., noise attenuation, echo return loss enhancement (ERLE), and perceptual evaluation of speech quality (PESQ) of the speech component in a wide range of test scenarios including double-talk. The derived scheme proves to be very useful for testing hands-free devices in practice but also for objective evaluation of sophisticated algorithms in science.
Tim Fingscheidt, Suhadi Suhadi, Kai Steinert
ICASSP1
2008 Hands-free system with low-delay subband acoustic echo control and noise reduction
abstract
Echo cancellation and noise reduction for hands-free systems are challenging tasks in speech signal processing. The presence of strong local speech and noise and a changing acoustical enclosing may severely impair the performance of the algorithms. Usually additional constraints such as a low signal delay are also requested for real time implementation. We present a hands-free system consisting of a delayless sub- band adaptive filter with a low-delay echo and noise suppression postfilter. All parameters are estimated in the subband domain, whereas the filtering takes place in the time domain. Thus, our system has a significantly lower processing delay than similar proposals. We compare its performance with respect to echo and noise attenuation and speech distortion with a state-of-the-art hands-free system in a simulated car environment.
Kai Steinert, Martin Schönle, Christophe Beaugeant, Tim Fingscheidt
ICASSP4
2008 Environment-Optimized Speech Enhancement
abstract
In this paper, we present a training-based approach to speech enhancement that exploits the spectral statistical characteristics of clean speech and noise in a specific environment. In contrast to many state-of-the-art approaches, we do not model the probability density function (pdf) of the clean speech and the noise spectra. Instead, subband-individual weighting rules for noisy speech spectral amplitudes are separately trained for speech presence and speech absence from noise recordings in the environment of interest. Weighting rules for a variety of cost functions are given; they are parameterized and stored as a table look-up. The speech enhancement system simply works by computing the weighting rules from the table look-up indexed by the a posteriori signal-to-noise ratio (SNR) and the a priori SNR for each subband computed on a Bark scale. Optimized for an automotive environment, our approach outperforms known-environment-independent-speech enhancement techniques, namely the a priori SNR-driven Wiener filter and the minimum mean square error (MMSE) log-spectral amplitude estimator, both in terms of speech distortion and noise attenuation.
Tim Fingscheidt, Suhadi Suhadi, Sorel Stan
IEEE Trans. Speech Audio Process.1
2007 Quality assessment of speech enhancement systems by separation of enhanced speech, noise, and echo
abstract
Abstract Quality assessment of speech enhancement systems is a non-trivial task, especially when (residual) noise and echo signalcomponents occur. We present a signal separation schemethat allows for a detailed analysis of unknown speech enhance-ment systems in a black box test scenario. Our approach sep-arates the speech, (residual) noise, and (residual) echo compo-nent of the speech enhancement system in the sending direc-tion (uplink direction). This makes it possible to independentlyjudge the speech degradation and the noise and echo attenua-tion/degradation. While state of the art tests always try to judgethe sending direction signal mixture, our new scheme allows amore reliable analysis in shorter time. It will be very usefulfor testing hands-free devices in practice as well as for testingspeech enhancement algorithms in research and development. Index Terms : objective signal quality assessment, non-blindsignal separation, speech enhancement, hands-free 1. Introduction In science, a comfortable way to evaluate speech enhancementalgorithms is to digitally add near-end speech and noise to theecho signal and thereby construct the microphone signal. Dur-ing the uplink processing of the speech enhancement (hands-free) system the operational influence on the noisy microphonesignal is then to be logged, and later applied individually to thespeech, echo, and noise components of the microphone signal(see, e.g., [1, 2, 3]). This presumes linear processing, as canbe found e.g. in frequency domain noise reduction, where again is applied to the spectral amplitudes. The strength of suchmethod is that one achieves three separate signals: The filteredspeechcomponent, thefilteredechocomponent, andthefilterednoise component, which represent the (slightly) distorted near-end talker’s speech signal, the suppressed echo signal, and theresidual noisesignal, respectively. Focusingonnoisereduction,e.g., aspects such as speech distortion, noise attenuation, andnoisedistortioncan thencomfortably bemeasured or auditivelyassessed.Thishowever isa
Tim Fingscheidt, Suhadi Suhadi
INTERSPEECH1
2007 Speech enhancement with improved a posteriori SNR computation
Suhadi Suhadi, Tim Fingscheidt
INTERSPEECH2
2006 A novel environment-dependent speech enhancement method with optimized memory footprint
Suhadi Suhadi, Sorel Stan, Tim Fingscheidt
INTERSPEECH3
2005 Overcoming the Statistical Independence Assumption w.r.t. Frequency in Speech Enhancement
abstract
In this paper, we give a solution on how to overcome the assumption of statistical independence of adjacent frequency bins in noise reduction techniques. We show that under relaxed assumptions the problem results in an a-priori SNR estimation problem, where all available noisy speech spectral amplitudes (observations) are exploited. Any state-of-the-art noise power spectral density (psd) estimation and weighting rule can be used - they do not need to be restated. In order to solve for an estimator well suited for real-time applications, we model the a-priori SNR values as Markov processes w.r.t. frequency. On the basis of the formulation by Ephraim and Malah, this leads to a new a-priori SNR estimator that yields fewer musical tones.
Tim Fingscheidt, Christophe Beaugeant, Suhadi Suhadi
ICASSP (1)1
2005 Robust speech recognition for mobile devices in car noise
Panji Setiawan, Suhadi Suhadi, Tim Fingscheidt, Sorel Stan
INTERSPEECH3
2004 Revisiting some model-based and data-driven denoising algorithms in Aurora 2 context
abstract
In this paper we evaluate some model-based and data-driven algorithms for robust speech recognition in noise, using the experimental framework provided by ETSI Aurora 2. Specifically, we focus on statistical linear approximation (SLA), sequential interacting multiple models (S-IMM), and histogram normalization (HN). As the baseline for the feature extraction scheme we use the ETSI front-end. Recognition tests on a subset of Aurora 2 show that SLA is approximately 4 % better than HN and that S-IMM is worse than HN by almost 3 % in terms of absolute word accuracy. A comparison with the ETSI advanced front-end (AFE) is also presented. While none of these algorithms outperforms AFE, we identify the reasons why this might have happened and point out potential directions for improvement.
Panji Setiawan, Sorel Stan, Tim Fingscheidt
INTERSPEECH3
2003 An evaluation of VTS and IMM for speaker verification in noise
Suhadi Suhadi, Sorel Stan, Tim Fingscheidt, Christophe Beaugeant
INTERSPEECH3
2002 Network-based vs. distributed speech recognition in adaptive multi-rate wireless systems
Tim Fingscheidt, Stefanie Aalburg, Sorel Stan, Christophe Beaugeant
INTERSPEECH1
2002 Joint source-channel (de-)coding for mobile communications
abstract
Real world source coding algorithms usually leave a certain amount of redundancy within the coded bit stream. Shannon (1948) already mentioned that this redundancy can be exploited at the receiver side to achieve a higher robustness against channel errors. We show how joint source-channel decoding can be performed in a way that is applicable to any mobile communication system standard. Considerable gains in terms of bit error rate or signal-to-noise ratio (SNR) are possible dependent on the amount of redundancy. However, an even better performance can be achieved by changing also the transmitter sided source and channel encoders. We propose an encoding concept employing low-dimensional quantization. Keeping the gross bit rate as well as the clean channel quality the same, it decreases the complexity of the source encoder and the decoder significantly. Finally, we give an application of our methods to spectral coefficient coding in speech transmission over a Rayleigh fading channel resulting in channel SNR gains of about 2 dB as compared to state-of-the-art (de-)coding and bad frame handling methods.
Tim Fingscheidt, Thomas Hindelang, Richard V. Cox, Nambi Seshadri
IEEE Trans. Commun.1
2001 A candidate proposal for a 3GPP adaptive multi-rate wideband speech codec
abstract
This paper describes an adaptive multi-rate wideband (AMR-WB) speech codec proposed for the GSM system and also for the evolving third generation (3G) mobile speech services. The speech codec is based on SB-CELP (splitband-code-excited linear prediction) with five modes operating bit rates from 24 kbit/s down to 9.1 kbit/s. The respective channel coding schemes are based on RSC (recursive systematic code) and UEP (unequal error protection). Both, source and channel codec are designed as homogenous as possible to guarantee robust transmission on current and future mobile radio channels.
Christoph Erdmann, Peter Vary, Kyrill A. Fischer, Matthias Marke, Tim Fingscheidt, Imre Varga, Markus Kaindl, Catherine Quinquis, Balázs Kövesi, Dominique Massaloux
ICASSP6
2001 Softbit speech decoding: a new approach to error concealment
abstract
In digital speech communication over noisy channels there is the need for reducing the subjective effects of residual bit errors which have not been eliminated by channel decoding. This task is usually called error concealment. We describe a new and generalizing approach to error concealment as part of a modified robust speech decoder. It can be applied to any speech codec standard and preserves bit exactness in the case of an error free channel. The proposed method requires bit reliability information provided by the demodulator or by the equalizer or specifically by the channel decoder and can exploit additionally a priori knowledge about codec parameters. We apply our algorithms to PCM, ADPCM, and GSM full-rate speech coding using AWGN, fading, and GSM channel models, respectively. It turns out that the speech quality is significantly enhanced, showing the desired inherent muting mechanism or graceful degradation behavior in the case of extreme adverse transmission conditions.
Tim Fingscheidt, Peter Vary
IEEE Trans. Speech Audio Process.1
2000 Combined Source/Channel (De-)Coding: Can a Priori Information be Used Twice?
abstract
In digital transmission of speech, audio, images and video signals residual redundancy is often left after source coding due to the complexity and delay constraints. This redundancy remains both inside one block or frame but also in a time correlation of subsequent frames. We describe an approach to improve channel and source decoding by using both kinds of correlation. Further on we consider the bit-mapping and the multiplexing in coding and its effect on decoding. The mutual information as measurement for the gain in decoding through a priori information is explained. In this work on combined source and channel decoding, we try to answer the following question: can a priori information that models the source parameters be used twice; first at the channel decoder and then at the source decoder. The channel decoder uses the a priori information that models the bit stream generated by the source coder. This does not capture all the details of the source parameter level statistics. By exploiting the a priori knowledge of parameters (once more) at the source decoder, we show that it is possible to achieve better reconstruction than if this information was used at either of the decoders.
Thomas Hindelang, Tim Fingscheidt, Nambi Seshadri, Richard V. Cox
ICC (3)2
1998 Robust speech decoding: can error concealment be better than error correction?
abstract
Digital speech transmission systems use source coding to reduce the bit rate and channel coding to correct transmission errors. Furthermore, in periods of a very poor channel quality error concealment of residual bit errors becomes necessary as channel decoding fails. However, if the channel is clear, channel coding would not be required at all and the speech quality could be improved by allowing a higher bit rate for source encoding. Usually a compromise is taken between speech quality in case of a clear channel and error robustness in case of poor channel quality. This paper addresses the problem of a joint optimization of error concealment and source/channel coding. Under the premise of a minimum mean square error criterion for signal reconstruction it turns out that error concealment instead of error correction may be the best choice if source coding leaves sufficient residual parameter correlations by less bit rate reduction.
Tim Fingscheidt, Peter Vary, Jesús A. Andonegui
ICASSP1
1997 Robust speech decoding: a universal approach to bit error concealment
abstract
In digital mobile communication systems there is the need for reducing the subjective effects of residual bit errors which have not been eliminated by channel decoding by the use of error concealment techniques. Due to the fact that most standards do not specify these algorithms bit exactly, there is room for new solutions to improve the speech quality. This article develops a new approach for optimum estimation of the speech codec parameters. It can be applied to any speech codec standard if bit reliability information is provided by the demodulator (e.g. DECT), or by the channel decoder (e.g. soft-output Viterbi algorithm-SOVA in GSM). The proposed method includes an inherent muting mechanism leading to a graceful degradation of speech quality in case of adverse transmission conditions. Particularly the additional exploitation of the residual source redundancy, i.e. some a priori knowledge about the codec parameters gives a significant enhancement of the output speech quality. In the case of an error free channel, bit exactness as required by the standards can be preserved.
Tim Fingscheidt, Peter Vary
ICASSP1
1997 Robust GSM speech decoding using the channel decoder's soft output
Tim Fingscheidt, Olaf Scheufen
EUROSPEECH1
1995 Implementation aspects of the GSM half-rate speech codec
Tim Fingscheidt, Thomas Wiechers, Eckhard Delfs
EUROSPEECH1