Dorothea Kolossa

dblp:44/3463 · DBLP profile ↗
← Back
70ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0003-0678-3053ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 40 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 37 · 4 first-author · 5 since 2021Security and privacy · 6 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Extending Information Bottleneck Attribution to Video Sequences for Deepfake Detection
Veronika Solopova, Lucas Schmidt, Vera Schmitt, Dorothea Kolossa
IDA4
2024 Generalization Boost in Bimodal Classification via Data Fusion Trained on Sparse Datasets
abstract
For internet users it is frequently challenging to properly navigate and interpret the overwhelming amount of information available to them. Multimodal artificial intelligence (AI) systems can aid in the process; for example, by helping to debunk misinformation, or by placing video content into a proper context. The training of commensurate AI systems requires the labelling of sizable multimodal datasets, a process that is very time-consuming and prone to being outpaced by the rapid changes of the information landscape. To address this problem, we are introducing a hybrid fusion algorithm, Bimodal Gaussian-Logits glObal stream Weighting (BiGLOW), which performs effective information integration in scenarios in which only a very limited amount of bimodally labelled training data is available. The method is motivated from Bayesian inference; we model a modality-weighted linear combination of differences in logits between class pairs through a Gaussian distribution. We evaluated the proposed method with two bimodal multiclass datasets: the Fakeddit dataset for fake news classification and the TAU Audio-Visual Urban Scenes 2021 for urban scene identification. The two datasets capture four of the most frequently used modalities on the web: text, images, audio, and video. Scenarios with scarce training data were created by reducing the respective size of the bimodally labelled set to various small fractions of the overall data size. Our experimental results indicate that the proposed hybrid fusion algorithm surpasses competing purely neural-network-based models in terms of the model complexity, average accuracy, robustness, and time requirements for both model training and evaluation. We attribute the improved performance of the proposed method to its implied, Bayesian inference inspired regularization.
Dorothea Kolossa, Robert M. Nickel
ICMI2
2024 DistriBlock: Identifying adversarial audio samples by leveraging characteristics of the output distribution
abstract
Adversarial attacks can mislead automatic speech recognition (ASR) systems into predicting an arbitrary target text, thus posing a clear security threat. To prevent such attacks, we propose DistriBlock, an efficient detection strategy applicable to any ASR system that predicts a probability distribution over output tokens in each time step. We measure a set of characteristics of this distribution: the median, maximum, and minimum over the output probabilities, the entropy of the distribution, as well as the Kullback-Leibler and the Jensen-Shannon divergence with respect to the distributions of the subsequent time step. Then, by leveraging the characteristics observed for both benign and adversarial data, we apply binary classifiers, including simple threshold-based classification, ensembles of such classifiers, and neural networks. Through extensive analysis across different state-of-the-art ASR systems and language data sets, we demonstrate the supreme performance of this approach, with a mean area under the receiver operating characteristic curve for distinguishing target adversarial examples against clean and noisy data of 99% and 97%, respectively. To assess the robustness of our method, we show that adaptive adversarial examples that can circumvent DistriBlock are much noisier, which makes them easier to detect through filtering and creates another avenue for preserving the system’s robustness.ics of this distribution: the median, maximum, and minimum over the output probabilities, the entropy of the distribution, as well as the Kullback-Leibler and the Jensen-Shannon divergence with respect to the distributions of the subsequent time step. Then, by leveraging the characteristics observed for both benign and adversarial data, we apply binary classifiers, including simple threshold-based classification, ensembles of such classifiers, and neural networks. Through extensive analysis across different state-of-the-art ASR systems and language data sets, we demonstrate the supreme performance of this approach, with a mean area under the receiver operating characteristic for distinguishing target adversarial examples against clean and noisy data of 99% and 97%, respectively. To assess the robustness of our method, we show that adaptive adversarial examples that can circumvent DistriBlock are much noisier, which makes them easier to detect through filtering and creates another avenue for preserving the system’s robustness.
Matias Pizarro, Dorothea Kolossa, Asja Fischer
UAI2
2022 Exploring accidental triggers of smart speakers
Lea Schönherr, Maximilian Golla, Thorsten Eisenhofer, Jan Wiele, Dorothea Kolossa, Thorsten Holz
Comput. Speech Lang.5
2022 Microscopic and Blind Prediction of Speech Intelligibility: Theory and Practice
abstract
Being able to estimate speech intelligibility without the need for listening tests would confer great benefits for a wide range of speech processing applications. Many attempts have therefore been made to introduce an objective, and ideally referencefree measure for this purpose. Most works analyze speech intelligibility prediction (SIP) methods from a macroscopic point of view, averaging over longer time spans. This paper, in contrast, presents a theoretical framework for the microscopic evaluation of SIP methods. Within our framework, a Statistically estimated Accuracy based on Theory (StAT) is derived, which numerically quantifies the statistical limitations inherent in microscopic SIP. A state-of-the-art approach to microscopic SIP, namely, the use of automatic speech recognition (ASR) to directly predict listening test results, is evaluated within this framework. The practical results are in good agreement with the theory. As the final contribution, a fully blind DIscriminative Speech intelligibility Predictor (DISP) is introduced and is also evaluated within the StAT framework. It is shown that this novel, blind estimator can predict intelligibility as well as-and often even with better accuracy than-the non-blind ASR-based approach, and that its results are again in good agreement with its theoretically derived performance potential.
Mahdie Karbasi, Steffen Zeiler, Dorothea Kolossa
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Data Fusion for Audiovisual Speaker Localization: Extending Dynamic Stream Weights to the Spatial Domain
abstract
Estimating the positions of multiple speakers can be helpful for tasks like automatic speech recognition or speaker diarization. Both applications benefit from a known speaker position when, for instance, applying beamforming or assigning unique speaker identities. Recently, several approaches utilizing acoustic signals augmented with visual data have been proposed for this task. However, both the acoustic and the visual modality may be corrupted in specific spatial regions, for instance due to poor lighting conditions or to the presence of background noise. This paper proposes a novel audiovisual data fusion framework for speaker localization by assigning individual dynamic stream weights to specific regions in the localization space. This fusion is achieved via a neural network, which combines the predictions of individual audio and video trackers based on their time- and location-dependent reliability. A performance evaluation using audiovisual recordings yields promising results, with the proposed fusion approach outperforming all baseline models.
Julio Wissing, Benedikt T. Boenninghoff, Dorothea Kolossa, Tsubasa Ochiai, Marc Delcroix, Keisuke Kinoshita, Tomohiro Nakatani, Shoko Araki, Christopher Schymura
ICASSP3
2021 Fusing Information Streams in End-to-End Audio-Visual Speech Recognition
abstract
End-to-end acoustic speech recognition has quickly gained widespread popularity and shows promising results in many studies. Specifically the joint transformer/CTC model pro-vides very good performance in many tasks. However, under noisy and distorted conditions, the performance still degrades notably. While audio-visual speech recognition can significantly improve the recognition rate of end-to-end models in such poor conditions, it is not obvious how to best utilize any available information on acoustic and visual signal quality and reliability in these models. We thus consider the question of how to optimally inform the transformer/CTC model of any time-variant reliability of the acoustic and visual in-formation streams. We propose a new fusion strategy, incorporating reliability information in a decision fusion net that considers the temporal effects of the attention mechanism. This approach yields significant improvements compared to a state-of-the-art baseline model on the Lip Reading Sentences 2 and 3 (LRS2 and LRS3) corpus. On average, the new sys-tem achieves a relative word error rate reduction of 43% com-pared to the audio-only setup and 31% compared to the audio-visual end-to-end baseline.
Steffen Zeiler, Dorothea Kolossa
ICASSP3
2021 Condition Monitoring for Power Converters via Deep One-Class Classification
abstract
We introduce a novel hybrid approach for the early detection of power converter faults, focusing on the use case of modular multilevel converters. The proposed method is based on training a deep one-class classifier, which learns the characteristics of the normal system operation and can hence recognize deviations even without any training on potential fault conditions of the system. In order to achieve robust and reliable performance, the diagnosis of the system state utilizes short sequences of observations, which are combined through a probabilistic model. The decision about the system state can then take the form of monitoring the T2test statistics, which allows us to control the maximum classification error. This proposed method, Reliability-guided One-Class Classification (ROCC) was tested on data recorded from a Modular Multilevel Converter. The approach is shown to be effective in all test cases, leading to reliable diagnostics even though the classifier is applied to a wide range of unseen conditions.
Nikola Markovic, Daniel Vahle, Volker Staudt, Dorothea Kolossa
ICMLA4
2021 Privacy-Preserving Feature Extraction for Cloud-Based Wake Word Verification
abstract
Wake word detection and verification systems often involve a local, on-device wake word detector and a cloud-based verification node. In such systems, the audio representation sent to the cloud-based server may exhibit sensitive information that might be intercepted by an eavesdropper. To improve privacy of cloud-based wake word verification (WWV) systems, we propose to use a privacy-preserving feature representation that minimizes the automatic speech recognition (ASR) capability of a potential attacker. The proposed approach employs an adversarial training schedule that aims to minimize an attacker’s word error rate (WER) while maintaining a high WWV performance. To this end, we apply an adaptive weighting factor in the combined loss function to control the balance between minimizing the WWV loss and maximizing the ASR loss. We show that the proposed training method significantly reduces possible privacy risks while maintaining a strong WWV performance.
Timm Koppelmann, Alexandru Nelus, Lea Schönherr, Dorothea Kolossa, Rainer Martin 0001
Interspeech4
2021 PILOT: Introducing Transformers for Probabilistic Sound Event Localization
abstract
Sound event localization aims at estimating the positions of sound sources in the environment with respect to an acoustic receiver (e.g. a microphone array). Recent advances in this domain most prominently focused on utilizing deep recurrent neural networks. Inspired by the success of transformer architectures as a suitable alternative to classical recurrent neural networks, this paper introduces a novel transformer-based sound event localization framework, where temporal dependencies in the received multi-channel audio signals are captured via self-attention mechanisms. Additionally, the estimated sound event positions are represented as multivariate Gaussian variables, yielding an additional notion of uncertainty, which many previously proposed deep learning-based systems designed for this application do not provide. The framework is evaluated on three publicly available multi-source sound event localization datasets and compared against state-of-the-art methods in terms of localization error and event detection accuracy. It outperforms all competing systems on all datasets with statistical significant differences in performance.
Christopher Schymura, Benedikt T. Boenninghoff, Tsubasa Ochiai, Marc Delcroix, Keisuke Kinoshita, Tomohiro Nakatani, Shoko Araki, Dorothea Kolossa
Interspeech8
2021 Large-vocabulary Audio-visual Speech Recognition in Noisy Environments
abstract
Audio-visual speech recognition (AVSR) can effectively and significantly improve the recognition rates of small-vocabulary systems, compared to their audio-only counterparts. For large-vocabulary systems, however, there are still many difficulties, such as unsatisfactory video recognition accuracies, that make it hard to improve over audio-only baselines. In this paper, we specifically consider such scenarios, focusing on the large-vocabulary task of the LRS2 database, where audio-only performance is far superior to video-only accuracies, making this an interesting and challenging setup for multi-modal integration.To address the inherent difficulties, we propose a new fusion strategy: a recurrent integration network is trained to fuse the state posteriors of multiple single-modality models, guided by a set of model-based and signal-based stream reliability measures. During decoding, this network is used for stream integration within a hybrid recognizer, where it can thus cope with the time-variant reliability and information content of its multiple feature inputs.We compare the results with end-to-end AVSR systems as well as with competitive hybrid baseline models, finding that the new fusion strategy shows superior results, on average even outperforming oracle dynamic stream weighting, which has so far marked the—realistically unachievable—upper bound for standard stream weighting. Even though the pure lipreading performance is low, audio-visual integration is helpful under all—clean, noisy, and reverberant—conditions. On average, the new system achieves a relative word error rate reduction of 42.18% compared to the audio-only model, pointing at a high effectiveness of the proposed integration approach.
Steffen Zeiler, Dorothea Kolossa
MMSP3
2021 Dompteur: Taming Audio Adversarial Examples
Thorsten Eisenhofer, Lea Schönherr, Joel Frank, Lars Speckemeier, Dorothea Kolossa, Thorsten Holz
USENIX Security Symposium5
2020 Imperio: Robust Over-the-Air Adversarial Examples for Automatic Speech Recognition Systems
abstract
Automatic speech recognition (ASR) systems can be fooled via targeted adversarial examples, which induce the ASR to produce arbitrary transcriptions in response to altered audio signals. However, state-of-the-art adversarial examples typically have to be fed into the ASR system directly, and are not successful when played in a room. Previously published over-the-air adversarial examples fall into one of three categories: they are either handcrafted examples, they are so conspicuous that human listeners can easily recognize the target transcription once they are alerted to its content, or they require precise information about the room where the attack takes place, and are hence not transferable to other rooms.
Lea Schönherr, Thorsten Eisenhofer, Steffen Zeiler, Thorsten Holz, Dorothea Kolossa
ACSAC5
2020 Variational Autoencoder with Embedded Student-t Mixture Model for Authorship Attribution
abstract
Traditional computational authorship attribution describes a classification task in a closed-set scenario.Given a finite set of candidate authors and corresponding labeled texts, the objective is to determine which of the authors has written another set of anonymous or disputed texts.In this work, we propose a probabilistic autoencoding framework to deal with this supervised classification task.Variational autoencoders (VAEs) have had tremendous success in learning latent representations.However, existing VAEs are currently still bound by limitations imposed by the assumed Gaussianity of the underlying probability distributions in the latent space.In this work, we are extending a VAE with an embedded Gaussian mixture model to a Student-t mixture model, which allows for an independent control of the "heaviness" of the respective tails of the implied probability densities.Experiments over an Amazon review dataset indicate superior performance of the proposed method.
Benedikt T. Boenninghoff, Steffen Zeiler, Robert M. Nickel, Dorothea Kolossa
COLING4
2020 A Dynamic Stream Weight Backprop Kalman Filter for Audiovisual Speaker Tracking
abstract
Audiovisual speaker tracking is an application that has been tackled by a wide range of classical approaches based on Gaussian filters, most notably the well-known Kalman filter. Recently, a specific Kalman filter implementation was proposed for this task, which incorporated dynamic stream weights to explicitly control the influence of acoustic and visual observations during estimation. Inspired by recent progress in the context of integrating uncertainty estimates into modern deep learning frameworks, this paper proposes a deep neural-network-based implementation of the Kalman filter with dynamic stream weights, whose parameters can be learned via standard backpropagation. This allows for jointly optimizing the parameters of the model and the dynamic stream weight estimator in a unified framework. An experimental study on audiovisual speaker tracking shows that the proposed model shows comparable performance to state-of-the-art recurrent neural networks with the additional advantage of requiring a smaller number of parameters and providing explicit uncertainty information.
Christopher Schymura, Tsubasa Ochiai, Marc Delcroix, Keisuke Kinoshita, Tomohiro Nakatani, Shoko Araki, Dorothea Kolossa
ICASSP7
2020 Leveraging Frequency Analysis for Deep Fake Image Recognition
abstract
Deep neural networks can generate images that are astonishingly realistic, so much so that it is often hard for humans to distinguish them from actual photos. These achievements have been largely made possible by Generative Adversarial Networks (GANs). While deep fake images have been thoroughly investigated in the image domain{—}a classical approach from the area of image forensics{—}an analysis in the frequency domain has been missing so far. In this paper,we address this shortcoming and our results reveal that in frequency space, GAN-generated images exhibit severe artifacts that can be easily identified. We perform a comprehensive analysis, showing that these artifacts are consistent across different neural network architectures, data sets, and resolutions. In a further investigation, we demonstrate that these artifacts are caused by upsampling operations found in all current GAN architectures, indicating a structural and fundamental problem in the way images are generated via GANs. Based on this analysis, we demonstrate how the frequency representation can be used to identify deep fake images in an automated way, surpassing state-of-the-art methods.
Joel Frank, Thorsten Eisenhofer, Lea Schönherr, Asja Fischer, Dorothea Kolossa, Thorsten Holz
ICML5
2020 Loss Functions for Deep Monaural Speech Enhancement
abstract
Deep neural networks have proven highly effective at speech enhancement, which makes them attractive not just as front-ends for machine listening and speech recognition, but also as enhancement models for the benefit of human listeners. They are, however, usually being trained on loss functions that only assess quality in terms of a minimum mean squared error. This is neglecting the fact that human audio perception functions in a manner far better described by logarithmic measures than linear ones, that psychoacoustic hearing thresholds limit the perceptibility of many signal components in a mixture, and that a degree of continuity of signals may also be expected. Hence, sudden changes in the gain of a system may be detrimental. In the following, we cast these properties of human perception into a form that can aid the optimization of a deep neural network speech enhancement system. We explore their effects on a range of model topologies, showing the efficacy of the proposed modifications.
Jan Freiwald, Lea Schönherr, Christopher Schymura, Steffen Zeiler, Dorothea Kolossa
IJCNN5
2020 Detecting Adversarial Examples for Speech Recognition via Uncertainty Quantification
abstract
Machine learning systems and also, specifically, automatic speech recognition (ASR) systems are vulnerable against adversarial attacks, where an attacker maliciously changes the input. In the case of ASR systems, the most interesting cases are targeted attacks, in which an attacker aims to force the system into recognizing given target transcriptions in an arbitrary audio sample. The increasing number of sophisticated, quasi imperceptible attacks raises the question of countermeasures. In this paper, we focus on hybrid ASR systems and compare four acoustic models regarding their ability to indicate uncertainty under attack: a feed-forward neural network and three neural networks specifically designed for uncertainty quantification, namely a Bayesian neural network, Monte Carlo dropout, and a deep ensemble. We employ uncertainty measures of the acoustic model to construct a simple one-class classification model for assessing whether inputs are benign or adversarial. Based on this approach, we are able to detect adversarial examples with an area under the receiving operator curve score of more than 0.99. The neural networks for uncertainty quantification simultaneously diminish the vulnerability to the attack, which is reflected in a lower recognition accuracy of the malicious target text in comparison to a standard hybrid ASR system.
Sina Däubener, Lea Schönherr, Asja Fischer, Dorothea Kolossa
INTERSPEECH4
2020 MyFixit: An Annotated Dataset, Annotation Tool, and Baseline Methods for Information Extraction from Repair Manuals
abstract
Text instructions are among the most widely used media for learning and teaching. Hence, to create assistance systems that are capable of supporting humans autonomously in new tasks, it would be immensely productive, if machines were enabled to extract task knowledge from such text instructions. In this paper, we, therefore, focus on information extraction (IE) from the instructional text in repair manuals. This brings with it the multiple challenges of information extraction from the situated and technical language in relatively long and often complex instructions. To tackle these challenges, we introduce a semi-structured dataset of repair manuals. The dataset is annotated in a large category of devices, with information that we consider most valuable for an automated repair assistant, including the required tools and the disassembled parts at each step of the repair progress. We then propose methods that can serve as baselines for this IE task: an unsupervised method based on a bags-of-n-grams similarity for extracting the needed tools in each repair step, and a deep-learning-based sequence labeling model for extracting the identity of disassembled parts. These baseline methods are integrated into a semi-automatic web-based annotator application that is also available along with the dataset.
Nima Nabizadeh, Dorothea Kolossa, Martin Heckmann
LREC2
2020 Audiovisual Speaker Tracking Using Nonlinear Dynamical Systems With Dynamic Stream Weights
abstract
Data fusion plays an important role in many technical applications that require efficient processing of multimodal sensory observations. A prominent example is audiovisual signal processing, which has gained increasing attention in automatic speech recognition, speaker localization and related tasks. If appropriately combined with acoustic information, additional visual cues can help to improve the performance in these applications, especially under adverse acoustic conditions. A dynamic weighting of acoustic and visual streams based on instantaneous sensor reliability measures is an efficient approach to data fusion in this context. This article presents a framework that extends the well-established theory of nonlinear dynamical systems with the notion of dynamic stream weights for an arbitrary number of sensory observations. It comprises a recursive state estimator based on the Gaussian filtering paradigm, which incorporates dynamic stream weights into a framework closely related to the extended Kalman filter. Additionally, a convex optimization approach to estimate oracle dynamic stream weights in fully observed dynamical systems utilizing a Dirichlet prior is presented. This serves as a basis for a generic parameter learning framework of dynamic stream weight estimators. The proposed system is application-independent and can be easily adapted to specific tasks and requirements. A study using audiovisual speaker tracking tasks is considered as an exemplary application in this work. An improved tracking performance of the dynamic stream-weight-based estimation framework over state-of-the-art methods is demonstrated in the experiments.
Christopher Schymura, Dorothea Kolossa
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Joining Sound Event Detection and Localization Through Spatial Segregation
abstract
Identification and localization of sounds are both integral parts of computational auditory scene analysis. Although each can be solved separately, the goal of forming coherent auditory objects and achieving a comprehensive spatial scene understanding suggests pursuing a joint solution of the two problems. This article presents an approach that robustly binds localization with the detection of sound events in a binaural robotic system. Both tasks are joined through the use of spatial stream segregation which produces probabilistic time-frequency masks for individual sources attributable to separate locations, enabling segregated sound event detection operating on these streams. We use simulations of a comprehensive suite of test scenes with multiple co-occurring sound sources, and propose performance measures for systematic investigation of the impact of scene complexity on this segregated detection of sound types. Analyzing the effect of spatial scene arrangement, we show how a robot could facilitate high performance through optimal head rotation. Furthermore, we investigate the performance of segregated detection given possible localization error as well as error in the estimation of number of active sources. Our analysis demonstrates that the proposed approach is an effective method to obtain joint sound event location and type information under a wide range of conditions.
Ivo Trowitzsch, Christopher Schymura, Dorothea Kolossa, Klaus Obermayer
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 Explainable Authorship Verification in Social Media via Attention-based Similarity Learning
abstract
Authorship verification is the task of analyzing the linguistic patterns of two or more texts to determine whether they were written by the same author or not. The analysis is traditionally performed by experts who consider linguistic features, which include spelling mistakes, grammatical inconsistencies, and stylistics for example. Machine learning algorithms, on the other hand, can be trained to accomplish the same, but have traditionally relied on so-called stylometric features. The disadvantage of such features is that their reliability is greatly diminished for short and topically varied social media texts. In this interdisciplinary work, we propose a substantial extension of a recently published hierarchical Siamese neural network approach, with which it is feasible to learn neural features and to visualize the decision-making process. For this purpose, a new large-scale corpus of short Amazon reviews for text comparison research is compiled and we show that the Siamese network topologies outperform state-of-the-art approaches that were built up on stylometric features. Our linguistic analysis of the internal attention weights of the network shows that the proposed method is indeed able to latch on to some traditional linguistic categories.
Benedikt T. Boenninghoff, Steffen Hessler, Dorothea Kolossa, Robert M. Nickel
IEEE BigData3
2019 Similarity Learning for Authorship Verification in Social Media
abstract
Authorship verification tries to answer the question if two documents with unknown authors were written by the same author or not. A range of successful technical approaches has been proposed for this task, many of which are based on traditional linguistic features such as n-grams. These algorithms achieve good results for certain types of written documents like books and novels. Forensic authorship verification for social media, however, is a much more challenging task since messages tend to be relatively short, with a large variety of different genres and topics. At this point, traditional methods based on features like n-grams have had limited success. In this work, we propose a new neural network topology for similarity learning that significantly improves the performance on the author verification task with such challenging data sets.
Benedikt T. Boenninghoff, Robert M. Nickel, Steffen Zeiler, Dorothea Kolossa
ICASSP4
2019 Learning Dynamic Stream Weights for Linear Dynamical Systems Using Natural Evolution Strategies
abstract
Multimodal data fusion is an important aspect of many object localization and tracking frameworks that rely on sensory observations from different sources. A prominent example is audiovisual speaker localization, where the incorporation of visual information has shown to benefit overall performance, especially in adverse acoustic conditions. Recently, the notion of dynamic stream weights as an efficient data fusion technique has been introduced into this field. Originally proposed in the context of audiovisual automatic speech recognition, dynamic stream weights allow for effective sensory-level data fusion on a per-frame basis, if reliability measures for the individual sensory streams are available. This study proposes a learning framework for dynamic stream weights based on natural evolution strategies, which does not require the explicit computation of oracle information. An experimental evaluation based on recorded audiovisual sequences shows that the proposed approach outperforms conventional methods based on supervised training in terms of localization performance.
Christopher Schymura, Dorothea Kolossa
ICASSP2
2019 Hybrid Condition Monitoring for Power Electronic Systems
abstract
This paper proposes a novel approach for condition monitoring of power electronic systems. When monitoring the state of a power system, reliability is crucial, as this type of system is usually operated continuously for long periods of time, and as both missed faults as well as false detections can easily become prohibitively expensive. Recently, machine-learning-based methods for fault detection of power systems have gained popularity, since they can overcome many of the constrains of model-based techniques. Most of these methods train classifiers for different states of the system under test, and thus, the problem of fault detection becomes a problem of classification. In this paper we compare two of such recent techniques. We show that despite good results, it cannot reasonably be expected that the state classification is solved perfectly for every instant of time, which makes the application of such classifiers infeasible in practical systems. In order to overcome these issues, we propose to re-formulate the task into one of hybrid—neural and statistical—cross-temporal hypothesis testing. This novel hybrid framework allows us to build upon the previous machine-learning-based classification approaches, and to achieve full reliability on a challenging dataset of fault monitoring measurements for a buck-converter.
Nikola Markovic, Thomas Stoetzel, Volker Staudt, Dorothea Kolossa
ICMLA4
2019 Adversarial Attacks Against Automatic Speech Recognition Systems via Psychoacoustic Hiding
Lea Schönherr, Katharina Kohls, Steffen Zeiler, Thorsten Holz, Dorothea Kolossa
NDSS5
2019 Speaker-adapted neural-network-based fusion for multimodal reference resolution
abstract
Humans use a variety of approaches to reference objects in the external world, including verbal descriptions, hand and head gestures, eye gaze or any combination of them.The amount of useful information from each modality, however, may vary depending on the specific person and on several other factors.For this reason, it is important to learn the correct combination of inputs for inferring the best-fitting reference.In this paper, we investigate speaker-dependent and independent fusion strategies in a multimodal reference resolution task.We show that without any change in the modality models, only through an optimized fusion technique, it is possible to reduce the error rate of the system on a reference resolution task by more than 50%.
Diana Kleingarn, Nima Nabizadeh, Martin Heckmann, Dorothea Kolossa
SIGdial4
2018 Potential-Field-Based Active Exploration for Acoustic Simultaneous Localization and Mapping
abstract
This paper presents a novel framework for active exploration in the context of acoustic simultaneous localization and mapping (SLAM) using a microphone array mounted on a mobile robotic agent. Acoustic SLAM aims at building a map of acoustic sources present in the environment and simultaneously estimating the agent's own trajectory and position within this map. Two important aspects of this task are robustness against disturbances arising from reverberation and sensor imperfections and an appropriate degree of exploration to achieve high map accuracy. Several approaches to the latter aspect using information-theoretic measures have recently been proposed. This study extends these approaches into a framework based on the potential field method, which is a widely used technique for robotic path planning and navigation. It allows to determine exploratory movement trajectories for the robotic agent via gradient descent, without requiring computationally expensive Monte Carlo simulations to predict the effects of specific trajectory choices. Furthermore, additional constraints like maintaining a safe distance to acoustic sources can easily be integrated into this framework. Experimental evaluation demonstrates that the proposed method yields adequate exploration strategies of the acoustic environment leading to accurate map estimates.
Christopher Schymura, Dorothea Kolossa
ICASSP2
2018 A Speech-Based On-Demand Intersection Assistant Prototype
abstract
We have recently proposed a speech-based on- demand intersection assistant which helps the driver to handle urban intersections by informing him of the traffic situation on the right hand side and recommending suitable gaps in traffic. In a previous user study, conducted in a simulator, we could show that the system is in general well accepted and preferred by drivers compared to driving without assistance or with only visual support. In this paper, we report on an implementation of this system and its evaluation in real urban traffic. We use LIDAR sensors for the perception of the traffic environment. A scene analyzer estimates the gaps between the vehicles in real time. The result of this analysis is provided to a dialog manager, which uses it to inform the driver of approaching vehicles and suitable gaps. While approaching the intersection, the driver can activate the system via a wake-up-word and control it with subsequent speech commands. The design of the data analyzer and dialog manager is based on evaluations at real intersections. The resulting system can provide suitable support to the driver in a wide range of traffic situations.
Dennis Orth, Bram Bolder, Nico Steinhardt, Mark Dunn, Dorothea Kolossa, Martin Heckmann
Intelligent Vehicles Symposium5
2017 Spoofing detection via simultaneous verification of audio-visual synchronicity and transcription
abstract
Acoustic speaker recognition systems are very vulnerable to spoofing attacks via replayed or synthesized utterances. One possible countermeasure is audio-visual speaker recognition. Nevertheless, the addition of the visual stream alone does not prevent spoofing attacks completely and only provides further information to assess the authenticity of the utterance. Many systems consider audio and video modalities independently and can easily be spoofed by imitating only a single modality or by a bimodal replay attack with a victim's photograph or video. Therefore, we propose the simultaneous verification of the data synchronicity and the transcription in a challenge-response setup. We use coupled hidden Markov models (CHMMs) for a text-dependent spoofing detection and introduce new features that provide information about the transcriptions of the utterance and the synchronicity of both streams. We evaluate the features for various spoofing scenarios and show that the combination of the features leads to a more robust recognition, also in comparison to the baseline method. Additionally, by evaluating the data on unseen speakers, we show the spoofing detection to be applicable in speaker-independent use-cases.
Lea Schönherr, Steffen Zeiler, Dorothea Kolossa
ASRU3
2017 Benefits of Personalization in the Context of a Speech-Based Left-Turn Assistant
abstract
We have previously introduced a novel Assistance On Demand (AOD) concept in the context of an urban speech-based left-turn assistant which supports the driver in monitoring and decision making by providing recommendations for suitable time gaps to enter the intersection. In a first user study participants showed a clear preference for the AOD system, yet also frequently mentioned that the recommended gaps did not fit their driving behavior. In the user study we present here, we investigate in how far the acceptance and efficiency of the AOD system can be increased by a personalization of the recommended gaps to the individual driver. For this purpose, we estimate individual drivers' gap acceptance from observations of their manual driving and use it to evaluate a default and a personalized variant of the AOD system. Results reveal a clear preference for the personalized assistant compared to the default one and to driving manually.
Dennis Orth, Nadja Schoemig, Christian Mark, Monika Jagiellowicz-Kaufmann, Dorothea Kolossa, Martin Heckmann
AutomotiveUI5
2017 Improving audio-visual speech recognition using deep neural networks with dynamic stream reliability estimates
abstract
Audio-visual speech recognition is a promising approach to tackling the problem of reduced recognition rates under adverse acoustic conditions. However, finding an optimal mechanism for combining multi-modal information remains a challenging task. Various methods are applicable for integrating acoustic and visual information in Gaussian-mixture-model-based speech recognition, e.g., via dynamic stream weighting. The recent advances of deep neural network (DNN)-based speech recognition promise improved performance when using audio-visual information. However, the question of how to optimally integrate acoustic and visual information remains. In this paper, we propose a state-based integration scheme that uses dynamic stream weights in DNN-based audio-visual speech recognition. The dynamic weights are obtained from a time-variant reliability estimate that is derived from the audio signal. We show that this state-based integration is superior to early integration of multi-modal features, even if early integration also includes the proposed reliability estimate. Furthermore, the proposed adaptive mechanism is able to outperform a fixed weighting approach that exploits oracle knowledge of the true signal-to-noise ratio.
Hendrik Meutzner, Ning Ma 0002, Robert M. Nickel, Christopher Schymura, Dorothea Kolossa
ICASSP5
2017 Speaker localization in reverberant rooms based on direct path dominance test statistics
abstract
Speaker localization using microphone arrays is typically based on the expected phase and amplitude differences between microphones as a function of the wave arrival direction. However, in rooms with significant reverberation, the direct sound is contaminated by reflections and localization often fails. Recently, a reverberation-robust localization method was proposed, which uses only the direct-path bins in the short-time Fourier transform (STFT) of the speech signals. The method is based on thresholding according to the ratio between the first two singular values of the spatial spectrum matrix. In this work, a confidence measure is developed based on this ratio, which is then used for speaker localization in a statistical estimation framework, based on a Gaussian mixture model. The paper presents the theory of the proposed method and simulation examples validating the advantages of the new approach.
Boaz Rafaely, Dorothea Kolossa
ICASSP2
2017 Monte Carlo exploration for active binaural localization
abstract
This study introduces a machine hearing system for robot audition, which enables a robotic agent to pro-actively minimize the uncertainty of sound source location estimates through motion. The proposed system is based on an active exploration approach, providing a means to model and predict effects of the agent's future motions on localization uncertainty in a probabilistic manner. Particle filtering is used to estimate the posterior probability density function of the source position from binaural measurements, enabling to jointly assess azimuth and distance of the source. The framework allows to infer and refine a policy to select appropriate actions via a Monte Carlo exploration approach. Experiments in simulated reverberant conditions are conducted, showing that active exploration and the incorporation of distance estimation significantly improve localization performance.
Christopher Schymura, Juan Diego Rios Grajales, Dorothea Kolossa
ICASSP3
2017 A maximum likelihood method for driver-specific critical-gap estimation
abstract
In this work, we introduce a maximum likelihood (ML) method to estimate the smallest accepted gap of a specific driver, the so-called critical gap. Previous methods, like Troutbeck's or Raff's method, are well known and widely used but require a consistently behaving driver, which is usually not met. The methods will be investigated for the personalization of an intersection assistant, which we are currently developing. We start with the hypothesis that the size that constitutes an acceptable gap is driver-dependent. To test this hypothesis, we have performed a simulator study with 9 participants. The results reveal that there is a significant inter-individual difference in the critical gap between the drivers, as postulated. Next we investigate how well we can predict which gap the driver will take when we use the previously estimated personalized critical gap versus a driver independent one. Using the personalized gap compared to a driver independent one reduces the error from 11.8 % to 9 8 %. These results are a clear indication that personalization can be utilized to increase the effectiveness and usability of an intersection assistant.
Dennis Orth, Dorothea Kolossa, Milton Orlando Sarria-Paja, Kersten Schaller, Andreas Pech, Martin Heckmann
Intelligent Vehicles Symposium2
2017 Toward Improved Audio CAPTCHAs Based on Auditory Perception and Language Understanding
abstract
A so-called completely automated public Turing test to tell computers and humans apart (CAPTCHA) represents a challenge-response test that is widely used on the Internet to distinguish human users from fraudulent computer programs, often referred to as bots. To enable access for visually impaired users, most Web sites utilize audio CAPTCHAs in addition to a conventional image-based scheme. Recent research has shown that most currently available audio CAPTCHAs are insecure, as they can be broken by means of machine learning at relatively low costs. Moreover, most audio CAPTCHAs suffer from low human success rates that arise from severe signal distortions. This article proposes two different audio CAPTCHA schemes that systematically exploit differences between humans and computers in terms of auditory perception and language understanding, yielding a better trade-off between usability and security as compared to currently available schemes. Furthermore, we provide an elaborate analysis of Google’s prominent reCAPTCHA that serves as a baseline setting when evaluating our proposed CAPTCHA designs.
Hendrik Meutzner, Santosh Gupta, Viet-Hung Nguyen, Thorsten Holz, Dorothea Kolossa
ACM Trans. Priv. Secur.5
2016 SkypeLine: Robust Hidden Data Transmission for VoIP
abstract
Internet censorship is used in many parts of the world to prohibit free access to online information. Different techniques such as IP address or URL blocking, DNS hijacking, or deep packet inspection are used to block access to specific content on the Internet. In response, several censorship circumvention systems were proposed that attempt to bypass existing filters. Especially systems that hide the communication in different types of cover protocols attracted a lot of attention. However, recent research results suggest that this kind of covert traffic can be easily detected by censors. In this paper, we present SkypeLine, a censorship circumvention system that leverages Direct-Sequence Spread Spectrum (DSSS) based steganography to hide information in Voice-over-IP (VoIP) communication. SkypeLine introduces two novel modulation techniques that hide data by modulating information bits on the voice carrier signal using pseudo-random, orthogonal noise sequences and repeating the spreading operation several times. Our design goals focus on undetectability in presence of a strong adversary and improved data rates. As a result, the hiding is inconspicuous, does not alter the statistical characteristics of the carrier signal, and is robust against alterations of the transmitted packets. We demonstrate the performance of SkypeLine based on two simulation studies that cover the theoretical performance and robustness. Our measurements demonstrate that the data rates achieved with our techniques substantially exceed existing DSSS approaches. Furthermore, we prove the real-world applicability of the presented system with an exemplary prototype for Skype.
Katharina Kohls, Thorsten Holz, Dorothea Kolossa, Christina Pöpper
AsiaCCS3
2016 Twin-HMM-based non-intrusive speech intelligibility prediction
abstract
Most of the objective measures employed for speech intelligibility prediction require a clean reference signal, which is not accessible in all realistic scenarios. In this paper, we propose to re-synthesize the relevant features of the clean signal using only the noisy speech signal and utilize them inside an intelligibility prediction framework which requires a reference. A statistical model called twin hidden Markov model (THMM) is used to synthesize the clean speech features. For the intelligibility prediction framework, the short-time objective intelligibility (STOI) measure is used as an accurate and well-known method. The experimental results show a high correlation between the twin-HMM-based STOI (THMMB-STOI) and the human speech recognition results, even slightly outperforming the conventional STOI predictions computed using the actual clean reference signals.
Mahdie Karbasi, Ahmed Hussen Abdelaziz, Dorothea Kolossa
ICASSP3
2016 Robust audiovisual speech recognition using noise-adaptive linear discriminant analysis
abstract
Automatic speech recognition (ASR) has become a widespread and convenient mode of human-machine interaction, but it is still not sufficiently reliable when used under highly noisy or reverberant conditions. One option for achieving far greater robustness is to include another modality that is unaffected by acoustic noise, such as video information. Currently the most successful approaches for such audiovisual ASR systems, coupled hidden Markov models (HMMs) and turbo decoding, both allow for slight asynchrony between audio and video features, and significantly improve recognition rates in this way. However, both typically still neglect residual errors in the estimation of audio features, so-called observation uncertainties. This paper compares two strategies for adding these observation uncertainties into the decoder, and shows that significant recognition rate improvements are achievable for both coupled HMMs and turbo decoding.
Steffen Zeiler, Robert Nicheli, Ning Ma 0002, Guy J. Brown, Dorothea Kolossa
ICASSP5
2016 Dynamic Stream Weighting for Turbo-Decoding-Based Audiovisual ASR
Sebastian Gergen, Steffen Zeiler, Ahmed Hussen Abdelaziz, Robert M. Nickel, Dorothea Kolossa
INTERSPEECH5
2016 Blind Non-Intrusive Speech Intelligibility Prediction Using Twin-HMMs
Mahdie Karbasi, Ahmed Hussen Abdelaziz, Hendrik Meutzner, Dorothea Kolossa
INTERSPEECH4
2016 Introducing the Turbo-Twin-HMM for Audio-Visual Speech Enhancement
Steffen Zeiler, Hendrik Meutzner, Ahmed Hussen Abdelaziz, Dorothea Kolossa
INTERSPEECH4
2016 Environmentally robust audio-visual speaker identification
abstract
To improve the accuracy of audio-visual speaker identification, we propose a new approach, which achieves an optimal combination of the different modalities on the score level. We use the i-vector method for the acoustics and the local binary pattern (LBP) for the visual speaker recognition. Regarding the input data of both modalities, multiple confidence measures are utilized to calculate an optimal weight for the fusion. Thus, oracle weights are chosen in such a way as to maximize the difference between the score of the genuine speaker and the person with the best competing score. Based on these oracle weights a mapping function for weight estimation is learned. To test the approach, various combinations of noise levels for the acoustic and visual data are considered. We show that the weighted multimodal identification is far less influenced by the presence of noise or distortions in acoustic or visual observations in comparison to an unweighted combination.
Lea Schönherr, Dennis Orth, Martin Heckmann, Dorothea Kolossa
SLT4
2016 Uncertain LDA: Including Observation Uncertainties in Discriminative Transforms
abstract
Linear discriminant analysis (LDA) is a powerful technique in pattern recognition to reduce the dimensionality of data vectors. It maximizes discriminability by retaining only those directions that minimize the ratio of within-class and between-class variance. In this paper, using the same principles as for conventional LDA, we propose to employ uncertainties of the noisy or distorted input data in order to estimate maximally discriminant directions. We demonstrate the efficiency of the proposed uncertain LDA on two applications using state-of-the-art techniques. First, we experiment with an automatic speech recognition task, in which the uncertainty of observations is imposed by real-world additive noise. Next, we examine a full-scale speaker recognition system, considering the utterance duration as the source of uncertainty in authenticating a speaker. The experimental results show that when employing an appropriate uncertainty estimation algorithm, uncertain LDA outperforms its conventional LDA counterpart.
Rahim Saeidi, Ramón Fernandez Astudillo, Dorothea Kolossa
IEEE Trans. Pattern Anal. Mach. Intell.3
2016 General hybrid framework for uncertainty-decoding-based automatic speech recognition systems
Ahmed Hussen Abdelaziz, Dorothea Kolossa
Speech Commun.2
2015 Constructing Secure Audio CAPTCHAs by Exploiting Differences between Humans and Machines
abstract
To prevent abuses of Internet services, CAPTCHAs are used to distinguish humans from programs where an audio-based scheme is beneficial to support visually impaired people. Previous studies show that most audio CAPTCHAs, albeit hard to solve for humans, are lacking security strength. In this work we propose an audio CAPTCHA that is far more robust against automated attacks than it is reported for current CAPTCHA schemes. The CAPTCHA exhibits a good trade-off between human usability and security. This is achieved by exploiting the fact that the human capabilities of language understanding and speech recognition are clearly superior compared to current machines. We evaluate the CAPTCHA security by using a state-of-the-art attack and assess the intelligibility by means of a large-scale listening experiment.
Hendrik Meutzner, Santosh Gupta, Dorothea Kolossa
CHI3
2015 Uncertainty propagation through deep neural networks
abstract
In order to improve the ASR performance in noisy environments, distorted speech is typically pre-processed by a speech enhancement algorithm, which usually results in a speech estimate containing residual noise and distortion.We may also have some measures of uncertainty or variance of the estimate.Uncertainty decoding is a framework that utilizes this knowledge of uncertainty in the input features during acoustic model scoring.Such frameworks have been well explored for traditional probabilistic models, but their optimal use for deep neural network (DNN)-based ASR systems is not yet clear.In this paper, we study the propagation of observation uncertainties through the layers of a DNN-based acoustic model.Since this is intractable due to the nonlinearities of the DNN, we employ approximate propagation methods, including Monte Carlo sampling, the unscented transform, and the piecewise exponential approximation of the activation function, to estimate the distribution of acoustic scores.Finally, the expected value of the acoustic score distribution is used for decoding, which is shown to further improve the ASR accuracy on the CHiME database, relative to a highly optimized DNN baseline.
Ahmed Hussen Abdelaziz, Shinji Watanabe 0001, John R. Hershey, Emmanuel Vincent 0001, Dorothea Kolossa
INTERSPEECH5
2015 Robust speech processing using observation uncertainty and uncertainty propagation: session and paper overview
Ramón Fernandez Astudillo, Shinji Watanabe 0001, Ahmed Hussen Abdelaziz, Dorothea Kolossa
INTERSPEECH4
2015 Binaural sound source localisation and tracking using a dynamic spherical head model
abstract
This paper introduces a binaural model for the localisation and tracking of a moving sound source’s azimuth in the horizontal plane. The model uses a nonlinear state space representation of the sound source dynamics including the current position of the listener’s head. The state is estimated via an unscented Kalman Filter by comparing the interaural level and time differences of the binaural signal with semi-analytically derived localisation cues from a spherical head model. The localisation performance of the model is evaluated in combination with two different head movement approaches based on openand closed-loop control strategies. The results show that adaptive strategies outperform non-adaptive ones and are able to compensate systematic deviations between the spherical head model and human heads.
Christopher Schymura, Fiete Winter, Dorothea Kolossa, Sascha Spors
INTERSPEECH3
2015 Learning Dynamic Stream Weights For Coupled-HMM-Based Audio-Visual Speech Recognition
abstract
With the increasing use of multimedia data in communication technologies, the idea of employing visual information in automatic speech recognition (ASR) has recently gathered momentum. In conjunction with the acoustical information, the visual data enhances the recognition performance and improves the robustness of ASR systems in noisy and reverberant environments. In audio-visual systems, dynamic weighting of audio and video streams according to their instantaneous confidence is essential for reliably and systematically achieving high performance. In this paper, we present a complete framework that allows blind estimation of dynamic stream weights for audio-visual speech recognition based on coupled hidden Markov models (CHMMs). As a stream weight estimator, we consider using multilayer perceptrons and logistic functions to map multidimensional reliability measure features to audiovisual stream weights. Training the parameters of the stream weight estimator requires numerous input-output tuples of reliability measure features and their corresponding stream weights. We estimate these stream weights based on oracle knowledge using an expectation maximization algorithm. We define 31-dimensional feature vectors that combine model-based and signal-based reliability measures as inputs to the stream weight estimator. During decoding, the trained stream weight estimator is used to blindly estimate stream weights. The entire framework is evaluated using the Grid audio-visual corpus and compared to state-of-the-art stream weight estimation strategies. The proposed framework significantly enhances the performance of the audio-visual ASR system in all examined test conditions.
Ahmed Hussen Abdelaziz, Steffen Zeiler, Dorothea Kolossa
IEEE ACM Trans. Audio Speech Lang. Process.3
2014 Using automatic speech recognition for attacking acoustic CAPTCHAs: the trade-off between usability and security
abstract
A common method to prevent automated abuses of Internet services is utilizing challenge-response tests that distinguish human users from machines. These tests are known as CAPTCHAs (Completely Automated Public Turing Tests to Tell Computers and Humans Apart) and should represent a task that is easy to solve for humans, but difficult for fraudulent programs. To enable access for visually impaired people, an acoustic CAPTCHA is typically provided in addition to the better-known visual CAPTCHAs. Recent security studies show that most acoustic CAPTCHAs, albeit difficult to solve for humans, can be broken via machine learning.
Hendrik Meutzner, Viet-Hung Nguyen, Thorsten Holz, Dorothea Kolossa
ACSAC4
2014 A newem estimationof dynamic stream weights for coupled-HMM-based audio-visual ASR
abstract
Mutually deploying visual and acoustical information in automatic speech recognition systems increases their robustness against acoustical environmental effects like additive noise and reverberation. Optimal fusion of the audio and video streams requires dynamic adaptation of the relative contribution of each modality. This can be achieved by weighting each stream according to its reliability by an appropriate stream weight. In this paper we propose a new expectation maximization algorithm that estimates oracle frame-dependent stream weights for coupled-HMM-based audio-visual speech recognition. Moreover, we introduce a greedy optimization approach that reasonably initializes this algorithm. The proposed approach is evaluated on the Grid audio-visual database and results in an average relative word error rate reduction of 38% and 58% compared to grid search and Bayes fusion, respectively. The estimated oracle stream weights can be used instead of the conventional global fixed stream weights to improve the supervised training of stream weight estimators.
Ahmed Hussen Abdelaziz, Steffen Zeiler, Dorothea Kolossa
ICASSP3
2014 Reducing the Cost of Breaking Audio CAPTCHAs by Active and Semi-supervised Learning
abstract
CAPTCHAs are challenge-response tests that are widely used in the Internet to distinguish human users from machines. In addition to the well-known visual CAPTCHAs, most Internet services also provide an audio-based scheme, e.g., To enable access for visually impaired users. Recent research has shown that most CAPTCHAs are vulnerable as they can be broken by machine learning techniques. However, such automated attacks come at a relatively high cost as they require human experts to create labels for the unlabeled CAPTCHA samples collected from a website in order to train an attacking system. In this work we utilize active and semi-supervised learning methods for breaking audio CAPTCHAs. We show that these methods can reduce the labeling costs considerably, resulting in an increased vulnerability of audio CAPTCHAs as automated attacks are rendered even more worthwhile. In addition, our findings give insight into improvements to the design of CAPTCHAs, helping to harden prospective audio CAPTCHA schemes against active learning attacks in the future.
Malte Darnstädt, Hendrik Meutzner, Dorothea Kolossa
ICMLA3
2014 Dynamic stream weight estimation in coupled-HMM-based audio-visual speech recognition using multilayer perceptrons
Ahmed Hussen Abdelaziz, Dorothea Kolossa
INTERSPEECH2
2014 Variational Bayesian Inference for Multichannel Dereverberation and Noise Reduction
abstract
Room reverberation and background noise severely degrade the quality of hands-free speech communication systems. In this work, we address the problem of combined speech dereverberation and noise reduction using a variational Bayesian (VB) inference approach. Our method relies on a multichannel state-space model for the acoustic channels that combines frame-based observation equations in the frequency domain with a first-order Markov model to describe the time-varying nature of the room impulse responses. By modeling the channels and the source signal as latent random variables, we formulate a lower bound on the log-likelihood function of the model parameters given the observed microphone signals and iteratively maximize it using an online expectation-maximization approach. Our derivation yields update equations to jointly estimate the channel and source posterior distributions and the remaining model parameters. An inspection of the resulting VB algorithm for blind equalization and channel identification (VB-BENCH) reveals that the presented framework includes previously proposed methods as special cases. Finally, we evaluate the performance of our approach in terms of speech quality, adaptation times, and speech recognition results to demonstrate its effectiveness for a wide range of reverberation and noise conditions.
Dominic Schmid, Gerald Enzner, Sarmad Malik, Dorothea Kolossa, Rainer Martin 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2013 Twin-HMM-based audio-visual speech enhancement
abstract
Most approaches for speech signal processing rely solely on acoustic input, which has the consequence that spectrum estimation becomes exceedingly difficult when the signal-to-noise ratio drops to values near 0 dB. However, alternative sources of information are becoming widely available with increasing use of multimedia data in everyday communication. In the following paper, we suggest to use video input as an auxiliary modality for speech processing by applying a new statistical model - the twin hidden Markov model. The resulting enhancement algorithm for audiovisual data greatly outperforms the standard audio-only log-MMSE estimator on all considered instrumental speech quality measures covering spectral and perceptual quality.
Ahmed Hussen Abdelaziz, Steffen Zeiler, Dorothea Kolossa
ICASSP3
2013 GMM-based significance decoding
abstract
The accuracy of automatic speech recognition systems in noisy and reverberant environments can be improved notably by exploiting the uncertainty of the estimated speech features using so-called uncertainty-of-observation techniques. In this paper, we introduce a new Bayesian decision rule that can serve as a mathematical framework from which both known and new uncertainty-of-observation techniques can be either derived or approximated. The new decision rule in its direct form leads to the new significance decoding approach for Gaussian mixture models, which results in better performance compared to standard uncertainty-of-observation techniques in different additive and convolutive noise scenarios.
Ahmed Hussen Abdelaziz, Steffen Zeiler, Dorothea Kolossa, Volker Leutnant, Reinhold Häb-Umbach
ICASSP3
2013 Using twin-HMM-based audio-visual speech enhancement as a front-end for robust audio-visual speech recognition
Ahmed Hussen Abdelaziz, Steffen Zeiler, Dorothea Kolossa
INTERSPEECH3
2013 Integration of beamforming and uncertainty-of-observation techniques for robust ASR in multi-source environments
Ramón Fernandez Astudillo, Dorothea Kolossa, Alberto Abad, Steffen Zeiler, Rahim Saeidi, Pejman Mowlaee, João Paulo da Silva Neto, Rainer Martin 0001
Comput. Speech Lang.2
2013 Noise-Adaptive LDA: A New Approach for Speech Recognition Under Observation Uncertainty
abstract
Automatic speech recognition (ASR) performance suffers severely from non-stationary noise, precluding widespread use of ASR in natural environments. Recently, so-termed uncertainty-of-observation techniques have helped to recover good performance. These consider the clean speech features as a hidden variable, of which the observable features are only an imperfect estimate. An estimated error variance of features is therefore used to further guide recognition. Based on the same idea, we introduce a new strategy: Reducing the speech feature dimensionality for optimal discriminance under observation uncertainty can yield significantly improved recognition performance, and is derived easily via Fisher's criterion of discriminant analysis.
Dorothea Kolossa, Steffen Zeiler, Rahim Saeidi, Ramón Fernandez Astudillo
IEEE Signal Process. Lett.1
2013 Corpus-Based Speech Enhancement With Uncertainty Modeling and Cepstral Smoothing
abstract
We present a new approach for corpus-based speech enhancement that significantly improves over a method published by Xiao and Nickel in 2010. Corpus-based enhancement systems do not merely filter an incoming noisy signal, but resynthesize its speech content via an inventory of pre-recorded clean signals. The goal of the procedure is to perceptually improve the sound of speech signals in background noise. The proposed new method modifies Xiao's method in four significant ways. Firstly, it employs a Gaussian mixture model (GMM) instead of a vector quantizer in the phoneme recognition front-end. Secondly, the state decoding of the recognition stage is supported with an uncertainty modeling technique. With the GMM and the uncertainty modeling it is possible to eliminate the need for noise dependent system training. Thirdly, the post-processing of the original method via sinusoidal modeling is replaced with a powerful cepstral smoothing operation. And lastly, due to the improvements of these modifications, it is possible to extend the operational bandwidth of the procedure from 4 kHz to 8 kHz. The performance of the proposed method was evaluated across different noise types and different signal-to-noise ratios. The new method was able to significantly outperform traditional methods, including the one by Xiao and Nickel, in terms of PESQ scores and other objective quality measures. Results of subjective CMOS tests over a smaller set of test samples support our claims.
Robert M. Nickel, Ramón Fernandez Astudillo, Dorothea Kolossa, Rainer Martin 0001
IEEE Trans. Speech Audio Process.3
2012 Inventory-style speech enhancement with uncertainty-of-observation techniques
abstract
We present a new method for inventory-style speech enhancement that significantly improves over earlier approaches [1]. Inventory-style enhancement attempts to resynthesize a clean speech signal from a noisy signal via corpus-based speech synthesis. The advantage of such an approach is that one is not bound to trade noise suppression against signal distortion in the same way that most traditional methods do. A significant improvement in perceptual quality is typically the result. Disadvantages of this new approach, however, include speaker dependency, increased processing delays, and the necessity of substantial system training. Earlier published methods relied on a-priori knowledge of the expected noise type during the training process [1]. In this paper we present a new method that exploits uncertainty-of-observation techniques to circumvent the need for noise specific training. Experimental results show that the new method is not only able to match, but outperform the earlier approaches in perceptual quality.
Robert M. Nickel, Ramón Fernandez Astudillo, Dorothea Kolossa, Steffen Zeiler, Rainer Martin 0001
ICASSP3
2012 Decoding of Uncertain Features Using the Posterior Distribution of the Clean Data for Robust Speech Recognition
Ahmed Hussen Abdelaziz, Dorothea Kolossa
INTERSPEECH2
2012 Inventory-Based Audio-Visual Speech Enhancement
Dorothea Kolossa, Robert M. Nickel, Steffen Zeiler, Rainer Martin 0001
INTERSPEECH1
2010 Efficient manycore CHMM speech recognition for audiovisual and multistream data
abstract
Robustness of speech recognition can be significantly improved by multi-stream and especially by audiovisual speech recognition. This is of interest for example for human-machine interaction in noisy reverberant environments, and for transcription of or search in multimedia data. The most robust implementations of audiovisual speech recognition often utilize Coupled Hidden Markov Models (CHMMs), which allow for both modalities to be asynchronous to a certain degree. In contrast to conventional speech recognition, this increases the search space significantly, so current implementations of CHMM systems are often not real-time capable. Thus, for real-time constrained applications such as online transcription of VoIP communication or responsive multi-modal human-machine interaction, using current multiprocessor computing capability is vital. This paper describes how general purpose graphics processors can be used to obtain a real-time implementation of audiovisual and multi-stream speech recognition. The design has been integrated both with a WFST-decoder and a token passing system, with parallelization leading to a maximum speedup factor of 32 and 25, respectively.
Dorothea Kolossa, Jike Chong, Steffen Zeiler, Kurt Keutzer
INTERSPEECH1
2010 WAPUSK20 - A Database for Robust Audiovisual Speech Recognition
Alexander Vorwerk, Dorothea Kolossa, Steffen Zeiler, Reinhold Orglmeister
LREC3
2009 Accounting for the uncertainty of speech estimates in the complex domain for minimum mean square error speech enhancement
Ramón Fernandez Astudillo, Dorothea Kolossa, Reinhold Orglmeister
INTERSPEECH2
2008 Missing feature speech recognition in a meeting situation with maximum SNR beamforming
abstract
Especially for tasks like automatic meeting transcription, it would be useful to automatically recognize speech also while multiple speakers are talking simultaneously. For this purpose, speech separation can be performed, for example by using maximum SNR beamforming. However, even when good interferer suppression is attained, the interfering speech will still be recognizable during those intervals, where the target speaker is silent. In order to avoid the consequential insertion errors, a new soft masking scheme is proposed, which works in the time domain by inducing a large damping on those temporal periods, where the observed direction of arrival does not correspond to that of the target speaker. Even though the masking scheme is aggressive, by means of missing feature recognition the recognition accuracy can be improved significantly, with relative error reductions in the order of 60% compared to maximum SNR beamforming alone, and it is successful also for three simultaneously active speakers. Results are reported based on the SOLON speech recognizer, NTT’s large vocabulary system [1], which is applied here for the recognition of artificially mixed data using real-room impulse responses and the entire clean test set of the Aurora 2 database.
Dorothea Kolossa, Shoko Araki, Marc Delcroix, Tomohiro Nakatani, Reinhold Orglmeister, Shoji Makino
ISCAS1
2003 Beamforming-based convolutive source separation
abstract
A robust independent component analysis (ICA) algorithm for blind separation of convolved mixtures of speech signals is introduced. It is based on two parallel frequency dependent beamforming stages, each of which cancels the signal from one interfering source by frequency dependent null-beamforming. The zero-directions of the beamforming stages are optimized to yield maximally independent outputs, which is achieved via second and higher order statistics. Optimization is carried out in the frequency domain for each frequency band separately, so that phase distortions caused by the room impulse responses are compensated. In contrast to other frequency domain source separation algorithms, this structure does not suffer from permutation of frequency bands, while retaining the major advantage of blind methods, that do not require an external estimate of the direction of arrival (DOA).
Wolf Baumann, Dorothea Kolossa, Reinhold Orglmeister
ICASSP (5)2
2002 Using time-stretched pulses for accurate splitting of speech utterances played back in noisy reverberant environments
Dorothea Kolossa, Qiang Huo
INTERSPEECH1