Tanja Schultz

dblp:s/TanjaSchultz · DBLP profile ↗
← Back
264ranked-venue papers
23as first author
48since 2021 · last 2026
0000-0002-9809-7028ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 197 · 18 first-author · 35 since 2021Artificial intelligence and machine learning · 155 · 14 first-author · 28 since 2021Human-computer interaction and ubiquitous computing · 30 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 3 since 2021Systems, architecture and hardware · 3Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Breathe with Me: Synchronizing Biosignals for User Embodiment in Robots
abstract
Embodiment of users within robotic systems has been explored in human-robot interaction, most often in telepresence and teleoperation. In these applications, synchronized visuomotor feedback can evoke a sense of body ownership and agency, contributing to the experience of embodiment. We extend this work by employing embreathment, the representation of the user's own breath in real time, as a means for enhancing user embodiment experience in robots. In a within-subjects experiment, participants controlled a robotic arm, while its movements were either synchronized or non-synchronized with their own breath. Synchrony was shown to significantly increase body ownership, and was preferred by most participants. We propose the representation of physiological signals as a novel interoceptive pathway for human–robot interaction, and discuss implications for telepresence, prosthetics, collaboration with robots, and shared autonomy.
Iddo Wald, Amber Maimon, Shiyao Zhang 0002, Dennis Küster, Robert Porzel, Tanja Schultz, Rainer Malaka
HRI6
2026 Leveraging Semi-Supervised Learning for Multimodal Hate Speech Data Annotation and Detection
Rathi Adarshi Rammohan, Zhao Ren, Dominik Puchala, Aleksandra Swiderska, Dennis Küster, Tanja Schultz
LREC6
2026 Spiking neural networks for EEG signal analysis: From theory to practice
Siqi Cai 0002, Zheyuan Lin, Wenjie Wei, Shuai Wang 0058, Malu Zhang, Tanja Schultz, Haizhou Li 0001
Neural Networks7
2025 Speech Separation for Low-Resource Languages
abstract
Speech separation aims to equip machines with the human ability of selective listening, i.e. to focus attention on specific information in spoken communication. Studies have shown that the language spoken in a cocktail party scenario matters. While the development of speech separation models can leverage extensive databases, for the majority of languages only very limited data is available. This work presents the very first study on speech separation for low-resource languages. We choose blind source separation as the task to be studied and analyze three strategies to overcome the data scarcity of two low-resource languages from the GlobalPhoneMS2 database. We show that data from other languages can be used to develop models that work for low-resource languages. Finetuning additionally boosts the performance, and training on multiple languages increases both performance and robustness. We show that dynamic mixing in the development helps to find a trade-off between performance and development time.
Marvin Borsdorf, Zexu Pan, Pascal Himmelmann, Haizhou Li 0001, Tanja Schultz
ICASSP5
2025 Shabdh: A multi lingual zero-shot voice cloning approach with speaker disentanglement
abstract
This paper presents a zero-shot voice cloning system leveraging the DIS-Vector framework, which disentangles and encodes key speech features: content, pitch, timbre, and rhythm. Using the YourTTS architecture, the system synthesizes high-quality speech with precise control over both speaker identity and speech characteristics. The approach integrates multilingual data from the LIMMITS-25 dataset. The system employs neural codec TTS and clustering techniques for efficient and personalized speech synthesis. By utilizing DIS-Vector embeddings, the system enables zero-shot voice cloning, allowing the synthesis of speech in the voice of any unseen speaker with high fidelity and adaptability across multiple languages.
Sreeram Manghat, Sreeja Manghat, Tanja Schultz
ICASSP3
2025 ATGnet: Adaptive Temporal Graph Network for EEG-enabled Sound Source Tracking in Cocktail Party Scenarios
abstract
Decoding selective auditory attention from electroencephalography (EEG) signals has gained considerable interest. However, few studies have looked into tracking the dynamic trajectory of moving sound source in complex auditory environments, e.g. with multiple moving speakers. We propose a novel model, namely Adaptive Temporal Graph Network (ATGnet), to continuously track the sound source trajectory using spatial-temporal EEG representations. ATGnet incorporates an adaptive graph topology to extract spatial features, and a graph-convolutional long short-term memory (GC-LSTM) network to capture spatial-temporal dependency. We evaluated ATGnet by performing within-subject leave-one-trial-out cross-validation on EEG signals from 10 participants. Experiment results indicate that ATGnet effectively overcomes the variation of signals across trials and subjects. They further confirm that ATGnet robustly tracks both attended and unattended sound sources, and significantly outperforms traditional methods. ATGnet offers a promising solution to continuous sound source tracking in dynamic conditions, with potential applications in neuro-steered hearing devices.
Saurav Pahuja, Gabriel Ivucic, Siqi Cai 0002, Dashanka De Silva, Tanja Schultz, Haizhou Li 0001
ICASSP5
2025 Breaking Resource Barriers in Speech Emotion Recognition via Data Distillation
Yi Chang 0004, Zhao Ren, Zhonghao Zhao, Thanh Tam Nguyen, Kun Qian 0003, Tanja Schultz, Björn W. Schuller
INTERSPEECH6
2025 Selective Auditory Attention Decoding in Naturalistic Conversations Using EEG-Based Speech Envelope Tracking in Multi-Speaker Environments
Gabriel Ivucic, Saurav Pahuja, Dashanka De Silva, Tanja Schultz
INTERSPEECH4
2025 GTAnet: Geometry-Guided Temporal Attention for EEG-Based Sound Source Tracking in Cocktail Party Scenarios
Saurav Pahuja, Gabriel Ivucic, Siqi Cai 0002, Dashanka De Silva, Haizhou Li 0001, Tanja Schultz
INTERSPEECH6
2025 DiffMV-ETS: Diffusion-based Multi-Voice Electromyography-to-Speech Conversion using Speaker-Independent Speech Training Targets
abstract
Electromyography (EMG) signals have been investigated for novel voice prostheses to enable speech communication with silent articulation.In this work, we propose DiffMV-ETS, a multi-voice, diffusion-based EMG-to-speech system that converts EMG signals to speech in selectable voices.We evaluate it for scenarios where no speech of the speaker wearing EMG sensors is used for training.For this purpose, we introduce EMG-VCTK, a dataset containing EMG and audio recordings of sentences from the Voice Conversion Tool Kit corpus.We compare EMG models trained with audio of the same speaker, of auxiliary speakers, and of text-to-speech systems.Experiments indicate that models retain their intelligibility and naturalness when trained with synthetic speech.DiffMV-ETS enhances the speech naturalness and similarity to unseen voices.To the best of our knowledge, this is the first work to train multi-voice EMG-to-speech systems with speaker-independent targets.
Kevin Scheck, Tom Dombeck, Zhao Ren, Peter Wu, Michael Wand 0002, Tanja Schultz
INTERSPEECH6
2025 NeuroSpex+: Dual-Task Training of Neuro-Guided Speaker Extraction with Speech Envelope and Waveform
Dashanka De Silva, Siqi Cai 0002, Saurav Pahuja, Tanja Schultz, Haizhou Li 0001
INTERSPEECH4
2025 STAA-Net: A Sparse and Transferable Adversarial Attack for Speech Emotion Recognition
abstract
Speech contains rich information on the emotions of humans, and Speech Emotion Recognition (SER) has been an important topic in the area of human-computer interaction. The robustness of SER models is crucial, particularly in privacy-sensitive and reliability-demanding domains like private healthcare. Recently, the vulnerability of deep neural networks in the audio domain to adversarial attacks has become a popular area of research. However, prior works on adversarial attacks in the audio domain primarily rely on iterative gradient-based techniques, which are time-consuming and prone to overfitting the specific threat model. Furthermore, the exploration of sparse perturbations, which have the potential for better stealthiness, remains limited in the audio domain. To address these challenges, we propose a generator-based attack method to generate sparse and transferable adversarial examples to deceive SER models in an end-to-end and efficient manner. We evaluate our method on two widely-used SER datasets, Database of Elicited Mood in Speech (DEMoS) and Interactive Emotional dyadic MOtion CAPture (IEMOCAP), and demonstrate its ability to generate successful sparse adversarial examples in an efficient manner. Moreover, our generated adversarial examples exhibit model-agnostic transferability, enabling effective adversarial attacks on advanced victim models.
Yi Chang 0004, Zhao Ren, Zixing Zhang 0001, Xin Jing 0001, Kun Qian 0003, Xi Shao, Bin Hu 0001, Tanja Schultz, Björn W. Schuller
IEEE Trans. Affect. Comput.8
2024 Uncovering the Full Potential of Visual Grounding Methods in VQA
abstract
Visual Grounding (VG) methods in Visual Question Answering (VQA) attempt to improve VQA performance by strengthening a model's reliance on question-relevant visual information.The presence of such relevant information in the visual input is typically assumed in training and testing.This assumption, however, is inherently flawed when dealing with imperfect image representations common in large-scale VQA, where the information carried by visual features frequently deviates from expected ground-truth contents.As a result, training and testing of VG-methods is performed with largely inaccurate data, which obstructs proper assessment of their potential benefits.In this study, we demonstrate that current evaluation schemes for VG-methods are problematic due to the flawed assumption of availability of relevant visual information.Our experiments show that these methods can be much more effective when evaluation conditions are corrected.Code is provided on GitHub 1 .
Daniel Reich, Tanja Schultz
ACL (1)2
2024 LSTM-MorA: Melody-Accompaniment Classification of MIDI Tracks
Hui Liu 0035, Leon Flaack, Shiyao Zhang 0002, Tanja Schultz
ICANN (9)4
2024 wTIMIT2mix: A Cocktail Party Mixtures Database to Study Target Speaker Extraction for Normal and Whispered Speech
Marvin Borsdorf, Zexu Pan, Haizhou Li 0001, Tanja Schultz
INTERSPEECH4
2024 Macro-descriptors for Alzheimer's disease detection using large language models
Catarina Botelho, John Mendonça, Anna Pompili, Tanja Schultz, Alberto Abad, Isabel Trancoso
INTERSPEECH4
2024 Does the Lombard Effect Matter in Speech Separation? Introducing the Lombard-GRID-2mix Dataset
Iva Ewert, Marvin Borsdorf, Haizhou Li 0001, Tanja Schultz
INTERSPEECH4
2024 Leveraging Graphic and Convolutional Neural Networks for Auditory Attention Detection with EEG
Saurav Pahuja, Gabriel Ivucic, Pascal Himmelmann, Siqi Cai 0002, Tanja Schultz, Haizhou Li 0001
INTERSPEECH5
2024 Investigating Effective Speaker Property Privacy Protection in Federated Learning for Speech Emotion Recognition
Sheng Li 0010, Yang Cao 0011, Zhao Ren, Tanja Schultz
MMAsia5
2024 Neurospex: Neuro-Guided Speaker Extraction With Cross-Modal Fusion
abstract
In the study of auditory attention, it has been revealed that there exists a robust correlation between attended speech and elicited neural responses, measurable through electroencephalography (EEG). Therefore, it is possible to use the attention information available within EEG signals to guide the extraction of the target speaker in a cocktail party computationally. In this paper, we present a neuro-guided speaker extraction model, i.e. NeuroSpex, using the EEG response of the listener as the sole auxiliary reference cue to extract attended speech from monaural speech mixtures. We propose a novel EEG signal encoder that captures the attention information. Additionally, we propose a cross-attention (CA) mechanism to enhance the speech feature representations, generating a speaker extraction mask. Experimental results on a publicly available dataset demonstrate that our proposed model outperforms two baseline models across various evaluation metrics.
Dashanka De Silva, Siqi Cai 0002, Saurav Pahuja, Tanja Schultz, Haizhou Li 0001
SLT4
2024 Virtual Lab - A VR Showroom for Biosignals Research
abstract
In this demo, we present Virtual Lab, a VR Showroom for Biosignals devices and experiments. In Virtual Lab, visitors can explore a digital twin of the real Biosignals Lab at the Cognitive Systems Lab at the University of Bremen, and interact with digital clones of real biosignals devices, including visualizations of their function and purpose.
Asmus Eike Eilks, Ahmed Seyit Kücük, Felix Putze, Tanja Schultz
VRST4
2024 NeuroHeed: Neuro-Steered Speaker Extraction Using EEG Signals
abstract
Humans possess the remarkable ability to selectively attend to a single speaker amidst competing voices and background noise, known asselective auditory attention. Recent studies in auditory neuroscience indicate a strong correlation between the attended speech signal and the corresponding brain's elicited neuronal activities. In this work, we study such brain activities measured using affordable and non-intrusive electroencephalography (EEG) devices. We present NeuroHeed, a speaker extraction model that leverages the listener's synchronized EEG signals to extract the attended speech signal in a cocktail party scenario, in which the extraction process is conditioned on a neuronal attractor encoded from the EEG signal. We propose both an offline and an online NeuroHeed, with the latter designed for real-time inference. In the online NeuroHeed, we additionally propose an autoregressive speaker encoder, which accumulates past extracted speech signals for self-enrollment of the attended speaker information into an auditory attractor, that retains the attentional momentum over time. Online NeuroHeed extracts the current window of the speech signals with guidance from both attractors. Experimental results on KUL dataset two-speaker scenario demonstrate that NeuroHeed effectively extracts brain-attended speech signals with an average scale-invariant signal-to-noise ratio improvement (SI-SDRi) of 14.3 dB and extraction accuracy of 90.8% in offline settings, and SI-SDRi of 11.2 dB and extraction accuracy of 85.1% in online settings.
Zexu Pan, Marvin Borsdorf, Siqi Cai 0002, Tanja Schultz, Haizhou Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Multi-Head Attention and GRU for Improved Match-Mismatch Classification of Speech Stimulus and EEG Response
abstract
This work is based on the participation by the HyperAttention team in the Auditory EEG Decoding Challenge, 2023 (ICASSP 2023 Signal Processing Grand Challenge) task 1, which deals with the match-mismatch classification of speech stimuli and EEG responses of human listeners. We demonstrate the benefits of using mel-spectrograms instead of speech envelopes as input features as well as the effectiveness of Multi-Head Attention and GRU for EEG and speech processing. With a total score of 79.05 %, we reach the second place in the challenge.
Marvin Borsdorf, Saurav Pahuja, Gabriel Ivucic, Siqi Cai 0002, Haizhou Li 0001, Tanja Schultz
ICASSP6
2023 Multi-Speaker Speech Synthesis from Electromyographic Signals by Soft Speech Unit Prediction
abstract
Electromyographic (EMG) signals of articulatory muscles reflect the speech production process even if the user is speaking silently i.e. moving the articulators without producing audible sound. We propose Speech-Unit-based EMG-to-Speech (SU-E2S), a system which relies on EMG to synthesize speech which contains the articulated content but is vocalized in another voice, determined by an acoustic reference utterance. It is based on a Voice Conversion (VC) system which decomposes acoustic speech into continuous soft speech units and a speaker embedding and then reconstructs acoustic features. SU-E2S performs speech synthesis by predicting soft speech units from EMG and using them as input to the VC system. Experiments show that the SU-E2S output is on par in terms of intelligibility of predicting acoustic features directly from EMG, but adds the functionality of synthesizing speech in other voices.
Kevin Scheck, Tanja Schultz
ICASSP2
2023 Towards Reference Speech Characterization for Health Applications
Catarina Botelho, Alberto Abad, Tanja Schultz, Isabel Trancoso
INTERSPEECH3
2023 STE-GAN: Speech-to-Electromyography Signal Conversion using Generative Adversarial Networks
Kevin Scheck, Tanja Schultz
INTERSPEECH2
2023 Enhancing Subject-Independent EEG-Based Auditory Attention Decoding with WGAN and Pearson Correlation Coefficient
abstract
Electroencephalography (EEG) related research faces a significant challenge of subject independence due to the variation in brain signals and responses among individuals. While deep learning models hold promise in addressing this challenge, their effectiveness depends on large datasets for training and generalization across participants. To overcome this limitation, we propose a solution to the above limitation by increasing the size and quality of training data for subject-independent auditory attention decoding (AAD) using EEG with deep learning. Specifically, our method employs a Wasserstein Generative Adversarial Network (WGAN) to generate synthetic data, with Pearson correlation filtering the most realistic samples. We evaluated this method on a publicly available dataset of selective auditory attention experiments and showed superior performance in subject-independent AAD performance. The mixed training set, consisting of both real and artificial data generated by the WGAN+Pearson Correlation Coefficient, demonstrated approximately 4% improvement in AAD accuracy for a 1-second window. These results demonstrate that deep learning remains a viable approach to overcoming data scarcity in subject-independent AAD tasks based on EEG. Moreover, the proposed method has the potential to improve the generalization and reliability of EEG classification tasks.
Saurav Pahuja, Gabriel Ivucic, Felix Putze, Siqi Cai 0002, Haizhou Li 0001, Tanja Schultz
SMC6
2022 Exploring Dementia Detection from Speech: Cross Corpus Analysis
abstract
In this work, we present a qualitative and quantitative analysis of speech and language features derived from two different corpora with the aim to predict early signs of dementia. One corpus consists of the Interdisciplinary Longitudinal Study on Adult Development and Aging (ILSE) designed to investigate satisfying and healthy aging. It consists of more than 6500 hours of biographic interviews from 1000 participants recorded over the course of 20 years. The other corpus is a cross sectional data set created for the ADReSS challenge 2020. In an experimental study we describe a large variety of acoustic and linguistic features that are automatically extracted from speech and corresponding transcriptions. We compare different traditional classifiers, i.e. Gaussian Mixture Models, Linear Discriminant Analysis, and Support Vector Machines. Our final performance results surpass the ADReSS benchmarks.
Ayimnisagul Ablimit, Catarina Botelho, Alberto Abad, Tanja Schultz, Isabel Trancoso
ICASSP4
2022 Towards Closed-Loop Speech Synthesis from Stereotactic EEG: A Unit Selection Approach
abstract
Neurological disorders can severely impact speech communication. Recently, neural speech prostheses have been proposed that reconstruct intelligible speech from neural signals recorded superficially on the cortex. Thus far, it has been unclear whether similar reconstruction is feasible from deeper brain structures, and whether audible speech can be directly synthesized from these reconstructions with low-latency, as required for a practical speech neuroprosthetic. The present study aims to address both challenges. First, we implement a low-latency unit selection based synthesizer that converts neural signals into audible speech. Second, we evaluate our approach on open-loop recordings from 5 patients implanted with stereotactic depth electrodes who conducted a read-aloud task of Dutch utterances. We achieve correlation coefficients significantly higher than chance level of up to 0.6 and an average computational cost of 6.6 ms for each 10 ms frames. While the current reconstructed utterances are not intelligible, our results indicate promising decoding and run-time capabilities that are suitable for investigations of speech processes in closed-loop experiments.
Miguel Angrick, Maarten C. Ottenhoff, Lorenz Diener, Darius Ivucic, Gabriel Ivucic, Sophocles Goulis, Albert J. Colon, G. Louis Wagner, Dean J. Krusienski, Pieter Leonard Kubben, Tanja Schultz, Christian Herff
ICASSP11
2022 Experts Versus All-Rounders: Target Language Extraction for Multiple Target Languages
abstract
Target language extraction (TLE) is a novel task in the field of selective auditory attention, which seeks to extract all speech signals that are spoken in a target language from other sources in a multilingual cocktail party. In our prior studies, a TLE model was trained to extract a predefined, single target language, referred to as Single-TLE. In this paper, we extend the Single-TLE framework to Multi-TLE. Multi-TLE models can also extract all speech signals of one specific target language, but they are optimized on a set of multiple target languages during training. As such, they learn the characteristics of several target languages and can replace multiple Single-TLE models without retraining. We perform experiments on the GlobalPhoneMCP database and incorporate a dynamic language mixing scheme for training. The Multi-TLE model does not only outperform Single-TLE models, but when given a language ID as additional input, it is also able to extract the speech of a specific target language from a mixture which contains multiple learned target languages.
Marvin Borsdorf, Kevin Scheck, Haizhou Li 0001, Tanja Schultz
ICASSP4
2022 Hybrid sub-word segmentation for handling long tail in morphologically rich low resource languages
abstract
Dealing with Out Of Vocabulary (OOV) words or unseen words is one of the main issues of Machine Translation (MT) as well as automatic speech recognition (ASR) systems. For morphologically rich languages having high type token ratio, the OOV percentage is also quite high. Sub-word segmentation has been found to be one of the major approaches in dealing with OOVs. In this paper we present a hybrid sub-word segmentation algorithm to deal with OOVs. A sub-word segmentation evaluation methodology is also presented. We also present results of our segmentation approach in comparison to some of the popular sub-word segmentation algorithms. Malayalam is a morphological rich low resource Indic language with very high type token ratio. All the experiments are done for conversational code-switched Malayalam-English corpus.
Sreeja Manghat, Sreeram Manghat, Tanja Schultz
ICASSP3
2022 An Overview of the FIRST ICASSP Special Session on Computer Audition for Healthcare
abstract
Audio has been increasingly used as a novel digital phenotype that carries important information of the subject’s health status. We can find tremendous efforts given to this young and promising field, i.e., computer audition for healthcare (CA4H), whereas the application scenarios have not been fully studied as compared to its counterpart in medical areas, computer vision. To this end, the first special session held at ICASSP 2020 was dedicated to the topic. In this overview paper, we at first summarise the invited high-quality contributions from leading scientists from a multi- disciplinary background. Then, we provide a detailed grouping of the contributions to several scenarios such as body sound analysis (e.g., heart sound), human speech analysis (e.g., stress detection), and artificial hearing technologies (e.g., cochlear implants). In addition to the collected works, we will compare them with other recent studies within the topic. Finally, we conclude the limitations and perspectives of the current stage. It is interesting and encouraging to find that the state-of-the-art machine learning and audio signal processing techniques have been successfully applied in the health domain, e.g., to fight with the global challenges of COVID-19 and ageing population.
Kun Qian 0003, Tanja Schultz, Björn W. Schuller
ICASSP2
2022 Deep Learning Approaches for Detecting Alzheimer's Dementia from Conversational Speech of ILSE Study
Ayimnisagul Ablimit, Karen Scholz, Tanja Schultz
INTERSPEECH3
2022 Blind Language Separation: Disentangling Multilingual Cocktail Party Voices by Language
Marvin Borsdorf, Kevin Scheck, Haizhou Li 0001, Tanja Schultz
INTERSPEECH4
2022 Challenges of using longitudinal and cross-domain corpora on studies of pathological speech
Catarina Botelho, Tanja Schultz, Alberto Abad, Isabel Trancoso
INTERSPEECH2
2022 Normalization of code-switched text for speech synthesis
Sreeram Manghat, Sreeja Manghat, Tanja Schultz
INTERSPEECH3
2022 SmartHelm: User Studies from Lab to Field for Attention Modeling
abstract
We present three user studies that gradually prepare our prototype system SmartHelm for use in the field, i.e. supporting cargo cyclists on public roads for cargo delivery. SmartHelm is an attention-sensitive smart helmet that integrates none-invasive brain and eye activity detection with hands-free Augmented Reality (AR) components in a speech-enabled outdoor assistance system. The described studies systematically increased in ecological validity from lab to field. The first study consisted of an Augmented Reality preparation examination in the lab. The second study then investigated simulated attention distraction modeling, whereas the third study examined real-world attention distraction modeling while cycling in traffic. During these three studies, multimodal data (EEG, eye-tracking, video, GPS and speech) has been collected synchronously and analyzed in offline and online experiments. Machine Learning models were trained and optimized for attention modeling.Results: Analyses of self-report and objective data during the simulation study show the plausibility of the simulated internal and external distractions. The analysis of behavioral data captured by multimodal biosignals recorded in the field study further shows that real visual attention distractions can be automatically identified using synchronized video and eye-tracking data. Machine Learning methods based on long short-term memory models (LSTMs) indicate that simulated attention distractions can be automatically detected from EEG data, with the best detection performance for mental distractions. Finally, the self-report data suggest that the comfort of the SmartHelm helmet should be further improved for permanent use in road traffic.
Mazen Salous, Dennis Küster, Kevin Scheck, Aytac Dikfidan, Tim Neumann, Felix Putze, Tanja Schultz
SMC7
2022 Evaluation of an Engagement-Aware Recommender System for People with Dementia
abstract
People with Dementia (PwD) and their caregivers can greatly benefit from regular cognitive and social activations. However, these activations need to be engaging and likeable to take effect and to maintain long-term motivation and wellbeing. Taking this into account, finding appropriate items in large activation content catalogues can be a challenging task, which can even lead to unhappiness (”Paradox of Choice”). User-centered Recommender Systems (RS) can help to overcome this obstacle and support PwD and their caregivers in finding engaging and likeable activation contents. In this study, we investigate a dataset collected from PwD and their (in)formal caregivers who jointly used a tablet-based activation system over multiple sessions in an unconstrained care setting. The system applies a content-based recommendation approach based on explicit ratings provided by the PwD and collects audiovisual data during usage. First, we evaluate the real-world user interactions with the RS to gain knowledge about suitable evaluation parameters for our offline analyses. Second, we train a recognition model for engagement based on the audiovisual data and enrich our dataset with the automatically detected information about the PwD’s level of engagement. Last, we apply an offline analysis and compare the RS performance based on different inputs. We show that considering PwD’s level of engagement can help to further improve the rating-based RS in terms of users’ needs and, thus, support them in the activations.
Lars Steinert, Fynn Linus Kölling, Felix Putze, Dennis Küster, Tanja Schultz
UMAP5
2022 Multilingual speech recognition for GlobalPhone languages
Martha Yifiru Tachbelie, Solomon Teferra Abate, Tanja Schultz
Speech Commun.3
2021 Target Language Extraction at Multilingual Cocktail Parties
abstract
Typically, target speaker extraction seeks to extract a target speaker's contribution according to his or her individual voice characteristics. In a “multilingual cocktail party” however, listeners may desire to extract speaker contributions spoken in a particular language, regard-less of the number of contributing speakers. In this paper, we pro-pose a novel task called “target language extraction” (TLE) which extracts voices based on the spoken language rather than on individ-ual speaker characteristics. We introduce a new database for TLE which simulates the multilingual cocktail party problem in mixtures of two and four speakers with German as the target language. The database is derived from the GlobalPhone 2000 Speaker Package and is called “GlobalPhone Multilingual Cocktail Party - German” (GlobalPhoneMCP-GE). Our experimental results show that our approach to TLE achieves very good performance regardless of the number of speakers in the mixture and that TLE generalizes well to unseen speakers and interfering languages. This work represents the first attempt at target language extraction.
Marvin Borsdorf, Haizhou Li 0001, Tanja Schultz
ASRU3
2021 End-to-End Multilingual Automatic Speech Recognition for Less-Resourced Languages: The Case of Four Ethiopian Languages
abstract
The End-to-End (E2E) approach, which maps a sequence of input features into a sequence of graphemes or words, to Automatic Speech Recognition (ASR) is a hot research agenda. It is interesting for less-resourced languages since it avoids the use of pronunciation dictionary, which is one of the major components in the traditional ASR systems. However, like any deep neural network (DNN) approaches, E2E is data greedy. This makes the application of E2E to less-resourced languages questionable. However, using data from other languages in a multilingual (ML) setup is being applied to solve the problem of data scarcity. We have, therefore, conducted ML E2E ASR experiments for four less-resourced Ethiopian languages using different language and acoustic modelling units. The results of our experiments show that relative Word Error Rate (WER) reductions (over the monolingual E2E systems) of up to 29.83% can be achieved by just using data of two related languages in E2E ASR system training. Moreover, we have also noticed that the use of data from less related languages also leads to E2E ASR performance improvement over the use of monolingual data.
Solomon Teferra Abate, Martha Yifiru Tachbelie, Tanja Schultz
ICASSP3
2021 3rd Workshop on Modeling Socio-Emotional and Cognitive Processes from Multimodal Data in the Wild
abstract
Modeling with multimodal data in the wild poses similar challenges in human-computer and human-robot interaction (HCI, HRI). This workshop series thus blends HCI and HRI to jointly address a broad range of current topics in multimodal modeling aimed at designing intelligent systems in the wild. From addressing data scarcity in multimodal user state recognition to emotion prediction from EEG while listening to music, our third workshop in this series aims to further stimulate this important multidisciplinary exchange.
Dennis Küster, Felix Putze, David St-Onge, Pascal E. Fortin, Nerea Urrestilla, Tanja Schultz
ICMI6
2021 Universal Speaker Extraction in the Presence and Absence of Target Speakers for Speech of One and Two Talkers
Marvin Borsdorf, Chenglin Xu, Haizhou Li 0001, Tanja Schultz
Interspeech4
2021 GlobalPhone Mix-To-Separate Out of 2: A Multilingual 2000 Speakers Mixtures Database for Speech Separation
Marvin Borsdorf, Chenglin Xu, Haizhou Li 0001, Tanja Schultz
Interspeech4
2021 Visual Speech for Obstructive Sleep Apnea Detection
Catarina Botelho, Alberto Abad, Tanja Schultz, Isabel Trancoso
Interspeech3
2021 Audio-Visual Recognition of Emotional Engagement of People with Dementia
Lars Steinert, Felix Putze, Dennis Küster, Tanja Schultz
Interspeech4
2021 Speech Activity Detection from Stereotactic EEG
abstract
Recent studies have shown promise for designing Brain-Computer Interfaces (BCIs) to restore speech communication for those suffering from neurological injury or disease. Numerous BCIs have been developed to reconstruct different aspects of speech, such as phonemes and words, from brain activity. However, many challenges remain toward the successful reconstruction of continuous speech from brain activity during speech imagery. Here, we investigate the potential of differentiating speech and non-speech using intracranial brain activity in different frequency bands acquired by stereotactic EEG. The results reveal statistically significant information in the alpha and theta bands for detecting voice activity, and that using a combination of multiple frequency bands further improves performance with over 92% accuracy. Furthermore, the model is causal and can be implemented with low-latency for future closed-loop experiments. These preliminary findings show the potential of cross-frequency brain signal features for detecting speech activity to enhance speech decoding and synthesis models.
Pedram Zanganeh Soroush, Miguel Angrick, Jerry J. Shih, Tanja Schultz, Dean J. Krusienski
SMC4
2021 Verbal fluency in normal aging and cognitive decline: Results of a longitudinal study
Claudia Frankenberg, Jochen Weiner, Maren Knebel, Ayimunishagu Abulimiti, Pablo Toro, Christina J. Herold, Tanja Schultz, Johannes Schröder
Comput. Speech Lang.7
2020 Platform for Studying Self-Repairing Auto-Corrections in Mobile Text Entry based on Brain Activity, Gaze, and Context
abstract
Auto-correction is a standard feature of mobile text entry. While the performance of state-of-the-art auto-correct methods is usually relatively high, any errors that occur are cumbersome to repair, interrupt the flow of text entry, and challenge the user's agency over the process. In this paper, we describe a system that aims to automatically identify and repair auto-correction errors. This system comprises a multi-modal classifier for detecting auto-correction errors from brain activity, eye gaze, and context information, as well as a strategy to repair such errors by replacing the erroneous correction or suggesting alternatives. We integrated both parts in a generic Android component and thus present a research platform for studying self-repairing end-to-end systems. To demonstrate its feasibility, we performed a user study to evaluate the classification performance and usability of our approach.
Felix Putze, Tilman Ihrig, Tanja Schultz, Wolfgang Stuerzlinger
CHI3
2020 Deep Neural Networks Based Automatic Speech Recognition for Four Ethiopian Languages
abstract
In this work, we present speech recognition systems for four Ethiopian languages: Amharic, Tigrigna, Oromo and Wolaytta. We have used comparable training corpora of about 20 to 29 hours speech and evaluation speech of about 1 hour for each of the languages. For Amharic and Tigrigna, lexical and language models of different vocabulary size have been developed. For Oromo and Wolaytta, the training lexicons have been used for decoding. We achieved relative word error rate (WER) reductions for all the languages by using Deep Neural Networks (DNN) based acoustic models that range from 15.1% to 31.45%. The relative improvement obtained for Wolaytta speech recognition system is much higher (31.45%) than the improvement achieved for the other languages. This attributes to the weaker language model and the bigger size of training speech we used for Wolaytta.
Solomon Teferra Abate, Martha Yifiru Tachbelie, Tanja Schultz
ICASSP3
2020 DNN-Based Speech Recognition for Globalphone Languages
abstract
This paper describes new reference benchmark results based on hybrid Hidden Markov Model and Deep Neural Networks (HMM-DNN) for the GlobalPhone (GP) multilingual text and speech database. GP is a multilingual database of high-quality read speech with corresponding transcriptions and pronunciation dictionaries in more than 20 languages. Moreover, we provide new results for five additional languages, namely, Amharic, Oromo, Tigrigna, Wolaytta, and Uyghur. Across the 22 languages considered, the hybrid HMM-DNN models outperform the HMM-GMM based models regardless of the size of the training speech used. Overall, we achieved relative improvements that range from 7.14% to 59.43%.
Martha Yifiru Tachbelie, Ayimunishagu Abulimiti, Solomon Teferra Abate, Tanja Schultz
ICASSP4
2020 Modeling Socio-Emotional and Cognitive Processes from Multimodal Data in the Wild
abstract
Detecting, modeling, and making sense of multimodal data from human users in the wild still poses numerous challenges. Starting from aspects of data quality and reliability of our measurement instruments, the multidisciplinary endeavor of developing intelligent adaptive systems in human-computer or human-robot interaction (HCI, HRI) requires a broad range of expertise and more integrative efforts to make such systems reliable, engaging, and user-friendly. At the same time, the spectrum of applications for machine learning and modeling of multimodal data in the wild keeps expanding. From the classroom to the robot-assisted operation theatre, our workshop aims to support a vibrant exchange about current trends and methods in the field of modeling multimodal data in the wild.
Dennis Küster, Felix Putze, Patrícia Alves-Oliveira, Maike Paetzel-Prüsmann, Tanja Schultz
ICMI5
2020 Towards Engagement Recognition of People with Dementia in Care Settings
abstract
Roughly 50 million people worldwide are currently suffering from dementia. This number is expected to triple by 2050. Dementia is characterized by a loss of cognitive function and changes in behaviour. This includes memory, language skills, and the ability to focus and pay attention. However, it has been shown that secondary therapy such as the physical, social and cognitive activation of People with Dementia (PwD) has significant positive effects. Activation impacts cognitive functioning and can help prevent the magnification of apathy, boredom, depression, and loneliness associated with dementia. Furthermore, activation can lead to higher perceived quality of life. We follow Cohen's argument that activation stimuli have to produce engagement to take effect and adopt his definition of engagement as "the act of being occupied or involved with an external stimulus".
Lars Steinert, Felix Putze, Dennis Küster, Tanja Schultz
ICMI4
2020 Multilingual Acoustic and Language Modeling for Ethio-Semitic Languages
Solomon Teferra Abate, Martha Yifiru Tachbelie, Tanja Schultz
INTERSPEECH3
2020 Automatic Speech Recognition for ILSE-Interviews: Longitudinal Conversational Speech Recordings Covering Aging and Cognitive Decline
Ayimunishagu Abulimiti, Jochen Weiner, Tanja Schultz
INTERSPEECH3
2020 Speech Spectrogram Estimation from Intracranial Brain Activity Using a Quantization Approach
abstract
Direct synthesis from intracranial brain activity into acoustic speech might provide an intuitive and natural communication means for speech-impaired users. In previous studies we have used logarithmic Mel-scaled speech spectrograms (logMels) as an intermediate representation in the decoding from ElectroCorticoGraphic (ECoG) recordings to an audible waveform. Mel-scaled speech spectrograms have a long tradition in acoustic speech processing and speech synthesis applications. In the past, we relied on regression approaches to find a mapping from brain activity to logMel spectral coefficients, due to the continuous feature space. However, regression tasks are unbounded and thus neuronal fluctuations in brain activity may result in abnormally high amplitudes in a synthesized acoustic speech signal. To mitigate these issues, we propose two methods for quantization of power values to discretize the feature space of logarithmic Mel-scaled spectral coefficients by using the median and the logistic formula, respectively, to reduce the complexity and restricting the number of intervals. We evaluate the practicability in a proof-of-concept with one participant through a simple classification based on linear discriminant analysis and compare the resulting waveform with the original speech. Reconstructed spectrograms achieve Pearson correlation coefficients with a mean of r=0.5 +/- 0.11 in a 5-fold cross validation.
Miguel Angrick, Christian Herff, Garett D. Johnson, Jerry J. Shih, Dean J. Krusienski, Tanja Schultz
INTERSPEECH6
2020 Toward Silent Paralinguistics: Speech-to-EMG - Retrieving Articulatory Muscle Activity from Speech
abstract
Electromyographic (EMG) signals recorded during speech production encode information on articulatory muscle activity and also on the facial expression of emotion, thus representing a speech-related biosignal with strong potential for paralinguistic applications.In this work, we estimate the electrical activity of the muscles responsible for speech articulation directly from the speech signal.To this end, we first perform a neural conversion of speech features into electromyographic time domain features, and then attempt to retrieve the original EMG signal from the time domain features.We propose a feed forward neural network to address the first step of the problem (speech features to EMG features) and a neural network composed of a convolutional block and a bidirectional long short-term memory block to address the second problem (true EMG features to EMG signal).We observe that four out of the five originally proposed time domain features can be estimated reasonably well from the speech signal.Further, the five time domain features are able to predict the original speech-related EMG signal with a concordance correlation coefficient of 0.663.We further compare our results with the ones achieved on the inverse problem of generating acoustic speech features from EMG features.
Catarina Botelho, Lorenz Diener, Dennis Küster, Kevin Scheck, Shahin Amiriparian, Björn W. Schuller, Tanja Schultz, Alberto Abad, Isabel Trancoso
INTERSPEECH7
2020 Towards Silent Paralinguistics: Deriving Speaking Mode and Speaker ID from Electromyographic Signals
abstract
Silent Computational Paralinguistics (SCP) -the assessment of speaker states and traits from non-audibly spoken communication -has rarely been targeted in the rich body of either Computational Paralinguistics or Silent Speech Processing.Here, we provide first steps towards this challenging but potentially highly rewarding endeavour: Paralinguistics can enrich spoken language interfaces, while Silent Speech Processing enables confidential and unobtrusive spoken communication for everybody, including mute speakers.We approach SCP by using speech-related biosignals stemming from facial muscle activities captured by surface electromyography (EMG).To demonstrate the feasibility of SCP, we select one speaker trait (speaker identity) and one speaker state (speaking mode).We introduce two promising strategies for SCP: (1) deriving paralinguistic speaker information directly from EMG of silently produced speech versus (2) first converting EMG into an audible speech signal followed by conventional computational paralinguistic methods.We compare traditional feature extraction and decision making approaches to more recent deep representation and transfer learning by convolutional and recurrent neural networks, using openly available EMG data.We find that paralinguistics can be assessed not only from acoustic speech but also from silent speech captured by EMG.
Lorenz Diener, Shahin Amiriparian, Catarina Botelho, Kevin Scheck, Dennis Küster, Isabel Trancoso, Björn W. Schuller, Tanja Schultz
INTERSPEECH8
2020 CSL-EMG_Array: An Open Access Corpus for EMG-to-Speech Conversion
Lorenz Diener, Mehrdad Roustay Vishkasougheh, Tanja Schultz
INTERSPEECH3
2020 Malayalam-English Code-Switched: Grapheme to Phoneme System
Sreeja Manghat, Sreeram Manghat, Tanja Schultz
INTERSPEECH3
2020 Development of Multilingual ASR Using GlobalPhone for Less-Resourced Languages: The Case of Ethiopian Languages
Martha Yifiru Tachbelie, Solomon Teferra Abate, Tanja Schultz
INTERSPEECH3
2020 From Human to Robot Everyday Activity
abstract
The Everyday Activities Science and Engineering (EASE) Collaborative Research Consortium's mission to enhance the performance of cognition-enabled robots establishes its foundation in the EASE Human Activities Data Analysis Pipeline. Through collection of diverse human activity information resources, enrichment with contextually relevant annotations, and subsequent multimodal analysis of the combined data sources, the pipeline described will provide a rich resource for robot planning researchers, through incorporation in the OpenEASE cloud platform.
Celeste Mason, Konrad Gadzicki, Moritz Meier, Florian Ahrens, Thorsten Kluss, Jaime Leonardo Maldonado Cañón, Felix Putze, Thorsten Fehr, Christoph Zetzsche, Manfred Herrmann, Kerstin Schill, Tanja Schultz
IROS12
2020 Automatic Speech Recognition for Uyghur through Multilingual Acoustic Modeling
abstract
Low-resource languages suffer from lower performance of Automatic Speech Recognition (ASR) system due to the lack of data. As a common approach, multilingual training has been applied to achieve more context coverage and has shown better performance over the monolingual training (Heigold et al., 2013). However, the difference between the donor language and the target language may distort the acoustic model trained with multilingual data, especially when much larger amount of data from donor languages is used for training the models of low-resource language. This paper presents our effort towards improving the performance of ASR system for the under-resourced Uyghur language with multilingual acoustic training. For the developing of multilingual speech recognition system for Uyghur, we used Turkish as donor language, which we selected from GlobalPhone corpus as the most similar language to Uyghur. By generating subsets of Uyghur training data, we explored the performance of multilingual speech recognition systems trained with different sizes of Uyghur and Turkish data. The best speech recognition system for Uyghur is achieved by multilingual training using all Uyghur data (10hours) and 17 hours of Turkish data and the WER is 19.17%, which corresponds to 4.95% relative improvement over monolingual training.
Ayimunishagu Abulimiti, Tanja Schultz
LREC2
2020 Analysis of GlobalPhone and Ethiopian Languages Speech Corpora for Multilingual ASR
abstract
In this paper, we present the analysis of GlobalPhone (GP) and speech corpora of Ethiopian languages (Amharic, Tigrigna, Oromo and Wolaytta). The aim of the analysis is to select speech data from GP for the development of multilingual Automatic Speech Recognition (ASR) system for the Ethiopian languages. To this end, phonetic overlaps among GP and Ethiopian languages have been analyzed. The result of our analysis shows that there is much phonetic overlap among Ethiopian languages although they are from three different language families. From GP, Turkish, Uyghur and Croatian are found to have much overlap with the Ethiopian languages. On the other hand, Korean has less phonetic overlap with the rest of the languages. Moreover, morphological complexity of the GP and Ethiopian languages, reflected by type to token ration (TTR) and out of vocabulary (OOV) rate, has been analyzed. Both metrics indicated the morphological complexity of the languages. Korean and Amharic have been identified as extremely morphologically complex compared to the other languages. Tigrigna, Russian, Turkish, Polish, etc. are also among the morphologically complex languages.
Martha Yifiru Tachbelie, Solomon Teferra Abate, Tanja Schultz
LREC3
2019 Improving Fundamental Frequency Generation in EMG-to-Speech Conversion Using a Quantization Approach
abstract
We present a novel approach to generating fundamental frequency (intonation and voicing) trajectories in an EMG-to-Speech conversion Silent Speech Interface, based on quantizing the EMG-to-F0mappings target values and thus turning a regression problem into a recognition problem. We present this method and evaluate its performance with regard to the accuracy of the voicing information obtained as well as the performance in generating plausible intonation trajectories within voiced sections of the signal. To this end, we also present a new measure for overall F0trajectory plausibility, the trajectory-label accuracy (TLAcc), and compare it with human evaluations. Our new F0generation method achieves a significantly better performance than a baseline approach in terms of voicing accuracy, correlation of voiced sections, trajectory-label accuracy and, most importantly, human evaluations.
Lorenz Diener, Tejas Umesh, Tanja Schultz
ASRU3
2019 Speech Reveals Future Risk of Developing Dementia: Predictive Dementia Screening from Biographic Interviews
abstract
Alzheimer's disease is a progressive incurable condition for which the success of any symptomatic therapy depends crucially on the starting time. Ideally it starts before the disease has caused any cognitive impairments. Our work aims at developing speech-based dementia screening methods that detect dementia as early as possible. Here, we aim to predict the outbreak even before clinical screening tests can diagnose the disease. Using the longitudinal ILSE study, we automatically extract features from biographic interviews and predict the development of dementia 5 and 12 years into the future. Our prediction system achieves results of 73.3% and 75.7% unweighted average recall (UAR), respectively, which clearly outperform a prediction based on prior diagnoses or disease prevalence. Thus, the automated analysis of spoken interviews offers a highly effective prediction procedure that allows for easy-to-use, cost-effective casual testing.
Jochen Weiner, Claudia Frankenberg, Johannes Schröder, Tanja Schultz
ASRU4
2019 Comparative Analysis of Think-Aloud Methods for Everyday Activities in the Context of Cognitive Robotics
Moritz Meier, Celeste Mason, Felix Putze, Tanja Schultz
INTERSPEECH4
2019 Biosignal Processing for Human-Machine Interaction
Tanja Schultz
INTERSPEECH1
2019 Augmented Reality Interface for Smart Home Control using SSVEP-BCI and Eye Gaze
abstract
We investigate the integration of eye-tracking and a Brain-Computer Interface into an Augmented Reality system to control a smart home environment. Through a head-mounted display, we present context-dependent control elements which the user selects by directing attention towards them. We show that the combination of both modalities leads to the most robust detection of selections and an interface which is accepted by its users.
Felix Putze, Dennis Weiß, Lisa-Marie Vortmann, Tanja Schultz
SMC4
2019 Visual and Memory-based HCI Obstacles: Behaviour-based Detection and User Interface Adaptations Analysis
abstract
Human Computer Interaction (HCI) performance can be impaired by several HCI obstacles. Cognitive adaptive systems should dynamically detect such obstacles and compensate them with suitable User Interface (UI) adaptation. In this paper, we discuss the detection of two main HCI obstacles: memory-based and visual obstacles. A sequential model based on Long-Short Term Memory (LSTM) is suggested for such a detection of HCI obstacles. UI adaptations for both types of obstacles are discussed and analyzed. We investigate the classification performance on data from a user study with 17 participants. Furthermore, we also investigate the influence of different adaptation mechanisms on performance and subjective assessment. Results show advantages of the proposed sequential LSTM model: on the one hand, the LSTM outperforms the baseline random guess and also a baseline static model LDA in the detection of visual obstacles with 70.6% as an average accuracy. On the other hand, the evaluation of HCI sessions impeded by obstacles but supported with different UI adaptations shows that LSTM results well match the subjective assessment as a plausible detector of behaviour changes.
Mazen Salous, Felix Putze, Markus Ihrig, Tanja Schultz
SMC4
2019 Interpretation of convolutional neural networks for speech spectrogram regression from intracranial recordings
Miguel Angrick, Christian Herff, Garett D. Johnson, Jerry J. Shih, Dean J. Krusienski, Tanja Schultz
Neurocomputing6
2018 interpretation of convolutional neural networks for speech regression from electrocorticography
Miguel Angrick, Christian Herff, Garett D. Johnson, Jerry J. Shih, Dean J. Krusienski, Tanja Schultz
ESANN6
2018 Modeling Cognitive Processes from Multimodal Signals
abstract
Multimodal signals allow us to gain insights into internal cognitive processes of a person, for example: speech and gesture analysis yields cues about hesitations, knowledgeability, or alertness, eye tracking yields information about a person's focus of attention, task, or cognitive state, EEG yields information about a person's cognitive load or information appraisal. Capturing cognitive processes is an important research tool to understand human behavior as well as a crucial part of a user model to an adaptive interactive system such as a robot or a tutoring system. As cognitive processes are often multifaceted, a comprehensive model requires the combination of multiple complementary signals. In this workshop at the ACM International Conference on Multimodal Interfaces (ICMI) conference in Boulder, Colorado, USA, we discussed the state-of-the-art in monitoring and modeling cognitive processes from multi-modal signals.
Felix Putze, Jutta Hild, Akane Sano, Enkelejda Kasneci, Erin Treacy Solovey, Tanja Schultz
ICMI6
2018 Investigating Objective Intelligibility in Real-Time EMG-to-Speech Conversion
abstract
This paper presents an analysis of the influence of various system parameters on the output quality of our neural network based real-time EMG-to-Speech conversion system. This EMG-to-Speech system allows for the direct conversion of facial surface electromyographic signals into audible speech in real time, allowing for a closed-loop setup where users get direct audio feedback. Such a setup opens new avenues for research and applications through co-adaptation approaches. In this paper, we evaluate the influence of several parameters on the output quality, such as time context, EMG-Audio delay, network-, training data- and Mel spectrogram size. The resulting output quality is evaluated based on the objective output quality measure STOI.
Lorenz Diener, Tanja Schultz
INTERSPEECH2
2018 Domain-Adversarial Training for Session Independent EMG-based Speech Recognition
Michael Wand 0002, Tanja Schultz, Jürgen Schmidhuber
INTERSPEECH2
2018 Investigating the Effect of Audio Duration on Dementia Detection Using Acoustic Features
abstract
This paper presents recent progress toward our goal to enable area-wide pre-screening methods for the early detection of dementia based on automatically processing conversational speech of a representative group of more than 200 subjects. We focus on conversational speech since it is the natural form of communication that can be recorded unobtrusively, without adding stress to subjects, and without the need of controlled clinical settings. We describe our unsupervised process chain consisting of voice activity detection and speaker diarization followed by extraction of features and detection of early signs of dementia. The unsupervised system achieves up to 0.645 unweighted average recall (UAR) and compares favorably to a system that was carefully designed on manually annotated data. To further lower the burden for subjects, we investigate UAR over speech duration, and find that about 12 minutes of interview are sufficient to achieve the best UAR.
Jochen Weiner, Miguel Angrick, Srinivasan Umesh, Tanja Schultz
INTERSPEECH4
2018 Detecting Memory-Based Interaction Obstacles with a Recurrent Neural Model of User Behavior
abstract
A memory-based interaction obstacle is a condition which impedes human memory during Human-Computer Interaction, for example a memory-loading secondary task. In this paper, we present an approach to detect the presence of such memory-based interaction obstacles from logged user behavior during system use. For this purpose, we use a recurrent neural network which models the resulting temporal sequences. To acquire a sufficient number of training episodes, we employ a cognitive user simulation. We evaluate the approach with data from a user test and on which we outperform a non-sequential baseline by up to 42% relative.
Felix Putze, Mazen Salous, Tanja Schultz
IUI3
2017 Automatic classification of auto-correction errors in predictive text entry based on EEG and context information
abstract
State-of-the-art auto-correction methods for predictive text entry systems work reasonably well, but can never be perfect due to the properties of human language. We present an approach for the automatic detection of erroneous auto-corrections based on brain activity and text-entry-based context features. We describe an experiment and a new system for the classification of human reactions to auto-correction errors. We show how auto-correction errors can be detected with an average accuracy of 85%.
Felix Putze, Maik Schünemann, Tanja Schultz, Wolfgang Stuerzlinger
ICMI3
2017 Manual and Automatic Transcriptions in Dementia Detection from Speech
abstract
As the population in developed countries is aging, larger numbers of people are at risk of developing dementia. In the near future there will be a need for time- and cost-efficient screening methods. Speech can be recorded and analyzed in this manner, and as speech and language are affected early on in the course of dementia, automatic speech processing can provide valuable support for such screening methods. We present two pipelines of feature extraction for dementia detection: the manual pipeline uses manual transcriptions while the fully automatic pipeline uses transcriptions created by automatic speech recognition (ASR). The acoustic and linguistic features that we extract need no language specific tools other than the ASR system. Using these two different feature extraction pipelines we automatically detect dementia. Our results show that the ASR system’s transcription quality is a good single feature and that the features extracted from automatic transcriptions perform similar or slightly better than the features extracted from the manual transcriptions.
Jochen Weiner, Mathis Engelbart, Tanja Schultz
INTERSPEECH3
2017 Introduction to the Special Issue on Biosignal-Based Spoken Communication
abstract
The papers in this special section focus on biosignal-based spoken communication. Speech production is a complex process resulting from human activities initiated in the brain, eventually leading to muscle activities that produce respiratory, laryngeal, and articulatory gestures which finally create acoustic signals. Traditional speech processing systems capture and interpret the acoustic signal of speech. However, speech is not only limited to acoustics – speech-related activities can be measured at each level of speech processing, including the central and peripheral nervous systems, muscular action potentials, and speech kinematics. Their measurement, obtained through recordings from variety of sensor technologies, results in speech-related “biosignals” that have been studied for decades to better understand the underlying mechanisms of human speech processing. However, there is more: speech-related biosignals have the potential to overcome limitations of traditional acoustic-based systems for spoken communication. Biosignals can be captured before the airborne acoustic signal and are thus less prone to environmental noise. Also, they do not rely on the production of audible speech - both features open up newtracks for “Biosignalbased Spoken Communication”. Examples of these tracks include Brain-Computer Interfaces allowing for communication by directly decoding cortical brain activity into speech representations, and Silent-Speech Interfaces, which offer a way to communicate privately without disturbing bystanders and to restore spoken communication for people who lost their voice due to severe speech impairments. Furthermore, biosignals could provide valuable articulatory biofeedback to speakers about their own voice production for increasing articulatory awareness in speech therapy or language learning.
Tanja Schultz, Thomas Hueber, Dean J. Krusienski, Jonathan S. Brumberg
IEEE ACM Trans. Audio Speech Lang. Process.1
2017 Biosignal-Based Spoken Communication: A Survey
abstract
Speech is a complex process involving a wide range of biosignals, including but not limited to acoustics. These biosignals-stemming from the articulators, the articulator muscle activities, the neural pathways, and the brain itself-can be used to circumvent limitations of conventional speech processing in particular, and to gain insights into the process of speech production in general. Research on biosignal-based speech processing is a wide and very active field at the intersection of various disciplines, ranging from engineering, computer science, electronics and machine learning to medicine, neuroscience, physiology, and psychology. Consequently, a variety of methods and approaches have been used to investigate the common goal of creating biosignal-based speech processing devices for communication applications in everyday situations and for speech rehabilitation, as well as gaining a deeper understanding of spoken communication. This paper gives an overview of the various modalities, research approaches, and objectives for biosignal-based spoken communication.
Tanja Schultz, Michael Wand 0002, Thomas Hueber, Dean J. Krusienski, Christian Herff, Jonathan S. Brumberg
IEEE ACM Trans. Audio Speech Lang. Process.1
2016 Intervention-free selection using EEG and eye tracking
abstract
In this paper, we show how recordings of gaze movements (via eye tracking) and brain activity (via electroencephalography) can be combined to provide an interface for implicit selection in a graphical user interface. This implicit selection works completely without manual intervention by the user. In our approach, we formulate implicit selection as a classification problem, describe the employed features and classification setup and introduce our experimental setup for collecting evaluation data. With a fully online-capable setup, we can achieve an F_0.2-score of up to 0.74 for temporal localization and a spatial localization accuracy of more than 0.95.
Felix Putze, Johannes Popp, Jutta Hild, Jürgen Beyerer, Tanja Schultz
ICMI5
2016 To be Defined
Tanja Schultz
ICPRAM1
2016 Speech-Based Detection of Alzheimer's Disease in Conversational German
abstract
The worldwide population is aging. With a larger population of elderly people, the numbers of people affected by cognitive impairment such as Alzheimer’s disease are growing. Unfortunately, there is no known cure for Alzheimer’s disease. The only way to alleviate it’s serious effects is to start therapy very early before the disease has wrought too much irreversible damage. Current diagnostic procedures are neither cost nor time efficient and therefore do not meet the demands for frequent mass screening required to mitigate the consequences of cognitive impairments on the global scale. We present an experiment to detect Alzheimer’s disease using spontaneous conversational speech. The speech data was recorded during biographic interviews in the Interdisciplinary Longitudinal Study on Adult Development and Aging (ILSE), a large data resource on healthy and satisfying aging in middle adulthood and later life in Germany. From these recordings we extract ten speech-based features using voice activity detection and transcriptions. In an experimental setup with 98 data samples we train a linear discriminant analysis classifier to distinguish subjects with Alzheimer’s disease from the control group. This setup results in an F-score of 0.8 for the detection of Alzheimer’s disease, clearly showing our approach detects dementia well.
Jochen Weiner, Christian Herff, Tanja Schultz
INTERSPEECH3
2016 Towards Automatic Transcription of ILSE ― an Interdisciplinary Longitudinal Study of Adult Development and Aging
Jochen Weiner, Claudia Frankenberg, Dominic Telaar, Britta Wendelstein, Johannes Schröder, Tanja Schultz
LREC6
2016 Word segmentation and pronunciation extraction from phoneme sequences through cross-lingual word-to-phoneme alignment
Felix Stahlberg, Tim Schlippe, Stephan Vogel, Tanja Schultz
Comput. Speech Lang.4
2015 Advancing Muscle-Computer Interfaces with High-Density Electromyography
abstract
In this paper we present our results on using electromyographic (EMG) sensor arrays for finger gesture recognition. Sensing muscle activity allows to capture finger motion without placing sensors directly at the hand or fingers and thus may be used to build unobtrusive body-worn interfaces. We use an electrode array with 192 electrodes to record a high-density EMG of the upper forearm muscles. We present in detail a baseline system for gesture recognition on our dataset, using a naive Bayes classifier to discriminate the 27 gestures. We recorded 25 sessions from 5 subjects. We report an average accuracy of 90% for the within-session scenario, showing the feasibility of the EMG approach to discriminate a large number of subtle gestures. We analyze the effect of the number of used electrodes on the recognition performance and show the benefit of using high numbers of electrodes. Cross-session recognition typically suffers from electrode position changes from session to session. We present two methods to estimate the electrode shift between sessions based on a small amount of calibration data and compare it to a baseline system with no shift compensation. The presented methods raise the accuracy from 59% baseline accuracy to 75% accuracy after shift compensation. The dataset is publicly available.
Christoph Amma, Thomas Krings, Jonas Böer, Tanja Schultz
CHI4
2015 Design and Evaluation of a Self-Correcting Gesture Interface based on Error Potentials from EEG
abstract
Any user interface which automatically interprets the user's input using natural modalities like gestures makes mistakes. System behavior depending on such mistakes will confuse the user and lead to an erroneous interaction flow. The automatic detection of error potentials in electroencephalographic data recorded from a user allows the system to detect such states of confusion and automatically bring the interaction back on track. In this work, we describe the design of such a self-correcting gesture interface, implement different strategies to deal with detected errors, use a simulation approach to analyze performance and costs of those strategies and execute a user study to evaluate user satisfaction. We show that self-correction significantly improves gesture recognition accuracy at lower costs and with higher acceptance than manual correction.
Felix Putze, Christoph Amma, Tanja Schultz
CHI3
2015 Cross-lingual lexical language discovery from audio data using multiple translations
abstract
Zero-resource Automatic Speech Recognition (ZR ASR) addresses target languages without given pronunciation dictionary, transcribed speech, and language model. Lexical discovery for ZR ASR aims to extract word-like chunks from speech. Lexical discovery benefits from the availability of written translations in another source language. In this paper, we improve lexical discovery even more by combining multiple source languages. We present a novel method for combining noisy word segmentations resulting in up to 11.2% relative F-score gain. When we extract word pronunciations from the combined segmentations to bootstrap an ASR system, we improve accuracy by 9.1% relative compared to the best system with only one translation, and by 50.1% compared to monolingual lexical discovery.
Felix Stahlberg, Tim Schlippe, Stephan Vogel, Tanja Schultz
ICASSP4
2015 Direct conversion from facial myoelectric signals to speech using Deep Neural Networks
abstract
This paper presents our first results using Deep Neural Networks for surface electromyographic (EMG) speech synthesis. The proposed approach enables a direct mapping from EMG signals captured from the articulatory muscle movements to the acoustic speech signal. Features are processed from multiple EMG channels and are fed into a feed forward neural network to achieve a mapping to the target acoustic speech output. We show that this approach is feasible to generate speech output from the input EMG signal and compare the results to a prior mapping technique based on Gaussian mixture models. The comparison is conducted via objective Mel-Cepstral distortion scores and subjective listening test evaluations. It shows that the proposed Deep Neural Network approach gives substantial improvements for both evaluation criteria.
Lorenz Diener, Matthias Janke, Tanja Schultz
IJCNN3
2015 Codebook clustering for unit selection based EMG-to-speech conversion
abstract
This paper reports on our recent advances in using Unit Selection to directly synthesize speech from facial surface electromyographic (EMG) signals generated by movement of the articulatory muscles during speech production. We achieve a robust Unit Selection mapping by using a more sophisticated unit codebook. This codebook is generated from a set of base units using a two stage unit clustering process. The units are first clustered based on the audio and afterwards on the EMG feature vectors they cover, and a new codebook is generated using these cluster assignments. We evaluate different cluster counts for both stages and revisit our evaluation of unit sizes in light of this clustering approach. Our final system achieves a significantly better Mel-Cepstral distortion score than the Unit Selection based EMG-to-Speech conversion system from our previous work while, due to the reduced codebook size, taking less time to perform the conversion.
Lorenz Diener, Matthias Janke, Tanja Schultz
INTERSPEECH3
2015 Continuous speech recognition from ECoG
abstract
Continuous speech production is a highly complex process involving many parts of the human brain. To date, no fundamental representation that allows for decoding of continuous speech from neural signals has been presented. Here we show that techniques from automatic speech recognition can be applied to decode a textual representation of spoken words from neural signals. We model phones as the fundamental unit of the speech process in invasively measured brain activity (intracranial electrocorticographic (ECoG)) recordings. These phone models give insights into timings and locations of neural processes associated with the continuous production of speech and can be used in a speech recognizer to decode the neural data into their textual representations. When restricting the dictionary to small subsets, Word Error Rates as low as 25% can be achieved. As the brain activity data sets are fairly small, alternative approaches to Gaussian models are investigated by relying on robust, regularized discriminative models.
Dominic Heger, Christian Herff, Adriana de Pesters, Dominic Telaar, Peter Brunner, Gerwin Schalk, Tanja Schultz
INTERSPEECH7
2015 Telemanipulation with force-based display of proximity fields
abstract
In this paper we show and evaluate the design of a novel telemanipulation system that maps proximity values, acquired inside of a gripper, to forces a user can feel through a haptic input device. The command console is complemented by input-devices that give the user an intuitive control over parameters relevant to the system. Furthermore, proximity sensors enable the autonomous alignment/centering of the gripper to objects in user-selected DoFs with the potential of aiding the user and lowering the workload. We evaluate our approach in a user study that shows that the telemanipulation system benefits from the supplementary proximity information and that the workload can indeed be reduced when the system operates with partial autonomy.
Stefan Escaida Navarro, Franz Heger, Felix Putze, Tim Beyl, Tanja Schultz, Björn Hein
IROS5
2015 Model-Based Evaluation of Playing Strategies in a Memo Game for Elderly Users
abstract
In this paper, we analyze game protocols for a Memo game for elderly users. We show how we can use generative statistical models to automatically reveal different playing strategies. We present a quantitative and qualitative evaluation of the approach on simulated and real data. We show that we can reliably detect different strategies and that we can use those strategy profiles to uncover relevant information on the players beyond pure performance measures.
Felix Putze, Tanja Schultz, Sonja Ehret, Heike Miller-Teynor, Andreas Kruse
SMC2
2015 Dummy Model Based Workload Modeling
abstract
In this paper, we show how a model of human cognition based on ACT-R can be improved to accurately predict cognitive performance under different workload levels. For this purpose, we propose a novel approach which uses an EEG-based workload model to (de-)activate a dummy model which runs in parallel to the actual task model. The dummy model consumes cognitive resources to reflect the effect of workload on behavior and performance. We evaluate the approach in two user studies with different tasks and show a significant reduction of prediction error.
Felix Putze, Tanja Schultz, Robert Pröpper
SMC2
2015 Syntactic and Semantic Features For Code-Switching Factored Language Models
abstract
This paper presents our latest investigations on different features for factored language models for Code-Switching speech and their effect on automatic speech recognition (ASR) performance. We focus on syntactic and semantic features which can be extracted from Code-Switching text data and integrate them into factored language models. Different possible factors, such as words, part-of-speech tags, Brown word clusters, open class words and clusters of open class word embeddings are explored. The experimental results reveal that Brown word clusters, part-of-speech tags and open-class words are the most effective at reducing the perplexity of factored language models on the Mandarin-English Code-Switching corpus SEAME. In ASR experiments, the model containing Brown word clusters and part-of-speech tags and the model also including clusters of open class word embeddings yield the best mixed error rate results. In summary, the best language model can significantly reduce the perplexity on the SEAME evaluation set by up to 10.8% relative and the mixed error rate by up to 3.4% relative.
Heike Adel, Ngoc Thang Vu, Katrin Kirchhoff, Dominic Telaar, Tanja Schultz
IEEE ACM Trans. Audio Speech Lang. Process.5
2014 Model-Based Identification of EEG Markers for Learning Opportunities in an Associative Learning Task with Delayed Feedback
Felix Putze, Daniel V. Holt, Tanja Schultz, Joachim Funke
ICANN3
2014 Connectivity based feature-level filtering for single-trial EEG BCIs
abstract
EEG-based Brain Computer interfaces (BCIs) often rely on power spectral density features to represent relevant aspects of brain activity. The information flow within human brain networks and the corresponding connectivity patterns may contain useful information to improve BCI performance, however they are typically not leveraged in current systems. In this paper, analyzes of information flow between independent sources of brain activity have been incorporated into the feature extraction stage of a BCI. For this purpose, connectivity measures based on multivariate autoregressive models have been estimated and are applied as filters to power spectral density based features. Two publicly available data sets have been used to evaluate the proposed feature extraction method: a two-back task and a motor imagery task. The results demonstrate significant performance improvements of the proposed method over band-power features and indicate that connectivity in brain networks can be used as powerful feature-level filters for BCIs.
Dominic Heger, Emiliyana Terziyska, Tanja Schultz
ICASSP3
2014 Fundamental frequency generation for whisper-to-audible speech conversion
abstract
In this work, we address the issues involved in whisper-to-audible speech conversion. Spectral mapping techniques using Gaussian mixture models or Artificial Neural Networks borrowed from voice conversion have been applied to transform whisper spectral features to normally phonated audible speech. However, the modeling and generation of fundamental frequency (F0) and its contour in the converted speech is a major issue. Whispered speech does not contain explicit voicing characteristics and hence it is hard to derive a suitable F0, making it difficult to generate a natural prosody after conversion. Our work addresses the F0 modeling in whisper-to-speech conversion. We show that F0 contours can be derived from the mapped spectral vectors, which can be used for the synthesis of a speech signal. We also present a hybrid unit selection approach for whisper-to-speech conversion. Unit selection is performed on the spectral vectors, where F0 and its contour can be obtained as a byproduct without any additional modeling.
Matthias Janke, Michael Wand 0002, Till Heistermann, Tanja Schultz, K. Prahallad
ICASSP4
2014 Multilingual deep neural network based acoustic modeling for rapid language adaptation
abstract
This paper presents a study on multilingual deep neural network (DNN) based acoustic modeling and its application to new languages. We investigate the effect of phone merging on multilingual DNN in context of rapid language adaptation. Moreover, the combination of multilingual DNNs with Kullback-Leibler divergence based acoustic modeling (KL-HMM) is explored. Using ten different languages from the Globalphone database, our studies reveal that crosslingual acoustic model transfer through multilingual DNNs is superior to unsupervised RBM pre-training and greedy layer-wise supervised training. We also found that KL-HMM based decoding consistently outperforms conventional hybrid decoding, especially in low-resource scenarios. Furthermore, the experiments indicate that multilingual DNN training equally benefits from simple phoneset concatenation and manually derived universal phonesets.
Ngoc Thang Vu, David Imseng, Daniel Povey, Petr Motlícek, Tanja Schultz, Hervé Bourlard
ICASSP5
2014 Compensation of recording position shifts for a myoelectric Silent Speech Recognizer
abstract
A myoelectric Silent Speech Recognizer is a system which recognizes speech by capturing the electrical activity of the human articulatory muscles, thus enabling the user to communicate silently. We recently devised a recording setup based on electrode arrays with multiple measuring points. In this study we show that this allows to compensate for shifts of the recording position, which happen when the array is removed and reattached between system training and application. We present a method which determines the amount of recording position shift; compensation is performed by linear interpolation. We evaluate our method by running recognition experiments across recording sessions and obtain a Word Error Rate improvement of 14.3% relative on the development set and 12.9% relative on the evaluation set, compared to using classical session adaptation.
Michael Wand 0002, Christopher Schulte, Matthias Janke, Tanja Schultz
ICASSP4
2014 Investigating Intrusiveness of Workload Adaptation
abstract
In this paper, we investigate how an automatic task assistant which can detect and react to a user's workload level is able to support the user in a complex, dynamic task. In a user study, we design a dispatcher scenario with low and high workload conditions and compare the effect of four support strategies with different levels of intrusiveness using objective and subjective metrics. We see that a more intrusive strategy results in higher efficiency and effectiveness, but is also less accepted by the participants. We also show that the benefit of supportive behavior depends on the user's workload level, i.e. adaptation to its changes are necessary. We describe and evaluate a Brain Computer Interface that is able to provide the necessary user state detection.
Felix Putze, Tanja Schultz
ICMI2
2014 Comparing approaches to convert recurrent neural networks into backoff language models for efficient decoding
abstract
In this paper, we investigate and compare three different possibilities to convert recurrent neural network language models (RNNLMs) into backoff language models (BNLM). While RNNLMs often outperform traditional n-gram approaches in the task of language modeling, their computational demands make them unsuitable for an efficient usage during decoding in an LVCSR system. It is, therefore, of interest to convert them into BNLMs in order to integrate their information into the decoding process. This paper compares three different approaches: a text based conversion, a probability based conversion and an iterative conversion. The resulting language models are evaluated in terms of perplexity and mixed error rate in the context of the Code-Switching data corpus SEAME. Although the best results are obtained by combining the results of all three approaches, the text based conversion approach alone leads to significant improvements on the SEAME corpus as well while offering the highest computational efficiency. In total, the perplexity can be reduced by 11.4% relative on the evaluation set and the mixed error rate by 3.0% relative on the same data set.
Heike Adel, Katrin Kirchhoff, Ngoc Thang Vu, Dominic Telaar, Tanja Schultz
INTERSPEECH5
2014 Combining recurrent neural networks and factored language models during decoding of code-Switching speech
abstract
In this paper, we present our latest investigations of language modeling for Code-Switching. Since there is only little text material for Code-Switching speech available, we integrate syntactic and semantic features into the language modeling process. In particular, we use part-of-speech tags, language identifiers, Brown word clusters and clusters of open class words. We develop factored language models and convert recurrent neural network language models into backoff language models for an efficient usage during decoding. A detailed error analysis reveals the strengths and weaknesses of the different language models. When we interpolate the models linearly, we reduce the perplexity by 15.6% relative on the SEAME evaluation set. This is even slightly better than the result of the unconverted recurrent neural network. We also combine the language models during decoding and obtain a mixed error rate reduction of 4.4% relative on the SEAME evaluation set.
Heike Adel, Dominic Telaar, Ngoc Thang Vu, Katrin Kirchhoff, Tanja Schultz
INTERSPEECH5
2014 Methods for efficient semi-automatic pronunciation dictionary bootstrapping
abstract
In this paper we propose efficient methods which contribute to a rapid and economic semi-automatic pronunciation dictionary development and evaluate them on English, German, Spanish, Vietnamese, Swahili, and Haitian Creole. First we determine optimal strategies for the word selection and the period for the grapheme-to-phoneme model retraining. In addition to the traditional concatenation of single phonemes most commonly associated with each grapheme, we show that web-derived pronunciations and cross-ligual grapheme-to-phoneme models can help to reduce the initial editing effort. Furthermore we show that our phoneme-level combination of the output of multiple grapheme-to-phoneme converters reduces the editing effort more than the best single converters. Totally, we report on average 15% relative editing effort reduction with our phonemelevel combination compared to conventional methods. An additional reduction of 6% relative was possible by including initial pronunciations from Wiktionary for English, German, and Spanish.
Tim Schlippe, Matthias Merz, Tanja Schultz
INTERSPEECH3
2014 BioKIT - real-time decoder for biosignal processing
abstract
We introduce BioKIT, a new Hidden Markov Model based toolkit to preprocess, model and interpret biosignals such as speech, motion, muscle and brain activities. The focus of this toolkit is to enable researchers from various communities to pursue their experiments and integrate real-time biosignal interpretation into their applications. BioKIT boosts a flexible two-layer structure with a modular C++ core that interfaces with a Python scripting layer, to facilitate development of new applications. BioKIT employs sequence-level parallelization and memory sharing across threads. Additionally, a fully integrated error blaming component facilitates in-depth analysis. A generic terminology keeps the barrier to entry for researchers from multiple fields to a minimum. We describe our onlinecapable dynamic decoder and report on initial experiments on three different tasks. The presented speech recognition experiments employ Kaldi [1] trained deep neural networks with the results set in relation to the real time factor needed to obtain them.
Dominic Telaar, Michael Wand 0002, Dirk Gehrig, Felix Putze, Christoph Amma, Dominic Heger, Ngoc Thang Vu, Mark Erhardt, Tim Schlippe, Matthias Janke, Christian Herff, Tanja Schultz
INTERSPEECH12
2014 Improving ASR performance on non-native speech using multilingual and crosslingual information
abstract
This paper presents our latest investigation of automatic speech recognition (ASR) on non-native speech. We first report on a non-native speech corpus an extension of the GlobalPhone database which contains English with Bulgarian, Chinese, German and Indian accent and German with Chinese accent. In this case, English is the spoken language (L2) and Bulgarian, Chinese, German and Indian are the mother tongues (L1) of the speakers. Afterwards, we investigate the effect of multilingual acoustic modeling on non-native speech. Our results reveal that a bilingual L1-L2 acoustic model significantly improves the ASR performance on non-native speech. For the case that L1 is unknown or L1 data is not available, a multilingual ASR system trained without L1 speech data consistently outperforms the monolingual L2 ASR system. Finally, we propose a method called crosslingual accent adaptation, which allows using English with Chinese accent to improve the German ASR on German with Chinese accent and vice versa. Without using any intra lingual adaptation data, we achieve 15.8% relative improvement in average over the baseline system.
Ngoc Thang Vu, Yuanfan Wang, Marten Klose, Zlatka Mihaylova, Tanja Schultz
INTERSPEECH5
2014 Investigating the learning effect of multilingual bottle-neck features for ASR
abstract
Deep neural networks (DNNs) have become state-of-the-art techniques of automatic speech recognition in the last few years. They can be used at the preprocessing level (Tandem or Bottle-Neck features) or at the acoustic model level (hybrid Hidden Markov Model/DNN). Moreover, they allow exploiting multilingual data to improve monolingual systems. This paper presents our investigation of the learning effect of neural networks in the context of multilingual Bottle-Neck features. For this, we perform a visual analysis of the output of the Bottle-Neck layer of a neural network using t-Distributed Stochastic Neighbor Embedding. Our results show that multilingual Bottle-Neck features seem to learn phoneme characteristics, such as the F1 and F2 formants which characterize different vowels, and other articulatory features, such as fricatives and nasals which characterize consonants. Furthermore, they seem to normalize language dependent variations and transfer the learned representation to unseen languages.
Ngoc Thang Vu, Jochen Weiner, Tanja Schultz
INTERSPEECH3
2014 The EMG-UKA corpus for electromyographic speech processing
abstract
This article gives an overview of the EMG-UKA corpus, a corpus of electromyographic (EMG) recordings of articulatory activity enabling speech processing (in particular speech recognition and synthesis) based on EMG signals, with the purpose of building Silent Speech interfaces. Data is available in multiple speaking modes, namely audibly spoken, whispered, and silently articulated speech. Besides the EMG data, synchronous acoustic data was additionally recorded to serve as a reference. The corpus comprises 63 recorded sessions from 8 speakers, the total amount of data is 7:32 hours. A trial subset, consisting of 1:52 hours of data, is freely available for download.
Michael Wand 0002, Matthias Janke, Tanja Schultz
INTERSPEECH3
2014 Towards real-life application of EMG-based speech recognition by using unsupervised adaptation
abstract
This paper deals with a Silent Speech Interface based on Surface Electromyography (EMG), where electrodes capture the electric activity generated by the articulatory muscles from a user’s face in order to decode the underlying speech, allowing speech to be recognized even when no sound is heard or created. So far, most EMG-based speech recognizers described in literature do not allow electrode reattachment between system training and usage, which we consider unsuitable for practical applications. In this study we report on our research on unsupervised session adaptation: A system is pre-trained with data from multiple recording sessions and then adapted towards the current recording session using data accruable during normal use, without requiring a time-consuming specific enrollment phase. We show that considerable accuracy improvements can be achieved with this method, paving the way towards real-life applications of the technology. Index Terms: Silent Speech Interfaces, EMG, EMG-based Speech Recognition, Unsupervised Adaptation
Michael Wand 0002, Tanja Schultz
INTERSPEECH2
2014 Conversion from facial myoelectric signals to speech: a unit selection approach
abstract
This paper reports on our recent research on surface electromyographic (EMG) speech synthesis: a direct conversion of the EMG signals of the articulatory muscle movements to the acoustic speech signal. In this work we introduce a unit selection approach which compares segments of the input EMG signal to a database of simultaneously recorded EMG/audio unit pairs and selects the best matching audio unit based on target and concatenation cost, which will be concatenated to synthesize an acoustic speech output. We show that this approach is feasible to generate a proper speech output from the input EMG signal. We evaluate different properties of the units and investigate what amount of data is necessary for an initial transformation. Prior work on EMG-to-speech conversion used a framebased approach from the voice conversion domain, which struggles with the generation of a natural $F_0$ contour. This problem may also be tackled by our unit selection approach.
Marlene Zahner, Matthias Janke, Michael Wand 0002, Tanja Schultz
INTERSPEECH4
2014 GlobalPhone: Pronunciation Dictionaries in 20 Languages
Tanja Schultz, Tim Schlippe
LREC1
2014 Airwriting: a wearable handwriting recognition system
Christoph Amma, Marcus Georgi, Tanja Schultz
Pers. Ubiquitous Comput.3
2014 Introduction to the special issue on processing under-resourced languages
Laurent Besacier, Etienne Barnard, Alexey Karpov 0001, Tanja Schultz
Speech Commun.4
2014 Automatic speech recognition for under-resourced languages: A survey
Laurent Besacier, Etienne Barnard, Alexey Karpov 0001, Tanja Schultz
Speech Commun.4
2014 Web-based tools and methods for rapid pronunciation dictionary creation
Tim Schlippe, Sebastian Ochs 0002, Tanja Schultz
Speech Commun.3
2013 Continuous Recognition of Affective States by Functional Near Infrared Spectroscopy Signals
abstract
Functional near infrared spectroscopy (fNIRS) is becoming more and more popular as an innovative imaging modality for brain computer interfaces. A continuous (i.e. asynchronous) affective state monitoring system using fNIRS signals would be highly relevant for numerous disciplines, including adaptive user interfaces, entertainment, biofeedback, and medical applications. However, only stimulus-locked emotion recognition systems have been proposed by now. fNRIS signals of eight subjects at eight prefrontal locations have been recorded in response to three different classes of affect induction by emotional audio-visual stimuli and a neutral class. Our system evaluates short windows of five seconds length to continuously recognize affective states. We analyze hemodynamic responses, present a careful evaluation of binary classification tasks and investigate classification accuracies over the time.
Dominic Heger, Reinhard Mutter, Christian Herff, Felix Putze, Tanja Schultz
ACII5
2013 Neighbour selection and adaptation for rapid speaker-dependent ASR
abstract
Speaker dependent (SD) ASR systems have significantly lower word error rates (WER) compared to speaker independent (SI) systems. However, SD systems require sufficient training data from the target speaker, which is impractical to collect in a short time. We present a technique for training SD models using just few minutes of speaker's data. We compensate for the lack of adequate speaker-specific data by selecting neighbours from a database of existing speakers who are acoustically close to the target speaker. These neighbours provide ample training data, which is used to adapt the SI model to obtain an initial SD model for the new speaker with significantly lower WER. We evaluate various neighbour selection algorithms on a large-scale medical transcription task and report significant reduction in WER using only 5 mins of speaker-specific data. We conduct a detailed analysis of various factors such as gender and accent in the neighbour selection. Finally, we study neighbour selection and adaptation in the context of discriminative objective functions.
Udhyakumar Nallasamy, Mark C. Fuhs, Monika Woszczyna, Florian Metze, Tanja Schultz
ASRU5
2013 Recurrent neural network language modeling for code switching conversational speech
abstract
Code-switching is a very common phenomenon in multilingual communities. In this paper, we investigate language modeling for conversational Mandarin-English code-switching (CS) speech recognition. First, we investigate the prediction of code switches based on textual features with focus on Part-of-Speech (POS) tags and trigger words. Second, we propose a structure of recurrent neural networks to predict code-switches. We extend the networks by adding POS information to the input layer and by factorizing the output layer into languages. The resulting models are applied to our task of code-switching language modeling. The final performance shows 10.8% relative improvement in perplexity on the SEAME development set which transforms into a 2% relative improvement in terms of Mixed Error Rate and a relative improvement of 16.9% in perplexity on the evaluation set which leads to a 2.7% relative improvement of MER.
Heike Adel, Ngoc Thang Vu, Franziska Kraus, Tim Schlippe, Haizhou Li 0001, Tanja Schultz
ICASSP6
2013 Rapid bootstrapping of a Ukrainian large vocabulary continuous speech recognition system
abstract
We report on our efforts toward an LVCSR system for the Slavic language Ukrainian. We describe the Ukrainian text and speech database recently collected as a part of our GlobalPhone corpus [1] with our Rapid Language Adaptation Toolkit [2]. The data was complemented by a large collection of text data crawled from various Ukrainian websites. For the production of the pronunciation dictionary, we investigate strategies using grapheme-to-phoneme (g2p) models derived from existing dictionaries of other languages, thereby reducing severely the necessary manual effort. Russian and Bulgarian g2p models even decrease the number of pronunciation rules to one fifth. We achieve significant improvement by applying state-of-the art techniques for acoustic modeling and our day-wise text collection and language model interpolation strategy [3]. Our best system achieves a word error rate of 11.21% on the test set on read newspaper speech.
Tim Schlippe, Mykola Volovyk, Kateryna Yurchenko, Tanja Schultz
ICASSP4
2013 Statistical machine translation based text normalization with crowdsourcing
abstract
In [1], we have proposed systems for text normalization based on statistical machine translation (SMT) methods which are constructed with the support of Internet users and evaluated those with French texts. Internet users normalize text displayed in a web interface in an annotation process, thereby providing a parallel corpus of normalized and non-normalized text. With this corpus, SMT models are generated to translate non-normalized into normalized text. In this paper, we analyze their efficiency for other languages. Additionally, we embedded the English annotation process for training data in Amazon Mechanical Turk and compare the quality of texts thoroughly annotated in our lab to those annotated by the Turkers. Finally, we investigate how to reduce the user effort by iteratively applying an SMT system to the next sentences to be edited, built from the sentences which have been annotated so far.
Tim Schlippe, Chenfei Zhu, Daniel Lemcke, Tanja Schultz
ICASSP4
2013 GlobalPhone: A multilingual text & speech database in 20 languages
abstract
This paper describes the advances in the multilingual text and speech database GlobalPhone, a multilingual database of high-quality read speech with corresponding transcriptions and pronunciation dictionaries in 20 languages. GlobalPhone was designed to be uniform across languages with respect to the amount of data, speech quality, the collection scenario, the transcription and phone set conventions. With more than 400 hours of transcribed audio data from more than 2000 native speakers GlobalPhone supplies an excellent basis for research in the areas of multilingual speech recognition, rapid deployment of speech processing systems to yet unsupported languages, language identification tasks, speaker recognition in multiple languages, multilingual speech synthesis, as well as monolingual speech recognition in a large variety of languages.
Tanja Schultz, Ngoc Thang Vu, Tim Schlippe
ICASSP1
2013 Experiments towards a better LVCSR system for tamil
abstract
This paper summarizes our latest efforts in the development of a Large Vocabulary Continuous Speech Recognition (LVCSR) system for Tamil at different levels: pronunciation dictionary, language modeling (LM) and front-end. Usually in Tamil there are not many word-pronunciation pairs to train data-driven grapheme-to-phoneme (G2P) converters. Therefore, we explore the correlation between the amount of training data and the performance of the grapheme-to-phoneme (G2P) conversion. To address the morphological complexity of Tamil, we investigate different levels of morphemes for language modeling including a comparison between our Dictionary Unit Merging Algorithm (DUMA) and Morfessor, followed by various experiments on hybrid systems using word and morpheme LMs. Finally, we integrate our multilingual bottle-neck features framework with Tamil LVCSR. The final best system produced 21.34% Syllable Error Rate (SyllER) on our Tamil test set.
Melvin Jose Johnson Premkumar, Ngoc Thang Vu, Tanja Schultz
INTERSPEECH3
2013 Unsupervised language model adaptation for automatic speech recognition of broadcast news using web 2.0
abstract
We improve the automatic speech recognition of broadcast news using paradigms from Web 2.0 to obtain timeand topicrelevant text data for language modeling. We elaborate an unsupervised text collection and decoding strategy that includes crawling appropriate texts from RSS Feeds, complementing it with texts from Twitter, language model and vocabulary adaptation, as well as a 2-pass decoding. The word error rates of the tested French broadcast news shows from Europe 1 are reduced by almost 32% relative with an underlying language model from the GlobalPhone project [1] and by almost 4% with an underlying language model from the Quaero project. The tools that we use for the text normalization, the collection of RSS Feeds together with the text on the related websites, a TF-IDF-based topic words extraction, as well as the opportunity for language model interpolation are available in our Rapid Language Adaptation Toolkit [2] [3].
Tim Schlippe, Lukasz Gren, Ngoc Thang Vu, Tanja Schultz
INTERSPEECH4
2013 Multilingual multilayer perceptron for rapid language adaptation between and across language families
abstract
In this paper, we present our latest investigations of multilingual Multilayer Perceptrons (MLPs) for rapid language adaptation between and across language families. We explore the impact of the amount of languages and data used for the multilingual MLP training process. We show that the overall system performance on the target language is significantly improved by initializing it with a multilingual MLP. Our experiments indicate that the more languages we use to train a multilingual MLP, the better is the initialization for MLP training. As a result, the ASR performance is improved, even if the target language and the source languages are not in the same language family. Our best results show an error rate improvement of up to 22.9% relative for different target languages (Czech, Hausa and Vietnamese) by using a multilingual MLP which has been trained with many different languages from the GlobalPhone corpus. In the case of very few training or adaptation data, an improvement of up to 24% relative in terms of error rate is observed.
Ngoc Thang Vu, Tanja Schultz
INTERSPEECH2
2013 Locating user attention using eye tracking and EEG for spatio-temporal event selection
abstract
In expert video analysis, the selection of certain events in a continuous video stream is a frequently occurring operation, e.g., in surveillance applications. Due to the dynamic and rich visual input, the constantly high attention and the required hand-eye coordination for mouse interaction, this is a very demanding and exhausting task. Hence, relevant events might be missed. We propose to use eye tracking and electroencephalography (EEG) as additional input modalities for event selection. From eye tracking, we derive the spatial location of a perceived event and from patterns in the EEG signal we derive its temporal location within the video stream. This reduces the amount of the required active user input in the selection process, and thus has the potential to reduce the user's workload. In this paper, we describe the employed methods for the localization processes and introduce the developed scenario in which we investigate the feasibility of this approach. Finally, we present and discuss results on the accuracy and the speed of the method and investigate how the modalities interact.
Felix Putze, Jutta Hild, Rainer Kärgel, Christian Herff, Alexander Redmann, Jürgen Beyerer, Tanja Schultz
IUI7
2012 Towards single pass discriminative training for speech recognition
abstract
This paper describes how we can combine our previously proposed fast extended Baum-Welch algorithm and generalized discriminative feature transformation to achieve single pass discriminative training, which we only process the data once. Compared to the state of the art training procedure, which uses feature space maximum mutual information (fMMI) and boosted maximum mutual information (BMMI), our proposed training procedure can achieve around 80% of the improvement available from discriminative training. We also show that if we are allowed to process the data twice, it is possible to achieve almost all of the improvement. We evaluate different training procedures on various large scale tasks using Iraqi and modern standard Arabic speech recognition systems.
Roger Hsiao, Tanja Schultz
ICASSP2
2012 Further investigations on EMG-to-speech conversion
abstract
Our study deals with a Silent Speech Interface based on mapping surface electromyographic (EMG) signals to speech waveforms. Electromyographic signals recorded from the facial muscles capture the activity of the human articulatory apparatus and therefore allow to retrace speech, even when no audible signal is produced. The mapping of EMG signals to speech is done via a Gaussian mixture model (GMM)-based conversion technique. In this paper, we follow the lead of EMG-based speech-to-text systems and apply two major recent technological advances to our system, namely, we consider session-independent systems, which are robust against electrode repositioning, and we show that mapping the EMG signal to whispered speech creates a better speech signal than a mapping to normally spoken speech. We objectively evaluate the performance of our systems using a spectral distortion measure.
Matthias Janke, Michael Wand 0002, Keigo Nakamura, Tanja Schultz
ICASSP4
2012 Grapheme-to-phoneme model generation for Indo-European languages
abstract
In this paper, we evaluate grapheme-to-phoneme (g2p) models among languages and of different quality. We created g2p models for Indo-European languages with word-pronunciation pairs from the GlobalPhone project and from Wiktionary [1]. Then we checked their quality in terms of consistency and complexity as well as their impact on Czech, English, French, Spanish, Polish, and German ASR. While the GlobalPhone dictionaries were manually cross-checked and have been used successfully in LVCSR, Wiktionary pronunciations have been provided by the Internet community and can be used to rapidely and economically create pronunciation dictionaries for new languages and domains.
Tim Schlippe, Sebastian Ochs 0002, Tanja Schultz
ICASSP3
2012 A first speech recognition system for Mandarin-English code-switch conversational speech
abstract
This paper presents first steps toward a large vocabulary continuous speech recognition system (LVCSR) for conversational Mandarin-English code-switching (CS) speech. We applied state-of-the-art techniques such as speaker adaptive and discriminative training to build the first baseline system on the SEAME corpus [1] (South East Asia Mandarin-English). For acoustic modeling, we applied different phone merging approaches based on the International Phonetic Alphabet (IPA) and Bhattacharyya distance in combination with discriminative training to improve accuracy. On language model level, we investigated statistical machine translation (SMT) - based text generation approaches for building code-switching language models. Furthermore, we integrated the provided information from a language identification system (LID) into the decoding process by using a multi-stream approach. Our best 2-pass system achieves a Mixed Error Rate (MER) of 36.6% on the SEAME development set.
Ngoc Thang Vu, Dau-Cheng Lyu, Jochen Weiner, Dominic Telaar, Tim Schlippe, Fabian Blaicher, Chng Eng Siong, Tanja Schultz, Haizhou Li 0001
ICASSP8
2012 Modeling gender dependency in the Subspace GMM framework
abstract
The Subspace GMM acoustic model has both globally shared parameters and parameters specific to acoustic states, and this makes it possible to do various kinds of tying. In the past we have investigated sharing the global parameters among systems with distinct acoustic states; this can be useful in a multilingual setting. In the current paper we investigate the reverse idea: to have different global parameters for different acoustic conditions (gender, in this case) while sharing the acoustic-state-specific parameters. We experiment with modeling gender dependency in this way, and show Word Error Rate improvements on a range of tasks and comparable results to the Vocal Tract Length Normalization (VTLN)-like technique Exponential Transform (ET).
Ngoc Thang Vu, Tanja Schultz, Daniel Povey
ICASSP2
2012 Vision-based handwriting recognition for unrestricted text input in mid-air
abstract
We propose a vision-based system that recognizes handwriting in mid-air. The system does not depend on sensors or markers attached to the users and allows unrestricted character and word input from any position. It is the result of combining handwriting recognition based on Hidden Markov Models with multi-camera 3D hand tracking. We evaluated the system for both quantitative and qualitative aspects. The system achieves recognition rates of 86.15% for character and 97.54% for small-vocabulary isolated word recognition. Limitations are due to slow and low-resolution cameras or physical strain. Overall, the proposed handwriting recognition system provides an easy-to-use and accurate text input modality without placing restrictions on the users.
Alexander Schick, Daniel Morlock, Christoph Amma, Tanja Schultz, Rainer Stiefelhagen
ICMI4
2012 Cross-Subject Classification of Speaking Modes Using fNIRS
Christian Herff, Dominic Heger, Felix Putze, Cuntai Guan, Tanja Schultz
ICONIP (2)5
2012 Enhanced Polyphone Decision Tree Adaptation for Accented Speech Recognition
abstract
State-of-the-art Automatic Speech Recognition (ASR) models struggle to handle accented speech, particularly if the target accent is under-represented in the training data. The acoustic variations presented by an unfamiliar accent, render the ASR polyphone decision tree (PDT) and its associated Gaussian mixture models (GMM) misfit to the test data. In this paper, we improve on the previous work of adapting the polyphone decision tree, using a semi-continuous model based approach to address the problem of data sparsity. We extend the existing PDT to introduce additional states with shared parameters, corresponding to the new contextual variations identified in the adaptation data, while still robustly estimating the state based parameters on a small adaptation set. We conduct ASR experiments on Arabic and English accents and show that our technique performs better than Maximum A-Posteriori (MAP) adaptation and a previous implementation of polyphone decision tree specialization (PDTS). Compared to MAP adaptation, we obtain 7% relative improvement for Dialectal Arabic and 13.8% relative improvement for Accented English.
Udhyakumar Nallasamy, Florian Metze, Tanja Schultz
INTERSPEECH3
2012 Automatic Error Recovery for Pronunciation Dictionaries
abstract
In this paper, we present our latest investigations on pronunciation modeling and its impact on ASR. We propose completely automatic methods to detect, remove, and substitute inconsistent or flawed entries in pronunciation dictionaries. The experiments were conducted on different tasks, namely (1) word-pronunciation pairs from the Czech, English, French, German, Polish, and Spanish Wiktionary [1], a multilingual wiki-based open content dictionary, (2) our GlobalPhone Hausa pronunciation dictionary [2], and (3) pronunciations to complement our Mandarin-English SEAME code-switch dictionary [3]. In the final results, we fairly observed on average an improvement of 2.0% relative in terms of word error rate and even 27.3% for the case of English Wiktionary word-pronunciation pairs.
Tim Schlippe, Sebastian Ochs 0002, Ngoc Thang Vu, Tanja Schultz
INTERSPEECH4
2012 Initialization Schemes for Multilayer Perceptron Training and their Impact on ASR Performance using Multilingual Data
Ngoc Thang Vu, Wojtek Breiter, Florian Metze, Tanja Schultz
INTERSPEECH4
2012 Airwriting: demonstrating mobile text input by 3D-space handwriting
abstract
We demonstrate our airwriting interface for mobile hands-free text entry. The interface enables a user to input text into a computer by writing in the air like on an imaginary blackboard. Hand motion is measured by an accelerometer and a gyroscope attached to the back of the hand and data is sent wirelessly to the processing computer. The system can continuously recognize arbitrary sentences based on a predefined vocabulary in real-time. The recognizer uses Hidden Markov Models (HMM) together with a statistical language model. We achieve a user-independent word error rate of 11% for a 8K vocabulary based on an experiment with nine users.
Christoph Amma, Tanja Schultz
IUI2
2012 Active learning for accent adaptation in Automatic Speech Recognition
abstract
We experiment with active learning for speech recognition in the context of accent adaptation. We adapt a source recognizer on the target accent by selecting a relatively small, matched subset of utterances from a large, untranscribed and multi-accented corpus for human transcription. Traditionally, active learning in speech recognition has relied on uncertainty based sampling to choose the most informative data for manual labeling. Such an approach doesn't include explicit relevance criterion during data selection, which is crucial for choosing utterances to match the target accent, from datasets with wide-ranging speakers of different accents. We formulate a cross-entropy based relevance measure to complement uncertainty based sampling for active learning to aid accent adaptation. We evaluate the algorithm on two different setups for Arabic and English accents and show that our approach performs favorably to conventional data selection. We analyze the results to show the effectiveness of our approach in finding the most relevant subset of utterances for improving the speech recognizer on the target accent.
Udhyakumar Nallasamy, Florian Metze, Tanja Schultz
SLT3
2012 Word segmentation through cross-lingual word-to-phoneme alignment
abstract
We present our new alignment model Model 3P for cross-lingual word-to-phoneme alignment, and show that unsupervised learning of word segmentation is more accurate when information of another language is used. Word segmentation with cross-lingual information is highly relevant to bootstrap pronunciation dictionaries from audio data for Automatic Speech Recognition, bypass the written form in Speech-to-Speech Translation or build the vocabulary of an unseen language, particularly in the context of under-resourced languages. Using Model 3P for the alignment between English words and Spanish phonemes outperforms a state-of-the-art monolingual word segmentation approach [1] on the BTEC corpus [2] by up to 42% absolute in F-Score on the phoneme level and a GIZA++ alignment based on IBM Model 3 by up to 17%.
Felix Stahlberg, Tim Schlippe, Stephan Vogel, Tanja Schultz
SLT4
2011 Online Recognition of Facial Actions for Natural EEG-Based BCI Applications
Dominic Heger, Felix Putze, Tanja Schultz
ACII (2)3
2011 Estimation of fundamental frequency from surface electromyographic data: EMG-to-F0
abstract
In this paper, we present our recent studies of F0estimation from the surface electromyographic (EMG) data us ing a Gaussian mixture model (GMM)-based voice con version (VC) technique, referred to as EMG-to-F0. In our approach, a support vector machine recognizes individual frames as unvoiced and voiced (U/V), and voiced F0contours are discriminated by the trained GMM based on the manner of minimum mean-square error. EMG-to-F0is experimentally evaluated using three data sets of different speakers. Each data set includes almost 500 utterances. Objective experiments demonstrate that we achieve a correlation coefficient of up to 0.49 between estimated and target F0contours with more than 84% U/V decision accuracy, although the results have large variations.
Keigo Nakamura, Matthias Janke, Michael Wand 0002, Tanja Schultz
ICASSP4
2011 Cross-language bootstrapping based on completely unsupervised training using multilingual A-stabil
abstract
This paper presents our work on rapid language adaptation of acoustic models based on multilingual cross-language bootstrapping and unsupervised training. We used Automatic Speech Recognition (ASR) systems in English, French, German, and Spanish to build a Czech ASR system from scratch. System building was performed without using any transcribed audio data by applying three consecutive steps, i.e. cross-language transfer, unsupervised training based on the "multilingual A-stabil" confidence score, and boot strapping. Based on the confidence score we selected 72% (16.6 hours) of the available audio data with a transcription WER of less than 14.5%. The cross-language bootstrap achieves a word error rate of 23.3% on the Czech development set and 22.4% on the evaluation set. These results are very promising as the performance compares favorably to the Czech ASR system which was trained on 23 hours of manually transcribed data (21.8% on the development set and 21.3% on the evaluation set).
Ngoc Thang Vu, Franziska Kraus, Tanja Schultz
ICASSP3
2011 Analysis of phone confusion in EMG-based speech recognition
abstract
In this paper we present a study on phone confusabilities based on phone recognition experiments from facial surface electromyographic (EMG) signals. In our study EMG captures the electrical potentials of the human articulatory muscles. This technology can be used to create Silent Speech Interfaces, where a user can communicate naturally without uttering any sound. This paper investigates to which extent different phone properties can be recognized from an EMG signal, shows which weaknesses have yet to be overcome, and compares the results to acoustic-based recognition of phones.
Michael Wand 0002, Tanja Schultz
ICASSP2
2011 Multimodal person independent recognition of workload related biosignal patterns
abstract
This paper presents an online multimodal person independent workload classification system using blood volume pressure, respiration measures, electrodermal activity and electroencephalography. For each modality a classifier based on linear discriminant analysis is trained. The classification results obtained on short data frames are fused using weighted majority voting. The system was trained and evaluated on a large training corpus of 152 participants, exposed to controlled and uncontrolled scenarios for inducing workload, including a driving task conducted in a realistic driving simulator. Using person dependent feature space normalization, we achieve a classification accuracy of up to 94% for discrimination of relaxed state vs. high workload.
Jan-Philip Jarvis, Felix Putze, Dominic Heger, Tanja Schultz
ICMI4
2011 Impact of Different Feedback Mechanisms in EMG-Based Speech Recognition
abstract
This paper reports on our recent research in the feedback effects of Silent Speech. Our technology is based on surface electromyography (EMG) which captures the electrical potentials of the human articulatory muscles rather than the acoustic speech signal. While recognition results are good for loudly articulated speech and when experienced users speak silently, novice users usually achieve far worse results when speaking silently. Since there is no acoustic feedback when speaking silently, we investigate different kinds of feedback modes: no additional feedback except the natural somatosensory feedback (like the touching of the lips), visual feedback using a mirror and indirect acoustic feedback by speaking simultaneously to a previously recorded audio signal. In addition we examine recorded EMG data when the subject speaks audibly and silently in a loud environment to see if the Lombard effect can be observed in Silent Speech, too.
Christian Herff, Matthias Janke, Michael Wand 0002, Tanja Schultz
INTERSPEECH4
2011 Generalized Baum-Welch Algorithm and its Implication to a New Extended Baum-Welch Algorithm
abstract
This paper describes how we can use the generalized Baum-Welch (GBW) algorithm to develop better extended Baum-Welch (EBW) algorithms. Based on GBW, we show that the backoff term in the EBW algorithm comes from KL-divergence which is used as a regularization function. This finding allows us to develop a fast EBW algorithm, which can reduce the time of model space discriminative training by half, without incurring any degradation on recognition accuracy. We compare the performance of the new EBW algorithm with the original one on various large scale systems including Farsi, Iraqi and modern standard Arabic ASR systems. Index Terms: speech recognition, discriminative training 1.
Roger Hsiao, Tanja Schultz
INTERSPEECH2
2011 Analysis of Dialectal Influence in Pan-Arabic ASR
abstract
In this paper, we analyze the impact of five Arabic dialects on the front-end and pronunciation dictionary component of an Automatic Speech Recognition (ASR) system.We use ASR"s phonetic decision tree as a diagnostic tool to compare the robustness of MFCC to MLP front-ends to dialectal variations in the speech data and found that MLP Bottle-Neck features are less robust to dialectal variation.We also perform a rulebased analysis of the pronunciation dictionary, which enables us to identify dialectal words in the vocabulary and automatically generate pronunciations for unseen words.We show that our technique produces pronunciations with an average phone error rate 9.2%.
Udhyakumar Nallasamy, Michael Garbus, Florian Metze, Qin Jin, Thomas Schaaf, Tanja Schultz
INTERSPEECH6
2011 Tue-SeA Real-Time Speech Command Detector for a Smart Control Room
abstract
In this work we present an online ASR system that is able to discriminate voice commands directed to an operationable screen from irrelevant speech segments. For classification of the sound segments we explored several features that are based on prosody as well as properties generated during the decoding process. For a vocabulary of 259 words and more than 10k possible commands, our realtime Verbal Command Detector managed to detect 88.3% of the commands in our evaluation data while maintaining a low False Positive Rate (FPR) of 1.5%. On an evaluation task using an episode of Star Trek, our system was able to detect 91.2% of all commands with a FPR of 1.8% with only minor adjustments. The system is part of and used in the Smart Control Room at the Fraunhofer IOSB in Karlsruhe [1], an experimental smart environment that uses multiple input modalities for crisis response.
Daniel Reich, Felix Putze, Dominic Heger, Joris IJsselmuiden, Rainer Stiefelhagen, Tanja Schultz
INTERSPEECH6
2011 Rapid Building of an ASR System for Under-Resourced Languages Based on Multilingual Unsupervised Training
abstract
This paper presents our work on rapid language adaptation of acoustic models based on multilingual cross-language bootstrapping and unsupervised training. We used Automatic Speech Recognition (ASR) systems in the six source languages English, French, German, Spanish, Bulgarian and Polish to build from scratch an ASR system for Vietnamese, an underresourced language. System building was performed without using any transcribed audio data by applying three consecutive steps, i.e. cross-language transfer, unsupervised training based on the “multilingual A-stabil ” confidence score [1], and bootstrapping. We investigated the correlation between performance of “multilingual A-stabil ” and the number of source languages and improved the performance of “multilingual A-stabil ” by applying it at the syllable level. Furthermore, we showed that increasing the amount of source language ASR systems for the multilingual framework results in better performance of the final ASR system in the target language Vietnamese. The final Vietnamese recognition system has a Syllable Error Rate (SyllER) of 16.8 % on the development set and 16.1 % on the evaluation set. Index Terms: rapid language adaptation of ASR, unsupervised training, multilingual A-Stabil
Ngoc Thang Vu, Franziska Kraus, Tanja Schultz
INTERSPEECH3
2011 Investigations on Speaking Mode Discrepancies in EMG-Based Speech Recognition
abstract
In this paper we present our recent study on the impact of speaking mode variabilities on speech recognition by surface electromyography (EMG). Surface electromyography captures the electric potentials of the human articulatory muscles, which enables a user to communicate naturally without making any audible sound. Our previous experiments have shown that the EMG signal varies greatly between different speaking modes, like audibly uttered speech and silently articulated speech. In this study we extend our previous research and quantify the impact of different speaking modes by investigating the amount of mode-specific leaves in phonetic decision trees. We show that this measure correlates highly with discrepancies in the spectral energy of the EMG signal, as well as with differences in the performance of a recognizer on different speaking modes. We furthermore present how EMG signal adaptation by spectral mapping decreases the effect of the speaking mode.
Michael Wand 0002, Matthias Janke, Tanja Schultz
INTERSPEECH3
2011 Investigation of Cross-Show Speaker Diarization
abstract
The goal of cross-show diarization is to index speech segments of speakers from a set of shows, with the particular challenge that reappearing speakers across shows have to be labeled with the same speaker identity. In this paper, we introduce three cross-show diarization systems namely Global-BIC-Seg, Global-BIC-Cluster, and Incremental. We compared the three systems on a set of 46 English scientific podcast shows. Among the three systems, the Global-BIC-Cluster achieves the best performance with 15.53% and 13.21% cross-show diarization error rate (DER) on the dev and test set, respectively. However, an incremental approach is more practical since data and shows are typically collected over time. By applying T-Norm on our incremental system, we obtain 13.18% and 10.97% relative improvements in terms of cross-show DER on dev and test set. We also investigate the impact of the show processing order on cross-show diarization for the incremental system. Index Terms: speaker diarization, cross-show diarization, conversational podcast shows
Qin Jin, Tanja Schultz
INTERSPEECH3
2011 Combined intention, activity, and motion recognition for a humanoid household robot
abstract
In this paper, a multi-level approach to intention, activity, and motion recognition for a humanoid robot is proposed. Our system processes images from a monocular camera and combines this information with domain knowledge. The recognition works on-line and in real-time, it is independent of the test person, but limited to predefined view-points. Main contributions of this paper are the extensible, multi-level modeling of the robot's vision system, the efficient activity and motion recognition, and the asynchronous information fusion based on generic processing of mid-level recognition results. The complementarity of the activity and motion recognition renders the approach robust against misclassifications. Experimental results on a real-world data set of complex kitchen tasks, e.g., Prepare Cereals or Lay Table, prove the performance and robustness of the multi-level recognition approach.
Dirk Gehrig, Peter Krauthausen, Lukas Rybok, Hilde Kuehne, Uwe D. Hanebeck, Tanja Schultz, Rainer Stiefelhagen
IROS6
2010 Speaker identification with distant microphone speech
abstract
The field of speaker identification has recently seen significant advancement, but improvements have tended to be benchmarked on near-field speech, ignoring the more realistic setting of far-field-instrumented speakers. In this work we present several findings on far-field speech from the MIXER5 Corpus, in the areas of feature extraction, speaker modeling, and multichannel score combination. First, we observe that minimum-variance distortionless response (MVDR) features outperform Mel-frequency cepstral coefficient (MFCC) features, and that fundamental frequency variation (FFV) features offer complimentary information to both MFCC and MVDR features. Second, we present evidence that factor analysis significantly improves system performance, compared to the more traditional GMM/UBM strategy. Third, we find that frame-based score competition significantly improves performance under mismatched conditions with multiple channels available.
Qin Jin, Runxin Li, Kornel Laskowski, Tanja Schultz
ICASSP5
2010 ICCHP Keynote: Recognizing Silent and Weak Speech Based on Electromyography
Tanja Schultz
ICCHP (1)1
2010 Multimodal Recognition of Cognitive Workload for Multitasking in the Car
abstract
This work describes the development and evaluation of a recognizer for different levels of cognitive workload in the car. We collected multiple biosignal streams (skin conductance, pulse, respiration, EEG) during an experiment in a driving simulator in which the drivers performed a primary driving task and several secondary tasks of varying difficulty. From this data, an SVM based workload classifier was trained and evaluated, yielding recognition rates of up to for three levels of workload.
Felix Putze, Jan-Philip Jarvis, Tanja Schultz
ICPR3
2010 Improvements to generalized discriminative feature transformation for speech recognition
abstract
Generalized Discriminative Feature Transformation (GDFT) is a feature space discriminative training algorithm for automatic speech recognition (ASR). GDFT uses Lagrange relaxation to transform the constrained maximum likelihood linear regression (CMLLR) algorithm for feature space discriminative training. This paper presents recent improvements on GDFT, which are achieved by regularization to the optimization problem. The resulting algorithm is called regularized GDFT (rGDFT) and we show that many regularization and smoothing techniques developed for model space discriminative training are also applicable to feature space training. We evaluated rGDFT on a real-time Iraqi ASR system and also on a large scale Arabic ASR task.
Roger Hsiao, Florian Metze, Tanja Schultz
INTERSPEECH3
2010 Impact of lack of acoustic feedback in EMG-based silent speech recognition
abstract
This paper presents our recent advances in speech recognition based on surface electromyography (EMG). This technology allows for Silent Speech Interfaces since EMG captures the electrical potentials of the human articulatory muscles rather than the acoustic speech signal. Our earlier experiments have shown that the EMG signal is greatly impacted by the mode of speaking. In this study we extend this line of research by comparing EMG signals from audible, whispered, and silent speaking mode. We distinguish between phonetic features like consonants and vowels and show that the lack of acoustic feedback in silent speech implies an increased focus on somatosensoric feedback, which is visible in the EMG signal. Based on this analysis we develop a spectral mapping method to compensate for these differences. Finally, we apply the spectral mapping to the front-end of our speech recognition system and show that recognition rates on silent speech improve by up to 11.59% relative. Index Terms: EMG, EMG-based speech recognition, Silent Speech Interfaces, somatosensoric feedback
Matthias Janke, Michael Wand 0002, Tanja Schultz
INTERSPEECH3
2010 The 2010 CMU GALE speech-to-text system
abstract
This paper describes the latest Speech-to-Text system developed for the Global Autonomous Language Exploitation ("GALE") domain by Carnegie Mellon University (CMU). This systems uses discriminative training, bottle-neck features and other techniques that were not used in previous versions of our system, and is trained on 1150 hours of data from a variety of Arabic speech sources. In this paper, we show how different lexica, pre-processing, and system combination techniques can be used to improve the final output, and provide analysis of the improvements achieved by the individual techniques.
Florian Metze, Roger Hsiao, Qin Jin, Udhyakumar Nallasamy, Tanja Schultz
INTERSPEECH5
2010 Utterance selection for speech acts in a cognitive tourguide scenario
abstract
Abstract This paper describes the integration of a cognitive memorymodel into a spoken dialog system for an in-car tourguide ap-plication. This memory model enhances the capabilities of thesystem and of the simulated user by estimating if and whichinformation is relevant and useful in a given situation. An eval-uation study with 15 human judges is performed to demonstratethe feasibility of the described approach. The results show thatthe proposed utterance selection strategy and the memory modelsignificantly improve the human-like interaction behavior of thespoken dialog system in terms of the amount and quality ofgiven information, relevance, manner, and naturalness of thespoken interaction.Index Terms: spoken interaction, cognition, memory model,workload, utterance selection 1. Introduction Spoken dialog systems (SDS) have matured to a point wherethey find their way into many real-world applications. How-ever, their application in very dynamic scenarios remains anopen and very challenging task. In our application, we imple-ment a virtual co-driver in the car that acts as a tourguide duringa ride. While a traditional SDS already offers an eyes-free andhands-free control for in-car information applications, the par-allel driving task uses the user’s cognitive capacity so we canno longer assume to deal with a fully attentive and perfect inter-action partner as in more static environments. Additionally, wehave to deal with an ever-changing context in the dynamic en-vironment. Therefore, we need to integrate components in ourdialog systems that are able to explicitly model, predict, andcope with this imperfect user and the varying focus to ensure aseamless and successful dialog experience.Especially in interaction scenarios which are not directlytask-driven like our tourguide scenario, utterance selection isnot trivial as while we still follow a clearly defined goal of pro-viding as much interesting information as possible, but have noclear order or priority of information chunks to present. Thesame is true if we want to simulate a user for evaluation or au-tomatic strategy learning. To create coherent user behavior, weneed to establish what the simulated users currently have ontheir mind. An explicit memory model aims for a detailed rep-resentation of human memory by dynamically modeling an in-dividual strength of activation for each chunk of information.This paper describes the implementation of the model and howit is applied to utterance selection for both the system and thesimulated user. The primary goal of the utterance selection is tofind an utterance that is most relevant in the current context andof most interest to the user.
Felix Putze, Tanja Schultz
INTERSPEECH2
2010 Wiktionary as a source for automatic pronunciation extraction
abstract
In this paper, we analyze whether dictionaries from the World Wide Web which contain phonetic notations, may support the rapid creation of pronunciation dictionaries within the speech recognition and speech synthesis system building process. As a representative dictionary, we selected Wiktionary [1] since it is at hand in multiple languages and, in addition to the definitions of the words, many phonetic notations in terms of the International Phonetic Alphabet (IPA) are available. Given word lists in four languages English, French, German, and Spanish, we calculated the percentage of words with phonetic notations in Wiktionary. Furthermore, two quality checks were performed: First, we compared pronunciations from Wiktionary to pronunciations from dictionaries based on the GlobalPhone project, which had been created in a rule-based fashion and were manually cross-checked [2]. Second, we analyzed the impact of Wiktionary pronunciations on automatic speech recognition (ASR) systems. French Wiktionary achieved the best pronunciation coverage, containing 92.58% phonetic notations for the French GlobalPhone word list as well as 76.12% and 30.16% for country and international city names. In our ASR systems evaluation, the Spanish system gained the most improvement from Wiktionary pronunciations with 7.22% relative word error rate reduction.
Tim Schlippe, Sebastian Ochs 0002, Tanja Schultz
INTERSPEECH3
2010 Text normalization based on statistical machine translation and internet user support
Tim Schlippe, Chenfei Zhu, Jan Gebhardt, Tanja Schultz
INTERSPEECH4
2010 Rapid bootstrapping of five eastern european languages using the rapid language adaptation toolkit
abstract
This paper presents our latest efforts toward LVCSR systems for five Eastern European languages such as Bulgarian, Croatian, Czech, Polish, and Russian using our Rapid Language Adaptation Toolkit (RLAT) [1]. We investigated the possibility of crawling large quantities of text material from the Internet, which is very cheap but also requires text post-processing steps due to the varying text quality. The goal of this study is to determine the best strategy for language model optimization on the given domain in a short time period with minimal human effort. Our results show that we can build an initial ASR system for these five languages in only twenty days using RLAT. On the multilingual GlobalPhone speech corpus [2], we achieved a word error rate (WER) of 16.9 % for Bulgarian, 32.8 % for
Ngoc Thang Vu, Tim Schlippe, Franziska Kraus, Tanja Schultz
INTERSPEECH4
2010 Multilingual a-stabil: A new confidence score for multilingual unsupervised training
abstract
This paper presents our work in Automatic Speech Recognition (ASR) in the context of multilingual unsupervised training with application to Czech. Starting without any transcribed acoustic training data we built a Czech ASR by combining cross-language bootstrapping and confidence based unsupervised training. We present our new method called “multilingual A-stabil” to compute confidence scores and explore the relative effectiveness of acoustic models from more than one language such as Russian, Bulgarian, Polish and Croatian for unsupervised training. While conventional confidence measures such as gamma and A-stabil work well with well-trained acoustic models but have problems with poorly estimated acoustic models, our new method works well in both cases. We describe our multilingual unsupervised training framework which gives very promising results in our experiments. We were able to select 80.5% of the audio training data (18.5 hours) with a transcription WER of 14.5% when using a small amount of untranscribed data (only about 23 hours). The final best WER on Czech is 23.6% on the development set and 22.9% on the evaluation set by using cross-lingual boostrapping, which is very close to the performance of the Czech ASR trained with 23 hours audio data with manual transcriptions (23.1% on the development set and 22.3% on the evaluation set).
Ngoc Thang Vu, Franziska Kraus, Tanja Schultz
SLT3
2010 Guest Editorial
Bruce Denby, Tanja Schultz, Kiyoshi Honda
Speech Commun.2
2010 Silent speech interfaces
Bruce Denby, Tanja Schultz, Kiyoshi Honda, Thomas Hueber, J. M. Gilbert, Jonathan S. Brumberg
Speech Commun.2
2010 Modeling coarticulation in EMG-based continuous speech recognition
Tanja Schultz, Michael Wand 0002
Speech Commun.1
2009 Speaker de-identification via voice transformation
abstract
It is a common feature of modern automated voice-driven applications and services to record and transmit a user's spoken request. At the same time, several domains and applications may require keeping the content of the user's request confidential and at the same time preserving the speaker's identity. This requires a technology that allows the speaker's voice to be de-identified in the sense that the voice sounds natural and intelligible but does not reveal the identity of the speaker. In this paper we investigate different voice transformation strategies on a large population of speakers to disguise the speakers' identities while preserving the intelligibility of the voices. We apply two automatic speaker identification approaches to verify the success of de-identification with voice transformation, a GMM-based and a phonetic approach. The evaluation based on the automatic speaker identification systems verifies that the proposed voice transformation technique enables transmission of the content of the users' spoken requests while successfully preserving their identities. Also, the results indicate that different speakers still sound distinct after the transformation. Furthermore, we carried out a human listening test that proved the transformed speech to be both intelligible and securely de-identified, as it hid the identity of the speakers even to listeners who knew the speakers very well.
Qin Jin, Arthur R. Toth, Tanja Schultz, Alan W. Black
ASRU3
2009 Rapid language adaptation tools for multilingual speech processing
abstract
Summary form only given: The performance of speech and language processing technologies has improved dramatically over the past decade, with an increasing number of systems being deployed in a large variety of applications, such as spoken dialog systems, speech summarization and information retrieval systems, and speech translation systems. Most efforts to date were focused on a very small number of languages with large number of speakers, economic potential, and information technology needs of the population. However, speech technology has a lot to contribute even to those languages that do not fall into this category. Languages with a small number of speakers and few linguistic resources may suddenly become of interest for humanitarian and military reasons. Furthermore, a large number of languages are in danger of becoming extinct, and ongoing projects for preserving them could benefit from speech technology. With more than 6900 languages in the world and the need to support multiple input and output languages, the most important challenge today is to port speech processing systems to new languages rapidly and at reasonable costs. In my talk I will introduce state-of-the-art techniques for rapid language adaptation and present solutions to overcome the ever-existing problem of data sparseness and the gap between language and technology expertise. I will describe the building process for speech and language processing components for new unsupported languages and introduce tools to do this rapidly and at lost costs. I describe the Rapid Language Adaptation Tools (RLAT) which built on existing projects like SPICE, GlobalPhone, and FestVox and enable users to develop speech processing components, to collect appropriate speech and text data for building and improving these components, and to evaluate the results allowing for iterative improvements.
Tanja Schultz
ASRU1
2009 A multiplatform speech recognition decoder based on weighted finite-state transducers
abstract
Speech recognition decoders based on static graphs have recently proven to significantly outperform the traditional approach of prefix tree expansion in terms of decoding speed. The reduced search effort makes static graph decoders an attractive alternative for tasks concerned with limited processing power or memory footprint on devices such as PDAs, internet tablets, and smart phones. In this paper we explore the benefits of decoding with an optimized speech recognition network over the fully task-optimized prefix-tree based decoder IBIS. We designed and implemented a new decoder called SWIFT (speedy weigthed finite-state transducer) based on WFSTs with its application to embedded platforms in mind. After describing the design, the network construction and storage process, we present evaluation results on a small task suitable for embedded applications, and on a large task, namely the European Parliament Plenary Sessions (EPPS) task from the TC-STAR project. The SWIFT Decoder is up to 50% faster than IBIS on both tasks. In addition, SWIFT achieves significant memory consumption reductions obtained by our innovative network specific storage layout optimization.
Emilian Stoimenov, Tanja Schultz
ASRU2
2009 Vietnamese large vocabulary continuous speech recognition
abstract
We report on our recent efforts toward a large vocabulary Vietnamese speech recognition system. In particular, we describe the Vietnamese text and speech database recently collected as part of our GlobalPhone corpus. The data was complemented by a large collection of text data crawled from various Vietnamese websites. To bootstrap the Vietnamese speech recognition system we used our Rapid Language Adaptation scheme applying a multilingual phone inventory. After initialization we investigated the peculiarities of the Vietnamese language and achieved significant improvements by implementing different tone modeling schemes, extended by pitch extraction, handling multiwords to address the monosyllable structure of Vietnamese, and featuring language modeling based on 5-grams. Furthermore, we addressed the issue of dialectal variations between South and North Vietnam by creating dialect dependent pronunciations and including dialect in the context decision tree of the recognizer. Our currently best recognition system achieves a word error rate of 11.7% on read newspaper speech.
Ngoc Thang Vu, Tanja Schultz
ASRU2
2009 Joint Learning of Preposition Senses and Semantic Roles of Prepositional Phrases
Daniel Dahlmeier, Hwee Tou Ng, Tanja Schultz
EMNLP3
2009 Detecting bandlimited audio in broadcast television shows
abstract
For TV and radio shows containing narrowband speech, Speech-to-text (STT) accuracy on the narrowband audio can be improved by using an acoustic model trained on acoustically matched data. To selectively apply it, one must first be able to accurately detect which audio segments are narrowband. The present paper explores two different bandwidth classification approaches: a traditional Gaussian mixture model (GMM) approach and a spline-based classifier that categorizes audio segments based on their power spectra. We focus on shows found in the DARPA GALE Mandarin training and test sets, where the ratio of wideband to narrowband shows is very large. In this setting, the spline-based classifier reduces the number of misclassified wideband segments by up to 95% relative to the GMM-based classifier for the same number of misclassified narrowband segments.
Mark C. Fuhs, Qin Jin, Tanja Schultz
ICASSP3
2009 Generalized Baum-Welch algorithm for discriminative training on large vocabulary continuous speech recognition system
abstract
We propose a new optimization algorithm called Generalized Baum Welch (GBW) algorithm for discriminative training on hidden Markov model (HMM). GBW is based on Lagrange relaxation on a transformed optimization problem. We show that both Baum-Welch (BW) algorithm for ML estimate of HMM parameters, and the popular extended Baum-Welch (EBW) algorithm for discriminative training are special cases of GBW.We compare the performance of GBW and EBW for Farsi large vocabulary continuous speech recognition (LVCSR).
Roger Hsiao, Yik-Cheung Tam, Tanja Schultz
ICASSP3
2009 Voice convergin: Speaker de-identification by voice transformation
abstract
Speaker identification might be a suitable answer to prevent unauthorized access to personal data. However we also need to provide solutions to secure transmission of spoken information. This challenge divides into two major aspects. First, the secure transmission of the content of the spoken input and second the secure transmission of the identity of the speaker. In this paper we concentrate on the latter, i.e. how to securely transmit information via voice without revealing the identity of the speaker to unauthorized listeners. In order to make the first steps toward solving this problem we study in this paper the potential of voice transformation for speaker de-identification. We use two speaker identification approaches to verify the success of de-identification with voice transformation, a GMM-based and a Phonetic approach, and study different voice transformation strategies to disguise speaker identity information while preserving understandability.
Qin Jin, Arthur R. Toth, Tanja Schultz, Alan W. Black
ICASSP3
2009 The I4U system in NIST 2008 speaker recognition evaluation
abstract
This paper describes the performance of the I4U speaker recognition system in the NIST 2008 Speaker Recognition Evaluation. The system consists of seven subsystems, each with different cepstral features and classifiers. We describe the I4U Primary system and report on its core test results as they were submitted, which were among the best-performing submissions. The I4U effort was led by the Institute for Infocomm Research, Singapore (IIR), with contributions from the University of Science and Technology of China (USTC), the University of New South Wales, Australia (UNSW), Nanyang Technological University, Singapore (NTU) and Carnegie Mellon University, USA (CMU).
Haizhou Li 0001, Bin Ma 0001, Kong-Aik Lee, Hanwu Sun, Donglai Zhu, Khe Chai Sim, Chang Huai You, Rong Tong, Ismo Kärkkäinen, Chien-Lin Huang, Vladimir Pervouchine, Wu Guo, Yijie Li 0001, Li-Rong Dai 0001, Mohaddeseh Nosratighods, Tharmarajah Thiruvaran, Julien Epps, Eliathamby Ambikairajah, Chng Eng Siong, Tanja Schultz, Qin Jin
ICASSP20
2009 Incorporating monolingual corpora into bilingual latent semantic analysis for crosslingual LM adaptation
abstract
The major limitation in bilingual latent semantic analysis (bLSA) is the requirement of parallel training corpora. Motivated by semi-supervised learning, we propose a clusterbased bLSA training approach to incorporate monolingual corpora. Treating each parallel document pair as centroids of the parallel document clusters, each monolingual document is associated to the closest centroid according to their topic similarity. The resulting parallel document clusters are used as constraints to enforce a one-to-one topic correspondence in variational EM. Slight performance improvement in crosslingual language model adaptation is observed compared to the baseline without monolingual corpora.
Yik-Cheung Tam, Tanja Schultz
ICASSP2
2009 Generalized discriminative feature transformation for speech recognition
abstract
We propose a new algorithm called Generalized Discriminative Feature Transformation (GDFT) for acoustic models in speech recognition. GDFT is based on Lagrange relaxation on a transformed optimization problem. We show that the existing discriminative feature transformation methods like feature space MMI/MPE (fMMI/MPE), region dependent linear transformation (RDLT), and a non-discriminative feature transformation, constrained maximum likelihood linear regression (CMLLR) are special cases of GDFT. We evaluate the performance of GDFT for Iraqi large vocabulary continuous speech recognition. Index Terms: speech recognition, discriminative training, feature transformation
Roger Hsiao, Tanja Schultz
INTERSPEECH2
2009 Improving speaker segmentation via speaker identification and text segmentation
abstract
Speaker segmentation is an essential part of a speaker diarization system. Common segmentation systems usually miss speaker change points when speakers switch fast. These errors seriously confuse the following speaker clustering step and result in high overall speaker diarization error rates. In this paper two methods are proposed to deal with this problem: The first approach uses speaker identification techniques to boost speaker segmentation. And the second approach applies text segmentation methods to improve the performance of speaker segmentation. Experiments on Quaero speaker diarization evaluation data shows that our methods achieve up to 45 % relative reduction in the speaker diarization error and 64 % relative increase in the speaker change detection recall rate over the baseline system. Moreover, both these two approaches can be considered as post-processing steps over the baseline segmentation, therefore, they can be applied in any speaker diarization systems. Index Terms: speaker diarization, speaker segmentation, speaker identification, text segmentation
Runxin Li, Tanja Schultz, Qin Jin
INTERSPEECH2
2009 Synthesizing speech from electromyography using voice transformation techniques
abstract
Surface electromyography (EMG) can be used to record the activation potentials of articulatory muscles while a person speaks. This technique could enable silent speech interfaces, as EMG signals are generated even when people pantomime speech without producing sound. Having effective silent speech interfaces would enable a number of compelling applications, allowing people to communicate in areas where they would not want to be overheard or where the background noise is so prevalent that they could not be heard. In order to use EMG signals in speech interfaces, however, there must be a relatively accurate method to map the signals to speech. Up to this point, it appears that most attempts to use EMG signals for speech interfaces have focused on Automatic Speech Recognition (ASR) based on features derived from EMG signals. Following the lead of other researchers who worked with Electro-Magnetic Articulograph (EMA) data and Non-Audible Murmur (NAM) speech, we explore the alternative idea of using Voice Transformation (VT) techniques to synthesize speech from EMG signals. With speech output, both ASR systems and human listeners can directly use EMG-based systems. We report the results of our preliminary studies, noting the difficulties we encountered and suggesting areas for future work. Index Terms: electromyography, silent speech, voice transformation, speech synthesis
Arthur R. Toth, Michael Wand 0002, Tanja Schultz
INTERSPEECH3
2009 Impact of different speaking modes on EMG-based speech recognition
abstract
We present our recent results on speech recognition by surface electromyography (EMG), which captures the electric potentials that are generated by the human articulatory muscles. This technique can be used to enable Silent Speech Interfaces, since EMG signals are generated even when people only articulate speech without producing any sound. Preliminary experiments have shown that the EMG signals created by audible and silent speech are quite distinct. In this paper we first compare various methods of initializing a silent speech EMG recognizer, showing that the performance of the recognizer substantially varies across different speakers. Based on this, we analyze EMG signals from audible and silent speech, present first results on how discrepancies between these speaking modes affect EMG recognizers, and suggest areas for future work. Index Terms: speech recognition, surface electromyography, silent speech, articulation
Michael Wand 0002, Szu-Chen Stan Jou, Arthur R. Toth, Tanja Schultz
INTERSPEECH4
2009 Speaker identification using warped MVDR cepstral features
abstract
It is common practice to use similar or even the same feature extraction methods for automatic speech recognition and speaker identification. While the front-end for the former requires to preserve phoneme discrimination and to compensate for speaker differences to some extend, the front-end for the latter has to preserve the unique characteristics of individual speakers. It seems, therefore, contradictory to use the same feature extraction methods for both tasks. Starting out from the common practice we propose to use warped minimum variance distortionless response (MVDR) cepstral coefficients, which have already been demonstrated to perform superior for automatic speech recognition in particular under adverse conditions. Replacing the widely used mel-frequency cepstral coefficients by WMVDR cepstral coefficients improves the speaker identification accuracy by up to 24 % relative. We found that the optimal choice of the model order within the WMVDR framework differs between speech recognition and speaker recognition, confirming our intuition that the two different tasks indeed require different feature extraction strategies. 1.
Matthias Wölfel, Qin Jin, Tanja Schultz
INTERSPEECH4
2009 Towards an EEG-based emotion recognizer for humanoid robots
abstract
In the field of interaction between humans and robots emotions have been disregarded for a long time. During the last few years interest in emotion research in this area has been constantly increasing as giving a robot the ability to react to the emotional state of the user can help to make the interaction more human-like and enhance the acceptance of the robots. In this paper we investigate a method to facilitate emotion recognition from electroencephalographic signals. For this purpose we developed a headband to measure electroencephalographic signals on the forehead. Using this headband we collected data from five subjects. To induce emotions we used 90 pictures from the International Affective Picture System (IAPS) belonging to the three categories pleasant, neutral, and unpleasant. For emotion recognition we developed a system based on support vector machines (SVMs). With this system an average recognition rate of 47.11% could be achieved on subject dependent recognition.
Kristina Schaaff, Tanja Schultz
RO-MAN2
2009 Introduction to the Special Issue on Processing Morphologically Rich Languages
abstract
The 12 papers in this special issue span a variety of speech and language processing applications highlighting the challenges and providing solutions for dealing with the morphological complexity of different languages.
Ruhi Sarikaya, Katrin Kirchhoff, Tanja Schultz, Dilek Hakkani-Tür
IEEE Trans. Speech Audio Process.3
2008 Is voice transformation a threat to speaker identification?
abstract
With the development of voice transformation and speech synthesis technologies, speaker identification systems are likely to face attacks from imposters who use voice transformed or synthesized speech to mimic a particular speaker. Therefore, we investigated in this paper how speaker identification systems perform on voice transformed speech. We conducted experiments with two different approaches, the classical GMM-based speaker identification system and the Phonetic speaker identification system. Our experimental results showed that current standard voice transformation techniques are able to fool the GMM-based system but not the Phonetic speaker identification system. These findings imply that future speaker identification systems should include idiosyncratic knowledge in order to successfully distinguish transformed speech from natural speech and thus be armed against imposter attacks.
Qin Jin, Arthur R. Toth, Alan W. Black, Tanja Schultz
ICASSP4
2008 Sentence segmentation and punctuation recovery for spoken language translation
abstract
Sentence segmentation and punctuation recovery are critical components for effective spoken language translation (SLT). In this paper we describe our recent work on sentence segmentation and punctuation recovery for three different language pairs, namely for English-to-Spanish, Arabic-to-English and Chinese-to-English. We show that the proposed approach works equally well in these very different language pairs. Furthermore, we introduce two features computed from the translation beam-search lattice that indicate if phrasal and target language model context is jeopardized when segmenting at a given word boundary. These features enable us to introduce short intra-sentence segments without degrading translation performance.
Matthias Paulik, Sharath Rao, Ian Lane, Stephan Vogel, Tanja Schultz
ICASSP5
2008 Selecting relevant features for human motion recognition
abstract
Recently, there is a growing interest in automatic recognition of human motion for applications, such as humanoid robots, human activity monitoring, and surveillance. In this paper we investigate motion recognition based on joint angle trajectories derived from marker-based video recordings. The goal of this paper is to improve the generalization and robustness of human motion recognition even if only limited amount of training data is available. We achieve this goal by significantly reducing the amount of input features. We leverage on recent studies in the area of neuroscience which indicate that human motions display only a few independent degrees of freedom (DOF). We examine which DOF are relevant for recognizing upper body human motions and to what extend the dimensionality of the feature vectors can be reduced in order to simplify the data acquisition and improve the robustness of the recognition process. Our final results indicate that careful selection of features proves to reduce the number of features by a factor of up to 3, while at the same time significantly improving the recognition performance.
Dirk Gehrig, Tanja Schultz
ICPR2
2008 The CMU-interACT 2008 Mandarin transcription system
abstract
We present our Mandarin BN/BC transcription system recently developed for the GALE07 evaluation. The system employs a 3-pass decoding strategy trained with over 1300 hours of quickly transcribed audio. We successfully apply discriminative training, dynamic unsupervised language model adaptation, and system combination techniques in our system. We furthermore achieve improvements by combining an Initial-Final system with a genre dependent phone system. On the GALE07 phase 2 retest evaluation, our system achieves a character error rate(CER) of 13.3 % on dev07 test set and 13.5 % on eval07 unsequestered test set. Our system also allows combination with other sites and in this paper, we investigate different system combination strategies which significantly improve the final recognition performance. Index Terms: Mandarin transcription system, broadcast news, broadcast conversation, GALE evaluation
Roger Hsiao, Mark C. Fuhs, Yik-Cheung Tam, Qin Jin, Tanja Schultz
INTERSPEECH5
2008 Robust far-field speaker identification under mismatched conditions
abstract
While speaker identification performance has improved dramatically over the past years, the presence of interfering noise and the variety of channel conditions pose a major obstacle. Particularly the mismatch between training and test condition leads to severe performance degradations. In this paper we investigate speaker identification based on data simultaneously recorded with multiple microphones in a farfield setup under different noise and reverberation conditions. Dramatic performance degradation is observed, especially when training and test conditions mismatch. To address this mismatch we apply our robust frame-based score competition approach in which we combine and compete models trained on multiple conditions. To further improve this approach we add simulated, i.e. artificially created training data on a variety of noise conditions for additional model training. Our experimental results show that the extended approach significantly improves speaker identification performance under adverse and mismatching conditions.
Qin Jin, Tanja Schultz
INTERSPEECH2
2008 Improving speech systems built from very little data
abstract
This paper studies two ways for helping non-specialist users develop speech systems from limited data for new languages. Focused web re-crawling finds additional examples of text matching the domain as specified by the user. This improves the language model and cuts word error rate nearly in half. Iterative voice building with interleaved lexicon construction uses the voice from a previous iteration to help construct an improved voice. 4.5 hours of the user’s time reduces transcription error rate from 32 % to 4%. 1.
John Kominek, Sameer Badaskar, Tanja Schultz, Alan W. Black
INTERSPEECH3
2008 Recovering participant identities in meetings from a probabilistic description of vocal interaction
abstract
An important decision in the design of automatic conversation understanding systems is the level at which information streams representing specific participants are merged. In the current work, we explore participant-dependence of low-level interactive aspects of conversation, namely the observed contextual preferences for talkspurt deployment. We argue that strong participant-dependence at this level gives cause for merging participant streams as early as possible. We demonstrate that our probabilistic description of talkspurt deployment preferences is strongly participant-dependent, and frequently predictive of participant identity.
Kornel Laskowski, Tanja Schultz
INTERSPEECH2
2008 NineOneOne: Recognizing and Classifying Speech for Handling Minority Language Emergency Calls
Udhyakumar Nallasamy, Alan W. Black, Tanja Schultz, Robert E. Frederking
LREC3
2008 Correlated Bigram LSA for Unsupervised Language Model Adaptation
abstract
We propose using correlated bigram LSA for unsupervised LM adaptation for automatic speech recognition. The model is trained using efficient variational EM and smoothed using the proposed fractional Kneser-Ney smoothing which handles fractional counts. Our approach can be scalable to large training corpora via bootstrapping of bigram LSA from unigram LSA. For LM adaptation, unigram and bigram LSA are integrated into the background N-gram LM via marginal adaptation and linear interpolation respectively. Experimental results show that applying unigram and bigram LSA together yields 6%--8% relative perplexity reduction and 0.6% absolute character error rates (CER) reduction compared to applying only unigram LSA on the Mandarin RT04 test set. Comparing with the unadapted baseline, our approach reduces the absolute CER by 1.2%.
Yik-Cheung Tam, Tanja Schultz
NIPS2
2008 Improving word segmentation for Thai speech translation
abstract
A vocabulary list and language model are primary components in a speech translation system. Generating both from plain text is a straightforward task for English. However, it is quite challenging for Chinese, Japanese, or Thai which provide no word segmentation, i.e. the text has no word boundary delimiter. For Thai word segmentation, maximal matching, a lexicon-based approach, is one of the popular methods. Nevertheless this method heavily relies on the coverage of the lexicon. When text contains an unknown word, this method usually produces a wrong boundary. When extracting words from this segmented text, some words will not be retrieved because of wrong segmentation. In this paper, we propose statistical techniques to tackle this problem. Based on different word segmentation methods we develop various speech translation systems and show that the proposed method can significantly improve the translation accuracy by about 6.42% BLEU points compared to the baseline system.
Paisarn Charoenpornsawat, Tanja Schultz
SLT2
2007 Bilingual-LSA Based LM Adaptation for Spoken Language Translation
Yik-Cheung Tam, Ian Lane, Tanja Schultz
ACL3
2007 Continuous Electromyographic Speech Recognition with a Multi-Stream Decoding Architecture
abstract
In our previous work, we reported a surface electromyographic (EMG) continuous speech recognition system with a novel EMG feature extraction method, E4, which is more robust to EMG noise than traditional spectral features. In this paper, we show that articulatory feature (AF) classifiers can also benefit from the E4 feature, which improve the F-score of the AF classifiers from 0.492 to 0.686. We also show that the E4 feature is less correlated across EMG channels and thus channel combination gains larger improvement in F-score. With a stream architecture, the AF classifiers are then integrated into the decoding framework and improve the word error rate by 11.8% relative from 33.9% to 29.9%.
Szu-Chen Stan Jou, Tanja Schultz, Alex Waibel
ICASSP (4)2
2007 Correlated Latent Semantic Model for Unsupervised LM Adaptation
abstract
We propose a latent Dirichlet-tree allocation (LDTA) model - a correlated latent semantic model - for unsupervised language model adaptation. The LDTA model extends the latent Dirichlet allocation (LDA) model by replacing a Dirichlet prior with a Dirichlet-tree prior over the topic proportions. Latent topics under the same subtree are expected to be more correlated than topics under different subtrees. The LDTA model falls back to the LDA model using a depth-one Dirichlet-tree, and the model fits to the variational Bayes inference framework employed in the LDA model. Empirical results show that the LDTA model has a faster training convergence than the LDA model with the same initial flat model. Experimental results show that LDTA-adapted LM performed better than LDA-adapted LM on the Mandarin RT04-eval set when the models were trained using a small text corpus, while both models had the same recognition performance when the models were trained using a big text corpus. We observed 0.4% absolute CER reduction after LM adaptation using LSA marginals.
Yik-Cheung Tam, Tanja Schultz
ICASSP (4)2
2007 Whispering Speaker Identification
abstract
This paper describes a study of automatically identifying whispering speakers. People usually whisper in order to avoid being identified or overheard by lowering their voices. The study compares performances between normal and whispered speech mode in clean and noisy environment under matched and mismatched training conditions, and describes the impact of feature warping and throat microphone on noise reduction. Score combination strategies are used when only little whisper data is available to improve performance. In sum, we achieved 8% to 33% relative improvements in identification accuracy with only 5 to 10 seconds noisy whispered speech data per speaker.
Qin Jin, Szu-Chen Stan Jou, Tanja Schultz
ICME3
2007 Handling OOV words in Arabic ASR via flexible morphological constraints
abstract
We propose a novel framework to detect and recognize outof-vocabulary (OOV) words in automated speech recognition (ASR).In the proposed framework a hybrid language model combining words and sub-word units is incorporated during ASR decoding then three different OOV words recognition methods are applied to generate OOV word hypotheses.Specifically, dictionary lookup, morphological composition, and direct phoneme-to-grapheme.The proposed approach successfully reduced WER by 1.9% and 1.6% for ASR systems with recognition vocabularies of 30K and 219K.Moreover, the proposed approach correctly recognized 5% of OOV words.
Nguyen Bach, Mohamed Noamany, Ian Lane, Tanja Schultz
INTERSPEECH4
2007 Optimizing sentence segmentation for spoken language translation
abstract
The conventional approach in text-based machine translation (MT) is to translate complete sentences, which are conveniently indicated by sentence boundary markers.However, since such boundary markers are not available for speech, new methods are required that define an optimal unit for translation.Our experimental results show that with a segment length optimized for a particular MT system, intrasentence segmentation can improve translation performance (measured in BLEU) by up to 11% for Arabic Broadcast Conversation (BC) and 6% for Arabic Broadcast News (BN).We show that acoustic segmentation that minimizes Word Error Rate (WER) may not give the best translation performance.We improve upon it by automatically resegmenting the ASR output in a way that is optimized for translation and argue that it might be necessary for different stages of a Spoken Language Translation (SLT) system to define their own optimal units.
Sharath Rao, Ian Lane, Tanja Schultz
INTERSPEECH3
2007 SPICE: web-based tools for rapid language adaptation in speech processing systems
abstract
In this paper we describe the design and implementation of a user interface for SPICE, a web-based toolkit for rapid prototyping of speech and language processing components.We report on the challenges and experiences gathered from testing these tools in an advanced graduate hands-on course, in which we created speech recognition, speech synthesis, and smalldomain translation components for 10 different languages within only 6 weeks.
Tanja Schultz, Alan W. Black, Sameer Badaskar, Matthew Hornyak, John Kominek
INTERSPEECH1
2007 Bilingual LSA-based translation lexicon adaptation for spoken language translation
abstract
We present a bilingual LSA (bLSA) framework for translation lexicon adaptation. The idea is to apply marginal adaptation on a translation lexicon so that the lexicon marginals match to in-domain marginals. In the framework of speech translation, the bLSA method transfers topic distributions from the source to the target side, such that the translation lexicon can be adapted before translation based on the source document. We evaluated the proposed approach on our Mandarin RT04 spoken language translation system. Results showed that the conditional likeli-hood on the test sentence pairs is improved significantly using an adapted translation lexicon compared to an unadapted base-line. The proposed approach showed improvement on BLEU-score in SMT. When both the target-side LM and the translation lexicon were adapted and applied simultaneously for SMT de-coding, the gain on BLEU-score was more than additive com-pared to the scenarios when the adapted models were individu-ally applied.
Yik-Cheung Tam, Tanja Schultz
INTERSPEECH2
2007 Wavelet-based front-end for electromyographic speech recognition
abstract
In this paper we present our investigations on the potential of wavelet-based preprocessing for surface electromyographic speech recognition.We implemented several variants of the Discrete Wavelet Transform and applied them to electromyographical data.First we examined different transforms with various filters and decomposition levels and found that the Redundant Discrete Wavelet Transform performs the best among all tested wavelet transforms.Furthermore, we compared the best wavelet transform to our EMG optimized spectral-and timedomain features.The results showed that the best wavelet transform slightly outperforms the optimized features with 30.9% word error rate compared to 32% for the optimized EMG spectral and time-domain features.Both numbers were achieved on a 108 word vocabulary test set using phone based acoustic models trained on continuously spoken speech captured by EMG.
Michael Wand 0002, Szu-Chen Stan Jou, Tanja Schultz
INTERSPEECH3
2007 Improving spoken language translation by automatic disfluency removal: evidence from conversational speech transcripts
Sharath Rao, Ian Lane, Tanja Schultz
MTSummit3
2007 Bilingual LSA-based adaptation for statistical machine translation
Yik-Cheung Tam, Ian Lane, Tanja Schultz
Mach. Transl.3
2007 Far-Field Speaker Recognition
abstract
In this paper, we study robust speaker recognition in far-field microphone situations. Two approaches are investigated to improve the robustness of speaker recognition in such scenarios. The first approach applies traditional techniques based on acoustic features. We introduce reverberation compensation as well as feature warping and gain significant improvements, even under mismatched training-testing conditions. In addition, we performed multiple channel combination experiments to make use of information from multiple distant microphones. Overall, we achieved up to 87.1% relative improvements on our Distant Microphone database and found that the gains hold across different data conditions and microphone settings. The second approach makes use of higher-level linguistic features. To capture speaker idiosyncrasies, we apply n-gram models trained on multilingual phone strings and show that higher-level features are more robust under mismatching conditions. Furthermore, we compared the performances between multilingual and multiengine systems, and examined the impact of a number of involved languages on recognition results. Our findings confirm the usefulness of language variety and indicate a language independent nature of this approach, which suggests that speaker recognition using multilingual phone strings could be successfully applied to any given language.
Qin Jin, Tanja Schultz, Alex Waibel
IEEE Trans. Speech Audio Process.2
2006 Far-Field Speaker Recognition
abstract
In this paper we study robust speaker recognition in far-field microphone situations such as meeting scenarios. By applying reverberation compensation and feature warping we achieved significant improvements under mismatched training-testing conditions. To capture useful information from multiple distant microphones, two approaches for multiple channel combination are investigated. This leads to 84.1% and 78.1% relative improvements on the Distant Microphone database. Furthermore, we tested the resulting system on the ICSI Meeting Corpus. The improvements are also very high on this task, which indicates that our system is robust to changing conditions in a remote microphone setting.
Qin Jin, Tanja Schultz
ICASSP (1)3
2006 Articulatory Feature Classification using Surface Electromyography
abstract
In this paper, we present an approach for articulatory feature classification based on surface electromyographic signals generated by the facial muscles. With parallel recorded audible speech and electromyographic signals, experiments are conducted to show the anticipatory behavior of electromyographic signals with respect to speech signals. On average, we found that the signals to be time delayed by 0.02 to 0.12 second. Furthermore, it is shown that different articulators have different anticipatory behavior. With offset-aligned signals, we improved the average F-score of the articulatory feature classifiers in our baseline system from 0.467 to 0.502.
Szu-Chen Stan Jou, Lena Maier-Hein, Tanja Schultz, Alex Waibel
ICASSP (1)3
2006 Unsupervised Learning of Overlapped Speech Model Parameters For Multichannel Speech Activity Detection in Meetings
abstract
The study of meetings, and multi-party conversation in general, is currently the focus of much attention, calling for more robust and more accurate speech activity detection systems. We present a novel multichannel speech activity detection algorithm, which explicitly models the overlap incurred by participants taking turns at speaking. Parameters for overlapped speech states are estimated during decoding by using and combining knowledge from other observed states in the same meeting, in an unsupervised manner. We demonstrate on the NIST Rich Transcription Spring 2004 data set that the new system almost halves the number of frames missed by a competitive algorithm within regions of overlapped speech. The overall speech detection error on unseen data is reduced by 36% relative
Kornel Laskowski, Tanja Schultz
ICASSP (1)2
2006 Acoustic-Phonetic Unit Similarities For Context Dependent Acoustic Model Portability
abstract
This paper addresses particularly the use of acoustic-phonetic unit similarities for portability of context dependent acoustic models to new languages. Since the IPA-based method is limited to a source/target phoneme mapping table construction, an estimation method of the similarity between two phonemes is proposed in this paper. Based on these phoneme similarities, some estimation methods for polyphone similarity and clustered polyphonic model similarity are investigated. For a new language, first a polyphonic decision tree is built with a small amount of speech data. Then, clustered models in the target language are duplicated from the nearest clustered models in the source language and adapted with limited data to the target language. Results obtained from the experiments demonstrate the feasibility of these methods.
Viet Bac Le, Laurent Besacier, Tanja Schultz
ICASSP (1)3
2006 Challenges with Rapid Adaptation of Speech Translation Systems to New Language Pairs
abstract
Although we have far from solved the issues in porting speech translation systems to new languages, we have gathered sufficient experience by now to identify a number of major challenges in the process. Although well-defined processes exist for building speech recognition, speech synthesis and statistical machine translation models, they still require both significant native speaker involvement and linguistic expertise. As the core technology improves we believe we will see increasing cultural and social issues in contributions from native speakers. This paper identifies some of these issues and presents our initial attempts to build tools that we hope will eventually allow linguistically naive native informants build complete speech translation systems
Tanja Schultz, Alan W. Black
ICASSP (5)1
2006 Example-based grapheme-to-phoneme conversion for Thai
abstract
Several characteristics of the Thai writing system make Thai grapheme-to-phoneme (G2P) conversion very challenging. In this paper, we propose an Example-Based Grapheme-to-Phoneme conversion approach. It generates the pronunciation of a word by selecting, modifying and combining pronunciations from syllables from training corpus. The best system achieves 80.99 % word accuracy and 94.19 % phone accuracy which significantly outperform previous approaches for Thai.
Paisarn Charoenpornsawat, Tanja Schultz
INTERSPEECH2
2006 Optimizing components for handheld two-way speech translation for an English-iraqi Arabic system
abstract
This paper described our handheld two-way speech translation system for English and Iraqi. The focus is on developing a field usable handheld device for speech-to-speech translation. The computation and memory limitations on the handheld impose critical constraints on the ASR, SMT, and TTS components. In this paper we discuss our approaches to optimize these components for the handheld device and present performance numbers from the evaluations that were an integral part of the project. Since one major aspect of the TransTac program is to build fieldable systems, we spent significant effort on developing an intuitive interface that minimizes the training time for users but also provides useful information such as back translations for translation quality feedback.
Roger Hsiao, Ashish Venugopal, Thilo Köhler, Ying Zhang 0048, Paisarn Charoenpornsawat, Andreas Zollmann, Stephan Vogel, Alan W. Black, Tanja Schultz, Alex Waibel
INTERSPEECH9
2006 Towards continuous speech recognition using surface electromyography
abstract
We present our research on continuous speech recognition of the surface electromyographic signals that are generated by the human articulatory muscles.Previous research on electromyographic speech recognition was limited to isolated word recognition because it was very difficult to train phoneme-based acoustic models for the electromyographic speech recognizer.In this paper, we demonstrate how to train the phoneme-based acoustic models with carefully designed electromyographic feature extraction methods.By decomposing the signal into different feature space, we successfully keep the useful information while reducing the noise.Additionally, we also model the anticipatory effect of the electromyographic signals compared to the speech signal.With a 108-word decoding vocabulary, the experimental results show that the word error rate improves from 86.8% to 32.0% by using our novel feature extraction methods.
Szu-Chen Stan Jou, Tanja Schultz, Matthias Walliczek, Florian Kraft, Alex Waibel
INTERSPEECH2
2006 Unsupervised language model adaptation using latent semantic marginals
abstract
We integrated the Latent Dirichlet Allocation (LDA) approach, a latent semantic analysis model, into unsupervised language model adaptation framework. We adapted a background language model by minimizing the Kullback-Leibler divergence between the adapted model and the background model subject to a constraint that the marginalized unigram probability distribution of the adapted model is equal to the corresponding distribution estimated by the LDA model – the latent semantic marginals. We evaluated our approach on the RT04 Mandarin Broadcast News test set and experimented with different LM training settings. Results showed that our approach reduces the perplexity and the character error rates using supervised and unsupervised adaptation.
Yik-Cheung Tam, Tanja Schultz
INTERSPEECH2
2006 Sub-word unit based non-audible speech recognition using surface electromyography
abstract
In this paper we present a novel approach for a surface electromyographic speech recognition system based on sub-word units. Rather than using full word models as integrated in our previous work we propose here smaller sub-word units as prerequisites for large vocabulary speech recognition. This allows the recognition of words not seen in the training set based on seen sub-word units. Therefore we report on experiments with syllables and phonemes as sub-word units. We also developed a new feature extraction method that gains significant improvement for words and sub-word units. Index Terms: silent speech, non-audible speech recognition, electromyography, sub-word unit comparison
Matthias Walliczek, Florian Kraft, Szu-Chen Stan Jou, Tanja Schultz, Alex Waibel
INTERSPEECH4
2006 Spontaneous Thai speech recognition
abstract
This paper expands previous work on Thai speech recognition, investigating pronunciation changes such as syllable and phoneme elisions as well as phoneme shifts in Thai spontaneous speech.We compare several approaches to model these effects in large vocabulary continuous speech recognition across multiple domains.This work includes experiments on two new speech databases that significantly alleviate the data sparseness problem of earlier publications.We found that given sufficient training data, a fully data driven approach using an allophone cluster tree yields the best results.Explicit modeling of pronunciation changes does not improve performance across domains.
Monika Woszczyna, Paisarn Charoenpornsawat, Tanja Schultz
INTERSPEECH3
2006 Thai Grapheme-Based Speech Recognition
Paisarn Charoenpornsawat, Sanjika Hewavitharana, Tanja Schultz
HLT-NAACL3
2006 Flexible speech translation systems
abstract
Speech translation research has made significant progress over the years with many high-visibility efforts showing that translation of spontaneously spoken speech from and to diverse languages is possible and applicable in a variety of domains. As language and domains continue to expand, practical concerns such as portability and reconfigurability of speech come into play: system maintenance becomes a key issue and data is never sufficient to cover the changing domains over varying languages. In this paper, we discuss strategies to overcome the limits of today's speech translation systems. In the first part, we describe our layered system architecture that allows for easy component integration, resource sharing across components, comparison of alternative approaches, and the migration toward hybrid desktop/PDA or stand-alone PDA systems. In the second part, we show how flexibility and reconfigurability is implemented by more radically relying on learning approaches and use our English–Thai two-way speech translation system as a concrete example.
Tanja Schultz, Alan W. Black, Stephan Vogel, Monika Woszczyna
IEEE Trans. Speech Audio Process.1
2005 Automatic Disfluency Removal on Recognized Spontaneous Speech - Rapid Adaptation to Speaker Dependent Disfluencies
abstract
In this paper, we investigate methods to adapt a system for disfluency removal to different data properties. A gradient descent algorithm for parameter optimization is presented which achieves 85.1% recall and 93.1% precision on the English Verbmobil corpus and 53.0% recall and 79.0% precision on the Mandarin Chinese CallHome corpus. This compares to the results produced with hand-optimization on the test set. Furthermore, we investigated the impact of cross-validation and training set selection on recognizer output. Finally, we examined speaker dependent disfluency production behavior and clustered training data accordingly in order to improve the overall system.
Matthias Honal, Tanja Schultz
ICASSP (1)2
2005 Whispery Speech Recognition using Adapted Articulatory Features
abstract
This paper describes our research on adaptation methods applied to articulatory feature detection on soft whispery speech recorded with a throat microphone. Since the amount of adaptation data is small and the testing data is very different from the training data, a series of adaptation methods is necessary. The adaptation methods include: maximum likelihood linear regression, feature-space adaptation, and re-training with downsampling, sigmoidal low-pass filter, and linear multivariate regression. Adapted articulatory feature detectors are used in parallel to standard senone-based HMM models in a stream architecture for decoding. With these adaptation methods, articulatory feature detection accuracy improves from 87.82% to 90.52% with corresponding F-measure from 0.504 to 0.617, while the final word error rate improves from 33.8% to 31.2%.
Szu-Chen Stan Jou, Tanja Schultz, Alex Waibel
ICASSP (1)2
2005 Thai Automatic Speech Recognition
abstract
We describe the development of a robust and flexible Thai speech recognizer as integrated into our English-Thai speech-to-speech translation system. We focus on the discussion of the rapid deployment of ASR for Thai under limited time and data resources, including rapid data collection issues, acoustic model bootstrap, and automatic generation of pronunciations. Issues relating to the translation and overall system will be reported elsewhere.
Sinaporn Suebvisai, Paisarn Charoenpornsawat, Alan W. Black, Monika Woszczyna, Tanja Schultz
ICASSP (1)5
2005 Document driven machine translation enhanced ASR
abstract
In human-mediated translation scenarios a human interpreter translates between a source and a target language using either a spoken or a written representation of the source language. In this paper we improve the recognition performance on the speech of the human translator spoken in the target language by taking advantage of the source language representations. We use machine translation techniques to translate between the source and target language resources and then bias the target language speech recognizer towards the gained knowledge, hence the name Machine Translation Enhanced Automatic Speech Recognition. We investigate several different techniques among which are restricting the search vocabulary, selecting hypotheses from n-best lists, applying cache and interpolation schemes to language modeling, and combining the most successful techniques into our final, iterative system. Overall we outperform the baseline system by a relative word error rate reduction of 37.6%.
Matthias Paulik, Christian Fügen, Sebastian Stüker, Tanja Schultz, Thomas Schaaf, Alex Waibel
INTERSPEECH4
2005 Dynamic language model adaptation using variational Bayes inference
abstract
We propose an unsupervised dynamic language model (LM) adaptation framework using long-distance latent topic mixtures.The framework employs the Latent Dirichlet Allocation model (LDA) which models the latent topics of a document collection in an unsupervised and Bayesian fashion.In the LDA model, each word is modeled as a mixture of latent topics.Varying topics within a context can be modeled by re-sampling the mixture weights of the latent topics from a prior Dirichlet distribution.The model can be trained using the variational Bayes Expectation Maximization algorithm.During decoding, mixture weights of the latent topics are adapted dynamically using the hypotheses of previously decoded utterances.In our work, the LDA model is combined with the trigram language model using linear interpolation.We evaluated the approach on the CCTV episode of the RT04 Mandarin Broadcast News test set.Results show that the proposed approach reduces the perplexity by up to 15.4% relative and the character error rate by 4.9% relative depending on the size and setup of the training set.
Yik-Cheung Tam, Tanja Schultz
INTERSPEECH2
2004 Towards language portability in statistical speech translation
abstract
Speech translation has made significant advances over the last years. We believe that we can overcome today's limits of language and domain portable conversational speech translation systems by relying more radically on learning approaches and by the use of multiple layers of reduction and transformation to extract the desired content in another language. Therefore, we cascade stochastic source-channel models that extract an underlying message from a corrupt observed output. The three models effectively translate: (1) speech to word lattices (automatic speech recognition, ASR); (2) ill-formed fragments of word strings into a compact well-formed sentence (Clean); (3) sentences in one language to sentences in another (machine translation, MT). We present results of our research efforts towards rapid language portability of all these components. The results on translation suggest that MT systems can be successfully constructed for any language pair by cascading multiple MT systems via English. Moreover, end-to-end performance can be improved, if the interlingua language is enriched with additional linguistic information that can be derived automatically and monolingually in a data-driven fashion.
Alex Waibel, Tanja Schultz, Stephan Vogel, Christian Fügen, Matthias Honal, Muntsin Kolss, Jürgen Reichert, Sebastian Stüker
ICASSP (3)2
2004 Identifying the addressee in human-human-robot interactions based on head pose and speech
abstract
In this work we investigate the power of acoustic and visual cues, and their combination, to identify the addressee in a human-human-robot interaction. Based on eighteen audio-visual recordings of two human beings and a (simulated) robot we discriminate the interaction of the two humans from the interaction of one human with the robot. The paper compares the result of three approaches. The first approach uses purely acoustic cues to find the addressees. Low level, feature based cues as well as higher-level cues are examined. In the second approach we test whether the human's head pose is a suitable cue. Our results show that visually estimated head pose is a more reliable cue for the identification of the addressee in the human-human-robot interaction. In the third approach we combine the acoustic and visual cues which results in significant improvements.
Michael Katzenmaier, Rainer Stiefelhagen, Tanja Schultz
ICMI3
2004 Speaker segmentation and clustering in meetings
abstract
This paper describes the issue of automatic speaker segmentation and clustering for natural, multi-speaker meeting conversations. Two systems were developed and evaluated in the NIST RT-04S Meeting Recognition Evaluation, the Multiple Distant Microphone (MDM) system and the Individual Headset Microphone (IHM) system. The MDM system achieved a speaker diarization performance of 28.17%. This system also aims to provide automatic speech segments and speaker grouping information for speech recognition, a necessary prerequisite for subsequent audio processing. A 44.5 % word error rate was achieved for speech recognition. The IHM system is based on the short-time crosscorrelation of all personal channel pairs. It requires no prior training and executes in one fifth real time on modern architectures. A 35.7 % word error rate was achieved for speech recognition when segmentation was provided by this system. 1.
Qin Jin, Tanja Schultz
INTERSPEECH2
2004 Adaptation for soft whisper recognition using a throat microphone
abstract
This paper describes various adaptation methods applied to recognizing soft whisper recorded with a throat microphone.Since the amount of adaptation data is small and the testing data is very different from the training data, a series of adaptation methods is necessary.The adaptation methods include: maximum likelihood linear regression, feature-space adaptation, and re-training with downsampling, sigmoidal low-pass filter, or linear multivariate regression.With these adaptation methods, the word error rate improves from 99.3% to 32.9%.
Szu-Chen Stan Jou, Tanja Schultz, Alex Waibel
INTERSPEECH2
2004 Crosscorrelation-based multispeaker speech activity detection
abstract
We propose an algorithm for segmenting multispeaker meeting audio, recorded with personal channel microphones, into speech and non-speech intervals for each microphone’s wearer. An algorithm of this type turns out to be necessary prior to subsequent audio processing because, in spite of close-talking microphones, the channels exhibit a high degree of crosstalk due to unbalanced calibration and small inter-speaker distance. The proposed algorithm is based on the short-time crosscorrelation of all channel pairs. It requires no prior training and executes in one fifth real time on modern architectures. Using meeting audio collected at several sites, we present error rates for the segmentation task which do not appear correlated with microphone type or number of speakers. We also present the resulting improvement in speech recognition accuracy when segmentation is provided by this algorithm.
Kornel Laskowski, Qin Jin, Tanja Schultz
INTERSPEECH3
2004 Issues in meeting transcription - the ISL meeting transcription system
abstract
This paper describes the Interactive Systems Lab’s Meeting transcription system, which performs segmentation, speaker clustering as well as transcriptions of conversational meeting speech. The system described here was evaluated in NIST’s RT-04S “Meeting” speech evaluation. This paper compares the performance of our Broadcast News and the most recent Switchboard system on the Meeting data and compares both with a newly-trained meeting recognizer. Furthermore we investigate the effects of automatic segmentation on adaptation. Our best meeting system achieves 44.5% on the MDM condition in NIST’s RT-04S evaluation.
Tanja Schultz, Qin Jin, Kornel Laskowski, Florian Metze, Christian Fügen
INTERSPEECH1
2004 Using word latice information for a tighter coupling in speech translation systems
abstract
In this paper we present first experiments towards a tighter coupling between Automatic Speech Recognition (ASR) and Statistical Machine Translation (SMT) to improve the overall performance of our speech translation system. In coventional speech translation systems, the recognizer outputs a single hypothesis which is then translated by the SMT system. This approach has the limitation of being largely dependent on the word error rate of the first best hypothesis. The word error rate is typically lowered by generating many alternative hypotheses in the form of a word lattice. The information in the word lattice and the scores from the recognizer can be used by the translation system to obtain better performance. In our experiments, by switching from the single best hypotheses to word lattices as the interface between ASR and SMT, and by introducing weighted acoustic scores in the translation system, the overall performance was increased by 16.22%.
Tanja Schultz, Szu-Chen Stan Jou, Stephan Vogel, Shirin Saleem
INTERSPEECH1
2003 Multilingual articulatory features
abstract
Speech recognition systems based on or aided by articulatory features, such as place and manner of articulation, have been shown to be useful under varying circumstances. Recognizers based on features better compensate channel and noise variability. We show that it is also possible to compensate for inter language variability using articulatory feature detectors. We come to the conclusion that articulatory features can be recognized across languages and that using detectors from many languages can improve the classification accuracy of the feature detectors on a single language. We further demonstrate how those multilingual and cross-lingual detectors can support an HMM based recognizer and thereby significantly reduce the word error rate by up to 12.3% relative. We expect that with the use of multilingual articulatory features it is possible to support the rapid deployment of recognition systems for new target languages.
Sebastian Stüker, Tanja Schultz, Florian Metze, Alex Waibel
ICASSP (1)2
2003 SMaRT: the Smart Meeting Room Task at ISL
abstract
As computational and communications systems become increasingly smaller, faster, more powerful, and more integrated, the goal of interactive, integrated meeting support rooms is slowly becoming reality. It is already possible, for instance, to rapidly locate task-related information during a meeting, filter it, and share it with remote users. Unfortunately, the technologies that provide such capabilities are as obstructive as they are useful - they force humans to focus on the tool rather than the task. Thus the veneer of utility often hides the true costs of use, which are longer, less focused human interactions. To address this issue, we present our current research efforts towards SMaRT: the Smart Meeting Room Task. The goal of SMaRT is to provide meeting support services that do not require explicit human-computer interaction. Instead, by monitoring the activities in the meeting room using both video and audio analysis, the room is able to react appropriately to users' needs and allow the users to focus on their own goals.
Alex Waibel, Tanja Schultz, Michael Bett, Matthias Denecke, Robert G. Malkin, Ivica Rogina, Rainer Stiefelhagen, Jie Yang 0001
ICASSP (4)2
2003 Comparison of acoustic model adaptation techniques on non-native speech
abstract
The performance of speech recognition systems is consistently poor on non-native speech. The challenge for non-native speech recognition is to maximize the recognition performance with a small amount of available non-native data. We report on acoustic modeling adaptation for the recognition of non-native speech. Using non-native data from German speakers, we investigate how bilingual models, speaker adaptation, acoustic model interpolation and polyphone decision tree specialization methods can help to improve the recognizer performance. Results obtained from the experiments demonstrate the feasibility of these methods.
Zhirong Wang, Tanja Schultz, Alex Waibel
ICASSP (1)2
2003 Correction of disfluencies in spontaneous speech using a noisy-channel approach
abstract
In this paper we present a system which automatically corrects disfluencies such as repairs and restarts typically occurring in spontaneously spoken speech. The system is based on a noisy-channel model and its development requires no linguistic knowledge, but only annotated texts. Therefore, it has large potential for rapid deployment and the adaptation to new target languages. The experiments were conducted on spontaneously spoken dialogs from the English VERBMOBIL corpus where a recall of 77.2% and a precision of 90.2% was obtained. To demonstrate the feasibility of rapid adaptation additional experiments on the spontaneous Mandarin Chinese CallHome corpus were performed achieving 49.4% recall and 76.8% precision.
Matthias Honal, Tanja Schultz
INTERSPEECH2
2003 Grapheme based speech recognition
abstract
Large vocabulary speech recognition systems traditionally represent words in terms of smaller subword units. During training and recognition they require a mapping table, called the dictionary, which maps words into sequences of these subword units. The performance of the speech recognition system depends critically on the definition of the subword units and the accuracy of the dictionary. In current large vocabulary speech recognition systems these components are often designed manually with the help of a language expert. This is time consuming and costly. In this Diplomarbeit graphemes are taken as subword units and the dictionary creation becomes a triviality.
Mirjam Killer, Sebastian Stüker, Tanja Schultz
INTERSPEECH3
2003 Integrating multilingual articulatory features into speech recognition
abstract
The use of articulatory features, such as place and manner of articulation, has been shown to reduce the word error rate of speech recognition systems under different conditions and in different settings.For example recognition systems based on features are more robust to noise and reverberation.In earlier work we showed that articulatory features can compensate for inter language variability and can be recognized across languages.In this paper we show that using cross-and multilingual detectors to support an HMM based speech recognition system significantly reduces the word error rate.By selecting and weighting the features in a discriminative way, we achieve an error rate reduction that lies in the same range as that seen when using language specific feature detectors.By combining feature detectors from many languages and training the weights discriminatively, we even outperform the case where only monolingual detectors are being used.
Sebastian Stüker, Florian Metze, Tanja Schultz, Alex Waibel
INTERSPEECH3
2003 Speechalator: two-way speech-to-speech translation on a consumer PDA
abstract
This paper describes a working two-way speech-to-speech translation system that runs in near real-time on a consumer handheld computer. It can translate from English to Arabic and Arabic to English in the domain of medical interviews. We describe the general architecture and frameworks within which we developed each of the components: HMM-based recognition, interlingua translation (both rule and statistically based), and unit selection synthesis.
Alex Waibel, Ahmed Badran, Alan W. Black, Robert E. Frederking, Donna Gates, Alon Lavie, Lori S. Levin, Kevin A. Lenzo, Laura Mayfield Tomokiyo, Jürgen Reichert, Tanja Schultz, Dorcas Wallace, Monika Woszczyna, Jing Zhang 0011
INTERSPEECH11
2003 Non-native spontaneous speech recognition through polyphone decision tree specialization
abstract
English, the fast and efficient adaptation to non-native English speech becomes a practical concern. The performance of speech recognition systems is consistently poor on non-native speech. The challenge for non-native speech recognition is to maximize the recognition performance with small amount of non-native data available. In this paper we report on the effectiveness of using polyphone decision tree specialization method for non-native speech adaptation and recognition. Several recognition results are presented by using non-native speech from German speakers. Results obtained from the experiments demonstrate the feasibility of this method.
Zhirong Wang, Tanja Schultz
INTERSPEECH2
2003 Enhanced tree clustering with single pronunciation dictionary for conversational speech recognition
abstract
Modeling pronunciation variation is key for recognizing conversational speech. Rather than being limited to dictionary modeling, we argue that triphone clustering is an integral part of pronunciation modeling. We propose a new approach called enhanced tree clustering.This approach, in contrast to traditional decision tree based state tying, allows parameter sharing across phonemes. We show that accurate pronunciation modeling can be achieved through efficient parameter sharing in the acoustic model. Combined with a single pronunciation dictionary, a 1.8% absolute word error rate improvement is achieved on Switchboard, a large vocabulary conversational speech recognition task.
Hua Yu 0008, Tanja Schultz
INTERSPEECH2
2003 Speechalator: Two-Way Speech-to-Speech Translation in Your Hand
Alex Waibel, Ahmed Badran, Alan W. Black, Robert E. Frederking, Donna Gates, Alon Lavie, Lori S. Levin, Kevin A. Lenzo, Laura Mayfield Tomokiyo, Jürgen Reichert, Tanja Schultz, Dorcas Wallace, Monika Woszczyna, Jing Zhang 0011
HLT-NAACL11
2003 Implicit Trajectory Modeling through Gaussian Transition Models for Speech Recognition
Hua Yu 0008, Tanja Schultz
HLT-NAACL2
2002 Speaker identification using multilingual phone strings
abstract
Far-field speaker identification is very challenging since varying recording conditions often result in un-matching training and testing situations. Although the widely used Gaussian Mixture Models (GMM) approach achieves reasonable good results when training and testing conditions match, its performance degrades dramatically under un-matching conditions. In this paper we propose a new approach for far-field speaker identification: the usage of multilingual phone strings derived from phone recognizers in eight different languages. The experiments are carried out on a database of 30 speakers recorded with eight different microphone distances. The results show that the multi-lingual phone string approach is robust against un-matching conditions and significantly outperforms the GMMs. On 10-second test chunks, the average closed-set identification performance achieves 96.7% on variable distance data.
Qin Jin, Tanja Schultz, Alex Waibel
ICASSP2
2002 Toward robust parametric trajectory segmental model for vowel recognition
abstract
In this paper we present a robust and discriminative segmental trajectory modeling for vowel recognition. We proposed two new approaches. One is using weighted least square estimation for the parametric trajectory parameter, which gives a much more robust performance over traditional least square estimation approach. The other is a specifically designed transformation matrix proposed to reduce the possible mismatch between the Gaussian modeling assumption and the trajectory feature's nature. Our experiments on the vowel classification using the mobile phone data of SpeechDAT(II) MDB showed significant improvement over both standard HMM and traditional segmental modeling.
Tanja Schultz
ICASSP2
2002 Towards Universal Speech Recognition
abstract
The increasing interest in multilingual applications like speech-to-speech translation systems is accompanied by the need for speech recognition front-ends in many languages that can also handle multiple input languages at the same time. We describe a universal speech recognition system that fulfills such needs. It is trained by sharing speech and text data across languages and thus reduces the number of parameters and overhead significantly at the cost of only slight accuracy loss. The final recognizer eases the burden of maintaining several monolingual engines, makes dedicated language identification obsolete and allows for code-switching within an utterance. To achieve these goals we developed new methods for constructing multilingual acoustic models and multilingual n-gram language models.
Zhirong Wang, Umut Topkara, Tanja Schultz, Alex Waibel
ICMI3
2002 Phonetic speaker identification
Qin Jin, Tanja Schultz, Alex Waibel
INTERSPEECH2
2002 Globalphone: a multilingual speech and text database developed at karlsruhe university
abstract
This paper describes the design, collection, and current status of the multilingual database GlobalPhone, an ongoing project since 1995 at Karlsruhe University. GlobalPhone is a highquality read speech and text database in a large variety of languages which is suitable for the development of large vocabulary speech recognition systems in many languages. It has already been successfully applied to language independent and language adaptive speech recognition. GlobalPhone currently covers 15 languages Arabic, Chinese (Mandarin and Shanghai), Croatian, Czech, French, German, Japanese, Korean, Portuguese, Russian, Spanish, Swedish, Tamil, and Turkish. The corpus contains more than 300 hours of transcribed speech spoken by more than 1500 native, adult speakers and will soon be available from ELRA.
Tanja Schultz
INTERSPEECH1
2001 Advances in automatic meeting record creation and access
abstract
Oral communication is transient, but many important decisions, social contracts and fact findings are first carried out in an oral setup, documented in written form and later retrieved. At Carnegie Mellon University's Interactive Systems Laboratories we have been experimenting with the documentation of meetings. The paper summarizes part of the progress that we have made in this test bed, specifically on the question of automatic transcription using large vocabulary continuous speech recognition, information access using non-keyword based methods, summarization and user interfaces. The system is capable of automatically constructing a searchable and browsable audio-visual database of meetings and provide access to these records.
Alex Waibel, Michael Bett, Florian Metze, Klaus Ries 0001, Thomas Schaaf, Tanja Schultz, Hagen Soltau, Hua Yu 0008, Klaus Zechner
ICASSP6
2001 Experiments on cross-language acoustic modeling
abstract
With the distribution of speech products all over the world, the portability to new target languages becomes a practical concern. As a consequence our research focuses on rapid transfer of LVCSR systems to other languages. In former studies we evaluated the performance if limited adaptation data is available. Particularly for very time constrained tasks and minority languages, it is even reasonable that no training data is available at all. In this paper we examine what performance can be expected in this scenario. All experiments are run in the framework of the GlobalPhone project which investigates LVCSR systems in 15 languages.
Tanja Schultz, Alex Waibel
INTERSPEECH1
2001 Language-independent and language-adaptive acoustic modeling for speech recognition
Tanja Schultz, Alex Waibel
Speech Commun.1
2000 Turkish LVCSR: towards better speech recognition for agglutinative languages
abstract
The Turkish language belongs to the Turkic family. All members of this family are close to one another in terms of linguistic structure. Typological similarities are vowel harmony, verb-final word order and agglutinative morphology. This latter property causes a very fast vocabulary growth resulting in a large number of out-of-vocabulary words. In this paper we describe our first experiments in a speaker independent LVCSR engine for Modern Standard Turkish. First results on our Turkish speech recognition system are presented. The currently best system shows very promising results achieving 16.9% word error rate. To overcome the OOV-problem we propose a morphem-based and the Hypothesis Driven Lexical Adaptation approach. The final Turkish system is integrated into the multilingual recognition engine of the GlobalPhone project.
Kenan Çarki, Petra Geutner, Tanja Schultz
ICASSP3
2000 Confidence measure based language identification
abstract
In this paper we present a new application for confidence measures in spoken language processing. In today's computerized dialogue systems, language identification (LID) is typically achieved via dedicated modules. In our approach, LID is integrated into the speech recognizer, therefore profiting from high-level linguistic knowledge at very little extra cost. Our new approach is based on a word lattice based confidence measure (Kemp and Schaaf, 1997), which was originally devised for unsupervised training. In this work, we show that the confidence based language identification algorithm outperforms conventional score based methods. Also, this method is less dependent on the acoustic characteristics of the transmission channel than score based methods. By introducing additional parameters, unknown languages can be rejected. The proposed method is compared to a score based approach on the Verbmobil database, a three language task.
Florian Metze, Thomas Kemp, Thomas Schaaf, Tanja Schultz, Hagen Soltau
ICASSP4
2000 Polyphone decision tree specialization for language adaptation
abstract
With the distribution of speech technology products all over the world, the fast and efficient portability to new target languages becomes a practical concern. The authors explore the relative effectiveness of adapting multilingual LVCSR systems to a new target language with limited adaptation data. For this purpose they introduce a polyphone decision tree specialization method. Several recognition results are presented based on mono- and multilingual recognizers. These recognizers are developed in the framework of the project GlobalPhone. In this project we investigate speech recognition in 15 languages: Arabic, Mandarin and Shanghai Chinese, Croatian, English, French, German, Japanese, Korean, Portuguese, Russian, Spanish, Swedish, Tamil, and Turkish.
Tanja Schultz, Alex Waibel
ICASSP1
2000 VERBMOBIL dialogues: multifaced analysis
abstract
This paper describes the outline of collecting and transcribing spontaneous spoken dialogues for VERBMOBIL, the German research project on multilingual processing of spontaneous speech. The method and conditions of data collection performed using the same scenario and the transliteration convention of spontaneous speech were described. The characteristics of VERBMOBIL corpus were presented in terms of the size of dialogues, turns, sentences, words, perplexity based on the linguistic analysis. 1
Akira Kurematsu, Youichi Akegami, Susanne Burger, Susanne Jekat, Brigitte Lause, Victoria MacLaren, Daniela Oppermann, Tanja Schultz
INTERSPEECH8
2000 Multilinguality in speech and spoken language systems
abstract
Building modern speech and language systems currently requires large data resources such as texts, voice recordings, pronunciation lexicons, morphological decomposition information and parsing grammars. Based on a study of the most important differences between language groups, we introduce approaches to efficiently deal with the enormous task of covering even a small percentage of the world's languages. For speech recognition, we have reduced the resource requirements by applying acoustic model combination, bootstrapping and adaption techniques. Similar algorithms have been applied to improve the recognition of foreign accents. Segmenting language into appropriate units reduces the amount of data required to robustly estimate statistical models. The underlying morphological principles are also used to automatically adapt the coverage of our speech recognition dictionaries with the Hypothesis-Driven Lexical Adaptation (HDLA) algorithm. This reduces the out-of-vocabulary problems encountered in agglutinative languages. Speech recognition results are reported for the read GlobalPhone database and some broadcast news data. For speech translation, using a task-oriented Interlingua allows to build a system with N languages with linear, rather than quadratic effort. We have introduced a modular grammar design to maximize reusability and portability. End-to-end translation results are reported on a travel-domain task in the framework of C-STAR.
Alex Waibel, Petra Geutner, Laura Mayfield Tomokiyo, Tanja Schultz, Monika Woszczyna
Proc. IEEE4
1999 Mandarin large vocabulary speech recognition using the globalphone database
Jürgen Reichert, Tanja Schultz, Alex Waibel
EUROSPEECH2
1998 Recognition of music types
abstract
This paper describes a music type recognition system that can be used to index and search in multimedia databases. A new approach to temporal structure modeling is supposed. The so called ETM-NN (explicit time modelling with neural network) method uses abstraction of acoustical events to the hidden units of a neural network. This new set of abstract features representing temporal structures, can be then learned via a traditional neural networks to discriminate between different types of music. The experiments show that this method outperforms HMMs significantly.
Hagen Soltau, Tanja Schultz, Martin Westphal, Alex Waibel
ICASSP2
1998 Language independent and language adaptive large vocabulary speech recognition
abstract
This paper describes the design of a multilingual speech recognizer using an LVCSR dictation database which has been collected under the project GlobalPhone. This project at the University of Karlsruhe investigates LVCSR systems in 15 languages of the world, namely Arabic, Chinese, Croatian, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish, Swedish, Tamil, and Turkish. Based on a global phoneme set we built different multilingual speech recognition systems for five of the 15 languages. Context dependent phoneme models are created data-driven by introducing questions about language and language groups to our polyphone clustering procedure. We apply the resulting multilingual models to unseen languages and present several recognition results in language independent and language adaptive setups. 1. Introduction As the demand for speech recognition systems in multiple languages grows, the development of multilingual systems which combine the phonetic inven...
Tanja Schultz, Alex Waibel
ICSLP1
1998 Linear discriminant - a new criterion for speaker normalization
abstract
In Vocal Tract Length Normalization (VTLN) a linear or nonlinear frequency transformation compensates for different vocal tract lengths. Finding good estimates for the speaker specific warp parameters is a critical issue. Despite good results using the Maximum Likelihood criterion to find parameters for a linear warping, there are concerns using this method. We searched for a new criterion that enhances the inter-class separability in addition to optimizing the distribution of each phonetic class. Using such a criterion, Linear Discriminant Analysis determines a linear transformation in a lower dimensional space. For VTLN, we keep the dimension constant and warp the training samples of each speaker such that the Linear Discriminant is optimized. Although that criterion depends on all training samples of all speakers it can iteratively provide speaker specific warp factors. We discuss how this approach can be applied in speech recognition and present first results on two different recog...
Martin Westphal, Tanja Schultz, Alex Waibel
ICSLP2
1997 Japanese LVCSR on the spontaneous scheduling task with JANUS-3
abstract
This paper presents our findings during the development of the recognition engine for the Japanese part of the VERBMOBIL speech-to-speech translation project. We describe an efficient method to bootstrap a large vocabulary speech recognizer for spontaneously spoken Japanese speech from a German recognizer and show that the amount of effort in developing the system could be reduced by using this rapid cross language bootstrapping technique. The Japanese recognizer is integrated into the VERBMOBIL system and shows very promising results achieving 9.3% word error rate. 1. INTRODUCTION The overall goal of the first phase of the VERBMOBIL project is to build a speech-to-speech translation system from both German and Japanese spontaneously spoken input speech to English, German and Japanese output in an appointment scenario [1]. The Japanese recognizer described in this paper is beeing designed to be part of this translation system. Unlike Japanese dictation systems [2] there is no need fo...
Tanja Schultz, Detlef Koll, Alex Waibel
EUROSPEECH1
1997 Fast bootstrapping of LVCSR systems with multilingual phoneme sets
abstract
In this paper we described an efficient method to bootstrap continuously spoken, large vocabulary speech recognition systems by multilingual phoneme sets. To evaluate this techniques we collected the multilingual database GlobalPhone which currently consists of 9 different languages. A multilingual recognizer (MULTI) based on the four languages German, English, Japanese and Spanish was developed to serve as a source system. Likewise this system is very useful for language identification and achieves 100% language identification rate. Based on the MULTI system we evaluated our bootstrap technique on such completely different languages as Chinese, Croatian, and Turkish. 1. INTRODUCTION As the demand for speech recognition and translation systems in multiple languages grows, the development of multilingual systems is of increasing concern. On the one hand a multilingual system can be used as a language independent speech recognition and translation system with integrated automatic langu...
Tanja Schultz, Alex Waibel
EUROSPEECH1
1996 LVCSR-based language identification
abstract
Automatic language identification is an important problem in building multilingual speech recognition and understanding systems. Building a language identification module for four languages we studied the influence of applying different levels of knowledge sources on a large vocabulary continuous speech recognition (LVCSR) approach, i.e. phonetic, phonotactic, lexical, and syntactic-semantic knowledge. The resulting language identification (LID) module can identify spontaneous speech input and can be used as a front end for the multilingual speech-to-speech translation system JANUS-II. A comparison of five LID systems showed that the incorporation of lexical and linguistic knowledge reduces the language identification error for the 2-language tests up to 50%. Based on these results we build a LID module for German, English, Spanish, and Japanese which yields 84% identification rate on the spontaneous scheduling task (SST).
Tanja Schultz, Ivica Rogina, Alex Waibel
ICASSP1
1995 Acoustic and language modeling of human and nonhuman noises for human-to-human spontaneous speech recognition
abstract
Several improvements of our speech-to-speech translation system JANUS on spontaneous human-to-human dialogs are presented. Common phenomena in spontaneous speech are described, followed by a classification of different types of noise. To handle the variety of spontaneous effects in human-to-human dialogs, special noise models are introduced representing both human and nonhuman noise, as well as word fragments. It is shown that both the acoustic and the language modeling of the noise increase the recognition performance significantly. In the experiments, a clustering of the noise classes is performed and the resulting cluster variants are compared, thus allowing one to determine the best tradeoff between the sensitivity and trainability of the models.
Tanja Schultz, Ivica Rogina
ICASSP1
1994 JANUS 93: towards spontaneous speech translation
abstract
We present first results from our efforts toward translation of spontaneously spoken speech. Improvements include increasing coverage, robustness, generality and speed of JANUS, the speech-to-speech translation system of Carnegie Mellon and Karlsruhe University. The recognition and machine translation engine have been upgraded to deal with requirements introduced by spontaneous human to human dialogs. To allow for development and evaluation of our system on adequate data, a large database with spontaneous scheduling dialogs is being gathered for English, German and Spanish.>
Monika Woszczyna, Naomi Aoki-Waibel, Finn Dag Buø, Noah Coccaro, Keiko Horiguchi, Thomas Kemp, Alon Lavie, Arthur E. McNair, Thomas Polzin, Ivica Rogina, Carolyn P. Rosé, Tanja Schultz, Bernhard Suhm, Masaru Tomita, Alex Waibel
ICASSP (1)12
1992 Stochastic modeling of syllable-based units for continuous speech recognition
Günther Ruske, Bernd Plannerer, Tanja Schultz
ICSLP3