VLDB 2026 Research / reviewers in the wild / expert
Ngoc Thang Vu
dblp:10/9231
· DBLP profile ↗
106ranked-venue papers
15as first author
48since 2021 · last 2026
0000-0001-7893-9147ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 74 · 9 first-author · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 64 · 14 first-author · 24 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Probing Discrete Speech Tokens of Spoken Language Models
Sven Naber, Julia Koch, Alberto Saponaro, Ioanna Karagianni, Ngoc Thang Vu |
LREC | 6 |
| 2025 | Evaluating the Reliability and Utility of GPT-4o as a Medical Expert Across Different Interaction ModesabstractGenerative Large Language Models (LLMs) such as GPT-4o are increasingly explored for use in clinical decision support, yet their real-world reliability and utility remain underexplored, especially in domains where factual accuracy is critical. This paper presents an empirical evaluation of GPT-40 in a medical expert role, focusing on its ability to provide clinically useful and reliable responses across varying degrees of external control. We evaluate three interaction strategies: Default Interaction, Hard Prompting, and Retrieval-augmented Generation (RAG), using a curated dataset of 70 authentic medical queries authored by practicing physicians. The practicing medical professionals evaluate the generated responses along two key dimensions: factual reliability and clinical utility. Our findings reveal critical trade-offs across the control spectrum: while RAG and hard prompting offer more constrained and verifiable responses, the default interaction approach achieved the highest combined ratings for reliability and clinical utility, showing the lowest abstention rate among the evaluated strategies. This study contributes a practical evaluation framework for assessing the potential of LLMs as medical experts and offers actionable insights for their deployment in healthcare contexts where factual precision, usability, and trust are paramount. Ufkun-Bayram Menderes, Dominik Morar, Ngoc Thang Vu |
BIBM | 3 |
| 2025 | A Survey of Code-switched Arabic NLP: Progress, Challenges, and Future DirectionsabstractLanguage in the Arab world presents a complex diglossic and multilingual setting, involving the use of Modern Standard Arabic, various dialects and sub-dialects, as well as multiple European languages. This diverse linguistic landscape has given rise to code-switching, both within Arabic varieties and between Arabic and foreign languages. The widespread occurrence of code-switching across the region makes it vital to address these linguistic needs when developing language technologies. In this paper, we provide a review of the current literature in the field of code-switched Arabic NLP, offering a broad perspective on ongoing efforts, challenges, research gaps, and recommendations for future research directions. Injy Hamed, Caroline Sabty, Slim Abdennadher, Ngoc Thang Vu, Thamar Solorio, Nizar Habash |
COLING | 4 |
| 2025 | Discrete Subgraph Sampling for Interpretable Graph based Visual Question AnsweringabstractExplainable artificial intelligence (XAI) aims to make machine learning models more transparent. While many approaches focus on generating explanations post-hoc, interpretable approaches, which generate the explanations intrinsically alongside the predictions, are relatively rare. In this work, we integrate different discrete subset sampling methods into a graph-based visual question answering system to compare their effectiveness in generating interpretable explanatory subgraphs intrinsically. We evaluate the methods on the dataset and show that the integrated methods effectively mitigate the performance trade-off between interpretability and answer accuracy, while also achieving strong co-occurrences between answer and question tokens. Furthermore, we conduct a human evaluation to assess the interpretability of the generated subgraphs using a comparative setting with the extended Bradley-Terry model, showing that the answer and question token co-occurrence metrics strongly correlate with human preferences. Our source code is publicly available. Pascal Tilli, Ngoc Thang Vu |
COLING | 2 |
| 2025 | It's What You Say and How You Say It: Investigating the Effect of Linguistic vs. Behavioral Adaptation in Task-Oriented ChatbotsabstractGiven the conflicting expectations users have for how a dialog agent should sound and behave, there is no one-size-fits-all option for dialog system design. Therefore, adaptation is critical to ensure successful and enjoyable interactions. However, it is not yet clear what the effects of behavioral (what the agent says) vs. linguistic adaptation (how the agent says this) are in terms of dialog success and user perception. In this work, we implement three different types of task-oriented dialog agents which can each vary their level of formality. We evaluate subjective and objective metrics of dialog success as well as user perceptions through a user study, comparing the collected data to that of (CITATION), where users interacted with the same three types of agents without linguistic adaptation. From this, we draw insights into which subjective and objective aspects of success and user perception are influenced by each type of adaptation. We additionally all code, user surveys, and dialog interaction logs. Lindsey Vanderlyn, Ngoc Thang Vu |
COLING | 2 |
| 2025 | High-Resolution Speech Restoration with Latent Diffusion ModelabstractTraditional speech enhancement methods often oversimplify the task of restoration by focusing on a single type of distortion. Generative models that handle multiple distortions frequently struggle with phone reconstruction and high-frequency harmonics, leading to breathing and gasping artifacts that reduce the intelligibility of reconstructed speech. These models are also computationally demanding, and many solutions are restricted to producing outputs in the wideband frequency range, which limits their suitability for professional applications. To address these challenges, we propose Hi-ResLDM, a novel generative model based on latent diffusion designed to remove multiple distortions and restore speech recordings to studio quality at a full-band sampling rate of 48kHz. Benchmarked against state-of-the-art methods that leverage GAN and Conditional Flow Matching (CFM) components, Hi-ResLDM demonstrates superior performance in regenerating high-frequency-band details. Hi-ResLDM not only excels in non-instrusive metrics but is also consistently preferred in human evaluation and performs competitively on intrusive evaluations, making it ideal for high-resolution speech restoration. Tushar Dhyani, Florian Lux, Michele Mancusi, Giorgio Fabbro, Fritz Hohl, Ngoc Thang Vu |
ICASSP | 6 |
| 2025 | What Affects the Performance of Fake Audio Detection? Analyzing Factors in a Continual Learning SettingabstractThe increasing sophistication of deepfake audio generation technologies makes it important to develop robust fake audio detection systems that can adapt over time. This study examines how various factors impact the performance of detection systems in a continual learning setting. We focus on factors such as attacker architectures, attackers’ training datasets, speaker diversity, and task order. We evaluate the performance of three detection models trained with four different strategies, including direct fine-tuning, one-class classification, random replay, and Learning without Forgetting. Results show that artifacts from the fake audios might arise from the attackers’ training datasets, and simply changing attacker architectures does not sufficiently challenge detection systems. Moreover, task order and speaker diversity can significantly influence performance, with varying degrees of sensitivity across different detection models and training strategies. These insights underline the need for careful consideration of these factors when developing robust detection systems. Yixuan Xiao, Ngoc Thang Vu |
ICASSP | 2 |
| 2025 | Investigating Stochastic Methods for Prosody Modeling in Speech Synthesis
Paul Mayer, Florian Lux, Alejandro Pérez González de Martos, Angelina Elizarova, Lindsey Vanderlyn, Dirk Väth, Ngoc Thang Vu |
INTERSPEECH | 7 |
| 2025 | First Steps Towards Voice Anonymization for Code-Switching Speech
Sarina Meyer, Ekaterina Kolos, Ngoc Thang Vu |
INTERSPEECH | 3 |
| 2025 | Layer-Wise Decision Fusion for Fake Audio Detection Using XLS-RabstractRecent fake audio detection methods often leverage large speech models to achieve robust speech representations. These models are typically very deep, providing multiple layer-wise representations. However, current works often rely solely on single layer representation or feature fusion to extract one utterance-level representation for decision making. These methods risk underutilizing rich information from multiple layers and might induce feature collapse. We propose a novel layer-wise decision fusion method that applies fusion after per-layer decision making and achieves the best cross-dataset performance on In-the-Wild dataset (EER 6.90%) compared to other strong baselines. Our model design also makes the model more transparent, allowing us to conduct detailed analysis to reveal the underlying mechanism of decision making. Yixuan Xiao, Ngoc Thang Vu |
INTERSPEECH | 2 |
| 2024 | Prompting-based Synthetic Data Generation for Few-Shot Question AnsweringabstractAlthough language models (LMs) have boosted the performance of Question Answering, they still need plenty of data. Data annotation, in contrast, is a time-consuming process. This especially applies to Question Answering, where possibly large documents have to be parsed and annotated with questions and their corresponding answers. Furthermore, Question Answering models often only work well for the domain they were trained on. Since annotation is costly, we argue that domain-agnostic knowledge from LMs, such as linguistic understanding, is sufficient to create a well-curated dataset. With this motivation, we show that using large language models can improve Question Answering performance on various datasets in the few-shot setting compared to state-of-the-art approaches. For this, we perform data generation leveraging the Prompting framework, suggesting that language models contain valuable task-agnostic knowledge that can be used beyond the common pre-training/fine-tuning scheme. As a result, we consistently outperform previous approaches on few-shot Question Answering. Andrea Bartezzaghi, Ngoc Thang Vu |
LREC/COLING | 3 |
| 2024 | Intrinsic Subgraph Generation for Interpretable Graph Based Visual Question AnsweringabstractThe large success of deep learning based methods in Visual Question Answering (VQA) has concurrently increased the demand for explainable methods. Most methods in Explainable Artificial Intelligence (XAI) focus on generating post-hoc explanations rather than taking an intrinsic approach, the latter characterizing an interpretable model. In this work, we introduce an interpretable approach for graph-based VQA and demonstrate competitive performance on the GQA dataset. This approach bridges the gap between interpretability and performance. Our model is designed to intrinsically produce a subgraph during the question-answering process as its explanation, providing insight into the decision making. To evaluate the quality of these generated subgraphs, we compare them against established post-hoc explainability methods for graph neural networks, and perform a human evaluation. Moreover, we present quantitative metrics that correlate with the evaluations of human assessors, acting as automatic metrics for the generated explanatory subgraphs. Our code will be made publicly available at link removed due to anonymity period. Pascal Tilli, Ngoc Thang Vu |
LREC/COLING | 2 |
| 2024 | Towards a Zero-Data, Controllable, Adaptive Dialog SystemabstractConversational Tree Search (Väth et al., 2023) is a recent approach to controllable dialog systems, where domain experts shape the behavior of a Reinforcement Learning agent through a dialog tree. The agent learns to efficiently navigate this tree, while adapting to information needs, e.g., domain familiarity, of different users. However, the need for additional training data hinders deployment in new domains. To address this, we explore approaches to generate this data directly from dialog trees. We improve the original approach, and show that agents trained on synthetic data can achieve comparable dialog success to models trained on human data, both when using a commercial Large Language Model for generation, or when using a smaller open-source model, running on a single GPU. We further demonstrate the scalability of our approach by collecting and testing on two new datasets: ONBOARD, a new domain helping foreign residents moving to a new city, and the medical domain DIAGNOSE, a subset of Wikipedia articles related to scalp and head symptoms. Finally, we perform human testing, where no statistically significant differences were found in either objective or subjective measures between models trained on human and generated data. Dirk Väth, Lindsey Vanderlyn, Ngoc Thang Vu |
LREC/COLING | 3 |
| 2024 | Explaining Pre-Trained Language Models with Attribution Scores: An Analysis in Low-Resource SettingsabstractAttribution scores indicate the importance of different input parts and can, thus, explain model behaviour. Currently, prompt-based models are gaining popularity, i.a., due to their easier adaptability in low-resource settings. However, the quality of attribution scores extracted from prompt-based models has not been investigated yet. In this work, we address this topic by analyzing attribution scores extracted from prompt-based models w.r.t. plausibility and faithfulness and comparing them with attribution scores extracted from fine-tuned models and large language models. In contrast to previous work, we introduce training size as another dimension into the analysis. We find that using the prompting paradigm (with either encoder-based or decoder-based models) yields more plausible explanations than fine-tuning the models in low-resource settings and Shapley Value Sampling consistently outperforms attention and Integrated Gradients in terms of leading to more plausible and faithful explanations. Wei Zhou 0067, Heike Adel, Hendrik Schuff, Ngoc Thang Vu |
LREC/COLING | 4 |
| 2024 | Combining Data Generation and Active Learning for Low-Resource Question Answering
Maximilian Kimmich, Andrea Bartezzaghi, Jasmina Bogojeska, Cristiano Malossi, Ngoc Thang Vu |
ICANN (7) | 5 |
| 2024 | Controlling Emotion in Text-to-Speech with Natural Language Prompts
Thomas Bott, Florian Lux, Ngoc Thang Vu |
INTERSPEECH | 3 |
| 2024 | Meta Learning Text-to-Speech Synthesis in over 7000 Languagesabstract4958 Florian Lux, Sarina Meyer, Lyonel Behringer, Frank Zalkow, Phat Do, Matt Coler, Emanuël A. P. Habets, Ngoc Thang Vu |
INTERSPEECH | 8 |
| 2024 | Probing the Feasibility of Multilingual Speaker Anonymization
Sarina Meyer, Florian Lux, Ngoc Thang Vu |
INTERSPEECH | 3 |
| 2023 | Ethical Considerations for Machine Translation of Indigenous Languages: Giving a Voice to the SpeakersabstractIn recent years machine translation has become very successful for high-resource language pairs.This has also sparked new interest in research on the automatic translation of lowresource languages, including Indigenous languages.However, the latter are deeply related to the ethnic and cultural groups that speak (or used to speak) them.The data collection, modeling and deploying machine translation systems thus result in new ethical questions that must be addressed.Motivated by this, we first survey the existing literature on ethical considerations for the documentation, translation, and general natural language processing for Indigenous languages.Afterward, we conduct and analyze an interview study to shed light on the positions of community leaders, teachers, and language activists regarding ethical concerns for the automatic translation of their languages.Our results show that the inclusion, at different degrees, of native speakers and community members is vital to performing better and more ethical research on Indigenous languages. Manuel Mager, Elisabeth Mager, Katharina Kann, Ngoc Thang Vu |
ACL (1) | 4 |
| 2023 | Leveraging Multilingual Self-Supervised Pretrained Models for Sequence-to-Sequence End-to-End Spoken Language UnderstandingabstractA number of methods have been proposed for End-to-End Spoken Language Understanding (E2E-SLU) using pretrained models, however their evaluation often lacks multilingual setup and tasks that require prediction of lexical fillers, such as slot filling. In this work, we propose a unified method that integrates multilingual pretrained speech and text models and performs E2E-SLU on six datasets in four languages in a generative manner, including the prediction of lexical fillers. We investigate how the proposed method can be improved by pretraining on widely available speech recognition data using several training objectives. Pretraining on 7000 hours of multilingual data allows us to outperform the state-of-the-art ultimately on two SLU datasets and partly on two more SLU datasets. Finally, we examine the cross-lingual capabilities of the proposed model and improve on the best known result on the PortMEDIA-Language dataset by almost half, achieving a Concept/Value Error Rate of 23.65%. Pavel Denisov, Ngoc Thang Vu |
ASRU | 2 |
| 2023 | HNC: Leveraging Hard Negative Captions towards Models with Fine-Grained Visual-Linguistic Comprehension CapabilitiesabstractImage-Text-Matching (ITM) is one of the defacto methods of learning generalized representations from a large corpus in Vision and Language (VL).However, due to the weak association between the web-collected image-text pairs, models fail to show a fine-grained understanding of the combined semantics of these modalities.To address this issue we propose Hard Negative Captions (HNC): an automatically created dataset containing foiled hard negative captions for ITM training towards achieving fine-grained cross-modal comprehension in VL.Additionally, we provide a challenging manually-created test set for benchmarking models on a fine-grained cross-modal mismatch task with varying levels of compositional complexity.Our results show the effectiveness of training on HNC by improving the models' zero-shot capabilities in detecting mismatches on diagnostic tasks and performing robustly under noisy visual input scenarios.Also, we demonstrate that HNC models yield a comparable or better initialization for fine-tuning.Our code and data are publicly available. Esra Dönmez, Pascal Tilli, Hsiu-Yu Yang, Ngoc Thang Vu, Carina Silberer |
CoNLL | 4 |
| 2023 | Exploring Segmentation Approaches for Neural Machine Translation of Code-Switched Egyptian Arabic-English TextabstractData sparsity is one of the main challenges posed by code-switching (CS), which is further exacerbated in the case of morphologically rich languages.For the task of machine translation (MT), morphological segmentation has proven successful in alleviating data sparsity in monolingual contexts; however, it has not been investigated for CS settings.In this paper, we study the effectiveness of different segmentation approaches on MT performance, covering morphology-based and frequency-based segmentation techniques.We experiment on MT from code-switched Arabic-English to English.We provide detailed analysis, examining a variety of conditions, such as data size and sentences with different degrees of CS.Empirical results show that morphology-aware segmenters perform the best in segmentation tasks but under-perform in MT.Nevertheless, we find that the choice of the segmentation setup to use for MT is highly dependent on the data size.For extreme low-resource scenarios, a combination of frequency and morphology-based segmentations is shown to perform the best.For more resourced settings, such a combination does not bring significant improvements over the use of frequency-based segmentation. Marwa Gaser, Manuel Mager, Injy Hamed, Nizar Habash, Slim Abdennadher, Ngoc Thang Vu |
EACL | 6 |
| 2023 | Conversational Tree Search: A New Hybrid Dialog TaskabstractConversational interfaces provide a flexible and easy way for users to seek information that may otherwise be difficult or inconvenient to obtain.However, existing interfaces generally fall into one of two categories: FAQs, where users must have a concrete question in order to retrieve a general answer, or dialogs, where users must follow a predefined path but may receive a personalized answer.In this paper, we introduce Conversational Tree Search (CTS) as a new task that bridges the gap between FAQ-style information retrieval and task-oriented dialog, allowing domain-experts to define dialog trees which can then be converted to an efficient dialog policy that learns only to ask the questions necessary to navigate a user to their goal.We collect a dataset for the travel reimbursement domain and demonstrate a baseline as well as a novel deep Reinforcement Learning architecture for this task.Our results show that the new architecture combines the positive aspects of both the FAQ and dialog system used in the baseline and achieves higher goal completion while skipping unnecessary questions. Dirk Väth, Lindsey Vanderlyn, Ngoc Thang Vu |
EACL | 3 |
| 2023 | Prosody Is Not Identity: A Speaker Anonymization Approach Using Prosody CloningabstractProsody is closely linked to the identity of a speaker, leading to individual pitch and intonation patterns. Therefore, it is challenging in speaker anonymization to generate speech utterances that both keep the original audio’s main prosodic structure and preserve the speaker’s privacy. In this paper, we present a system that extends a speech-to-text-to-speech anonymization pipeline with prosody cloning and show how to control the cloning by multiplying pitch and energy sequences with random offset values. Using automatic and human evaluation, we find this combination to successfully overcome the privacy-utility trade-off for prosody by achieving high privacy and high pitch correlation scores. At the same time, the anonymized utterances prove to reproduce the original voice distinctiveness and content with high intelligibility and only a small loss in naturalness, making them suitable for downstream applications. Sarina Meyer, Florian Lux, Julia Koch, Pavel Denisov, Pascal Tilli, Ngoc Thang Vu |
ICASSP | 6 |
| 2023 | Regularisation for Efficient Softmax Parameter Generation in Low-Resource Text ClassifiersabstractMeta-learning has made tremendous progress in recent years and was demonstrated to be particularly suitable in low-resource settings where training data is very limited. However, meta-learning models still require large amounts of training tasks to achieve good generalisation. Since labelled training data may be sparse, self-supervision-based approaches are able to further improve performance on downstream tasks. Although no labelled data is necessary for this training, a large corpus of unlabelled text needs to be available. In this paper, we improve on recent advances in meta-learning for natural language models that allow training on a diverse set of training tasks for few-shot, low-resource target tasks. We introduce a way to generate new training data with the need for neither more supervised nor unsupervised datasets. We evaluate the method on a diverse set of NLP tasks and show that the model decreases in performance when trained on this data without further adjustments. Therefore, we introduce and evaluate two methods for regularising the training process and show that they not only improve performance when used in conjunction with the new training data but also improve average performance when training only on the original data, compared to the baseline. Daniel Grießhaber, Johannes Maucher, Ngoc Thang Vu |
IJCAI | 3 |
| 2023 | Controllable Generation of Artificial Speaker Embeddings through Discovery of Principal DirectionsabstractCustomizing voice and speaking style in a speech synthesis system with intuitive and fine-grained controls is challenging, given that little data with appropriate labels is available. Furthermore, editing an existing human's voice also comes with ethical concerns. In this paper, we propose a method to generate artificial speaker embeddings that cannot be linked to a real human while offering intuitive and fine-grained control over the voice and speaking style of the embeddings, without requiring any labels for speaker or style. The artificial and controllable embeddings can be fed to a speech synthesis system, conditioned on embeddings of real humans during training, without sacrificing privacy during inference. Florian Lux, Pascal Tilli, Sarina Meyer, Ngoc Thang Vu |
INTERSPEECH | 4 |
| 2023 | Visual Analysis of Scene-Graph-Based Visual Question AnsweringabstractScene-graph-based Visual Question Answering (VQA) has emerged as a burgeoning field in Deep Learning research, with a growing demand for robust and interpretable VQA systems. In this paper, we present a novel visual analysis approach that addresses two critical objectives in VQA: identifying and correcting prediction issues and providing insights into model decision-making processes through visualizing internal information. Our approach builds on the GraphVQA framework, which uses graph neural networks to process scene graphs representing images and which was trained on the widely-used GQA dataset. Our analysis tool aims at users familiar with the basics of graph-based VQA. By leveraging query-based scene analysis and visualization of crucial internal states, we are able to detect and pinpoint reasons for inaccurate predictions, facilitating model refinement and dataset curation. Identifying expressive internal states is a challenge. Through rigorous computer-based evaluations and presentation of a use case, we demonstrate the effectiveness of our analysis tool and model state visualization. Noel Schäfer, Sebastian Künzel, Tanja Munz-Körner, Pascal Tilli, Sandeep Vidyapu, Ngoc Thang Vu, Daniel Weiskopf |
VINCI | 6 |
| 2023 | Ethical Awareness in Paralinguistics: A Taxonomy of ApplicationsabstractSince the end of the last century, the automatic processing of paralinguistics has been investigated widely and put into practice in many applications, on wearables, smartphones, and computers. In this contribution, we address ethical awareness for paralinguistic applications, by establishing taxonomies for data representations, system designs for and a typology of applications, and users/test sets and subject areas. These are related to an “ethical grid” consisting of the most relevant ethical cornerstones, based on principalism. The characteristics of and the interdependencies between these taxonomies are described and exemplified. This makes it possible to assess more or less critical “ethical constellations.” To the best of our knowledge, this is the first attempt of its kind. Anton Batliner, Michael Neumann 0001, Felix Burkhardt, Alice Baird, Sarina Meyer, Ngoc Thang Vu, Björn W. Schuller |
Int. J. Hum. Comput. Interact. | 6 |
| 2023 | How to do human evaluation: A brief introduction to user studies in NLPabstractAbstract Many research topics in natural language processing (NLP), such as explanation generation, dialog modeling, or machine translation, require evaluation that goes beyond standard metrics like accuracy or F1score toward a more human-centered approach. Therefore, understanding how to design user studies becomes increasingly important. However, few comprehensive resources exist on planning, conducting, and evaluating user studies for NLP, making it hard to get started for researchers without prior experience in the field of human evaluation. In this paper, we summarize the most important aspects of user studies and their design and evaluation, providing direct links to NLP tasks and NLP-specific challenges where appropriate. We (i) outline general study design, ethical considerations, and factors to consider for crowdsourcing, (ii) discuss the particularities of user studies in NLP, and provide starting points to select questionnaires, experimental designs, and evaluation methods that are tailored to the specific NLP tasks. Additionally, we offer examples with accompanying statistical evaluation code, to bridge the gap between theoretical guidelines and practical applications. Hendrik Schuff, Lindsey Vanderlyn, Heike Adel, Ngoc Thang Vu |
Nat. Lang. Eng. | 4 |
| 2022 | AmericasNLI: Evaluating Zero-shot Natural Language Understanding of Pretrained Multilingual Models in Truly Low-resource LanguagesabstractAbteen Ebrahimi, Manuel Mager, Arturo Oncevay, Vishrav Chaudhary, Luis Chiruzzo, Angela Fan, John Ortega, Ricardo Ramos, Annette Rios, Ivan Vladimir Meza Ruiz, Gustavo Giménez-Lugo, Elisabeth Mager, Graham Neubig, Alexis Palmer, Rolando Coto-Solano, Thang Vu, Katharina Kann. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Abteen Ebrahimi, Manuel Mager, Arturo Oncevay, Vishrav Chaudhary, Luis Chiruzzo, Angela Fan, John E. Ortega, Ricardo Ramos, Annette Rios, Iván V. Meza, Gustavo Giménez Lugo, Elisabeth Mager, Graham Neubig, Alexis Palmer, Rolando Coto-Solano, Ngoc Thang Vu, Katharina Kann |
ACL (1) | 16 |
| 2022 | Language-Agnostic Meta-Learning for Low-Resource Text-to-Speech with Articulatory FeaturesabstractWhile neural text-to-speech systems perform remarkably well in high-resource scenarios, they cannot be applied to the majority of the over 6,000 spoken languages in the world due to a lack of appropriate training data.In this work, we use embeddings derived from articulatory vectors rather than embeddings derived from phoneme identities to learn phoneme representations that hold across languages.In conjunction with language agnostic meta learning, this enables us to fine-tune a high-quality textto-speech model on just 30 minutes of data in a previously unseen language spoken by a previously unseen speaker. Florian Lux, Ngoc Thang Vu |
ACL (1) | 2 |
| 2022 | ESPnet-SLU: Advancing Spoken Language Understanding Through ESPnetabstractAs Automatic Speech Processing (ASR) systems are getting better, there is an increasing interest of using the ASR output to do downstream Natural Language Processing (NLP) tasks. However, there are few open source toolkits that can be used to generate reproducible results on different Spoken Language Understanding (SLU) benchmarks. Hence, there is a need to build an open source standard that can be used to have a faster start into SLU research. We present ESPnet-SLU, which is designed for quick development of spoken language understanding in a single framework. ESPnet-SLU is a project inside end-to-end speech processing toolkit, ESPnet, which is a widely used open-source standard for various speech processing tasks like ASR, Text to Speech (TTS) and Speech Translation (ST). We enhance the toolkit to provide implementations for various SLU benchmarks that enable researchers to seamlessly mix-and-match different ASR and NLU models. We also provide pretrained models with intensively tuned hyper-parameters that can match or even outperform the current state-of-the-art performances. The toolkit is publicly available at https://github.com/espnet/espnet. Siddhant Arora, Siddharth Dalmia, Pavel Denisov, Xuankai Chang, Yushi Ueda, Yifan Peng 0003, Yuekai Zhang, Sujay Kumar, Karthik Ganesan 0003, Brian Yan, Ngoc Thang Vu, Alan W. Black, Shinji Watanabe 0001 |
ICASSP | 11 |
| 2022 | PoeticTTS - Controllable Poetry Reading for Literary StudiesabstractSpeech synthesis for poetry is challenging due to specific intonation patterns inherent to poetic speech.In this work, we propose an approach to synthesise poems with almost human like naturalness in order to enable literary scholars to systematically examine hypotheses on the interplay between text, spoken realisation, and the listener's perception of poems.To meet these special requirements for literary studies, we resynthesise poems by cloning prosodic values from a human reference recitation, and afterwards make use of fine-grained prosody control to manipulate the synthetic speech in a human-in-the-loop setting to alter the recitation w.r.t.specific phenomena.We find that finetuning our TTS model on poetry captures poetic intonation patterns to a large extent which is beneficial for prosody cloning and manipulation and verify the success of our approach both in an objective evaluation as well as in human studies. Julia Koch, Florian Lux, Nadja Schauffler, Toni Bernhart, Felix Dieterle, Jonas Kuhn, Sandra Richter, Gabriel Viehhauser, Ngoc Thang Vu |
INTERSPEECH | 9 |
| 2022 | Speaker Anonymization with Phonetic Intermediate Representations
Sarina Meyer, Florian Lux, Pavel Denisov, Julia Koch, Pascal Tilli, Ngoc Thang Vu |
INTERSPEECH | 6 |
| 2022 | »textklang« - Towards a Multi-Modal Exploration Platform for German PoetryabstractWe present the steps taken towards an exploration platform for a multi-modal corpus of German lyric poetry from the Romantic era developed in the project »textklang«. This interdisciplinary project develops a mixed-methods approach for the systematic investigation of the relationship between written text (here lyric poetry) and its potential and actual sonic realisation (in recitations, musical performances etc.). The multi-modal »textklang« platform will be designed to technically and analytically combine three modalities: the poetic text, the audio signal of a recorded recitation and, at a later stage, music scores of a musical setting of a poem. The methodological workflow will enable scholars to develop hypotheses about the relationship between textual form and sonic/prosodic realisation based on theoretical considerations, text interpretation and evidence from recorded recitations. The full workflow will support hypothesis testing either through systematic corpus analysis alone or with addtional contrastive perception experiments. For the experimental track, researchers will be enabled to manipulate prosodic parameters in (re-)synthesised variants of the original recordings. The focus of this paper is on the design of the base corpus and on tools for systematic exploration – placing special emphasis on our response to challenges stemming from multi-modality and the methodologically diverse interdisciplinary setup. Nadja Schauffler, Toni Bernhart, André Blessing, Gunilla Eschenbach, Markus Gärtner, Kerstin Jung, Anna Kinder, Julia Koch, Sandra Richter, Gabriel Viehhauser, Ngoc Thang Vu, Lorenz Wesemann, Jonas Kuhn |
LREC | 11 |
| 2022 | Improving Semi-Supervised End-To-End Automatic Speech Recognition Using Cyclegan and Inter-Domain LossesabstractWe propose a novel method that combines CycleGAN and inter-domain losses for semi-supervised end-to-end automatic speech recognition. Inter-domain loss targets the extraction of an intermediate shared representation of speech and text inputs using a shared network. CycleGAN uses cycle-consistent loss and the identity mapping loss to preserve relevant characteristics of the input feature after converting from one domain to another. As such, both approaches are suitable to train end-to-end models on unpaired speech-text inputs. In this paper, we exploit the advantages from both inter-domain loss and CycleGAN to achieve better shared representation of unpaired speech and text inputs and thus improve the speech-to-text mapping. Our experimental results on the WSJ eval92 and Voxforge (non English) show$8\sim 8.5\%$character error rate reduction over the baseline, and the results on LibriSpeech test_clean also show noticeable improvement. Chia-Yu Li, Ngoc Thang Vu |
SLT | 2 |
| 2022 | Combining Contrastive and Non-Contrastive Losses for Fine-Tuning Pretrained Models in Speech AnalysisabstractEmbedding paralinguistic properties is a challenging task as there are only a few hours of training data available for domains such as emotional speech. One solution to this problem is to pretrain a general self-supervised speech representation model on large amounts of unlabeled speech. This pretrained model is then finetuned to a specific task. Paralinguistic properties however have notoriously high class variance, making the finetuning ineffective. In this work, we propose a two step approach to this. First we improve the embedding space, then we train an adapter to bridge the gap from the embedding space to a classification task. In order to improve the class invariance we use a combination of contrastive and non-contrastive losses to explicitly optimize for class invariant, yet discriminative features. Our approach consistently outperforms baselines that are finetuned end-to-end on multiple tasks and surpasses a benchmark on state-of-the-art emotion classification. Florian Lux, Ching-Yi Chen, Ngoc Thang Vu |
SLT | 3 |
| 2022 | Exact Prosody Cloning in Zero-Shot Multispeaker Text-to-SpeechabstractThe cloning of a speaker's voice using an untranscribed reference sample is one of the great advances of modern neural text-to-speech (TTS) methods. Approaches for mimicking the prosody of a transcribed reference audio have also been proposed recently. In this work, we bring these two tasks together for the first time through utterance level normalization in conjunction with an utterance level speaker embedding. We further introduce a lightweight aligner for extracting fine-grained prosodic features, that can be finetuned on individual samples within seconds. We show that it is possible to clone the voice of a speaker as well as the prosody of a spoken reference independently without any degradation in quality and high similarity to both original voice and prosody, as our objective evaluation and human study show. All of our code and trained models are available, alongside static and interactive demos. Florian Lux, Julia Koch, Ngoc Thang Vu |
SLT | 3 |
| 2022 | Anonymizing Speech with Generative Adversarial Networks to Preserve Speaker PrivacyabstractIn order to protect the privacy of speech data, speaker anonymization aims for hiding the identity of a speaker by changing the voice in speech recordings. This typically comes with a privacy-utility trade-off between protection of individuals and usability of the data for downstream applications. One of the challenges in this context is to create non-existent voices that sound as natural as possible. In this work, we propose to tackle this issue by generating speaker embeddings using a generative adversarial network with Wasserstein distance as cost function. By incorporating these artificial embeddings into a speech-to-text-to-speech pipeline, we outperform previous approaches in terms of privacy and utility. According to standard objective metrics and human evaluation, our approach generates intelligible and content-preserving yet privacy-protecting versions of the original recordings. Sarina Meyer, Pascal Tilli, Pavel Denisov, Florian Lux, Julia Koch, Ngoc Thang Vu |
SLT | 6 |
| 2022 | Visualization-based improvement of neural machine translation
Tanja Munz-Körner, Dirk Väth, Paul Kuznecov, Ngoc Thang Vu, Daniel Weiskopf |
Comput. Graph. | 4 |
| 2022 | Investigations on speech recognition systems for low-resource dialectal Arabic-English code-switching speech
Injy Hamed, Pavel Denisov, Chia-Yu Li, Mohamed Elmahdy 0001, Slim Abdennadher, Ngoc Thang Vu |
Comput. Speech Lang. | 6 |
| 2021 | Improving Speech Recognition on Noisy Speech via Speech Enhancement with Multi-Discriminators CycleGANabstractThis paper presents our latest investigations on improving automatic speech recognition for noisy speech via speech enhancement. We propose a novel method named Multi-discriminators CycleGAN to reduce noise of input speech and therefore improve the automatic speech recognition performance. Our proposed method leverages the CycleGAN framework for speech enhancement without any parallel data and improve it by introducing multiple discriminators that check different frequency areas. Furthermore, we show that training multiple generators on homogeneous subset of the training data is better than training one generator on all the training data. We evaluate our method on CHiME-3 data set and observe up to 10.03 % relatively WER improvement on the development set and up to 14.09 % on the evaluation set. Chia-Yu Li, Ngoc Thang Vu |
ASRU | 2 |
| 2021 | "It seemed like an annoying woman": On the Perception and Ethical Considerations of Affective Language in Text-Based Conversational AgentsabstractPrevious research has found that task-oriented conversational agents are perceived more positively by users when they provide information in an empathetic manner compared to a plain, emotionless information exchange.However, users' perception and ethical considerations related to a dialog systems' response language style have received comparatively little attention in the field of human-computer interaction.To bridge this gap, we explored these ethical implications through a scenario-based user study.127 participants interacted with one of three variants of an affective, task-oriented conversational agent, each variant providing responses in a different language style.After the interaction, participants filled out a survey about their feelings during the experiment and their perception of various aspects of the chatbot.Based on statistical and qualitative analysis of the responses, we found language style played an important role in how humanlike participants perceived a dialog agent as well as how likable.Language style also had a direct effect on how users perceived the use of personal pronouns 'I' and 'You' and how they projected gender onto the chatbot.Finally, we identify and discuss ethical implications.In particular we focus on what factors/stereotypes influenced participants' impressions of gender, and what trade-offs a more human-like chatbot brings. Lindsey Vanderlyn, Gianna Weber, Michael Neumann 0001, Dirk Väth, Sarina Meyer, Ngoc Thang Vu |
CoNLL | 6 |
| 2021 | "It's our fault!": Insights Into Users' Understanding and Interaction With an Explanatory Collaborative Dialog SystemabstractHuman-AI collaboration, a long standing goal in AI, refers to a partnership where a human and artificial intelligence work together towards a shared goal. Collaborative dialog allows human-AI teams to communicate and leverage strengths from both partners. To design collaborative dialog systems, it is important to understand what mental models users form about their AI-dialog partners, however, how users perceive these systems is not fully understood. In this study, we designed a novel, collaborative, communication-based puzzle game and explanatory dialog system. We created a public corpus from 117 conversations and post-surveys and used this to analyze what mental models users formed. Key takeaways include: Even when users were not engaged in the game, they perceived the AI-dialog partner as intelligent and likeable, implying they saw it as a partner separate from the game. This was further supported by users often overestimating the system’s abilities and projecting human-like attributes which led to miscommunications. We conclude that creating shared mental models between users and AI systems is important to achieving successful dialogs. We propose that our insights on mental models and miscommunication, the game, and our corpus provide useful tools for designing collaborative dialog systems. Katharina Weitz, Lindsey Vanderlyn, Ngoc Thang Vu, Elisabeth André |
CoNLL | 3 |
| 2021 | Few-shot Learning for Slot Tagging with Attentive Relational NetworkabstractMetric-based learning is a well-known family of methods for few-shot learning, especially in computer vision.Recently, they have been used in many natural language processing applications but not for slot tagging.In this paper, we explore metric-based learning methods in the slot tagging task and propose a novel metric-based learning architecture -Attentive Relational Network.Our proposed method extends relation networks, making them more suitable for natural language processing applications in general, by leveraging pretrained contextual embeddings such as ELMO and BERT and by using attention mechanism.The results on SNIPS data show that our proposed method outperforms other state of the art metric-based learning methods. Cennet Oguz, Ngoc Thang Vu |
EACL | 2 |
| 2021 | Visual-Interactive Neural Machine TranslationabstractWe introduce a novel visual analytics approach for analyzing, understanding, and correcting neural machine translation. Our system supports users in automatically translating documents using neural machine translation and identifying and correcting possible erroneous translations. User corrections can then be used to fine-tune the neural machine translation model and automatically improve the whole document. While translation results of neural machine translation can be impressive, there are still many challenges such as overand under-translation, domain-specific terminology, and handling long sentences, making it necessary for users to verify translation results; our system aims at supporting users in this task. Our visual analytics approach combines several visualization techniques in an interactive system. A parallel coordinates plot with multiple metrics related to translation quality can be used to find, filter, and select translations that might contain errors. An interactive beam search visualization and graph visualization for attention weights can be used for post-editing and understanding machine-generated translations. The machine translation model is updated from user corrections to improve the translation quality of the whole document. We designed our approach for an LSTM-based translation model and extended it to also include the Transformer architecture. We show for representative examples possible mistranslations and how to use our system to deal with them. A user study revealed that many participants favor such a system over manual text-based translation, especially for translating large documents. Tanja Munz-Körner, Dirk Väth, Paul Kuznecov, Ngoc Thang Vu, Daniel Weiskopf |
Graphics Interface | 4 |
| 2021 | Meta-Learning for Improving Rare Word Recognition in End-to-End ASRabstractIn this work we take on the challenge of rare word recognition in end-to-end (E2E) automatic speech recognition (ASR) by integrating a meta learning mechanism into an E2E ASR system, enabling few-shot adaptation. We propose a novel method of generating embeddings for speech, changes to four meta learning approaches, enabling them to perform keyword spotting and an approach to using their outcomes in an E2E ASR system. We verify the functionality of each of our three contributions in two experiments exploring their performance for different amounts of classes (N-way) and examples per class (k-shot) in a few-shot setting. We find that the information encoded in the speech embeddings suffices to allow the modified meta learning approaches to perform continuous signal spotting. Despite the simplicity of the interface between keyword spotting and speech recognition, we are able to consistently improve word error rate by up to 5%. Florian Lux, Ngoc Thang Vu |
ICASSP | 2 |
| 2021 | Investigations on audiovisual emotion recognition in noisy conditionsabstractIn this paper we explore audiovisual emotion recognition under noisy acoustic conditions with a focus on speech features. We attempt to answer the following research questions: (i) How does speech emotion recognition perform on noisy data? and (ii) To what extend does a multimodal approach improve the accuracy and compensate for potential performance degradation at different noise levels? We present an analytical investigation on two emotion datasets with superimposed noise at different signal-to-noise ratios, comparing three types of acoustic features. Visual features are incorporated with a hybrid fusion approach: The first neural network layers are separate modality-specific ones, followed by at least one shared layer before the final prediction. The results show a significant performance decrease when a model trained on clean audio is applied to noisy data and that the addition of visual features alleviates this effect. Michael Neumann 0001, Ngoc Thang Vu |
SLT | 2 |
| 2020 | Fast and Accurate Non-Projective Dependency Tree LinearizationabstractWe propose a graph-based method to tackle the dependency tree linearization task. We formulate the task as a Traveling Salesman Problem (TSP), and use a biaffine attention model to calculate the edge costs. We facilitate the decoding by solving the TSP for each subtree and combining the solution into a projective tree. We then design a transition system as post-processing, inspired by non-projective transition-based parsing, to obtain non-projective sentences. Our proposed method outperforms the state-of-the-art linearizer while being 10 times faster in training and decoding. Simon Tannert, Ngoc Thang Vu, Jonas Kuhn |
ACL | 3 |
| 2020 | ClaVis: An Interactive Visual Comparison System for ClassifiersabstractWe propose ClaVis, a visual analytics system for comparative analysis of classification models. ClaVis allows users to visually compare the performance and behavior of tens to hundreds of classifiers trained with different hyperparameter configurations. Our approach is plugin-based and classifier-agnostic and allows users to add their own datasets and classifier implementations. It provides multiple visualizations, including a multivariate ranking, a similarity map, a scatterplot that reveals correlations between parameters and scores, and a training history chart. We demonstrate the effectivity of our approach in multiple case studies for training classification models in the domain of natural language processing. Frank Heyen, Tanja Munz-Körner, Michael Neumann 0001, Daniel Ortega, Ngoc Thang Vu, Daniel Weiskopf, Michael Sedlmair |
AVI | 5 |
| 2020 | Fine-tuning BERT for Low-Resource Natural Language Understanding via Active LearningabstractRecently, leveraging pre-trained Transformer based language models in down stream, task specific models has advanced state of the art results in natural language understanding tasks.However, only a little research has explored the suitability of this approach in low resource settings with less than 1,000 training data points.In this work, we explore fine-tuning methods of BERT -a pre-trained Transformer based language model -by utilizing pool-based active learning to speed up training while keeping the cost of labeling new data constant.Our experimental results on the GLUE data set show an advantage in model performance by maximizing the approximate knowledge gain of the model when querying from the pool of unlabeled data.Finally, we demonstrate and analyze the benefits of freezing layers of the language model during fine-tuning to reduce the number of trainable parameters, making it more suitable for low-resource settings. Daniel Grießhaber, Johannes Maucher, Ngoc Thang Vu |
COLING | 3 |
| 2020 | Interpreting Attention Models with Human Visual Attention in Machine Reading ComprehensionabstractWhile neural networks with attention mechanisms have achieved superior performance on many natural language processing tasks, it remains unclear to which extent learned attention resembles human visual attention.In this paper, we propose a new method that leverages eye-tracking data to investigate the relationship between human visual attention and neural attention in machine reading comprehension.To this end, we introduce a novel 23 participant eye tracking dataset -MQA-RC, in which participants read movie plots and answered pre-defined questions.We compare state of the art networks based on long shortterm memory (LSTM), convolutional neural models (CNN) and XLNet Transformer architectures.We find that higher similarity to human attention and performance significantly correlates to the LSTM and CNN models.However, we show this relationship does not hold true for the XLNet models -despite the fact that the XLNet performs best on this challenging task.Our results suggest that different architectures seem to learn rather different neural attention strategies and similarity of neural to human attention does not guarantee best performance. Ekta Sood, Simon Tannert, Diego Frassinelli, Andreas Bulling, Ngoc Thang Vu |
CoNLL | 5 |
| 2020 | F1 is Not Enough! Models and Evaluation Towards User-Centered Explainable Question AnsweringabstractExplainable question answering systems predict an answer together with an explanation showing why the answer has been selected. The goal is to enable users to assess the correctness of the system and understand its reasoning process. However, we show that current models and evaluation settings have shortcomings regarding the coupling of answer and explanation which might cause serious issues in user experience. As a remedy, we propose a hierarchical model and a new regularization term to strengthen the answer-explanation coupling as well as two evaluation scores to quantify the coupling. We conduct experiments on the HOTPOTQA benchmark data set and perform a user study. The user study shows that our models increase the ability of the users to judge the correctness of the system and that scores like F1 are not enough to estimate the usefulness of a model in a practical setting with human users. Our scores are better aligned with user experience, making them promising candidates for model selection. Hendrik Schuff, Heike Adel, Ngoc Thang Vu |
EMNLP (1) | 3 |
| 2020 | OH, JEEZ! or UH-HUH? A Listener-Aware Backchannel Predictor on ASR TranscriptionsabstractThis paper presents our latest investigation on modeling backchannel in conversations. Motivated by a proactive backchanneling theory, we aim at developing a system which acts as a proactive listener by inserting backchannels, such as continuers and assessment, to influence speakers. Our model takes into account not only lexical and acoustic cues, but also introduces the simple and novel idea of using listener embeddings to mimic different backchanneling behaviours. Our experimental results on the Switchboard benchmark dataset reveal that acoustic cues are more important than lexical cues in this task and their combination with listener embeddings works best on both, manual transcriptions and automatically generated transcriptions. Daniel Ortega, Chia-Yu Li, Ngoc Thang Vu |
ICASSP | 3 |
| 2020 | Pretrained Semantic Speech Embeddings for End-to-End Spoken Language Understanding via Cross-Modal Teacher-Student LearningabstractSpoken language understanding is typically based on pipeline architectures including speech recognition and natural language understanding steps. These components are optimized independently to allow usage of available data, but the overall system suffers from error propagation. In this paper, we propose a novel training method that enables pretrained contextual embeddings to process acoustic features. In particular, we extend it with an encoder of pretrained speech recognition systems in order to construct end-to-end spoken language understanding systems. Our proposed method is based on the teacher-student framework across speech and text modalities that aligns the acoustic and the semantic latent spaces. Experimental results in three benchmarks show that our system reaches the performance comparable to the pipeline architecture without using any training data and outperforms it after fine-tuning with ten examples per class on two out of three benchmarks. Pavel Denisov, Ngoc Thang Vu |
INTERSPEECH | 2 |
| 2020 | Improving Code-Switching Language Modeling with Artificially Generated Texts Using Cycle-Consistent Adversarial NetworksabstractThis paper presents our latest effort on improving Codeswitching language models that suffer from data scarcity.We investigate methods to augment Code-switching training text data by artificially generating them.Concretely, we propose a cycle-consistent adversarial networks based framework to transfer monolingual text into Code-switching text, considering Code-switching as a speaking style.Our experimental results on the SEAME corpus show that utilizing artificially generated Code-switching text data improves consistently the language model as well as the automatic speech recognition performance. Chia-Yu Li, Ngoc Thang Vu |
INTERSPEECH | 2 |
| 2020 | Cairo Student Code-Switch (CSCS) Corpus: An Annotated Egyptian Arabic-English CorpusabstractCode-switching has become a prevalent phenomenon across many communities. It poses a challenge to NLP researchers, mainly due to the lack of available data needed for training and testing applications. In this paper, we introduce a new resource: a corpus of Egyptian- Arabic code-switch speech data that is fully tokenized, lemmatized and annotated for part-of-speech tags. Beside the corpus itself, we provide annotation guidelines to address the unique challenges of annotating code-switch data. Another challenge that we address is the fact that Egyptian Arabic orthography and grammar are not standardized. Mohamed Balabel, Injy Hamed, Slim Abdennadher, Ngoc Thang Vu, Özlem Çetinoglu |
LREC | 4 |
| 2020 | ArzEn: A Speech Corpus for Code-switched Egyptian Arabic-EnglishabstractIn this paper, we present our ArzEn corpus, an Egyptian Arabic-English code-switching (CS) spontaneous speech corpus. The corpus is collected through informal interviews with 38 Egyptian bilingual university students and employees held in a soundproof room. A total of 12 hours are recorded, transcribed, validated and sentence segmented. The corpus is mainly designed to be used in Automatic Speech Recognition (ASR) systems, however, it also provides a useful resource for analyzing the CS phenomenon from linguistic, sociological, and psychological perspectives. In this paper, we first discuss the CS phenomenon in Egypt and the factors that gave rise to the current language. We then provide a detailed description on how the corpus was collected, giving an overview on the participants involved. We also present statistics on the CS involved in the corpus, as well as a summary to the effort exerted in the corpus development, in terms of number of hours required for transcription, validation, segmentation and speaker annotation. Finally, we discuss some factors contributing to the complexity of the corpus, as well as Arabic-English CS behaviour that could pose potential challenges to ASR systems. Injy Hamed, Ngoc Thang Vu, Slim Abdennadher |
LREC | 2 |
| 2020 | Low-resource text classification using domain-adversarial learning
Daniel Grießhaber, Ngoc Thang Vu, Johannes Maucher |
Comput. Speech Lang. | 2 |
| 2020 | Acoustic and temporal representations in convolutional neural network models of prosodic events
Sabrina Stehwien, Antje Schweitzer, Ngoc Thang Vu |
Speech Commun. | 3 |
| 2019 | Improving Speech Emotion Recognition with Unsupervised Representation Learning on Unlabeled SpeechabstractIn this paper we present our findings on how representation learning on large unlabeled speech corpora can be beneficially utilized for speech emotion recognition (SER). Prior work on representation learning for SER mostly focused on the relatively small emotional speech datasets without making use of additional unlabeled speech data. We show that integrating representations learnt by an unsupervised autoencoder into a CNN-based emotion classifier improves the recognition accuracy. To gain insights about what those models learn, we analyze visualizations of the different representations using t-distributed neighbor embeddings (t-SNE). We evaluate our approach on IEMOCAP and MSP-IMPROV by means of within- and cross-corpus testing. Michael Neumann 0001, Ngoc Thang Vu |
ICASSP | 2 |
| 2019 | Context-aware Neural-based Dialog Act Classification on Automatically Generated TranscriptionsabstractThis paper presents our latest investigations on dialog act (DA) classification on automatically generated transcriptions. We propose a novel approach that combines convolutional neural networks (CNNs) and conditional random fields (CRFs) for context modeling in DA classification. We explore the impact of transcriptions generated from different automatic speech recognition systems such as hybrid TDNN/HMM and End-to-End systems on the final performance. Experimental results on two benchmark datasets (MRDA and SwDA) show that the combination CNN and CRF improves consistently the accuracy. Furthermore, they show that although the word error rates are comparable, End-to-End ASR system seems to be more suitable for DA classification. Daniel Ortega, Chia-Yu Li, Gisela Vallejo, Pavel Denisov, Ngoc Thang Vu |
ICASSP | 5 |
| 2019 | Head-First Linearization with Tree-Structured RepresentationabstractWe present a dependency tree linearization model with two novel components: (1) a tree-structured encoder based on bidirectional Tree-LSTM that propagates information first bottom-up then top-down, which allows each token to access information from the entire tree; and (2) a linguistically motivated headfirst decoder that emphasizes the central role of the head and linearizes the subtree by incrementally attaching the dependents on both sides of the head.With the new encoder and decoder, we reach state-of-the-art performance on the Surface Realization Shared Task 2018 dataset, outperforming not only the shared tasks participants, but also previous state-ofthe-art systems (Bohnet et al., 2011;Puduppully et al., 2016).Furthermore, we analyze the power of the tree-structured encoder with a probing task and show that it is able to recognize the topological relation between any pair of tokens in a tree. Agnieszka Falenska, Ngoc Thang Vu, Jonas Kuhn |
INLG | 3 |
| 2019 | Automatic Compression of Subtitles with Neural Networks and its Effect on User Experience
Katrin Angerbauer, Heike Adel, Ngoc Thang Vu |
INTERSPEECH | 3 |
| 2019 | CycleGAN-Based Emotion Style Transfer as Data Augmentation for Speech Emotion Recognition
Fang Bao, Michael Neumann 0001, Ngoc Thang Vu |
INTERSPEECH | 3 |
| 2019 | End-to-End Multi-Speaker Speech Recognition Using Speaker Embeddings and Transfer LearningabstractThis paper presents our latest investigation on end-to-end automatic speech recognition (ASR) for overlapped speech. We propose to train an end-to-end system conditioned on speaker embeddings and further improved by transfer learning from clean speech. This proposed framework does not require any parallel non-overlapped speech materials and is independent of the number of speakers. Our experimental results on overlapped speech datasets show that joint conditioning on speaker embeddings and transfer learning significantly improves the ASR performance. Pavel Denisov, Ngoc Thang Vu |
INTERSPEECH | 2 |
| 2019 | Multimodal Articulation-Based Pronunciation Error Detection with Spectrogram and Acoustic Features
Sabrina Jenne, Ngoc Thang Vu |
INTERSPEECH | 2 |
| 2019 | To Combine or Not To Combine? A Rainbow Deep Reinforcement Learning Agent for Dialog PoliciesabstractWe explore state-of-the-art deep reinforcement learning methods such as prioritized experience replay, double deep Q-Networks, dueling network architectures, distributional learning methods for dialog policy.Our main findings show that each individual method improves the rewards and the task success rate but combining these methods in a Rainbow agent, which performs best across tasks and environments, is a non-trivial task.We, therefore, provide insights about the influence of each method on the combination and how to combine them to form the Rainbow agent. Dirk Väth, Ngoc Thang Vu |
SIGdial | 2 |
| 2018 | Comparing Attention-Based Convolutional and Recurrent Neural Networks: Success and Limitations in Machine Reading ComprehensionabstractWe propose a machine reading comprehension model based on the compare-aggregate framework with two-staged attention that achieves state-of-the-art results on the MovieQA question answering dataset.To investigate the limitations of our model as well as the behavioral difference between convolutional and recurrent neural networks, we generate adversarial examples to confuse the model and compare to human performance.Furthermore, we assess the generalizability of our model by analyzing its differences to human inference, drawing upon insights from cognitive science. Matthias Blohm, Glorianna Jagfeld, Ekta Sood, Ngoc Thang Vu |
CoNLL | 5 |
| 2018 | CRoss-lingual and Multilingual Speech Emotion Recognition on English and FrenchabstractResearch on multilingual speech emotion recognition faces the problem that most available speech corpora differ from each other in important ways, such as annotation methods or interaction scenarios. These inconsistencies complicate building a multilingual system. We present results for cross-lingual and multilingual emotion recognition on English and French speech data with similar characteristics in terms of interaction (human-human conversations). Further, we explore the possibility of fine-tuning a pre-trained cross-lingual model with only a small number of samples from the target language, which is of great interest for low-resource languages. To gain more insights in what is learned by the deployed convolutional neural network, we perform an analysis on the attention mechanism inside the network. Michael Neumann 0001, Ngoc Thang Vu |
ICASSP | 2 |
| 2018 | Lexico-Acoustic Neural-Based Models for Dialog Act ClassificationabstractRecent works have proposed neural models for dialog act classification in spoken dialogs. However, they have not explored the role and the usefulness of acoustic information. We propose a neural model that processes both lexical and acoustic features for classification. Our results on two benchmark datasets reveal that acoustic features are helpful in improving the overall accuracy. Finally, a deeper analysis shows that acoustic features are valuable in three cases: when a dialog act has sufficient data, when lexical information is limited and when strong lexical cues are not present. Daniel Ortega, Ngoc Thang Vu |
ICASSP | 2 |
| 2018 | Investigations on End- to-End Audiovisual FusionabstractAudiovisual speech recognition (AVSR) is a method to alleviate the adverse effect of noise in the acoustic signal. Leveraging recent developments in deep neural network-based speech recognition, we present an AVSR neural network architecture which is trained end-to-end, without the need to separately model the process of decision fusion as in conventional (e.g. HMM-based) systems. The fusion system outperforms single-modality recognition under all noise conditions. Investigation of the saliency of the input features shows that the neural network automatically adapts to different noise levels in the acoustic signal. Michael Wand 0002, Jürgen Schmidhuber, Ngoc Thang Vu |
ICASSP | 3 |
| 2018 | Sequence-to-Sequence Models for Data-to-Text Natural Language Generation: Word- vs. Character-based Processing and Output DiversityabstractWe present a comparison of word-based and character-based sequence-to-sequence models for data-to-text natural language generation, which generate natural language descriptions for structured inputs.On the datasets of two recent generation challenges, our models achieve comparable or better automatic evaluation results than the best challenge submissions.Subsequent detailed statistical and human analyses shed light on the differences between the two input representations and the diversity of the generated texts.In a controlled experiment with synthetic training data generated from templates, we demonstrate the ability of neural models to learn novel combinations of the templates and thereby generalize beyond the linguistic structures they were trained on. Glorianna Jagfeld, Sabrina Jenne, Ngoc Thang Vu |
INLG | 3 |
| 2017 | Distinguishing Antonyms and Synonyms in a Pattern-based Neural NetworkabstractDistinguishing between antonyms and synonyms is a key task to achieve high performance in NLP systems.While they are notoriously difficult to distinguish by distributional co-occurrence models, pattern-based methods have proven effective to differentiate between the relations.In this paper, we present a novel neural network model AntSynNET that exploits lexico-syntactic patterns from syntactic parse trees.In addition to the lexical and syntactic information, we successfully integrate the distance between the related words along the syntactic path as a new pattern feature.The results from classification experiments show that AntSyn-NET improves the performance over prior pattern-based methods. Kim Anh Nguyen 0001, Sabine Schulte im Walde, Ngoc Thang Vu |
EACL (1) | 3 |
| 2017 | Hierarchical Embeddings for Hypernymy Detection and DirectionalityabstractWe present a novel neural model HyperVec to learn hierarchical embeddings for hypernymy detection and directionality.While previous embeddings have shown limitations on prototypical hypernyms, HyperVec represents an unsupervised measure where embeddings are learned in a specific order and capture the hypernym-hyponym distributional hierarchy.Moreover, our model is able to generalize over unseen hypernymy pairs, when using only small sets of training data, and by mapping to other languages.Results on benchmark datasets show that HyperVec outperforms both state-of-theart unsupervised measures and embedding models on hypernymy detection and directionality, and on predicting graded lexical entailment. Kim Anh Nguyen 0001, Maximilian Köper, Sabine Schulte im Walde, Ngoc Thang Vu |
EMNLP | 4 |
| 2017 | Attentive Convolutional Neural Network Based Speech Emotion Recognition: A Study on the Impact of Input Features, Signal Length, and Acted SpeechabstractSpeech emotion recognition is an important and challenging task in the realm of human-computer interaction. Prior work proposed a variety of models and feature sets for training a system. In this work, we conduct extensive experiments using an attentive convolutional neural network with multi-view learning objective function. We compare system performance using different lengths of the input signal, different types of acoustic features and different types of emotion speech (improvised/scripted). Our experimental results on the Interactive Emotional Motion Capture (IEMOCAP) database reveal that the recognition performance strongly depends on the type of speech data independent of the choice of input features. Furthermore, we achieved state-of-the-art results on the improvised speech data of IEMOCAP. Michael Neumann 0001, Ngoc Thang Vu |
INTERSPEECH | 2 |
| 2017 | Prosodic Event Recognition Using Convolutional Neural Networks with Context InformationabstractThis paper demonstrates the potential of convolutional neural networks (CNN) for detecting and classifying prosodic events on words, specifically pitch accents and phrase boundary tones, from frame-based acoustic features. Typical approaches use not only feature representations of the word in question but also its surrounding context. We show that adding position features indicating the current word benefits the CNN. In addition, this paper discusses the generalization from a speaker-dependent modelling approach to a speaker-independent setup. The proposed method is simple and efficient and yields strong results not only in speaker-dependent but also speaker-independent cases. Sabrina Stehwien, Ngoc Thang Vu |
INTERSPEECH | 2 |
| 2017 | Neural-based Context Representation Learning for Dialog Act ClassificationabstractWe explore context representation learning methods in neural-based models for dialog act classification.We propose and compare extensively different methods which combine recurrent neural network architectures and attention mechanisms (AMs) at different context levels.Our experimental results on two benchmark datasets show consistent improvements compared to the models without contextual information and reveal that the most suitable AM in the architecture depends on the nature of the dataset. Daniel Ortega, Ngoc Thang Vu |
SIGDIAL Conference | 2 |
| 2016 | Neural-based Noise Filtering from Word EmbeddingsabstractWord embeddings have been demonstrated to benefit NLP tasks impressively. Yet, there is room for improvements in the vector representations, because current word embeddings typically contain unnecessary information, i.e., noise. We propose two novel models to improve word embeddings by unsupervised learning, in order to yield word denoising embeddings. The word denoising embeddings are obtained by strengthening salient information and weakening noise in the original word embeddings, based on a deep feed-forward neural network filter. Results from benchmark tasks show that the filtered word denoising embeddings outperform the original word embeddings. Kim Anh Nguyen 0001, Sabine Schulte im Walde, Ngoc Thang Vu |
COLING | 3 |
| 2016 | Bi-directional recurrent neural network with ranking loss for spoken language understandingabstractThis paper presents our latest investigation of recurrent neural networks for the slot filling task of spoken language understanding. We implement a bi-directional Elman-type recurrent neural network which takes the information not only from the past but also from the future context to predict the semantic label of the target word. Furthermore, we propose to use ranking loss function to train the model. This improves the performance over the cross entropy loss function. On the ATIS benchmark data set, we achieve a new state-of-the-art result of 95.56% F1-score without using any additional knowledge or data sources. Ngoc Thang Vu, Pankaj Gupta 0003, Heike Adel, Hinrich Schütze |
ICASSP | 1 |
| 2016 | Cross-Gender and Cross-Dialect Tone Recognition for Vietnamese
Antje Schweitzer, Ngoc Thang Vu |
INTERSPEECH | 2 |
| 2016 | Exploring the Correlation of Pitch Accents and Semantic Slots for Spoken Language Understanding
Sabrina Stehwien, Ngoc Thang Vu |
INTERSPEECH | 2 |
| 2016 | Sequential Convolutional Neural Networks for Slot Filling in Spoken Language UnderstandingabstractWe investigate the usage of convolutional neural networks (CNNs) for the slot filling task in spoken language understanding. We propose a novel CNN architecture for sequence labeling which takes into account the previous context words with preserved order information and pays special attention to the current word with its surrounding context. Moreover, it combines the information from the past and the future words for classification. Our proposed CNN architecture outperforms even the previously best ensembling recurrent neural network model and achieves state-of-the-art results with an F1-score of 95.61% on the ATIS benchmark dataset without using any additional linguistic knowledge and resources. Ngoc Thang Vu |
INTERSPEECH | 1 |
| 2016 | Combining Recurrent and Convolutional Neural Networks for Relation ClassificationabstractNgoc Thang Vu, Heike Adel, Pankaj Gupta, Hinrich Schütze. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Ngoc Thang Vu, Heike Adel, Pankaj Gupta 0003, Hinrich Schütze |
HLT-NAACL | 1 |
| 2015 | Syntactic and Semantic Features For Code-Switching Factored Language ModelsabstractThis paper presents our latest investigations on different features for factored language models for Code-Switching speech and their effect on automatic speech recognition (ASR) performance. We focus on syntactic and semantic features which can be extracted from Code-Switching text data and integrate them into factored language models. Different possible factors, such as words, part-of-speech tags, Brown word clusters, open class words and clusters of open class word embeddings are explored. The experimental results reveal that Brown word clusters, part-of-speech tags and open-class words are the most effective at reducing the perplexity of factored language models on the Mandarin-English Code-Switching corpus SEAME. In ASR experiments, the model containing Brown word clusters and part-of-speech tags and the model also including clusters of open class word embeddings yield the best mixed error rate results. In summary, the best language model can significantly reduce the perplexity on the SEAME evaluation set by up to 10.8% relative and the mixed error rate by up to 3.4% relative. Heike Adel, Ngoc Thang Vu, Katrin Kirchhoff, Dominic Telaar, Tanja Schultz |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Multilingual deep neural network based acoustic modeling for rapid language adaptationabstractThis paper presents a study on multilingual deep neural network (DNN) based acoustic modeling and its application to new languages. We investigate the effect of phone merging on multilingual DNN in context of rapid language adaptation. Moreover, the combination of multilingual DNNs with Kullback-Leibler divergence based acoustic modeling (KL-HMM) is explored. Using ten different languages from the Globalphone database, our studies reveal that crosslingual acoustic model transfer through multilingual DNNs is superior to unsupervised RBM pre-training and greedy layer-wise supervised training. We also found that KL-HMM based decoding consistently outperforms conventional hybrid decoding, especially in low-resource scenarios. Furthermore, the experiments indicate that multilingual DNN training equally benefits from simple phoneset concatenation and manually derived universal phonesets. Ngoc Thang Vu, David Imseng, Daniel Povey, Petr Motlícek, Tanja Schultz, Hervé Bourlard |
ICASSP | 1 |
| 2014 | Comparing approaches to convert recurrent neural networks into backoff language models for efficient decodingabstractIn this paper, we investigate and compare three different possibilities to convert recurrent neural network language models (RNNLMs) into backoff language models (BNLM). While RNNLMs often outperform traditional n-gram approaches in the task of language modeling, their computational demands make them unsuitable for an efficient usage during decoding in an LVCSR system. It is, therefore, of interest to convert them into BNLMs in order to integrate their information into the decoding process. This paper compares three different approaches: a text based conversion, a probability based conversion and an iterative conversion. The resulting language models are evaluated in terms of perplexity and mixed error rate in the context of the Code-Switching data corpus SEAME. Although the best results are obtained by combining the results of all three approaches, the text based conversion approach alone leads to significant improvements on the SEAME corpus as well while offering the highest computational efficiency. In total, the perplexity can be reduced by 11.4% relative on the evaluation set and the mixed error rate by 3.0% relative on the same data set. Heike Adel, Katrin Kirchhoff, Ngoc Thang Vu, Dominic Telaar, Tanja Schultz |
INTERSPEECH | 3 |
| 2014 | Combining recurrent neural networks and factored language models during decoding of code-Switching speechabstractIn this paper, we present our latest investigations of language modeling for Code-Switching. Since there is only little text material for Code-Switching speech available, we integrate syntactic and semantic features into the language modeling process. In particular, we use part-of-speech tags, language identifiers, Brown word clusters and clusters of open class words. We develop factored language models and convert recurrent neural network language models into backoff language models for an efficient usage during decoding. A detailed error analysis reveals the strengths and weaknesses of the different language models. When we interpolate the models linearly, we reduce the perplexity by 15.6% relative on the SEAME evaluation set. This is even slightly better than the result of the unconverted recurrent neural network. We also combine the language models during decoding and obtain a mixed error rate reduction of 4.4% relative on the SEAME evaluation set. Heike Adel, Dominic Telaar, Ngoc Thang Vu, Katrin Kirchhoff, Tanja Schultz |
INTERSPEECH | 3 |
| 2014 | BioKIT - real-time decoder for biosignal processingabstractWe introduce BioKIT, a new Hidden Markov Model based toolkit to preprocess, model and interpret biosignals such as speech, motion, muscle and brain activities. The focus of this toolkit is to enable researchers from various communities to pursue their experiments and integrate real-time biosignal interpretation into their applications. BioKIT boosts a flexible two-layer structure with a modular C++ core that interfaces with a Python scripting layer, to facilitate development of new applications. BioKIT employs sequence-level parallelization and memory sharing across threads. Additionally, a fully integrated error blaming component facilitates in-depth analysis. A generic terminology keeps the barrier to entry for researchers from multiple fields to a minimum. We describe our onlinecapable dynamic decoder and report on initial experiments on three different tasks. The presented speech recognition experiments employ Kaldi [1] trained deep neural networks with the results set in relation to the real time factor needed to obtain them. Dominic Telaar, Michael Wand 0002, Dirk Gehrig, Felix Putze, Christoph Amma, Dominic Heger, Ngoc Thang Vu, Mark Erhardt, Tim Schlippe, Matthias Janke, Christian Herff, Tanja Schultz |
INTERSPEECH | 7 |
| 2014 | Improving ASR performance on non-native speech using multilingual and crosslingual informationabstractThis paper presents our latest investigation of automatic speech recognition (ASR) on non-native speech. We first report on a non-native speech corpus an extension of the GlobalPhone database which contains English with Bulgarian, Chinese, German and Indian accent and German with Chinese accent. In this case, English is the spoken language (L2) and Bulgarian, Chinese, German and Indian are the mother tongues (L1) of the speakers. Afterwards, we investigate the effect of multilingual acoustic modeling on non-native speech. Our results reveal that a bilingual L1-L2 acoustic model significantly improves the ASR performance on non-native speech. For the case that L1 is unknown or L1 data is not available, a multilingual ASR system trained without L1 speech data consistently outperforms the monolingual L2 ASR system. Finally, we propose a method called crosslingual accent adaptation, which allows using English with Chinese accent to improve the German ASR on German with Chinese accent and vice versa. Without using any intra lingual adaptation data, we achieve 15.8% relative improvement in average over the baseline system. Ngoc Thang Vu, Yuanfan Wang, Marten Klose, Zlatka Mihaylova, Tanja Schultz |
INTERSPEECH | 1 |
| 2014 | Investigating the learning effect of multilingual bottle-neck features for ASRabstractDeep neural networks (DNNs) have become state-of-the-art techniques of automatic speech recognition in the last few years. They can be used at the preprocessing level (Tandem or Bottle-Neck features) or at the acoustic model level (hybrid Hidden Markov Model/DNN). Moreover, they allow exploiting multilingual data to improve monolingual systems. This paper presents our investigation of the learning effect of neural networks in the context of multilingual Bottle-Neck features. For this, we perform a visual analysis of the output of the Bottle-Neck layer of a neural network using t-Distributed Stochastic Neighbor Embedding. Our results show that multilingual Bottle-Neck features seem to learn phoneme characteristics, such as the F1 and F2 formants which characterize different vowels, and other articulatory features, such as fricatives and nasals which characterize consonants. Furthermore, they seem to normalize language dependent variations and transfer the learned representation to unseen languages. Ngoc Thang Vu, Jochen Weiner, Tanja Schultz |
INTERSPEECH | 1 |
| 2013 | Recurrent neural network language modeling for code switching conversational speechabstractCode-switching is a very common phenomenon in multilingual communities. In this paper, we investigate language modeling for conversational Mandarin-English code-switching (CS) speech recognition. First, we investigate the prediction of code switches based on textual features with focus on Part-of-Speech (POS) tags and trigger words. Second, we propose a structure of recurrent neural networks to predict code-switches. We extend the networks by adding POS information to the input layer and by factorizing the output layer into languages. The resulting models are applied to our task of code-switching language modeling. The final performance shows 10.8% relative improvement in perplexity on the SEAME development set which transforms into a 2% relative improvement in terms of Mixed Error Rate and a relative improvement of 16.9% in perplexity on the evaluation set which leads to a 2.7% relative improvement of MER. Heike Adel, Ngoc Thang Vu, Franziska Kraus, Tim Schlippe, Haizhou Li 0001, Tanja Schultz |
ICASSP | 2 |
| 2013 | GlobalPhone: A multilingual text & speech database in 20 languagesabstractThis paper describes the advances in the multilingual text and speech database GlobalPhone, a multilingual database of high-quality read speech with corresponding transcriptions and pronunciation dictionaries in 20 languages. GlobalPhone was designed to be uniform across languages with respect to the amount of data, speech quality, the collection scenario, the transcription and phone set conventions. With more than 400 hours of transcribed audio data from more than 2000 native speakers GlobalPhone supplies an excellent basis for research in the areas of multilingual speech recognition, rapid deployment of speech processing systems to yet unsupported languages, language identification tasks, speaker recognition in multiple languages, multilingual speech synthesis, as well as monolingual speech recognition in a large variety of languages. Tanja Schultz, Ngoc Thang Vu, Tim Schlippe |
ICASSP | 2 |
| 2013 | Experiments towards a better LVCSR system for tamilabstractThis paper summarizes our latest efforts in the development of a Large Vocabulary Continuous Speech Recognition (LVCSR) system for Tamil at different levels: pronunciation dictionary, language modeling (LM) and front-end. Usually in Tamil there are not many word-pronunciation pairs to train data-driven grapheme-to-phoneme (G2P) converters. Therefore, we explore the correlation between the amount of training data and the performance of the grapheme-to-phoneme (G2P) conversion. To address the morphological complexity of Tamil, we investigate different levels of morphemes for language modeling including a comparison between our Dictionary Unit Merging Algorithm (DUMA) and Morfessor, followed by various experiments on hybrid systems using word and morpheme LMs. Finally, we integrate our multilingual bottle-neck features framework with Tamil LVCSR. The final best system produced 21.34% Syllable Error Rate (SyllER) on our Tamil test set. Melvin Jose Johnson Premkumar, Ngoc Thang Vu, Tanja Schultz |
INTERSPEECH | 2 |
| 2013 | Unsupervised language model adaptation for automatic speech recognition of broadcast news using web 2.0abstractWe improve the automatic speech recognition of broadcast news using paradigms from Web 2.0 to obtain timeand topicrelevant text data for language modeling. We elaborate an unsupervised text collection and decoding strategy that includes crawling appropriate texts from RSS Feeds, complementing it with texts from Twitter, language model and vocabulary adaptation, as well as a 2-pass decoding. The word error rates of the tested French broadcast news shows from Europe 1 are reduced by almost 32% relative with an underlying language model from the GlobalPhone project [1] and by almost 4% with an underlying language model from the Quaero project. The tools that we use for the text normalization, the collection of RSS Feeds together with the text on the related websites, a TF-IDF-based topic words extraction, as well as the opportunity for language model interpolation are available in our Rapid Language Adaptation Toolkit [2] [3]. Tim Schlippe, Lukasz Gren, Ngoc Thang Vu, Tanja Schultz |
INTERSPEECH | 3 |
| 2013 | Multilingual multilayer perceptron for rapid language adaptation between and across language familiesabstractIn this paper, we present our latest investigations of multilingual Multilayer Perceptrons (MLPs) for rapid language adaptation between and across language families. We explore the impact of the amount of languages and data used for the multilingual MLP training process. We show that the overall system performance on the target language is significantly improved by initializing it with a multilingual MLP. Our experiments indicate that the more languages we use to train a multilingual MLP, the better is the initialization for MLP training. As a result, the ASR performance is improved, even if the target language and the source languages are not in the same language family. Our best results show an error rate improvement of up to 22.9% relative for different target languages (Czech, Hausa and Vietnamese) by using a multilingual MLP which has been trained with many different languages from the GlobalPhone corpus. In the case of very few training or adaptation data, an improvement of up to 24% relative in terms of error rate is observed. Ngoc Thang Vu, Tanja Schultz |
INTERSPEECH | 1 |
| 2012 | Generating exact lattices in the WFST frameworkabstractWe describe a lattice generation method that is exact, i.e. it satisfies all the natural properties we would want from a lattice of alternative transcriptions of an utterance. This method does not introduce substantial overhead above one-best decoding. Our method is most directly applicable when using WFST decoders where the WFST is “fully expanded”, i.e. where the arcs correspond to HMM transitions. It outputs lattices that include HMM-state-level alignments as well as word labels. The general idea is to create a state-level lattice during decoding, and to do a special form of determinization that retains only the best-scoring path for each word sequence. This special determinization algorithm is a solution to the following problem: Given a WFST A, compute a WFST B that, for each input-symbol-sequence of A, contains just the lowest-cost path through A. Daniel Povey, Mirko Hannemann, Gilles Boulianne, Lukás Burget, Arnab Ghoshal, Milos Janda, Martin Karafiát, Stefan Kombrink, Petr Motlícek, Yanmin Qian, Korbinian Riedhammer, Karel Veselý, Ngoc Thang Vu |
ICASSP | 13 |
| 2012 | A first speech recognition system for Mandarin-English code-switch conversational speechabstractThis paper presents first steps toward a large vocabulary continuous speech recognition system (LVCSR) for conversational Mandarin-English code-switching (CS) speech. We applied state-of-the-art techniques such as speaker adaptive and discriminative training to build the first baseline system on the SEAME corpus [1] (South East Asia Mandarin-English). For acoustic modeling, we applied different phone merging approaches based on the International Phonetic Alphabet (IPA) and Bhattacharyya distance in combination with discriminative training to improve accuracy. On language model level, we investigated statistical machine translation (SMT) - based text generation approaches for building code-switching language models. Furthermore, we integrated the provided information from a language identification system (LID) into the decoding process by using a multi-stream approach. Our best 2-pass system achieves a Mixed Error Rate (MER) of 36.6% on the SEAME development set. Ngoc Thang Vu, Dau-Cheng Lyu, Jochen Weiner, Dominic Telaar, Tim Schlippe, Fabian Blaicher, Chng Eng Siong, Tanja Schultz, Haizhou Li 0001 |
ICASSP | 1 |
| 2012 | Modeling gender dependency in the Subspace GMM frameworkabstractThe Subspace GMM acoustic model has both globally shared parameters and parameters specific to acoustic states, and this makes it possible to do various kinds of tying. In the past we have investigated sharing the global parameters among systems with distinct acoustic states; this can be useful in a multilingual setting. In the current paper we investigate the reverse idea: to have different global parameters for different acoustic conditions (gender, in this case) while sharing the acoustic-state-specific parameters. We experiment with modeling gender dependency in this way, and show Word Error Rate improvements on a range of tasks and comparable results to the Vocal Tract Length Normalization (VTLN)-like technique Exponential Transform (ET). Ngoc Thang Vu, Tanja Schultz, Daniel Povey |
ICASSP | 1 |
| 2012 | Automatic Error Recovery for Pronunciation DictionariesabstractIn this paper, we present our latest investigations on pronunciation modeling and its impact on ASR. We propose completely automatic methods to detect, remove, and substitute inconsistent or flawed entries in pronunciation dictionaries. The experiments were conducted on different tasks, namely (1) word-pronunciation pairs from the Czech, English, French, German, Polish, and Spanish Wiktionary [1], a multilingual wiki-based open content dictionary, (2) our GlobalPhone Hausa pronunciation dictionary [2], and (3) pronunciations to complement our Mandarin-English SEAME code-switch dictionary [3]. In the final results, we fairly observed on average an improvement of 2.0% relative in terms of word error rate and even 27.3% for the case of English Wiktionary word-pronunciation pairs. Tim Schlippe, Sebastian Ochs 0002, Ngoc Thang Vu, Tanja Schultz |
INTERSPEECH | 3 |
| 2012 | Initialization Schemes for Multilayer Perceptron Training and their Impact on ASR Performance using Multilingual Data
Ngoc Thang Vu, Wojtek Breiter, Florian Metze, Tanja Schultz |
INTERSPEECH | 1 |
| 2011 | Cross-language bootstrapping based on completely unsupervised training using multilingual A-stabilabstractThis paper presents our work on rapid language adaptation of acoustic models based on multilingual cross-language bootstrapping and unsupervised training. We used Automatic Speech Recognition (ASR) systems in English, French, German, and Spanish to build a Czech ASR system from scratch. System building was performed without using any transcribed audio data by applying three consecutive steps, i.e. cross-language transfer, unsupervised training based on the "multilingual A-stabil" confidence score, and boot strapping. Based on the confidence score we selected 72% (16.6 hours) of the available audio data with a transcription WER of less than 14.5%. The cross-language bootstrap achieves a word error rate of 23.3% on the Czech development set and 22.4% on the evaluation set. These results are very promising as the performance compares favorably to the Czech ASR system which was trained on 23 hours of manually transcribed data (21.8% on the development set and 21.3% on the evaluation set). Ngoc Thang Vu, Franziska Kraus, Tanja Schultz |
ICASSP | 1 |
| 2011 | Rapid Building of an ASR System for Under-Resourced Languages Based on Multilingual Unsupervised TrainingabstractThis paper presents our work on rapid language adaptation of acoustic models based on multilingual cross-language bootstrapping and unsupervised training. We used Automatic Speech Recognition (ASR) systems in the six source languages English, French, German, Spanish, Bulgarian and Polish to build from scratch an ASR system for Vietnamese, an underresourced language. System building was performed without using any transcribed audio data by applying three consecutive steps, i.e. cross-language transfer, unsupervised training based on the “multilingual A-stabil ” confidence score [1], and bootstrapping. We investigated the correlation between performance of “multilingual A-stabil ” and the number of source languages and improved the performance of “multilingual A-stabil ” by applying it at the syllable level. Furthermore, we showed that increasing the amount of source language ASR systems for the multilingual framework results in better performance of the final ASR system in the target language Vietnamese. The final Vietnamese recognition system has a Syllable Error Rate (SyllER) of 16.8 % on the development set and 16.1 % on the evaluation set. Index Terms: rapid language adaptation of ASR, unsupervised training, multilingual A-Stabil Ngoc Thang Vu, Franziska Kraus, Tanja Schultz |
INTERSPEECH | 1 |
| 2010 | Rapid bootstrapping of five eastern european languages using the rapid language adaptation toolkitabstractThis paper presents our latest efforts toward LVCSR systems for five Eastern European languages such as Bulgarian, Croatian, Czech, Polish, and Russian using our Rapid Language Adaptation Toolkit (RLAT) [1]. We investigated the possibility of crawling large quantities of text material from the Internet, which is very cheap but also requires text post-processing steps due to the varying text quality. The goal of this study is to determine the best strategy for language model optimization on the given domain in a short time period with minimal human effort. Our results show that we can build an initial ASR system for these five languages in only twenty days using RLAT. On the multilingual GlobalPhone speech corpus [2], we achieved a word error rate (WER) of 16.9 % for Bulgarian, 32.8 % for Ngoc Thang Vu, Tim Schlippe, Franziska Kraus, Tanja Schultz |
INTERSPEECH | 1 |
| 2010 | Multilingual a-stabil: A new confidence score for multilingual unsupervised trainingabstractThis paper presents our work in Automatic Speech Recognition (ASR) in the context of multilingual unsupervised training with application to Czech. Starting without any transcribed acoustic training data we built a Czech ASR by combining cross-language bootstrapping and confidence based unsupervised training. We present our new method called “multilingual A-stabil” to compute confidence scores and explore the relative effectiveness of acoustic models from more than one language such as Russian, Bulgarian, Polish and Croatian for unsupervised training. While conventional confidence measures such as gamma and A-stabil work well with well-trained acoustic models but have problems with poorly estimated acoustic models, our new method works well in both cases. We describe our multilingual unsupervised training framework which gives very promising results in our experiments. We were able to select 80.5% of the audio training data (18.5 hours) with a transcription WER of 14.5% when using a small amount of untranscribed data (only about 23 hours). The final best WER on Czech is 23.6% on the development set and 22.9% on the evaluation set by using cross-lingual boostrapping, which is very close to the performance of the Czech ASR trained with 23 hours audio data with manual transcriptions (23.1% on the development set and 22.3% on the evaluation set). Ngoc Thang Vu, Franziska Kraus, Tanja Schultz |
SLT | 1 |
| 2009 | Vietnamese large vocabulary continuous speech recognitionabstractWe report on our recent efforts toward a large vocabulary Vietnamese speech recognition system. In particular, we describe the Vietnamese text and speech database recently collected as part of our GlobalPhone corpus. The data was complemented by a large collection of text data crawled from various Vietnamese websites. To bootstrap the Vietnamese speech recognition system we used our Rapid Language Adaptation scheme applying a multilingual phone inventory. After initialization we investigated the peculiarities of the Vietnamese language and achieved significant improvements by implementing different tone modeling schemes, extended by pitch extraction, handling multiwords to address the monosyllable structure of Vietnamese, and featuring language modeling based on 5-grams. Furthermore, we addressed the issue of dialectal variations between South and North Vietnam by creating dialect dependent pronunciations and including dialect in the context decision tree of the recognizer. Our currently best recognition system achieves a word error rate of 11.7% on read newspaper speech. Ngoc Thang Vu, Tanja Schultz |
ASRU | 1 |