VLDB 2026 Research / reviewers in the wild / expert
Dietrich Klakow
dblp:00/1846
· DBLP profile ↗
173ranked-venue papers
7as first author
68since 2021 · last 2026
0000-0002-4147-9690ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 134 · 3 first-author · 60 since 2021Graphics, computer vision, multimedia, augmented reality and games · 70 · 7 first-author · 21 since 2021Databases, data management, data science and information retrieval · 14 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Uhura: A Benchmark for Evaluating Scientific Question Answering and Truthfulness in Low-Resource African LanguagesabstractEvaluations of Large Language Models (LLMs) on knowledge-intensive tasks and factual accuracy often focus on high-resource languages primarily because datasets for low-resource languages (LRLs) are scarce. In this paper, we present Uhura -- a new benchmark that focuses on two tasks in six typologically-diverse African languages, created via human translation of existing English benchmarks. The first dataset, Uhura-ARC-Easy, is composed of multiple-choice science questions. The second, Uhura-TruthfulQA, is a safety benchmark testing the truthfulness of models on topics including health, law, finance, and politics. We highlight the challenges creating benchmarks with highly technical content for LRLs and outline mitigation strategies. Our evaluation reveals a significant performance gap between proprietary models such as GPT-4o and o1-preview, and Claude models, and open-source models like Meta's LLaMA and Google's Gemma. Additionally, all models perform better in English than in African languages. These results indicate that LMs struggle with answering scientific questions and are more prone to generating false claims in low-resource African languages. Our findings underscore the necessity for continuous improvement of multilingual LM capabilities in LRL settings to ensure safe and reliable use in real-world contexts. We open-source the Uhura Benchmark and Uhura Platform to foster further research and development in NLP for LRLs. Edward Bayes, Israel Abebe Azime, Jesujoba O. Alabi, Jonas Kgomo, Tyna Eloundou, Elizabeth Proehl, Imaan Khadir, Naome A. Etori, Shamsuddeen Hassan Muhammad, Choice Mpanza, Igneciah Pocia Thete, Dietrich Klakow, David Ifeoluwa Adelani |
LREC | 13 |
| 2026 | Audio-Lyrics Alignment Dataset for Italian AriasabstractAligning song lyrics with sung audio is challenging, especially for languages and music styles where annotated datasets are scarce. We address this gap by presenting the first dataset of Italian opera arias annotated with lyrics and time-stamps per word. The dataset comprises of 24 arias drawn from well-known operas of the 18th to 20th centuries with a total audio duration of nearly two hours. We benchmark both music alignment models and speech forced alignment models and show that existing methods face significant challenges on this dataset, with performance dropping by 45% compared to other datasets. Multilingual and speech-based models exhibit relatively better performance on this dataset. We also evaluate few-shot fine-tuning of these models on the new dataset and find that, while it yields only marginal overall improvement, it produces localized gains on specific arias, suggesting that limited exposure helps the model adapt to some patterns but cannot fully overcome differences in language or musical style. Pushkar Jajoria, Arianna Graciotti, Giovanna Casali, Jesujoba O. Alabi, Rodolfo Delmonte, Angelo Pompilio, Rocco Tripodi, James McDermott, Dietrich Klakow |
LREC | 9 |
| 2026 | Modeling the Memory-Surprisal Trade-Off over Time: Communicative Efficiency Decreases with Lexico-Grammatical Change in Scientific English
Julius Steuer, Marie-Pauline Krielke, Stefania Degaetano-Ortlieb, Elke Teich, Dietrich Klakow |
LREC | 5 |
| 2026 | Aligned Probing: Relating Toxic Behavior and Model InternalsabstractAbstract Warning: This paper contains offensive text. We introduce aligned probing, a novel interpretability framework that aligns the behavior of language models (LMs), based on their outputs, and their internal representations (internals). Using this framework, we examine over 20 OLMo, Llama, and Mistral models, bridging behavioral and internal perspectives for toxicity for the first time. Our results show that LMs strongly encode information about the toxicity level of inputs and subsequent outputs, particularly in lower layers. Focusing on how unique LMs differ offers both correlative and causal evidence that they generate less toxic output when strongly encoding information about the input toxicity. We also highlight the heterogeneity of toxicity, as model behavior and internals vary across unique attributes such as Threat. Finally, four case studies analyzing detoxification, multi-prompt evaluations, model quantization, and pre-training dynamics underline the practical impact of aligned probing with further concrete insights. Our findings contribute to a more holistic understanding of LMs, both within and beyond the context of toxicity. alignedprobing.github.io Andreas Waldis, Vagrant Gautam, Anne Lauscher, Dietrich Klakow, Iryna Gurevych |
Trans. Assoc. Comput. Linguistics | 4 |
| 2026 | Learning Program Behavioral Models from Synthesized Input-Output PairsabstractWe introduce Modelizer —a novel framework that, given a black-box program, learns a model from its input/output behavior using neural machine translation algorithms. The resulting model mocks the original program: Given an input, the model predicts the output that would have been produced by the program. However, the model is also reversible —that is, the model can predict the input that would have produced a given output. Finally, the model is differentiable and can be efficiently restricted to predict only a certain aspect of the program behavior. Modelizer uses grammars to synthesize and inputs and unsupervised tokenizers to decompose the resulting outputs, allowing it to learn sequence-to-sequence associations between token streams. Other than input grammars, Modelizer only requires the ability to execute the program. The resulting models are small, requiring fewer than 6.3 million parameters for languages such as Markdown or HTML; and they are accurate, achieving up to 95.4% accuracy and a BLEU score of 0.98 with standard error of 0.04 in mocking real-world applications. As it learns from and predicts executions rather than code, Modelizer departs from the LLM-centric research trend, opening new opportunities for program-specific models that are fully tuned toward individual programs. Indeed, we foresee several applications of these models, especially as the output of the program can be any aspect of program behavior. Beyond mocking and predicting program behavior, the models can also synthesize inputs that are likely to produce a particular behavior, such as failures or coverage, thus assisting in program understanding and maintenance. Tural Mammadov, Dietrich Klakow, Alexander Koller, Andreas Zeller |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2025 | INJONGO: A Multicultural Intent Detection and Slot-filling Dataset for 16 African LanguagesabstractSlot-filling and intent detection are well-established tasks in Conversational AI. However, current large-scale benchmarks for these tasks often exclude evaluations of low-resource languages and rely on translations from English benchmarks, thereby predominantly reflecting Western-centric concepts. In this paper, we introduce “INJONGO” - a multicultural, open-source benchmark dataset for 16 African languages with utterances generated by native speakers across diverse domains, including banking, travel, home, and dining. Through extensive experiments, we benchmark fine-tuning multilingual transformer models and prompting large language models (LLMs), and show the advantage of leveraging African-cultural utterances over Western-centric utterances for improving cross-lingual transfer from the English language. Experimental results reveal that current LLMs struggle with the slot-filling task, with GPT-4o achieving an average performance of 26 F1. In contrast, intent detection performance is notably better, with an average accuracy of 70.6%, though it still falls short of fine-tuning baselines. When compared to the English language, GPT-4o and fine-tuning baselines perform similarly on intent detection, achieving an accuracy of approximately 81%. Our findings suggest that LLMs performance is still behind for many low-resource African languages, and more work is needed to further improve their downstream performance. Jesujoba O. Alabi, Andiswa Bukula, Jian Yun Zhuang, Annie En-Shiun Lee, Tadesse Kebede Guge, Israel Abebe Azime, Happy Buzaaba, Blessing K. Sibanda, Godson Kalipe, Jonathan Mukiibi, Salomon Kabongo, Mmasibidi Setaka, Lolwethu Ndolela, Nkiruka Odu, Rooweither Mabuya, Shamsuddeen Hassan Muhammad, Salomey Osei, Sokhar Samb, Dietrich Klakow, David Ifeoluwa Adelani |
ACL (1) | 20 |
| 2025 | It's Not a Walk in the Park! Challenges of Idiom Translation in Speech-to-text SystemsabstractIuliia Zaitova, Badr M. Abdullah, Wei Xue, Dietrich Klakow, Bernd Möbius, Tania Avgustinova. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Iuliia Zaitova, Badr Abdullah, Dietrich Klakow, Bernd Möbius, Tania Avgustinova |
ACL (1) | 4 |
| 2025 | Evaluating the Capabilities of Large Language Models for Multi-label Emotion UnderstandingabstractLarge Language Models (LLMs) show promising learning and reasoning abilities. Compared to other NLP tasks, multilingual and multi-label emotion evaluation tasks are under-explored in LLMs. In this paper, we present EthioEmo, a multi-label emotion classification dataset for four Ethiopian languages, namely, Amharic (amh), Afan Oromo (orm), Somali (som), and Tigrinya (tir). We perform extensive experiments with an additional English multi-label emotion dataset from SemEval 2018 Task 1. Our evaluation includes encoder-only, encoder-decoder, and decoder-only language models. We compare zero and few-shot approaches of LLMs to fine-tuning smaller language models. The results show that accurate multi-label emotion classification is still insufficient even for high-resource languages such as English, and there is a large gap between the performance of high-resource and low-resource languages. The results also show varying performance levels depending on the language and model type. EthioEmo is available publicly to further improve the understanding of emotions in language models and how people convey emotions through various languages. Tadesse Destaw Belay, Israel Abebe Azime, Abinew Ali Ayele, Grigori Sidorov, Dietrich Klakow, Philipp Slusallek, Olga Kolesnikova, Seid Muhie Yimam |
COLING | 5 |
| 2025 | AFRIDOC-MT: Document-level MT Corpus for African LanguagesabstractJesujoba Oluwadara Alabi, Israel Abebe Azime, Miaoran Zhang, Cristina España-Bonet, Rachel Bawden, Dawei Zhu, David Ifeoluwa Adelani, Clement Oyeleke Odoje, Idris Akinade, Iffat Maab, Davis David, Shamsuddeen Hassan Muhammad, Neo Putini, David O. Ademuyiwa, Andrew Caines, Dietrich Klakow. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Jesujoba O. Alabi, Israel Abebe Azime, Miaoran Zhang, Cristina España-Bonet, Rachel Bawden, David Ifeoluwa Adelani, Clement Odoje, Idris Akinade, Iffat Maab, Davis David, Shamsuddeen Hassan Muhammad, Neo Putini, David O. Ademuyiwa, Andrew Caines, Dietrich Klakow |
EMNLP | 16 |
| 2025 | Charting the Landscape of African NLP: Mapping Progress and Shaping the Road AheadabstractWith over 2,000 languages and potentially millions of speakers, Africa represents one of the richest linguistic regions in the world.Yet, this diversity is scarcely reflected in state-of-the-art natural language processing (NLP) systems and large language models (LLMs), which predominantly support a narrow set of high-resource languages.This exclusion not only limits the reach and utility of modern NLP technologies but also risks widening the digital divide across linguistic communities.Nevertheless, NLP research on African languages is active and growing.In recent years, there has been a surge of interest in this area, driven by several factors-including the creation of multilingual language resources, the rise of community-led initiatives, and increased support through funding programs.In this survey, we analyze 884 research papers on NLP for African languages published over the past five years, offering a comprehensive overview of recent progress across core tasks.We identify key trends shaping the field and conclude by outlining promising directions to foster more inclusive and sustainable NLP research for African languages.1 27,593 papers (All papers, except expert recommendations, were collected using keyword matching via the Semantic Scholar API.) Jesujoba O. Alabi, Michael A. Hedderich, David Ifeoluwa Adelani, Dietrich Klakow |
EMNLP | 4 |
| 2025 | PricingLogic: Evaluating LLMs Reasoning on Complex Tourism Pricing TasksabstractWe present PricingLogic, the first benchmark that probes whether Large Language Models (LLMs) can reliably automate tourism-related prices when multiple, overlapping fare rules apply.Travel agencies are eager to offload this error-prone task onto AI systems; however, deploying LLMs without verified reliability could result in significant financial losses and erode customer trust.PricingLogic comprises 300 natural-language questions based on booking requests derived from 42 real-world pricing policies, spanning two levels of difficulty: (i) basic customer-type pricing and (ii) bundled-tour calculations involving interacting discounts.Evaluations of a line of LLMs reveal a steep performance drop on the harder tier, exposing systematic failures in rule interpretation and arithmetic reasoning.These results highlight that, despite their general capabilities, today's LLMs remain unreliable in revenuecritical applications without further safeguards or domain adaptation. Yunuo Liu, Zena Al-Khalili, Dai Cheng, Yanjun Chen 0001, Dietrich Klakow, Wei Zhang 0185, Xiaoyu Shen 0001 |
EMNLP | 6 |
| 2025 | Improving Semantic Understanding in Speech Language Models via Brain-tuningabstractSpeech language models align with human brain responses to natural language to an impressive degree. However, current models rely heavily on low-level speech features, indicating they lack brain-relevant semantics which limits their utility as model organisms of semantic processing in the brain. In this work, we address this limitation by inducing brain-relevant bias directly into the models via fine-tuning with fMRI recordings of people listening to natural stories--a process we name brain-tuning. After testing it on 3 different pretrained model families, we show that brain-tuning not only improves overall alignment with new brain recordings in semantic language regions, but also reduces the reliance on low-level speech features for this alignment. Excitingly, we further show that brain-tuning leads to 1) consistent improvements in performance on semantic downstream tasks and 2) a representational space with increased semantic preference. Our results provide converging evidence, for the first time, that incorporating brain signals into the training of language models improves the models’ semantic understanding. We make the code available at https://github.com/bridge-ai-neuro/brain-tuning. Omer Moussa, Dietrich Klakow, Mariya Toneva |
ICLR | 2 |
| 2025 | Voice Conversion Improves Cross-Domain Robustness for Spoken Arabic Dialect Identification
Badr Abdullah, Matthew Baas, Bernd Möbius, Dietrich Klakow |
INTERSPEECH | 4 |
| 2025 | AfriHuBERT: A self-supervised speech representation model for African languages
Jesujoba O. Alabi, Xuechen Liu 0001, Dietrich Klakow, Junichi Yamagishi |
INTERSPEECH | 3 |
| 2024 | What Are the Rules? Discovering Constraints from DataabstractConstraint programming and AI planning are powerful tools for solving assignment, optimization, and scheduling problems. They require, however, the rarely available combination of domain knowledge and mathematical modeling expertise. Learning constraints from exemplary solutions can close this gap and alleviate the effort of modeling. Existing approaches either require extensive user interaction, need exemplary invalid solutions that must be generated by experts at great expense, or show high noise-sensitivity. We aim to find constraints from potentially noisy solutions, without the need of user interaction. To this end, we formalize the problem in terms of the Minimum Description Length (MDL) principle, by which we select the model with the best lossless compression of the data. Solving the problem involves model counting, which is #P-hard to approximate. We therefore propose the greedy URPILS algorithm to find high-quality constraints in practice. Extensive experiments on constraint programming and AI planning benchmark data show URPILS not only finds more accurate and succinct constraints, but also is more robust to noise, and has lower sample complexity than the state of the art. Boris Wiegand, Dietrich Klakow, Jilles Vreeken |
AAAI | 2 |
| 2024 | The Hidden Space of Transformer Language AdaptersabstractWe analyze the operation of transformer language adapters, which are small modules trained on top of a frozen language model to adapt its predictions to new target languages.We show that adapted predictions mostly evolve in the source language the model was trained on, while the target language becomes pronounced only in the very last layers of the model.Moreover, the adaptation process is gradual and distributed across layers, where it is possible to skip small groups of adapters without decreasing adaptation performance.Last, we show that adapters operate on top of the model's frozen representation space while largely preserving its structure, rather than on an "isolated" subspace.Our findings provide a deeper view into the adaptation process of language models to new languages, showcasing the constraints imposed on it by the underlying model and introduces practical implications to enhance its efficiency. 1 Jesujoba O. Alabi, Marius Mosbach, Matan Eyal, Dietrich Klakow, Mor Geva |
ACL (1) | 4 |
| 2024 | Learning Part-whole Hierarchies from the Sequence of Handwriting
David Schlangen, Dietrich Klakow |
CogSci | 3 |
| 2024 | Who Did You Blame When Your Project Failed? Designing a Corpus for Presupposition Generation in Cross-Examination DialoguesabstractThis paper introduces the corpus for the novel task of presupposition generation - a natural language generation problem where a model produces a list of presuppositions carried by the given input sentence, in the context of the presented research - given the cross-examination question. Two datasets, PECaN (Presupposition, Entailment, Contradiction and Neutral) and PGen (Presuppostion Generation), are designed to fine-tune existing BERT (CITATION) and T5 (CITATION) models for classification and generation tasks. Various corpora construction methods are proposed ranging from manual annotations, prompting the GPT 3.0 model, to augmenting data from the existing corpora. The fine-tuned models achieved high accuracy on the novel Presupposition as Natural Language Inference (PNLI) task which extends the traditional Natural Language Inference (NLI) incorporating instances of presupposition into classification. T5 outperforms BERT by broad margin achieving an overall accuracy of 84.35% compared to 71.85% of BERT, and specifically when classifying presuppositions (93% vs 73% respectively). Regarding presupposition generation, we observed that despite the limited amount of data used for fine-tuning, the model displays an emerging proficiency in generation presuppositions reaching ROUGE scores of 43.47, adhering to systematic patterns that mirror valid strategies for presupposition generation, although failed to generate the complete lists. Maria Francis, Julius Steuer, Dietrich Klakow, Volha Petukhova |
LREC/COLING | 3 |
| 2024 | Annotating Customer-Oriented Behaviour in Call Centre Sales DialoguesabstractCustomer-oriented behaviour (COB) plays an important role in call centre interactions, particularly in the context of successful sales negotiation. However, the evaluation of COB in customer-agent conversations often lacks clarity in its definition and robust computational assessment methods. This paper addresses these challenges by presenting a comprehensive conceptual and empirical framework. We conducted multidimensional dialogue act annotations on authentic call centre interactions using the ISO 24617-2 taxonomy, capturing the multifaceted nature of these interactions. This process led to the identification of relevant dialogue act categories, proposed extensions concerning relationship-building aspects, and derived corpus statistics. The findings highlight specific facets of COB that positively impact on Customer Satisfaction (CS), as determined through correlation analysis. Additionally, we delved into the dependencies between COB and feedback acts, leveraging the hierarchical structure of the DIT++ model. This framework improves our understanding of the dynamics shaping sales strategies in call centres and holds promise for practical applications in optimising customer-agent interactions. Jutta Stock, Volha Petukhova, Dietrich Klakow |
LREC/COLING | 3 |
| 2024 | EthioLLM: Multilingual Large Language Models for Ethiopian Languages with Task EvaluationabstractLarge language models (LLMs) have gained popularity recently due to their outstanding performance in various downstream Natural Language Processing (NLP) tasks. However, low-resource languages are still lagging behind current state-of-the-art (SOTA) developments in the field of NLP due to insufficient resources to train LLMs. Ethiopian languages exhibit remarkable linguistic diversity, encompassing a wide array of scripts, and are imbued with profound religious and cultural significance. This paper introduces EthioLLM – multilingual large language models for five Ethiopian languages (Amharic, Ge’ez, Afan Oromo, Somali, and Tigrinya) and English, and Ethiobenchmark – a new benchmark dataset for various downstream NLP tasks. We evaluate the performance of these models across five downstream NLP tasks. We open-source our multilingual language models, new benchmark datasets for various downstream tasks, and task-specific fine-tuned language models and discuss the performance of the models. Our dataset and models are available at the https://huggingface.co/EthioNLP repository. Atnafu Lambebo Tonja, Israel Abebe Azime, Tadesse Destaw Belay, Mesay Gemeda Yigezu, Moges Ahmed Mehamed, Abinew Ali Ayele, Ebrahim Chekol Jibril, Michael Melese Woldeyohannis, Olga Kolesnikova, Philipp Slusallek, Dietrich Klakow, Seid Muhie Yimam |
LREC/COLING | 11 |
| 2024 | From Insights to Actions: The Impact of Interpretability and Analysis Research on NLPabstractInterpretability and analysis (IA) research is a growing subfield within NLP with the goal of developing a deeper understanding of the behavior or inner workings of NLP systems and methods.Despite growing interest in the subfield, a criticism of this work is that it lacks actionable insights and therefore has little impact on NLP.In this paper, we seek to quantify the impact of IA research on the broader field of NLP.We approach this with a mixed-methods analysis 1 of: (1) a citation graph of 185K+ papers built from all papers published at ACL and EMNLP conferences from 2018 to 2023, and their references and citations, and (2) a survey of 138 members of the NLP community.Our quantitative results show that IA work is well-cited outside of IA, and central in the NLP citation graph.Through qualitative analysis of survey responses and manual annotation of 556 papers, we find that NLP researchers build on findings from IA work and perceive it as important for progress in NLP, multiple subfields, and rely on its findings and terminology for their own work.Many novel methods are proposed based on IA findings and highly influenced by them, but highly influential non-IA work cites IA findings without being driven by them.We end by summarizing what is missing in IA work today and provide a call to action, to pave the way for a more impactful future of IA research. Marius Mosbach, Vagrant Gautam, Tomás Vergara Browne, Dietrich Klakow, Mor Geva |
EMNLP | 4 |
| 2024 | Understanding "Democratization" in NLP and ML ResearchabstractRecent improvements in natural language processing (NLP) and machine learning (ML) and increased mainstream adoption have led to researchers frequently discussing the "democratization" of artificial intelligence.In this paper, we seek to clarify how democratization is understood in NLP and ML publications, through large-scale mixed-methods analyses of papers using the keyword "democra*" published in NLP and adjacent venues.We find that democratization is most frequently used to convey (ease of) access to or use of technologies, without meaningfully engaging with theories of democratization, while research using other invocations of "democra*" tends to be grounded in theories of deliberation and debate.Based on our findings, we call for researchers to enrich their use of the term democratization with appropriate theory, towards democratic technologies beyond superficial access. 1 Arjun Subramonian, Vagrant Gautam, Dietrich Klakow, Zeerak Talat |
EMNLP | 3 |
| 2024 | Fine-Tuning Large Language Models to Translate: Will a Touch of Noisy Data in Misaligned Languages Suffice?abstractTraditionally, success in multilingual machine translation can be attributed to three key factors in training data: large volume, diverse translation directions, and high quality.In the current practice of fine-tuning large language models (LLMs) for translation, we revisit the importance of these factors.We find that LLMs display strong translation capability after being fine-tuned on as few as 32 parallel sentences and that fine-tuning on a single translation direction enables translation in multiple directions.However, the choice of direction is critical: fine-tuning LLMs with only English on the target side can lead to task misinterpretation, which hinders translation into non-English languages.Problems also arise when noisy synthetic data is placed on the target side, especially when the target language is wellrepresented in LLM pre-training.Yet interestingly, synthesized data in an under-represented language has a less pronounced effect.Our findings suggest that when adapting LLMs to translation, the requirement on data quantity can be eased but careful considerations are still crucial to prevent an LLM from exploiting unintended data biases. Pinzhen Chen, Miaoran Zhang, Barry Haddow, Xiaoyu Shen 0001, Dietrich Klakow |
EMNLP | 6 |
| 2024 | Self-Supervised Adaptive Pre-Training of Multilingual Speech Models for Language and Dialect IdentificationabstractTransformer-based, pre-trained speech models have shown striking performance when fine-tuned on various downstream tasks such as automatic speech recognition and spoken language identification (SLID). However, the problem of domain mismatch remains a challenge in this area, where the domain of the pre-training data might differ from that of the downstream labeled data used for fine-tuning. In multilingual tasks such as SLID, the pre-trained speech model may not support all the languages in the downstream task. To address this challenge, we propose self-supervised adaptive pre-training (SAPT) to adapt the pre-trained model to the target domain and languages of the downstream task. We apply SAPT to the XLSR-128 model and investigate the effectiveness of this approach for the SLID task. First, we demonstrate that SAPT improves XLSR performance on the FLEURS benchmark with substantial gains up to 40.1% for under-represented languages. Second, we apply SAPT on four different datasets in a few-shot learning setting, showing that our approach improves the sample efficiency of XLSR during fine-tuning. Our experiments provide strong empirical evidence that continual adaptation via self-supervision improves downstream performance for multilingual speech models. Mohammed Maqsood Shaik, Dietrich Klakow, Badr Abdullah |
ICASSP | 2 |
| 2024 | Wave to Interlingua: Analyzing Representations of Multilingual Speech Transformers for Spoken Language Translation
Badr Abdullah, Mohammed Maqsood Shaik, Dietrich Klakow |
INTERSPEECH | 3 |
| 2024 | Joint vs Sequential Speaker-Role Detection and Automatic Speech Recognition for Air-traffic Control
Alexander Blatt, Aravind Krishnan, Dietrich Klakow |
INTERSPEECH | 3 |
| 2024 | On the Encoding of Gender in Transformer-based ASR Representations
Aravind Krishnan, Badr Abdullah, Dietrich Klakow |
INTERSPEECH | 3 |
| 2024 | Cross-Linguistic Intelligibility of Non-Compositional Expressions in Spoken Context
Iuliia Zaitova, Irina Stenger, Tania Avgustinova, Bernd Möbius, Dietrich Klakow |
INTERSPEECH | 6 |
| 2024 | A Preference-driven Paradigm for Enhanced Translation with Large Language ModelsabstractDawei Zhu, Sony Trenous, Xiaoyu Shen, Dietrich Klakow, Bill Byrne, Eva Hasler. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Sony Trenous, Xiaoyu Shen 0001, Dietrich Klakow, William J. Byrne, Eva Hasler |
NAACL-HLT | 4 |
| 2024 | Robust Pronoun Fidelity with English LLMs: Are they Reasoning, Repeating, or Just Biased?abstractAbstract Robust, faithful, and harm-free pronoun use for individuals is an important goal for language model development as their use increases, but prior work tends to study only one or two of these characteristics at a time. To measure progress towards the combined goal, we introduce the task of pronoun fidelity: Given a context introducing a co-referring entity and pronoun, the task is to reuse the correct pronoun later. We present RUFF, a carefully designed dataset of over 5 million instances to measure robust pronoun fidelity in English, and we evaluate 37 model variants from nine popular families, across architectures (encoder-only, decoder-only, and encoder-decoder) and scales (11M-70B parameters). When an individual is introduced with a pronoun, models can mostly faithfully reuse this pronoun in the next sentence, but they are significantly worse with she/her/her, singular they, and neopronouns. Moreover, models are easily distracted by non-adversarial sentences discussing other people; even one sentence with a distractor pronoun causes accuracy to drop on average by 34 percentage points. Our results show that pronoun fidelity is not robust, in a simple, naturalistic setting where humans achieve nearly 100% accuracy. We encourage researchers to bridge the gaps we find and to carefully evaluate reasoning in settings where superficial repetition might inflate perceptions of model performance. Vagrant Gautam, Eileen Bingert, Anne Lauscher, Dietrich Klakow |
Trans. Assoc. Comput. Linguistics | 5 |
| 2023 | MasakhaPOS: Part-of-Speech Tagging for Typologically Diverse African languagesabstractCheikh M. Bamba Dione, David Ifeoluwa Adelani, Peter Nabende, Jesujoba Alabi, Thapelo Sindane, Happy Buzaaba, Shamsuddeen Hassan Muhammad, Chris Chinenye Emezue, Perez Ogayo, Anuoluwapo Aremu, Catherine Gitau, Derguene Mbaye, Jonathan Mukiibi, Blessing Sibanda, Bonaventure F. P. Dossou, Andiswa Bukula, Rooweither Mabuya, Allahsera Auguste Tapo, Edwin Munkoh-Buabeng, Victoire Memdjokam Koagne, Fatoumata Ouoba Kabore, Amelia Taylor, Godson Kalipe, Tebogo Macucwa, Vukosi Marivate, Tajuddeen Gwadabe, Mboning Tchiaze Elvis, Ikechukwu Onyenwe, Gratien Atindogbe, Tolulope Adelani, Idris Akinade, Olanrewaju Samuel, Marien Nahimana, Théogène Musabeyezu, Emile Niyomutabazi, Ester Chimhenga, Kudzai Gotosa, Patrick Mizha, Apelete Agbolo, Seydou Traore, Chinedu Uchechukwu, Aliyu Yusuf, Muhammad Abdullahi, Dietrich Klakow. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Cheikh M. Bamba Dione, David Ifeoluwa Adelani, Peter Nabende, Jesujoba O. Alabi, Thapelo Sindane, Happy Buzaaba, Shamsuddeen Hassan Muhammad, Chris C. Emezue, Perez Ogayo, Aremu Anuoluwapo, Catherine Gitau, Derguene Mbaye, Jonathan Mukiibi, Blessing K. Sibanda, Bonaventure F. P. Dossou, Andiswa Bukula, Rooweither Mabuya, Allahsera Tapo, Edwin Munkoh-Buabeng, Victoire Memdjokam Koagne, Fatoumata Ouoba Kabore, Amelia V. Taylor, Godson Kalipe, Tebogo Macucwa, Vukosi Marivate, Tajuddeen Rabiu Gwadabe, Elvis Mboning, Ikechukwu E. Onyenwe, Gratien Atindogbe, Tolulope Anu Adelani, Idris Akinade, Samuel Olanrewaju, Marien Nahimana, Théogène Musabeyezu, Emile Niyomutabazi, Ester Chimhenga, Kudzai Gotosa, Patrick Mizha, Apelete Agbolo, Seydou Traore, Chinedu Uchechukwu, Aliyu Yusuf, Muhammad Abdullahi, Dietrich Klakow |
ACL (1) | 44 |
| 2023 | Weaker Than You Think: A Critical Look at Weakly Supervised LearningabstractWeakly supervised learning is a popular approach for training machine learning models in low-resource settings.Instead of requesting high-quality yet costly human annotations, it allows training models with noisy annotations obtained from various weak sources.Recently, many sophisticated approaches have been proposed for robust training under label noise, reporting impressive results.In this paper, we revisit the setup of these approaches and find that the benefits brought by these approaches are significantly overestimated.Specifically, we find that the success of existing weakly supervised learning approaches heavily relies on the availability of clean validation samples which, as we show, can be leveraged much more efficiently by simply training on them.After using these clean labels in training, the advantages of using these sophisticated approaches are mostly wiped out.This remains true even when reducing the size of the available clean data to just five samples per class, making these approaches impractical.To understand the true value of weakly supervised learning, we thoroughly analyze diverse NLP datasets and tasks to ascertain when and why weakly supervised approaches work.Based on our findings, we provide recommendations for future research.1 Xiaoyu Shen 0001, Marius Mosbach, Andreas Stephan, Dietrich Klakow |
ACL (1) | 5 |
| 2023 | Enabling Noisy Label Usage for Out-of-Airspace Data in Read-Back Error DetectionabstractDeveloping language understanding (NLU) methods for low ressource domains is an ongoing challenge. The air-traffic control (ATC) domain is a paragon of this. There is a high pressure for automatized solutions to ease the workload of air-traffic controllers (ATCOs), but a low availability of open-source datasets. The available datasets contain mostly unlabeled transcripts, targeting automatic speech recognition (ASR) and cover just one or a few airspaces. Models trained on these airspaces might fail on unseen target airspace. We evaluate different methods to overcome this problem on the task of read-back error detection (RED), which uncovers mistakes in ATCO-pilot communication to prevent incidents. We generate noisy labels for our two stage RED approach, that combines data augmentation and noisy labels. This allows the use of unlabeled data of non-target airspaces to increase the performance on the target airspaces with a relative improvement of 35 % over the baseline method. Lakshmi Rajendram Bashyam, Alexander Blatt, Dietrich Klakow |
ASRU | 3 |
| 2023 | Ending the Blind Flight: Analyzing the Impact of Acoustic and Lexical Factors on WAV2VEC 2.0 in Air-Traffic ControlabstractTransformer neural networks have shown remarkable success on standard automatic speech recognition (ASR) benchmarks. However, they are known to be less robust against domain mismatch, particularly with air traffic control (ATC) speech data. In the ATC domain, transformer-based ASR systems do usually not transfer across different datasets. The reasons for poor transferability across ATC datasets remain unclear. Our study investigates the influence of acoustic variability and lexical differences on the ASR performance across various ATC datasets. By fine-tuning and evaluating wav2vec 2.0 on synthetic ATC datasets, we examine the effect of acoustic variability on the model performance. Furthermore, we assess the effect of lexical differences by correlating language model perplexity with performance. Our findings reveal that a combination of acoustic and lexical mismatch causes the bad inter-dataset transferability and give insights on how to improve future ASR models for ATC. Alexander Blatt, Badr Abdullah, Dietrich Klakow |
ASRU | 3 |
| 2023 | Multilingual Normalization of Temporal Expressions with Masked Language ModelsabstractThe detection and normalization of temporal expressions is an important task and preprocessing step for many applications.However, prior work on normalization is rule-based, which severely limits the applicability in realworld multilingual settings, due to the costly creation of new rules.We propose a novel neural method for normalizing temporal expressions based on masked language modeling.Our multilingual method outperforms prior rule-based systems in many languages, and in particular, for low-resource languages with performance improvements of up to 33 F 1 on average compared to the state of the art. Lukas Lange, Jannik Strötgen, Heike Adel, Dietrich Klakow |
EACL | 4 |
| 2023 | Meta Self-Refinement for Robust Learning with Weak SupervisionabstractTraining deep neural networks (DNNs) under weak supervision has attracted increasing research attention as it can significantly reduce the annotation cost.However, labels from weak supervision can be noisy, and the high capacity of DNNs enables them to easily overfit the label noise, resulting in poor generalization.Recent methods leverage self-training to build noiseresistant models, in which a teacher trained under weak supervision is used to provide highly confident labels for teaching the students.Nevertheless, the teacher derived from such frameworks may have fitted a substantial amount of noise and therefore produce incorrect pseudolabels with high confidence, leading to severe error propagation.In this work, we propose Meta Self-Refinement (MSR), a noise-resistant learning framework, to effectively combat label noise from weak supervision.Instead of relying on a fixed teacher trained with noisy labels, we encourage the teacher to refine its pseudolabels.At each training step, MSR performs a meta gradient descent on the current mini-batch to maximize the student performance on a clean validation set.Extensive experimentation on eight NLP benchmarks demonstrates that MSR is robust against label noise in all settings and outperforms state-of-the-art methods by up to 11.4% in accuracy and 9.26% in F1 score. Xiaoyu Shen 0001, Michael A. Hedderich, Dietrich Klakow |
EACL | 4 |
| 2023 | An Information-Theoretic Analysis of Self-supervised Discrete Representations of Speech
Badr Abdullah, Mohammed Maqsood Shaik, Bernd Möbius, Dietrich Klakow |
INTERSPEECH | 4 |
| 2023 | On the N-gram Approximation of Pre-trained Language Models
Aravind Krishnan, Jesujoba O. Alabi, Dietrich Klakow |
INTERSPEECH | 3 |
| 2023 | Why Are We Waiting? Discovering Interpretable Models for Predicting Sojourn and Waiting TimesabstractQueueing models explain waiting times, predict sojourn times and help to identify and avoid bottlenecks. Domain experts usually create these models by intensive handcrafting, often resulting in idealized models not fitting the actual process behavior well. Discovering queueing models from data can alleviate this effort, but existing methods do not suffice as they are unable to model complex queueing behaviors. We propose a novel approach to discover queueing models for interpretable waiting time prediction using a rich modeling language to fit complex processes. We formalize the problem in terms of the Minimum Description Length (MDL) principle, by which the best model gives the best lossless compression. The resulting optimization problem is computationally hard, and hence we propose the greedy CueMin algorithm to efficiently find good queueing models from data. Through an extensive set of experiments including a case study on call center data, we show it discovers inherently interpretable models, which explain and predict behavior of waiting lines better than the state of the art. Boris Wiegand, Dietrich Klakow, Jilles Vreeken |
SDM | 2 |
| 2022 | Discovering Interpretable Data-to-Sequence GeneratorsabstractWe study the problem of predicting an event sequence given some meta data. In particular, we are interested in learning easily interpretable models that can accurately generate a sequence based on an attribute vector. To this end, we propose to learn a sparse event-flow graph over the training sequences, and statistically robust rules that use meta data to determine which paths to follow. We formalize the problem in terms of the Minimum Description Length (MDL) principle, by which we identify the best model as the one that compresses the data best. As the resulting optimization problem is NP-hard, we propose the efficient ConSequence algorithm to discover good event-flow graphs from data. Through an extensive set of experiments including a case study, we show that it ably discovers compact, interpretable and accurate models for the generation and prediction of event sequences from data, has a low sample complexity, and is particularly robust against noise. Boris Wiegand, Dietrich Klakow, Jilles Vreeken |
AAAI | 2 |
| 2022 | Adapting Pre-trained Language Models to African Languages via Multilingual Adaptive Fine-TuningabstractMultilingual pre-trained language models (PLMs) have demonstrated impressive performance on several downstream tasks for both high-resourced and low-resourced languages. However, there is still a large performance drop for languages unseen during pre-training, especially African languages. One of the most effective approaches to adapt to a new language is language adaptive fine-tuning (LAFT) — fine-tuning a multilingual PLM on monolingual texts of a language using the pre-training objective. However, adapting to target language individually takes large disk space and limits the cross-lingual transfer abilities of the resulting models because they have been specialized for a single language. In this paper, we perform multilingual adaptive fine-tuning on 17 most-resourced African languages and three other high-resource languages widely spoken on the African continent to encourage cross-lingual transfer learning. To further specialize the multilingual PLM, we removed vocabulary tokens from the embedding layer that corresponds to non-African writing scripts before MAFT, thus reducing the model size by around 50%. Our evaluation on two multilingual PLMs (AfriBERTa and XLM-R) and three NLP tasks (NER, news topic classification, and sentiment classification) shows that our approach is competitive to applying LAFT on individual languages while requiring significantly less disk space. Additionally, we show that our adapted PLM also improves the zero-shot cross-lingual transfer abilities of parameter efficient fine-tuning methods. Jesujoba O. Alabi, David Ifeoluwa Adelani, Marius Mosbach, Dietrich Klakow |
COLING | 4 |
| 2022 | MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity RecognitionabstractDavid Adelani, Graham Neubig, Sebastian Ruder, Shruti Rijhwani, Michael Beukman, Chester Palen-Michel, Constantine Lignos, Jesujoba Alabi, Shamsuddeen Muhammad, Peter Nabende, Cheikh M. Bamba Dione, Andiswa Bukula, Rooweither Mabuya, Bonaventure F. P. Dossou, Blessing Sibanda, Happy Buzaaba, Jonathan Mukiibi, Godson Kalipe, Derguene Mbaye, Amelia Taylor, Fatoumata Kabore, Chris Chinenye Emezue, Anuoluwapo Aremu, Perez Ogayo, Catherine Gitau, Edwin Munkoh-Buabeng, Victoire Memdjokam Koagne, Allahsera Auguste Tapo, Tebogo Macucwa, Vukosi Marivate, Mboning Tchiaze Elvis, Tajuddeen Gwadabe, Tosin Adewumi, Orevaoghene Ahia, Joyce Nakatumba-Nabende, Neo Lerato Mokono, Ignatius Ezeani, Chiamaka Chukwuneke, Mofetoluwa Oluwaseun Adeyemi, Gilles Quentin Hacheme, Idris Abdulmumin, Odunayo Ogundepo, Oreen Yousuf, Tatiana Moteu, Dietrich Klakow. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. David Ifeoluwa Adelani, Graham Neubig, Sebastian Ruder, Shruti Rijhwani, Michael Beukman, Chester Palen-Michel, Constantine Lignos, Jesujoba O. Alabi, Shamsuddeen Hassan Muhammad, Peter Nabende, Cheikh M. Bamba Dione, Andiswa Bukula, Rooweither Mabuya, Bonaventure F. P. Dossou, Blessing K. Sibanda, Happy Buzaaba, Jonathan Mukiibi, Godson Kalipe, Derguene Mbaye, Amelia V. Taylor, Fatoumata Ouoba Kabore, Chris C. Emezue, Aremu Anuoluwapo, Perez Ogayo, Catherine Gitau, Edwin Munkoh-Buabeng, Victoire Memdjokam Koagne, Allahsera Tapo, Tebogo Macucwa, Vukosi Marivate, Elvis Mboning, Tajuddeen Rabiu Gwadabe, Tosin P. Adewumi, Orevaoghene Ahia, Joyce Nakatumba-Nabende, Neo L. Mokono, Ignatius Ezeani, Chiamaka Ijeoma Chukwuneke, Mofe Adeyemi, Gilles Hacheme, Idris Abdulmumin, Odunayo Ogundepo, Oreen Yousuf, Tatiana Moteu Ngoli, Dietrich Klakow |
EMNLP | 45 |
| 2022 | Call-Sign Recognition and Understanding for Noisy Air-Traffic Transcripts Using Surveillance InformationabstractAir traffic control (ATC) relies on communication via speech between pilot and air-traffic controller (ATCO). The call-sign, as unique identifier for each flight, is used to address a specific pilot by the ATCO. Extracting the call-sign from the communication is a challenge because of the noisy ATC voice channel and the additional noise introduced by the receiver. A low signal-to-noise ratio (SNR) in the speech leads to high word error rate (WER) transcripts. We propose a new call-sign recognition and understanding (CRU) system that addresses this issue. The recognizer is trained to identify call-signs in noisy ATC transcripts and convert them into the standard International Civil Aviation Organization (ICAO) format. By incorporating surveillance information, we can multiply the call-sign accuracy (CSA) up to a factor of four. The introduced data augmentation adds additional performance on high WER transcripts and allows the adaptation of the model to unseen airspaces. Alexander Blatt, Martin Kocour, Karel Veselý, Igor Szöke, Dietrich Klakow |
ICASSP | 5 |
| 2022 | Label-Descriptive Patterns and Their Application to Characterizing Classification ErrorsabstractState-of-the-art deep learning methods achieve human-like performance on many tasks, but make errors nevertheless. Characterizing these errors in easily interpretable terms gives insight into whether a classifier is prone to making systematic errors, but also gives a way to act and improve the classifier. We propose to discover those feature-value combinations (i.e., patterns) that strongly correlate with correct resp. erroneous predictions to obtain a global and interpretable description for arbitrary classifiers. We show this is an instance of the more general label description problem, which we formulate in terms of the Minimum Description Length principle. To discover a good pattern set, we develop the efficient Premise algorithm. Through an extensive set of experiments we show it performs very well in practice on both synthetic and real-world data. Unlike existing solutions, it ably recovers ground truth patterns, even on highly imbalanced data over many features. Through two case studies on Visual Question Answering and Named Entity Recognition, we confirm that Premise gives clear and actionable insight into the systematic errors made by modern NLP classifiers. Michael A. Hedderich, Jonas Fischer, Dietrich Klakow, Jilles Vreeken |
ICML | 3 |
| 2022 | Integrating Form and Meaning: A Multi-Task Learning Model for Acoustic Word Embeddings
Badr Abdullah, Bernd Möbius, Dietrich Klakow |
INTERSPEECH | 3 |
| 2022 | Placing M-Phasis on the Plurality of Hate: A Feature-Based Corpus of Hate OnlineabstractEven though hate speech (HS) online has been an important object of research in the last decade, most HS-related corpora over-simplify the phenomenon of hate by attempting to label user comments as “hate” or “neutral”. This ignores the complex and subjective nature of HS, which limits the real-life applicability of classifiers trained on these corpora. In this study, we present the M-Phasis corpus, a corpus of ~9k German and French user comments collected from migration-related news articles. It goes beyond the “hate”-“neutral” dichotomy and is instead annotated with 23 features, which in combination become descriptors of various types of speech, ranging from critical comments to implicit and explicit expressions of hate. The annotations are performed by 4 native speakers per language and achieve high (0.77 <= k <= 1) inter-annotator agreements. Besides describing the corpus creation and presenting insights from a content, error and domain analysis, we explore its data characteristics by training several classification baselines. Dana Ruiter, Liane Reiners, Ashwin Geet D'Sa, Thomas Kleinbauer, Dominique Fohr, Irina Illina, Dietrich Klakow, Christian Schemer, Angeliki Monnier |
LREC | 7 |
| 2022 | Modeling the Impact of Syntactic Distance and Surprisal on Cross-Slavic Text ComprehensionabstractWe focus on the syntactic variation and measure syntactic distances between nine Slavic languages (Belarusian, Bulgarian, Croatian, Czech, Polish, Slovak, Slovene, Russian, and Ukrainian) using symmetric measures of insertion, deletion and movement of syntactic units in the parallel sentences of the fable “The North Wind and the Sun”. Additionally, we investigate phonetic and orthographic asymmetries between selected languages by means of the information theoretical notion of surprisal. Syntactic distance and surprisal are, thus, considered as potential predictors of mutual intelligibility between related languages. In spoken and written cloze test experiments for Slavic native speakers, the presented predictors will be validated as to whether variations in syntax lead to a slower or impeded intercomprehension of Slavic texts. Irina Stenger, Philip Georgis, Tania Avgustinova, Bernd Möbius, Dietrich Klakow |
LREC | 5 |
| 2022 | Adapting Language Models When Training on Privacy-Transformed DataabstractIn recent years, voice-controlled personal assistants have revolutionized the interaction with smart devices and mobile applications. The collected data are then used by system providers to train language models (LMs). Each spoken message reveals personal information, hence removing private information from the input sentences is necessary. Our data sanitization process relies on recognizing and replacing named entities by other words from the same class. However, this may harm LM training because privacy-transformed data is unlikely to match the test distribution. This paper aims to fill the gap by focusing on the adaptation of LMs initially trained on privacy-transformed sentences using a small amount of original untransformed data. To do so, we combine class-based LMs, which provide an effective approach to overcome data sparsity in the context of n-gram LMs, and neural LMs, which handle longer contexts and can yield better predictions. Our experiments show that training an LM on privacy-transformed data result in a relative 11% word error rate (WER) increase compared to training on the original untransformed data, and adapting that model on a limited amount of original untransformed data leads to a relative 8% WER improvement over the model trained solely on privacy-transformed data. M. A. Tugtekin Turan, Dietrich Klakow, Emmanuel Vincent 0001, Denis Jouvet |
LREC | 2 |
| 2022 | A Few Thousand Translations Go a Long Way! Leveraging Pre-trained Models for African News TranslationabstractDavid Adelani, Jesujoba Alabi, Angela Fan, Julia Kreutzer, Xiaoyu Shen, Machel Reid, Dana Ruiter, Dietrich Klakow, Peter Nabende, Ernie Chang, Tajuddeen Gwadabe, Freshia Sackey, Bonaventure F. P. Dossou, Chris Emezue, Colin Leong, Michael Beukman, Shamsuddeen Muhammad, Guyo Jarso, Oreen Yousuf, Andre Niyongabo Rubungo, Gilles Hacheme, Eric Peter Wairagala, Muhammad Umair Nasir, Benjamin Ajibade, Tunde Ajayi, Yvonne Gitau, Jade Abbott, Mohamed Ahmed, Millicent Ochieng, Anuoluwapo Aremu, Perez Ogayo, Jonathan Mukiibi, Fatoumata Ouoba Kabore, Godson Kalipe, Derguene Mbaye, Allahsera Auguste Tapo, Victoire Memdjokam Koagne, Edwin Munkoh-Buabeng, Valencia Wagner, Idris Abdulmumin, Ayodele Awokoya, Happy Buzaaba, Blessing Sibanda, Andiswa Bukula, Sam Manthalu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. David Ifeoluwa Adelani, Jesujoba O. Alabi, Angela Fan, Julia Kreutzer, Xiaoyu Shen 0001, Machel Reid, Dana Ruiter, Dietrich Klakow, Peter Nabende, Ernie Chang, Tajuddeen Rabiu Gwadabe, Sackey Freshia, Bonaventure F. P. Dossou, Chris C. Emezue, Colin Leong, Michael Beukman, Shamsuddeen Hassan Muhammad, Guyo Dub Jarso, Oreen Yousuf, Rubungo Andre Niyongabo, Gilles Hacheme, Eric Peter Wairagala, Muhammad Umair Nasir, Benjamin Ajibade, Tunde Ajayi, Yvonne Wambui, Jade Z. Abbott, Millicent Ochieng, Aremu Anuoluwapo, Perez Ogayo, Jonathan Mukiibi, Fatoumata Ouoba Kabore, Godson Kalipe, Derguene Mbaye, Allahsera Tapo, Victoire Memdjokam Koagne, Edwin Munkoh-Buabeng, Valencia Wagner, Idris Abdulmumin, Ayodele Awokoya, Happy Buzaaba, Blessing K. Sibanda, Andiswa Bukula, Sam Manthalu |
NAACL-HLT | 8 |
| 2022 | MCSE: Multimodal Contrastive Learning of Sentence EmbeddingsabstractMiaoran Zhang, Marius Mosbach, David Adelani, Michael Hedderich, Dietrich Klakow. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Miaoran Zhang, Marius Mosbach, David Ifeoluwa Adelani, Michael A. Hedderich, Dietrich Klakow |
NAACL-HLT | 5 |
| 2022 | A Data-Driven Investigation of Noise-Adaptive Utterance Generation with Linguistic ModificationabstractIn noisy environments, speech can be hard to understand for humans. Spoken dialog systems can help to enhance the intelligibility of their output, either by modifying the speech synthesis (e.g., imitate Lombard speech) or by optimizing the language generation. We here focus on the second type of approach, by which an intended message is realized with words that are more intelligible in a specific noisy environment. By conducting a speech perception experiment, we created a dataset of 900 paraphrases in babble noise, perceived by native English speakers with normal hearing. We find that careful selection of paraphrases can improve intelligibility by 33% at SNR -5 dB. Our analysis of the data shows that the intelligibility differences between paraphrases are mainly driven by noise-robust acoustic cues. Furthermore, we propose an intelligibility-aware paraphrase ranking model, which outperforms baseline models with a relative improvement of 31.37% at SNR -5 dB. Anupama Chingacham, Vera Demberg, Dietrich Klakow |
SLT | 3 |
| 2022 | CLIN-X: pre-trained language models and a study on cross-task transfer for concept extraction in the clinical domainabstractMOTIVATION: The field of natural language processing (NLP) has recently seen a large change toward using pre-trained language models for solving almost any task. Despite showing great improvements in benchmark datasets for various tasks, these models often perform sub-optimal in non-standard domains like the clinical domain where a large gap between pre-training documents and target documents is observed. In this article, we aim at closing this gap with domain-specific training of the language model and we investigate its effect on a diverse set of downstream tasks and settings. RESULTS: We introduce the pre-trained CLIN-X (Clinical XLM-R) language models and show how CLIN-X outperforms other pre-trained transformer models by a large margin for 10 clinical concept extraction tasks from two languages. In addition, we demonstrate how the transformer model can be further improved with our proposed task- and language-agnostic model architecture based on ensembles over random splits and cross-sentence context. Our studies in low-resource and transfer settings reveal stable model performance despite a lack of annotated data with improvements of up to 47 F1 points when only 250 labeled sentences are available. Our results highlight the importance of specialized language models, such as CLIN-X, for concept extraction in non-standard domains, but also show that our task-agnostic model architecture is robust across the tested tasks and languages so that domain- or task-specific adaptations are not required. AVAILABILITY AND IMPLEMENTATION: The CLIN-X language models and source code for fine-tuning and transferring the model are publicly available at https://github.com/boschresearch/clin_x/ and the huggingface model hub. Lukas Lange, Heike Adel, Jannik Strötgen, Dietrich Klakow |
Bioinform. | 4 |
| 2021 | Analysing the Noise Model Error for Realistic Noisy Label DataabstractDistant and weak supervision allow to obtain large amounts of labeled training data quickly and cheaply, but these automatic annotations tend to contain a high amount of errors. A popular technique to overcome the negative effects of these noisy labels is noise modelling where the underlying noise process is modelled. In this work, we study the quality of these estimated noise models from the theoretical side by deriving the expected error of the noise model. Apart from evaluating the theoretical results on commonly used synthetic noise, we also publish NoisyNER, a new noisy label dataset from the NLP domain that was obtained through a realistic distant supervision technique. It provides seven sets of labels with differing noise patterns to evaluate different noise levels on the same instances. Parallel, clean labels are available making it possible to study scenarios where a small amount of gold-standard data can be leveraged. Our theoretical results and the corresponding experiments give insights into the factors that influence the noise model estimation like the noise distribution and the sampling technique. Michael A. Hedderich, Dietrich Klakow |
AAAI | 3 |
| 2021 | SoloFinger: Robust Microgestures while Grasping Everyday ObjectsabstractUsing microgestures, prior work has successfully enabled gestural interactions while holding objects. Yet, these existing methods are prone to false activations caused by natural finger movements while holding or manipulating the object. We address this issue with SoloFinger, a novel concept that allows design of microgestures that are robust against movements that naturally occur during primary activities. Using a data-driven approach, we establish that single-finger movements are rare in everyday hand-object actions and infer a single-finger input technique resilient to false activation. We demonstrate this concept’s robustness using a white-box classifier on a pre-existing dataset comprising 36 everyday hand-object actions. Our findings validate that simple SoloFinger gestures can relieve the need for complex finger configurations or delimiting gestures and that SoloFinger is applicable to diverse hand-object actions. Finally, we demonstrate SoloFinger’s high performance on commodity hardware using random forest classifiers. Adwait Sharma, Michael A. Hedderich, Divyanshu Bhardwaj 0001, Bruno Fruchard, Jess McIntosh, Aditya Shekhar Nittala, Dietrich Klakow, Daniel Ashbrook, Jürgen Steimle |
CHI | 7 |
| 2021 | Preventing Author Profiling through Zero-Shot Multilingual Back-TranslationabstractDocuments as short as a single sentence may inadvertently reveal sensitive information about their authors, including e.g.their gender or ethnicity.Style transfer is an effective way of transforming texts in order to remove any information that enables author profiling.However, for a number of current state-of-theart approaches the improved privacy is accompanied by an undesirable drop in the downstream utility of the transformed data.In this paper, we propose a simple, zero-shot way to effectively lower the risk of author profiling through multilingual back-translation using off-the-shelf translation models.We compare our models with five representative text style transfer models on three datasets across different domains.Results from both an automatic and a human evaluation show that our approach achieves the best overall performance while requiring no training data.We are able to lower the adversarial prediction of gender and race by up to 22% while retaining 95% of the original utility on downstream tasks. David Ifeoluwa Adelani, Miaoran Zhang, Xiaoyu Shen 0001, Ali Davody, Thomas Kleinbauer, Dietrich Klakow |
EMNLP (1) | 6 |
| 2021 | FAME: Feature-Based Adversarial Meta-Embeddings for Robust Input RepresentationsabstractCombining several embeddings typically improves performance in downstream tasks as different embeddings encode different information.It has been shown that even models using embeddings from transformers still benefit from the inclusion of standard word embeddings.However, the combination of embeddings of different types and dimensions is challenging.As an alternative to attention-based meta-embeddings, we propose feature-based adversarial meta-embeddings (FAME) with an attention function that is guided by features reflecting word-specific properties, such as shape and frequency, and show that this is beneficial to handle subword-based embeddings.In addition, FAME uses adversarial training to optimize the mappings of differently-sized embeddings to the same space.We demonstrate that FAME works effectively across languages and domains for sequence labeling and sentence classification, in particular in lowresource settings.FAME sets the new state of the art for POS tagging in 27 languages, various NER settings and question classification in different domains. Lukas Lange, Heike Adel, Jannik Strötgen, Dietrich Klakow |
EMNLP (1) | 4 |
| 2021 | To Share or not to Share: Predicting Sets of Sources for Model Transfer LearningabstractIn low-resource settings, model transfer can help to overcome a lack of labeled data for many tasks and domains.However, predicting useful transfer sources is a challenging problem, as even the most similar sources might lead to unexpected negative transfer results.Thus, ranking methods based on task and text similarity -as suggested in prior workmay not be sufficient to identify promising sources.To tackle this problem, we propose a new approach to automatically determine which and how many sources should be exploited.For this, we study the effects of model transfer on sequence labeling across various domains and tasks and show that our methods based on model similarity and support vector machines are able to predict promising sources, resulting in performance increases of up to 24 F 1 points. Lukas Lange, Jannik Strötgen, Heike Adel, Dietrich Klakow |
EMNLP (1) | 4 |
| 2021 | Estimating Formulas for Model Performance Under Noisy Labels Using Symbolic RegressionabstractWe present a generic formula characterizing the learning of our model under a variety of label-noise settings.This is achieved by using the symbolic regressor model, a genetic programming algorithm, from which we learn functions based on a large set of performance evaluations.Equipped with the knowledge from the regressor, we find a universal formula governing the model performance with respect to noise.This result from our empirical approach could have qualitative applications in mitigating the performance of real-world noisy data and could complement certain noise-robust models. Fech Scen Khoo, Michael A. Hedderich, Dietrich Klakow |
ESANN | 4 |
| 2021 | On the Stability of Fine-tuning BERT: Misconceptions, Explanations, and Strong Baselines
Marius Mosbach, Maksym Andriushchenko, Dietrich Klakow |
ICLR | 3 |
| 2021 | Do Acoustic Word Embeddings Capture Phonological Similarity? An Empirical StudyabstractSeveral variants of deep neural networks have been successfully employed for building parametric models that project variable-duration spoken word segments onto fixed-size vector representations, or acoustic word embeddings (AWEs). However, it remains unclear to what degree we can rely on the distance in the emerging AWE space as an estimate of word-form similarity. In this paper, we ask: does the distance in the acoustic embedding space correlate with phonological dissimilarity? To answer this question, we empirically investigate the performance of supervised approaches for AWEs with different neural architectures and learning objectives. We train AWE models in controlled settings for two languages (German and Czech) and evaluate the embeddings on two tasks: word discrimination and phonological similarity. Our experiments show that (1) the distance in the embedding space in the best cases only moderately correlates with phonological distance, and (2) improving the performance on the word discrimination task does not necessarily yield models that better reflect word phonological similarity. Our findings highlight the necessity to rethink the current intrinsic evaluations for AWEs. Badr Abdullah, Marius Mosbach, Iuliia Zaitova, Bernd Möbius, Dietrich Klakow |
Interspeech | 5 |
| 2021 | Exploring the Potential of Lexical Paraphrases for Mitigating Noise-Induced Comprehension ErrorsabstractListening in noisy environments can be difficult even for individuals with a normal hearing thresholds. The speech signal can be masked by noise, which may lead to word misperceptions on the side of the listener, and overall difficulty to understand the message. To mitigate hearing difficulties on listeners, a co-operative speaker utilizes voice modulation strategies like Lombard speech to generate noise-robust utterances, and similar solutions have been developed for speech synthesis systems. In this work, we propose an alternate solution of choosing noise-robust lexical paraphrases to represent an intended meaning. Our results show that lexical paraphrases differ in their intelligibility in noise. We evaluate the intelligibility of synonyms in context and find that choosing a lexical unit that is less risky to be misheard than its synonym introduced an average gain in comprehension of 37% at SNR -5 dB and 21% at SNR 0 dB for babble noise. Anupama Chingacham, Vera Demberg, Dietrich Klakow |
Interspeech | 3 |
| 2021 | Boosting of Contextual Information in ASR for Air-Traffic Call-Sign RecognitionabstractContextual adaptation of ASR can be very beneficial for multi-accent and often noisy Air-Traffic Control (ATC) speech. Our focus is call-sign recognition, which can be used to track conversations of ATC operators with individual airplanes. We developed a two-stage boosting strategy, consisting of HCLG boosting and Lattice boosting. Both are implemented as WFST compositions and the contextual information is specific to each utterance. In HCLG boosting we give score discounts to individual words, while in Lattice boosting the score discounts are given to word sequences. The context data have origin in surveillance database of OpenSky Network. From this, we obtain lists of call-signs that are made more likely to appear in the best hypothesis of ASR. This also improves the accuracy of the NLU module that recognizes the call-signs from the best hypothesis of ASR. Martin Kocour, Karel Veselý, Alexander Blatt, Juan Zuluaga-Gomez, Igor Szöke, Jan Cernocký, Dietrich Klakow, Petr Motlícek |
Interspeech | 7 |
| 2021 | Phonetic Distance and Surprisal in Multilingual Priming: Evidence from Slavic
Jacek Kudera, Philip Georgis, Bernd Möbius, Tania Avgustinova, Dietrich Klakow |
Interspeech | 5 |
| 2021 | Integrating Unsupervised Data Generation into Self-Supervised Neural Machine Translation for Low-Resource LanguagesabstractFor most language combinations and parallel data is either scarce or simply unavailable. To address this and unsupervised machine translation (UMT) exploits large amounts of monolingual data by using synthetic data generation techniques such as back-translation and noising and while self-supervised NMT (SSNMT) identifies parallel sentences in smaller comparable data and trains on them. To this date and the inclusion of UMT data generation techniques in SSNMT has not been investigated. We show that including UMT techniques into SSNMT significantly outperforms SSNMT (up to +4.3 BLEU and af2en) as well as statistical (+50.8 BLEU) and hybrid UMT (+51.5 BLEU) baselines on related and distantly-related and unrelated language pairs. Dana Ruiter, Dietrich Klakow, Josef van Genabith, Cristina España-Bonet |
MTSummit (1) | 2 |
| 2021 | A Survey on Recent Approaches for Natural Language Processing in Low-Resource ScenariosabstractMichael A. Hedderich, Lukas Lange, Heike Adel, Jannik Strötgen, Dietrich Klakow. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Michael A. Hedderich, Lukas Lange, Heike Adel, Jannik Strötgen, Dietrich Klakow |
NAACL-HLT | 5 |
| 2021 | Mining Easily Understandable Models from Complex Event Logs
Boris Wiegand, Dietrich Klakow, Jilles Vreeken |
SDM | 2 |
| 2021 | Trends in Integration of Vision and Language Research: A Survey of Tasks, Datasets, and MethodsabstractInterest in Artificial Intelligence (AI) and its applications has seen unprecedented growth in the last few years. This success can be partly attributed to the advancements made in the sub-fields of AI such as machine learning, computer vision, and natural language processing. Much of the growth in these fields has been made possible with deep learning, a sub-area of machine learning that uses artificial neural networks. This has created significant interest in the integration of vision and language. In this survey, we focus on ten prominent tasks that integrate language and vision by discussing their problem formulation, methods, existing datasets, evaluation measures, and compare the results obtained with corresponding state-of-the-art methods. Our efforts go beyond earlier surveys which are either task-specific or concentrate only on one type of visual content, i.e., image or video. Furthermore, we also provide some potential future directions in this field of research with an anticipation that this survey stimulates innovative thoughts and ideas to address the existing challenges and build new applications. Aditya Mogadala, Marimuthu Kalimuthu, Dietrich Klakow |
J. Artif. Intell. Res. | 3 |
| 2021 | Image manipulation with natural language using Two-sided Attentive Conditional Generative Adversarial Network
Aditya Mogadala, Dietrich Klakow |
Neural Networks | 3 |
| 2020 | Neural Data-to-Text Generation via Jointly Learning the Segmentation and CorrespondenceabstractThe neural attention model has achieved great success in data-to-text generation tasks. Though usually excelling at producing fluent text, it suffers from the problem of information missing, repetition and "hallucination". Due to the black-box nature of the neural attention architecture, avoiding these problems in a systematic way is non-trivial. To address this concern, we propose to explicitly segment target text into fragment units and align them with their data correspondences. The segmentation and correspondence are jointly learned as latent variables without any human annotations. We further impose a soft statistical constraint to regularize the segmental granularity. The resulting architecture maintains the same expressive power as neural attention models, while being able to generate fully interpretable outputs with several times less computational cost. On both E2E and WebNLG benchmarks, we show the proposed model consistently outperforms its neural attention counterparts. Xiaoyu Shen 0001, Ernie Chang, Hui Su, Cheng Niu, Dietrich Klakow |
ACL | 5 |
| 2020 | A Closer Look at Linguistic Knowledge in Masked Language Models: The Case of Relative Clauses in American EnglishabstractTransformer-based language models achieve high performance on various tasks, but we still lack understanding of the kind of linguistic knowledge they learn and rely on.We evaluate three models (BERT, RoBERTa, and ALBERT), testing their grammatical and semantic knowledge by sentence-level probing, diagnostic cases, and masked prediction tasks.We focus on relative clauses (in American English) as a complex phenomenon needing contextual information and antecedent identification to be resolved.Based on a naturalistic dataset, probing shows that all three models indeed capture linguistic knowledge about grammaticality, achieving high performance.Evaluation on diagnostic cases and masked prediction tasks considering fine-grained linguistic knowledge, however, shows pronounced model-specific weaknesses especially on semantic knowledge, strongly impacting models' performance.Our results highlight the importance of (a) model comparison in evaluation task and (b) building up claims of model performance and the linguistic knowledge they capture beyond purely probing-based evaluations. Marius Mosbach, Stefania Degaetano-Ortlieb, Marie-Pauline Krielke, Badr Abdullah, Dietrich Klakow |
COLING | 5 |
| 2020 | Transfer Learning and Distant Supervision for Multilingual Transformer Models: A Study on African LanguagesabstractMultilingual transformer models like mBERT and XLM-RoBERTa have obtained great improvements for many NLP tasks on a variety of languages.However, recent works also showed that results from high-resource languages could not be easily transferred to realistic, low-resource scenarios.In this work, we study trends in performance for different amounts of available resources for the three African languages Hausa, isiXhosa and Yorùbá on both NER and topic classification.We show that in combination with transfer learning or distant supervision, these models can achieve with as little as 10 or 100 labeled sentences the same performance as baselines with much more supervised training data.However, we also find settings where this does not hold.Our discussions and additional experiments on assumptions such as time and hardware restrictions highlight challenges and opportunities in low-resource learning. Michael A. Hedderich, David Ifeoluwa Adelani, Jesujoba O. Alabi, Udia Markus, Dietrich Klakow |
EMNLP (1) | 6 |
| 2020 | On the Security Relevance of Initial Weights in Deep Neural Networks
Kathrin Grosse, Thomas Alexander Trost, Marius Mosbach, Michael Backes 0001, Dietrich Klakow |
ICANN (1) | 5 |
| 2020 | Cross-Domain Adaptation of Spoken Language Identification for Related Languages: The Curious Case of Slavic LanguagesabstractState-of-the-art spoken language identification (LID) systems, which are based on end-to-end deep neural networks, have shown remarkable success not only in discriminating between distant languages but also between closely-related languages or even different spoken varieties of the same language. However, it is still unclear to what extent neural LID models generalize to speech samples with different acoustic conditions due to domain shift. In this paper, we present a set of experiments to investigate the impact of domain mismatch on the performance of neural LID systems for a subset of six Slavic languages across two domains (read speech and radio broadcast) and examine two low-level signal descriptors (spectral and cepstral features) for this task. Our experiments show that (1) out-of-domain speech samples severely hinder the performance of neural LID models, and (2) while both spectral and cepstral features show comparable performance within-domain, spectral features show more robustness under domain mismatch. Moreover, we apply unsupervised domain adaptation to minimize the discrepancy between the two domains in our study. We achieve relative accuracy improvements that range from 9% to 77% depending on the diversity of acoustic conditions in the source domain. Badr Abdullah, Tania Avgustinova, Bernd Möbius, Dietrich Klakow |
INTERSPEECH | 4 |
| 2020 | Privacy Guarantees for De-Identifying Text TransformationsabstractMachine Learning approaches to Natural Language Processing tasks benefit from a comprehensive collection of real-life user data. At the same time, there is a clear need for protecting the privacy of the users whose data is collected and processed. For text collections, such as, e.g., transcripts of voice interactions or patient records, replacing sensitive parts with benign alternatives can provide de-identification. However, how much privacy is actually guaranteed by such text transformations, and are the resulting texts still useful for machine learning? In this paper, we derive formal privacy guarantees for general text transformation-based de-identification methods on the basis of Differential Privacy. We also measure the effect that different ways of masking private information in dialog transcripts have on a subsequent machine learning task. To this end, we formulate different masking strategies and compare their privacy-utility trade-offs. In particular, we compare a simple redact approach with more sophisticated word-by-word replacement using deep learning models on multiple natural language understanding tasks like named entity recognition, intent detection, and dialog act classification. We find that only word-by-word replacement is robust against performance drops in various tasks. David Ifeoluwa Adelani, Ali Davody, Thomas Kleinbauer, Dietrich Klakow |
INTERSPEECH | 4 |
| 2020 | ATC-ANNO: Semantic Annotation for Air Traffic Control with Assistive Auto-AnnotationabstractIn air traffic control, assistant systems support air traffic controllers in their work. To improve the reactivity and accuracy of the assistant, automatic speech recognition can monitor the commands uttered by the controller. However, to provide sufficient training data for the speech recognition system, many hours of air traffic communications have to be transcribed and semantically annotated. For this purpose we developed the annotation tool ATC-ANNO. It provides a number of features to support the annotator in their task, such as auto-complete suggestions for semantic tags, access to preliminary speech recognition predictions, syntax highlighting and consistency indicators. Its core assistive feature, however, is its ability to automatically generate semantic annotations. Although it is based on a simple hand-written finite state grammar, it is also able to annotate sentences that deviate from this grammar. We evaluate the impact of different features on annotator efficiency and find that automatic annotation allows annotators to cover four times as many utterances in the same time. Marc Schulder, Johannah O'Mahony, Yury Bakanouski, Dietrich Klakow |
LREC | 4 |
| 2019 | Feature-Dependent Confusion Matrices for Low-Resource NER Labeling with Noisy LabelsabstractLukas Lange, Michael A. Hedderich, Dietrich Klakow. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Lukas Lange, Michael A. Hedderich, Dietrich Klakow |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Select and Attend: Towards Controllable Content Selection in Text GenerationabstractXiaoyu Shen, Jun Suzuki, Kentaro Inui, Hui Su, Dietrich Klakow, Satoshi Sekine. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xiaoyu Shen 0001, Jun Suzuki 0001, Kentaro Inui, Hui Su, Dietrich Klakow, Satoshi Sekine |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Improving Latent Alignment in Text Summarization by Generalizing the Pointer GeneratorabstractXiaoyu Shen, Yang Zhao, Hui Su, Dietrich Klakow. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xiaoyu Shen 0001, Hui Su, Dietrich Klakow |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Modelling Shared Decision Making in Medical Negotiations: Interactive Training with Cognitive Agents
Volha Petukhova, Firuza Sharifullaeva, Dietrich Klakow |
PRIMA | 3 |
| 2019 | A Formal Proof of the Expressiveness of Deep LearningabstractDeep learning has had a profound impact on computer science in recent years, with applications to image recognition, language processing, bioinformatics, and more. Recently, Cohen et al. provided theoretical evidence for the superiority of deep learning over shallow learning. We formalized their mathematical proof using Isabelle/HOL. The Isabelle development simplifies and generalizes the original proof, while working around the limitations of the HOL type system. To support the formalization, we developed reusable libraries of formalized mathematics, including results about the matrix rank, the Borel measure, and multivariate polynomials as well as a library for tensor analysis. Alexander Bentkamp, Jasmin Blanchette, Dietrich Klakow |
J. Autom. Reason. | 3 |
| 2018 | Long-Span Language Models for Query-Focused Unsupervised Extractive Text Summarization
Mittul Singh, Arunav Mishra, Youssef Oualil, Klaus Berberich, Dietrich Klakow |
ECIR | 5 |
| 2018 | Nexus Network: Connecting the Preceding and the Following in Dialogue GenerationabstractSequence-to-Sequence (seq2seq) models have become overwhelmingly popular in building end-to-end trainable dialogue systems.Though highly efficient in learning the backbone of human-computer communications, they suffer from the problem of strongly favoring short generic responses.In this paper, we argue that a good response should smoothly connect both the preceding dialogue history and the following conversations.We strengthen this connection through mutual information maximization.To sidestep the nondifferentiability of discrete natural language tokens, we introduce an auxiliary continuous code space and map such code space to a learnable prior distribution for generation purpose.Experiments on two dialogue datasets validate the effectiveness of our model, where the generated responses are closely related to the dialogue context and lead to more interactive conversations.* Indicates equal contribution.X. Shen focuses on algorithm and H. Su is responsible for experiments. Xiaoyu Shen 0001, Hui Su, Wenjie Li 0002, Dietrich Klakow |
EMNLP | 4 |
| 2018 | Toward Bayesian Synchronous Tree Substitution Grammars for Sentence PlanningabstractDeveloping conventional natural language generation systems requires extensive attention from human experts in order to craft complex sets of sentence planning rules.We propose a Bayesian nonparametric approach to learn sentence planning rules by inducing synchronous tree substitution grammars for pairs of text plans and morphosyntactically-specified dependency trees.Our system is able to learn rules which can be used to generate novel texts after training on small datasets. David M. Howcroft, Dietrich Klakow, Vera Demberg |
INLG | 2 |
| 2018 | The Metalogue Debate Trainee Corpus: Data Collection and Annotations
Volha Petukhova, Andrei Malchanau, Youssef Oualil, Dietrich Klakow, Saturnino Luz, Fasih Haider, Nick Campbell 0001, Dimitris Koryzis, Dimitris Spiliotopoulos, Pierre Albert, Nicklas Linz, Jan Alexandersson |
LREC | 4 |
| 2017 | A context-aware speech recognition and understanding system for air traffic control domainabstractAutomatic Speech Recognition and Understanding (ASRU) systems can generally use temporal and situational context information to improve their performance for a given task. This is typically done by rescoring the ASR hypotheses or by dynamically adapting the ASR models. For some domains, such as Air Traffic Control (ATC), this context information can be, however, small in size, partial and available only as abstract concepts (e.g. airline codes), which are difficult to map into full possible spoken sentences to perform rescoring or adaptation. This paper presents a multi-modal ASRU system, which dynamically integrates partial temporal and situational ATC context information to improve its performance. This is done either by 1) extracting word sequences which carry relevant ATC information from ASR N-best Lists and then perform a context-based rescoring on the extracted ATC segments or 2) by a partial adaptation of the language model. Experiments conducted on 4 hours of test data from Prague and Vienna approach (arrivals) showed a relative reduction of the ATC command error rate metric by 30% to 50%. Youssef Oualil, Dietrich Klakow, György Szaszák, Ajay Srinivasamurthy, Hartmut Helmke, Petr Motlícek |
ASRU | 2 |
| 2017 | A Neural Network approach for mixing language modelsabstractThe performance of Neural Network (NN)-based language models is steadily improving due to the emergence of new architectures, which are able to learn different natural language characteristics. This paper presents a novel framework, which shows that a significant improvement can be achieved by combining different existing heterogeneous models in a single architecture. This is done through 1) a feature layer, which separately learns different NN-based models and 2) a mixture layer, which merges the resulting model features. In doing so, this architecture benefits from the learning capabilities of each model with no noticeable increase in the number of model parameters or the training time. Extensive experiments conducted on the Penn Treebank (PTB) and the Large Text Compression Benchmark (LTCB) corpus showed a significant reduction of the perplexity when compared to state-of-the-art feedforward as well as recurrent neural network architectures. Youssef Oualil, Dietrich Klakow |
ICASSP | 2 |
| 2017 | Wake-Sleep Variational Autoencoders for Language Modeling
Xiaoyu Shen 0001, Hui Su, Shuzi Niu, Dietrich Klakow |
ICONIP (1) | 4 |
| 2017 | Incremental Dialogue Act Recognition: Token- vs Chunk-Based Classification
Eustace Ebhotemhen, Volha Petukhova, Dietrich Klakow |
INTERSPEECH | 3 |
| 2017 | The Extended SPaRKy Restaurant Corpus: Designing a Corpus with Variable Information Density
David M. Howcroft, Dietrich Klakow, Vera Demberg |
INTERSPEECH | 2 |
| 2017 | A Batch Noise Contrastive Estimation Approach for Training Large Vocabulary Language ModelsabstractTraining large vocabulary Neural Network Language Models (NNLMs) is a difficult task due to the explicit requirement of the output layer normalization, which typically involves the evaluation of the full softmax function over the complete vocabulary. This paper proposes a Batch Noise Contrastive Estimation (B-NCE) approach to alleviate this problem. This is achieved by reducing the vocabulary, at each time step, to the target words in the batch and then replacing the softmax by the noise contrastive estimation approach, where these words play the role of targets and noise samples at the same time. In doing so, the proposed approach can be fully formulated and implemented using optimal dense matrix operations. Applying B-NCE to train different NNLMs on the Large Text Compression Benchmark (LTCB) and the One Billion Word Benchmark (OBWB) shows a significant reduction of the training time with no noticeable degradation of the models performance. This paper also presents a new baseline comparative study of different standard NNLMs on the large OBWB on a single Titan-X GPU. Youssef Oualil, Dietrich Klakow |
INTERSPEECH | 2 |
| 2017 | Estimation of Gap Between Current Language Models and Human Performance
Xiaoyu Shen 0001, Youssef Oualil, Clayton Greenberg, Mittul Singh, Dietrich Klakow |
INTERSPEECH | 5 |
| 2017 | Approximated and Domain-Adapted LSTM Language Models for First-Pass Decoding in Speech Recognition
Mittul Singh, Youssef Oualil, Dietrich Klakow |
INTERSPEECH | 3 |
| 2017 | A Formal Proof of the Expressiveness of Deep Learning
Alexander Bentkamp, Jasmin Blanchette, Dietrich Klakow |
ITP | 3 |
| 2016 | Sub-Word Similarity based Search for Embeddings: Inducing Rare-Word Embeddings for Word Similarity Tasks and Language ModellingabstractTraining good word embeddings requires large amounts of data. Out-of-vocabulary words will still be encountered at test-time, leaving these words without embeddings. To overcome this lack of embeddings for rare words, existing methods leverage morphological features to generate embeddings. While the existing methods use computationally-intensive rule-based (Soricut and Och, 2015) or tool-based (Botha and Blunsom, 2014) morphological analysis to generate embeddings, our system applies a computationally-simpler sub-word search on words that have existing embeddings. Embeddings of the sub-word search results are then combined using string similarity functions to generate rare word embeddings. We augmented pre-trained word embeddings with these novel embeddings and evaluated on a rare word similarity task, obtaining up to 3 times improvement in correlation over the original set of embeddings. Applying our technique to embeddings trained on larger datasets led to on-par performance with the existing state-of-the-art for this task. Additionally, while analysing augmented embeddings in a log-bilinear language model, we observed up to 50% reduction in rare word perplexity in comparison to other more complex language models. Mittul Singh, Clayton Greenberg, Youssef Oualil, Dietrich Klakow |
COLING | 4 |
| 2016 | Long-Short Range Context Neural Networks for Language ModelingabstractThe goal of language modeling techniques is to capture the statistical and structural properties of natural languages from training corpora.This task typically involves the learning of short range dependencies, which generally model the syntactic properties of a language and/or long range dependencies, which are semantic in nature.We propose in this paper a new multi-span architecture, which separately models the short and long context information while it dynamically merges them to perform the language modeling task.This is done through a novel recurrent Long-Short Range Context (LSRC) network, which explicitly models the local (short) and global (long) context using two separate hidden states that evolve in time.This new architecture is an adaptation of the Long-Short Term Memory network (LSTM) to take into account the linguistic properties.Extensive experiments conducted on the Penn Treebank (PTB) and the Large Text Compression Benchmark (LTCB) corpus showed a significant reduction of the perplexity when compared to state-of-the-art language modeling techniques. Youssef Oualil, Mittul Singh, Clayton Greenberg, Dietrich Klakow |
EMNLP | 4 |
| 2016 | Event participant modelling with neural networksabstractA common problem in cognitive modelling is lack of access to accurate broad-coverage models of event-level surprisal.As shown in, e.g., Bicknell et al. (2010), event-level knowledge does affect human expectations for verbal arguments.For example, the model should be able to predict that mechanics are likely to check tires, while journalists are more likely to check typos.Similarly, we would like to predict what locations are likely for playing football or playing flute in order to estimate the surprisal of actually-encountered locations.Furthermore, such a model can be used to provide a probability distribution over fillers for a thematic role which is not mentioned in the text at all.To this end, we train two neural network models (an incremental one and a non-incremental one) on large amounts of automatically rolelabelled text.Our models are probabilistic and can handle several roles at once, which also enables them to learn interactions between different role fillers.Evaluation shows a drastic improvement over current state-of-the-art systems on modelling human thematic fit judgements, and we demonstrate via a sentence similarity task that the system learns highly useful embeddings. Ottokar Tilk, Vera Demberg, Asad B. Sayeed, Dietrich Klakow, Stefan Thater |
EMNLP | 4 |
| 2016 | Sequential Recurrent Neural Networks for Language ModelingabstractFeedforward Neural Network (FNN)-based language models estimate the probability of the next word based on the history of the last N words, whereas Recurrent Neural Networks (RNN) perform the same task based only on the last word and some context information that cycles in the network. This paper presents a novel approach, which bridges the gap between these two categories of networks. In particular, we propose an architecture which takes advantage of the explicit, sequential enumeration of the word history in FNN structure while enhancing each word representation at the projection layer through recurrent context information that evolves in the network. The context integration is performed using an additional word-dependent weight matrix that is also learned during the training. Extensive experiments conducted on the Penn Treebank (PTB) and the Large Text Compression Benchmark (LTCB) corpus showed a significant reduction of the perplexity when compared to state-of-the-art feedforward as well as recurrent neural network architectures. Youssef Oualil, Clayton Greenberg, Mittul Singh, Dietrich Klakow |
INTERSPEECH | 4 |
| 2016 | Creating Annotated Dialogue Resources: Cross-domain Dialogue Act Classification
Dilafruz Amanova, Volha Petukhova, Dietrich Klakow |
LREC | 3 |
| 2016 | Orthographic and Morphological Correspondences between Related Slavic Languages as a Base for Modeling of Mutual Intelligibility
Andrea K. Fischer, Klara Jagrova, Irina Stenger, Tania Avgustinova, Dietrich Klakow, Roland Marti |
LREC | 5 |
| 2015 | Real-time integration of dynamic context information for improving automatic speech recognitionabstractThe use of prior situational/contextual knowledge about a given \ntask can significantly improve automatic speech recognition \n(ASR) performance. This is typically done through adaptation \nof acoustic or language models if data is available or using \nknowledge-based rescoring. The main adaptation techniques, \nhowever, are either domain-specific, which makes them inadequate \nfor other tasks, or static and offline, and therefore cannot \ndeal with dynamic knowledge. To circumvent this problem, \nwe propose a real-time system which dynamically integrates \nsituational context into ASR. The context integration is done \neither post-recognition, in which case a weighted Levenshtein \ndistance between the ASR hypotheses and the context information \nbased on the ASR confidence scores is proposed to extract \nthe most likely sequence of spoken words, or pre-recognition, \nwhere the search space is adjusted to the new situational knowledge \nthrough adaptation of the finite state machine modeling \nthe spoken language. Experiments conducted on 3 hours of \nAir Traffic Control (ATC) data achieved a 51% reduction of \nthe Command Error Rate (CmdER) which is used as evaluation \nmetric in the ATC domain. Youssef Oualil, Marc Schulder, Hartmut Helmke, Anna Schmidt, Dietrich Klakow |
INTERSPEECH | 5 |
| 2015 | Combining Pattern-Based and Distributional Similarity for Graph-Based Noun Categorization
Michael Wiegand, Benjamin Roth 0001, Dietrich Klakow |
NLDB | 3 |
| 2015 | Bridging the vocabulary gap between questions and answer sentences
Saeedeh Momtazi, Dietrich Klakow |
Inf. Process. Manag. | 2 |
| 2014 | Separating Brands from Types: an Investigation of Different Features for the Food Domain
Michael Wiegand, Dietrich Klakow |
COLING | 2 |
| 2014 | Unsupervised Parsing for Generating Surface-Based Relation Extraction PatternsabstractFinding the right features and patterns for identifying relations in natural language is one of the most pressing research questions for relation extraction.In this paper, we compare patterns based on supervised and unsupervised syntactic parsing and present a simple method for extracting surface patterns from a parsed training set.Results show that the use of surfacebased patterns not only increases extraction speed, but also improves the quality of the extracted relations.We find that, in this setting, unsupervised parsing, besides requiring less resources, compares favorably in terms of extraction quality. Jens Illig, Benjamin Roth 0001, Dietrich Klakow |
EACL | 3 |
| 2014 | RelationFactory: A Fast, Modular and Effective System for Knowledge Base PopulationabstractBenjamin Roth, Tassilo Barth, Grzegorz Chrupała, Martin Gropp, Dietrich Klakow. Proceedings of the Demonstrations at the 14th Conference of the European Chapter of the Association for Computational Linguistics. 2014. Benjamin Roth 0001, Tassilo Barth, Grzegorz Chrupala, Martin Gropp, Dietrich Klakow |
EACL | 5 |
| 2014 | Automatic Food Categorization from Large Unlabeled Corpora and Its Impact on Relation ExtractionabstractWe present a weakly-supervised induction method to assign semantic information to food items.We consider two tasks of categorizations being food-type classification and the distinction of whether a food item is composite or not.The categorizations are induced by a graph-based algorithm applied on a large unlabeled domain-specific corpus.We show that the usage of a domain-specific corpus is vital.We do not only outperform a manually designed open-domain ontology but also prove the usefulness of these categorizations in relation extraction, outperforming state-of-the-art features that include syntactic information and Brown clustering. Michael Wiegand, Benjamin Roth 0001, Dietrich Klakow |
EACL | 3 |
| 2014 | Accurate client-server based speech recognition keeping personal data on the clientabstractIn this paper, a novel technique is proposed that recognizes speech on a server but all private knowledge is processed on the client. Private knowledge could be address book entries, calendar entries or medical patient data. The technique combines the advantage of a powerful server with almost unlimited memory and the advantage using locally available user dependent knowledge. A dynamic language model is used to recognize speech with the help of content dependent acoustic fillers on a server. The result is then recognized including user dependent knowledge on a client, e.g., a smart phone. We achieved a word error rate reduction of 17% on the Wall Street Journal Corpus. Munir Georges, Stephan Kanthak, Dietrich Klakow |
ICASSP | 3 |
| 2014 | Multiple concurrent speaker short-term tracking using a Kalman filter bankabstractThis paper presents a novel filtering approach for tracking multiple concurrent speakers with a microphone array. In this framework, a Kalman filter bank that evolves in time according to a temporal Hidden Markov Model (HMM) is proposed. This approach was designed to overcome two major problems that occur in spontaneous speech; namely, 1) the speaker overlap. This problem is solved using a bank of parallel Kalman filters that track multiple simultaneous speakers, and 2) the high discontinuity of spontaneous speech caused by short breaks and silences. This is solved using an HMM that allows speakers to change their state (speaking, silent, etc.) over time. The actual active speakers number and locations are extracted from the active filters using a second Kalman filter. Experiments on the AV16.3 showed an average tracking rate improvement of 8% compared to a short-term clustering approach, while being 7 times faster. Youssef Oualil, Dietrich Klakow |
ICASSP | 2 |
| 2014 | Metalogue: A Multiperspective Multimodal Dialogue System with Metacognitive Abilities for Highly Adaptive and Flexible Dialogue ManagementabstractThis poster paper presents a high-level description of the Metalogue project that is developing a multi-modal dialogue system that is able to implement interactive behaviors that seem natural to users and is flexible enough to exploit the full potential of multimodal interaction. We provide an outline of the initial work undertaken to define a an open architecture for the integrated Metalogue system. This system includes components that are necessary for the implementation of the processing stages for a variety of application domains: initialization, training, information gathering, orchestration, multimodality, dialogue management, speech recognition, speech synthesis and user modelling. Jan Alexandersson, Maria Aretoulaki, Nick Campbell 0001, Michael Gardner, Andrey Girenko, Dietrich Klakow, Dimitris Koryzis, Volha Petukhova, Marcus Specht, Dimitris Spiliotopoulos, Alexander Stricker, Niels Taatgen |
Intelligent Environments | 6 |
| 2014 | The DBOX Corpus Collection of Spoken Human-Human and Human-Machine Dialogues
Volha Petukhova, Martin Gropp, Dietrich Klakow, Gregor Eigner, Mario Topf, Stefan Srb, Petr Motlícek, Blaise Potard, John Dines, Olivier Deroo, Ronny Egeler, Uwe Meinz, Steffen Liersch, Anna Schmidt |
LREC | 3 |
| 2014 | Context-based recognition network adaptation for improving on-line ASR in Air Traffic ControlabstractThis paper presents an approach for incorporating situational context information into an on-line Automatic Speech Recognition (ASR) component of an Air Traffic Control (ATC) assistance system to improve recognition performance. Context information is treated as prior information to reduce the search space for recognition. It is integrated in the ASR pipeline by continually updating the recognition network. This is achieved by automatically adapting the underlying grammar whenever new situational knowledge becomes available. The context-dependent recognition network is then re-created and substituted for recognition based on these context-dependent grammars. As a result, the recognizer's search space is constantly being limited to that subset of hypotheses that are deemed plausible in the current situation. Since recognition and adaptation tasks can be easily performed by two separate parallel processes, on-line capabilities of the system are maintained, and response times do not increase as a result of context integration. Experiments conducted on about two hours of ATC data show a reduction in command error rate by a factor of three when context is used. Anna Schmidt, Youssef Oualil, Oliver Ohneiser, Matthias Kleinert, Marc Schulder, Arif Khan 0004, Hartmut Helmke, Dietrich Klakow |
SLT | 8 |
| 2014 | A keyword search system using open source softwareabstractProvides an overview of a speech-to-text (STT) and keyword search (KWS) system architecture build primarily on the top of the Kaldi toolkit and expands on a few highlights. The system was developed as a part of the research efforts of the Radical team while participating in the IARPA Babel program. Our aim was to develop a general system pipeline which could be easily and rapidly deployed in any language, independently on the language script and phonological and linguistic features of the language. Jan Trmal, Guoguo Chen, Daniel Povey, Sanjeev Khudanpur, Pegah Ghahremani, Xiaohui Zhang 0007, Vimal Manohar, Chunxi Liu, Aren Jansen, Dietrich Klakow, David Yarowsky, Florian Metze |
SLT | 10 |
| 2013 | Feature-based models for improving the quality of noisy training data for relation extractionabstractSupervised relation extraction from text relies on annotated data. Distant supervision is a scheme to obtain noisy training data by using a knowledge base of relational tuples as the ground truth and finding entity pair matches in a text corpus. We propose and evaluate two feature-based models for increasing the quality of distant supervision extraction patterns. Benjamin Roth 0001, Dietrich Klakow |
CIKM | 2 |
| 2013 | Combining Generative and Discriminative Model Scores for Distant SupervisionabstractDistant supervision is a scheme to generate noisy training data for relation extraction by aligning entities of a knowledge base with text.In this work we combine the output of a discriminative at-least-one learner with that of a generative hierarchical topic model to reduce the noise in distant supervision data.The combination significantly increases the ranking quality of extracted facts and achieves state-of-the-art extraction performance in an end-to-end setting.A simple linear interpolation of the model scores performs better than a parameter-free scheme based on nondominated sorting. Benjamin Roth 0001, Dietrich Klakow |
EMNLP | 2 |
| 2013 | A probabilistic framework for multiple speaker localizationabstractThis paper presents a novel probabilistic framework for localizing multiple speakers with a microphone array. In this framework, the generalized cross correlation function (GCC) of each microphone pair is interpreted as a probability distribution of the time difference of arrival (TDOA) and subsequently approximated as a Gaussian mixture. The distribution parameters are estimated with a weighted expectation maximization algorithm. Then, the joint distribution of the TDOA Gaussian mixtures is mapped to a multimodal distribution in the location space, where each mode represents a potential source location. The approach taken here performs the localization by 1) reducing the search space to some regions that are likely to contain a source and then 2) extracting the actual speaker locations with a numerical optimization algorithm. The effectiveness of the proposed approach is shown using the AV16.3 corpus. Youssef Oualil, Mathew Magimai-Doss, Friedrich Faubel, Dietrich Klakow |
ICASSP | 4 |
| 2013 | Comparing RNNs and log-linear interpolation of improved skip-model on four Babel languages: Cantonese, Pashto, Tagalog, TurkishabstractRecurrent neural networks (RNNs) are a very recent technique to model long range dependencies in natural languages. They have clearly outperformed trigrams and other more advanced language modeling techniques by using non-linearly modeling long range dependencies. An alternative is to use log-linear interpolation of skip models (i.e. skip bigrams and skip trigrams). The method as such has been published earlier. In this paper we investigate the impact of different smoothing techniques on the skip models as a measure of their overall performance. One option is to use automatically trained distance clusters (both hard and soft) to increase robustness and to combat sparseness in the skip model. We also investigate alternative smoothing techniques on word level. For skip bigrams when skipping a small number of words Kneser-Ney smoothing (KN) is advantageous. For a larger number of words being skipped Dirichlet smoothing performs better. In order to exploit the advantages of both KN and Dirichlet smoothing we propose a new unified smoothing technique. Experiments are performed on four Babel languages: Cantonese, Pashto, Tagalog and Turkish. RNNs and log-linearly interpolated skip models are on par if the skip models are trained with standard smoothing techniques. Using the improved smoothing of the skip models along with distance clusters, we can clearly outperform RNNs by about 8-11 % in perplexity across all four languages. Mittul Singh, Dietrich Klakow |
ICASSP | 2 |
| 2013 | Towards Contextual Healthiness Classification of Food Items - A Linguistic Approach
Michael Wiegand, Dietrich Klakow |
IJCNLP | 2 |
| 2013 | Transducer-based speech recognition with dynamic language models
Munir Georges, Stephan Kanthak, Dietrich Klakow |
INTERSPEECH | 3 |
| 2013 | An unsupervised Bayesian classifier for multiple speaker detection and localizationabstractMultiple speaker localization algorithms generally require a bi-nary detector, which performs the source/noise classification of the location estimates. This is mainly due to the unknown time-varying number of sources, and to the presence of noise and reverberation. In this paper, we propose an unsupervised learn-ing approach based on a naive Bayesian classifier. The proposed approach couples two speaker location features, namely, 1) the steered response power introduced at the location estimate, and 2) the corresponding maximum likelihood error, which charac-terizes the variance of the estimate. The latter is experimentally shown to be highly correlated with the steered power at the loca-tion estimate. The proposed method is further extended to con-trol the misclassification rate through the use of a loss function. This approach is general, and can be easily extended to integrate more speaker/speech features. Experiments on the AV16.3 cor-pus show the effectiveness of the proposed approach. Index Terms: microphone arrays, multiple speaker localiza-tion, source detection, Bayesian classification. Youssef Oualil, Friedrich Faubel, Dietrich Klakow |
INTERSPEECH | 3 |
| 2013 | Predicative Adjectives: An Unsupervised Criterion to Extract Subjective Adjectives
Michael Wiegand, Josef Ruppenhofer, Dietrich Klakow |
HLT-NAACL | 3 |
| 2012 | Generalization Methods for In-Domain and Cross-Domain Opinion Holder Extraction
Michael Wiegand, Dietrich Klakow |
EACL | 2 |
| 2012 | Knowledge-Based Word Lattice Rescoring in a Dynamic ContextabstractRecent advances in automatic speech recognition (ASR) technology continue to be based heavily on data-driven methods, meaning that the full benefits of such research are often not enjoyed in domains for which there is little training data. Moreover, tractability is often an issue with these methods when conditioning for long-distance dependencies, entailing that many higher-level knowledge sources such as situational knowledge cannot be easily utilized in classification. This paper describes an effort to circumvent this problem by using dynamic contextual knowledge to rescore ASR lattice output using a dynamic weighted constraint satisfaction function. With this method, it was possible to achieve a roughly 80 % reduction in WER for ASR in the context of an air traffic control scenario. Index Terms: lattice rescoring, knowledge-based, contextsensitivity 1. Todd Shore, Friedrich Faubel, Hartmut Helmke, Dietrich Klakow |
INTERSPEECH | 4 |
| 2012 | Mobile texting: can post-ASR correction solve the issues? an experimental study on gain vs. costsabstractThe next big step in embedded, mobile speech recognition will be to allow completely free input as it is needed for messaging like SMS or email. However, unconstrained dictation remains error-prone, especially when the environment is noisy. In this paper, we compare different methods for improving a given free-text dictation system used to enter textbased messages in embedded mobile scenarios, where distraction, interaction cost, and hardware limitations enforce strict constraints over traditional scenarios. We present a corpus-based evaluation, measuring the trade-off between improvement of the word error rate versus the interaction steps that are required under various parameters. Results show that by post-processing the output of a "black box" speech recognizer (e.g. a web-based speech recognition service), a reduction of word error rate by 55% (10.3% abs.) can be obtained. For further error reduction, however, a richer representation of the original hypotheses (e.g. lattice) is necessary. Michael Feld, Saeedeh Momtazi, Farina Freigang, Dietrich Klakow, Christian Müller 0014 |
IUI | 4 |
| 2012 | Task-Driven Linguistic Analysis based on an Underspecified Features Representation
Stasinos Konstantopoulos, Valia Kordoni, Nicola Cancedda, Vangelis Karkaletsis, Dietrich Klakow, Jean-Michel Renders |
LREC | 5 |
| 2012 | A Gold Standard for Relation Extraction in the Food Domain
Michael Wiegand, Benjamin Roth 0001, Eva Lasarcyk, Stephanie Köser, Dietrich Klakow |
LREC | 5 |
| 2012 | Web-Based Relation Extraction for the Food Domain
Michael Wiegand, Benjamin Roth 0001, Dietrich Klakow |
NLDB | 3 |
| 2011 | Trained trigger language model for sentence retrieval in QA: bridging the vocabulary gapabstractWe propose a novel language model for sentence retrieval in Question Answering (QA) systems called trained trigger language model. This model addresses the word mismatch problem in information retrieval. The proposed model captures pairs of trigger and target words while training on a large corpus. The word pairs are extracted based on both unsupervised and supervised approaches while different notions of triggering are used. In addition, we study the impact of corpus size and domain for a supervised model. All notions of the trained trigger model are finally used in a language model-based sentence retrieval framework. Our experiments on TREC QA collection verify that the proposed model significantly improves the sentence retrieval performance compared to the state-of-the-art translation model and class model which address the same problem. Saeedeh Momtazi, Dietrich Klakow |
CIKM | 2 |
| 2011 | A Model-Based Spectral Envelope Wiener Filter for Perceptually Motivated Speech EnhancementabstractIn this work, we present a model-based Wiener filter whose frequency response is optimized in the dimensionally reduced log-Mel domain. That is achieved by making use of a reasonably novel speech feature enhancement approach that has originally been developed in the area of speech recognition. Its combination with Wiener filtering is motivated by the fact that signal reconstruction from log-Mel features sounds very unnatural. Hence, we correct only the spectral envelope and preserve the fine spectral structure of the noisy signal. Experiments on a Wall Street Journal corpus showed a relative improvement of up to 24 % relative in PESQ and 45 % relative in log spectral distance (LSD), compared to Ephraim and Mallah’s log spectral amplitude estimator. Najib Hadir, Friedrich Faubel, Dietrich Klakow |
INTERSPEECH | 3 |
| 2010 | An adaptive level of detail approach to nonlinear estimationabstractIn this work, we present a general method for approximating non-linear transformations of Gaussian mixture random variables. It is based on transforming the individual Gaussians with the unscented transform. The level of detail is adapted by iteratively splitting those components of the initial mixture that exhibited a high degree of nonlinearity during transformation. After each splitting operation, the affected components are re-transformed. This procedure gives more accurate results in cases where a Gaussian fit does not well represent the true distribution. Hence, it is of interest in a number of signal processing fields, ranging from nonlinear adaptive filtering to speech feature enhancement. In simulations, the proposed approach achieved a 48-fold reduction of the approximation error, compared to a single unscented transform. Friedrich Faubel, Dietrich Klakow |
ICASSP | 2 |
| 2010 | On expectation maximization based channel and noise estimation beyond the vector Taylor series expansionabstractIn this work, we show how expectation maximization based simultaneous channel and noise estimation can be derived without a vector Taylor series expansion. The central idea is to approximate the distribution of all the random variables involved - that is noisy speech, clean speech, channel and noise - as one large, joint Gaussian distribution. Consequently, instantaneous estimates of the noise and channel distribution parameters can be obtained by conditioning the joint distribution on observed, noisy speech spectra. This approach allows for the combination of expectation maximization based channel and noise estimation with the unscented transform. Friedrich Faubel, John W. McDonough, Dietrich Klakow |
ICASSP | 3 |
| 2010 | Estimating noise from noisy speech features with a monte carlo variant of the expectation maximization algorithmabstractIn this work, we derive a Monte Carlo expectation maximization algorithm for estimating noise from a noisy utterance. In contrast to earlier approaches, where the distribution of noise was estimated based on a vector Taylor series expansion, we use a combination of importance sampling and Parzen-window density estimation to numerically approximate the occurring integrals with the Monte Carlo method. Experimental results show that the proposed algorithm has superior convergence properties, compared to previous implementations of the EM algorithm. Its application to speech feature enhancement reduced the word error rate by over 30% on a phone number recognition task recorded in a (real) noisy car environment. Index Terms: robust speech recognition, noise estimation, Monte Carlo Friedrich Faubel, Dietrich Klakow |
INTERSPEECH | 2 |
| 2010 | Within and across sentence boundary language modelabstractIn this paper, we propose two different language modeling approaches, namely skip trigram and across sentence boundary, to capture the long range dependencies. The skip trigram model is able to cover more predecessor words of the present word compared to the normal trigram while the same memory space is required. The across sentence boundary model uses the word distribution of the previous sentences to calculate the unigram probability which is applied as the emission probability in the word and the class model frameworks. Our experiments on the Penn Treebank [1] show that each of our proposed models and also their combination significantly outperform the baseline for both the word and the class models and their linear interpolation. The linear interpolation of the word and the class models with the proposed skip trigram and across sentence boundary models achieves 118.4 perplexity while the best state-of-the-art language model has a perplexity of 137.2 on the same dataset. 1. Saeedeh Momtazi, Friedrich Faubel, Dietrich Klakow |
INTERSPEECH | 3 |
| 2010 | A Named Entity Labeler for German: Exploiting Wikipedia and Distributional Clusters
Grzegorz Chrupala, Dietrich Klakow |
LREC | 2 |
| 2010 | Predictive Features for Detecting Indefinite Polar Sentences
Michael Wiegand, Dietrich Klakow |
LREC | 2 |
| 2010 | Paragraph Acquisition and Selection for List Question Using Amazon's Mechanical Turk
Dietrich Klakow |
LREC | 2 |
| 2010 | A Comparative Study of Word Co-occurrence for Term Clustering in Language Model-based Sentence Retrieval
Saeedeh Momtazi, Sanjeev Khudanpur, Dietrich Klakow |
HLT-NAACL | 3 |
| 2010 | Convolution Kernels for Opinion Holder Extraction
Michael Wiegand, Dietrich Klakow |
HLT-NAACL | 2 |
| 2010 | Hierarchical pitman-yor language model for information retrievalabstractIn this paper, we propose a new application of Bayesian language model based on Pitman-Yor process for information retrieval. This model is a generalization of the Dirichlet distribution. The Pitman-Yor process creates a power-law distribution which is one of the statistical properties of word frequency in natural language. Our experiments on Robust04 indicate that this model improves the document retrieval performance compared to the commonly used Dirichlet prior and absolute discounting smoothing techniques. Saeedeh Momtazi, Dietrich Klakow |
SIGIR | 2 |
| 2010 | Cross-language retrieval using link-based language modelsabstractWe propose a cross-language retrieval model that is solely based on Wikipedia as a training corpus. The main contributions of our work are: 1. A translation model based on linked text in Wikipedia and a term weighting method associated with it. 2. A combination scheme to interpolate the link translation model with retrieval based on Latent Dirichlet Allocation. On the CLEF 2000 data we achieve improvement with respect to the best German-English system at the bilingual track (non-significant) and improvement against a baseline based on machine translation (significant). Benjamin Roth 0001, Dietrich Klakow |
SIGIR | 2 |
| 2009 | A word clustering approach for language model-based sentence retrieval in question answering systemsabstractIn this paper we propose a term clustering approach to improve the performance of sentence retrieval in Question Answering (QA) systems. As the search in question answering is conducted over smaller segments of data than in a document retrieval task, the problems of data sparsity and exact matching become more critical. In this paper we propose Language Modeling (LM) techniques to overcome such problems and improve the sentence retrieval performance. Saeedeh Momtazi, Dietrich Klakow |
CIKM | 2 |
| 2009 | Bounded conditional mean imputation with Gaussian mixture models: A reconstruction approach to partly occluded featuresabstractIn this work we show how conditional mean imputation can be bounded through the use of box-truncated Gaussian distributions. That is of interest when signals or features are partly occluded by a superimposed interference, as then the noisy observation poses an upper bound. Unfortunately, the occurring integrals are not analytic. Hence an approximate solution has to be used. In the experimental section we apply the bounded approach to the reconstruction of partly occluded speech spectra and demonstrate its superiority over the unbounded case with respect to automatic speech recognition performance. Friedrich Faubel, John W. McDonough, Dietrich Klakow |
ICASSP | 3 |
| 2009 | Impact of novel sources on content-based image and video retrievalabstractThe problem of content-based image and video retrieval with textual queries is often posed as that of visual concept classification, where classifiers for a set of predetermined visual concepts are trained using a set of manually annotated images. Such a formulation implicitly assumes that the training data has similar distributional characteristics as that of the data which need to be indexed. In this paper we demonstrate empirically that even within the relatively narrow domain of news videos collected from a variety of news programs and broadcasters, the assumption of distributional similarity of visual features does not hold across programs from different broadcasters. This is manifested in considerable degradation of ranked retrieval performance on novel sources. We observe that concepts whose spatial locations remain relatively fixed between various sources are also more robust to source mismatches, and vice versa. We also show that a simple averaging of multiple visual detectors is more robust than any of the individual detectors. Furthermore, we show that for certain sources using only 20% of the available annotated data can bridge roughly 80% of the performance drop, while others can require larger amounts of annotated data. Arnab Ghoshal, Sanjeev Khudanpur, Dietrich Klakow |
ICASSP | 3 |
| 2009 | A Combined Query Expansion Technique for Retrieving Opinions from BlogsabstractIn this paper, we discuss the the role of the retrieval component in an TREC style opinion question answering system. Since blog retrieval differs from traditional ad-hoc document retrieval, we need to work on dedicated retrieval methods. In particular we focus on a new query expansion technique to retrieve people's opinions from blog posts. We propose a combined approach for expanding queries while considering two aspects: finding more relevant data, and finding more opinionative data. We introduce a method to select opinion bearing terms for query expansion based on a chi-squared test and use this new query expansion to combine it in a liner weighting scheme with the original query terms and relevant feedback terms from Web. We report our experiments on the TREC 2006 and TREC 2007 queries from the blog retrieval track. The results show that the methods investigated here enhanced mean average precision of document retrieval from 17.91% to 25.20% on TREC 2006 and from 22.28% to 32.61% on TREC 2007 queries. Saeedeh Momtazi, Stefan Kazalski, Dietrich Klakow |
ISDA | 3 |
| 2009 | The Split and Merge Unscented Gaussian Mixture FilterabstractIn this work we present a novel approach to nonlinear, non-Gaussian tracking problems based on splitting and merging Gaussian filters in order to increase the level of detail of the filtering density in likely regions of the state space and reduce it in unlikely ones. As this is only effective in the presence of nonlinearities, we describe a split control technique that prevents filters from being split if they operate in linear regions of state space. In simulations with polar measurements, the new algorithm reduced the mean square error by nearly 50% compared to the unscented Kalman filter. Friedrich Faubel, John W. McDonough, Dietrich Klakow |
IEEE Signal Process. Lett. | 3 |
| 2009 | Beamforming With a Maximum Negentropy CriterionabstractIn this paper, we address a beamforming application based on the capture of far-field speech data from a single speaker in a real meeting room. After the position of the speaker is estimated by a speaker tracking system, we construct a subband-domain beamformer in generalized sidelobe canceller (GSC) configuration. In contrast to conventional practice, we then optimize the active weight vectors of the GSC so as to obtain an output signal with maximum negentropy (MN). This implies the beamformer output should be as non-Gaussian as possible. For calculating negentropy, we consider the Gamma and the generalized Gaussian (GG) pdfs. After MN beamforming, Zelinski postfiltering is performed to further enhance the speech by removing residual noise. Our beamforming algorithm can suppress noise and reverberation without the signal cancellation problems encountered in the conventional beamforming algorithms. We demonstrate this fact through a set of acoustic simulations. Moreover, we show the effectiveness of our proposed technique through a series of far-field automatic speech recognition experiments on the Multi-Channel Wall Street Journal Audio Visual Corpus (MC-WSJ-AV), a corpus of data captured with real far-field sensors, in a realistic acoustic environment, and spoken by real speakers. On the MC-WSJ-AV evaluation data, the delay-and-sum beamformer with postfiltering achieved a word error rate (WER) of 16.5%. MN beamforming with the Gamma pdf achieved a 15.8% WER, which was further reduced to 13.2% with the GG pdf, whereas the simple delay-and-sum beamformer provided a WER of 17.8%. To the best of our knowledge, no lower error rates at present have been reported in the literature on this automatic speech recognition (ASR) task. Ken'ichi Kumatani, John W. McDonough, Barbara Rauch 0001, Dietrich Klakow, Philip N. Garner, Weifeng Li 0001 |
IEEE Trans. Speech Audio Process. | 4 |
| 2008 | Optimizing Language Models for Polarity Classification
Michael Wiegand, Dietrich Klakow |
ECIR | 2 |
| 2008 | Filter bank design based on minimization of individual aliasing terms for minimum mutual information subband adaptive beamformingabstractThis paper presents new filter bank design methods for sub- band adaptive beamforming. In this work, we design analysis and synthesis prototypes for modulated filter banks so as to minimize each aliasing term individually. We then drive the total response error to null by constraining these prototypes to be Nyquist(M) filters. Thereafter those modulated filter banks are applied to a speech separation system which extracts a target speech signal. In our system, speech signals are first transformed into the subband domain with our filter banks, and the subband components are then processed with a beamforming algorithm. Following beamforming, post-filtering and binary masking are further performed to remove residual noises. We show that our filter banks can suppress the residual aliasing distortion more than conventional ones. Furthermore, we demonstrate the effectiveness of our design techniques through a set of automatic speech recognition experiments on the multi-channel speech data from the PASCAL Speech Separation Challenge. The experimental results prove that our beamforming system with the proposed filter banks achieves the best recognition performance, a 39.6 % word error rate (WER), with half the amount of computation of that of the conventional filter banks while the perfect reconstruction filter banks provided a 44.4 % WER. Ken'ichi Kumatani, John W. McDonough, S. Schachl, Dietrich Klakow, Philip N. Garner, Weifeng Li 0001 |
ICASSP | 4 |
| 2008 | A phase-averaged model for the relationship between noisy speech, clean speech and noise in the log-mel domainabstractIn this work, we demonstrate that the most widely-used model for the relationship between noisy speech, clean speech and noise in the log-Mel domain is inaccurate due to its disregard of the phase. Moreover, we show how a more exact model can be derived by averaging over the phase in the log-Mel domain, and how this can profitably be applied to particle filter based sequential noise compensation. Experimental results confirm the superiority of the phase-averaged model for both clean speech estimation in general and the particle filter in particular. Reductions in word error rate of up to 17 % relative were obtained on a large vocabulary task. Index Terms: model, relative phase, noise compensation 1. Friedrich Faubel, John W. McDonough, Dietrich Klakow |
INTERSPEECH | 3 |
| 2008 | Cost-Sensitive Learning in Answer Extraction
Michael Wiegand, Jochen L. Leidner, Dietrich Klakow |
LREC | 3 |
| 2007 | An algorithm for fast composition of weighted finite-state transducersabstractIn automatic speech recognition based on weighted-finite transducers, a static decoding graph HC o L o G is typically constructed. In this work, we first show how the size of the decoding graph can be reduced and the necessity of determinizing it can be eliminated by removing the ambiguity associated with transitions to the backoff state or states in G. We then show how the static construction can be avoided entirely by performing fast on-the-fly composition of HC and L o G. We demonstrate that speech recognition based on this on-the-fly composition approximately 80% more run-time than recognition based on the statically-expanded network R, which makes it competitive compared with other dynamic expansion algorithms that have appeared in the literature. Moreover, the dynamic algorithm requires a factor of approximately seven less main memory as the recognition based on the static decoding graph. John W. McDonough, Emilian Stoimenov, Dietrich Klakow |
ASRU | 3 |
| 2007 | Language Model Based Query Classification
Andreas Merkel, Dietrich Klakow |
ECIR | 2 |
| 2007 | Improved methods for language model based question classificationabstractIn this paper, we propose a language model based approach to classify user questions in the context of question answering systems. As categorization paradigm, a Bayes classifier is used to determine a corresponding semantic class. We present experiments with state-of-the-art smoothing methods as well as with some improved language models. Our results indicate that the techniques proposed here provide performance superior to the standard methods, including support vector machines. Andreas Merkel, Dietrich Klakow |
INTERSPEECH | 2 |
| 2007 | Combining term-based and event-based matching for question answeringabstractIn question answering, two main kinds of matching methods for finding answer sentences for a question are term-based approaches -- which are simple, efficient, effective, and yield high recall -- and event-based approaches that take syntactic and semantic information into account. The latter often sacrifice recall for increased precision, but actually capture the meaning of the events denoted by the textual units of a passage or sentence. We propose a robust, data-driven method that learns the mapping between questions and answers using logistic regression and show that combining term-based and event-based approaches significantly outperforms the individual methods. Michael Wiegand, Jochen L. Leidner, Dietrich Klakow |
SIGIR | 3 |
| 2006 | Exploring Correlation of Dependency Relation Paths for Answer ExtractionabstractIn this paper, we explore correlation of dependency relation paths to rank candidate answers in answer extraction.Using the correlation measure, we compare dependency relations of a candidate answer and mapped question phrases in sentence with the corresponding relations in question.Different from previous studies, we propose an approximate phrase mapping algorithm and incorporate the mapping score into the correlation measure.The correlations are further incorporated into a Maximum Entropy-based ranking model which estimates path weights from training.Experimental results show that our method significantly outperforms state-ofthe-art syntactic relation-based methods by up to 20% in MRR. Dan Shen 0001, Dietrich Klakow |
ACL | 2 |
| 2006 | Using Regional Information in Language Model Based Automatic Concept Annotation and Retrieval Of VideoabstractIn this paper we propose the use of regional information for the TREC video retrieval task of retrieving key-frames showing specific concepts in the image. We observe that the presence of a certain visual feature (e.g. color) is a strong indicator for a concept in particular if it occurs in a certain region of the image. We use this observation in a language model based image retrieval framework. This approach improves mean average precision from 0.187 (HMM based approach by Ghoshal et al.) to 0.221 Dietrich Klakow |
ICASSP (2) | 1 |
| 2006 | Language model adaptation for tiny adaptation corpora
Dietrich Klakow |
INTERSPEECH | 1 |
| 2006 | Building an Evaluation Corpus for German Question Answering by Harvesting Wikipedia
Irene M. Cramer, Jochen L. Leidner, Dietrich Klakow |
LREC | 3 |
| 2005 | A Statistical Classification Approach to Question Answering using Web DataabstractIn this paper we treat question answering (QA) as a classification problem. Our motivation is to build systems for many languages without the need for highly tuned linguistic modules. Consequently, word tokens and Web data are used extensively but no explicit linguistic knowledge is incorporated. A mathematical model for answer retrieval, answer classification and answer length prediction is derived. The TREC 2002 QA task is used for system development where 33% of questions are answered correctly. Performance is then evaluated on the factoid questions of the TREC 2003 QA task where 23% of questions were answered correctly, which would rank the system in the top 10 of contemporary QA systems on the same task Edward W. D. Whittaker, Sadaoki Furui, Dietrich Klakow |
CW | 3 |
| 2005 | Exploring Syntactic Relation Patterns for Question Answering
Dan Shen 0001, Geert-Jan M. Kruijff, Dietrich Klakow |
IJCNLP | 3 |
| 2005 | Joint visual-text modeling for automatic retrieval of multimedia documentsabstractIn this paper we describe a novel approach for jointly modeling the text and the visual components of multimedia documents for the purpose of information retrieval(IR). We propose a novel framework where individual components are developed to model different relationships between documents and queries and then combined into a joint retrieval framework. In the state-of-the-art systems, a late combination between two independent systems, one analyzing just the text part of such documents, and the other analyzing the visual part without leveraging any knowledge acquired in the text processing, is the norm. Such systems rarely exceed the performance of any single modality (i.e. text or video) in information retrieval tasks. Our experiments indicate that allowing a rich interaction between the modalities results in significant improvement in performance over any single modality. We demonstrate these results using the TRECVID03 corpus, which comprises 120 hours of broadcast news videos. Our results demonstrate over 14 % improvement in IR performance over the best reported text-only baseline and ranks amongst the best results reported on this corpus. Giridharan Iyengar, Pinar Duygulu, Shaolei Feng 0001, Pavel Ircing, Sanjeev Khudanpur, Dietrich Klakow, M. R. Krause, R. Manmatha, Harriet J. Nock, D. Petkova, Brock Pytlik, Paola Virga |
ACM Multimedia | 6 |
| 2003 | Information retrieval based call classification
Jan Kneissler, Anne K. Kienappel, Dietrich Klakow |
INTERSPEECH | 3 |
| 2002 | Efficient construction of long-range language models using log-linear interpolationabstractIn this paper we examine the construction of long-range language models using log-linear interpolation and how this can be achieved effectively. Particular attention is paid to the efficient computation of the normalisation in the models. Using the Penn Treebank for experiments we argue that the perplexity performance demonstrated recently in the literature using grammar-based approaches can actually be achieved with an appropriately smoothed 4-gram language model. Using such a model as the baseline, we demonstrate how further improvements can be obtained using loglinear interpolation to combine distance word and class models. We also examine the performance of similar model combinations for rescoring word lattices on a medium-sized vocabulary Wall Street Journal task. 1. Edward W. D. Whittaker, Dietrich Klakow |
INTERSPEECH | 2 |
| 2002 | Large vocabulary continuous speech recognition of Broadcast News - The Philips/RWTH approach
Peter Beyerlein, Xavier L. Aubert, Reinhold Häb-Umbach, Matthew Harris, Dietrich Klakow, Andreas Wendemuth, Sirko Molau, Hermann Ney, Michael Pitz, Achim Sixtus |
Speech Commun. | 5 |
| 2002 | Testing the correlation of word error rate and perplexity
Dietrich Klakow, Jochen Peters |
Speech Commun. | 1 |
| 2001 | Generation and expansion of word graphs using long span context informationabstractAn algorithm for the generation of word graphs in a cross-word decoder that uses long span m-gram language models (LMs) is presented. The generation of word hypotheses within the graph relies on the word m-tuple-based boundary optimization. The graphs contain the full word history knowledge information since the graph structure reflects all LM constraints used during the search. This results in better word boundaries and in enhanced capabilities to prune the graphs. Furthermore, the memory costs for expanding the m-gram constrained word graphs to apply very long span LMs (e.g. ten-grams that are constructed by log linear LM combination) are considerably reduced. Experiments for lattice generation and rescoring have been carried out on the 5K-word WSJ task and the 64K-word NAB task. Christoph Neukirchen, Dietrich Klakow, Xavier L. Aubert |
ICASSP | 2 |
| 2001 | Speech recognition for huge vocabularies by using optimized sub-word unitsabstractThis paper describes approaches for decomposing words of huge vocabularies (up to 2 million) into smaller particles that are suitable for a recognition lexicon. Results on a Finnish dictation task and a flat list of German street names are given. Jan Kneissler, Dietrich Klakow |
INTERSPEECH | 2 |
| 2000 | Selecting articles from the language model training corpusabstractThe paper suggests the use of a log-likelihood based criterion to select articles from a training corpus that are suitable to reduce perplexity on a specific task defined by a small target corpus. This method is not only efficient as an adaptation technique reducing perplexity by 32% and OOV rate from 4.2% to 2.7% but also as a pruning technique, decreasing the language model size by a factor of 3 at the same time. Dietrich Klakow |
ICASSP | 1 |
| 2000 | Long range language models for free spelling recognitionabstractHighly accurate spelling recognizers are essential in many commercially relevant applications. Examples are directory assistance, address taking in ordering services or help desks, and input of difficult or unknown words in dictation. Language models whose context length is flexibly configured by incorporating automatically determined letter groups code the structure of the recognition items in traditional bi- or trigrams. This gives a powerful method trading off computational demands versus recognition accuracy, comparing favorably with the standard word-list constraint. In addition, the automatic modeling of new words proves its benefits in applications with high out-of-vocabulary rates. For the task of spelling German last names over the telephone, letter error rates could be improved from 12.5% using a standard bigram to 3.6% with a trigram on a set of 2782 letter groups, giving the additional benefit of recognizing about 40% of the names not seen before in the language model training corpus. Frank Thiele, Bernhard Rüber, Dietrich Klakow |
ICASSP | 3 |
| 1999 | The philips/RWTH system for transcription of broadcast newsabstractThis paper contains a description of the Philips/RWTH 1998 HUB4 system which has been build in a joint eort of Philips Research Laboratories Aachen and Aachen University o f T echnology.We will focus our discussion on recent improvements compared to the original 1997 HUB4 system and evaluate them on the HUB4'97 evaluation data.The paper will deal with 1. a rough system overview including feature extraction, acoustic training, audio stream segmentation, and decoding 2. log-linear interpolation of distance-language models, 3. and the integration of various acoustic and language models via Discriminative Model Combination (DMC).The performance of the described system is 23% (relative) better than the performance of the 1997 Philips HUB4 system.A w ord error rate of 17.9% was achieved on the 1997 HUB4 evaluation set, compared to 23.5% using the original 1997 system. Peter Beyerlein, Xavier L. Aubert, Reinhold Häb-Umbach, Matthew Harris, Dietrich Klakow, Andreas Wendemuth, Sirko Molau, Michael Pitz, Achim Sixtus |
EUROSPEECH | 5 |
| 1999 | OOV-detection in large vocabulary system using automatically defined word-fragments as fillersabstractThe problem of unknown words has been addressed using automatically generated ller fragments which augment the lexicon and are incorporated in the language model. These fragments are used to reduce the damage on in-vocabulary words, to detect OOV regions and to provide a phonetic transcription for these regions. The performance of this technique has been evaluated in terms of damage reduction error rate and OOV tagging rate. Signi cant improvements are reported on both measures. In particular, the in uence of an appropriate tuning of the language model factor and word penalties is demonstrated as well as the usefulness of using cross-word triphones over fragments boundaries. Dietrich Klakow, Georg Rose, Xavier L. Aubert |
EUROSPEECH | 1 |
| 1998 | Language-model optimization by mapping of corporaabstractIt is questionable whether words are really the best basic units for the estimation of stochastic language models-grouping frequent word sequences to phrases can improve language models. More generally, we have investigated various coding schemes for a corpus. In this paper, it is applied to optimize the perplexity of n-gram language models. In tests on two large corpora (WSJ and BNA) the bigram perplexity was reduced by up to 29%. Furthermore, this approach allows to tackle the problem of an open vocabulary with no unknown word. Dietrich Klakow |
ICASSP | 1 |
| 1998 | Log-linear interpolation of language modelsabstractA new method to combine language models is derived. This method of log-linear interpolation (LLI) is used for adaptation and for combining models of dierent context length. In both cases LLI is better than linear interpolation. 1 Dietrich Klakow |
ICSLP | 1 |
| 1997 | Language model adaptation using dynamic marginalsabstractA new method is presented to quickly adapt a given language model to local text characteristics. The basic approach is to choose the adaptive models as close as possible to the background estimates while constraining them to respect the locally estimated unigram probabilities. Several means are investigated to speed up the calculations. We measure both perplexity and word error rate to gauge the quality of our model. Reinhard Kneser, Jochen Peters, Dietrich Klakow |
EUROSPEECH | 3 |