Heike Adel

dblp:132/6980 · DBLP profile ↗
← Back
34ranked-venue papers
11as first author
13since 2021 · last 2025
0000-0003-2787-0084ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 10 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Lost in Multilinguality: Dissecting Cross-lingual Factual Inconsistency in Transformer Language Models
abstract
Mingyang Wang, Heike Adel, Lukas Lange, Yihong Liu, Ercong Nie, Jannik Strötgen, Hinrich Schuetze. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Mingyang Wang 0003, Heike Adel, Lukas Lange, Yihong Liu 0001, Ercong Nie, Jannik Strötgen, Hinrich Schütze
ACL (1)2
2025 Language Mixing in Reasoning Language Models: Patterns, Impact, and Internal Causes
abstract
Reasoning language models (RLMs) excel at complex tasks by leveraging a chain-of-thought process to generate structured intermediate steps.However, language mixing, i.e., reasoning steps containing tokens from languages other than the prompt, has been observed in their outputs and shown to affect performance, though its impact remains debated.We present the first systematic study of language mixing in RLMs, examining its patterns, impact, and internal causes across 15 languages, 7 task difficulty levels, and 18 subject areas, and show how all three factors influence language mixing.Moreover, we demonstrate that the choice of reasoning language significantly affects performance: forcing models to reason in Latin or Han scripts via constrained decoding notably improves accuracy.Finally, we show that the script composition of reasoning traces closely aligns with that of the model's internal representations, indicating that language mixing reflects latent processing preferences in RLMs.Our findings provide actionable insights for optimizing multilingual reasoning and open new directions for controlling reasoning languages to build more interpretable and adaptable RLMs. 1 4 This overthinking behavior is also observed in prior work such as Cuadron et al. (2025).
Mingyang Wang 0003, Lukas Lange, Heike Adel, Yunpu Ma, Jannik Strötgen, Hinrich Schütze
EMNLP3
2024 Explaining Pre-Trained Language Models with Attribution Scores: An Analysis in Low-Resource Settings
abstract
Attribution scores indicate the importance of different input parts and can, thus, explain model behaviour. Currently, prompt-based models are gaining popularity, i.a., due to their easier adaptability in low-resource settings. However, the quality of attribution scores extracted from prompt-based models has not been investigated yet. In this work, we address this topic by analyzing attribution scores extracted from prompt-based models w.r.t. plausibility and faithfulness and comparing them with attribution scores extracted from fine-tuned models and large language models. In contrast to previous work, we introduce training size as another dimension into the analysis. We find that using the prompting paradigm (with either encoder-based or decoder-based models) yields more plausible explanations than fine-tuning the models in low-resource settings and Shapley Value Sampling consistently outperforms attention and Integrated Gradients in terms of leading to more plausible and faithful explanations.
Wei Zhou 0067, Heike Adel, Hendrik Schuff, Ngoc Thang Vu
LREC/COLING2
2024 FREB-TQA: A Fine-Grained Robustness Evaluation Benchmark for Table Question Answering
abstract
Wei Zhou, Mohsen Mesgar, Heike Adel, Annemarie Friedrich. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Wei Zhou 0067, Mohsen Mesgar, Heike Adel, Annemarie Friedrich
NAACL-HLT3
2024 SciOL and MuLMS-Img: Introducing A Large-Scale Multimodal Scientific Dataset and Models for Image-Text Tasks in the Scientific Domain
abstract
In scientific publications, a substantial part of the information is expressed via figures containing images and diagrams. Hence, the retrieval of relevant figures given a natural language query is an important real-world task. However, due to the lack of training and evaluation data, most existing approaches are either limited to one modality or focus on non-scientific domains, making their application to scientific publications challenging.In this paper, we address this gap by introducing two novel datasets: (1) SciOL, the largest openly-licensed pre-training corpus for multimodal models in the scientific domain, covering multiple sciences including materials science, physics, and computer science, and (2) MuLMS-Img, a high-quality dataset in the materials science domain, manually annotated for various image-text tasks. Our experiments show that pre-training large-scale vision-language models on SciOL increases performance considerably across a broad variety of image-text tasks including figure type classification, optical character recognition, captioning, and figure retrieval. Using MuLMS-Img, we show that integrating text-based features extracted via a fine-tuned model for a specific domain can boost cross-modal scientific figure retrieval performance by up to 50%.
Tim Tarsi, Heike Adel, Jan Hendrik Metzen, Matteo Finco, Annemarie Friedrich
WACV2
2023 SwitchPrompt: Learning Domain-Specific Gated Soft Prompts for Classification in Low-Resource Domains
abstract
Prompting pre-trained language models leads to promising results across natural language processing tasks but is less effective when applied in low-resource domains, due to the domain gap between the pre-training data and the downstream task.In this work, we bridge this gap with a novel and lightweight prompting methodology called SwitchPrompt for the adaptation of language models trained on datasets from the general domain to diverse low-resource domains.Using domain-specific keywords with a trainable gated prompt, Switch-Prompt offers domain-oriented prompting, that is, effective guidance on the target domains for general-domain language models.Our fewshot experiments on three text classification benchmarks demonstrate the efficacy of the general-domain pre-trained language models when used with SwitchPrompt.They often even outperform their domain-specific counterparts trained with baseline state-of-the-art prompting methods by up to 10.7% performance increase in accuracy.This result indicates that SwitchPrompt effectively reduces the need for domain-specific language model pre-training.
Koustava Goswami, Lukas Lange, Jun Araki, Heike Adel
EACL4
2023 Multilingual Normalization of Temporal Expressions with Masked Language Models
abstract
The detection and normalization of temporal expressions is an important task and preprocessing step for many applications.However, prior work on normalization is rule-based, which severely limits the applicability in realworld multilingual settings, due to the costly creation of new rules.We propose a novel neural method for normalizing temporal expressions based on masked language modeling.Our multilingual method outperforms prior rule-based systems in many languages, and in particular, for low-resource languages with performance improvements of up to 33 F 1 on average compared to the state of the art.
Lukas Lange, Jannik Strötgen, Heike Adel, Dietrich Klakow
EACL3
2023 GradSim: Gradient-Based Language Grouping for Effective Multilingual Training
abstract
Most languages of the world pose low-resource challenges to natural language processing models.With multilingual training, knowledge can be shared among languages.However, not all languages positively influence each other and it is an open research question how to select the most suitable set of languages for multilingual training and avoid negative interference among languages whose characteristics or data distributions are not compatible.In this paper, we propose GradSim, a language grouping method based on gradient similarity.Our experiments on three diverse multilingual benchmark datasets show that it leads to the largest performance gains compared to other similarity measures and it is better correlated with cross-lingual model performance.As a result, we set the new state of the art on AfriSenti, a benchmark dataset for sentiment analysis on low-resource African languages.In our extensive analysis, we further reveal that besides linguistic features, the topics of the datasets play an important role for language grouping and that lower layers of transformer models encode language-specific features while higher layers capture task-specific information.
Mingyang Wang 0003, Heike Adel, Lukas Lange, Jannik Strötgen, Hinrich Schütze
EMNLP2
2023 How to do human evaluation: A brief introduction to user studies in NLP
abstract
Abstract Many research topics in natural language processing (NLP), such as explanation generation, dialog modeling, or machine translation, require evaluation that goes beyond standard metrics like accuracy or F1score toward a more human-centered approach. Therefore, understanding how to design user studies becomes increasingly important. However, few comprehensive resources exist on planning, conducting, and evaluating user studies for NLP, making it hard to get started for researchers without prior experience in the field of human evaluation. In this paper, we summarize the most important aspects of user studies and their design and evaluation, providing direct links to NLP tasks and NLP-specific challenges where appropriate. We (i) outline general study design, ethical considerations, and factors to consider for crowdsourcing, (ii) discuss the particularities of user studies in NLP, and provide starting points to select questionnaires, experimental designs, and evaluation methods that are tailored to the specific NLP tasks. Additionally, we offer examples with accompanying statistical evaluation code, to bridge the gap between theoretical guidelines and practical applications.
Hendrik Schuff, Lindsey Vanderlyn, Heike Adel, Ngoc Thang Vu
Nat. Lang. Eng.3
2022 CLIN-X: pre-trained language models and a study on cross-task transfer for concept extraction in the clinical domain
abstract
MOTIVATION: The field of natural language processing (NLP) has recently seen a large change toward using pre-trained language models for solving almost any task. Despite showing great improvements in benchmark datasets for various tasks, these models often perform sub-optimal in non-standard domains like the clinical domain where a large gap between pre-training documents and target documents is observed. In this article, we aim at closing this gap with domain-specific training of the language model and we investigate its effect on a diverse set of downstream tasks and settings. RESULTS: We introduce the pre-trained CLIN-X (Clinical XLM-R) language models and show how CLIN-X outperforms other pre-trained transformer models by a large margin for 10 clinical concept extraction tasks from two languages. In addition, we demonstrate how the transformer model can be further improved with our proposed task- and language-agnostic model architecture based on ensembles over random splits and cross-sentence context. Our studies in low-resource and transfer settings reveal stable model performance despite a lack of annotated data with improvements of up to 47 F1 points when only 250 labeled sentences are available. Our results highlight the importance of specialized language models, such as CLIN-X, for concept extraction in non-standard domains, but also show that our task-agnostic model architecture is robust across the tested tasks and languages so that domain- or task-specific adaptations are not required. AVAILABILITY AND IMPLEMENTATION: The CLIN-X language models and source code for fine-tuning and transferring the model are publicly available at https://github.com/boschresearch/clin_x/ and the huggingface model hub.
Lukas Lange, Heike Adel, Jannik Strötgen, Dietrich Klakow
Bioinform.2
2021 FAME: Feature-Based Adversarial Meta-Embeddings for Robust Input Representations
abstract
Combining several embeddings typically improves performance in downstream tasks as different embeddings encode different information.It has been shown that even models using embeddings from transformers still benefit from the inclusion of standard word embeddings.However, the combination of embeddings of different types and dimensions is challenging.As an alternative to attention-based meta-embeddings, we propose feature-based adversarial meta-embeddings (FAME) with an attention function that is guided by features reflecting word-specific properties, such as shape and frequency, and show that this is beneficial to handle subword-based embeddings.In addition, FAME uses adversarial training to optimize the mappings of differently-sized embeddings to the same space.We demonstrate that FAME works effectively across languages and domains for sequence labeling and sentence classification, in particular in lowresource settings.FAME sets the new state of the art for POS tagging in 27 languages, various NER settings and question classification in different domains.
Lukas Lange, Heike Adel, Jannik Strötgen, Dietrich Klakow
EMNLP (1)2
2021 To Share or not to Share: Predicting Sets of Sources for Model Transfer Learning
abstract
In low-resource settings, model transfer can help to overcome a lack of labeled data for many tasks and domains.However, predicting useful transfer sources is a challenging problem, as even the most similar sources might lead to unexpected negative transfer results.Thus, ranking methods based on task and text similarity -as suggested in prior workmay not be sufficient to identify promising sources.To tackle this problem, we propose a new approach to automatically determine which and how many sources should be exploited.For this, we study the effects of model transfer on sequence labeling across various domains and tasks and show that our methods based on model similarity and support vector machines are able to predict promising sources, resulting in performance increases of up to 24 F 1 points.
Lukas Lange, Jannik Strötgen, Heike Adel, Dietrich Klakow
EMNLP (1)3
2021 A Survey on Recent Approaches for Natural Language Processing in Low-Resource Scenarios
abstract
Michael A. Hedderich, Lukas Lange, Heike Adel, Jannik Strötgen, Dietrich Klakow. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Michael A. Hedderich, Lukas Lange, Heike Adel, Jannik Strötgen, Dietrich Klakow
NAACL-HLT3
2020 The SOFC-Exp Corpus and Neural Approaches to Information Extraction in the Materials Science Domain
abstract
Annemarie Friedrich, Heike Adel, Federico Tomazic, Johannes Hingerl, Renou Benteau, Anika Marusczyk, Lukas Lange. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Annemarie Friedrich, Heike Adel, Federico Tomazic, Johannes Hingerl, Renou Benteau, Anika Marusczyk, Lukas Lange
ACL2
2020 Closing the Gap: Joint De-Identification and Concept Extraction in the Clinical Domain
abstract
Exploiting natural language processing in the clinical domain requires de-identification, i.e., anonymization of personal information in texts. However, current research considers de-identification and downstream tasks, such as concept extraction, only in isolation and does not study the effects of de-identification on other tasks. In this paper, we close this gap by reporting concept extraction performance on automatically anonymized data and investigating joint models for de-identification and concept extraction. In particular, we propose a stacked model with restricted access to privacy-sensitive information and a multitask model. We set the new state of the art on benchmark datasets in English (96.1% F1 for de-identification and 88.9% F1 for concept extraction) and Spanish (91.4% F1 for concept extraction).
Lukas Lange, Heike Adel, Jannik Strötgen
ACL2
2020 An Analysis of Simple Data Augmentation for Named Entity Recognition
abstract
Simple yet effective data augmentation techniques have been proposed for sentence-level and sentence-pair natural language processing tasks.Inspired by these efforts, we design and compare data augmentation for named entity recognition, which is usually modeled as a token-level sequence labeling problem.Through experiments on two data sets from the biomedical and materials science domains (i2b2-2010 and MaSciP), we show that simple augmentation can boost performance for both recurrent and transformer-based models, especially for small training sets.
Xiang Dai 0001, Heike Adel
COLING2
2020 F1 is Not Enough! Models and Evaluation Towards User-Centered Explainable Question Answering
abstract
Explainable question answering systems predict an answer together with an explanation showing why the answer has been selected. The goal is to enable users to assess the correctness of the system and understand its reasoning process. However, we show that current models and evaluation settings have shortcomings regarding the coupling of answer and explanation which might cause serious issues in user experience. As a remedy, we propose a hierarchical model and a new regularization term to strengthen the answer-explanation coupling as well as two evaluation scores to quantify the coupling. We conduct experiments on the HOTPOTQA benchmark data set and perform a user study. The user study shows that our models increase the ability of the users to judge the correctness of the system and that scores like F1 are not enough to estimate the usefulness of a model in a practical setting with human users. Our scores are better aligned with user experience, making them promising candidates for model selection.
Hendrik Schuff, Heike Adel, Ngoc Thang Vu
EMNLP (1)2
2020 ExCut: Explainable Embedding-Based Clustering over Knowledge Graphs
Mohamed H. Gad-Elrab, Daria Stepanova 0001, Trung Kien Tran, Heike Adel, Gerhard Weikum
ISWC (1)4
2020 Tackling challenges of neural purchase stage identification from imbalanced twitter data
abstract
Abstract Twitter and other social media platforms are often used for sharing interest in products. The identification of purchase decision stages, such as in the AIDA model (Awareness, Interest, Desire, and Action), can enable more personalized e-commerce services and a finer-grained targeting of advertisements than predicting purchase intent only. In this paper, we propose and analyze neural models for identifying the purchase stage of single tweets in a user’s tweet sequence. In particular, we identify three challenges of purchase stage identification: imbalanced label distribution with a high number of non-purchase-stage instances, limited amount of training data, and domain adaptation with no or only little target domain data. Our experiments reveal that the imbalanced label distribution is the main challenge for our models. We address it with ranking loss and perform detailed investigations of the performance of our models on the different output classes. In order to improve the generalization of the models and augment the limited amount of training data, we examine the use of sentiment analysis as a complementary, secondary task in a multitask framework. For applying our models to tweets from another product domain, we consider two scenarios: for the first scenario without any labeled data in the target product domain, we show that learning domain-invariant representations with adversarial training is most promising, while for the second scenario with a small number of labeled target examples, fine-tuning the source model weights performs best. Finally, we conduct several analyses, including extracting attention weights and representative phrases for the different purchase stages. The results suggest that the model is learning features indicative of purchase stages and that the confusion errors are sensible.
Heike Adel, Francine Chen 0001, Yan-Ying Chen
Nat. Lang. Eng.1
2019 Automatic Compression of Subtitles with Neural Networks and its Effect on User Experience
Katrin Angerbauer, Heike Adel, Ngoc Thang Vu
INTERSPEECH2
2019 Type-aware Convolutional Neural Networks for Slot Filling
abstract
The slot filling task aims at extracting answers for queries about entities from text, such as "Who founded Apple". In this paper, we focus on the relation classification component of a slot filling system. We propose type-aware convolutional neural networks to benefit from the mutual dependencies between entity and relation classification. In particular, we explore different ways of integrating the named entity types of the relation arguments into a neural network for relation classification, including a joint training and a structured prediction approach. To the best of our knowledge, this is the first study on type-aware neural networks for slot filling. The type-aware models lead to the best results of our slot filling pipeline. Joint training performs comparable to structured prediction. To understand the impact of the different components of the slot filling pipeline, we perform a recall analysis, a manual error analysis and several ablation studies. Such analyses are of particular importance to other slot filling researchers since the official slot filling evaluations only assess pipeline outputs. The analyses show that especially coreference resolution and our convolutional neural networks have a large positive impact on the final performance of the slot filling pipeline. The presented models, the source code of our system as well as our coreference resource is publicly available.
Heike Adel, Hinrich Schütze
J. Artif. Intell. Res.1
2018 Corpus-Level Fine-Grained Entity Typing
abstract
Extracting information about entities remains an important research area. This paper addresses the problem of corpus-level entity typing, i.e., inferring from a large corpus that an entity is a member of a class, such as "food" or "artist". The application of entity typing we are interested in is knowledge base completion, specifically, to learn which classes an entity is a member of. We propose FIGMENT to tackle this problem. FIGMENT is embedding-based and combines (i) a global model that computes scores based on global information of an entity and (ii) a context model that first evaluates the individual occurrences of an entity and then aggregates the scores. Each of the two proposed models has specific properties. For the global model, learning high-quality entity representations is crucial because it is the only source used for the predictions. Therefore, we introduce representations using the name and contexts of entities on the three levels of entity, word, and character. We show that each level provides complementary information and a multi-level representation performs best. For the context model, we need to use distant supervision since there are no context-level labels available for entities. Distantly supervised labels are noisy and this harms the performance of models. Therefore, we introduce and apply new algorithms for noise mitigation using multi-instance learning. We show the effectiveness of our models on a large entity typing dataset built from Freebase.
Yadollah Yaghoobzadeh, Heike Adel, Hinrich Schütze
J. Artif. Intell. Res.2
2017 Overview of Character-Based Models for Natural Language Processing
Heike Adel, Ehsaneddin Asgari, Hinrich Schütze
CICLing (1)1
2017 Exploring Different Dimensions of Attention for Uncertainty Detection
abstract
Neural networks with attention have proven effective for many natural language processing tasks.In this paper, we develop attention mechanisms for uncertainty detection.In particular, we generalize standardly used attention mechanisms by introducing external attention and sequence-preserving attention.These novel architectures differ from standard approaches in that they use external resources to compute attention weights and preserve sequence information.We compare them to other configurations along different dimensions of attention.Our novel architectures set the new state of the art on a Wikipedia benchmark dataset and perform similar to the state-of-the-art model on a biomedical benchmark which uses a large set of linguistic features.
Heike Adel, Hinrich Schütze
EACL (1)1
2017 Noise Mitigation for Neural Entity Typing and Relation Extraction
abstract
In this paper, we address two different types of noise in information extraction models: noise from distant supervision and noise from pipeline input features.Our target tasks are entity typing and relation extraction.For the first noise type, we introduce multi-instance multi-label learning algorithms using neural network models, and apply them to fine-grained entity typing for the first time.Our model outperforms the state-of-the-art supervised approach which uses global embeddings of entities.For the second noise type, we propose ways to improve the integration of noisy entity type predictions into relation extraction.Our experiments show that probabilistic predictions are more robust than discrete predictions and that joint training of the two tasks performs best.
Yadollah Yaghoobzadeh, Heike Adel, Hinrich Schütze
EACL (1)2
2017 Global Normalization of Convolutional Neural Networks for Joint Entity and Relation Classification
abstract
We introduce globally normalized convolutional neural networks for joint entity classification and relation extraction.In particular, we propose a way to utilize a linear-chain conditional random field output layer for predicting entity types and relations between entities at the same time.Our experiments show that global normalization outperforms a locally normalized softmax layer on a benchmark dataset.
Heike Adel, Hinrich Schütze
EMNLP1
2016 Bi-directional recurrent neural network with ranking loss for spoken language understanding
abstract
This paper presents our latest investigation of recurrent neural networks for the slot filling task of spoken language understanding. We implement a bi-directional Elman-type recurrent neural network which takes the information not only from the past but also from the future context to predict the semantic label of the target word. Furthermore, we propose to use ranking loss function to train the model. This improves the performance over the cross entropy loss function. On the ATIS benchmark data set, we achieve a new state-of-the-art result of 95.56% F1-score without using any additional knowledge or data sources.
Ngoc Thang Vu, Pankaj Gupta 0003, Heike Adel, Hinrich Schütze
ICASSP3
2016 Comparing Convolutional Neural Networks to Traditional Models for Slot Filling
abstract
We address relation classification in the context of slot filling, the task of finding and evaluating fillers like "Steve Jobs" for the slot X in "X founded Apple".We propose a convolutional neural network which splits the input sentence into three parts according to the relation arguments and compare it to state-ofthe-art and traditional approaches of relation classification.Finally, we combine different methods and show that the combination is better than individual approaches.We also analyze the effect of genre differences on performance.
Heike Adel, Benjamin Roth 0001, Hinrich Schütze
HLT-NAACL1
2016 Combining Recurrent and Convolutional Neural Networks for Relation Classification
abstract
Ngoc Thang Vu, Heike Adel, Pankaj Gupta, Hinrich Schütze. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Ngoc Thang Vu, Heike Adel, Pankaj Gupta 0003, Hinrich Schütze
HLT-NAACL2
2015 Syntactic and Semantic Features For Code-Switching Factored Language Models
abstract
This paper presents our latest investigations on different features for factored language models for Code-Switching speech and their effect on automatic speech recognition (ASR) performance. We focus on syntactic and semantic features which can be extracted from Code-Switching text data and integrate them into factored language models. Different possible factors, such as words, part-of-speech tags, Brown word clusters, open class words and clusters of open class word embeddings are explored. The experimental results reveal that Brown word clusters, part-of-speech tags and open-class words are the most effective at reducing the perplexity of factored language models on the Mandarin-English Code-Switching corpus SEAME. In ASR experiments, the model containing Brown word clusters and part-of-speech tags and the model also including clusters of open class word embeddings yield the best mixed error rate results. In summary, the best language model can significantly reduce the perplexity on the SEAME evaluation set by up to 10.8% relative and the mixed error rate by up to 3.4% relative.
Heike Adel, Ngoc Thang Vu, Katrin Kirchhoff, Dominic Telaar, Tanja Schultz
IEEE ACM Trans. Audio Speech Lang. Process.1
2014 Using Mined Coreference Chains as a Resource for a Semantic Task
abstract
We propose to use coreference chains ex-tracted from a large corpus as a resource for semantic tasks. We extract three mil-lion coreference chains and train word embeddings on them. Then, we com-pare these embeddings to word vectors de-rived from raw text data and show that coreference-based word embeddings im-prove F1 on the task of antonym classifi-cation by up to.09. 1
Heike Adel, Hinrich Schütze
EMNLP1
2014 Comparing approaches to convert recurrent neural networks into backoff language models for efficient decoding
abstract
In this paper, we investigate and compare three different possibilities to convert recurrent neural network language models (RNNLMs) into backoff language models (BNLM). While RNNLMs often outperform traditional n-gram approaches in the task of language modeling, their computational demands make them unsuitable for an efficient usage during decoding in an LVCSR system. It is, therefore, of interest to convert them into BNLMs in order to integrate their information into the decoding process. This paper compares three different approaches: a text based conversion, a probability based conversion and an iterative conversion. The resulting language models are evaluated in terms of perplexity and mixed error rate in the context of the Code-Switching data corpus SEAME. Although the best results are obtained by combining the results of all three approaches, the text based conversion approach alone leads to significant improvements on the SEAME corpus as well while offering the highest computational efficiency. In total, the perplexity can be reduced by 11.4% relative on the evaluation set and the mixed error rate by 3.0% relative on the same data set.
Heike Adel, Katrin Kirchhoff, Ngoc Thang Vu, Dominic Telaar, Tanja Schultz
INTERSPEECH1
2014 Combining recurrent neural networks and factored language models during decoding of code-Switching speech
abstract
In this paper, we present our latest investigations of language modeling for Code-Switching. Since there is only little text material for Code-Switching speech available, we integrate syntactic and semantic features into the language modeling process. In particular, we use part-of-speech tags, language identifiers, Brown word clusters and clusters of open class words. We develop factored language models and convert recurrent neural network language models into backoff language models for an efficient usage during decoding. A detailed error analysis reveals the strengths and weaknesses of the different language models. When we interpolate the models linearly, we reduce the perplexity by 15.6% relative on the SEAME evaluation set. This is even slightly better than the result of the unconverted recurrent neural network. We also combine the language models during decoding and obtain a mixed error rate reduction of 4.4% relative on the SEAME evaluation set.
Heike Adel, Dominic Telaar, Ngoc Thang Vu, Katrin Kirchhoff, Tanja Schultz
INTERSPEECH1
2013 Recurrent neural network language modeling for code switching conversational speech
abstract
Code-switching is a very common phenomenon in multilingual communities. In this paper, we investigate language modeling for conversational Mandarin-English code-switching (CS) speech recognition. First, we investigate the prediction of code switches based on textual features with focus on Part-of-Speech (POS) tags and trigger words. Second, we propose a structure of recurrent neural networks to predict code-switches. We extend the networks by adding POS information to the input layer and by factorizing the output layer into languages. The resulting models are applied to our task of code-switching language modeling. The final performance shows 10.8% relative improvement in perplexity on the SEAME development set which transforms into a 2% relative improvement in terms of Mixed Error Rate and a relative improvement of 16.9% in perplexity on the evaluation set which leads to a 2.7% relative improvement of MER.
Heike Adel, Ngoc Thang Vu, Franziska Kraus, Tim Schlippe, Haizhou Li 0001, Tanja Schultz
ICASSP1