Benjamin Roth 0001

dblp:63/8171-1 · DBLP profile ↗
← Back
37ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0003-0362-0267ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 5 first-author · 12 since 2021Databases, data management, data science and information retrieval · 10 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Influential Training Data Retrieval for Explaining Verbalized Confidence of LLMs
Yuxi Xia, Loris Schoenegger, Benjamin Roth 0001
ECIR (1)3
2026 The Impact of Graph Structure, Cluster Centroid and Text Review Embeddings on Recommendation Methods
abstract
It is generally accepted that collaborative information is important for the performance of recommender systems. It is also generally accepted that if this information is sparser, it impacts recommendation systems negatively. Various approaches have tried to lift this problem by employing side information. However, global patterns that can be provided by clusters of similar items and users or even additional information such as text are often not used together with collaborative information. We study the impact of integrating clustering embeddings, review embeddings, and their combinations with embeddings obtained by a recommender system. We study the performance of this approach across various state-of-the-art recommender system algorithms including graph-based methods. We highlight that graph structures are important with sparser datasets and both, in knowledge graphs with side information as well as in collaborative bipartite graphs. In less sparse datasets, a collaborative bipartite graph is usually sufficient. We also highlight that the improvement of recommendation performance through clustering, particularly evident when combined with review embeddings is most visible on sparser data, while on less sparse data incorporating review embeddings may be sufficient when combined with one of the graph-based methods, or otherwise when combined with clustering in other methods.
Peter Dolog, Sergio David Rico Torres, Yllka Velaj, Ylli Sadikaj, Andreas Stephan, Benjamin Roth 0001, Claudia Plant
Trans. Recomm. Syst.6
2025 Influences on LLM Calibration: A Study of Response Agreement, Loss Functions, and Prompt Styles
abstract
Calibration, the alignment between model confidence and prediction accuracy, is critical for the reliable deployment of large language models (LLMs).Existing works neglect to measure the generalization of their methods to other prompt styles and different sizes of LLMs.To address this, we define a controlled experimental setting covering 12 LLMs and four prompt styles.We additionally investigate if incorporating the response agreement of multiple LLMs and an appropriate loss function can improve calibration performance.Concretely, we build Calib-n, a novel framework that trains an auxiliary model for confidence estimation that aggregates responses from multiple LLMs to capture inter-model agreement.To optimize calibration, we integrate focal and AUC surrogate losses alongside binary cross-entropy.Experiments across four datasets demonstrate that both response agreement and focal loss improve calibration from baselines.We find that few-shot prompts are the most effective for auxiliary model-based methods, and auxiliary models demonstrate robust calibration performance across accuracy variations, outperforming LLMs' internal probabilities and verbalized confidences. 1 Sigmoid ...
Yuxi Xia, Pedro Henrique Luz de Araujo, Klim Zaporojets, Benjamin Roth 0001
ACL (1)4
2025 Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance
abstract
Expert persona prompting-assigning roles such as expert in math to language models-is widely used for task improvement.However, prior work shows mixed results on its effectiveness, and does not consider when and why personas should improve performance.We analyze the literature on persona prompting for task improvement and distill three desiderata: 1) performance advantage of expert personas, 2) robustness to irrelevant persona attributes, and 3) fidelity to persona attributes.We then evaluate 9 state-of-the-art LLMs across 27 tasks with respect to these desiderata.We find that expert personas usually lead to positive or non-significant performance changes.Surprisingly, models are highly sensitive to irrelevant persona details, with performance drops of almost 30 percentage points.In terms of fidelity, we find that while higher education, specialization, and domain-relatedness can boost performance, their effects are often inconsistent or negligible across tasks.We propose mitigation strategies to improve robustness-but find they only work for the largest, most capable models.Our findings underscore the need for more careful persona design and for evaluation schemes that reflect the intended effects of persona usage.
Pedro Henrique Luz de Araujo, Paul Röttger, Dirk Hovy, Benjamin Roth 0001
EMNLP4
2024 Text-Guided Image Clustering
abstract
Andreas Stephan, Lukas Miklautz, Kevin Sidak, Jan Philip Wahle, Bela Gipp, Claudia Plant, Benjamin Roth. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Andreas Stephan, Lukas Miklautz, Kevin Sidak, Jan Philip Wahle, Bela Gipp, Claudia Plant, Benjamin Roth 0001
EACL (1)7
2024 Counterfactual Reasoning with Knowledge Graph Embeddings
abstract
Knowledge graph embeddings (KGEs) were originally developed to infer true but missing facts in incomplete knowledge repositories.In this paper, we link knowledge graph completion and counterfactual reasoning via our new task CFKGR.We model the original world state as a knowledge graph, hypothetical scenarios as edges added to the graph, and plausible changes to the graph as inferences from logical rules.We create corresponding benchmark datasets, which contain diverse hypothetical scenarios with plausible changes to the original knowledge graph and facts that should be retained.We develop COULDD, a general method for adapting existing knowledge graph embeddings given a hypothetical premise, and evaluate it on our benchmark.Our results indicate that KGEs learn patterns in the graph without explicit training.We further observe that KGEs adapted with COULDD solidly detect plausible counterfactual changes to the graph that follow these patterns.An evaluation on human-annotated data reveals that KGEs adapted with COULDD are mostly unable to recognize changes to the graph that do not follow learned inference rules.In contrast, Chat-GPT mostly outperforms KGEs in detecting plausible changes to the graph but has poor knowledge retention.In summary, CFKGR connects two previously distinct areas, namely KG completion and counterfactual reasoning.
Lena Zellinger, Andreas Stephan, Benjamin Roth 0001
EACL (1)3
2023 Seeing through the mess: evolutionary dynamics of lexical polysemy
abstract
Evidently, words can have multiple senses.For example, the word mess refers to a place to have food or to a confusing situation.How exactly multiple senses emerge is less clear.In this work, we propose and analyze a mathematical model of the evolution of lexical meaning to investigate mechanisms leading to polysemy.This model features factors that have been discussed to impact the semantic processing and transmission of words: word frequency, nonconformism, and semantic discriminability.We formally derive conditions under which a sense of a word tends to diversify itself into multiple senses that coexist stably.The model predicts that diversification is promoted by low frequency, a strong bias for nonconformist usage, and high semantic discriminability.We statistically validate these predictions with historical language data covering semantic developments of a set of English words.Multiple alternative measures are used to operationalize each variable involved, and we confirm the predicted tendencies for twelve combinations of measures.
Andreas Stephan, Benjamin Roth 0001
EMNLP3
2023 ULF: Unsupervised Labeling Function Correction using Cross-Validation for Weak Supervision
abstract
A cost-effective alternative to manual data labeling is weak supervision (WS), where data samples are automatically annotated using a predefined set of labeling functions (LFs), rulebased mechanisms that generate artificial labels for the associated classes.In this work, we investigate noise reduction techniques for WS based on the principle of k-fold crossvalidation.We introduce a new algorithm ULF for Unsupervised Labeling Function correction, which denoises WS data by leveraging models trained on all but some LFs to identify and correct biases specific to the held-out LFs.Specifically, ULF refines the allocation of LFs to classes by re-estimating this assignment on highly reliable cross-validated samples.Evaluation on multiple datasets confirms ULF's effectiveness in enhancing WS learning without the need for manual labeling.1
Anastasiya Sedova, Benjamin Roth 0001
EMNLP2
2023 MemeGraphs: Linking Memes to Knowledge Graphs
Vasiliki Kougia, Simon Fetzel, Thomas Kirchmair, Erion Çano, Sina Moayed Baharlou, Sahand Sharifzadeh, Benjamin Roth 0001
ICDAR (1)7
2023 Learning with Noisy Labels by Adaptive Gradient-Based Outlier Removal
Anastasiya Sedova, Lena Zellinger, Benjamin Roth 0001
ECML/PKDD (1)3
2023 Cross-functional Analysis of Generalization in Behavioral Learning
abstract
Abstract In behavioral testing, system functionalities underrepresented in the standard evaluation setting (with a held-out test set) are validated through controlled input-output pairs. Optimizing performance on the behavioral tests during training (behavioral learning) would improve coverage of phenomena not sufficiently represented in the i.i.d. data and could lead to seemingly more robust models. However, there is the risk that the model narrowly captures spurious correlations from the behavioral test suite, leading to overestimation and misrepresentation of model performance—one of the original pitfalls of traditional evaluation. In this work, we introduce BeLUGA, an analysis method for evaluating behavioral learning considering generalization across dimensions of different granularity levels. We optimize behavior-specific loss functions and evaluate models on several partitions of the behavioral test suite controlled to leave out specific phenomena. An aggregate score measures generalization to unseen functionalities (or overfitting). We use BeLUGA to examine three representative NLP tasks (sentiment analysis, paraphrase identification, and reading comprehension) and compare the impact of a diverse set of regularization and domain generalization methods on generalization performance.1
Pedro Henrique Luz de Araujo, Benjamin Roth 0001
Trans. Assoc. Comput. Linguistics2
2021 KnowMAN: Weakly Supervised Multinomial Adversarial Networks
abstract
The absence of labeled data for training neural models is often addressed by leveraging knowledge about the specific task, resulting in heuristic but noisy labels.The knowledge is captured in labeling functions, which detect certain regularities or patterns in the training samples and annotate corresponding labels for training.This process of weakly supervised training may result in an over-reliance on the signals captured by the labeling functions and hinder models to exploit other signals or to generalize well.We propose KnowMAN, an adversarial scheme that enables to control influence of signals associated with specific labeling functions.KnowMAN forces the network to learn representations that are invariant to those signals and to pick up other signals that are more generally associated with an output label.KnowMAN strongly improves results compared to direct weakly supervised learning with a pre-trained transformer language model and a feature-based baseline.
Luisa März, Ehsaneddin Asgari, Fabienne Braune, Franziska Zimmermann, Benjamin Roth 0001
EMNLP (1)5
2021 Data Centric Domain Adaptation for Historical Text with OCR Errors
Luisa März, Stefan Schweter, Nina Pörner, Benjamin Roth 0001, Hinrich Schütze
ICDAR (2)4
2021 Python for Linguists
abstract
Teaching programming skills is a hard task.It is even harder if one targets an audience with no or little mathematical background.Although there are books on programming that target such groups, they often fail to raise or maintain interest due to artificial examples that lack reference to the professional issues that the audience typically face.This book fills the gap by addressing linguistics, a profession and academic subject for which basic knowledge of script programming is becoming more and more important.The book Python for Linguists by Michael Hammond is an introductory Python course targeted at linguists with no prior programming background.It succeeds previous books for Perl (Hammond 2008) and Java (Hammond 2002) by the same author, and reflects the current de facto prevalence of Python when it comes to adoption and available packages for natural language processing.We feel it necessary to clarify that the book aims at (general) linguists in the broad sense rather than computational linguists.Its aim is to teach linguists the fundamental concepts of programming using typical examples from linguistics.The book should not be mistaken as a course for learning basic algorithms in computational linguistics.We acknowledge that the author nowhere makes such a claim; however, given the thematic proximity to computational linguistics, one should have the right expectation before working with the book.Chapters 1-5 lay the foundations of the Python programming language, introducing the most important language constructs but deferring object oriented programming to a later part of the book.The focus in Chapters 1 and 2 covers the basic data types (numbers, strings, dictionaries), with a particular emphasis on simple string operations, and introduces some more advanced concepts such as mutability.Chapters 3-5 introduce control structures, input-output operations, and modules.The book goes at great length to visualize the program flow and the state of different variables for different steps in a program execution, which is certainly very helpful for learners with no prior programming experience.The book also guides the learner to understand certain error types that frequently occur in computer programming (but might be unintuitive for beginners).For example, when discussing function calls, much care is devoted to pointing out the unintended consequences stemming from mutability and side effects.
Benjamin Roth 0001, Michael Wiegand
Comput. Linguistics1
2020 UniSent: Universal Adaptable Sentiment Lexica for 1000+ Languages
abstract
In this paper, we introduce UniSent universal sentiment lexica for 1000+ languages. Sentiment lexica are vital for sentiment analysis in absence of document-level annotations, a very common scenario for low-resource languages. To the best of our knowledge, UniSent is the largest sentiment resource to date in terms of the number of covered languages, including many low resource ones. In this work, we use a massively parallel Bible corpus to project sentiment information from English to other languages for sentiment analysis on Twitter data. We introduce a method called DomDrift to mitigate the huge domain mismatch between Bible and Twitter by a confidence weighting scheme that uses domain-specific embeddings to compare the nearest neighbors for a candidate sentiment word in the source (Bible) and target (Twitter) domain. We evaluate the quality of UniSent in a subset of languages for which manually created ground truth was available, Macedonian, Czech, German, Spanish, and French. We show that the quality of UniSent is comparable to manually created sentiment resources when it is used as the sentiment seed for the task of word sentiment prediction on top of embedding representations. In addition, we show that emoticon sentiments could be reliably predicted in the Twitter domain using only UniSent and monolingual embeddings in German, Spanish, French, and Italian. With the publication of this paper, we release the UniSent sentiment lexica at http://language-lab.info/unisent.
Ehsaneddin Asgari, Fabienne Braune, Benjamin Roth 0001, Christoph Ringlstetter, Mohammad R. K. Mofrad
LREC3
2020 Dirichlet-Smoothed Word Embeddings for Low-Resource Settings
abstract
Nowadays, classical count-based word embeddings using positive pointwise mutual information (PPMI) weighted co-occurrence matrices have been widely superseded by machine-learning-based methods like word2vec and GloVe. But these methods are usually applied using very large amounts of text data. In many cases, however, there is not much text data available, for example for specific domains or low-resource languages. This paper revisits PPMI by adding Dirichlet smoothing to correct its bias towards rare words. We evaluate on standard word similarity data sets and compare to word2vec and the recent state of the art for low-resource settings: Positive and Unlabeled (PU) Learning for word embeddings. The proposed method outperforms PU-Learning for low-resource settings and obtains competitive results for Maltese and Luxembourgish.
Jakob Jungmaier, Nora Kassner, Benjamin Roth 0001
LREC3
2020 Intent Recognition in Doctor-Patient Interviews
abstract
Learning to interview patients to find out their disease is an essential part of the training of medical students. The practical part of this training has traditionally relied on paid actors that play the role of a patient to be interviewed. This process is expensive and severely limits the amount of practice per student. In this work, we present a novel data set and methods based on Natural Language Processing, for making progress towards modern applications and e-learning tools that support this training by providing language-based user interfaces with virtual patients. A data set of german transcriptions from live doctor-patient interviews was collected. These transcriptions are based on audio recordings of exercise sessions within the university and only the doctor’s utterances could be transcribed. We annotated each utterance with an intent inventory characterizing the purpose of the question or statement. For some intent classes, the data only contains a few samples, and we apply Information Retrieval and Deep Learning methods that are robust with respect to small amounts of training data for recognizing the intent of an utterance and providing the correct response. Our results show that the models are effective and they provide baseline performance scores on the data set for further research.
Robin Rojowiec, Benjamin Roth 0001, Maximilian Fink
LREC2
2019 Interpretable Question Answering on Knowledge Bases and Text
abstract
Interpretability of machine learning (ML) models becomes more relevant with their increasing adoption.In this work, we address the interpretability of ML based question answering (QA) models on a combination of knowledge bases (KB) and text documents.We adapt post hoc explanation methods such as LIME and input perturbation (IP) and compare them with the self-explanatory attention mechanism of the model.For this purpose, we propose an automatic evaluation paradigm for explanation methods in the context of QA.We also conduct a study with human annotators to evaluate whether explanations help them identify better QA models.Our results suggest that IP provides better explanations than LIME or attention, according to both automatic and human evaluation.We obtain the same ranking of methods in both experiments, which supports the validity of our automatic evaluation paradigm.
Alona Sydorova, Nina Pörner, Benjamin Roth 0001
ACL (1)3
2019 Neural architectures for open-type relation argument extraction
abstract
Abstract In this work, we focus on the task of open-type relation argument extraction (ORAE): given a corpus, a query entity Q, and a knowledge base relation (e.g., “Q authored notable work with title X”), the model has to extract an argument of non-standard entity type (entities that cannot be extracted by a standard named entity tagger, for example, X: the title of a book or a work of art) from the corpus. We develop and compare a wide range of neural models for this task yielding large improvements over a strong baseline obtained with a neural question answering system. The impact of different sentence encoding architectures and answer extraction methods is systematically compared. An encoder based on gated recurrent units combined with a conditional random fields tagger yields the best results. We release a data set to train and evaluate ORAE, based on Wikidata and obtained by distant supervision.
Benjamin Roth 0001, Costanza Conforti, Nina Pörner, Sanjeev Kumar Karn, Hinrich Schütze
Nat. Lang. Eng.1
2018 Evaluating neural network explanation methods using hybrid documents and morphosyntactic agreement
abstract
The behavior of deep neural networks (DNNs) is hard to understand.This makes it necessary to explore post hoc explanation methods.We conduct the first comprehensive evaluation of explanation methods for NLP.To this end, we design two novel evaluation paradigms that cover two important classes of NLP problems: small context and large context problems.Both paradigms require no manual annotation and are therefore broadly applicable.We also introduce LIMSSE, an explanation method inspired by LIME that is designed for NLP.We show empirically that LIMSSE, LRP and DeepLIFT are the most effective explanation methods and recommend them for explaining DNNs in NLP.
Nina Pörner, Hinrich Schütze, Benjamin Roth 0001
ACL (1)3
2018 Joint Aspect and Polarity Classification for Aspect-based Sentiment Analysis with End-to-End Neural Networks
abstract
In this work, we propose a new model for aspect-based sentiment analysis.In contrast to previous approaches, we jointly model the detection of aspects and the classification of their polarity in an end-to-end trainable neural network.We conduct experiments with different neural architectures and word representations on the recent GermEval 2017 dataset.We were able to show considerable performance gains by using the joint modeling approach in all settings compared to pipeline approaches.The combination of a convolutional neural network and fasttext embeddings outperformed the best submission of the shared task in 2017, establishing a new state of the art.
Martin Schmitt, Simon Steinheber, Konrad Schreiber, Benjamin Roth 0001
EMNLP4
2018 Joint Bootstrapping Machines for High Confidence Relation Extraction
abstract
Pankaj Gupta, Benjamin Roth, Hinrich Schütze. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Pankaj Gupta 0003, Benjamin Roth 0001, Hinrich Schütze
NAACL-HLT2
2017 Towards Bootstrapping a Polarity Shifter Lexicon using Linguistic Features
abstract
We present a major step towards the creation of the first high-coverage lexicon of polarity shifters. In this work, we bootstrap a lexicon of verbs by exploiting various linguistic features. Polarity shifters, such as “abandon”, are similar to negations (e.g. “not”) in that they move the polarity of a phrase towards its inverse, as in “abandon all hope”. While there exist lists of negation words, creating comprehensive lists of polarity shifters is far more challenging due to their sheer number. On a sample of manually annotated verbs we examine a variety of linguistic features for this task. Then we build a supervised classifier to increase coverage. We show that this approach drastically reduces the annotation effort while ensuring a high-precision lexicon. We also show that our acquired knowledge of verbal polarity shifters improves phrase-level sentiment analysis.
Marc Schulder, Michael Wiegand, Josef Ruppenhofer, Benjamin Roth 0001
IJCNLP(1)4
2016 Finding Relevant Relations in Relevant Documents
Michael Schuhmacher, Benjamin Roth 0001, Simone Paolo Ponzetto, Laura Dietz
ECIR2
2016 Comparing Convolutional Neural Networks to Traditional Models for Slot Filling
abstract
We address relation classification in the context of slot filling, the task of finding and evaluating fillers like "Steve Jobs" for the slot X in "X founded Apple".We propose a convolutional neural network which splits the input sentence into three parts according to the relation arguments and compare it to state-ofthe-art and traditional approaches of relation classification.Finally, we combine different methods and show that the combination is better than individual approaches.We also analyze the effect of genre differences on performance.
Heike Adel, Benjamin Roth 0001, Hinrich Schütze
HLT-NAACL2
2016 Multilingual Relation Extraction using Compositional Universal Schema
abstract
Patrick Verga, David Belanger, Emma Strubell, Benjamin Roth, Andrew McCallum. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Patrick Verga, David Belanger 0002, Emma Strubell, Benjamin Roth 0001, Andrew McCallum
HLT-NAACL4
2015 Compositional Vector Space Models for Knowledge Base Completion
abstract
Arvind Neelakantan, Benjamin Roth, Andrew McCallum. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Arvind Neelakantan, Benjamin Roth 0001, Andrew McCallum
ACL (1)2
2015 Combining Pattern-Based and Distributional Similarity for Graph-Based Noun Categorization
Michael Wiegand, Benjamin Roth 0001, Dietrich Klakow
NLDB2
2014 Unsupervised Parsing for Generating Surface-Based Relation Extraction Patterns
abstract
Finding the right features and patterns for identifying relations in natural language is one of the most pressing research questions for relation extraction.In this paper, we compare patterns based on supervised and unsupervised syntactic parsing and present a simple method for extracting surface patterns from a parsed training set.Results show that the use of surfacebased patterns not only increases extraction speed, but also improves the quality of the extracted relations.We find that, in this setting, unsupervised parsing, besides requiring less resources, compares favorably in terms of extraction quality.
Jens Illig, Benjamin Roth 0001, Dietrich Klakow
EACL2
2014 RelationFactory: A Fast, Modular and Effective System for Knowledge Base Population
abstract
Benjamin Roth, Tassilo Barth, Grzegorz Chrupała, Martin Gropp, Dietrich Klakow. Proceedings of the Demonstrations at the 14th Conference of the European Chapter of the Association for Computational Linguistics. 2014.
Benjamin Roth 0001, Tassilo Barth, Grzegorz Chrupala, Martin Gropp, Dietrich Klakow
EACL1
2014 Automatic Food Categorization from Large Unlabeled Corpora and Its Impact on Relation Extraction
abstract
We present a weakly-supervised induction method to assign semantic information to food items.We consider two tasks of categorizations being food-type classification and the distinction of whether a food item is composite or not.The categorizations are induced by a graph-based algorithm applied on a large unlabeled domain-specific corpus.We show that the usage of a domain-specific corpus is vital.We do not only outperform a manually designed open-domain ontology but also prove the usefulness of these categorizations in relation extraction, outperforming state-of-the-art features that include syntactic information and Brown clustering.
Michael Wiegand, Benjamin Roth 0001, Dietrich Klakow
EACL2
2013 Feature-based models for improving the quality of noisy training data for relation extraction
abstract
Supervised relation extraction from text relies on annotated data. Distant supervision is a scheme to obtain noisy training data by using a knowledge base of relational tuples as the ground truth and finding entity pair matches in a text corpus. We propose and evaluate two feature-based models for increasing the quality of distant supervision extraction patterns.
Benjamin Roth 0001, Dietrich Klakow
CIKM1
2013 Combining Generative and Discriminative Model Scores for Distant Supervision
abstract
Distant supervision is a scheme to generate noisy training data for relation extraction by aligning entities of a knowledge base with text.In this work we combine the output of a discriminative at-least-one learner with that of a generative hierarchical topic model to reduce the noise in distant supervision data.The combination significantly increases the ranking quality of extracted facts and achieves state-of-the-art extraction performance in an end-to-end setting.A simple linear interpolation of the model scores performs better than a parameter-free scheme based on nondominated sorting.
Benjamin Roth 0001, Dietrich Klakow
EMNLP1
2012 A Gold Standard for Relation Extraction in the Food Domain
Michael Wiegand, Benjamin Roth 0001, Eva Lasarcyk, Stephanie Köser, Dietrich Klakow
LREC2
2012 Web-Based Relation Extraction for the Food Domain
Michael Wiegand, Benjamin Roth 0001, Dietrich Klakow
NLDB2
2010 Topic Models for Word Sense Disambiguation and Token-Based Idiom Detection
Linlin Li 0001, Benjamin Roth 0001, Caroline Sporleder
ACL2
2010 Cross-language retrieval using link-based language models
abstract
We propose a cross-language retrieval model that is solely based on Wikipedia as a training corpus. The main contributions of our work are: 1. A translation model based on linked text in Wikipedia and a term weighting method associated with it. 2. A combination scheme to interpolate the link translation model with retrieval based on Latent Dirichlet Allocation. On the CLEF 2000 data we achieve improvement with respect to the best German-English system at the bilingual track (non-significant) and improvement against a baseline based on machine translation (significant).
Benjamin Roth 0001, Dietrich Klakow
SIGIR1