EDBT 2026 Demo / reviewers in the wild / expert
Alexander Fraser 0001
dblp:145/8377 · also Alexander M. Fraser
· DBLP profile ↗
67ranked-venue papers
6as first author
27since 2021 · last 2026
0000-0003-4891-682XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 65 · 6 first-author · 26 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Generic Responses: Target-Aware Strategies for Countering Hate Speech
Yen-Yu Chang, Daryna Dementieva, Alexander Fraser 0001 |
LREC | 3 |
| 2026 | Parallel Sentence Filtering for Low-Resource Language Pairs: A Case Study for Upper Sorbian, German, and Czech
Ruiyang Jiang, Shu Okabe, Alexander Fraser 0001 |
LREC | 3 |
| 2026 | Explainable Semantic Textual Similarity via Dissimilar Span DetectionabstractSemantic Textual Similarity (STS) is a crucial component of many Natural Language Processing (NLP) applications. However, existing approaches typically reduce semantic nuances to a single score, limiting interpretability. To address this, we introduce the task of Dissimilar Span Detection (DSD), which aims to identify semantically differing spans between pairs of texts. This can help users understand which particular words or tokens negatively affect the similarity score, or be used to improve performance in STS-dependent downstream tasks. Furthermore, we release a new dataset suitable for the task, the Span Similarity Dataset (SSD), developed through a semi-automated pipeline combining large language models (LLMs) with human verification. We propose and evaluate different baseline methods for DSD, both unsupervised, based on LIME, SHAP, LLMs, and our own method, as well as an additional supervised approach. While LLMs and supervised models achieve the highest performance, overall results remain low, highlighting the complexity of the task. Finally, we set up an additional experiment that shows how DSD can lead to increased performance in the specific task of paraphrase detection. Diego Miguel Lozano, Daryna Dementieva, Alexander Fraser 0001 |
LREC | 3 |
| 2026 | Literally Concrete or Figuratively Abstract? Multilingual Concreteness Norms for Verb-Object ExpressionsabstractAbstract While existing concreteness norms primarily target words in isolation, little attention has been paid to concreteness in context. To address this, we systematically collect multilingual concreteness ratings using Best-Worst Scaling (BWS) for 5,814 verb-direct object noun expressions in three languages with different degrees of resource availability: English, German, and Slovene. We identify consistent patterns where the concreteness of verb-noun combinations is more strongly influenced by the nominal object than the verb. Through comparative analyses on an English subset, we demonstrate that BWS guarantees more reliable concreteness judgments than traditional rating scales. Expanding beyond our human-generated data, we use traditional and LLM-based automatic extrapolation methods to generate a large-scale multilingual resource of over 430,000 expressions. Additionally, we conduct a study examining the interaction between concreteness and literal vs. figurative judgments for a subset of 1,800 expressions in all three languages, along with example usage sentences. Our findings show that lower concreteness ratings correlate with figurative language, thus reinforcing the link between abstractness and figurativeness. All resources are available from https://github.com/urbikn/multilingual-concreteness-vo. Urban Knuples, Diego Frassinelli, Alexander Fraser 0001, Sabine Schulte im Walde |
Trans. Assoc. Comput. Linguistics | 3 |
| 2025 | Multilingual Text-to-Image Generation Magnifies Gender StereotypesabstractText-to-image (T2I) generation models have achieved great results in image quality, flexibility, and text alignment, leading to widespread use. Through improvements in multilingual abilities, a larger community can access this technology. Yet, we show that multilingual models suffer from substantial gender bias. Furthermore, the expectation that results should be similar across languages does not hold. We introduce MAGBIG, a controlled benchmark designed to study gender bias in multilingual T2I models, and use it to assess the impact of multilingualism on gender bias. To this end, we construct a set of multilingual prompts that offers a carefully controlled setting accounting for the complex grammatical differences influencing gender across languages. Our results show strong gender biases and notable language-specific differences across models. While we explore prompt engineering strategies to mitigate these biases, we find them largely ineffective and sometimes even detrimental to text-to-image alignment. Our analysis highlights the need for research on diverse language representations and greater control over bias in T2I models. Felix Friedrich, Katharina Hämmerl, Patrick Schramowski, Manuel Brack, Jindrich Libovický, Alexander Fraser 0001, Kristian Kersting |
ACL (1) | 6 |
| 2025 | Positional Overload: Positional Debiasing and Context Window Extension for Large Language Models using Set EncodingabstractLarge Language Models (LLMs) typically track the order of tokens using positional encoding, which causes the following problems: positional bias, where the model is influenced by an ordering within the prompt, and a fixed context window, as models struggle to generalize to positions beyond those encountered during training. To address these limitations, we developed a novel method called \textit{set encoding}. This method allows multiple pieces of text to be encoded in the same position, thereby eliminating positional bias entirely. Another promising use case for set encoding is to increase the size of the input an LLM can handle. Our experiments demonstrate that set encoding allows an LLM to solve tasks with far more tokens than without set encoding. To our knowledge, set encoding is the first technique to effectively extend an LLM’s context window without requiring any additional training. Lukas Kinder, Lukas Edman, Alexander Fraser 0001, Tobias Käfer |
ACL (1) | 3 |
| 2025 | LLM Sensitivity Challenges in Abusive Language Detection: Instruction-Tuned vs. Human FeedbackabstractThe capacity of large language models (LLMs) to understand and distinguish socially unacceptable texts enables them to play a promising role in abusive language detection. However, various factors can affect their sensitivity. In this work, we test whether LLMs have an unintended bias in abusive language detection, i.e., whether they predict more or less of a given abusive class than expected in zero-shot settings. Our results show that instruction-tuned LLMs tend to under-predict positive classes, since datasets used for tuning are dominated by the negative class. On the contrary, models fine-tuned with human feedback tend to be overly sensitive. In an exploratory approach to mitigate these issues, we show that label frequency in the prompt helps with the significant over-prediction. Viktor Hangya, Alexander Fraser 0001 |
COLING | 3 |
| 2025 | Data-Efficient Hate Speech Detection via Cross-Lingual Nearest Neighbor Retrieval with Limited Labeled DataabstractConsidering the importance of detecting hateful content, labeled hate speech data is expensive and time-consuming to collect and annotate, particularly for low-resource languages.Prior work has demonstrated the effectiveness of cross-lingual transfer learning and data augmentation in improving performance on tasks with limited labeled data.To develop an efficient and scalable cross-lingual transfer learning approach, we leverage nearest-neighbor retrieval to augment minimal labeled data in the target language, thereby enhancing detection performance.Specifically, we assume access to a small set of labeled training instances in the target language and use these to retrieve the most relevant labeled examples from a large multilingual hate speech detection pool.We evaluate our approach on eight languages and demonstrate that it consistently outperforms models trained solely on the target language data.Furthermore, in most cases, our method surpasses the current state-of-the-art.Notably, our approach is highly data-efficient, retrieving as few as 200 instances in some cases while maintaining superior performance.Moreover, it is scalable, as the retrieval pool can be easily expanded, and the method can be readily adapted to new languages and tasks.We also apply maximum marginal relevance to mitigate redundancy and filter out highly similar retrieved instances, resulting in improvements in some languages. Faeze Ghorbanpour, Daryna Dementieva, Alexander Fraser 0001 |
EMNLP | 3 |
| 2025 | From Unaligned to Aligned: Scaling Multilingual LLMs with Multi-Way Parallel CorporaabstractContinued pretraining and instruction tuning on large-scale multilingual data have proven to be effective in scaling large language models (LLMs) to low-resource languages.However, the unaligned nature of such data limits its ability to effectively capture cross-lingual semantics.In contrast, multi-way parallel data, where identical content is aligned across multiple languages, provides stronger cross-lingual consistency and offers greater potential for improving multilingual performance.In this paper, we introduce a large-scale, high-quality multiway parallel corpus, TED2025, based on TED Talks.The corpus spans 113 languages, with up to 50 languages aligned in parallel, ensuring extensive multilingual coverage.Using this dataset, we investigate best practices for leveraging multi-way parallel data to enhance LLMs, including strategies for continued pretraining, instruction tuning, and the analysis of key influencing factors.Experiments on six multilingual benchmarks show that models trained on multiway parallel data consistently outperform those trained on unaligned multilingual data. Yingli Shen, Wen Lai, Shuo Wang 0013, Kangyang Luo, Alexander Fraser 0001, Maosong Sun 0001 |
EMNLP | 6 |
| 2025 | Extracting Linguistic Information from Large Language Models: Syntactic Relations and Derivational KnowledgeabstractThis paper presents a study of the linguistic knowledge and generalization capabilities of Large Language Models (LLMs), focusing on their morphosyntactic competence.We design three diagnostic tasks: (i) labeling syntactic information at the sentence level -identifying subjects, objects, and indirect objects; (ii) derivational decomposition at the word level -identifying morpheme boundaries and labeling the decomposed sequence; and (iii) in-depth study of morphological decomposition in German and Amharic.We evaluate prompting strategies in GPT-4o and LLaMA 3.3-70B to extract different types of linguistic structures for typologically diverse languages.Our results show that GPT-4o consistently outperforms LLaMA in all tasks; however, both models exhibit limitations and show little evidence of abstract morphological rule learning.Importantly, we show strong evidence that the models fail to learn underlying morphological structures.Therefore, raising important doubts about their ability to generalize. Tsedeniya Kinfe Temesgen, Marion Di Marco, Alexander Fraser 0001 |
EMNLP | 3 |
| 2025 | Joint Localization and Activation Editing for Low-Resource Fine-TuningabstractParameter-efficient fine-tuning (PEFT) methods, such as LoRA, are commonly used to adapt LLMs. However, the effectiveness of standard PEFT methods is limited in low-resource scenarios with only a few hundred examples. Recent advances in interpretability research have inspired the emergence of activation editing (or steering) techniques, which modify the activations of specific model components. Due to their extremely small parameter counts, these methods show promise for small datasets. However, their performance is highly dependent on identifying the correct modules to edit and often lacks stability across different datasets. In this paper, we propose Joint Localization and Activation Editing (JoLA), a method that jointly learns (1) which heads in the Transformer to edit (2) whether the intervention should be additive, multiplicative, or both and (3) the intervention parameters themselves - the vectors applied as additive offsets or multiplicative scalings to the head output. Through evaluations on three benchmarks spanning commonsense reasoning, natural language understanding, and natural language generation, we demonstrate that JoLA consistently outperforms existing methods. Wen Lai, Alexander Fraser 0001, Ivan Titov 0001 |
ICML | 2 |
| 2025 | Fine-Grained Transfer Learning for Harmful Content Detection through Label-Specific Soft Prompt TuningabstractFaeze Ghorbanpour, Viktor Hangya, Alexander Fraser. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Faeze Ghorbanpour, Viktor Hangya, Alexander Fraser 0001 |
NAACL (Long Papers) | 3 |
| 2025 | DCAD-2000: A Multilingual Dataset across 2000+ Languages with Data Cleaning as Anomaly DetectionabstractThe rapid development of multilingual large language models (LLMs) highlights the need for high-quality, diverse, and well-curated multilingual datasets. In this paper, we introduce DCAD-2000 (Data Cleaning as Anomaly Detection), a large-scale multilingual corpus constructed from newly extracted Common Crawl data and existing multilingual sources. DCAD-2000 covers 2,282 languages, 46.72TB of text, and 8.63 billion documents, spanning 155 high- and medium-resource languages and 159 writing scripts. To overcome the limitations of existing data cleaning approaches, which rely on manually designed heuristic thresholds, we reframe data cleaning as an anomaly detection problem. This dynamic filtering paradigm substantially improves data quality by automatically identifying and removing noisy or anomalous content. By fine-tuning LLMs on DCAD-2000, we demonstrate notable improvements in data quality, robustness of the cleaning pipeline, and downstream performance, particularly for low-resource languages across multiple multilingual benchmarks. Yingli Shen, Wen Lai, Shuo Wang 0013, Xueren Zhang, Kangyang Luo, Alexander Fraser 0001, Maosong Sun 0001 |
NeurIPS | 6 |
| 2024 | How to Solve Few-Shot Abusive Content Detection Using the Data We Actually HaveabstractDue to the broad range of social media platforms, the requirements of abusive language detection systems are varied and ever-changing. Already a large set of annotated corpora with different properties and label sets were created, such as hate or misogyny detection, but the form and targets of abusive speech are constantly evolving. Since, the annotation of new corpora is expensive, in this work we leverage datasets we already have, covering a wide range of tasks related to abusive language detection. Our goal is to build models cheaply for a new target label set and/or language, using only a few training examples of the target domain. We propose a two-step approach: first we train our model in a multitask fashion. We then carry out few-shot adaptation to the target requirements. Our experiments show that using already existing datasets and only a few-shots of the target task the performance of models improve both monolingually and across languages. Our analysis also shows that our models acquire a general understanding of abusive language, since they improve the prediction of labels which are present only in the target dataset and can benefit from knowledge about labels which are not directly used for the target task. Viktor Hangya, Alexander Fraser 0001 |
LREC/COLING | 2 |
| 2024 | Analyzing the Understanding of Morphologically Complex Words in Large Language ModelsabstractWe empirically study the ability of a Large Language Model (gpt-3.5-turbo-instruct) to understand morphologically complex words. In our experiments, we looked at a variety of tasks to analyse German compounds with regard to compositional word formation and derivation, such as identifying the head noun of existing and novel compounds, identifying the shared verb stem between two words, or recognizing words constructed with inappropriately used derivation morphemes as invalid. Our results show that the language model is generally capable of solving most tasks, except for the task of identifying ill-formed word forms. While the model demonstrated a good overall understanding of complex words and their word-internal structure, the results also suggest that there is no formal knowledge of derivational rules, but rather an interpretation of the observed word parts to derive the meaning of a word. Marion Di Marco, Alexander Fraser 0001 |
LREC/COLING | 2 |
| 2024 | Reconstruction of Cuneiform Literary Texts as Text MatchingabstractAncient Mesopotamian literature is riddled with gaps, caused by the decay and fragmentation of its writing material, clay tablets. The discovery of overlaps between fragments allows reconstruction to advance, but it is a slow and unsystematic process. Since new pieces are found and digitized constantly, NLP techniques can help to identify fragments and match them with existing text collections to restore complete literary works. We compare a number of approaches and determine that a character-level n-gram-based similarity matching approach works well for this problem, leading to a large speed-up for researchers in Assyriology. Fabian Simonjetz, Jussi Laasonen, Yunus Cobanoglu, Alexander Fraser 0001, Enrique Jiménez |
LREC/COLING | 4 |
| 2024 | CUTE: Measuring LLMs' Understanding of Their TokensabstractLarge Language Models (LLMs) show remarkable performance on a wide variety of tasks.Most LLMs split text into multi-character tokens and process them as atomic units without direct access to individual characters.This raises the question: To what extent can LLMs learn orthographic information?To answer this, we propose a new benchmark, CUTE, which features a collection of tasks designed to test the orthographic knowledge of LLMs.We evaluate popular LLMs on CUTE, finding that most of them seem to know the spelling of their tokens, yet fail to use this information effectively to manipulate text, calling into question how much of this knowledge is generalizable. Lukas Edman, Helmut Schmid, Alexander Fraser 0001 |
EMNLP | 3 |
| 2024 | Style-Specific Neurons for Steering LLMs in Text Style TransferabstractText style transfer (TST) aims to modify the style of a text without altering its original meaning.Large language models (LLMs) demonstrate superior performance across multiple tasks, including TST.However, in zero-shot setups, they tend to directly copy a significant portion of the input text to the output without effectively changing its style.To enhance the stylistic variety and fluency of the text, we present sNeuron-TST, a novel approach for steering LLMs using style-specific neurons in TST.Specifically, we identify neurons associated with the source and target styles and deactivate source-style-only neurons to give target-style words a higher probability, aiming to enhance the stylistic diversity of the generated text.However, we find that this deactivation negatively impacts the fluency of the generated text, which we address by proposing an improved contrastive decoding method that accounts for rapid token probability shifts across layers caused by deactivated source-style neurons.Empirical experiments demonstrate the effectiveness of the proposed method on six benchmarks, encompassing formality, toxicity, politics, politeness, authorship, and sentiment 1 . Wen Lai, Viktor Hangya, Alexander Fraser 0001 |
EMNLP | 3 |
| 2024 | Subword Segmentation in LLMs: Looking at Inflection and ConsistencyabstractThe role of subword segmentation in relation to capturing morphological patterns in LLMs is currently not well explored.Ideally, one would train large models like GPT using various segmentations and evaluate how well word meanings are captured.Since this is not computationally feasible, we group words according to their segmentation properties and compare how well a model can solve a linguistic task for these groups.We study two criteria: (i) adherence to morpheme boundaries and (ii) the segmentation consistency of the different inflected forms of a lemma.We select word forms with high and low values for these criteria and carry out experiments on GPT-4o's ability to capture verbal inflection for 10 languages.Our results indicate that in particular the criterion of segmentation consistency can help to predict the model's ability to recognize and generate the lemma from an inflected form, providing evidence that subword segmentation is relevant. Marion Di Marco, Alexander Fraser 0001 |
EMNLP | 2 |
| 2024 | Hate Personified: Investigating the role of LLMs in content moderationabstractFor subjective tasks such as hate detection, where people perceive hate differently, the Large Language Model's (LLM) ability to represent diverse groups is unclear.By including additional context in prompts, we comprehensively analyze LLM's sensitivity to geographical priming, persona attributes, and numerical information to assess how well the needs of various groups are reflected.Our findings on two LLMs, five languages, and six datasets reveal that mimicking persona-based attributes leads to annotation variability.Meanwhile, incorporating geographical signals leads to better regional alignment.We also find that the LLMs are sensitive to numerical anchors, indicating the ability to leverage community-based flagging efforts and exposure to adversaries.Our work provides preliminary guidelines and highlights the nuances of applying LLMs in culturally sensitive cases. 1 Sarah Masud, Sahajpreet Singh, Viktor Hangya, Alexander Fraser 0001, Tanmoy Chakraborty 0002 |
EMNLP | 4 |
| 2023 | A Survey of Methods for Addressing Class Imbalance in Deep-Learning Based Natural Language ProcessingabstractMany natural language processing (NLP) tasks are naturally imbalanced, as some target categories occur much more frequently than others in the real world.In such scenarios, current NLP models tend to perform poorly on less frequent classes.Addressing class imbalance in NLP is an active research topic, yet, finding a good approach for a particular task and imbalance scenario is difficult.In this survey, the first overview on class imbalance in deep-learning based NLP, we first discuss various types of controlled and realworld class imbalance.Our survey then covers approaches that have been explicitly proposed for class-imbalanced NLP tasks or, originating in the computer vision community, have been evaluated on them.We organize the methods by whether they are based on sampling, data augmentation, choice of loss function, staged learning, or model design.Finally, we discuss open problems and how to move forward. Sophie Henning, William Beluch, Alexander Fraser 0001, Annemarie Friedrich |
EACL | 3 |
| 2023 | A Study on Accessing Linguistic Information in Pre-Trained Language Models by Using PromptsabstractWe study whether linguistic information in pretrained multilingual language models can be accessed by human language: So far, there is no easy method to directly obtain linguistic information and gain insights into the linguistic principles encoded in such models.We use the technique of prompting and formulate linguistic tasks to test the LM's access to explicit grammatical principles and study how effective this method is at providing access to linguistic features.Our experiments on German, Icelandic and Spanish show that some linguistic properties can in fact be accessed through prompting, whereas others are harder to capture. Marion Di Marco, Katharina Hämmerl, Alexander Fraser 0001 |
EMNLP | 3 |
| 2022 | Improving Both Domain Robustness and Domain Adaptability in Machine TranslationabstractWe consider two problems of NMT domain adaptation using meta-learning. First, we want to reach domain robustness, i.e., we want to reach high quality on both domains seen in the training data and unseen domains. Second, we want our systems to be adaptive, i.e., making it possible to finetune systems with just hundreds of in-domain parallel sentences. We study the domain adaptability of meta-learning when improving the domain robustness of the model. In this paper, we propose a novel approach, RMLNMT (Robust Meta-Learning Framework for Neural Machine Translation Domain Adaptation), which improves the robustness of existing meta-learning models. More specifically, we show how to use a domain classifier in curriculum learning and we integrate the word-level domain mixing model into the meta-learning framework with a balanced sampling strategy. Experiments on English-German and English-Chinese translation show that RMLNMT improves in terms of both domain robustness and domain adaptability in seen and unseen domains. Wen Lai, Jindrich Libovický, Alexander Fraser 0001 |
COLING | 3 |
| 2022 | Improving Low-Resource Languages in Pre-Trained Multilingual Language ModelsabstractPre-trained multilingual language models are the foundation of many NLP approaches, including cross-lingual transfer solutions.However, languages with small available monolingual corpora are often not well-supported by these models leading to poor performance.We propose an unsupervised approach to improve the cross-lingual representations of lowresource languages by bootstrapping word translation pairs from monolingual corpora and using them to improve language alignment in pretrained language models.We perform experiments on nine languages, using contextual word retrieval and zero-shot named entity recognition to measure both intrinsic cross-lingual word representation quality and downstream task performance, showing improvements on both tasks.Our results show that it is possible to improve pre-trained multilingual language models by relying only on non-parallel resources. Viktor Hangya, Hossain Shaikh Saadi, Alexander Fraser 0001 |
EMNLP | 3 |
| 2022 | Demonstrating CAT: Synthesizing Data-Aware Conversational Agents for Transactional DatabasesabstractDatabases for OLTP are often the backbone for applications such as hotel room or cinema ticket booking applications. However, developing a conversational agent (i.e., a chatbot-like interface) to allow end-users to interact with an application using natural language requires both immense amounts of training data and NLP expertise. This motivates CAT , which can be used to easily create conversational agents for transactional databases. The main idea is that, for a given OLTP database, CAT uses weak supervision to synthesize the required training data to train a state-of-the-art conversational agent, allowing users to interact with the OLTP database. Furthermore, CAT provides an out-of-the-box integration of the resulting agent with the database. As a major difference to existing conversational agents, agents synthesized by CAT are data-aware. This means that the agent decides which information should be requested from the user based on the current data distributions in the database, which typically results in markedly more efficient dialogues compared with non-data-aware agents. We publish the code for CAT as open source. Marius Gassen, Benjamin Hättasch, Benjamin Hilprecht, Nadja Geisler, Alexander Fraser 0001, Carsten Binnig |
Proc. VLDB Endow. | 5 |
| 2021 | A Comparison of Sentence-Weighting Techniques for NMTabstractSentence weighting is a simple and powerful domain adaptation technique. We carry out domain classification for computing sentence weights with 1) language model cross entropy difference 2) a convolutional neural network 3) a Recursive Neural Tensor Network. We compare these approaches with regard to domain classification accuracy and and study the posterior probability distributions. Then we carry out NMT experiments in the scenario where we have no in-domain parallel corpora and and only very limited in-domain monolingual corpora. Here and we use the domain classifier to reweight the sentences of our out-of-domain training corpus. This leads to improvements of up to 2.1 BLEU for German to English translation. Simon Riess, Matthias Huck, Alexander Fraser 0001 |
MTSummit (1) | 3 |
| 2021 | Improving the Lexical Ability of Pretrained Language Models for Unsupervised Neural Machine TranslationabstractAlexandra Chronopoulou, Dario Stojanovski, Alexander Fraser. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Alexandra Chronopoulou, Dario Stojanovski, Alexander Fraser 0001 |
NAACL-HLT | 3 |
| 2020 | Modeling Word Formation in English-German Neural Machine TranslationabstractThis paper studies strategies to model word formation in NMT using rich linguistic information, namely a word segmentation approach that goes beyond splitting into substrings by considering fusional morphology.Our linguistically sound segmentation is combined with a method for target-side inflection to accommodate modeling word formation.The best system variants employ source-side morphological analysis and model complex target-side words, improving over a standard system. Marion Di Marco, Alexander Fraser 0001 |
ACL | 2 |
| 2020 | Combining Word Embeddings with Bilingual Orthography Embeddings for Bilingual Dictionary InductionabstractBilingual dictionary induction (BDI) is the task of accurately translating words to the target language.It is of great importance in many low-resource scenarios where cross-lingual training data is not available.To perform BDI, bilingual word embeddings (BWEs) are often used due to their low bilingual training signal requirements.They achieve high performance, but problematic cases still remain, such as the translation of rare words or named entities, which often need to be transliterated.In this paper, we enrich BWE-based BDI with transliteration information by using Bilingual Orthography Embeddings (BOEs).BOEs represent source and target language transliteration word pairs with similar vectors.A key problem in our BDI setup is to decide which information source -BWEs (or semantics) vs. BOEs (or orthography) -is more reliable for a particular word pair.We propose a novel classification-based BDI system that uses BWEs, BOEs and a number of other features to make this decision.We test our system on English-Russian BDI and show improved performance.In addition, we show the effectiveness of our BOEs by successfully using them for transliteration mining based on cosine similarity. Silvia Severini, Viktor Hangya, Alexander Fraser 0001, Hinrich Schütze |
COLING | 3 |
| 2020 | ContraCAT: Contrastive Coreference Analytical Templates for Machine TranslationabstractRecent high scores on pronoun translation using context-aware neural machine translation have suggested that current approaches work well.ContraPro is a notable example of a contrastive challenge set for English→German pronoun translation.The high scores achieved by transformer models may suggest that they are able to effectively model the complicated set of inferences required to carry out pronoun translation.This entails the ability to determine which entities could be referred to, identify which entity a sourcelanguage pronoun refers to (if any), and access the target-language grammatical gender for that entity.We first show through a series of targeted adversarial attacks that in fact current approaches are not able to model all of this information well.Inserting small amounts of distracting information is enough to strongly reduce scores, which should not be the case.We then create a new template test set Contracat, designed to individually assess the ability to handle the specific steps necessary for successful pronoun translation.Our analyses show that current approaches to context-aware nmt rely on a set of surface heuristics, which break down when translations require real reasoning.We also propose an approach for augmenting the training data, with some improvements. Dario Stojanovski, Benno Krojer, Denis Peskov, Alexander Fraser 0001 |
COLING | 4 |
| 2020 | Reusing a Pretrained Language Model on Languages with Limited Corpora for Unsupervised NMTabstractUsing a language model (LM) pretrained on two languages with large monolingual data in order to initialize an unsupervised neural machine translation (UNMT) system yields stateof-the-art results.When limited data is available for one language, however, this method leads to poor translations.We present an effective approach that reuses an LM that is pretrained only on a high-resource language.The monolingual LM is fine-tuned on both languages and is then used to initialize a UNMT model.To reuse the pretrained LM, we have to modify its predefined vocabulary, to account for the new language.We therefore propose a novel vocabulary extension method.Our approach, RE-LM, outperforms a competitive cross-lingual pretraining model (XLM) in English-Macedonian (En-Mk) and English-Albanian (En-Sq), yielding more than +8.3 BLEU points for all four translation directions. Alexandra Chronopoulou, Dario Stojanovski, Alexander Fraser 0001 |
EMNLP (1) | 3 |
| 2020 | Towards Reasonably-Sized Character-Level Transformer NMT by Finetuning Subword SystemsabstractApplying the Transformer architecture on the character level usually requires very deep architectures that are difficult and slow to train.These problems can be partially overcome by incorporating a segmentation into tokens in the model.We show that by initially training a subword model and then finetuning it on characters, we can obtain a neural machine translation model that works at the character level without requiring token segmentation.We use only the vanilla 6-layer Transformer Base architecture.Our character-level models better capture morphological phenomena and show more robustness to noise at the expense of somewhat worse overall translation quality.Our study is a significant step towards highperformance and easy to train character-based models that are not extremely large. Jindrich Libovický, Alexander Fraser 0001 |
EMNLP (1) | 2 |
| 2020 | Exploring Bilingual Word Embeddings for Hiligaynon, a Low-Resource LanguageabstractThis paper investigates the use of bilingual word embeddings for mining Hiligaynon translations of English words. There is very little research on Hiligaynon, an extremely low-resource language of Malayo-Polynesian origin with over 9 million speakers in the Philippines (we found just one paper). We use a publicly available Hiligaynon corpus with only 300K words, and match it with a comparable corpus in English. As there are no bilingual resources available, we manually develop a English-Hiligaynon lexicon and use this to train bilingual word embeddings. But we fail to mine accurate translations due to the small amount of data. To find out if the same holds true for a related language pair, we simulate the same low-resource setup on English to German and arrive at similar results. We then vary the size of the comparable English and German corpora to determine the minimum corpus size necessary to achieve competitive results. Further, we investigate the role of the seed lexicon. We show that with the same corpus size but with a smaller seed lexicon, performance can surpass results of previous studies. We release the lexicon of 1,200 English-Hiligaynon word pairs we created to encourage further investigation. Leah Michel, Viktor Hangya, Alexander Fraser 0001 |
LREC | 3 |
| 2019 | Unsupervised Parallel Sentence Extraction with Parallel Segment Detection Helps Machine TranslationabstractMining parallel sentences from comparable corpora is important.Most previous work relies on supervised systems, which are trained on parallel data, thus their applicability is problematic in low-resource scenarios.Recent developments in building unsupervised bilingual word embeddings made it possible to mine parallel sentences based on cosine similarities of source and target language words.We show that relying only on this information is not enough, since sentences often have similar words but different meanings.We detect continuous parallel segments in sentence pair candidates and rely on them when mining parallel sentences.We show better mining accuracy on three language pairs in a standard shared task on artificial data.We also provide the first experiments showing that parallel sentences mined from real life sources improve unsupervised MT.Our code is available, we hope it will be used to support low-resource MT research. Viktor Hangya, Alexander Fraser 0001 |
ACL (1) | 2 |
| 2019 | Better OOV Translation with Bilingual Terminology MiningabstractUnseen words, also called out-of-vocabulary words (OOVs), are difficult for machine translation.In neural machine translation, byte-pair encoding can be used to represent OOVs, but they are still often incorrectly translated.We improve the translation of OOVs in NMT using easy-to-obtain monolingual data.We look for OOVs in the text to be translated and translate them using simple-to-construct bilingual word embeddings (BWEs).In our MT experiments we take the 5-best candidates, which is motivated by intrinsic mining experiments.Using all five of the proposed target language words as queries we mine target-language sentences.We then back-translate, forcing the back-translation of each of the five proposed target-language OOV-translation-candidates to be the original source-language OOV.We show that by using this synthetic data to finetune our system the translation of OOVs can be dramatically improved.In our experiments we use a system trained on Europarl and mine sentences containing medical terms from monolingual data. Matthias Huck, Viktor Hangya, Alexander Fraser 0001 |
ACL (1) | 3 |
| 2019 | Improving Anaphora Resolution in Neural Machine Translation Using Curriculum Learning
Dario Stojanovski, Alexander Fraser 0001 |
MTSummit (1) | 2 |
| 2018 | Two Methods for Domain Adaptation of Bilingual Tasks: Delightfully Simple and Broadly ApplicableabstractBilingual tasks, such as bilingual lexicon induction and cross-lingual classification, are crucial for overcoming data sparsity in the target language.Resources required for such tasks are often out-of-domain, thus domain adaptation is an important problem here.We make two contributions.First, we test a delightfully simple method for domain adaptation of bilingual word embeddings.We evaluate these embeddings on two bilingual tasks involving different domains: cross-lingual twitter sentiment classification and medical bilingual lexicon induction.Second, we tailor a broadly applicable semi-supervised classification method from computer vision to these tasks.We show that this method also helps in low-resource setups.Using both methods together we achieve large improvements over our baselines, by using only additional unlabeled data. Viktor Hangya, Fabienne Braune, Alexander Fraser 0001, Hinrich Schütze |
ACL (1) | 3 |
| 2018 | Embedding Learning Through Multilingual Concept InductionabstractWe present a new method for estimating vector space representations of words: embedding learning by concept induction.We test this method on a highly parallel corpus and learn semantic representations of words in 1259 different languages in a single common space.An extensive experimental evaluation on crosslingual word similarity and sentiment analysis indicates that concept-based multilingual embedding learning performs better than previous approaches. Philipp Dufter, Martin Schmitt, Alexander Fraser 0001, Hinrich Schütze |
ACL (1) | 4 |
| 2017 | Statistical Models for Unsupervised, Semi-Supervised, and Supervised Transliteration MiningabstractWe present a generative model that efficiently mines transliteration pairs in a consistent fashion in three different settings: unsupervised, semi-supervised, and supervised transliteration mining. The model interpolates two sub-models, one for the generation of transliteration pairs and one for the generation of non-transliteration pairs (i.e., noise). The model is trained on noisy unlabeled data using the EM algorithm. During training the transliteration sub-model learns to generate transliteration pairs and the fixed non-transliteration model generates the noise pairs. After training, the unlabeled data is disambiguated based on the posterior probabilities of the two sub-models. We evaluate our transliteration mining system on data from a transliteration mining shared task and on parallel corpora. For three out of four language pairs, our system outperforms all semi-supervised and supervised systems that participated in the NEWS 2010 shared task. On word pairs extracted from parallel corpora with fewer than 2% transliteration pairs, our system achieves up to 86.7% F-measure with 77.9% precision and 97.8% recall. Hassan Sajjad 0001, Helmut Schmid, Alexander Fraser 0001, Hinrich Schütze |
Comput. Linguistics | 3 |
| 2016 | Target-Side Context for Discriminative Models in Statistical Machine TranslationabstractDiscriminative translation models utilizing source context have been shown to help statistical machine translation performance.We propose a novel extension of this work using target context information.Surprisingly, we show that this model can be efficiently integrated directly in the decoding process.Our approach scales to large training data sizes and results in consistent improvements in translation quality on four language pairs.We also provide an analysis comparing the strengths of the baseline source-context model with our extended source-context and targetcontext model and we show that our extension allows us to better capture morphological coherence.Our work is freely available as part of Moses. Ales Tamchyna, Alexander Fraser 0001, Ondrej Bojar, Marcin Junczys-Dowmunt |
ACL (1) | 2 |
| 2015 | Labeled Morphological Segmentation with Semi-Markov ModelsabstractWe present labeled morphological segmentation-an alternative view of morphological processing that unifies several tasks.We introduce a new hierarchy of morphotactic tagsets and CHIPMUNK, a discriminative morphological segmentation system that, contrary to previous work, explicitly models morphotactics.We show improved performance on three tasks for all six languages: (i) morphological segmentation, (ii) stemming and (iii) morphological tag classification.For morphological segmentation our method shows absolute improvements of 2-6 points F 1 over a strong baseline. Ryan Cotterell, Thomas Müller 0009, Alexander Fraser 0001, Hinrich Schütze |
CoNLL | 3 |
| 2015 | Target-Side Generation of Prepositions for SMT
Marion Di Marco, Alexander Fraser 0001, Sabine Schulte im Walde |
EAMT | 2 |
| 2015 | Joint Lemmatization and Morphological Tagging with LemmingabstractWe present LEMMING, a modular loglinear model that jointly models lemmatization and tagging and supports the integration of arbitrary global features.It is trainable on corpora annotated with gold standard tags and lemmata and does not rely on morphological dictionaries or analyzers.LEMMING sets the new state of the art in token-based statistical lemmatization on six languages; e.g., for Czech lemmatization, we reduce the error by 60%, from 4.05 to 1.58.We also give empirical evidence that jointly modeling morphological tags and lemmata is mutually beneficial. Thomas Müller 0009, Ryan Cotterell, Alexander Fraser 0001, Hinrich Schütze |
EMNLP | 3 |
| 2015 | Rule Selection with Soft Syntactic Features for String-to-Tree Statistical Machine TranslationabstractIn syntax-based machine translation, rule selection is the task of choosing the correct target side of a translation rule among rules with the same source side.We define a discriminative rule selection model for systems that have syntactic annotation on the target language side (stringto-tree).This is a new and clean way to integrate soft source syntactic constraints into string-to-tree systems as features of the rule selection model.We release our implementation as part of Moses. Fabienne Braune, Nina Seemann, Alexander Fraser 0001 |
EMNLP | 3 |
| 2015 | The Operation Sequence Model - Combining N-Gram-Based and Phrase-Based Statistical Machine TranslationabstractIn this article, we present a novel machine translation model, the Operation Sequence Model (OSM), which combines the benefits of phrase-based and N-gram-based statistical machine translation (SMT) and remedies their drawbacks. The model represents the translation process as a linear sequence of operations. The sequence includes not only translation operations but also reordering operations. As in N-gram-based SMT, the model is: (i) based on minimal translation units, (ii) takes both source and target information into account, (iii) does not make a phrasal independence assumption, and (iv) avoids the spurious phrasal segmentation problem. As in phrase-based SMT, the model (i) has the ability to memorize lexical reordering triggers, (ii) builds the search graph dynamically, and (iii) decodes with large translation units during search. The unique properties of the model are (i) its strong coupling of reordering and translation where translation and reordering decisions are conditioned on n previous translation and reordering decisions, and (ii) the ability to model local and long-range reorderings consistently. Using BLEU as a metric of translation accuracy, we found that our system performs significantly better than state-of-the-art phrase-based systems (Moses and Phrasal) and N-gram-based systems (Ncode) on standard translation tasks. We compare the reordering component of the OSM to the Moses lexical reordering model by integrating it into Moses. Our results show that OSM outperforms lexicalized reordering on all translation tasks. The translation quality is shown to be improved further by learning generalized representations with a POS-based OSM. Nadir Durrani, Helmut Schmid, Alexander Fraser 0001, Philipp Koehn, Hinrich Schütze |
Comput. Linguistics | 3 |
| 2014 | Investigating the Usefulness of Generalized Word Representations in SMT
Nadir Durrani, Philipp Koehn, Helmut Schmid, Alexander Fraser 0001 |
COLING | 4 |
| 2014 | How to Produce Unseen Teddy Bears: Improved Morphological Processing of Compounds in SMTabstractCompounding in morphologically rich languages is a highly productive process which often causes SMT approaches to fail because of unseen words.We present an approach for translation into a compounding language that splits compounds into simple words for training and, due to an underspecified representation, allows for free merging of simple words into compounds after translation.In contrast to previous approaches, we use features projected from the source language to predict compound mergings.We integrate our approach into end-to-end SMT and show that many compounds matching the reference translation are produced which did not appear in the training data.Additional manual evaluations support the usefulness of generalizing compound formation in SMT. Fabienne Cap, Alexander Fraser 0001, Marion Di Marco, Aoife Cahill |
EACL | 2 |
| 2014 | Combining bilingual terminology mining and morphological modeling for domain adaptation in SMT
Marion Di Marco, Alexander Fraser 0001, Ulrich Heid |
EAMT | 2 |
| 2013 | Using subcategorization knowledge to improve case prediction for translation to German
Marion Di Marco, Alexander Fraser 0001, Sabine Schulte im Walde |
ACL (1) | 2 |
| 2013 | Model With Minimal Translation Units, But Decode With Phrases
Nadir Durrani, Alexander Fraser 0001, Helmut Schmid |
HLT-NAACL | 2 |
| 2013 | Knowledge Sources for Constituent Parsing of German, a Morphologically Rich and Less-Configurational LanguageabstractWe study constituent parsing of German, a morphologically rich and less-configurational language. We use a probabilistic context-free grammar treebank grammar that has been adapted to the morphologically rich properties of German by markovization and special features added to its productions. We evaluate the impact of adding lexical knowledge. Then we examine both monolingual and bilingual approaches to parse reranking. Our reranking parser is the new state of the art in constituency parsing of the TIGER Treebank. We perform an analysis, concluding with lessons learned, which apply to parsing other morphologically rich and less-configurational languages. Alexander Fraser 0001, Helmut Schmid, Richárd Farkas, Renjing Wang, Hinrich Schütze |
Comput. Linguistics | 1 |
| 2012 | A Statistical Model for Unsupervised and Semi-supervised Transliteration Mining
Hassan Sajjad 0001, Alexander Fraser 0001, Helmut Schmid |
ACL (1) | 2 |
| 2012 | Modeling Inflection and Word-Formation in SMT
Alexander Fraser 0001, Marion Di Marco, Aoife Cahill, Fabienne Cap |
EACL | 1 |
| 2012 | Determining the placement of German verbs in English-to-German SMT
Anita Gojun, Alexander Fraser 0001 |
EACL | 2 |
| 2012 | Long-distance reordering during search for hierarchical phrase-based SMT
Fabienne Braune, Anita Gojun, Alexander Fraser 0001 |
EAMT | 3 |
| 2011 | A Joint Sequence Translation Model with Integrated Reordering
Nadir Durrani, Helmut Schmid, Alexander Fraser 0001 |
ACL | 3 |
| 2011 | An Algorithm for Unsupervised Transliteration Mining with an Application to Word Alignment
Hassan Sajjad 0001, Alexander Fraser 0001, Helmut Schmid |
ACL | 2 |
| 2011 | Comparing Two Techniques for Learning Transliteration Models Using a Parallel Corpus
Hassan Sajjad 0001, Nadir Durrani, Helmut Schmid, Alexander Fraser 0001 |
IJCNLP | 4 |
| 2010 | Hindi-to-Urdu Machine Translation through Transliteration
Nadir Durrani, Hassan Sajjad 0001, Alexander Fraser 0001, Helmut Schmid |
ACL | 3 |
| 2010 | Bitext-Based Resolution of German Subject-Object Ambiguities
Florian Schwarck, Alexander Fraser 0001, Hinrich Schütze |
HLT-NAACL | 2 |
| 2009 | Rich Bitext Projection Features for Parse Reranking
Alexander Fraser 0001, Renjing Wang, Hinrich Schütze |
EACL | 1 |
| 2007 | Getting the Structure Right for Word Alignment: LEAF
Alexander Fraser 0001, Daniel Marcu |
EMNLP-CoNLL | 1 |
| 2007 | Measuring Word Alignment Quality for Statistical Machine TranslationabstractAutomatic word alignment plays a critical role in statistical machine translation. Unfortunately, the relationship between alignment quality and statistical machine translation performance has not been well understood. In the recent literature, the alignment task has frequently been decoupled from the translation task and assumptions have been made about measuring alignment quality for machine translation which, it turns out, are not justified. In particular, none of the tens of papers published over the last five years has shown that significant decreases in alignment error rate (AER) result in significant increases in translation performance. This paper explains this state of affairs and presents steps towards measuring alignment quality in a way which is predictive of statistical machine translation performance. Alexander Fraser 0001, Daniel Marcu |
Comput. Linguistics | 1 |
| 2006 | Semi-Supervised Training for Statistical Word AlignmentabstractWe introduce a semi-supervised approach to training for statistical machine translation that alternates the traditional Expectation Maximization step that is applied on a large training corpus with a discriminative step aimed at increasing word-alignment quality on a small, manually word-aligned sub-corpus. We show that our algorithm leads not only to improved alignments but also to machine translation outputs of higher quality. Alexander Fraser 0001, Daniel Marcu |
ACL | 1 |
| 2004 | Improved Machine Translation Performance via Parallel Sentence Extraction from Comparable Corpora
Dragos Stefan Munteanu, Alexander Fraser 0001, Daniel Marcu |
HLT-NAACL | 2 |
| 2004 | A Smorgasbord of Features for Statistical Machine Translation
Franz Josef Och, Daniel Gildea, Sanjeev Khudanpur, Anoop Sarkar, Kenji Yamada, Alexander Fraser 0001, Shankar Kumar, Libin Shen, Katherine Eng, Viren Jain, Zhen Jin 0007, Dragomir R. Radev |
HLT-NAACL | 6 |
| 2002 | Empirical studies in strategies for Arabic retrievalabstractThis work evaluates a few search strategies for Arabic monolingual and cross-lingual retrieval, using the TREC Arabic corpus as the test-bed. The release by NIST in 2001 of an Arabic corpus of nearly 400k documents with both monolingual and cross-lingual queries and relevance judgments has been a new enabler for empirical studies. Experimental results show that spelling normalization and stemming can significantly improve Arabic monolingual retrieval. Character tri-grams from stems improved retrieval modestly on the test corpus, but the improvement is not statistically significant. To further improve retrieval, we propose a novel thesaurus-based technique. Different from existing approaches to thesaurus-based retrieval, ours formulates word synonyms as probabilistic term translations that can be automatically derived from a parallel corpus. Retrieval results show that the thesaurus can significantly improve Arabic monolingual retrieval. For cross-lingual retrieval (CLIR), we found that spelling normalization and stemming have little impact. Jinxi Xu, Alexander Fraser 0001, Ralph M. Weischedel |
SIGIR | 2 |