EDBT 2026 Demo / reviewers in the wild / expert
Arianna Bisazza
dblp:32/10934
· DBLP profile ↗
47ranked-venue papers
9as first author
27since 2021 · last 2026
0000-0003-1270-3048ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 8 first-author · 27 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CAIT: A Syntactic Parsing Toolkit for Child-Adult InTeractionsabstractFrancesca Padovani, Xiulin Yang, Bastian Bunzeck, Jaap Jumelet, Yevgen Matusevych, Nathan Schneider, Arianna Bisazza. Proceedings of the 30th Conference on Computational Natural Language Learning. 2026. Francesca Padovani, Xiulin Yang, Bastian Bunzeck, Jaap Jumelet, Yevgen Matusevych, Nathan Schneider 0001, Arianna Bisazza |
CoNLL | 7 |
| 2026 | MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal PairsabstractAbstract We introduce MultiBLiMP 1.0, a massively multilingual benchmark of linguistic minimal pairs, covering 101 languages and 2 types of subject-verb agreement, containing more than 128,000 minimal pairs. Our minimal pairs are created using a fully automated pipeline, leveraging the large-scale linguistic resources of Universal Dependencies and UniMorph. MultiBLiMP 1.0 evaluates abilities of LLMs at an unprecedented multilingual scale, and highlights the shortcomings of the current state-of-the-art in modelling low-resource languages.1 Jaap Jumelet, Leonie Weissweiler, Joakim Nivre, Arianna Bisazza |
Trans. Assoc. Comput. Linguistics | 4 |
| 2025 | Simulating the Emergence of Differential Case Marking with Communicating Neural-Network Agents
Yuchen Lian, Arianna Bisazza, Tessa Verhoef |
CogSci | 2 |
| 2025 | TurBLiMP: A Turkish Benchmark of Linguistic Minimal PairsabstractWe introduce TurBLiMP, the first Turkish benchmark of linguistic minimal pairs, designed to evaluate the linguistic abilities of monolingual and multilingual language models (LMs).Covering 16 linguistic phenomena with 1000 minimal pairs each, TurBLiMP fills an important gap in linguistic evaluation resources for Turkish.In designing the benchmark, we give extra attention to two properties of Turkish that remain understudied in current syntactic evaluations of LMs, namely word order flexibility and subordination through morphological processes.Our experiments on a wide range of LMs and a newly collected set of human acceptability judgments reveal that even cutting-edge Large LMs still struggle with grammatical phenomena that are not challenging for humans, and may also exhibit different sensitivities to word order and morphological complexity compared to humans. Ezgi Basar, Francesca Padovani, Jaap Jumelet, Arianna Bisazza |
EMNLP | 4 |
| 2025 | Reading Between the Prompts: How Stereotypes Shape LLM's Implicit PersonalizationabstractGenerative Large Language Models (LLMs) infer user's demographic information from subtle cues in the conversation -a phenomenon called implicit personalization.Prior work has shown that such inferences can lead to lower quality responses for users assumed to be from minority groups, even when no demographic information is explicitly provided.In this work, we systematically explore how LLMs respond to stereotypical cues using controlled synthetic conversations, by analyzing the models' latent user representations through both model internals and generated answers to targeted user questions.Our findings reveal that LLMs do infer demographic attributes based on these stereotypical signals, which for a number of groups even persists when the user explicitly identifies with a different demographic group.Finally, we show that this form of stereotypedriven implicit personalization can be effectively mitigated by intervening on the model's internal representations using a trained linear probe to steer them toward the explicitly stated identity.Our results highlight the need for greater transparency and control in how LLMs represent user identity. Vera Neplenbroek, Arianna Bisazza, Raquel Fernández |
EMNLP | 2 |
| 2025 | Child-Directed Language Does Not Consistently Boost Syntax Learning in Language ModelsabstractSeminal work by Huebner et al. (2021) showed that language models (LMs) trained on English Child-Directed Language (CDL) can reach similar syntactic abilities as LMs trained on much larger amounts of adult-directed written text, suggesting that CDL could provide more effective LM training material than the commonly used internet-crawled data.However, the generalizability of these results across languages, model types, and evaluation settings remains unclear.We test this by comparing models trained on CDL vs. Wikipedia across two LM objectives (masked and causal), three languages (English, French, German), and three syntactic minimal-pair benchmarks.Our results on these benchmarks show inconsistent benefits of CDL, which in most cases is outperformed by Wikipedia models.We then identify various shortcomings in previous benchmarks, and introduce a novel testing methodology, FIT-CLAMS, which uses a frequency-controlled design to enable balanced comparisons across training corpora.Through minimal pair evaluations and regression analysis we show that training on CDL does not yield stronger generalizations for acquiring syntax and highlight the importance of controlling for frequency effects when evaluating syntactic ability. 1 Francesca Padovani, Jaap Jumelet, Yevgen Matusevych, Arianna Bisazza |
EMNLP | 4 |
| 2025 | Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreementabstractWord-level quality estimation (WQE) aims to automatically identify fine-grained error spans in machine-translated outputs and has found many uses, including assisting translators during post-editing.Modern WQE techniques are often expensive, involving prompting of large language models or ad-hoc training on large amounts of human-labeled data.In this work, we investigate efficient alternatives exploiting recent advances in language model interpretability and uncertainty quantification to identify translation errors from the inner workings of translation models.In our evaluation spanning 14 metrics across 12 translation directions, we quantify the impact of human label variation on metric performance by using multiple sets of human labels.Our results highlight the untapped potential of unsupervised metrics, the shortcomings of supervised methods when faced with label uncertainty, and the brittleness of single-annotator evaluation practices. Gabriele Sarti, Vilém Zouhar, Malvina Nissim, Arianna Bisazza |
EMNLP | 4 |
| 2025 | On the reliability of feature attribution methods for speech classificationabstractAs the capabilities of large-scale pre-trained models evolve, understanding the determinants of their outputs becomes more important. Feature attribution aims to reveal which parts of the input elements contribute the most to model outputs. In speech processing, the unique characteristics of the input signal make the application of feature attribution methods challenging. We study how factors such as input type and aggregation and perturbation timespan impact the reliability of standard feature attribution methods, and how these factors interact with characteristics of each classification task. We find that standard approaches to feature attribution are generally unreliable when applied to the speech domain, with the exception of word-aligned perturbation methods when applied to word-based classification tasks. Gaofei Shen, Hosein Mohebbi, Arianna Bisazza, Afra Alishahi, Grzegorz Chrupala |
INTERSPEECH | 3 |
| 2025 | Pointwise Mutual Information as a Performance Gauge for Retrieval-Augmented GenerationabstractTianyu Liu, Jirui Qi, Paul He, Arianna Bisazza, Mrinmaya Sachan, Ryan Cotterell. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Tianyu Liu 0004, Jirui Qi, Paul He 0001, Arianna Bisazza, Mrinmaya Sachan, Ryan Cotterell |
NAACL (Long Papers) | 4 |
| 2025 | QE4PE: Word-level Quality Estimation for Human Post-Editing
Gabriele Sarti, Vilém Zouhar, Grzegorz Chrupala, Ana Guerberof Arenas, Malvina Nissim, Arianna Bisazza |
Trans. Assoc. Comput. Linguistics | 6 |
| 2024 | Neural-agent Language Learning and Communication: Emergence of Dependency Length Minimization
Yuqing Zhang 0003, Tessa Verhoef, Gertjan van Noord, Arianna Bisazza |
CogSci | 4 |
| 2024 | Endowing Neural Language Learners with Human-like Biases: A Case Study on Dependency Length MinimizationabstractNatural languages show a tendency to minimize the linear distance between heads and their dependents in a sentence, known as dependency length minimization (DLM). Such a preference, however, has not been consistently replicated with neural agent simulations. Comparing the behavior of models with that of human learners can reveal which aspects affect the emergence of this phenomenon. In this work, we investigate the minimal conditions that may lead neural learners to develop a DLM preference. We add three factors to the standard neural-agent language learning and communication framework to make the simulation more realistic, namely: (i) the presence of noise during listening, (ii) context-sensitivity of word use through non-uniform conditional word distributions, and (iii) incremental sentence processing, or the extent to which an utterance’s meaning can be guessed before hearing it entirely. While no preference appears in production, we show that the proposed factors can contribute to a small but significant learning advantage of DLM for listeners of verb-initial languages. Yuqing Zhang 0003, Tessa Verhoef, Gertjan van Noord, Arianna Bisazza |
LREC/COLING | 4 |
| 2024 | Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented GenerationabstractEnsuring the verifiability of model answers is a fundamental challenge for retrieval-augmented generation (RAG) in the question answering (QA) domain.Recently, self-citation prompting was proposed to make large language models (LLMs) generate citations to supporting documents along with their answers.However, self-citing LLMs often struggle to match the required format, refer to non-existent sources, and fail to faithfully reflect LLMs' context usage throughout the generation.In this work, we present MIRAGE -Model Internals-based RAG Explanations -a plug-and-play approach using model internals for faithful answer attribution in RAG applications.MIRAGE detects context-sensitive answer tokens and pairs them with retrieved documents contributing to their prediction via saliency methods.We evaluate our proposed approach on a multilingual extractive QA dataset, finding high agreement with human answer attribution.On open-ended QA, MIRAGE achieves citation quality and efficiency comparable to self-citation while also allowing for a finer-grained control of attribution parameters.Our qualitative evaluation highlights the faithfulness of MIRAGE's attributions and underscores the promising application of model internals for RAG answer attribution. 1 Jirui Qi, Gabriele Sarti, Raquel Fernández, Arianna Bisazza |
EMNLP | 4 |
| 2024 | Quantifying the Plausibility of Context Reliance in Neural Machine TranslationabstractEstablishing whether language models can use contextual information in a human-plausible way is important to ensure their safe adoption in real-world settings. However, the questions of $\textit{when}$ and $\textit{which parts}$ of the context affect model generations are typically tackled separately, and current plausibility evaluations are practically limited to a handful of artificial benchmarks. To address this, we introduce $\textbf{P}$lausibility $\textbf{E}$valuation of $\textbf{Co}$ntext $\textbf{Re}$liance (PECoRe), an end-to-end interpretability framework designed to quantify context usage in language models' generations. Our approach leverages model internals to (i) contrastively identify context-sensitive target tokens in generated texts and (ii) link them to contextual cues justifying their prediction. We use PECoRe to quantify the plausibility of context-aware machine translation models, comparing model rationales with human annotations across several discourse-level phenomena. Finally, we apply our method to unannotated model translations to identify context-mediated predictions and highlight instances of (im)plausible context usage throughout generation. Gabriele Sarti, Grzegorz Chrupala, Malvina Nissim, Arianna Bisazza |
ICLR | 4 |
| 2024 | Encoding of lexical tone in self-supervised models of spoken languageabstractGaofei Shen, Michaela Watkins, Afra Alishahi, Arianna Bisazza, Grzegorz Chrupała. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Gaofei Shen, Michaela Watkins, Afra Alishahi, Arianna Bisazza, Grzegorz Chrupala |
NAACL-HLT | 4 |
| 2024 | Are Character-level Translations Worth the Wait? Comparing ByT5 and mT5 for Machine TranslationabstractAbstract Pretrained character-level and byte-level language models have been shown to be competitive with popular subword models across a range of Natural Language Processing tasks. However, there has been little research on their effectiveness for neural machine translation (NMT), particularly within the popular pretrain-then-finetune paradigm. This work performs an extensive comparison across multiple languages and experimental conditions of character- and subword-level pretrained models (ByT5 and mT5, respectively) on NMT. We show the effectiveness of character-level modeling in translation, particularly in cases where fine-tuning data is limited. In our analysis, we show how character models’ gains in translation quality are reflected in better translations of orthographically similar words and rare words. While evaluating the importance of source texts in driving model predictions, we highlight word-level patterns within ByT5, suggesting an ability to modulate word-level and character-level information during generation. We conclude by assessing the efficiency tradeoff of byte models, suggesting their usage in non-time-critical scenarios to boost translation quality. Lukas Edman, Gabriele Sarti, Antonio Toral, Gertjan van Noord, Arianna Bisazza |
Trans. Assoc. Comput. Linguistics | 5 |
| 2023 | The importance of communicative success for simulating the emergence of a Word Order/Case Marking trade-off with Neural Agents
Yuchen Lian, Arianna Bisazza, Tessa Verhoef |
CogSci | 2 |
| 2023 | Cross-Lingual Consistency of Factual Knowledge in Multilingual Language ModelsabstractMultilingual large-scale Pretrained Language Models (PLMs) have been shown to store considerable amounts of factual knowledge, but large variations are observed across languages.With the ultimate goal of ensuring that users with different language backgrounds obtain consistent feedback from the same model, we study the cross-lingual consistency (CLC) of factual knowledge in various multilingual PLMs.To this end, we propose a Rankingbased Consistency (RankC) metric to evaluate knowledge consistency across languages independently from accuracy.Using this metric, we conduct an in-depth analysis of the determining factors for CLC, both at model level and at language-pair level.Among other results, we find that increasing model size leads to higher factual probing accuracy in most languages, but does not improve cross-lingual consistency.Finally, we conduct a case study on CLC when new factual associations are inserted in the PLMs via model editing.Results on a small sample of facts inserted in English reveal a clear pattern whereby the new piece of knowledge transfers only to languages with which English has a high RankC score. 1 Jirui Qi, Raquel Fernández, Arianna Bisazza |
EMNLP | 3 |
| 2023 | Wave to Syntax: Probing spoken language models for syntaxabstractUnderstanding which information is encoded in deep models of spoken and written language has been the focus of much research in recent years, as it is crucial for debugging and improving these architectures. Most previous work has focused on probing for speaker characteristics, acoustic and phonological information in models of spoken language, and for syntactic information in models of written language. Here we focus on the encoding of syntax in several self-supervised and visually grounded models of spoken language. We employ two complementary probing methods, combined with baselines and reference representations to quantify the degree to which syntactic structure is encoded in the activations of the target models. We show that syntax is captured most prominently in the middle layers of the networks, and more explicitly within models with more parameters. Gaofei Shen, Afra Alishahi, Arianna Bisazza, Grzegorz Chrupala |
INTERSPEECH | 3 |
| 2023 | Communication Drives the Emergence of Language Universals in Neural Agents: Evidence from the Word-order/Case-marking Trade-offabstractAbstract Artificial learners often behave differently from human learners in the context of neural agent-based simulations of language emergence and change. A common explanation is the lack of appropriate cognitive biases in these learners. However, it has also been proposed that more naturalistic settings of language learning and use could lead to more human-like results. We investigate this latter account, focusing on the word-order/case-marking trade-off, a widely attested language universal that has proven particularly hard to simulate. We propose a new Neural-agent Language Learning and Communication framework (NeLLCom) where pairs of speaking and listening agents first learn a miniature language via supervised learning, and then optimize it for communication via reinforcement learning. Following closely the setup of earlier human experiments, we succeed in replicating the trade-off with the new framework without hard-coding specific biases in the agents. We see this as an essential step towards the investigation of language universals with neural learners. Yuchen Lian, Arianna Bisazza, Tessa Verhoef |
Trans. Assoc. Comput. Linguistics | 2 |
| 2022 | InDeep $\times$ NMT: Empowering Human Translators via Interpretable Neural Machine Translation
Gabriele Sarti, Arianna Bisazza |
EAMT | 2 |
| 2022 | DivEMT: Neural Machine Translation Post-Editing Effort Across Typologically Diverse LanguagesabstractWe introduce DivEMT, the first publicly available post-editing study of Neural Machine Translation (NMT) over a typologically diverse set of target languages.Using a strictly controlled setup, 18 professional translators were instructed to translate or post-edit the same set of English documents into Arabic, Dutch, Italian, Turkish, Ukrainian, and Vietnamese.During the process, their edits, keystrokes, editing times and pauses were recorded, enabling an in-depth, cross-lingual evaluation of NMT quality and post-editing effectiveness.Using this new dataset, we assess the impact of two state-of-the-art NMT systems, Google Translate and the multilingual mBART-50 model, on translation productivity.We find that post-editing is consistently faster than translation from scratch.However, the magnitude of productivity gains varies widely across systems and languages, highlighting major disparities in post-editing effectiveness for languages at different degrees of typological relatedness to English, even when controlling for system architecture and training data size.We publicly release the complete dataset 1 including all collected behavioral data, to foster new research on the translation capabilities of NMT systems for typologically diverse languages. Gabriele Sarti, Arianna Bisazza, Ana Guerberof Arenas, Antonio Toral |
EMNLP | 2 |
| 2022 | Hyper-X: A Unified Hypernetwork for Multi-Task Multilingual TransferabstractMassively multilingual models are promising for transfer learning across tasks and languages.However, existing methods are unable to fully leverage training data when it is available in different task-language combinations.To exploit such heterogeneous supervision, we propose Hyper-X, a single hypernetwork that unifies multi-task and multilingual learning with efficient adaptation.This model generates weights for adapter modules conditioned on both tasks and language embeddings.By learning to combine task and language-specific knowledge, our model enables zero-shot transfer for unseen languages and task-language combinations.Our experiments on a diverse set of languages demonstrate that Hyper-X achieves the best or competitive gain when a mixture of multiple resources is available, while being on par with strong baselines in the standard scenario.Hyper-X is also considerably more efficient in terms of parameters and resources compared to methods that train separate adapters.Finally, Hyper-X consistently produces strong results in few-shot scenarios for new languages, showing the versatility of our approach beyond zero-shot transfer.1 NER en Pre-trained Model Pre-trained Model Fine-tuned Model ar tr Fine-tuned Model ar tr POS en Single-Task Pre-trained Model Fine-tuned Model ar tr ar tr NER POS en en Multi-Task Pre-trained Model Fine-tuned Model tr NER POS ar tr NER ar POS en en Mixed-Language Multi-Task Anne Lauscher, Vinit Ravishankar, Ivan Vulić, and Goran Glavaš.2020a.From Zero to Hero: On the Limitations of Zero-Shot Ahmet Üstün, Arianna Bisazza, Gosse Bouma, Gertjan van Noord, Sebastian Ruder |
EMNLP | 2 |
| 2022 | Evaluating Pre-training Objectives for Low-Resource Translation into Morphologically Rich LanguagesabstractThe scarcity of parallel data is a major limitation for Neural Machine Translation (NMT) systems, in particular for translation into morphologically rich languages (MRLs). An important way to overcome the lack of parallel data is to leverage target monolingual data, which is typically more abundant and easier to collect. We evaluate a number of techniques to achieve this, ranging from back-translation to random token masking, on the challenging task of translating English into four typologically diverse MRLs, under low-resource settings. Additionally, we introduce Inflection Pre-Training (or PT-Inflect), a novel pre-training objective whereby the NMT system is pre-trained on the task of re-inflecting lemmatized target sentences before being trained on standard source-to-target language translation. We conduct our evaluation on four typologically diverse target MRLs, and find that PT-Inflect surpasses NMT systems trained only on parallel data. While PT-Inflect is outperformed by back-translation overall, combining the two techniques leads to gains in some of the evaluated language pairs. Prajit Dhar, Arianna Bisazza, Gertjan van Noord |
LREC | 2 |
| 2022 | UDapter: Typology-based Language Adapters for Multilingual Dependency Parsing and Sequence LabelingabstractAbstract Recent advances in multilingual language modeling have brought the idea of a truly universal parser closer to reality. However, such models are still not immune to the “curse of multilinguality”: Cross-language interference and restrained model capacity remain major obstacles. To address this, we propose a novel language adaptation approach by introducing contextual language adapters to a multilingual parser. Contextual language adapters make it possible to learn adapters via language embeddings while sharing model parameters across languages based on contextual parameter generation. Moreover, our method allows for an easy but effective integration of existing linguistic typology features into the parsing model. Because not all typological features are available for every language, we further combine typological feature prediction with parsing in a multi-task model that achieves very competitive parsing performance without the need for an external prediction system for missing features. The resulting parser, UDapter, can be used for dependency parsing as well as sequence labeling tasks such as POS tagging, morphological tagging, and NER. In dependency parsing, it outperforms strong monolingual and multilingual baselines on the majority of both high-resource and low-resource (zero-shot) languages, showing the success of the proposed adaptation approach. In sequence labeling tasks, our parser surpasses the baseline on high resource languages, and performs very competitively in a zero-shot setting. Our in-depth analyses show that adapter generation via typological features of languages is key to this success.1 Ahmet Üstün, Arianna Bisazza, Gosse Bouma, Gertjan van Noord |
Comput. Linguistics | 2 |
| 2021 | The Effect of Efficient Messaging and Input Variability on Neural-Agent Iterated Language LearningabstractNatural languages display a trade-off among different strategies to convey syntactic structure, such as word order or inflection.This trade-off, however, has not appeared in recent simulations of iterated language learning with neural network agents (Chaabouni et al., 2019b).We re-evaluate this result in light of three factors that play an important role in comparable experiments from the Language Evolution field: (i) speaker bias towards efficient messaging, (ii) non systematic input languages, and (iii) learning bottleneck.Our simulations show that neural agents mainly strive to maintain the utterance type distribution observed during learning, instead of developing a more efficient or systematic language. Yuchen Lian, Arianna Bisazza, Tessa Verhoef |
EMNLP (1) | 2 |
| 2021 | On the Difficulty of Translating Free-Order Case-Marking LanguagesabstractAbstract Identifying factors that make certain languages harder to model than others is essential to reach language equality in future Natural Language Processing technologies. Free-order case-marking languages, such as Russian, Latin, or Tamil, have proved more challenging than fixed-order languages for the tasks of syntactic parsing and subject-verb agreement prediction. In this work, we investigate whether this class of languages is also more difficult to translate by state-of-the-art Neural Machine Translation (NMT) models. Using a variety of synthetic languages and a newly introduced translation challenge set, we find that word order flexibility in the source language only leads to a very small loss of NMT quality, even though the core verb arguments become impossible to disambiguate in sentences without semantic cues. The latter issue is indeed solved by the addition of case marking. However, in medium- and low-resource settings, the overall NMT quality of fixed-order languages remains unmatched. Arianna Bisazza, Ahmet Üstün, Stephan Sportel |
Trans. Assoc. Comput. Linguistics | 1 |
| 2020 | UDapter: Language Adaptation for Truly Universal Dependency ParsingabstractRecent advances in multilingual dependency parsing have brought the idea of a truly universal parser closer to reality.However, crosslanguage interference and restrained model capacity remain major obstacles.To address this, we propose a novel multilingual task adaptation approach based on contextual parameter generation and adapter modules.This approach enables to learn adapters via language embeddings while sharing model parameters across languages.It also allows for an easy but effective integration of existing linguistic typology features into the parsing network.The resulting parser, UDapter, outperforms strong monolingual and multilingual baselines on the majority of both high-resource and lowresource (zero-shot) languages, showing the success of the proposed adaptation approach.Our in-depth analyses show that soft parameter sharing via typological features is key to this success.1 Ahmet Üstün, Arianna Bisazza, Gosse Bouma, Gertjan van Noord |
EMNLP (1) | 2 |
| 2019 | Compositionality in emerging multi-agent languages: Marrying Language Evolution and Natural Language Processing
Kees Sommer, Jae Perris, Arianna Bisazza, Tessa Verhoef |
CogSci | 3 |
| 2018 | The Lazy Encoder: A Fine-Grained Analysis of the Role of Morphology in Neural Machine TranslationabstractNeural sequence-to-sequence models have proven very effective for machine translation, but at the expense of model interpretability.To shed more light into the role played by linguistic structure in the process of neural machine translation, we perform a fine-grained analysis of how various source-side morphological features are captured at different levels of the NMT encoder while varying the target language.Differently from previous work, we find no correlation between the accuracy of source morphology encoding and translation quality.We do find that morphological features are only captured in context and only to the extent that they are directly transferable to the target words.FR: Les arbres pl sont hauts. Arianna Bisazza, Clara Tump |
EMNLP | 1 |
| 2018 | The importance of Being Recurrent for Modeling Hierarchical StructureabstractRecent work has shown that recurrent neural networks (RNNs) can implicitly capture and exploit hierarchical information when trained to solve common natural language processing tasks (Blevins et al., 2018) such as language modeling (Linzen et al., 2016;Gulordava et al., 2018) and neural machine translation (Shi et al., 2016).In contrast, the ability to model structured data with non-recurrent neural networks has received little attention despite their success in many NLP tasks (Gehring et al., 2017;Vaswani et al., 2017).In this work, we compare the two architectures-recurrent versus non-recurrent-with respect to their ability to model hierarchical structure and find that recurrency is indeed important for this purpose.The code and data used in our experiments is available at https://github.com/ ketranm/fan_vs_rnn Ke M. Tran, Arianna Bisazza, Christof Monz |
EMNLP | 2 |
| 2018 | Examining the Tip of the Iceberg: A Data Set for Idiom Translation
Marzieh Fadaee, Arianna Bisazza, Christof Monz |
LREC | 2 |
| 2018 | Evaluation of Machine Translation Performance Across Multiple Genres and Languages
Marlies van der Wees, Arianna Bisazza, Christof Monz |
LREC | 2 |
| 2018 | Neural versus phrase-based MT quality: An in-depth analysis on English-German and English-French
Luisa Bentivogli, Arianna Bisazza, Mauro Cettolo, Marcello Federico |
Comput. Speech Lang. | 2 |
| 2017 | Dynamic Data Selection for Neural Machine TranslationabstractIntelligent selection of training data has proven a successful technique to simultaneously increase training efficiency and translation performance for phrase-based machine translation (PBMT).With the recent increase in popularity of neural machine translation (NMT), we explore in this paper to what extent and how NMT can also benefit from data selection.While state-of-the-art data selection (Axelrod et al., 2011) consistently performs well for PBMT, we show that gains are substantially lower for NMT.Next, we introduce dynamic data selection for NMT, a method in which we vary the selected subset of training data between different training epochs.Our experiments show that the best results are achieved when applying a technique we call gradual fine-tuning, with improvements up to +2.6 BLEU over the original data selection approach and up to +3.1 BLEU over a general baseline. Marlies van der Wees, Arianna Bisazza, Christof Monz |
EMNLP | 2 |
| 2016 | Measuring the Effect of Conversational Aspects on Machine Translation QualityabstractResearch in statistical machine translation (SMT) is largely driven by formal translation tasks, while translating informal text is much more challenging. In this paper we focus on SMT for the informal genre of dialogues, which has rarely been addressed to date. Concretely, we investigate the effect of dialogue acts, speakers, gender, and text register on SMT quality when translating fictional dialogues. We first create and release a corpus of multilingual movie dialogues annotated with these four dialogue-specific aspects. When measuring translation performance for each of these variables, we find that BLEU fluctuations between their categories are often significantly larger than randomly expected. Following this finding, we hypothesize and show that SMT of fictional dialogues benefits from adaptation towards dialogue acts and registers. Finally, we find that male speakers are harder to translate and use more vulgar language than female speakers, and that vulgarity is often not preserved during translation. Marlies van der Wees, Arianna Bisazza, Christof Monz |
COLING | 2 |
| 2016 | Neural versus Phrase-Based Machine Translation Quality: a Case StudyabstractWithin the field of Statistical Machine Translation (SMT), the neural approach (NMT) has recently emerged as the first technology able to challenge the long-standing dominance of phrase-based approaches (PBMT).In particular, at the IWSLT 2015 evaluation campaign, NMT outperformed well established state-ofthe-art PBMT systems on English-German, a language pair known to be particularly hard because of morphology and syntactic differences.To understand in what respects NMT provides better translation quality than PBMT, we perform a detailed analysis of neural vs. phrase-based SMT outputs, leveraging high quality post-edits performed by professional translators on the IWSLT data.For the first time, our analysis provides useful insights on what linguistic phenomena are best modeled by neural models -such as the reordering of verbs -while pointing out other aspects that remain to be improved. Luisa Bentivogli, Arianna Bisazza, Mauro Cettolo, Marcello Federico |
EMNLP | 2 |
| 2016 | Recurrent Memory Networks for Language ModelingabstractRecurrent Neural Networks (RNNs) have obtained excellent result in many natural language processing (NLP) tasks.However, understanding and interpreting the source of this success remains a challenge.In this paper, we propose Recurrent Memory Network (RMN), a novel RNN architecture, that not only amplifies the power of RNN but also facilitates our understanding of its internal functioning and allows us to discover underlying patterns in data.We demonstrate the power of RMN on language modeling and sentence completion tasks.On language modeling, RMN outperforms Long Short-Term Memory (LSTM) network on three large German, Italian, and English dataset.Additionally we perform indepth analysis of various linguistic dimensions that RMN captures.On Sentence Completion Challenge, for which it is essential to capture sentence coherence, our RMN obtains 69.2% accuracy, surpassing the previous state of the art by a large margin. 1 Ke M. Tran, Arianna Bisazza, Christof Monz |
HLT-NAACL | 2 |
| 2016 | A Survey of Word Reordering in Statistical Machine Translation: Computational Models and Language PhenomenaabstractWord reordering is one of the most difficult aspects of statistical machine translation (SMT), and an important factor of its quality and efficiency. Despite the vast amount of research published to date, the interest of the community in this problem has not decreased, and no single method appears to be strongly dominant across language pairs. Instead, the choice of the optimal approach for a new translation task still seems to be mostly driven by empirical trials. To orient the reader in this vast and complex research area, we present a comprehensive survey of word reordering viewed as a statistical modeling challenge and as a natural language phenomenon. The survey describes in detail how word reordering is modeled within different string-based and tree-based SMT frameworks and as a stand-alone task, including systematic overviews of the literature in advanced reordering modeling. We then question why some approaches are more successful than others in different language pairs. We argue that besides measuring the amount of reordering, it is important to understand which kinds of reordering occur in a given language pair. To this end, we conduct a qualitative analysis of word reordering phenomena in a diverse sample of language pairs, based on a large collection of linguistic knowledge. Empirical results in the SMT literature are shown to support the hypothesis that a few linguistic facts can be very useful to anticipate the reordering characteristics of a language pair and to select the SMT framework that best suits them. Arianna Bisazza, Marcello Federico |
Comput. Linguistics | 1 |
| 2015 | A distributed inflection model for translating into morphologically rich languages
Ke M. Tran, Arianna Bisazza, Christof Monz |
MTSummit | 2 |
| 2014 | Class-Based Language Modeling for Translating into Morphologically Rich Languages
Arianna Bisazza, Christof Monz |
COLING | 1 |
| 2014 | Word Translation Prediction for Morphologically Rich Languages with Bilingual Neural NetworksabstractTranslating into morphologically rich lan-guages is a particularly difficult problem in machine translation due to the high de-gree of inflectional ambiguity in the tar-get language, often only poorly captured by existing word translation models. We present a general approach that exploits source-side contexts of foreign words to improve translation prediction accuracy. Our approach is based on a probabilistic neural network which does not require lin-guistic annotation nor manual feature en-gineering. We report significant improve-ments in word translation prediction accu-racy for three morphologically rich target languages. In addition, preliminary results for integrating our approach into a large-scale English-Russian statistical machine translation system show small but statisti-cally significant improvements in transla-tion quality. 1 Ke M. Tran, Arianna Bisazza, Christof Monz |
EMNLP | 2 |
| 2013 | Dynamically Shaping the Reordering Search Space of Phrase-Based Statistical Machine TranslationabstractDefining the reordering search space is a crucial issue in phrase-based SMT between distant languages. In fact, the optimal trade-off between accuracy and complexity of decoding is nowadays reached by harshly limiting the input permutation space. We propose a method to dynamically shape such space and, thus, capture long-range word movements without hurting translation quality nor decoding time. The space defined by loose reordering constraints is dynamically pruned through a binary classifier that predicts whether a given input word should be translated right after another. The integration of this model into a phrase-based decoder improves a strong Arabic-English baseline already including state-of-the-art early distortion cost (Moore and Quirk, 2007) and hierarchical phrase orientation models (Galley and Manning, 2008). Significant improvements in the reordering of verbs are achieved by a system that is notably faster than the baseline, while bleu and meteor remain stable, or even increase, at a very high distortion limit. Arianna Bisazza, Marcello Federico |
Trans. Assoc. Comput. Linguistics | 1 |
| 2012 | Modified Distortion Matrices for Phrase-Based Statistical Machine Translation
Arianna Bisazza, Marcello Federico |
ACL (1) | 1 |
| 2012 | Cutting the Long Tail: Hybrid Language Models for Translation Style Adaptation
Arianna Bisazza, Marcello Federico |
EACL | 1 |
| 2012 | Chunk-lattices for verb reordering in Arabic-English statistical machine translation - Special issues on machine translation for Arabic
Arianna Bisazza, Daniele Pighin, Marcello Federico |
Mach. Transl. | 1 |
| 2008 | Semantic annotations for conversational speech: From speech transcriptions to predicate argument structuresabstractIn this paper, we describe the semantic content, which can be automatically generated, for the design of advanced dialog systems. Since the latter will be based on machine learning approaches, we created training data by annotating a corpus with the needed content. Given a sentence of our transcribed corpus, domain concepts and other linguistic levels ranging from basic ones, i.e. part-of-speech tagging and constituent chunking level, to more advanced ones, i.e. syntactic and predicate argument structure (PAS) levels are annotated. In particular, the proposed PAS and taxonomy of dialog acts appear to be promising for the design of more complex dialog systems. Statistics about our semantic annotation are reported. Arianna Bisazza, Marco Dinarelli, Silvia Quarteroni, Sara Tonelli, Alessandro Moschitti, Giuseppe Riccardi |
SLT | 1 |