VLDB 2026 Research / reviewers in the wild / expert
Edoardo Maria Ponti
dblp:178/8829 · also Edoardo M. Ponti
· DBLP profile ↗
52ranked-venue papers
9as first author
39since 2021 · last 2026
0000-0002-6308-1050ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 52 · 9 first-author · 39 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Measuring the Effects of Visual Salience in Human and AI Descriptions with Image EditingabstractHow does our perception of the world influence the way we talk about it? Psycholinguistic studies have investigated whether visual salience correlates with entity mention and ordering, but often disregarded its effect on grammar or relied on simplistic images or artificial cues. In this study, we explore the use of generative AI to better control for salience in visual stimuli while keeping them realistic, and to serve as a proxy for human participants in studying how different types of salience impact image descriptions.We consider three salience types: perceptual (e.g. relative size in the image), inherent (e.g. animacy), and relational (e.g. human–object interaction). We first analyze human- and AI-generated captions for natural images to examine how salience correlates with how early, and in what grammatical role, an entity is mentioned. We find strong correlations between models and humans in this observational study, justifying the use of AI models alone in a further causal study. For this second study, we created datasets composed of pairs of images, where we used an image-editing model to intervene on the salience of a target entity. We show that relational and perceptual salience lead to the entity being mentioned earlier in captions and being mapped to more prominent grammatical roles. The magnitude of this effect varies across entity types, with animate entities (high inherent salience) showing a particularly distinct pattern. Nina Gregorio, Edoardo Maria Ponti, Sharon Goldwater |
CoNLL | 2 |
| 2025 | The Cross-linguistic Role of Animacy in Grammar StructuresabstractAnimacy is a semantic feature of nominals and follows a hierarchy: personal pronouns > human > animate > inanimate.In several languages, animacy imposes hard constraints on grammar.While it has been argued that these constraints may emerge from universal soft tendencies, it has been difficult to provide empirical evidence for this conjecture due to the lack of data annotated with animacy classes.In this work, we first propose a method to reliably classify animacy classes of nominals in 11 languages from 5 families, leveraging multilingual large language models (LLMs) and word sense disambiguation datasets.Then, through this newly acquired data, we verify that animacy displays consistent cross-linguistic tendencies in terms of preferred morphosyntactic constructions, although not always in line with received wisdom: animacy in nouns correlates with the alignment role of agent, early positions in a clause, and syntactic pivot (e.g., for relativisation), but not necessarily with grammatical subjecthood.Furthermore, the behaviour of personal pronouns in the hierarchy is idiosyncratic as they are rarely plural and relativised, contrary to high-animacy nouns. Nina Gregorio, Matteo Gay, Sharon Goldwater, Edoardo Maria Ponti |
ACL (1) | 4 |
| 2025 | Mixtures of In-Context LearnersabstractIn-context learning (ICL) adapts LLMs by providing demonstrations without fine-tuning the model parameters; however, it is very sensitive to the choice of in-context demonstrations, and processing many demonstrations can be computationally demanding.We propose Mixtures of In-Context Learners (MOICL), a novel approach that uses subsets of demonstrations to train a set of experts via ICL and learns a weighting function to merge their output distributions via gradient-based optimisation.In our experiments, we show performance improvements on 5 out of 7 classification datasets compared to a set of strong baselines (e.g., up to +13% compared to ICL and LENS).Moreover, we improve the Pareto frontier of ICL by reducing the inference time needed to achieve the same performance with fewer demonstrations.Finally, MOICL is more robust to out-ofdomain (up to +11%), imbalanced (up to +49%) and perturbed demonstrations (up to +38%). 1 Giwon Hong, Emile van Krieken, Edoardo Maria Ponti, Nikolay Malkin, Pasquale Minervini |
ACL (1) | 3 |
| 2025 | Post-hoc Reward Calibration: A Case Study on Length BiasabstractReinforcement Learning from Human Feedback aligns the outputs of Large Language Models with human values and preferences. Central to this process is the reward model (RM), which translates human feedback into training signals for optimising LLM behaviour. However, RMs can develop biases by exploiting spurious correlations in their training data, such as favouring outputs based on length or
style rather than true quality. These biases can lead to incorrect output rankings, sub-optimal model evaluations, and the amplification of undesirable behaviours in LLMs alignment. This paper addresses the challenge of correcting such biases without additional data and training, introducing the concept of Post-hoc Reward Calibration. We first propose to use local average reward to estimate the bias term
and, thus, remove it to approximate the underlying true reward. We then extend the approach to a more general and robust form with the Locally Weighted Regression. Focusing on the prevalent length bias, we validate our proposed approaches across three experimental settings, demonstrating consistent improvements: (1) a 3.11 average performance gain across 33 reward models on the RewardBench
dataset; (2) improved agreement of RM produced rankings with GPT-4 evaluations and human preferences based on the AlpacaEval benchmark; and (3) improved Length-Controlled win rate (Dubois et al., 2024) of the RLHF process in multiple LLM–RM combinations. According to our experiments, our method is computationally efficient and generalisable to other types of bias and RMs, offering a scalable and robust solution for mitigating biases in LLM alignment and evaluation. Zihan Qiu, Edoardo Maria Ponti, Ivan Titov 0001 |
ICLR | 4 |
| 2025 | Cross-Lingual and Cross-Cultural Variation in Image DescriptionsabstractUri Berger, Edoardo Ponti. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Uri Berger, Edoardo Maria Ponti |
NAACL (Long Papers) | 2 |
| 2025 | A Grounded Typology of Word ClassesabstractColeman Haley, Sharon Goldwater, Edoardo Ponti. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Coleman Haley, Sharon Goldwater, Edoardo Maria Ponti |
NAACL (Long Papers) | 3 |
| 2025 | Fine-Tuning Large Language Models with Sequential InstructionsabstractHanxu Hu, Simon Yu, Pinzhen Chen, Edoardo Ponti. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Hanxu Hu, Simon Yu, Pinzhen Chen, Edoardo Maria Ponti |
NAACL (Long Papers) | 4 |
| 2025 | Neurosymbolic Reasoning Shortcuts under the Independence AssumptionabstractThe ubiquitous independence assumption among symbolic concepts in neurosymbolic (NeSy) predictors is a convenient simplification: NeSy predictors use it to speed up probabilistic reasoning. Recent works like van Krieken et al. (2024) and Marconato et al. (2024) argued that the independence assumption can hinder learning of NeSy predictors and, more crucially, prevent them from correctly modelling uncertainty. There is, however, scepticism in the NeSy community around the scenarios in which the independence assumption actually limits NeSy systems (Faronius and Dos Martires, 2025). In this work, we settle this question by formally showing that assuming independence among symbolic concepts entails that a model can never represent uncertainty over certain concept combinations. Thus, the model fails to be aware of _reasoning shortcuts_, i.e., the pathological behaviour of NeSy predictors that predict correct downstream tasks but for the wrong reasons. Emile van Krieken, Pasquale Minervini, Edoardo Maria Ponti, Antonio Vergari |
NeSy | 3 |
| 2025 | MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts SystemsabstractThe sparse Mixture-of-Experts (MoE) architecture is increasingly favored for scaling Large Language Models (LLMs) efficiently, but it depends on heterogeneous compute and memory resources. These factors jointly affect system Cost, Accuracy, and Performance (CAP), making trade-offs inevitable. Existing benchmarks often fail to capture these trade-offs accurately, complicating practical deployment decisions. To address this, we introduce MoE-CAP, a benchmark specifically designed for MoE systems. Our analysis reveals that achieving an optimal balance across CAP is difficult with current hardware; MoE systems typically optimize two of the three dimensions at the expense of the third—a dynamic we term the MoE-CAP trade-off. To visualize this, we propose the CAP Radar Diagram. We further introduce sparsity-aware performance metrics—Sparse Memory Bandwidth Utilization (S-MBU) and Sparse Model FLOPS Utilization (S-MFU)—to enable accurate performance benchmarking of MoE systems across diverse hardware platforms and deployment scenarios. This benchmark is available on Github: https://github.com/sparse-generative-ai/MoE-CAP. Yinsicheng Jiang, Yao Fu 0013, Yeqi Huang, Ping Nie, Zhan Lu, Leyang Xue, Congjie He, Man-Kit Sit, Jilong Xue, Ziming Miao, Dayou Du, Tairan Xu, Edoardo Maria Ponti, Luo Mai |
NeurIPS | 15 |
| 2025 | Neurosymbolic Diffusion ModelsabstractNeurosymbolic (NeSy) predictors combine neural perception with symbolic reasoning to solve tasks like visual reasoning. However, standard NeSy predictors assume conditional independence between the symbols they extract, thus limiting their ability to model interactions and uncertainty --- often leading to overconfident predictions and poor out-of-distribution generalisation. To overcome the limitations of the independence assumption, we introduce _neurosymbolic diffusion models_ (NeSyDMs), a new class of NeSy predictors that use discrete diffusion to model dependencies between symbols. Our approach reuses the independence assumption from NeSy predictors at each step of the diffusion process, enabling scalable learning while capturing symbol dependencies and uncertainty quantification. Across both synthetic and real-world benchmarks — including high-dimensional visual path planning and rule-based autonomous driving — NeSyDMs achieve state-of-the-art accuracy among NeSy predictors and demonstrate strong calibration. Emile van Krieken, Pasquale Minervini, Edoardo Maria Ponti, Antonio Vergari |
NeurIPS | 3 |
| 2025 | Inference-Time Hyper-Scaling with KV Cache CompressionabstractInference-time scaling trades efficiency for increased reasoning accuracy by generating longer or more parallel sequences. However, in Transformer LLMs, generation cost is bottlenecked by the size of the key–value (KV) cache, rather than the number of generated tokens. Hence, we explore inference-time hyper-scaling: by compressing the KV cache, we can generate more tokens within the same compute budget and further improve the accuracy of scaled inference. The success of this approach, however, hinges on the ability of compression methods to preserve accuracy even at high compression ratios. To make hyper-scaling practical, we introduce Dynamic Memory Sparsification (DMS), a novel method for sparsifying KV caches that only requires 1K training steps to achieve 8× compression, while maintaining better accuracy than training-free sparse attention. Instead of prematurely discarding cached tokens, DMS delays token eviction, implicitly merging representations and preserving critical information. We demonstrate the effectiveness of inference-time hyper-scaling with DMS on multiple families of LLMs, showing that it boosts accuracy for comparable inference latency and memory load. For instance, we enhance Qwen-R1 32B by 9.1 points on AIME 24, 7.6 on GPQA, and 9.6 on LiveCodeBench on average for an equivalent number of memory reads. Adrian Lancucki, Konrad Staniszewski, Piotr Nawrot, Edoardo Maria Ponti |
NeurIPS | 4 |
| 2025 | Universal Cross-Tokenizer Distillation via Approximate Likelihood MatchingabstractDistillation has shown remarkable success in transferring knowledge from a Large Language Model (LLM) teacher to a student LLM. However, current distillation methods require similar tokenizers between the teacher and the student, restricting their applicability to only a small subset of teacher--student pairs. In this work, we develop a principled cross-tokenizer distillation method to solve this crucial deficiency. Our method is the first to enable effective distillation across fundamentally different tokenizers, while also substantially outperforming prior methods in all other cases. We verify the efficacy of our method on three distinct use cases. First, we show that viewing tokenizer transfer as self-distillation enables unprecedentedly effective transfer across tokenizers, including rapid transfer of subword models to the byte-level. Transferring different models to the same tokenizer also enables ensembling to boost performance. Secondly, we distil a large maths-specialised LLM into a small general-purpose model with a different tokenizer, achieving competitive maths problem-solving performance. Thirdly, we use our method to train state-of-the-art embedding prediction hypernetworks for training-free tokenizer transfer. Our results unlock an expanded range of teacher--student pairs for distillation, enabling new ways to adapt and enhance interaction between LLMs. Benjamin Minixhofer, Ivan Vulic, Edoardo Maria Ponti |
NeurIPS | 3 |
| 2024 | Model Merging by Uncertainty-Based Gradient MatchingabstractModels trained on different datasets can be merged by a weighted-averaging of their parameters, but why does it work and when can it fail? Here, we connect the inaccuracy of weighted-averaging to mismatches in the gradients and propose a new uncertainty-based scheme to improve the performance by reducing the mismatch. The connection also reveals implicit assumptions in other schemes such as averaging, task arithmetic, and Fisher-weighted averaging. Our new method gives consistent improvements for large language models and vision transformers, both in terms of performance and robustness to hyperparameters. Nico Daheim, Thomas Möllenhoff, Edoardo Maria Ponti, Iryna Gurevych, Mohammad Emtiyaz Khan |
ICLR | 3 |
| 2024 | On the Independence Assumption in Neurosymbolic LearningabstractState-of-the-art neurosymbolic learning systems use probabilistic reasoning to guide neural networks towards predictions that conform to logical constraints. Many such systems assume that the probabilities of the considered symbols are conditionally independent given the input to simplify learning and reasoning. We study and criticise this assumption, highlighting how it can hinder optimisation and prevent uncertainty quantification. We prove that loss functions bias conditionally independent neural networks to become overconfident in their predictions. As a result, they are unable to represent uncertainty over multiple valid options. Furthermore, we prove that the minima of such loss functions are usually highly disconnected and non-convex, and thus difficult to optimise. Our theoretical analysis gives the foundation for replacing the conditional independence assumption and designing more expressive neurosymbolic probabilistic models. Emile van Krieken, Pasquale Minervini, Edoardo Maria Ponti, Antonio Vergari |
ICML | 3 |
| 2024 | Dynamic Memory Compression: Retrofitting LLMs for Accelerated InferenceabstractTransformers have emerged as the backbone of large language models (LLMs). However, generation remains inefficient due to the need to store in memory a cache of key–value representations for past tokens, whose size scales linearly with the input sequence length and batch size. As a solution, we propose Dynamic Memory Compression (DMC), a method for on-line key–value cache compression at inference time. Most importantly, the model learns to apply different compression ratios in different heads and layers. We retrofit pre-trained LLMs such as Llama 2 (7B, 13B and 70B) into DMC Transformers, achieving up to $\sim 3.7 \times$ throughput increase during auto-regressive inference on an NVIDIA H100 GPU. DMC is applied via continued pre-training on a negligible percentage of the original data without adding any extra parameters. We find that DMC preserves the original downstream performance with up to 4$\times$ cache compression, outperforming up-trained grouped-query attention (GQA) and key–value eviction policies (H$_2$O, TOVA). GQA and DMC can be even combined to obtain compounded gains. As a result DMC fits longer contexts and larger batches within any given memory budget. We release the DMC code and models at https://github.com/NVIDIA/Megatron-LM/tree/DMC. Piotr Nawrot, Adrian Lancucki, Marcin Chochowski, David Tarjan, Edoardo Maria Ponti |
ICML | 5 |
| 2024 | Towards Modular LLMs by Building and Reusing a Library of LoRAsabstractGiven the increasing number of parameter-efficient adapters of large language models (LLMs), how can we reuse them to improve LLM performance on new tasks? We study how to best build a library of adapters given multi-task data and devise techniques for both zero-shot and supervised task generalization through routing in such library. We benchmark existing approaches to build this library and introduce model-based clustering, $\texttt{MBC}$, a method that groups tasks based on the similarity of their adapter parameters, indirectly optimizing for transfer across the multi-task dataset. In order to reuse the library, we present a novel zero-shot routing mechanism, $\texttt{Arrow}$, which enables dynamic selection of the most relevant adapters for new inputs without the need for retraining. We experiment with several LLMs, such as Phi-2 and Mistral, on a wide array of held-out tasks, verifying that MBC-based adapters and Arrow routing lead to superior generalization to new tasks. Thus, we make steps towards creating modular, adaptable LLMs that can match or outperform traditional joint training. Oleksiy Ostapenko, Zhan Su 0002, Edoardo Maria Ponti, Laurent Charlin, Nicolas Le Roux, Lucas Caccia, Alessandro Sordoni |
ICML | 3 |
| 2024 | Elastic Weight Removal for Faithful and Abstractive Dialogue GenerationabstractNico Daheim, Nouha Dziri, Mrinmaya Sachan, Iryna Gurevych, Edoardo Ponti. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Nico Daheim, Nouha Dziri, Mrinmaya Sachan, Iryna Gurevych, Edoardo Maria Ponti |
NAACL-HLT | 5 |
| 2024 | Are Large Language Model Temporally Grounded?abstractYifu Qiu, Zheng Zhao, Yftah Ziser, Anna Korhonen, Edoardo Ponti, Shay Cohen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Yifu Qiu, Zheng Zhao 0005, Yftah Ziser, Anna Korhonen, Edoardo Maria Ponti, Shay B. Cohen |
NAACL-HLT | 5 |
| 2024 | Zero-Shot Tokenizer TransferabstractLanguage models (LMs) are bound to their tokenizer, which maps raw text to a sequence of vocabulary items (tokens). This restricts their flexibility: for example, LMs trained primarily on English may still perform well in other natural and programming languages, but have vastly decreased efficiency due to their English-centric tokenizer. To mitigate this, we should be able to swap the original LM tokenizer with an arbitrary one, on the fly, without degrading performance. Hence, in this work we define a new problem: Zero-Shot Tokenizer Transfer (ZeTT). The challenge at the core of ZeTT is finding embeddings for the tokens in the vocabulary of the new tokenizer. Since prior heuristics for initializing embeddings often perform at chance level in a ZeTT setting, we propose a new solution: we train a hypernetwork taking a tokenizer as input and predicting the corresponding embeddings. We empirically demonstrate that the hypernetwork generalizes to new tokenizers both with encoder (e.g., XLM-R) and decoder LLMs (e.g., Mistral-7B). Our method comes close to the original models' performance in cross-lingual and coding tasks while markedly reducing the length of the tokenized sequence. We also find that the remaining gap can be quickly closed by continued training on less than 1B tokens. Finally, we show that a ZeTT hypernetwork trained for a base (L)LM can also be applied to fine-tuned variants without extra training. Overall, our results make substantial strides toward detaching LMs from their tokenizer. Benjamin Minixhofer, Edoardo Maria Ponti, Ivan Vulic |
NeurIPS | 2 |
| 2024 | Spectral Editing of Activations for Large Language Model AlignmentabstractLarge language models (LLMs) often exhibit undesirable behaviours, such as generating untruthful or biased content. Editing their internal representations has been shown to be effective in mitigating such behaviours on top of the existing alignment methods. We propose a novel inference-time editing method, namely spectral editing of activations (SEA), to project the input representations into directions with maximal covariance with the positive demonstrations (e.g., truthful) while minimising covariance with the negative demonstrations (e.g., hallucinated). We also extend our method to non-linear editing using feature functions. We run extensive experiments on benchmarks concerning truthfulness and bias with six open-source LLMs of different sizes and model families. The results demonstrate the superiority of SEA in effectiveness, generalisation to similar tasks, as well as computation and data efficiency. We also show that SEA editing only has a limited negative impact on other model capabilities. Yifu Qiu, Zheng Zhao 0005, Yftah Ziser, Anna Korhonen, Edoardo Maria Ponti, Shay B. Cohen |
NeurIPS | 5 |
| 2023 | Efficient Transformers with Dynamic Token PoolingabstractTransformers achieve unrivalled performance in modelling language, but remain inefficient in terms of memory and time complexity.A possible remedy is to reduce the sequence length in the intermediate layers by pooling fixed-length segments of tokens.Nevertheless, natural units of meaning, such as words or phrases, display varying sizes.To address this mismatch, we equip language models with a dynamic-pooling mechanism, which predicts segment boundaries in an autoregressive fashion.We compare several methods to infer boundaries, including end-to-end learning through stochastic re-parameterisation, supervised learning (based on segmentations from subword tokenizers or spikes in conditional entropy), as well as linguistically motivated boundaries.We perform character-level evaluation on texts from multiple datasets and morphologically diverse languages.The results demonstrate that dynamic pooling, which jointly segments and models language, is both faster and more accurate than vanilla Transformers and fixed-length pooling within the same computational budget. Piotr Nawrot, Jan Chorowski, Adrian Lancucki, Edoardo Maria Ponti |
ACL (1) | 4 |
| 2023 | Combining Parameter-efficient Modules for Task-level GeneralisationabstractA modular design encourages neural models to disentangle and recombine different facets of knowledge to generalise more systematically to new tasks.In this work, we assume that each task is associated with a subset of latent skills from an (arbitrary size) inventory.In turn, each skill corresponds to a parameter-efficient (sparse / low-rank) model adapter.By jointly learning adapters and a routing function that allocates skills to each task, the full network is instantiated as the average of the parameters of active skills.We propose several inductive biases that encourage re-usage and composition of the skills, including variable-size skill allocation and a dual-speed learning rate.We evaluate our latent-skill model in two main settings: 1) multitask reinforcement learning for instruction following on 8 levels of the BabyAI platform; and 2) few-shot fine-tuning of language models on 160 NLP tasks of the CrossFit benchmark.We find that the modular design of our network enhances sample efficiency in reinforcement learning and few-shot generalisation in supervised learning, compared to a series of baselines.These include models where parameters are fully shared, task-specific, or conditionally generated (HyperFormer), as well as sparse mixture-of-experts (Task-MoE). Edoardo Maria Ponti, Alessandro Sordoni, Yoshua Bengio, Siva Reddy |
EACL | 1 |
| 2023 | Probing Cross-Lingual Lexical Knowledge from Multilingual Sentence EncodersabstractIvan Vulić, Goran Glavaš, Fangyu Liu, Nigel Collier, Edoardo Maria Ponti, Anna Korhonen. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Ivan Vulic, Goran Glavas, Fangyu Liu 0001, Nigel Collier, Edoardo Maria Ponti, Anna Korhonen |
EACL | 5 |
| 2023 | Unifying Cross-Lingual Transfer across Scenarios of Resource ScarcityabstractThe scarcity of data in many of the world's languages necessitates the transfer of knowledge from other, resource-rich languages.However, the level of scarcity varies significantly across multiple dimensions, including: i) the amount of task-specific data available in the source and target languages; ii) the amount of monolingual and parallel data available for both languages; and iii) the extent to which they are supported by pretrained multilingual and translation models.Prior work has largely treated these dimensions and the various techniques for dealing with them separately; in this paper, we offer a more integrated view by exploring how to deploy the arsenal of cross-lingual transfer tools across a range of scenarios, especially the most challenging, low-resource ones.To this end, we run experiments on the Americas-NLI and NusaX benchmarks over 20 languages, simulating a range of few-shot settings.The best configuration in our experiments employed parameter-efficient language and task adaptation of massively multilingual Transformers, trained simultaneously on source language data and both machine-translated and natural data for multiple target languages.In addition, we show that pre-trained translation models can be easily adapted to unseen languages, thus extending the range of our hybrid technique and translation-based transfer more broadly.Beyond new insights into the mechanisms of cross-lingual transfer, we hope our work will provide practitioners with a toolbox to integrate multiple techniques for different real-world scenarios.Our code is available at https: //github.com/parovicm/unified-xlt. Alan Ansell, Marinela Parovic, Ivan Vulic, Anna Korhonen, Edoardo Maria Ponti |
EMNLP | 5 |
| 2023 | Detecting and Mitigating Hallucinations in Multilingual SummarisationabstractHallucinations pose a significant challenge to the reliability of neural models for abstractive summarisation.While automatically generated summaries may be fluent, they often lack faithfulness to the original document.This issue becomes even more pronounced in lowresource languages, where summarisation requires cross-lingual transfer.With the existing faithful metrics focusing on English, even measuring the extent of this phenomenon in crosslingual settings is hard.To address this, we first develop a novel metric, mFACT, evaluating the faithfulness of non-English summaries, leveraging translation-based transfer from multiple English faithfulness metrics.Through extensive experiments in multiple languages, we demonstrate that mFACT is best suited to detect hallucinations compared to alternative metrics.With mFACT, we assess a broad range of multilingual large language models, and find that they all tend to hallucinate often in languages different from English.We then propose a simple but effective method to reduce hallucinations in cross-lingual transfer, which weighs the loss of each training example by its faithfulness score.This method drastically increases both performance and faithfulness according to both automatic and human evaluation when compared to strong baselines for cross-lingual transfer such as MAD-X. Yifu Qiu, Yftah Ziser, Anna Korhonen, Edoardo Maria Ponti, Shay B. Cohen |
EMNLP | 4 |
| 2023 | Multi-Head Adapter Routing for Cross-Task GeneralizationabstractParameter-efficient fine-tuning (PEFT) for cross-task generalization consists in pre-training adapters on a multi-task training set before few-shot adaptation to test tasks. Polytropon [Ponti et al., 2023] ($\texttt{Poly}$) jointly learns an inventory of adapters and a *routing* function that selects a (variable-size) subset of adapters for each task during both pre-training and few-shot adaptation. In this paper, we investigate the role that adapter routing plays in its success and design new variants based on our findings.
First, we build on the intuition that finer-grained routing provides more expressivity. Hence,
we propose $\texttt{MHR}$ (Multi-Head Routing) which combines *subsets* of adapter parameters and outperforms $\texttt{Poly}$ under a comparable parameter budget; by only fine-tuning the routing function and not the adapters ($\texttt{MHR}$-$z$) we achieve competitive performance with extreme parameter efficiency. Second, we find that $\texttt{Poly}$/$\texttt{MHR}$ performance is a result of better multi-task optimization, rather than modular inductive biases that facilitate adapter recombination and local adaptation, as previously hypothesized. In fact, we find that $\texttt{MHR}$ exhibits high gradient alignment between training tasks. We find that routing is most beneficial during multi-task pre-training rather than during few-shot adaptation and propose $\texttt{MHR}$-$\mu$, which discards routing and fine-tunes the average of the pre-trained adapters on each downstream tasks. This establishes $\texttt{MHR}$-$\mu$ as an effective method for single-adapter fine-tuning. We also show that $\texttt{MHR}$-$\mu$ can be used as an effective zero-shot transfer method by training the average of the pre-trained adapters for a few additional steps on the multi-task training set: this yields gains up to 3\% on absolute accuracy w.r.t. the baselines. Code is available at <https://github.com/microsoft/mttl>. Lucas Caccia, Edoardo Maria Ponti, Zhan Su 0002, Matheus Pereira, Nicolas Le Roux, Alessandro Sordoni |
NeurIPS | 2 |
| 2023 | Cross-Lingual Dialogue Dataset Creation via Outline-Based GenerationabstractAbstract Multilingual task-oriented dialogue (ToD) facilitates access to services and information for many (communities of) speakers. Nevertheless, its potential is not fully realized, as current multilingual ToD datasets—both for modular and end-to-end modeling—suffer from severe limitations. 1) When created from scratch, they are usually small in scale and fail to cover many possible dialogue flows. 2) Translation-based ToD datasets might lack naturalness and cultural specificity in the target language. In this work, to tackle these limitations we propose a novel outline-based annotation process for multilingual ToD datasets, where domain-specific abstract schemata of dialogue are mapped into natural language outlines. These in turn guide the target language annotators in writing dialogues by providing instructions about each turn’s intents and slots. Through this process we annotate a new large-scale dataset for evaluation of multilingual and cross-lingual ToD systems. Our Cross-lingual Outline-based Dialogue dataset (cod) enables natural language understanding, dialogue state tracking, and end-to-end dialogue evaluation in 4 diverse languages: Arabic, Indonesian, Russian, and Kiswahili. Qualitative and quantitative analyses of cod versus an equivalent translation-based dataset demonstrate improvements in data quality, unlocked by the outline-based approach. Finally, we benchmark a series of state-of-the-art systems for cross-lingual ToD, setting reference scores for future work and demonstrating that cod prevents over-inflated performance, typically met with prior translation-based ToD datasets. Olga Majewska, Evgeniia Razumovskaia, Edoardo Maria Ponti, Ivan Vulic, Anna Korhonen |
Trans. Assoc. Comput. Linguistics | 3 |
| 2022 | Composable Sparse Fine-Tuning for Cross-Lingual TransferabstractFine-tuning the entire set of parameters of a large pretrained model has become the mainstream approach for transfer learning.To increase its efficiency and prevent catastrophic forgetting and interference, techniques like adapters and sparse fine-tuning have been developed.Adapters are modular, as they can be combined to adapt a model towards different facets of knowledge (e.g., dedicated language and/or task adapters).Sparse finetuning is expressive, as it controls the behavior of all model components.In this work, we introduce a new fine-tuning method with both these desirable properties.In particular, we learn sparse, real-valued masks based on a simple variant of the Lottery Ticket Hypothesis.Task-specific masks are obtained from annotated data in a source language, and languagespecific masks from masked language modeling in a target language.Both these masks can then be composed with the pretrained model.Unlike adapter-based fine-tuning, this method neither increases the number of parameters at inference time nor alters the original model architecture.Most importantly, it outperforms adapters in zero-shot cross-lingual transfer by a large margin in a series of multilingual benchmarks, including Universal Dependencies, MasakhaNER, and AmericasNLI.Based on an in-depth analysis, we additionally find that sparsity is crucial to prevent both 1) interference between the fine-tunings to be composed and 2) overfitting.We release the code and models at https://github.com/ cambridgeltl/composable-sft. Alan Ansell, Edoardo Maria Ponti, Anna Korhonen, Ivan Vulic |
ACL (1) | 2 |
| 2022 | Image Retrieval from Contextual DescriptionsabstractBenno Krojer, Vaibhav Adlakha, Vibhav Vineet, Yash Goyal, Edoardo Ponti, Siva Reddy. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Benno Krojer, Vaibhav Adlakha, Vibhav Vineet, Yash Goyal, Edoardo Maria Ponti, Siva Reddy |
ACL (1) | 5 |
| 2022 | IGLUE: A Benchmark for Transfer Learning across Modalities, Tasks, and LanguagesabstractReliable evaluation benchmarks designed for replicability and comprehensiveness have driven progress in machine learning. Due to the lack of a multilingual benchmark, however, vision-and-language research has mostly focused on English language tasks. To fill this gap, we introduce the Image-Grounded Language Understanding Evaluation benchmark. IGLUE brings together{—}by both aggregating pre-existing datasets and creating new ones{—}visual question answering, cross-modal retrieval, grounded reasoning, and grounded entailment tasks across 20 diverse languages. Our benchmark enables the evaluation of multilingual multimodal models for transfer learning, not only in a zero-shot setting, but also in newly defined few-shot learning setups. Based on the evaluation of the available state-of-the-art models, we find that translate-test transfer is superior to zero-shot transfer and that few-shot learning is hard to harness for many tasks. Moreover, downstream performance is partially explained by the amount of available unlabelled textual data for pretraining, and only weakly by the typological distance of target{–}source languages. We hope to encourage future research efforts in this area by releasing the benchmark to the community. Emanuele Bugliarello, Fangyu Liu 0001, Jonas Pfeiffer, Siva Reddy, Desmond Elliott, Edoardo Maria Ponti, Ivan Vulic |
ICML | 6 |
| 2022 | UniMorph 4.0: Universal MorphologyabstractThe Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological inflection tables for hundreds of diverse world languages. The project comprises two major thrusts: a language-independent feature schema for rich morphological annotation, and a type-level resource of annotated data in diverse languages realizing that schema. This paper presents the expansions and improvements on several fronts that were made in the last couple of years (since McCarthy et al. (2020)). Collaborative efforts by numerous linguists have added 66 new languages, including 24 endangered languages. We have implemented several improvements to the extraction pipeline to tackle some issues, e.g., missing gender and macrons information. We have amended the schema to use a hierarchical structure that is needed for morphological phenomena like multiple-argument agreement and case stacking, while adding some missing morphological features to make the schema more inclusive. In light of the last UniMorph release, we also augmented the database with morpheme segmentation for 16 languages. Lastly, this new release makes a push towards inclusion of derivational morphology in UniMorph by enriching the data and annotation schema with instances representing derivational processes from MorphyNet. Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieras, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina J. Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Lane 0002, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóga, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer C. White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo Maria Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar 0002, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Tucker Prud'hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, Ekaterina Vylomova |
LREC | 73 |
| 2022 | Same Neurons, Different Languages: Probing Morphosyntax in Multilingual Pre-trained ModelsabstractKarolina Stanczak, Edoardo Ponti, Lucas Torroba Hennigen, Ryan Cotterell, Isabelle Augenstein. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Karolina Stanczak, Edoardo Maria Ponti, Lucas Torroba Hennigen, Ryan Cotterell, Isabelle Augenstein |
NAACL-HLT | 2 |
| 2022 | Crossing the Conversational Chasm: A Primer on Natural Language Processing for Multilingual Task-Oriented Dialogue SystemsabstractIn task-oriented dialogue (ToD), a user holds a conversation with an artificial agent with the aim of completing a concrete task. Although this technology represents one of the central objectives of AI and has been the focus of ever more intense research and development efforts, it is currently limited to a few narrow domains (e.g., food ordering, ticket booking) and a handful of languages (e.g., English, Chinese). This work provides an extensive overview of existing methods and resources in multilingual ToD as an entry point to this exciting and emerging field. We find that the most critical factor preventing the creation of truly multilingual ToD systems is the lack of datasets in most languages for both training and evaluation. In fact, acquiring annotations or human feedback for each component of modular systems or for data-hungry end-to-end systems is expensive and tedious. Hence, state-of-the-art approaches to multilingual ToD mostly rely on (zero- or few-shot) cross-lingual transfer from resource-rich languages (almost exclusively English), either by means of (i) machine translation or (ii) multilingual representations. These approaches are currently viable only for typologically similar languages and languages with parallel / monolingual corpora available. On the other hand, their effectiveness beyond these boundaries is doubtful or hard to assess due to the lack of linguistically diverse benchmarks (especially for natural language generation and end-to-end evaluation). To overcome this limitation, we draw parallels between components of the ToD pipeline and other NLP tasks, which can inspire solutions for learning in low-resource scenarios. Finally, we list additional challenges that multilinguality poses for related areas (such as speech, fluency in generated text, and human-centred evaluation), and indicate future directions that hold promise to further expand language coverage and dialogue capabilities of current ToD systems. Evgeniia Razumovskaia, Goran Glavas, Olga Majewska, Edoardo Maria Ponti, Anna Korhonen, Ivan Vulic |
J. Artif. Intell. Res. | 4 |
| 2022 | FaithDial: A Faithful Benchmark for Information-Seeking DialogueabstractAbstract The goal of information-seeking dialogue is to respond to seeker queries with natural language utterances that are grounded on knowledge sources. However, dialogue systems often produce unsupported utterances, a phenomenon known as hallucination. To mitigate this behavior, we adopt a data-centric solution and create FaithDial, a new benchmark for hallucination-free dialogues, by editing hallucinated responses in the Wizard of Wikipedia (WoW) benchmark. We observe that FaithDial is more faithful than WoW while also maintaining engaging conversations. We show that FaithDial can serve as training signal for: i) a hallucination critic, which discriminates whether an utterance is faithful or not, and boosts the performance by 12.8 F1 score on the BEGIN benchmark compared to existing datasets for dialogue coherence; ii) high-quality dialogue generation. We benchmark a series of state-of-the-art models and propose an auxiliary contrastive objective that achieves the highest level of faithfulness and abstractiveness based on several automated metrics. Further, we find that the benefits of FaithDial generalize to zero-shot transfer on other datasets, such as CMU-Dog and TopicalChat. Finally, human evaluation reveals that responses generated by models trained on FaithDial are perceived as more interpretable, cooperative, and engaging. Nouha Dziri, Ehsan Kamalloo, Sivan Milton, Osmar R. Zaïane, Mo Yu, Edoardo Maria Ponti, Siva Reddy |
Trans. Assoc. Comput. Linguistics | 6 |
| 2021 | Verb Knowledge Injection for Multilingual Event ProcessingabstractOlga Majewska, Ivan Vulić, Goran Glavaš, Edoardo Maria Ponti, Anna Korhonen. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Olga Majewska, Ivan Vulic, Goran Glavas, Edoardo Maria Ponti, Anna Korhonen |
ACL/IJCNLP (1) | 4 |
| 2021 | LexFit: Lexical Fine-Tuning of Pretrained Language ModelsabstractIvan Vulić, Edoardo Maria Ponti, Anna Korhonen, Goran Glavaš. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ivan Vulic, Edoardo Maria Ponti, Anna Korhonen, Goran Glavas |
ACL/IJCNLP (1) | 2 |
| 2021 | Visually Grounded Reasoning across Languages and CulturesabstractThe design of widespread vision-and-language datasets and pre-trained encoders directly adopts, or draws inspiration from, the concepts and images of ImageNet. While one can hardly overestimate how much this benchmark contributed to progress in computer vision, it is mostly derived from lexical databases and image queries in English, resulting in source material with a North American or Western European bias. Therefore, we devise a new protocol to construct an ImageNet-style hierarchy representative of more languages and cultures. In particular, we let the selection of both concepts and images be entirely driven by native speakers, rather than scraping them automatically. Specifically, we focus on a typologically diverse set of languages, namely, Indonesian, Mandarin Chinese, Swahili, Tamil, and Turkish. On top of the concepts and images obtained through this new protocol, we create a multilingual dataset for Multicultural Reasoning over Vision and Language (MaRVL) by eliciting statements from native speaker annotators about pairs of images. The task consists of discriminating whether each grounded statement is true or false. We establish a series of baselines using state-of-the-art models and find that their cross-lingual transfer performance lags dramatically behind supervised performance in English. These results invite us to reassess the robustness and accuracy of current state-of-the-art models beyond a narrow domain, but also open up new exciting challenges for the development of truly multilingual and multicultural systems. Fangyu Liu 0001, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy, Nigel Collier, Desmond Elliott |
EMNLP (1) | 3 |
| 2021 | AM2iCo: Evaluating Word Meaning in Context across Low-Resource Languages with Adversarial ExamplesabstractCapturing word meaning in context and distinguishing between correspondences and variations across languages is key to building successful multilingual and cross-lingual text representation models.However, existing multilingual evaluation datasets that evaluate lexical semantics "in-context" have various limitations.In particular, 1) their language coverage is restricted to high-resource languages and skewed in favor of only a few language families and areas, 2) a design that makes the task solvable via superficial cues, which results in artificially inflated (and sometimes super-human) performances of pretrained encoders, and 3) no support for crosslingual evaluation.In order to address these gaps, we present AM 2 ICO (Adversarial and Multilingual Meaning in Context), a widecoverage cross-lingual and multilingual evaluation set; it aims to faithfully assess the ability of state-of-the-art (SotA) representation models to understand the identity of word meaning in cross-lingual contexts for 14 language pairs.We conduct a series of experiments in a wide range of setups and demonstrate the challenging nature of AM 2 ICO.The results reveal that current SotA pretrained encoders substantially lag behind human performance, and the largest gaps are observed for low-resource languages and languages dissimilar to English. Qianchu Liu, Edoardo Maria Ponti, Diana McCarthy, Ivan Vulic, Anna Korhonen |
EMNLP (1) | 2 |
| 2021 | Parameter Space Factorization for Zero-Shot Learning across Tasks and LanguagesabstractAbstract Most combinations of NLP tasks and language varieties lack in-domain examples for supervised training because of the paucity of annotated data. How can neural models make sample-efficient generalizations from task–language combinations with available data to low-resource ones? In this work, we propose a Bayesian generative model for the space of neural parameters. We assume that this space can be factorized into latent variables for each language and each task. We infer the posteriors over such latent variables based on data from seen task–language combinations through variational inference. This enables zero-shot classification on unseen combinations at prediction time. For instance, given training data for named entity recognition (NER) in Vietnamese and for part-of-speech (POS) tagging in Wolof, our model can perform accurate predictions for NER in Wolof. In particular, we experiment with a typologically diverse sample of 33 languages from 4 continents and 11 families, and show that our model yields comparable or better results than state-of-the-art, zero-shot cross-lingual transfer methods. Our code is available at github.com/cambridgeltl/parameter-factorization. Edoardo Maria Ponti, Ivan Vulic, Ryan Cotterell, Marinela Parovic, Roi Reichart, Anna Korhonen |
Trans. Assoc. Comput. Linguistics | 1 |
| 2020 | Specializing Unsupervised Pretraining Models for Word-Level Semantic SimilarityabstractUnsupervised pretraining models have been shown to facilitate a wide range of downstream NLP applications. These models, however, retain some of the limitations of traditional static word embeddings. In particular, they encode only the distributional knowledge available in raw text corpora, incorporated through language modeling objectives. In this work, we complement such distributional knowledge with external lexical knowledge, that is, we integrate the discrete knowledge on word-level semantic similarity into pretraining. To this end, we generalize the standard BERT model to a multi-task learning setting where we couple BERT’s masked language modeling and next sentence prediction objectives with an auxiliary task of binary word relation classification. Our experiments suggest that our "Lexically Informed” BERT (LIBERT), specialized for the word-level semantic similarity, yields better performance than the lexically blind “vanilla” BERT on several language understanding tasks. Concretely, LIBERT outperforms BERT in 9 out of 10 tasks of the GLUE benchmark and is on a par with BERT in the remaining one. Moreover, we show consistent gains on 3 benchmarks for lexical simplification, a task where knowledge about word-level semantic similarity is paramount, as well as large gains on lexical reasoning probes. Anne Lauscher, Ivan Vulic, Edoardo Maria Ponti, Anna Korhonen, Goran Glavas |
COLING | 3 |
| 2020 | Emergent Communication Pretraining for Few-Shot Machine TranslationabstractWhile state-of-the-art models that rely upon massively multilingual pretrained encoders achieve sample efficiency in downstream applications, they still require abundant amounts of unlabelled text.Nevertheless, most of the world's languages lack such resources.Hence, we investigate a more radical form of unsupervised knowledge transfer in the absence of linguistic data.In particular, for the first time we pretrain neural networks via emergent communication from referential games.Our key assumption is that grounding communication on images-as a crude approximation of real-world environments-inductively biases the model towards learning natural languages.On the one hand, we show that this substantially benefits machine translation in few-shot settings.On the other hand, this also provides an extrinsic evaluation protocol to probe the properties of emergent languages ex vitro.Intuitively, the closer they are to natural languages, the higher the gains from pretraining on them should be.For instance, in this work we measure the influence of communication success and maximum sequence length on downstream performances.Finally, we introduce a customised adapter layer and annealing strategies for the regulariser of maximum-a-posteriori inference during fine-tuning.These turn out to be crucial to facilitate knowledge transfer and prevent catastrophic forgetting.Compared to a recurrent baseline, our method yields gains of 59.0%∼147.6% in BLEU score with only 500 NMT training instances and 65.1%∼196.7%with 1, 000 NMT training instances across four language pairs.These proofof-concept results reveal the potential of emergent communication pretraining for both natural language processing tasks in resource-poor settings and extrinsic evaluation of artificial languages. Yaoyiran Li, Edoardo Maria Ponti, Ivan Vulic, Anna Korhonen |
COLING | 2 |
| 2020 | XCOPA: A Multilingual Dataset for Causal Commonsense ReasoningabstractIn order to simulate human language capacity, natural language processing systems must be able to reason about the dynamics of everyday situations, including their possible causes and effects. Moreover, they should be able to generalise the acquired world knowledge to new languages, modulo cultural differences. Advances in machine reasoning and cross-lingual transfer depend on the availability of challenging evaluation benchmarks. Motivated by both demands, we introduce Cross-lingual Choice of Plausible Alternatives (XCOPA), a typologically diverse multilingual dataset for causal commonsense reasoning in 11 languages, which includes resource-poor languages like Eastern Apurímac Quechua and Haitian Creole. We evaluate a range of state-of-the-art models on this novel dataset, revealing that the performance of current methods based on multilingual pretraining and zero-shot fine-tuning falls short compared to translation-based transfer. Finally, we propose strategies to adapt multilingual models to out-of-sample resource-lean languages where only a small corpus or a bilingual dictionary is available, and report substantial improvements over the random baseline. The XCOPA dataset is freely available at github.com/cambridgeltl/xcopa Edoardo Maria Ponti, Goran Glavas, Olga Majewska, Qianchu Liu, Ivan Vulic, Anna Korhonen |
EMNLP (1) | 1 |
| 2020 | Probing Pretrained Language Models for Lexical SemanticsabstractThe success of large pretrained language models (LMs) such as BERT and RoBERTa has sparked interest in probing their representations, in order to unveil what types of knowledge they implicitly capture.While prior research focused on morphosyntactic, semantic, and world knowledge, it remains unclear to which extent LMs also derive lexical type-level knowledge from words in context.In this work, we present a systematic empirical analysis across six typologically diverse languages and five different lexical tasks, addressing the following questions: 1) How do different lexical knowledge extraction strategies (monolingual versus multilingual source LM, out-ofcontext versus in-context encoding, inclusion of special tokens, and layer-wise averaging) impact performance?How consistent are the observed effects across tasks and languages?2) Is lexical knowledge stored in few parameters, or is it scattered throughout the network?3) How do these representations fare against traditional static word vectors in lexical tasks?4) Does the lexical information emerging from independently trained monolingual LMs display latent similarities?Our main results indicate patterns and best practices that hold universally, but also point to prominent variations across languages and tasks.Moreover, we validate the claim that lower Transformer layers carry more type-level lexical knowledge, but also show that this knowledge is distributed across multiple layers. Ivan Vulic, Edoardo Maria Ponti, Robert Litschko, Goran Glavas, Anna Korhonen |
EMNLP (1) | 2 |
| 2020 | Multi-SimLex: A Large-Scale Evaluation of Multilingual and Crosslingual Lexical Semantic SimilarityabstractWe introduce Multi-SimLex, a large-scale lexical resource and evaluation benchmark covering data sets for 12 typologically diverse languages, including major languages (e.g., Mandarin Chinese, Spanish, Russian) as well as less-resourced ones (e.g., Welsh, Kiswahili). Each language data set is annotated for the lexical relation of semantic similarity and contains 1,888 semantically aligned concept pairs, providing a representative coverage of word classes (nouns, verbs, adjectives, adverbs), frequency ranks, similarity intervals, lexical fields, and concreteness levels. Additionally, owing to the alignment of concepts across languages, we provide a suite of 66 crosslingual semantic similarity data sets. Because of its extensive size and language coverage, Multi-SimLex provides entirely novel opportunities for experimental evaluation and analysis. On its monolingual and crosslingual benchmarks, we evaluate and analyze a wide array of recent state-of-the-art monolingual and crosslingual representation models, including static and contextualized word embeddings (such as fastText, monolingual and multilingual BERT, XLM), externally informed lexical representations, as well as fully unsupervised and (weakly) supervised crosslingual word embeddings. We also present a step-by-step data set creation protocol for creating consistent, Multi-Simlex–style resources for additional languages. We make these contributions—the public release of Multi-SimLex data sets, their creation protocol, strong baseline results, and in-depth analyses which can be helpful in guiding future developments in multilingual lexical semantics and representation learning—available via a Web site that will encourage community effort in further expansion of Multi-Simlex to many more languages. Such a large-scale semantic resource could inspire significant further advances in NLP across languages. Ivan Vulic, Simon Baker, Edoardo Maria Ponti, Ulla Petti, Ira Leviant, Kelly Wing, Olga Majewska, Eden Bar, Matt Malone, Thierry Poibeau, Roi Reichart, Anna Korhonen |
Comput. Linguistics | 3 |
| 2019 | Towards Zero-shot Language ModelingabstractEdoardo Maria Ponti, Ivan Vulić, Ryan Cotterell, Roi Reichart, Anna Korhonen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Edoardo Maria Ponti, Ivan Vulic, Ryan Cotterell, Roi Reichart, Anna Korhonen |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Cross-lingual Semantic Specialization via Lexical Relation InductionabstractEdoardo Maria Ponti, Ivan Vulić, Goran Glavaš, Roi Reichart, Anna Korhonen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Edoardo Maria Ponti, Ivan Vulic, Goran Glavas, Roi Reichart, Anna Korhonen |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Modeling Language Variation and Universals: A Survey on Typological Linguistics for Natural Language ProcessingabstractLinguistic typology aims to capture structural and semantic variation across the world’s languages. A large-scale typology could provide excellent guidance for multilingual Natural Language Processing (NLP), particularly for languages that suffer from the lack of human labeled resources. We present an extensive literature survey on the use of typological information in the development of NLP techniques. Our survey demonstrates that to date, the use of information in existing typological databases has resulted in consistent but modest improvements in system performance. We show that this is due to both intrinsic limitations of databases (in terms of coverage and feature granularity) and under-utilization of the typological features included in them. We advocate for a new approach that adapts the broad and discrete nature of typological categories to the contextual and continuous nature of machine learning algorithms used in contemporary NLP. In particular, we suggest that such an approach could be facilitated by recent developments in data-driven induction of typological knowledge. Edoardo Maria Ponti, Helen O'Horan, Yevgeni Berzak, Ivan Vulic, Roi Reichart, Thierry Poibeau, Ekaterina Shutova, Anna Korhonen |
Comput. Linguistics | 1 |
| 2018 | Isomorphic Transfer of Syntactic Structures in Cross-Lingual NLPabstractThe transfer or share of knowledge between languages is a popular solution to resource scarcity in NLP.However, the effectiveness of cross-lingual transfer can be challenged by variation in syntactic structures.Frameworks such as Universal Dependencies (UD) are designed to be cross-lingually consistent, but even in carefully designed resources trees representing equivalent sentences may not always overlap.In this paper, we measure cross-lingual syntactic variation, or anisomorphism, in the UD treebank collection, considering both morphological and structural properties.We show that reducing the level of anisomorphism yields consistent gains in cross-lingual transfer tasks.We introduce a source language selection procedure that facilitates effective cross-lingual parser transfer, and propose a typologically driven method for syntactic tree processing which reduces anisomorphism.Our results show the effectiveness of this method for both machine translation and cross-lingual sentence similarity, demonstrating the importance of syntactic structure compatibility for boosting cross-lingual transfer in NLP. Edoardo Maria Ponti, Roi Reichart, Anna Korhonen, Ivan Vulic |
ACL (1) | 1 |
| 2018 | On the Relation between Linguistic Typology and (Limitations of) Multilingual Language ModelingabstractA key challenge in cross-lingual NLP is developing general language-independent architectures that are equally applicable to any language.However, this ambition is largely hampered by the variation in structural and semantic properties, i.e. the typological profiles of the world's languages.In this work, we analyse the implications of this variation on the language modeling (LM) task.We present a largescale study of state-of-the art n-gram based and neural language models on 50 typologically diverse languages covering a wide variety of morphological systems.Operating in the full vocabulary LM setup focused on wordlevel prediction, we demonstrate that a coarse typology of morphological systems is predictive of absolute LM performance.Moreover, fine-grained typological features such as exponence, flexivity, fusion, and inflectional synthesis are borne out to be responsible for the proliferation of low-frequency phenomena which are organically difficult to model by statistical architectures, or for the meaning ambiguity of character n-grams.Our study strongly suggests that these features have to be taken into consideration during the construction of nextlevel language-agnostic LM architectures, capable of handling morphologically complex languages such as Tamil or Korean. Daniela Gerz, Ivan Vulic, Edoardo Maria Ponti, Roi Reichart, Anna Korhonen |
EMNLP | 3 |
| 2018 | Adversarial Propagation and Zero-Shot Cross-Lingual Transfer of Word Vector SpecializationabstractSemantic specialization is a process of finetuning pre-trained distributional word vectors using external lexical knowledge (e.g., Word-Net) to accentuate a particular semantic relation in the specialized vector space.While post-processing specialization methods are applicable to arbitrary distributional vectors, they are limited to updating only the vectors of words occurring in external lexicons (i.e., seen words), leaving the vectors of all other words unchanged.We propose a novel approach to specializing the full distributional vocabulary.Our adversarial post-specialization method propagates the external lexical knowledge to the full distributional space.We exploit words seen in the resources as training examples for learning a global specialization function.This function is learned by combining a standard L 2 -distance loss with a adversarial loss: the adversarial component produces more realistic output vectors.We show the effectiveness and robustness of the proposed method across three languages and on three tasks: word similarity, dialog state tracking, and lexical simplification.We report consistent improvements over distributional word vectors and vectors specialized by other state-of-the-art specialization frameworks.Finally, we also propose a cross-lingual transfer method for zero-shot specialization which successfully specializes a full target distributional space without any lexical knowledge in the target language and without any bilingual data. Edoardo Maria Ponti, Ivan Vulic, Goran Glavas, Nikola Mrksic, Anna Korhonen |
EMNLP | 1 |
| 2018 | Language Modeling for Morphologically Rich Languages: Character-Aware Modeling for Word-Level PredictionabstractNeural architectures are prominent in the construction of language models (LMs). However, word-level prediction is typically agnostic of subword-level information (characters and character sequences) and operates over a closed vocabulary, consisting of a limited word set. Indeed, while subword-aware models boost performance across a variety of NLP tasks, previous work did not evaluate the ability of these models to assist next-word prediction in language modeling tasks. Such subword-level informed models should be particularly effective for morphologically-rich languages (MRLs) that exhibit high type-to-token ratios. In this work, we present a large-scale LM study on 50 typologically diverse languages covering a wide variety of morphological systems, and offer new LM benchmarks to the community, while considering subword-level information. The main technical contribution of our work is a novel method for injecting subword-level information into semantic word vectors, integrated into the neural language modeling training, to facilitate word-level prediction. We conduct experiments in the LM setting where the number of infrequent words is large, and demonstrate strong perplexity gains across our 50 languages, especially for morphologically-rich languages. Our code and data sets are publicly available. Daniela Gerz, Ivan Vulic, Edoardo Maria Ponti, Jason Naradowsky, Roi Reichart, Anna Korhonen |
Trans. Assoc. Comput. Linguistics | 3 |
| 2016 | Differentia compositionem facit. A Slower-Paced and Reliable Parser for Latin
Edoardo Maria Ponti, Marco Passarotti |
LREC | 1 |