EDBT 2026 Demo / reviewers in the wild / expert
Ivan Vulic
dblp:77/9768
· DBLP profile ↗
137ranked-venue papers
28as first author
62since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 120 · 23 first-author · 59 since 2021Databases, data management, data science and information retrieval · 15 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Value of Information: A Framework for Human-Agent CommunicationabstractYijiang River Dong, Tiancheng Hu, Zheng Hui, Caiqi Zhang, Ivan Vulić, Andreea Bobu, Nigel Collier. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yijiang River Dong, Tiancheng Hu, Zheng Hui, Caiqi Zhang, Ivan Vulic, Andreea Bobu, Nigel Collier |
ACL (1) | 5 |
| 2025 | Language Fusion for Parameter-Efficient Cross-lingual TransferabstractLimited availability of multilingual text corpora for training language models often leads to poor performance on downstream tasks due to undertrained representation spaces for languages other than English.This 'under-representation' has motivated recent cross-lingual transfer methods to leverage the English representation space by e.g.mixing English and 'non-English' tokens at the input level or extending model parameters to accommodate new languages.However, these approaches often come at the cost of increased computational complexity.We propose Fusion for Language Representations (FLARE) in adapters, a novel method that enhances representation quality and downstream performance for languages other than English while maintaining parameter efficiency.FLARE integrates source and target language representations within low-rank (LoRA) adapters using lightweight linear transformations, maintaining parameter efficiency while improving transfer performance.A series of experiments across representative crosslingual natural language understanding tasks, including natural language inference, questionanswering and sentiment analysis, demonstrate FLARE's effectiveness.FLARE achieves performance improvements of 4.9% for Llama 3.1 and 2.2% for Gemma 2 compared to standard LoRA fine-tuning on question-answering tasks, as measured by the exact match metric. 1 method for cross-lingual language understanding. Philipp Borchert, Ivan Vulic, Marie-Francine Moens, Jochen De Weerdt |
ACL (1) | 2 |
| 2025 | Retrofitting Large Language Models with Dynamic TokenizationabstractCurrent language models (LMs) use a fixed, static subword tokenizer.This default choice typically results in degraded efficiency and language capabilities, especially in languages other than English.To address this issue, we challenge the static design and propose retrofitting LMs with dynamic tokenization: a way to dynamically decide on token boundaries based on the input text via a subwordmerging algorithm inspired by byte-pair encoding.We merge frequent subword sequences in a batch, then apply a pre-trained embeddingprediction hypernetwork to compute the token embeddings on-the-fly.For encoder-style models (e.g., XLM-R), this on average reduces token sequence lengths by >20% across 14 languages while degrading performance by less than 2%.The same method applied to prefilling and scoring in decoder-style models (e.g., Mistral-7B) results in minimal performance degradation at up to 17% reduction in sequence length.Overall, we find that dynamic tokenization can mitigate the limitations of static tokenization by substantially improving inference speed and promoting fairness across languages, enabling more equitable and adaptable LMs. Darius Feher, Ivan Vulic, Benjamin Minixhofer |
ACL (1) | 2 |
| 2025 | Quantifying Language Disparities in Multilingual Large Language ModelsabstractResults reported in large-scale multilingual evaluations are often fragmented and confounded by factors such as target languages, differences in experimental setups, and model choices.We propose a framework that disentangles these confounding variables and introduces three interpretable metrics-the performance realisation ratio, its coefficient of variation, and language potential-enabling a finergrained and more insightful quantification of actual performance disparities across both (i) models and (ii) languages.Through a case study of 13 model variants on 11 multilingual datasets, we demonstrate that our framework provides a more reliable measurement of model performance and language disparities, particularly for low-resource languages, which have so far proven challenging to evaluate.Importantly, our results reveal that higher overall model performance does not necessarily imply greater fairness across languages. Songbo Hu, Ivan Vulic, Anna Korhonen |
EMNLP | 2 |
| 2025 | Imagine While Reasoning in Space: Multimodal Visualization-of-ThoughtabstractChain-of-Thought (CoT) prompting has proven highly effective for enhancing complex reasoning in Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs). Yet, it struggles in complex spatial reasoning tasks. Nonetheless, human cognition extends beyond language alone, enabling the remarkable capability to think in both words and images. Inspired by this mechanism, we propose a new reasoning paradigm, Multimodal Visualization-of-Thought (MVoT). It enables visual thinking in MLLMs by generating image visualizations of their reasoning traces. To ensure high-quality visualization, we introduce token discrepancy loss into autoregressive MLLMs. This innovation significantly improves both visual coherence and fidelity. We validate this approach through several dynamic spatial reasoning tasks. Experimental results reveal that MVoT demonstrates competitive performance across tasks. Moreover, it exhibits robust and reliable improvements in the most challenging scenarios where CoT fails. Ultimately, MVoT establishes new possibilities for complex reasoning tasks where visual thinking can effectively complement verbal reasoning. Chengzu Li, Wenshan Wu, Huanyu Zhang 0002, Yan Xia 0005, Shaoguang Mao, Li Dong 0004, Ivan Vulic, Furu Wei |
ICML | 7 |
| 2025 | Aligning with Logic: Measuring, Evaluating and Improving Logical Preference Consistency in Large Language ModelsabstractLarge Language Models (LLMs) are expected to be predictable and trustworthy to support reliable decision-making systems. Yet current LLMs often show inconsistencies in their judgments. In this work, we examine \textit{logical preference consistency} as a foundational requirement for building more dependable LLM systems, ensuring stable and coherent decision-making while minimizing erratic or contradictory outputs.
To quantify the logical preference consistency, we propose a universal evaluation framework based on three fundamental properties: *transitivity*, *commutativity* and *negation invariance*.
Through extensive experimentation across diverse LLMs, we demonstrate that these properties serve as strong indicators of judgment robustness.
Furthermore, we introduce a data refinement and augmentation technique, REPAIR, that enhances logical consistency while maintaining alignment with human preferences. Finally, we show that improving consistency leads to better performance in LLM-driven logic-based algorithms, reinforcing stability and coherence in decision-making systems. Yinhong Liu, Zhijiang Guo, Tianya Liang, Ehsan Shareghi, Ivan Vulic, Nigel Collier |
ICML | 5 |
| 2025 | UNDIAL: Self-Distillation with Adjusted Logits for Robust Unlearning in Large Language ModelsabstractYijiang River Dong, Hongzhou Lin, Mikhail Belkin, Ramon Huerta, Ivan Vulić. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yijiang River Dong, Hongzhou Lin, Mikhail Belkin, Ramón Huerta, Ivan Vulic |
NAACL (Long Papers) | 5 |
| 2025 | Universal Cross-Tokenizer Distillation via Approximate Likelihood MatchingabstractDistillation has shown remarkable success in transferring knowledge from a Large Language Model (LLM) teacher to a student LLM. However, current distillation methods require similar tokenizers between the teacher and the student, restricting their applicability to only a small subset of teacher--student pairs. In this work, we develop a principled cross-tokenizer distillation method to solve this crucial deficiency. Our method is the first to enable effective distillation across fundamentally different tokenizers, while also substantially outperforming prior methods in all other cases. We verify the efficacy of our method on three distinct use cases. First, we show that viewing tokenizer transfer as self-distillation enables unprecedentedly effective transfer across tokenizers, including rapid transfer of subword models to the byte-level. Transferring different models to the same tokenizer also enables ensembling to boost performance. Secondly, we distil a large maths-specialised LLM into a small general-purpose model with a different tokenizer, achieving competitive maths problem-solving performance. Thirdly, we use our method to train state-of-the-art embedding prediction hypernetworks for training-free tokenizer transfer. Our results unlock an expanded range of teacher--student pairs for distillation, enabling new ways to adapt and enhance interaction between LLMs. Benjamin Minixhofer, Ivan Vulic, Edoardo Maria Ponti |
NeurIPS | 2 |
| 2025 | Knowledge-aware audio-grounded generative slot filling for limited annotated dataabstractManually annotating fine-grained slot-value labels for task-oriented dialogue (ToD) systems is an expensive and time-consuming endeavour. This motivates research into slot-filling methods that operate with limited amounts of labelled data. Moreover, the majority of current work on ToD is based solely on text as the input modality, neglecting the additional challenges of imperfect automatic speech recognition (ASR) when working with spoken language. In this work, we propose a Knowledge-Aware Audio-Grounded generative slot filling framework, termed KA2G, that focuses on few-shot and zero-shot slot filling for ToD with speech input. KA2G achieves robust and data-efficient slot filling for speech-based ToD by (1) framing it as a text generation task, (2) grounding text generation additionally in the audio modality, and (3) conditioning on available external knowledge ( e.g. a predefined list of possible slot values). We show that combining both modalities within the KA2G framework improves the robustness against ASR errors. Further, the knowledge-aware slot-value generator in KA2G, implemented via a pointer generator mechanism, particularly benefits few-shot and zero-shot learning. Experiments, conducted on the standard speech-based single-turn SLURP dataset and a multi-turn dataset extracted from a commercial ToD system , display strong and consistent gains over prior work, especially in few-shot and zero-shot setups. Guangzhi Sun, Chao Zhang 0031, Ivan Vulic, Pawel Budzianowski, Philip C. Woodland |
Comput. Speech Lang. | 3 |
| 2025 | Analyzing and Adapting Large Language Models for Few-Shot Multilingual NLU: Are We There Yet?abstractAbstract Supervised fine-tuning (SFT), supervised instruction tuning (SIT), and in-context learning (ICL) are three alternative, de facto standard approaches to few-shot learning. ICL has gained popularity recently with the advent of LLMs due to its versatile simplicity and sample efficiency. Prior research has conducted only limited investigation into how these approaches work for multilingual few-shot learning, and the focus so far has been mostly on their performance. In this work, we present an extensive and systematic comparison of the three approaches, testing them on a variety of high- and low-resource languages over five different NLU tasks, and a myriad of language and domain setups. Importantly, performance is only one aspect of the comparison, where we also analyze and discuss the approaches through the optics of their computational, inference and financial costs. Some of the highlighted findings concern an excellent trade-off between performance and resource requirements/cost for SIT. We further analyze the impact of target language adaptation of pretrained LLMs and find that the standard adaptation approaches can (superficially) improve target language generation capabilities, but language understanding elicited through ICL does not improve accordingly and remains limited, especially for low-resource languages. Evgeniia Razumovskaia, Ivan Vulic, Anna Korhonen |
Trans. Assoc. Comput. Linguistics | 2 |
| 2025 | DARE: Diverse Visual Question Answering with Robustness EvaluationabstractAbstract Vision Language Models (VLMs) extend remarkable capabilities of text-only large language models and vision-only models, being able to learn from and process multi-modal vision-text input. While modern VLMs perform well on a number of standard image classification and image-text matching tasks, they still struggle with a number of crucial vision-language (VL) reasoning abilities such as counting and spatial reasoning. Moreover, while they might be very brittle to small variations in instructions and/or evaluation protocols, existing benchmarks fail to evaluate their robustness (or rather the lack of it). In order to couple challenging VL scenarios with comprehensive robustness evaluation, we introduce DARE, Diverse Visual Question Answering with Robustness Evaluation, a carefully created and curated multiple-choice VQA benchmark. DARE evaluates VLM performance on five diverse categories and includes four robustness-oriented evaluations based on the variations of prompts, the subsets of answer options, the output format, and the number of correct answers. Among a spectrum of other findings, we report that state-of-the-art VLMs still struggle with questions in most categories and are unable to consistently deliver their peak performance across the tested robustness evaluations. Consequently, our work calls for the systematic addition of robustness evaluations in future VLM research. Hannah Sterz, Jonas Pfeiffer, Ivan Vulic |
Trans. Assoc. Comput. Linguistics | 3 |
| 2024 | Reranking Overgenerated Responses for End-to-End Task-Oriented Dialogue SystemsabstractEnd-to-end task-oriented dialogue systems are prone to fall into the so-called ‘likelihood trap’, resulting in generated responses which are dull, repetitive, and often inconsistent with dialogue history. Comparing ranked lists of multiple generated responses against the ‘gold response’ reveals a wide diversity in quality, with many good responses placed lower in the ranked list. The main challenge addressed in this work is how to reach beyond greedily generated system responses, that is, how to obtain and select high-quality responses from the list of overgenerated responses at inference without the availability of the gold response. To this end, we propose a simple yet effective reranking method to select high-quality items from the lists of initially overgenerated responses. The idea is to use any sequence-level scoring function to divide the semantic space of responses into high-scoring versus low-scoring partitions. At training, the high-scoring partition comprises all generated responses whose similarity to the gold response is higher than the similarity of the greedy response to the gold response. At inference, the aim is to estimate the probability that each overgenerated response belongs to the high-scoring partition. We evaluate our proposed method on the standard MultiWOZ dataset, the BiTOD dataset, and with human evaluation. Songbo Hu, Ivan Vulic, Fangyu Liu 0001, Anna Korhonen |
LREC/COLING | 2 |
| 2024 | Segment Any Text: A Universal Approach for Robust, Efficient and Adaptable Sentence SegmentationabstractSegmenting text into sentences plays an early and crucial role in many NLP systems.This is commonly achieved by using rule-based or statistical methods relying on lexical features such as punctuation.Although some recent works no longer exclusively rely on punctuation, we find that no prior method achieves all of (i) robustness to missing punctuation, (ii) effective adaptability to new domains, and (iii) high efficiency.We introduce a new model -Segment any Text (SAT) -to solve this problem.To enhance robustness, we propose a new pretraining scheme that ensures less reliance on punctuation.To address adaptability, we introduce an extra stage of parameter-efficient fine-tuning, establishing state-of-the-art performance in distinct domains such as verses from lyrics and legal documents.Along the way, we introduce architectural modifications that result in a threefold gain in speed over the previous state of the art and solve spurious reliance on context far in the future.Finally, we introduce a variant of our model with fine-tuning on a diverse, multilingual mixture of sentence-segmented data, acting as a drop-in replacement and enhancement for existing segmentation tools.Overall, our contributions provide a universal approach for segmenting any text.Our method outperforms all baselines -including strong LLMs -across 8 corpora spanning diverse domains and languages, especially in practically relevant situations where text is poorly formatted.1 Markus Frohmann, Igor Sterner, Ivan Vulic, Benjamin Minixhofer, Markus Schedl |
EMNLP | 3 |
| 2024 | TopViewRS: Vision-Language Models as Top-View Spatial ReasonersabstractTop-view perspective denotes a typical way in which humans read and reason over different types of maps, and it is vital for localization and navigation of humans as well as of 'non-human' agents, such as the ones backed by large Vision-Language Models (VLMs).Nonetheless, spatial reasoning capabilities of modern VLMs in this setup remain unattested and underexplored.In this work, we study their capability to understand and reason over spatial relations from the top view.The focus on top view also enables controlled evaluations at different granularity of spatial reasoning; we clearly disentangle different abilities (e.g., recognizing particular objects versus understanding their relative positions).We introduce the TOPVIEWRS (Top-View Reasoning in Space) dataset, consisting of 11,384 multiple-choice questions with either realistic or semantic top-view map as visual input.We then use it to study and evaluate VLMs across 4 perception and reasoning tasks with different levels of complexity.Evaluation of 10 representative open-and closedsource VLMs reveals the gap of more than 50% compared to average human performance, and it is even lower than the random baseline in some cases.Although additional experiments show that Chain-of-Thought reasoning can boost model capabilities by 5.82% on average, the overall performance of VLMs remains limited.Our findings underscore the critical need for enhanced model capability in top-view spatial reasoning and set a foundation for further research towards human-level proficiency of VLMs in real-world multimodal tasks. Chengzu Li, Caiqi Zhang, Han Zhou 0010, Nigel Collier, Anna Korhonen, Ivan Vulic |
EMNLP | 6 |
| 2024 | Fairer Preferences Elicit Improved Human-Aligned Large Language Model JudgmentsabstractLarge language models (LLMs) have shown promising abilities as cost-effective and reference-free evaluators for assessing language generation quality.In particular, pairwise LLM evaluators, which compare two generated texts and determine the preferred one, have been employed in a wide range of applications.However, LLMs exhibit preference biases and worrying sensitivity to prompt designs.In this work, we first reveal that the predictive preference of LLMs can be highly brittle and skewed, even with semantically equivalent instructions.We find that fairer predictive preferences from LLMs consistently lead to judgments that are better aligned with humans.Motivated by this phenomenon, we propose an automatic Zero-shot Evaluation-oriented Prompt Optimization framework, ZEPO, which aims to produce fairer preference decisions and improve the alignment of LLM evaluators with human judgments.To this end, we propose a zeroshot learning objective based on the preference decision fairness.ZEPO demonstrates substantial performance improvements over stateof-the-art LLM evaluators, without requiring labeled data, on representative meta-evaluation benchmarks.Our findings underscore the critical correlation between preference fairness and human alignment, positioning ZEPO as an efficient prompt optimizer for bridging the gap between LLM evaluators and human judgments.* Now at Google.Code is available at https://github. com/cambridgeltl/zepo.Generate new prompts Zero-shot Fairness ZEPO Biased Preference Fairer Preference Optimized Prompt Initial Prompt LLM Optimizer LLM Evaluator Which summary candidate has better coherence?If the candidate A is better, please return 'A'.If the candidate B is better, please return 'B'.Which one exhibits better coherence?Return 'A' for the rst summary or 'B' for the second.Only provide the letter of your choice. Han Zhou 0010, Xingchen Wan, Yinhong Liu, Nigel Collier, Ivan Vulic, Anna Korhonen |
EMNLP | 5 |
| 2024 | FUN with Fisher: Improving Generalization of Adapter-Based Cross-lingual Transfer with Scheduled UnfreezingabstractChen Cecilia Liu, Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jonas Pfeiffer, Ivan Vulic, Iryna Gurevych |
NAACL-HLT | 3 |
| 2024 | SQATIN: Supervised Instruction Tuning Meets Question Answering for Improved Dialogue NLUabstractEvgeniia Razumovskaia, Goran Glavaš, Anna Korhonen, Ivan Vulić. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Evgeniia Razumovskaia, Goran Glavas, Anna Korhonen, Ivan Vulic |
NAACL-HLT | 4 |
| 2024 | Zero-Shot Tokenizer TransferabstractLanguage models (LMs) are bound to their tokenizer, which maps raw text to a sequence of vocabulary items (tokens). This restricts their flexibility: for example, LMs trained primarily on English may still perform well in other natural and programming languages, but have vastly decreased efficiency due to their English-centric tokenizer. To mitigate this, we should be able to swap the original LM tokenizer with an arbitrary one, on the fly, without degrading performance. Hence, in this work we define a new problem: Zero-Shot Tokenizer Transfer (ZeTT). The challenge at the core of ZeTT is finding embeddings for the tokens in the vocabulary of the new tokenizer. Since prior heuristics for initializing embeddings often perform at chance level in a ZeTT setting, we propose a new solution: we train a hypernetwork taking a tokenizer as input and predicting the corresponding embeddings. We empirically demonstrate that the hypernetwork generalizes to new tokenizers both with encoder (e.g., XLM-R) and decoder LLMs (e.g., Mistral-7B). Our method comes close to the original models' performance in cross-lingual and coding tasks while markedly reducing the length of the tokenized sequence. We also find that the remaining gap can be quickly closed by continued training on less than 1B tokens. Finally, we show that a ZeTT hypernetwork trained for a base (L)LM can also be applied to fine-tuned variants without extra training. Overall, our results make substantial strides toward detaching LMs from their tokenizer. Benjamin Minixhofer, Edoardo Maria Ponti, Ivan Vulic |
NeurIPS | 3 |
| 2024 | CALRec: Contrastive Alignment of Generative LLMs for Sequential RecommendationabstractTraditional recommender systems such as matrix factorization methods have primarily focused on learning a shared dense embedding space to represent both items and user preferences. Subsequently, sequence models such as RNN, GRUs, and, recently, Transformers have emerged and excelled in the task of sequential recommendation. This task requires understanding the sequential structure present in users’ historical interactions to predict the next item they may like. Building upon the success of Large Language Models (LLMs) in a variety of tasks, researchers have recently explored using LLMs that are pretrained on vast corpora of text for sequential recommendation. To use LLMs for sequential recommendation, both the history of user interactions and the model’s prediction of the next item are expressed in text form. We propose CALRec, a two-stage LLM finetuning framework that finetunes a pretrained LLM in a two-tower fashion using a mixture of two contrastive losses and a language modeling loss: the LLM is first finetuned on a data mixture from multiple domains followed by another round of target domain finetuning. Our model significantly outperforms many state-of-the-art baselines (+37% in Recall@1 and +24% in NDCG@10) and our systematic ablation studies reveal that (i) both stages of finetuning are crucial, and, when combined, we achieve improved performance, and (ii) contrastive alignment is effective among the target domains explored in our experiments. Yaoyiran Li, Xiang Zhai, Moustafa Farid Alzantot, Keyi Yu, Ivan Vulic, Anna Korhonen, Mohamed Hammad |
RecSys | 5 |
| 2024 | AutoPEFT: Automatic Configuration Search for Parameter-Efficient Fine-TuningabstractAbstract Large pretrained language models are widely used in downstream NLP tasks via task- specific fine-tuning, but such procedures can be costly. Recently, Parameter-Efficient Fine-Tuning (PEFT) methods have achieved strong task performance while updating much fewer parameters than full model fine-tuning (FFT). However, it is non-trivial to make informed design choices on the PEFT configurations, such as their architecture, the number of tunable parameters, and even the layers in which the PEFT modules are inserted. Consequently, it is highly likely that the current, manually designed configurations are suboptimal in terms of their performance-efficiency trade-off. Inspired by advances in neural architecture search, we propose AutoPEFT for automatic PEFT configuration selection: We first design an expressive configuration search space with multiple representative PEFT modules as building blocks. Using multi-objective Bayesian optimization in a low-cost setup, we then discover a Pareto-optimal set of configurations with strong performance-cost trade-offs across different numbers of parameters that are also highly transferable across different tasks. Empirically, on GLUE and SuperGLUE tasks, we show that AutoPEFT-discovered configurations significantly outperform existing PEFT methods and are on par or better than FFT without incurring substantial training efficiency costs. Han Zhou 0010, Xingchen Wan, Ivan Vulic, Anna Korhonen |
Trans. Assoc. Comput. Linguistics | 3 |
| 2023 | Translation-Enhanced Multilingual Text-to-Image GenerationabstractResearch on text-to-image generation (TTI) still predominantly focuses on the English language due to the lack of annotated imagecaption data in other languages; in the long run, this might widen inequitable access to TTI technology.In this work, we thus investigate multilingual TTI (termed mTTI) and the current potential of neural machine translation (NMT) to bootstrap mTTI systems.We provide two key contributions.1) Relying on a multilingual multi-modal encoder, we provide a systematic empirical study of standard methods used in cross-lingual NLP when applied to mTTI: TRANSLATE TRAIN, TRANS-LATE TEST, and ZERO-SHOT TRANSFER.2) We propose Ensemble Adapter (ENSAD), a novel parameter-efficient approach that learns to weigh and consolidate the multilingual text knowledge within the mTTI framework, mitigating the language gap and thus improving mTTI performance.Our evaluations on standard mTTI datasets COCO-CN, Multi30K Task2, and LAION-5B demonstrate the potential of translation-enhanced mTTI systems and also validate the benefits of the proposed EN-SAD which derives consistent gains across all datasets.Further investigations on model variants, ablation studies, and qualitative analyses provide additional insights on the inner workings of the proposed mTTI approaches. Yaoyiran Li, Ching-Yun Chang, Stephen Rawls, Ivan Vulic, Anna Korhonen |
ACL (1) | 4 |
| 2023 | Where's the Point? Self-Supervised Multilingual Punctuation-Agnostic Sentence SegmentationabstractMany NLP pipelines split text into sentences as one of the crucial preprocessing steps.Prior sentence segmentation tools either rely on punctuation or require a considerable amount of sentence-segmented training data: both central assumptions might fail when porting sentence segmenters to diverse languages on a massive scale.In this work, we thus introduce a multilingual punctuation-agnostic sentence segmentation method, currently covering 85 languages, trained in a self-supervised fashion on unsegmented text, by making use of newline characters which implicitly perform segmentation into paragraphs.We further propose an approach that adapts our method to the segmentation in a given corpus by using only a small number (64-256) of sentence-segmented examples.The main results indicate that our method outperforms all the prior best sentence-segmentation tools by an average of 6.1% F1 points.Furthermore, we demonstrate that proper sentence segmentation has a point: the use of a (powerful) sentence segmenter makes a considerable difference for a downstream application such as machine translation (MT).By using our method to match sentence segmentation to the segmentation used during training of MT models, we achieve an average improvement of 2.3 BLEU points over the best prior segmentation tool, as well as massive gains over a trivial segmenter that splits text into equally sized blocks. Benjamin Minixhofer, Jonas Pfeiffer, Ivan Vulic |
ACL (1) | 3 |
| 2023 | Free Lunch: Robust Cross-Lingual Transfer via Model Checkpoint AveragingabstractMassively multilingual language models have displayed strong performance in zero-shot (ZS-XLT) and few-shot (FS-XLT) cross-lingual transfer setups, where models fine-tuned on task data in a source language are transferred without any or with only a few annotated instances to the target language(s).However, current work typically overestimates model performance as fine-tuned models are frequently evaluated at model checkpoints that generalize best to validation instances in the target languages.This effectively violates the main assumptions of 'true' ZS-XLT and FS-XLT.Such XLT setups require robust methods that do not depend on labeled target language data for validation and model selection.In this work, aiming to improve the robustness of 'true' ZS-XLT and FS-XLT, we propose a simple and effective method that averages different checkpoints (i.e., model snapshots) during task fine-tuning.We conduct exhaustive ZS-XLT and FS-XLT experiments across higher-level semantic tasks (NLI, extractive QA) and lower-level token classification tasks (NER, POS).The results indicate that averaging model checkpoints yields systematic and consistent performance gains across diverse target languages in all tasks.Importantly, it simultaneously substantially desensitizes XLT to varying hyperparameter choices in the absence of target language validation.We also show that checkpoint averaging benefits performance when further combined with run averaging (i.e., averaging the parameters of models fine-tuned over independent runs). Fabian David Schmidt, Ivan Vulic, Goran Glavas |
ACL (1) | 2 |
| 2023 | Probing Cross-Lingual Lexical Knowledge from Multilingual Sentence EncodersabstractIvan Vulić, Goran Glavaš, Fangyu Liu, Nigel Collier, Edoardo Maria Ponti, Anna Korhonen. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Ivan Vulic, Goran Glavas, Fangyu Liu 0001, Nigel Collier, Edoardo Maria Ponti, Anna Korhonen |
EACL | 1 |
| 2023 | Can Pretrained Language Models (Yet) Reason Deductively?abstractAcquiring factual knowledge with Pretrained Language Models (PLMs) has attracted increasing attention, showing promising performance in many knowledge-intensive tasks.Their good performance has led the community to believe that the models do possess a modicum of reasoning competence rather than merely memorising the knowledge.In this paper, we conduct a comprehensive evaluation of the learnable deductive (also known as explicit) reasoning capability of PLMs.Through a series of controlled experiments, we posit two main findings.(i) PLMs inadequately generalise learned logic rules and perform inconsistently against simple adversarial surface form edits. (ii) While the deductive reasoning fine-tuning of PLMs does improve their performance on reasoning over unseen knowledge facts, it results in catastrophically forgetting the previously learnt knowledge.Our main results suggest that PLMs cannot yet perform reliable deductive reasoning, demonstrating the importance of controlled examinations and probing of PLMs' deductive reasoning abilities; we reach beyond (misleading) task performance, revealing that PLMs are still far from robust reasoning capabilities, even for simple deductive tasks. Zhangdie Yuan, Songbo Hu, Ivan Vulic, Anna Korhonen, Zaiqiao Meng |
EACL | 3 |
| 2023 | Unifying Cross-Lingual Transfer across Scenarios of Resource ScarcityabstractThe scarcity of data in many of the world's languages necessitates the transfer of knowledge from other, resource-rich languages.However, the level of scarcity varies significantly across multiple dimensions, including: i) the amount of task-specific data available in the source and target languages; ii) the amount of monolingual and parallel data available for both languages; and iii) the extent to which they are supported by pretrained multilingual and translation models.Prior work has largely treated these dimensions and the various techniques for dealing with them separately; in this paper, we offer a more integrated view by exploring how to deploy the arsenal of cross-lingual transfer tools across a range of scenarios, especially the most challenging, low-resource ones.To this end, we run experiments on the Americas-NLI and NusaX benchmarks over 20 languages, simulating a range of few-shot settings.The best configuration in our experiments employed parameter-efficient language and task adaptation of massively multilingual Transformers, trained simultaneously on source language data and both machine-translated and natural data for multiple target languages.In addition, we show that pre-trained translation models can be easily adapted to unseen languages, thus extending the range of our hybrid technique and translation-based transfer more broadly.Beyond new insights into the mechanisms of cross-lingual transfer, we hope our work will provide practitioners with a toolbox to integrate multiple techniques for different real-world scenarios.Our code is available at https: //github.com/parovicm/unified-xlt. Alan Ansell, Marinela Parovic, Ivan Vulic, Anna Korhonen, Edoardo Maria Ponti |
EMNLP | 3 |
| 2023 | A Systematic Study of Performance Disparities in Multilingual Task-Oriented Dialogue SystemsabstractAchieving robust language technologies that can perform well across the world's many languages is a central goal of multilingual NLP.In this work, we take stock of and empirically analyse task performance disparities that exist between multilingual task-oriented dialogue (TOD) systems.We first define new quantitative measures of absolute and relative equivalence in system performance, capturing disparities across languages and within individual languages.Through a series of controlled experiments, we demonstrate that performance disparities depend on a number of factors: the nature of the TOD task at hand, the underlying pretrained language model, the target language, and the amount of TOD annotated data.We empirically prove the existence of the adaptation bias and intrinsic biases in current TOD systems: e.g., TOD systems trained for Arabic or Turkish using annotated TOD data fully parallel to English TOD data still exhibit diminished TOD task performance.Beyond providing a series of insights into the performance disparities of TOD systems in different languages, our analyses offer practical tips on how to approach TOD data collection and system development for new languages. Songbo Hu, Han Zhou 0010, Moy Yuan, Milan Gritta, Guchun Zhang, Ignacio Iacobacci, Anna Korhonen, Ivan Vulic |
EMNLP | 8 |
| 2023 | On Bilingual Lexicon Induction with Large Language ModelsabstractBilingual Lexicon Induction (BLI) is a core task in multilingual NLP that still, to a large extent, relies on calculating cross-lingual word representations.Inspired by the global paradigm shift in NLP towards Large Language Models (LLMs), we examine the potential of the latest generation of LLMs for the development of bilingual lexicons.We ask the following research question: Is it possible to prompt and fine-tune multilingual LLMs (mLLMs) for BLI, and how does this approach compare against and complement current BLI approaches?To this end, we systematically study 1) zero-shot prompting for unsupervised BLI and 2) fewshot in-context prompting with a set of seed translation pairs, both without any LLM finetuning, as well as 3) standard BLI-oriented finetuning of smaller LLMs.We experiment with 18 open-source text-to-text mLLMs of different sizes (from 0.3B to 13B parameters) on two standard BLI benchmarks covering a range of typologically diverse languages.Our work is the first to demonstrate strong BLI capabilities of text-to-text mLLMs.The results reveal that few-shot prompting with in-context examples from nearest neighbours achieves the best performance, establishing new state-of-the-art BLI scores for many language pairs.We also conduct a series of in-depth analyses and ablation studies, providing more insights on BLI with (m)LLMs, also along with their limitations. Yaoyiran Li, Anna Korhonen, Ivan Vulic |
EMNLP | 3 |
| 2023 | CompoundPiece: Evaluating and Improving Decompounding Performance of Language ModelsabstractWhile many languages possess processes of joining two or more words to create compound words, previous studies have been typically limited only to languages with excessively productive compound formation (e.g., German, Dutch) and there is no public dataset containing compound and non-compound words across a large number of languages.In this work, we systematically study decompounding, the task of splitting compound words into their constituents, at a wide scale.We first address the data gap by introducing a dataset of 255k compound and noncompound words across 56 diverse languages obtained from Wiktionary.We then use this dataset to evaluate an array of Large Language Models (LLMs) on the decompounding task.We find that LLMs perform poorly, especially on words which are tokenized unfavorably by subword tokenization.We thus introduce a novel methodology to train dedicated models for decompounding.The proposed two-stage procedure relies on a fully self-supervised objective in the first stage, while the second, supervised learning stage optionally fine-tunes the model on the annotated Wiktionary data.Our self-supervised models outperform the prior best unsupervised decompounding models by 13.9% accuracy on average.Our fine-tuned models outperform all prior (language-specific) decompounding tools.Furthermore, we use our models to leverage decompounding during the creation of a subword tokenizer, which we refer to as CompoundPiece.CompoundPiece tokenizes compound words more favorably on average, leading to improved performance on decompounding over an otherwise equivalent model using SentencePiece tokenization. Benjamin Minixhofer, Jonas Pfeiffer, Ivan Vulic |
EMNLP | 3 |
| 2023 | Transfer-Free Data-Efficient Multilingual Slot LabelingabstractSlot labeling (SL) is a core component of taskoriented dialogue (TOD) systems, where slots and corresponding values are usually language-, task-and domain-specific.Therefore, extending the system to any new language-domaintask configuration requires (re)running an expensive and resource-intensive data annotation process.To mitigate the inherent data scarcity issue, current research on multilingual ToD assumes that sufficient English-language annotated data are always available for particular tasks and domains, and thus operates in a standard cross-lingual transfer setup.In this work, we depart from this often unrealistic assumption.We examine challenging scenarios where such transfer-enabling English annotated data cannot be guaranteed, and focus on bootstrapping multilingual data-efficient slot labelers in transfer-free scenarios directly in the target languages without any English-ready data.We propose a two-stage slot labeling approach (termed TWOSL) which transforms standard multilingual sentence encoders into effective slot labelers.In Stage 1, relying on SL-adapted contrastive learning with only a handful of SLannotated examples, we turn sentence encoders into task-specific span encoders.In Stage 2, we recast SL from a token classification into a simpler, less data-intensive span classification task.Our results on two standard multilingual TOD datasets and across diverse languages confirm the effectiveness and robustness of TWOSL.It is especially effective for the most challenging transfer-free few-shot setups, paving the way for quick and data-efficient bootstrapping of multilingual slot labelers for TOD. Evgeniia Razumovskaia, Ivan Vulic, Anna Korhonen |
EMNLP | 2 |
| 2023 | Multi 3 WOZ: A Multilingual, Multi-Domain, Multi-Parallel Dataset for Training and Evaluating Culturally Adapted Task-Oriented Dialog SystemsabstractAbstract Creating high-quality annotated data for task-oriented dialog (ToD) is known to be notoriously difficult, and the challenges are amplified when the goal is to create equitable, culturally adapted, and large-scale ToD datasets for multiple languages. Therefore, the current datasets are still very scarce and suffer from limitations such as translation-based non-native dialogs with translation artefacts, small scale, or lack of cultural adaptation, among others. In this work, we first take stock of the current landscape of multilingual ToD datasets, offering a systematic overview of their properties and limitations. Aiming to reduce all the detected limitations, we then introduce Multi3WOZ, a novel multilingual, multi-domain, multi-parallel ToD dataset. It is large-scale and offers culturally adapted dialogs in 4 languages to enable training and evaluation of multilingual and cross-lingual ToD systems. We describe a complex bottom–up data collection process that yielded the final dataset, and offer the first sets of baseline scores across different ToD-related tasks for future reference, also highlighting its challenging nature. Songbo Hu, Han Zhou 0010, Mete Hergul, Milan Gritta, Guchun Zhang, Ignacio Iacobacci, Ivan Vulic, Anna Korhonen |
Trans. Assoc. Comput. Linguistics | 7 |
| 2023 | Cross-Lingual Dialogue Dataset Creation via Outline-Based GenerationabstractAbstract Multilingual task-oriented dialogue (ToD) facilitates access to services and information for many (communities of) speakers. Nevertheless, its potential is not fully realized, as current multilingual ToD datasets—both for modular and end-to-end modeling—suffer from severe limitations. 1) When created from scratch, they are usually small in scale and fail to cover many possible dialogue flows. 2) Translation-based ToD datasets might lack naturalness and cultural specificity in the target language. In this work, to tackle these limitations we propose a novel outline-based annotation process for multilingual ToD datasets, where domain-specific abstract schemata of dialogue are mapped into natural language outlines. These in turn guide the target language annotators in writing dialogues by providing instructions about each turn’s intents and slots. Through this process we annotate a new large-scale dataset for evaluation of multilingual and cross-lingual ToD systems. Our Cross-lingual Outline-based Dialogue dataset (cod) enables natural language understanding, dialogue state tracking, and end-to-end dialogue evaluation in 4 diverse languages: Arabic, Indonesian, Russian, and Kiswahili. Qualitative and quantitative analyses of cod versus an equivalent translation-based dataset demonstrate improvements in data quality, unlocked by the outline-based approach. Finally, we benchmark a series of state-of-the-art systems for cross-lingual ToD, setting reference scores for future work and demonstrating that cod prevents over-inflated performance, typically met with prior translation-based ToD datasets. Olga Majewska, Evgeniia Razumovskaia, Edoardo Maria Ponti, Ivan Vulic, Anna Korhonen |
Trans. Assoc. Comput. Linguistics | 4 |
| 2022 | Composable Sparse Fine-Tuning for Cross-Lingual TransferabstractFine-tuning the entire set of parameters of a large pretrained model has become the mainstream approach for transfer learning.To increase its efficiency and prevent catastrophic forgetting and interference, techniques like adapters and sparse fine-tuning have been developed.Adapters are modular, as they can be combined to adapt a model towards different facets of knowledge (e.g., dedicated language and/or task adapters).Sparse finetuning is expressive, as it controls the behavior of all model components.In this work, we introduce a new fine-tuning method with both these desirable properties.In particular, we learn sparse, real-valued masks based on a simple variant of the Lottery Ticket Hypothesis.Task-specific masks are obtained from annotated data in a source language, and languagespecific masks from masked language modeling in a target language.Both these masks can then be composed with the pretrained model.Unlike adapter-based fine-tuning, this method neither increases the number of parameters at inference time nor alters the original model architecture.Most importantly, it outperforms adapters in zero-shot cross-lingual transfer by a large margin in a series of multilingual benchmarks, including Universal Dependencies, MasakhaNER, and AmericasNLI.Based on an in-depth analysis, we additionally find that sparsity is crucial to prevent both 1) interference between the fine-tunings to be composed and 2) overfitting.We release the code and models at https://github.com/ cambridgeltl/composable-sft. Alan Ansell, Edoardo Maria Ponti, Anna Korhonen, Ivan Vulic |
ACL (1) | 4 |
| 2022 | Improving Word Translation via Two-Stage Contrastive LearningabstractWord translation or bilingual lexicon induction (BLI) is a key cross-lingual task, aiming to bridge the lexical gap between different languages.In this work, we propose a robust and effective two-stage contrastive learning framework for the BLI task.At Stage C1, we propose to refine standard cross-lingual linear maps between static word embeddings (WEs) via a contrastive learning objective; we also show how to integrate it into the self-learning procedure for even more refined cross-lingual maps.In Stage C2, we conduct BLI-oriented contrastive fine-tuning of mBERT, unlocking its word translation capability.We also show that static WEs induced from the 'C2-tuned' mBERT complement static WEs from Stage C1.Comprehensive experiments on standard BLI datasets for diverse languages and different experimental setups demonstrate substantial gains achieved by our framework.While the BLI method from Stage C1 already yields substantial gains over all state-of-the-art BLI methods in our comparison, even stronger improvements are met with the full two-stage framework: e.g., we report gains for 112/112 BLI setups, spanning 28 language pairs. Yaoyiran Li, Fangyu Liu 0001, Nigel Collier, Anna Korhonen, Ivan Vulic |
ACL (1) | 5 |
| 2022 | Prix-LM: Pretraining for Multilingual Knowledge Base ConstructionabstractKnowledge bases (KBs) contain plenty of structured world and commonsense knowledge.As such, they often complement distributional text-based information and facilitate various downstream tasks.Since their manual construction is resource-and timeintensive, recent efforts have tried leveraging large pretrained language models (PLMs) to generate additional monolingual knowledge facts for KBs.However, such methods have not been attempted for building and enriching multilingual KBs.Besides wider application, such multilingual KBs can provide richer combined knowledge than monolingual (e.g., English) KBs.Knowledge expressed in different languages may be complementary and unequally distributed: this implies that the knowledge available in high-resource languages can be transferred to low-resource ones.To achieve this, it is crucial to represent multilingual knowledge in a shared/unified space.To this end, we propose a unified representation model, Prix-LM , for multilingual KB construction and completion.We leverage two types of knowledge, monolingual triples and cross-lingual links, extracted from existing multilingual KBs, and tune a multilingual language encoder XLM-R via a causal language modeling objective.Prix-LM integrates useful multilingual and KB-based factual knowledge into a single model.Experiments on standard entity-related tasks, such as link prediction in multiple languages, cross-lingual entity linking and bilingual lexicon induction, demonstrate its effectiveness, with gains reported over strong task-specialised baselines. Wenxuan Zhou 0002, Fangyu Liu 0001, Ivan Vulic, Nigel Collier, Muhao Chen 0001 |
ACL (1) | 3 |
| 2022 | Parameter-Efficient Neural Reranking for Cross-Lingual and Multilingual RetrievalabstractState-of-the-art neural (re)rankers are notoriously data-hungry which – given the lack of large-scale training data in languages other than English – makes them rarely used in multilingual and cross-lingual retrieval settings. Current approaches therefore commonly transfer rankers trained on English data to other languages and cross-lingual setups by means of multilingual encoders: they fine-tune all parameters of pretrained massively multilingual Transformers (MMTs, e.g., multilingual BERT) on English relevance judgments, and then deploy them in the target language(s). In this work, we show that two parameter-efficient approaches to cross-lingual transfer, namely Sparse Fine-Tuning Masks (SFTMs) and Adapters, allow for a more lightweight and more effective zero-shot transfer to multilingual and cross-lingual retrieval tasks. We first train language adapters (or SFTMs) via Masked Language Modelling and then train retrieval (i.e., reranking) adapters (SFTMs) on top, while keeping all other parameters fixed. At inference, this modular design allows us to compose the ranker by applying the (re)ranking adapter (or SFTM) trained with source language data together with the language adapter (or SFTM) of a target language. We carry out a large scale evaluation on the CLEF-2003 and HC4 benchmarks and additionally, as another contribution, extend the former with queries in three new languages: Kyrgyz, Uyghur and Turkish. The proposed parameter-efficient methods outperform standard zero-shot transfer with full MMT fine-tuning, while being more modular and reducing training times. The gains are particularly pronounced for low-resource languages, where our approaches also substantially outperform the competitive machine translation-based rankers. Robert Litschko, Ivan Vulic, Goran Glavas |
COLING | 2 |
| 2022 | Don't Stop Fine-Tuning: On Training Regimes for Few-Shot Cross-Lingual Transfer with Multilingual Language ModelsabstractA large body of recent work highlights the fallacies of zero-shot cross-lingual transfer (ZS-XLT) with large multilingual language models.Namely, their performance varies substantially for different target languages and is the weakest where needed the most: for low-resource languages distant to the source language.One remedy is few-shot transfer (FS-XLT), where leveraging only a few task-annotated instances in the target language(s) may yield sizable performance gains.However, FS-XLT also succumbs to large variation, as models easily overfit to the small datasets.In this work, we present a systematic study focused on a spectrum of FS-XLT fine-tuning regimes, analyzing key properties such as effectiveness, (in)stability, and modularity.We conduct extensive experiments on both higher-level (NLI, paraphrasing) and lowerlevel tasks (NER, POS), presenting new FS-XLT strategies that yield both improved and more stable FS-XLT across the board.Our findings challenge established FS-XLT methods: e.g., we propose to replace sequential fine-tuning with joint fine-tuning on source and target language instances, offering consistent gains with different number of shots (including resourcerich scenarios).We also show that further gains can be achieved with multi-stage FS-XLT training in which joint multilingual fine-tuning precedes the bilingual source-target specialization. Fabian David Schmidt, Ivan Vulic, Goran Glavas |
EMNLP | 2 |
| 2022 | SLICER: Sliced Fine-Tuning for Low-Resource Cross-Lingual Transfer for Named Entity RecognitionabstractLarge multilingual language models generally demonstrate impressive results in zero-shot cross-lingual transfer, yet often fail to successfully transfer to low-resource languages, even for token-level prediction tasks like named entity recognition (NER).In this work, we introduce a simple yet highly effective approach for improving zero-shot transfer for NER to lowresource languages.We observe that NER finetuning in the source language decontextualizes token representations, i.e., tokens increasingly attend to themselves.This increased reliance on token information itself, we hypothesize, triggers a type of overfitting to properties that NE tokens within the source languages share, but are generally not present in NE mentions of target languages.As a remedy, we propose a simple yet very effective sliced fine-tuning for NER (SLICER) that forces stronger token contextualization in the Transformer: we divide the transformed token representations and classifier into disjoint slices that are then independently classified during training.We evaluate SLICER on two standard benchmarks for NER that involve low-resource languages, WikiANN and MasakhaNER, and show that it (i) indeed reduces decontextualization (i.e., extent to which NE tokens attend to themselves), consequently (ii) yielding consistent transfer gains, especially prominent for low-resource target languages distant from the source language. Fabian David Schmidt, Ivan Vulic, Goran Glavas |
EMNLP | 2 |
| 2022 | Multi-Label Intent Detection via Contrastive Task Specialization of Sentence EncodersabstractDeploying task-oriented dialog (TOD) systems for new domains and tasks requires natural language understanding models that are 1) resource-efficient and work under low-data regimes; 2) adaptable, efficient, and quickto-train; 3) expressive and can handle complex TOD scenarios with multiple user intents in a single utterance.Motivated by these requirements, we introduce a novel framework for multi-label intent detection (mID): MULTI-CONVFIT (Multi-Label Intent Detection via Contrastive Conversational Fine-Tuning).While previous work on efficient single-label intent detection learns a classifier on top of a fixed sentence encoder (SE), we propose to 1) transform general-purpose SEs into task-specialized SEs via contrastive fine-tuning on annotated multi-label data, 2) where task specialization knowledge can be stored into lightweight adapter modules without updating the original parameters of the input SE, and then 3) we build improved mID classifiers stacked on top of fixed specialized SEs.Our main results indicate that MULTI-CONVFIT yields effective mID models, with large gains over non-specialized SEs reported across a spectrum of different mID datasets, both in low-data and high-data regimes. Ivan Vulic, Iñigo Casanueva, Georgios Spithourakis, Avishek Mondal, Tsung-Hsien Wen, Pawel Budzianowski |
EMNLP | 1 |
| 2022 | IGLUE: A Benchmark for Transfer Learning across Modalities, Tasks, and LanguagesabstractReliable evaluation benchmarks designed for replicability and comprehensiveness have driven progress in machine learning. Due to the lack of a multilingual benchmark, however, vision-and-language research has mostly focused on English language tasks. To fill this gap, we introduce the Image-Grounded Language Understanding Evaluation benchmark. IGLUE brings together{—}by both aggregating pre-existing datasets and creating new ones{—}visual question answering, cross-modal retrieval, grounded reasoning, and grounded entailment tasks across 20 diverse languages. Our benchmark enables the evaluation of multilingual multimodal models for transfer learning, not only in a zero-shot setting, but also in newly defined few-shot learning setups. Based on the evaluation of the available state-of-the-art models, we find that translate-test transfer is superior to zero-shot transfer and that few-shot learning is hard to harness for many tasks. Moreover, downstream performance is partially explained by the amount of available unlabelled textual data for pretraining, and only weakly by the typological distance of target{–}source languages. We hope to encourage future research efforts in this area by releasing the benchmark to the community. Emanuele Bugliarello, Fangyu Liu 0001, Jonas Pfeiffer, Siva Reddy, Desmond Elliott, Edoardo Maria Ponti, Ivan Vulic |
ICML | 7 |
| 2022 | Multi2WOZ: A Robust Multilingual Dataset and Conversational Pretraining for Task-Oriented DialogabstractChia-Chien Hung, Anne Lauscher, Ivan Vulić, Simone Ponzetto, Goran Glavaš. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Chia-Chien Hung, Anne Lauscher, Ivan Vulic, Simone Paolo Ponzetto, Goran Glavas |
NAACL-HLT | 3 |
| 2022 | BAD-X: Bilingual Adapters Improve Zero-Shot Cross-Lingual TransferabstractMarinela Parović, Goran Glavaš, Ivan Vulić, Anna Korhonen. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Marinela Parovic, Goran Glavas, Ivan Vulic, Anna Korhonen |
NAACL-HLT | 3 |
| 2022 | On cross-lingual retrieval with multilingual text encodersabstractAbstract Pretrained multilingual text encoders based on neural transformer architectures , such as multilingual BERT (mBERT) and XLM, have recently become a default paradigm for cross-lingual transfer of natural language processing models, rendering cross-lingual word embedding spaces (CLWEs) effectively obsolete. In this work we present a systematic empirical study focused on the suitability of the state-of-the-art multilingual encoders for cross-lingual document and sentence retrieval tasks across a number of diverse language pairs. We first treat these models as multilingual text encoders and benchmark their performance in unsupervised ad-hoc sentence- and document-level CLIR. In contrast to supervised language understanding, our results indicate that for unsupervised document-level CLIR—a setup with no relevance judgments for IR-specific fine-tuning—pretrained multilingual encoders on average fail to significantly outperform earlier models based on CLWEs. For sentence-level retrieval, we do obtain state-of-the-art performance: the peak scores, however, are met by multilingual encoders that have been further specialized, in a supervised fashion, for sentence understanding tasks, rather than using their vanilla ‘off-the-shelf’ variants. Following these results, we introduce localized relevance matching for document-level CLIR, where we independently score a query against document sections. In the second part, we evaluate multilingual encoders fine-tuned in a supervised fashion (i.e., we learn to rank ) on English relevance data in a series of zero-shot language and domain transfer CLIR experiments. Our results show that, despite the supervision, and due to the domain and language shift, supervised re-ranking rarely improves the performance of multilingual transformers as unsupervised base rankers. Finally, only with in-domain contrastive fine-tuning (i.e., same domain, only language transfer), we manage to improve the ranking quality. We uncover substantial empirical differences between cross-lingual retrieval results and results of (zero-shot) cross-lingual transfer for monolingual retrieval in target languages, which point to “monolingual overfitting” of retrieval models trained on monolingual (English) data, even if they are based on multilingual transformers. Robert Litschko, Ivan Vulic, Simone Paolo Ponzetto, Goran Glavas |
Inf. Retr. J. | 2 |
| 2022 | Crossing the Conversational Chasm: A Primer on Natural Language Processing for Multilingual Task-Oriented Dialogue SystemsabstractIn task-oriented dialogue (ToD), a user holds a conversation with an artificial agent with the aim of completing a concrete task. Although this technology represents one of the central objectives of AI and has been the focus of ever more intense research and development efforts, it is currently limited to a few narrow domains (e.g., food ordering, ticket booking) and a handful of languages (e.g., English, Chinese). This work provides an extensive overview of existing methods and resources in multilingual ToD as an entry point to this exciting and emerging field. We find that the most critical factor preventing the creation of truly multilingual ToD systems is the lack of datasets in most languages for both training and evaluation. In fact, acquiring annotations or human feedback for each component of modular systems or for data-hungry end-to-end systems is expensive and tedious. Hence, state-of-the-art approaches to multilingual ToD mostly rely on (zero- or few-shot) cross-lingual transfer from resource-rich languages (almost exclusively English), either by means of (i) machine translation or (ii) multilingual representations. These approaches are currently viable only for typologically similar languages and languages with parallel / monolingual corpora available. On the other hand, their effectiveness beyond these boundaries is doubtful or hard to assess due to the lack of linguistically diverse benchmarks (especially for natural language generation and end-to-end evaluation). To overcome this limitation, we draw parallels between components of the ToD pipeline and other NLP tasks, which can inspire solutions for learning in low-resource scenarios. Finally, we list additional challenges that multilinguality poses for related areas (such as speech, fluency in generated text, and human-centred evaluation), and indicate future directions that hold promise to further expand language coverage and dialogue capabilities of current ToD systems. Evgeniia Razumovskaia, Goran Glavas, Olga Majewska, Edoardo Maria Ponti, Anna Korhonen, Ivan Vulic |
J. Artif. Intell. Res. | 6 |
| 2022 | Retrieve Fast, Rerank Smart: Cooperative and Joint Approaches for Improved Cross-Modal RetrievalabstractAbstract Current state-of-the-art approaches to cross- modal retrieval process text and visual input jointly, relying on Transformer-based architectures with cross-attention mechanisms that attend over all words and objects in an image. While offering unmatched retrieval performance, such models: 1) are typically pretrained from scratch and thus less scalable, 2) suffer from huge retrieval latency and inefficiency issues, which makes them impractical in realistic applications. To address these crucial gaps towards both improved and efficient cross- modal retrieval, we propose a novel fine-tuning framework that turns any pretrained text-image multi-modal model into an efficient retrieval model. The framework is based on a cooperative retrieve-and-rerank approach that combines: 1) twin networks (i.e., a bi-encoder) to separately encode all items of a corpus, enabling efficient initial retrieval, and 2) a cross-encoder component for a more nuanced (i.e., smarter) ranking of the retrieved small set of items. We also propose to jointly fine- tune the two components with shared weights, yielding a more parameter-efficient model. Our experiments on a series of standard cross-modal retrieval benchmarks in monolingual, multilingual, and zero-shot setups, demonstrate improved accuracy and huge efficiency benefits over the state-of-the-art cross- encoders.1 Gregor Geigle, Jonas Pfeiffer, Nils Reimers 0001, Ivan Vulic, Iryna Gurevych |
Trans. Assoc. Comput. Linguistics | 4 |
| 2021 | Analogy Training Multilingual EncodersabstractLanguage encoders encode words and phrases in ways that capture their local semantic relatedness, but are known to be globally inconsistent. Global inconsistency can seemingly be corrected for, in part, by leveraging signals from knowledge bases, but previous results are partial and limited to monolingual English encoders. We extract a large-scale multilingual, multi-word analogy dataset from Wikidata for diagnosing and correcting for global inconsistencies, and then implement a four-way Siamese BERT architecture for grounding multilingual BERT (mBERT) in Wikidata through analogy training. We show that analogy training not only improves the global consistency of mBERT, as well as the isomorphism of language-specific subspaces, but also leads to consistent gains on downstream tasks such as bilingual dictionary induction and sentence retrieval. Nicolas Garneau, Mareike Hartmann, Anders Sandholm 0001, Sebastian Ruder, Ivan Vulic, Anders Søgaard |
AAAI | 5 |
| 2021 | RedditBias: A Real-World Resource for Bias Evaluation and Debiasing of Conversational Language ModelsabstractSoumya Barikeri, Anne Lauscher, Ivan Vulić, Goran Glavaš. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Soumya Barikeri, Anne Lauscher, Ivan Vulic, Goran Glavas |
ACL/IJCNLP (1) | 3 |
| 2021 | Verb Knowledge Injection for Multilingual Event ProcessingabstractOlga Majewska, Ivan Vulić, Goran Glavaš, Edoardo Maria Ponti, Anna Korhonen. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Olga Majewska, Ivan Vulic, Goran Glavas, Edoardo Maria Ponti, Anna Korhonen |
ACL/IJCNLP (1) | 2 |
| 2021 | How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language ModelsabstractPhillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, Iryna Gurevych. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Phillip Rust, Jonas Pfeiffer, Ivan Vulic, Sebastian Ruder, Iryna Gurevych |
ACL/IJCNLP (1) | 3 |
| 2021 | LexFit: Lexical Fine-Tuning of Pretrained Language ModelsabstractIvan Vulić, Edoardo Maria Ponti, Anna Korhonen, Goran Glavaš. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ivan Vulic, Edoardo Maria Ponti, Anna Korhonen, Goran Glavas |
ACL/IJCNLP (1) | 1 |
| 2021 | A Closer Look at Few-Shot Crosslingual Transfer: The Choice of Shots MattersabstractMengjie Zhao, Yi Zhu, Ehsan Shareghi, Ivan Vulić, Roi Reichart, Anna Korhonen, Hinrich Schütze. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ehsan Shareghi, Ivan Vulic, Roi Reichart, Anna Korhonen, Hinrich Schütze |
ACL/IJCNLP (1) | 4 |
| 2021 | MirrorWiC: On Eliciting Word-in-Context Representations from Pretrained Language ModelsabstractRecent work indicated that pretrained language models (PLMs) such as BERT and RoBERTa can be transformed into effective sentence and word encoders even via simple self-supervised techniques.Inspired by this line of work, in this paper we propose a fully unsupervised approach to improving word-in-context (WiC) representations in PLMs, achieved via a simple and efficient WiC-targeted fine-tuning procedure: MIRROR-WIC.The proposed method leverages only raw texts sampled from Wikipedia, assuming no sense-annotated data, and learns contextaware word representations within a standard contrastive learning setup.We experiment with a series of standard and comprehensive WiC benchmarks across multiple languages.Our proposed fully unsupervised MIRROR-WIC models obtain substantial gains over offthe-shelf PLMs across all monolingual, multilingual and cross-lingual setups.Moreover, on some standard WiC benchmarks, MIRROR-WIC is even on-par with supervised models fine-tuned with in-task data and sense labels. Qianchu Liu, Fangyu Liu 0001, Nigel Collier, Anna Korhonen, Ivan Vulic |
CoNLL | 5 |
| 2021 | Is Supervised Syntactic Parsing Beneficial for Language Understanding Tasks? An Empirical InvestigationabstractTraditional NLP has long held (supervised) syntactic parsing necessary for successful higher-level semantic language understanding (LU).The recent advent of end-to-end neural models, self-supervised via language modeling (LM), and their success on a wide range of LU tasks, however, questions this belief.In this work, we empirically investigate the usefulness of supervised parsing for semantic LU in the context of LM-pretrained transformer networks.Relying on the established fine-tuning paradigm, we first couple a pretrained transformer with a biaffine parsing head, aiming to infuse explicit syntactic knowledge from Universal Dependencies treebanks into the transformer.We then fine-tune the model for LU tasks and measure the effect of the intermediate parsing training (IPT) on downstream LU task performance.Results from both monolingual English and zero-shot language transfer experiments (with intermediate target-language parsing) show that explicit formalized syntax, injected into transformers through IPT, has very limited and inconsistent effect on downstream LU performance.Our results, coupled with our analysis of transformers' representation spaces before and after intermediate parsing, make a significant step towards providing answers to an essential question: how (un)availing is supervised parsing for high-level semantic natural language understanding in the era of large neural models?1 Disclaimer 1: In this work, we make a clear distinction between Computational Linguistics (CL), i.e., the area of linguistics leveraging computational methods for analyses of human languages and NLP, the area of artificial intelligence tackling human language in order to perform intelligent tasks.This work scrutinizes the usefulness of supervised parsing and explicit syntax only for the latter.We find the usefulness of explicit syntax in CL to be self-evident.2 Disclaimer 2: The purpose of this work is definitely not to invalidate the admirable efforts on syntactic annotation and modeling, but rather to make an empirically driven step towards a deeper understanding of the relationship between LU and formalised syntactic knowledge, and the extent of its impact to modern semantic LU and applications. Goran Glavas, Ivan Vulic |
EACL | 2 |
| 2021 | Evaluating Multilingual Text Encoders for Unsupervised Cross-Lingual Retrieval
Robert Litschko, Ivan Vulic, Simone Paolo Ponzetto, Goran Glavas |
ECIR (1) | 2 |
| 2021 | Fast, Effective, and Self-Supervised: Transforming Masked Language Models into Universal Lexical and Sentence EncodersabstractPrevious work has indicated that pretrained Masked Language Models (MLMs) are not effective as universal lexical and sentence encoders off-the-shelf, i.e., without further taskspecific fine-tuning on NLI, sentence similarity, or paraphrasing tasks using annotated task data.In this work, we demonstrate that it is possible to turn MLMs into effective lexical and sentence encoders even without any additional data, relying simply on self-supervision.We propose an extremely simple, fast, and effective contrastive learning technique, termed Mirror-BERT, which converts MLMs (e.g., BERT and RoBERTa) into such encoders in 20-30 seconds with no access to additional external knowledge.Mirror-BERT relies on identical and slightly modified string pairs as positive (i.e., synonymous) fine-tuning examples, and aims to maximise their similarity during "identity fine-tuning".We report huge gains over off-the-shelf MLMs with Mirror-BERT both in lexical-level and in sentencelevel tasks, across different domains and different languages.Notably, in sentence similarity (STS) and question-answer entailment (QNLI) tasks, our self-supervised Mirror-BERT model even matches the performance of the Sentence-BERT models from prior work which rely on annotated task data.Finally, we delve deeper into the inner workings of MLMs, and suggest some evidence on why this simple Mirror-BERT fine-tuning approach can yield effective universal lexical and sentence encoders. Fangyu Liu 0001, Ivan Vulic, Anna Korhonen, Nigel Collier |
EMNLP (1) | 2 |
| 2021 | Multilingual and Cross-Lingual Intent Detection from Spoken DataabstractDaniela Gerz, Pei-Hao Su, Razvan Kusztos, Avishek Mondal, Michał Lis, Eshan Singhal, Nikola Mrkšić, Tsung-Hsien Wen, Ivan Vulić. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Daniela Gerz, Pei-hao Su, Razvan Kusztos, Avishek Mondal, Michal Lis, Eshan Singhal, Nikola Mrksic, Tsung-Hsien Wen, Ivan Vulic |
EMNLP (1) | 9 |
| 2021 | AM2iCo: Evaluating Word Meaning in Context across Low-Resource Languages with Adversarial ExamplesabstractCapturing word meaning in context and distinguishing between correspondences and variations across languages is key to building successful multilingual and cross-lingual text representation models.However, existing multilingual evaluation datasets that evaluate lexical semantics "in-context" have various limitations.In particular, 1) their language coverage is restricted to high-resource languages and skewed in favor of only a few language families and areas, 2) a design that makes the task solvable via superficial cues, which results in artificially inflated (and sometimes super-human) performances of pretrained encoders, and 3) no support for crosslingual evaluation.In order to address these gaps, we present AM 2 ICO (Adversarial and Multilingual Meaning in Context), a widecoverage cross-lingual and multilingual evaluation set; it aims to faithfully assess the ability of state-of-the-art (SotA) representation models to understand the identity of word meaning in cross-lingual contexts for 14 language pairs.We conduct a series of experiments in a wide range of setups and demonstrate the challenging nature of AM 2 ICO.The results reveal that current SotA pretrained encoders substantially lag behind human performance, and the largest gaps are observed for low-resource languages and languages dissimilar to English. Qianchu Liu, Edoardo Maria Ponti, Diana McCarthy, Ivan Vulic, Anna Korhonen |
EMNLP (1) | 4 |
| 2021 | UNKs Everywhere: Adapting Multilingual Language Models to New ScriptsabstractMassively multilingual language models such as multilingual BERT offer state-of-the-art cross-lingual transfer performance on a range of NLP tasks.However, due to limited capacity and large differences in pretraining data sizes, there is a profound performance gap between resource-rich and resource-poor target languages.The ultimate challenge is dealing with under-resourced languages not covered at all by the models and written in scripts unseen during pretraining.In this work, we propose a series of novel data-efficient methods that enable quick and effective adaptation of pretrained multilingual models to such lowresource languages and unseen scripts.Relying on matrix factorization, our methods capitalize on the existing latent knowledge about multiple languages already available in the pretrained model's embedding matrix.Furthermore, we show that learning of the new dedicated embedding matrix in the target language can be improved by leveraging a small number of vocabulary items (i.e., the so-called lexically overlapping tokens) shared between mBERT's and target language vocabulary.Our adaptation techniques offer substantial performance gains for languages with unseen scripts.We also demonstrate that they can yield improvements for low-resource languages written in scripts covered by the pretrained model. Jonas Pfeiffer, Ivan Vulic, Iryna Gurevych, Sebastian Ruder |
EMNLP (1) | 2 |
| 2021 | ConvFiT: Conversational Fine-Tuning of Pretrained Language ModelsabstractIvan Vulić, Pei-Hao Su, Samuel Coope, Daniela Gerz, Paweł Budzianowski, Iñigo Casanueva, Nikola Mrkšić, Tsung-Hsien Wen. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Ivan Vulic, Pei-hao Su, Sam Coope, Daniela Gerz, Pawel Budzianowski, Iñigo Casanueva, Nikola Mrksic, Tsung-Hsien Wen |
EMNLP (1) | 1 |
| 2021 | ConVEx: Data-Efficient and Few-Shot Slot LabelingabstractWe propose ConVEx (Conversational Value Extractor), an efficient pretraining and finetuning neural approach for slot-labeling dialog tasks.Instead of relying on more general pretraining objectives from prior work (e.g., language modeling, response selection), Con-VEx's pretraining objective, a novel pairwise cloze task using Reddit data, is well aligned with its intended usage on sequence labeling tasks.This enables learning domain-specific slot labelers by simply fine-tuning decoding layers of the pretrained general-purpose sequence labeling model, while the majority of the pretrained model's parameters are kept frozen.We report state-of-the-art performance of ConVEx across a range of diverse domains and data sets for dialog slot-labeling, with the largest gains in the most challenging, few-shot setups.We believe that ConVEx's reduced pretraining times (i.e., only 18 hours on 12 GPUs) and cost, along with its efficient finetuning and strong performance, promise wider portability and scalability for data-efficient sequence-labeling tasks in general. Matthew Henderson, Ivan Vulic |
NAACL-HLT | 2 |
| 2021 | Semantic Data Set Construction from Human Clustering and Spatial ArrangementabstractAbstract Research into representation learning models of lexical semantics usually utilizes some form of intrinsic evaluation to ensure that the learned representations reflect human semantic judgments. Lexical semantic similarity estimation is a widely used evaluation method, but efforts have typically focused on pairwise judgments of words in isolation, or are limited to specific contexts and lexical stimuli. There are limitations with these approaches that either do not provide any context for judgments, and thereby ignore ambiguity, or provide very specific sentential contexts that cannot then be used to generate a larger lexical resource. Furthermore, similarity between more than two items is not considered. We provide a full description and analysis of our recently proposed methodology for large-scale data set construction that produces a semantic classification of a large sample of verbs in the first phase, as well as multi-way similarity judgments made within the resultant semantic classes in the second phase. The methodology uses a spatial multi-arrangement approach proposed in the field of cognitive neuroscience for capturing multi-way similarity judgments of visual stimuli. We have adapted this method to handle polysemous linguistic stimuli and much larger samples than previous work. We specifically target verbs, but the method can equally be applied to other parts of speech. We perform cluster analysis on the data from the first phase and demonstrate how this might be useful in the construction of a comprehensive verb resource. We also analyze the semantic information captured by the second phase and discuss the potential of the spatially induced similarity judgments to better reflect human notions of word similarity. We demonstrate how the resultant data set can be used for fine-grained analyses and evaluation of representation learning models on the intrinsic tasks of semantic clustering and semantic similarity. In particular, we find that stronger static word embedding methods still outperform lexical representations emerging from more recent pre-training methods, both on word-level similarity and clustering. Moreover, thanks to the data set’s vast coverage, we are able to compare the benefits of specializing vector representations for a particular type of external knowledge by evaluating FrameNet- and VerbNet-retrofitted models on specific semantic domains such as “Heat” or “Motion.” Olga Majewska, Diana McCarthy, Jasper J. F. van den Bosch, Nikolaus Kriegeskorte, Ivan Vulic, Anna Korhonen |
Comput. Linguistics | 5 |
| 2021 | Parameter Space Factorization for Zero-Shot Learning across Tasks and LanguagesabstractAbstract Most combinations of NLP tasks and language varieties lack in-domain examples for supervised training because of the paucity of annotated data. How can neural models make sample-efficient generalizations from task–language combinations with available data to low-resource ones? In this work, we propose a Bayesian generative model for the space of neural parameters. We assume that this space can be factorized into latent variables for each language and each task. We infer the posteriors over such latent variables based on data from seen task–language combinations through variational inference. This enables zero-shot classification on unseen combinations at prediction time. For instance, given training data for named entity recognition (NER) in Vietnamese and for part-of-speech (POS) tagging in Wolof, our model can perform accurate predictions for NER in Wolof. In particular, we experiment with a typologically diverse sample of 33 languages from 4 continents and 11 families, and show that our model yields comparable or better results than state-of-the-art, zero-shot cross-lingual transfer methods. Our code is available at github.com/cambridgeltl/parameter-factorization. Edoardo Maria Ponti, Ivan Vulic, Ryan Cotterell, Marinela Parovic, Roi Reichart, Anna Korhonen |
Trans. Assoc. Comput. Linguistics | 2 |
| 2020 | A General Framework for Implicit and Explicit Debiasing of Distributional Word Vector SpacesabstractDistributional word vectors have recently been shown to encode many of the human biases, most notably gender and racial biases, and models for attenuating such biases have consequently been proposed. However, existing models and studies (1) operate on under-specified and mutually differing bias definitions, (2) are tailored for a particular bias (e.g., gender bias) and (3) have been evaluated inconsistently and non-rigorously. In this work, we introduce a general framework for debiasing word embeddings. We operationalize the definition of a bias by discerning two types of bias specification: explicit and implicit. We then propose three debiasing models that operate on explicit or implicit bias specifications and that can be composed towards more robust debiasing. Finally, we devise a full-fledged evaluation framework in which we couple existing bias metrics with newly proposed ones. Experimental findings across three embedding methods suggest that the proposed debiasing models are robust and widely applicable: they often completely remove the bias both implicitly and explicitly without degradation of semantic information encoded in any of the input distributional spaces. Moreover, we successfully transfer debiasing models, by means of cross-lingual embedding spaces, and remove or attenuate biases in distributional word vector spaces of languages that lack readily available bias specifications. Anne Lauscher, Goran Glavas, Simone Paolo Ponzetto, Ivan Vulic |
AAAI | 4 |
| 2020 | Span-ConveRT: Few-shot Span Extraction for Dialog with Pretrained Conversational RepresentationsabstractWe introduce Span-ConveRT, a light-weight model for dialog slot-filling which frames the task as a turn-based span extraction task.This formulation allows for a simple integration of conversational knowledge coded in large pretrained conversational models such as Con-veRT (Henderson et al., 2019a).We show that leveraging such knowledge in Span-ConveRT is especially useful for few-shot learning scenarios: we report consistent gains over 1) a span extractor that trains representations from scratch in the target domain, and 2) a BERTbased span extractor.In order to inspire more work on span extraction for the slot-filling task, we also release RESTAURANTS-8K, a new challenging data set of 8,198 utterances, compiled from actual conversations in the restaurant booking domain. Sam Coope, Tyler Farghly, Daniela Gerz, Ivan Vulic, Matthew Henderson |
ACL | 4 |
| 2020 | Multidirectional Associative Optimization of Function-Specific Word RepresentationsabstractWe present a neural framework for learning associations between interrelated groups of words such as the ones found in Subject-Verb-Object (SVO) structures.Our model induces a joint function-specific word vector space, where vectors of e.g.plausible SVO compositions lie close together.The model retains information about word group membership even in the joint space, and can thereby effectively be applied to a number of tasks reasoning over the SVO structure.We show the robustness and versatility of the proposed framework by reporting state-of-the-art results on the tasks of estimating selectional preference and event similarity.The results indicate that the combinations of representations learned with our task-independent model outperform task-specific architectures from prior work, while reducing the number of parameters by up to 95%. Daniela Gerz, Ivan Vulic, Marek Rei, Roi Reichart, Anna Korhonen |
ACL | 2 |
| 2020 | Non-Linear Instance-Based Cross-Lingual Mapping for Non-Isomorphic Embedding SpacesabstractWe present INSTAMAP, an instance-based method for learning projection-based crosslingual word embeddings.Unlike prior work, it deviates from learning a single global linear projection.INSTAMAP is a non-parametric model that learns a non-linear projection by iteratively: (1) finding a globally optimal rotation of the source embedding space relying on the Kabsch algorithm, and then (2) moving each point along an instance-specific translation vector estimated from the translation vectors of the point's nearest neighbours in the training dictionary.We report performance gains with INSTAMAP over four representative state-of-the-art projection-based models on bilingual lexicon induction across a set of 28 diverse language pairs.We note prominent improvements, especially for more distant language pairs (i.e., languages with nonisomorphic monolingual spaces). Goran Glavas, Ivan Vulic |
ACL | 2 |
| 2020 | Classification-Based Self-Learning for Weakly Supervised Bilingual Lexicon InductionabstractEffective projection-based cross-lingual word embedding (CLWE) induction critically relies on the iterative self-learning procedure.It gradually expands the initial small seed dictionary to learn improved cross-lingual mappings.In this work, we present CLASSYMAP, a classification-based approach to self-learning, yielding a more robust and a more effective induction of projection-based CLWEs.Unlike prior self-learning methods, our approach allows for integration of diverse features into the iterative process.We show the benefits of CLASSYMAP for bilingual lexicon induction: we report consistent improvements in a weakly supervised setup (500 seed translation pairs) on a benchmark with 28 language pairs. Mladen Karan, Ivan Vulic, Anna Korhonen, Goran Glavas |
ACL | 2 |
| 2020 | XHate-999: Analyzing and Detecting Abusive Language Across Domains and LanguagesabstractWe present XHATE-999, a multi-domain and multilingual evaluation data set for abusive language detection.By aligning test instances across six typologically diverse languages, XHATE-999 for the first time allows for disentanglement of the domain transfer and language transfer effects in abusive language detection.We conduct a series of domain-and language-transfer experiments with state-of-the-art monolingual and multilingual transformer models, setting strong baseline results and profiling XHATE-999 as a comprehensive evaluation resource for abusive language detection.Finally, we show that domain-and language-adaptation, via intermediate masked language modeling on abusive corpora in the target language, can lead to substantially improved abusive language detection in the target language in the zero-shot transfer setups. Goran Glavas, Mladen Karan, Ivan Vulic |
COLING | 3 |
| 2020 | Specializing Unsupervised Pretraining Models for Word-Level Semantic SimilarityabstractUnsupervised pretraining models have been shown to facilitate a wide range of downstream NLP applications. These models, however, retain some of the limitations of traditional static word embeddings. In particular, they encode only the distributional knowledge available in raw text corpora, incorporated through language modeling objectives. In this work, we complement such distributional knowledge with external lexical knowledge, that is, we integrate the discrete knowledge on word-level semantic similarity into pretraining. To this end, we generalize the standard BERT model to a multi-task learning setting where we couple BERT’s masked language modeling and next sentence prediction objectives with an auxiliary task of binary word relation classification. Our experiments suggest that our "Lexically Informed” BERT (LIBERT), specialized for the word-level semantic similarity, yields better performance than the lexically blind “vanilla” BERT on several language understanding tasks. Concretely, LIBERT outperforms BERT in 9 out of 10 tasks of the GLUE benchmark and is on a par with BERT in the remaining one. Moreover, we show consistent gains on 3 benchmarks for lexical simplification, a task where knowledge about word-level semantic similarity is paramount, as well as large gains on lexical reasoning probes. Anne Lauscher, Ivan Vulic, Edoardo Maria Ponti, Anna Korhonen, Goran Glavas |
COLING | 2 |
| 2020 | Emergent Communication Pretraining for Few-Shot Machine TranslationabstractWhile state-of-the-art models that rely upon massively multilingual pretrained encoders achieve sample efficiency in downstream applications, they still require abundant amounts of unlabelled text.Nevertheless, most of the world's languages lack such resources.Hence, we investigate a more radical form of unsupervised knowledge transfer in the absence of linguistic data.In particular, for the first time we pretrain neural networks via emergent communication from referential games.Our key assumption is that grounding communication on images-as a crude approximation of real-world environments-inductively biases the model towards learning natural languages.On the one hand, we show that this substantially benefits machine translation in few-shot settings.On the other hand, this also provides an extrinsic evaluation protocol to probe the properties of emergent languages ex vitro.Intuitively, the closer they are to natural languages, the higher the gains from pretraining on them should be.For instance, in this work we measure the influence of communication success and maximum sequence length on downstream performances.Finally, we introduce a customised adapter layer and annealing strategies for the regulariser of maximum-a-posteriori inference during fine-tuning.These turn out to be crucial to facilitate knowledge transfer and prevent catastrophic forgetting.Compared to a recurrent baseline, our method yields gains of 59.0%∼147.6% in BLEU score with only 500 NMT training instances and 65.1%∼196.7%with 1, 000 NMT training instances across four language pairs.These proofof-concept results reveal the potential of emergent communication pretraining for both natural language processing tasks in resource-poor settings and extrinsic evaluation of artificial languages. Yaoyiran Li, Edoardo Maria Ponti, Ivan Vulic, Anna Korhonen |
COLING | 3 |
| 2020 | Towards Instance-Level Parser Selection for Cross-Lingual Transfer of Dependency ParsersabstractCurrent methods of cross-lingual parser transfer focus on predicting the best parser for a lowresource target language globally, that is, "at treebank level".In this work, we propose and argue for a novel cross-lingual transfer paradigm: instance-level parser selection (ILPS), and present a proof-of-concept study focused on instance-level selection in the framework of delexicalized parser transfer.Our work is motivated by an empirical observation that different source parsers are the best choice for different Universal POS-sequences (i.e., UPOS sentences) in the target language.We then propose to predict the best parser at the instance level.To this end, we train a supervised regression model, based on the Transformer architecture, to predict parser accuracies for individual POS-sequences.We compare ILPS against two strong single-best parser selection baselines (SBPS): (1) a model that compares POS n-gram distributions between the source and target languages (KL) and ( 2) a model that selects the source based on the similarity between manually created language vectors encoding syntactic properties of languages (L2V).The results from our extensive evaluation, coupling 42 source parsers and 20 diverse low-resource test languages, show that ILPS outperforms KL and L2V on 13/20 and 14/20 test languages, respectively.Further, we show that by predicting the best parser "at treebank level" (SBPS), using the aggregation of predictions from our instance-level model, we outperform the same baselines on 17/20 and 16/20 test languages. Robert Litschko, Ivan Vulic, Zeljko Agic, Goran Glavas |
COLING | 2 |
| 2020 | Manual Clustering and Spatial Arrangement of Verbs for Multilingual Evaluation and Typology AnalysisabstractWe present the first evaluation of the applicability of a spatial arrangement method (SpAM) to a typologically diverse language sample, and its potential to produce semantic evaluation resources to support multilingual NLP, with a focus on verb semantics.We demonstrate SpAM's utility in allowing for quick bottom-up creation of large-scale evaluation datasets that balance cross-lingual alignment with language specificity.Starting from a shared sample of 825 English verbs, translated into Chinese, Japanese, Finnish, Polish, and Italian, we apply a two-phase annotation process which produces (i) semantic verb classes and (ii) fine-grained similarity scores for nearly 130 thousand verb pairs.We use the two types of verb data to (a) examine cross-lingual similarities and variation, and (b) evaluate the capacity of static and contextualised representation models to accurately reflect verb semantics, contrasting the performance of large language-specific pretraining models with their multilingual equivalent on semantic clustering and lexical similarity, across different domains of verb meaning.We release the data from both phases as a large-scale multilingual resource, comprising 85 verb classes and nearly 130k pairwise similarity scores, offering a wealth of possibilities for further evaluation and research on multilingual verb semantics. Olga Majewska, Ivan Vulic, Diana McCarthy, Anna Korhonen |
COLING | 2 |
| 2020 | The Secret is in the Spectra: Predicting Cross-lingual Task Performance with Spectral Similarity MeasuresabstractPerformance in cross-lingual NLP tasks is impacted by the (dis)similarity of languages at hand: e.g., previous work has suggested there is a connection between the expected success of bilingual lexicon induction (BLI) and the assumption of (approximate) isomorphism between monolingual embedding spaces.In this work we present a large-scale study focused on the correlations between monolingual embedding space similarity and task performance, covering thousands of language pairs and four different tasks: BLI, parsing, POS tagging and MT.We hypothesize that statistics of the spectrum of each monolingual embedding space indicate how well they can be aligned.We then introduce several isomorphism measures between two embedding spaces, based on the relevant statistics of their individual spectra.We empirically show that 1) language similarity scores derived from such spectral isomorphism measures are strongly associated with performance observed in different crosslingual tasks, and 2) our spectral-based measures consistently outperform previous standard isomorphism measures, while being computationally more tractable and easier to interpret.Finally, our measures capture complementary information to typologically driven language distance measures, and the combination of measures from the two families yields even higher task performance correlations. Haim Dubossarsky, Ivan Vulic, Roi Reichart, Anna Korhonen |
EMNLP (1) | 2 |
| 2020 | From Zero to Hero: On the Limitations of Zero-Shot Language Transfer with Multilingual Transformersabstractbak probing, til beskrivelser av transformerarkitekturen og debatten rundt tolkbarhet av oppmerksomhet. I hvert introduksjonskapittel prver vi derfor kontekstualisere vrt eget arbeid. Anne Lauscher, Vinit Ravishankar, Ivan Vulic, Goran Glavas |
EMNLP (1) | 3 |
| 2020 | MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual TransferabstractThe main goal behind state-of-the-art pretrained multilingual models such as multilingual BERT and XLM-R is enabling and bootstrapping NLP applications in low-resource languages through zero-shot or few-shot crosslingual transfer.However, due to limited model capacity, their transfer performance is the weakest exactly on such low-resource languages and languages unseen during pretraining.We propose MAD-X, an adapter-based framework that enables high portability and parameter-efficient transfer to arbitrary tasks and languages by learning modular language and task representations.In addition, we introduce a novel invertible adapter architecture and a strong baseline method for adapting a pretrained multilingual model to a new language.MAD-X outperforms the state of the art in cross-lingual transfer across a representative set of typologically diverse languages on named entity recognition and causal commonsense reasoning, and achieves competitive results on question answering.Our code and adapters are available at AdapterHub.ml. Jonas Pfeiffer, Ivan Vulic, Iryna Gurevych, Sebastian Ruder |
EMNLP (1) | 2 |
| 2020 | XCOPA: A Multilingual Dataset for Causal Commonsense ReasoningabstractIn order to simulate human language capacity, natural language processing systems must be able to reason about the dynamics of everyday situations, including their possible causes and effects. Moreover, they should be able to generalise the acquired world knowledge to new languages, modulo cultural differences. Advances in machine reasoning and cross-lingual transfer depend on the availability of challenging evaluation benchmarks. Motivated by both demands, we introduce Cross-lingual Choice of Plausible Alternatives (XCOPA), a typologically diverse multilingual dataset for causal commonsense reasoning in 11 languages, which includes resource-poor languages like Eastern Apurímac Quechua and Haitian Creole. We evaluate a range of state-of-the-art models on this novel dataset, revealing that the performance of current methods based on multilingual pretraining and zero-shot fine-tuning falls short compared to translation-based transfer. Finally, we propose strategies to adapt multilingual models to out-of-sample resource-lean languages where only a small corpus or a bilingual dictionary is available, and report substantial improvements over the random baseline. The XCOPA dataset is freely available at github.com/cambridgeltl/xcopa Edoardo Maria Ponti, Goran Glavas, Olga Majewska, Qianchu Liu, Ivan Vulic, Anna Korhonen |
EMNLP (1) | 5 |
| 2020 | Probing Pretrained Language Models for Lexical SemanticsabstractThe success of large pretrained language models (LMs) such as BERT and RoBERTa has sparked interest in probing their representations, in order to unveil what types of knowledge they implicitly capture.While prior research focused on morphosyntactic, semantic, and world knowledge, it remains unclear to which extent LMs also derive lexical type-level knowledge from words in context.In this work, we present a systematic empirical analysis across six typologically diverse languages and five different lexical tasks, addressing the following questions: 1) How do different lexical knowledge extraction strategies (monolingual versus multilingual source LM, out-ofcontext versus in-context encoding, inclusion of special tokens, and layer-wise averaging) impact performance?How consistent are the observed effects across tasks and languages?2) Is lexical knowledge stored in few parameters, or is it scattered throughout the network?3) How do these representations fare against traditional static word vectors in lexical tasks?4) Does the lexical information emerging from independently trained monolingual LMs display latent similarities?Our main results indicate patterns and best practices that hold universally, but also point to prominent variations across languages and tasks.Moreover, we validate the claim that lower Transformer layers carry more type-level lexical knowledge, but also show that this knowledge is distributed across multiple layers. Ivan Vulic, Edoardo Maria Ponti, Robert Litschko, Goran Glavas, Anna Korhonen |
EMNLP (1) | 1 |
| 2020 | Are All Good Word Vector Spaces Isomorphic?abstractExisting algorithms for aligning cross-lingual word vector spaces assume that vector spaces are approximately isomorphic.As a result, they perform poorly or fail completely on nonisomorphic spaces.Such non-isomorphism has been hypothesised to result from typological differences between languages.In this work, we ask whether non-isomorphism is also crucially a sign of degenerate word vector spaces.We present a series of experiments across diverse languages which show that variance in performance across language pairs is not only due to typological differences, but can mostly be attributed to the size of the monolingual resources available, and to the properties and duration of monolingual training (e.g."under-training"). Ivan Vulic, Sebastian Ruder, Anders Søgaard |
EMNLP (1) | 1 |
| 2020 | Spatial Multi-Arrangement for Clustering and Multi-way Similarity Dataset ConstructionabstractWe present a novel methodology for fast bottom-up creation of large-scale semantic similarity resources to support development and evaluation of NLP systems. Our work targets verb similarity, but the methodology is equally applicable to other parts of speech. Our approach circumvents the bottleneck of slow and expensive manual development of lexical resources by leveraging semantic intuitions of native speakers and adapting a spatial multi-arrangement approach from cognitive neuroscience, used before only with visual stimuli, to lexical stimuli. Our approach critically obtains judgments of word similarity in the context of a set of related words, rather than of word pairs in isolation. We also handle lexical ambiguity as a natural consequence of a two-phase process where verbs are placed in broad semantic classes prior to the fine-grained spatial similarity judgments. Our proposed design produces a large-scale verb resource comprising 17 relatedness-based classes and a verb similarity dataset containing similarity scores for 29,721 unique verb pairs and 825 target verbs, which we release with this paper. Olga Majewska, Diana McCarthy, Jasper J. F. van den Bosch, Nikolaus Kriegeskorte, Ivan Vulic, Anna Korhonen |
LREC | 5 |
| 2020 | Multi-SimLex: A Large-Scale Evaluation of Multilingual and Crosslingual Lexical Semantic SimilarityabstractWe introduce Multi-SimLex, a large-scale lexical resource and evaluation benchmark covering data sets for 12 typologically diverse languages, including major languages (e.g., Mandarin Chinese, Spanish, Russian) as well as less-resourced ones (e.g., Welsh, Kiswahili). Each language data set is annotated for the lexical relation of semantic similarity and contains 1,888 semantically aligned concept pairs, providing a representative coverage of word classes (nouns, verbs, adjectives, adverbs), frequency ranks, similarity intervals, lexical fields, and concreteness levels. Additionally, owing to the alignment of concepts across languages, we provide a suite of 66 crosslingual semantic similarity data sets. Because of its extensive size and language coverage, Multi-SimLex provides entirely novel opportunities for experimental evaluation and analysis. On its monolingual and crosslingual benchmarks, we evaluate and analyze a wide array of recent state-of-the-art monolingual and crosslingual representation models, including static and contextualized word embeddings (such as fastText, monolingual and multilingual BERT, XLM), externally informed lexical representations, as well as fully unsupervised and (weakly) supervised crosslingual word embeddings. We also present a step-by-step data set creation protocol for creating consistent, Multi-Simlex–style resources for additional languages. We make these contributions—the public release of Multi-SimLex data sets, their creation protocol, strong baseline results, and in-depth analyses which can be helpful in guiding future developments in multilingual lexical semantics and representation learning—available via a Web site that will encourage community effort in further expansion of Multi-Simlex to many more languages. Such a large-scale semantic resource could inspire significant further advances in NLP across languages. Ivan Vulic, Simon Baker, Edoardo Maria Ponti, Ulla Petti, Ira Leviant, Kelly Wing, Olga Majewska, Eden Bar, Matt Malone, Thierry Poibeau, Roi Reichart, Anna Korhonen |
Comput. Linguistics | 1 |
| 2019 | JW300: A Wide-Coverage Parallel Corpus for Low-Resource LanguagesabstractViable cross-lingual transfer critically depends on the availability of parallel texts.Shortage of such resources imposes a development and evaluation bottleneck in multilingual processing.We introduce JW300, a parallel corpus of over 300 languages with around 100 thousand parallel sentences per language pair on average.In this paper, we present the resource and showcase its utility in experiments with crosslingual word embedding induction and multisource part-of-speech projection. Zeljko Agic, Ivan Vulic |
ACL (1) | 2 |
| 2019 | How to (Properly) Evaluate Cross-Lingual Word Embeddings: On Strong Baselines, Comparative Analyses, and Some MisconceptionsabstractCross-lingual word embeddings (CLEs) enable multilingual modeling of meaning and facilitate cross-lingual transfer of NLP models. Despite their ubiquitous usage in downstream tasks, recent increasingly popular projection-based CLE models are almost exclusively evaluated on a single task only: bilingual lexicon induction (BLI). Even BLI evaluations vary greatly, hindering our ability to correctly interpret performance and properties of different CLE models. In this work, we make the first step towards a comprehensive evaluation of cross-lingual word embeddings. We thoroughly evaluate both supervised and unsupervised CLE models on a large number of language pairs in the BLI task and three downstream tasks, providing new insights concerning the ability of cutting-edge CLE models to support cross-lingual NLP. We empirically demonstrate that the performance of CLE models largely depends on the task at hand and that optimizing CLE models for BLI can result in deteriorated downstream performance. We indicate the most robust supervised and unsupervised CLE models and emphasize the need to reassess existing baselines, which still display competitive performance across the board. We hope that our work will catalyze further work on CLE evaluation and model analysis. Goran Glavas, Robert Litschko, Sebastian Ruder, Ivan Vulic |
ACL (1) | 4 |
| 2019 | Generalized Tuning of Distributional Word Vectors for Monolingual and Cross-Lingual Lexical EntailmentabstractLexical entailment (LE; also known as hyponymy-hypernymy or is-a relation) is a core asymmetric lexical relation that supports tasks like taxonomy induction and text generation.In this work, we propose a simple and effective method for fine-tuning distributional word vectors for LE.Our Generalized Lexical ENtailment model (GLEN) is decoupled from the word embedding model and applicable to any distributional vector space.Yet -unlike existing retrofitting models -it captures a general specialization function allowing for LE-tuning of the entire distributional space and not only the vectors of words seen in lexical constraints.Coupled with a multilingual embedding space, GLEN seamlessly enables cross-lingual LE detection.We demonstrate the effectiveness of GLEN in graded LE and report large improvements (over 20% in accuracy) over state-ofthe-art in cross-lingual LE detection. Goran Glavas, Ivan Vulic |
ACL (1) | 2 |
| 2019 | Training Neural Response Selection for Task-Oriented Dialogue SystemsabstractMatthew Henderson, Ivan Vulić, Daniela Gerz, Iñigo Casanueva, Paweł Budzianowski, Sam Coope, Georgios Spithourakis, Tsung-Hsien Wen, Nikola Mrkšić, Pei-Hao Su. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Matthew Henderson, Ivan Vulic, Daniela Gerz, Iñigo Casanueva, Pawel Budzianowski, Sam Coope, Georgios Spithourakis, Tsung-Hsien Wen, Nikola Mrksic, Pei-hao Su |
ACL (1) | 2 |
| 2019 | Multilingual and Cross-Lingual Graded Lexical EntailmentabstractGrounded in cognitive linguistics, graded lexical entailment (GR-LE) is concerned with finegrained assertions regarding the directional hierarchical relationships between concepts on a continuous scale.In this paper, we present the first work on cross-lingual generalisation of GR-LE relation.Starting from Hyper-Lex, the only available GR-LE dataset in English, we construct new monolingual GR-LE datasets for three other languages, and combine those to create a set of six cross-lingual GR-LE datasets termed CL-HYPERLEX.We next present a novel method dubbed CLEAR (Cross-Lingual Lexical Entailment Attract-Repel) for effectively capturing graded (and binary) LE, both monolingually in different languages as well as across languages (i.e., on CL-HYPERLEX).Coupled with a bilingual dictionary, CLEAR leverages taxonomic LE knowledge in a resource-rich language (e.g., English) and propagates it to other languages.Supported by cross-lingual LE transfer, CLEAR sets competitive baseline performance on three new monolingual GR-LE datasets and six cross-lingual GR-LE datasets.In addition, we show that CLEAR outperforms current state-ofthe-art on binary cross-lingual LE detection by a wide margin for diverse language pairs.x y z en beagle es perro es mamífero en animal es organismo en roadster es coche en vehicle es transporte Ivan Vulic, Simone Paolo Ponzetto, Goran Glavas |
ACL (1) | 1 |
| 2019 | Investigating Cross-Lingual Alignment Methods for Contextualized Embeddings with Token-Level EvaluationabstractIn this paper, we present a thorough investigation on methods that align pre-trained contextualized embeddings into shared crosslingual context-aware embedding space, providing strong reference benchmarks for future context-aware crosslingual models.We propose a novel and challenging task, Bilingual Token-level Sense Retrieval (BTSR).It specifically evaluates the accurate alignment of words with the same meaning in crosslingual non-parallel contexts, currently not evaluated by existing tasks such as Bilingual Contextual Word Similarity and Sentence Retrieval.We show how the proposed BTSR task highlights the merits of different alignment methods.In particular, we find that using context average type-level alignment is effective in transferring monolingual contextualized embeddings cross-lingually especially in non-parallel contexts, and at the same time improves the monolingual space.Furthermore, aligning independently trained models yields better performance than aligning multilingual embeddings with shared vocabulary. Qianchu Liu, Diana McCarthy, Ivan Vulic, Anna Korhonen |
CoNLL | 3 |
| 2019 | On the Importance of Subword Information for Morphological Tasks in Truly Low-Resource LanguagesabstractRecent work has validated the importance of subword information for word representation learning. Since subwords increase parameter sharing ability in neural models, their value should be even more pronounced in low-data regimes. In this work, we therefore provide a comprehensive analysis focused on the usefulness of subwords for word representation learning in truly low-resource scenarios and for three representative morphological tasks: fine-grained entity typing, morphological tagging, and named entity recognition. We conduct a systematic study that spans several dimensions of comparison: 1) type of data scarcity which can stem from the lack of task-specific training data, or even from the lack of unannotated data required to train word embeddings, or both; 2) language type by working with a sample of 16 typologically diverse languages including some truly low-resource ones (e.g. Rusyn, Buryat, and Zulu); 3) the choice of the subword-informed word representation method. Our main results show that subword-informed models are universally useful across all language types, with large gains over subword-agnostic embeddings. They also suggest that the effective use of subwords largely depends on the language (type) and the task at hand, as well as on the amount of available data for training the embeddings and task-based models, where having sufficient in-task data is a more critical requirement. Benjamin Heinzerling, Ivan Vulic, Michael Strube 0001, Roi Reichart, Anna Korhonen |
CoNLL | 3 |
| 2019 | Zero-Shot Language Transfer for Cross-Lingual Sentence Retrieval Using Bidirectional Attention Model
Goran Glavas, Ivan Vulic |
ECIR (1) | 2 |
| 2019 | Towards Zero-shot Language ModelingabstractEdoardo Maria Ponti, Ivan Vulić, Ryan Cotterell, Roi Reichart, Anna Korhonen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Edoardo Maria Ponti, Ivan Vulic, Ryan Cotterell, Roi Reichart, Anna Korhonen |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Cross-lingual Semantic Specialization via Lexical Relation InductionabstractEdoardo Maria Ponti, Ivan Vulić, Goran Glavaš, Roi Reichart, Anna Korhonen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Edoardo Maria Ponti, Ivan Vulic, Goran Glavas, Roi Reichart, Anna Korhonen |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Do We Really Need Fully Unsupervised Cross-Lingual Embeddings?abstractIvan Vulić, Goran Glavaš, Roi Reichart, Anna Korhonen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Ivan Vulic, Goran Glavas, Roi Reichart, Anna Korhonen |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Evaluating Resource-Lean Cross-Lingual Embedding Models in Unsupervised RetrievalabstractCross-lingual embeddings (CLE) facilitate cross-lingual natural language processing and information retrieval. Recently, a wide variety of resource-lean projection-based models for inducing CLEs has been introduced, requiring limited or no bilingual supervision. Despite potential usefulness in downstream IR and NLP tasks, these CLE models have almost exclusively been evaluated on word translation tasks. In this work, we provide a comprehensive comparative evaluation of projection-based CLE models for both sentence-level and document-level cross-lingual Information Retrieval (CLIR). We show that in some settings resource-lean CLE-based CLIR models may outperform resource-intensive models using full-blown machine translation (MT). We hope our work serves as a guideline for choosing the right model for CLIR practitioners. Robert Litschko, Goran Glavas, Ivan Vulic, Laura Dietz |
SIGIR | 3 |
| 2019 | Modeling Language Variation and Universals: A Survey on Typological Linguistics for Natural Language ProcessingabstractLinguistic typology aims to capture structural and semantic variation across the world’s languages. A large-scale typology could provide excellent guidance for multilingual Natural Language Processing (NLP), particularly for languages that suffer from the lack of human labeled resources. We present an extensive literature survey on the use of typological information in the development of NLP techniques. Our survey demonstrates that to date, the use of information in existing typological databases has resulted in consistent but modest improvements in system performance. We show that this is due to both intrinsic limitations of databases (in terms of coverage and feature granularity) and under-utilization of the typological features included in them. We advocate for a new approach that adapts the broad and discrete nature of typological categories to the contextual and continuous nature of machine learning algorithms used in contemporary NLP. In particular, we suggest that such an approach could be facilitated by recent developments in data-driven induction of typological knowledge. Edoardo Maria Ponti, Helen O'Horan, Yevgeni Berzak, Ivan Vulic, Roi Reichart, Thierry Poibeau, Ekaterina Shutova, Anna Korhonen |
Comput. Linguistics | 4 |
| 2019 | Classification as an approach to public key infrastructure requirements analysisabstractClassification schemes and taxonomy of request have focused only on business or software requirements defined by general non‐functional requests. Public key infrastructure (PKI), as a complex system, requires special attention in the process of identification and definition of requests in all parts of the system. Requests efficiently identified and defined are the basis for successful PKI development and implementation. This study proposes a new classification scheme of requirements that enable effective identification of PKI requests in all system development life cycle phases. The proposed classification scheme has been designed in a way that can be used both by organisations just introducing a PKI system in their business operations, as well as organisations developing and implementing the PKI. Although there are more classification schemes and taxonomy of requests, the proposed classification scheme is the optimum solution for identification and definition of PKI system requests, as well as a good basis for testing, verification and validation of the PKI final product. Radomir Prodanovic, Ivan Vulic |
IET Softw. | 2 |
| 2019 | A Survey of Cross-lingual Word Embedding ModelsabstractCross-lingual representations of words enable us to reason about word meaning in multilingual contexts and are a key facilitator of cross-lingual transfer when developing natural language processing models for low-resource languages. In this survey, we provide a comprehensive typology of cross-lingual word embedding models. We compare their data requirements and objective functions. The recurring theme of the survey is that many of the models presented in the literature optimize for the same objectives, and that seemingly different models are often equivalent, modulo optimization strategies, hyper-parameters, and such. We also discuss the different ways cross-lingual word embeddings are evaluated, as well as future challenges and research horizons. Sebastian Ruder, Ivan Vulic, Anders Søgaard |
J. Artif. Intell. Res. | 2 |
| 2018 | Isomorphic Transfer of Syntactic Structures in Cross-Lingual NLPabstractThe transfer or share of knowledge between languages is a popular solution to resource scarcity in NLP.However, the effectiveness of cross-lingual transfer can be challenged by variation in syntactic structures.Frameworks such as Universal Dependencies (UD) are designed to be cross-lingually consistent, but even in carefully designed resources trees representing equivalent sentences may not always overlap.In this paper, we measure cross-lingual syntactic variation, or anisomorphism, in the UD treebank collection, considering both morphological and structural properties.We show that reducing the level of anisomorphism yields consistent gains in cross-lingual transfer tasks.We introduce a source language selection procedure that facilitates effective cross-lingual parser transfer, and propose a typologically driven method for syntactic tree processing which reduces anisomorphism.Our results show the effectiveness of this method for both machine translation and cross-lingual sentence similarity, demonstrating the importance of syntactic structure compatibility for boosting cross-lingual transfer in NLP. Edoardo Maria Ponti, Roi Reichart, Anna Korhonen, Ivan Vulic |
ACL (1) | 4 |
| 2018 | Bridging Languages through Images with Deep Partial Canonical Correlation AnalysisabstractWe present a deep neural network that leverages images to improve bilingual text embeddings.Relying on bilingual image tags and descriptions, our approach conditions text embedding induction on the shared visual information for both languages, producing highly correlated bilingual embeddings.In particular, we propose a novel model based on Partial Canonical Correlation Analysis (PCCA).While the original PCCA finds linear projections of two views in order to maximize their canonical correlation conditioned on a shared third variable, we introduce a non-linear Deep PCCA (DPCCA) model, and develop a new stochastic iterative algorithm for its optimization.We evaluate PCCA and DPCCA on multilingual word similarity and cross-lingual image description retrieval.Our models outperform a large variety of previous methods, despite not having access to any visual signal during test time inference.1 Guy Rotman, Ivan Vulic, Roi Reichart |
ACL (1) | 2 |
| 2018 | On the Limitations of Unsupervised Bilingual Dictionary InductionabstractUnsupervised machine translation-i.e., not assuming any cross-lingual supervision signal, whether a dictionary, translations, or comparable corpora-seems impossible, but nevertheless, Lample et al. (2018a) recently proposed a fully unsupervised machine translation (MT) model.The model relies heavily on an adversarial, unsupervised alignment of word embedding spaces for bilingual dictionary induction (Conneau et al., 2018), which we examine here.Our results identify the limitations of current unsupervised MT: unsupervised bilingual dictionary induction performs much worse on morphologically rich languages that are not dependent marking, when monolingual corpora from different domains or different embedding algorithms are used.We show that a simple trick, exploiting a weak supervision signal from identical words, enables more robust induction, and establish a near-perfect correlation between unsupervised bilingual dictionary induction performance and a previously unexplored graph similarity metric. Anders Søgaard, Sebastian Ruder, Ivan Vulic |
ACL (1) | 3 |
| 2018 | Explicit Retrofitting of Distributional Word VectorsabstractSemantic specialization of distributional word vectors, referred to as retrofitting, is a process of fine-tuning word vectors using external lexical knowledge in order to better embed some semantic relation.Existing retrofitting models integrate linguistic constraints directly into learning objectives and, consequently, specialize only the vectors of words from the constraints.In this work, in contrast, we transform external lexico-semantic relations into training examples which we use to learn an explicit retrofitting model (ER).The ER model allows us to learn a global specialization function and specialize the vectors of words unobserved in the training data as well.We report large gains over original distributional vector spaces in (1) intrinsic word similarity evaluation and on (2) two downstream tasks -lexical simplification and dialog state tracking.Finally, we also successfully specialize vector spaces of new languages (i.e., unseen in the training data) by coupling ER with shared multilingual distributional vector spaces. Goran Glavas, Ivan Vulic |
ACL (1) | 2 |
| 2018 | On the Relation between Linguistic Typology and (Limitations of) Multilingual Language ModelingabstractA key challenge in cross-lingual NLP is developing general language-independent architectures that are equally applicable to any language.However, this ambition is largely hampered by the variation in structural and semantic properties, i.e. the typological profiles of the world's languages.In this work, we analyse the implications of this variation on the language modeling (LM) task.We present a largescale study of state-of-the art n-gram based and neural language models on 50 typologically diverse languages covering a wide variety of morphological systems.Operating in the full vocabulary LM setup focused on wordlevel prediction, we demonstrate that a coarse typology of morphological systems is predictive of absolute LM performance.Moreover, fine-grained typological features such as exponence, flexivity, fusion, and inflectional synthesis are borne out to be responsible for the proliferation of low-frequency phenomena which are organically difficult to model by statistical architectures, or for the meaning ambiguity of character n-grams.Our study strongly suggests that these features have to be taken into consideration during the construction of nextlevel language-agnostic LM architectures, capable of handling morphologically complex languages such as Tamil or Korean. Daniela Gerz, Ivan Vulic, Edoardo Maria Ponti, Roi Reichart, Anna Korhonen |
EMNLP | 2 |
| 2018 | Adversarial Propagation and Zero-Shot Cross-Lingual Transfer of Word Vector SpecializationabstractSemantic specialization is a process of finetuning pre-trained distributional word vectors using external lexical knowledge (e.g., Word-Net) to accentuate a particular semantic relation in the specialized vector space.While post-processing specialization methods are applicable to arbitrary distributional vectors, they are limited to updating only the vectors of words occurring in external lexicons (i.e., seen words), leaving the vectors of all other words unchanged.We propose a novel approach to specializing the full distributional vocabulary.Our adversarial post-specialization method propagates the external lexical knowledge to the full distributional space.We exploit words seen in the resources as training examples for learning a global specialization function.This function is learned by combining a standard L 2 -distance loss with a adversarial loss: the adversarial component produces more realistic output vectors.We show the effectiveness and robustness of the proposed method across three languages and on three tasks: word similarity, dialog state tracking, and lexical simplification.We report consistent improvements over distributional word vectors and vectors specialized by other state-of-the-art specialization frameworks.Finally, we also propose a cross-lingual transfer method for zero-shot specialization which successfully specializes a full target distributional space without any lexical knowledge in the target language and without any bilingual data. Edoardo Maria Ponti, Ivan Vulic, Goran Glavas, Nikola Mrksic, Anna Korhonen |
EMNLP | 2 |
| 2018 | Acquiring Verb Classes Through Bottom-Up Semantic Verb Clustering
Olga Majewska, Diana McCarthy, Ivan Vulic, Anna Korhonen |
LREC | 3 |
| 2018 | Post-Specialisation: Retrofitting Vectors of Words Unseen in Lexical ResourcesabstractIvan Vulić, Goran Glavaš, Nikola Mrkšić, Anna Korhonen. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Ivan Vulic, Goran Glavas, Nikola Mrksic, Anna Korhonen |
NAACL-HLT | 1 |
| 2018 | Specialising Word Vectors for Lexical EntailmentabstractWe present LEAR (Lexical Entailment Attract-Repel), a novel post-processing method that transforms any input word vector space to emphasise the asymmetric relation of lexical entailment (LE), also known as the IS-A or hyponymy-hypernymy relation.By injecting external linguistic constraints (e.g., WordNet links) into the initial vector space, the LE specialisation procedure brings true hyponymyhypernymy pairs closer together in the transformed Euclidean space.The proposed asymmetric distance measure adjusts the norms of word vectors to reflect the actual WordNetstyle hierarchy of concepts.Simultaneously, a joint objective enforces semantic similarity using the symmetric cosine distance, yielding a vector space specialised for both lexical relations at once.LEAR specialisation achieves state-of-the-art performance in the tasks of hypernymy directionality, hypernymy detection, and graded lexical entailment, demonstrating the effectiveness and robustness of the proposed asymmetric specialisation model. Ivan Vulic, Nikola Mrksic |
NAACL-HLT | 1 |
| 2018 | Unsupervised Cross-Lingual Information Retrieval Using Monolingual Data OnlyabstractWe propose a fully unsupervised framework for ad-hoc cross-lingual information retrieval (CLIR) which requires no bilingual data at all. The framework leverages shared cross-lingual word embedding spaces in which terms, queries, and documents can be represented, irrespective of their actual language. The shared embedding spaces are induced solely on the basis of monolingual corpora in two languages through an iterative process based on adversarial neural networks. Our experiments on the standard CLEF CLIR collections for three language pairs of varying degrees of language similarity (English-Dutch/Italian/Finnish) demonstrate the usefulness of the proposed fully unsupervised approach. Our CLIR models with unsupervised cross-lingual embeddings outperform baselines that utilize cross-lingual embeddings induced relying on word-level and document-level alignments. We then demonstrate that further improvements can be achieved by unsupervised ensemble CLIR models. We believe that the proposed framework is the first step towards development of effective CLIR models for language pairs and domains where parallel data are scarce or non-existent. Robert Litschko, Goran Glavas, Simone Paolo Ponzetto, Ivan Vulic |
SIGIR | 4 |
| 2018 | Bio-SimVerb and Bio-SimLex: wide-coverage evaluation sets of word similarity in biomedicineabstractBACKGROUND: Word representations support a variety of Natural Language Processing (NLP) tasks. The quality of these representations is typically assessed by comparing the distances in the induced vector spaces against human similarity judgements. Whereas comprehensive evaluation resources have recently been developed for the general domain, similar resources for biomedicine currently suffer from the lack of coverage, both in terms of word types included and with respect to the semantic distinctions. Notably, verbs have been excluded, although they are essential for the interpretation of biomedical language. Further, current resources do not discern between semantic similarity and semantic relatedness, although this has been proven as an important predictor of the usefulness of word representations and their performance in downstream applications. RESULTS: We present two novel comprehensive resources targeting the evaluation of word representations in biomedicine. These resources, Bio-SimVerb and Bio-SimLex, address the previously mentioned problems, and can be used for evaluations of verb and noun representations respectively. In our experiments, we have computed the Pearson's correlation between performances on intrinsic and extrinsic tasks using twelve popular state-of-the-art representation models (e.g. word2vec models). The intrinsic-extrinsic correlations using our datasets are notably higher than with previous intrinsic evaluation benchmarks such as UMNSRS and MayoSRS. In addition, when evaluating representation models for their abilities to capture verb and noun semantics individually, we show a considerable variation between performances across all models. CONCLUSION: Bio-SimVerb and Bio-SimLex enable intrinsic evaluation of word representations. This evaluation can serve as a predictor of performance on various downstream tasks in the biomedical domain. The results on Bio-SimVerb and Bio-SimLex using standard word representation models highlight the importance of developing dedicated evaluation resources for NLP in biomedicine for particular word classes (e.g. verbs). These are needed to identify the most accurate methods for learning class-specific representations. Bio-SimVerb and Bio-SimLex are publicly available. Billy Chiu, Sampo Pyysalo, Ivan Vulic, Anna Korhonen |
BMC Bioinform. | 3 |
| 2018 | A deep learning approach to bilingual lexicon induction in the biomedical domainabstractBACKGROUND: Bilingual lexicon induction (BLI) is an important task in the biomedical domain as translation resources are usually available for general language usage, but are often lacking in domain-specific settings. In this article we consider BLI as a classification problem and train a neural network composed of a combination of recurrent long short-term memory and deep feed-forward networks in order to obtain word-level and character-level representations. RESULTS: The results show that the word-level and character-level representations each improve state-of-the-art results for BLI and biomedical translation mining. The best results are obtained by exploiting the synergy between these word-level and character-level representations in the classification model. We evaluate the models both quantitatively and qualitatively. CONCLUSIONS: Translation of domain-specific biomedical terminology benefits from the character-level representations compared to relying solely on word-level representations. It is beneficial to take a deep learning approach and learn character-level representations rather than relying on handcrafted representations that are typically used. Our combined model captures the semantics at the word level while also taking into account that specialized terminology often originates from a common root form (e.g., from Greek or Latin). Geert Heyman, Ivan Vulic, Marie-Francine Moens |
BMC Bioinform. | 2 |
| 2018 | Language Modeling for Morphologically Rich Languages: Character-Aware Modeling for Word-Level PredictionabstractNeural architectures are prominent in the construction of language models (LMs). However, word-level prediction is typically agnostic of subword-level information (characters and character sequences) and operates over a closed vocabulary, consisting of a limited word set. Indeed, while subword-aware models boost performance across a variety of NLP tasks, previous work did not evaluate the ability of these models to assist next-word prediction in language modeling tasks. Such subword-level informed models should be particularly effective for morphologically-rich languages (MRLs) that exhibit high type-to-token ratios. In this work, we present a large-scale LM study on 50 typologically diverse languages covering a wide variety of morphological systems, and offer new LM benchmarks to the community, while considering subword-level information. The main technical contribution of our work is a novel method for injecting subword-level information into semantic word vectors, integrated into the neural language modeling training, to facilitate word-level prediction. We conduct experiments in the LM setting where the number of infrequent words is large, and demonstrate strong perplexity gains across our 50 languages, especially for morphologically-rich languages. Our code and data sets are publicly available. Daniela Gerz, Ivan Vulic, Edoardo Maria Ponti, Jason Naradowsky, Roi Reichart, Anna Korhonen |
Trans. Assoc. Comput. Linguistics | 2 |
| 2017 | Morph-fitting: Fine-Tuning Word Vector Spaces with Simple Language-Specific RulesabstractIvan Vulić, Nikola Mrkšić, Roi Reichart, Diarmuid Ó Séaghdha, Steve Young, Anna Korhonen. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017. Ivan Vulic, Nikola Mrksic, Roi Reichart, Diarmuid Ó Séaghdha, Steve J. Young, Anna Korhonen |
ACL (1) | 1 |
| 2017 | Automatic Selection of Context Configurations for Improved Class-Specific Word RepresentationsabstractThis paper is concerned with identifying contexts useful for training word representation models for different word classes such as adjectives (A), verbs (V), and nouns (N).We introduce a simple yet effective framework for an automatic selection of class-specific context configurations.We construct a context configuration space based on universal dependency relations between words, and efficiently search this space with an adapted beam search algorithm.In word similarity tasks for each word class, we show that our framework is both effective and efficient.Particularly, it improves the Spearman's ρ correlation with human scores on SimLex-999 over the best previously proposed class-specific contexts by 6 (A), 6 (V) and 5 (N) ρ points.With our selected context configurations, we train on only 14% (A), 26.2% (V), and 33.6% (N) of all dependency-based contexts, resulting in a reduced training time.Our results generalise: we show that the configurations our algorithm learns for one English training setup outperform previously proposed context types in another training setup for English.Moreover, basing the configuration space on universal dependencies, it is possible to transfer the learned configurations to German and Italian.We also demonstrate improved per-class results over other context types in these two languages. Ivan Vulic, Roy Schwartz 0001, Ari Rappoport, Roi Reichart, Anna Korhonen |
CoNLL | 1 |
| 2017 | Evaluation by Association: A Systematic Study of Quantitative Word Association EvaluationabstractRecent work on evaluating representation learning architectures in NLP has established a need for evaluation protocols based on subconscious cognitive measures rather than manually tailored intrinsic similarity and relatedness tasks.In this work, we propose a novel evaluation framework that enables large-scale evaluation of such architectures in the free word association (WA) task, which is firmly grounded in cognitive theories of human semantic representation.This evaluation is facilitated by the existence of large manually constructed repositories of word association data.In this paper, we (1) present a detailed analysis of the new quantitative WA evaluation protocol, (2) suggest new evaluation metrics for the WA task inspired by its direct analogy with information retrieval problems, (3) evaluate various state-of-the-art representation models on this task, and (4) discuss the relationship between WA and prior evaluations of semantic representation with well-known similarity and relatedness evaluation sets.We have made the WA evaluation toolkit publicly available. Ivan Vulic, Douwe Kiela, Anna Korhonen |
EACL (1) | 1 |
| 2017 | Bilingual Lexicon Induction by Learning to Combine Word-Level and Character-Level RepresentationsabstractWe study the problem of bilingual lexicon induction (BLI) in a setting where some translation resources are available, but unknown translations are sought for certain, possibly domain-specific terminology.We frame BLI as a classification problem for which we design a neural network based classification architecture composed of recurrent long short-term memory and deep feed forward networks.The results show that word-and character-level representations each improve state-of-the-art results for BLI, and the best results are obtained by exploiting the synergy between these wordand character-level representations in the classification model. Geert Heyman, Ivan Vulic, Marie-Francine Moens |
EACL (1) | 2 |
| 2017 | Cross-Lingual Induction and Transfer of Verb Classes Based on Word Vector Space SpecialisationabstractExisting approaches to automatic VerbNetstyle verb classification are heavily dependent on feature engineering and therefore limited to languages with mature NLP pipelines.In this work, we propose a novel cross-lingual transfer method for inducing VerbNets for multiple languages.To the best of our knowledge, this is the first study which demonstrates how the architectures for learning word embeddings can be applied to this challenging syntactic-semantic task.Our method uses cross-lingual translation pairs to tie each of the six target languages into a bilingual vector space with English, jointly specialising the representations to encode the relational information from English VerbNet.A standard clustering algorithm is then run on top of the VerbNet-specialised representations, using vector dimensions as features for learning verb classes.Our results show that the proposed cross-lingual transfer approach sets new state-of-the-art verb classification performance across all six target languages explored in this work. Ivan Vulic, Nikola Mrksic, Anna Korhonen |
EMNLP | 1 |
| 2017 | HyperLex: A Large-Scale Evaluation of Graded Lexical EntailmentabstractWe introduce HyperLex—a data set and evaluation resource that quantifies the extent of the semantic category membership, that is, type-of relation, also known as hyponymy–hypernymy or lexical entailment (LE) relation between 2,616 concept pairs. Cognitive psychology research has established that typicality and category/class membership are computed in human semantic memory as a gradual rather than binary relation. Nevertheless, most NLP research and existing large-scale inventories of concept category membership (WordNet, DBPedia, etc.) treat category membership and LE as binary. To address this, we asked hundreds of native English speakers to indicate typicality and strength of category membership between a diverse range of concept pairs on a crowdsourcing platform. Our results confirm that category membership and LE are indeed more gradual than binary. We then compare these human judgments with the predictions of automatic systems, which reveals a huge gap between human performance and state-of-the-art LE, distributional and representation learning models, and substantial differences between the models themselves. We discuss a pathway for improving semantic models to overcome this discrepancy, and indicate future application areas for improved graded LE systems. Ivan Vulic, Daniela Gerz, Douwe Kiela, Felix Hill, Anna Korhonen |
Comput. Linguistics | 1 |
| 2017 | Semantic Specialization of Distributional Word Vector Spaces using Monolingual and Cross-Lingual ConstraintsabstractWe present Attract-Repel, an algorithm for improving the semantic quality of word vectors by injecting constraints extracted from lexical resources. Attract-Repel facilitates the use of constraints from mono- and cross-lingual resources, yielding semantically specialized cross-lingual vector spaces. Our evaluation shows that the method can make use of existing cross-lingual lexicons to construct high-quality vector spaces for a plethora of different languages, facilitating semantic transfer from high- to lower-resource ones. The effectiveness of our approach is demonstrated with state-of-the-art results on semantic similarity datasets in six languages. We next show that Attract-Repel-specialized vectors boost performance in the downstream task of dialogue state tracking (DST) across multiple languages. Finally, we show that cross-lingual vector spaces produced by our algorithm facilitate the training of multilingual DST models, which brings further performance improvements. Nikola Mrksic, Ivan Vulic, Diarmuid Ó Séaghdha, Ira Leviant, Roi Reichart, Milica Gasic, Anna Korhonen, Steve J. Young |
Trans. Assoc. Comput. Linguistics | 2 |
| 2016 | On the Role of Seed Lexicons in Learning Bilingual Word EmbeddingsabstractA shared bilingual word embedding space (SBWES) is an indispensable resource in a variety of cross-language NLP and IR tasks.A common approach to the SB-WES induction is to learn a mapping function between monolingual semantic spaces, where the mapping critically relies on a seed word lexicon used in the learning process.In this work, we analyze the importance and properties of seed lexicons for the SBWES induction across different dimensions (i.e., lexicon source, lexicon size, translation method, translation pair reliability).On the basis of our analysis, we propose a simple but effective hybrid bilingual word embedding (BWE) model.This model (HYBWE) learns the mapping between two monolingual embedding spaces using only highly reliable symmetric translation pairs from a seed document-level embedding space.We perform bilingual lexicon learning (BLL) with 3 language pairs and show that by carefully selecting reliable translation pairs our new HYBWE model outperforms benchmarking BWE learning models, all of which use more expensive bilingual signals.Effectively, we demonstrate that a SBWES may be induced by leveraging only a very weak bilingual signal (document alignments) along with monolingual data. Ivan Vulic, Anna Korhonen |
ACL (1) | 1 |
| 2016 | Survey on the Use of Typological Information in Natural Language ProcessingabstractIn recent years linguistic typologies, which classify the world’s languages according to their functional and structural properties, have been widely used to support multilingual NLP. While the growing importance of typologies in supporting multilingual tasks has been recognised, no systematic survey of existing typological resources and their use in NLP has been published. This paper provides such a survey as well as discussion which we hope will both inform and inspire future work in the area. Helen O'Horan, Yevgeni Berzak, Ivan Vulic, Roi Reichart, Anna Korhonen |
COLING | 3 |
| 2016 | SimVerb-3500: A Large-Scale Evaluation Set of Verb SimilarityabstractVerbs play a critical role in the meaning of sentences, but these ubiquitous words have received little attention in recent distributional semantics research. We introduce SimVerb-3500, an evaluation resource that provides human ratings for the similarity of 3,500 verb pairs. SimVerb-3500 covers all normed verb types from the USF free-association database, providing at least three examples for every VerbNet class. This broad coverage facilitates detailed analyses of how syntactic and semantic phenomena together influence human understanding of verb meaning. Further, with significantly larger development and test sets than existing benchmarks, SimVerb-3500 enables more robust evaluation of representation learning architectures and promotes the development of methods tailored to verbs. We hope that SimVerb-3500 will enable a richer understanding of the diversity and complexity of verb semantics and guide the development of systems that can effectively represent and interpret this meaning. Daniela Gerz, Ivan Vulic, Felix Hill, Roi Reichart, Anna Korhonen |
EMNLP | 2 |
| 2016 | C-BiLDA extracting cross-lingual topics from non-parallel texts by distinguishing shared from unshared content
Geert Heyman, Ivan Vulic, Marie-Francine Moens |
Data Min. Knowl. Discov. | 2 |
| 2016 | Latent Dirichlet allocation for linking user-generated content and e-commerce data
Susana Zoghbi, Ivan Vulic, Marie-Francine Moens |
Inf. Sci. | 2 |
| 2016 | Bilingual Distributed Word Representations from Document-Aligned Comparable DataabstractWe propose a new model for learning bilingual word representations from non-parallel document-aligned data. Following the recent advances in word representation learning, our model learns dense real-valued word vectors, that is, bilingual word embeddings (BWEs). Unlike prior work on inducing BWEs which heavily relied on parallel sentence-aligned corpora and/or readily available translation resources such as dictionaries, the article reveals that BWEs may be learned solely on the basis of document-aligned comparable data without any additional lexical resources nor syntactic information. We present a comparison of our approach with previous state-of-the-art models for learning bilingual word representations from comparable data that rely on the framework of multilingual probabilistic topic modeling (MuPTM), as well as with distributional local context-counting models. We demonstrate the utility of the induced BWEs in two semantic tasks: (1) bilingual lexicon extraction, (2) suggesting word translations in context for polysemous words. Our simple yet effective BWE-based models significantly outperform the MuPTM-based and context-counting representation models from comparable data as well as prior BWE-based models, and acquire the best reported results on both tasks for all three tested language pairs. Ivan Vulic, Marie-Francine Moens |
J. Artif. Intell. Res. | 1 |
| 2015 | Semantic Role Labeling of Speech Transcripts
Niraj Shrestha, Ivan Vulic, Marie-Francine Moens |
CICLing (2) | 2 |
| 2015 | Visual Bilingual Lexicon Induction with Transferred ConvNet FeaturesabstractThis paper is concerned with the task of bilingual lexicon induction using imagebased features.By applying features from a convolutional neural network (CNN), we obtain state-of-the-art performance on a standard dataset, obtaining a 79% relative improvement over previous work which uses bags of visual words based on SIFT features.The CNN image-based approach is also compared with state-of-the-art linguistic approaches to bilingual lexicon induction, even outperforming these for one of three language pairs on another standard dataset.Furthermore, we shed new light on the type of visual similarity metric to use for genuine similarity versus relatedness tasks, and experiment with using multiple layers from the same network in an attempt to improve performance. Douwe Kiela, Ivan Vulic, Stephen Clark |
EMNLP | 2 |
| 2015 | Monolingual and Cross-Lingual Information Retrieval Models Based on (Bilingual) Word EmbeddingsabstractWe propose a new unified framework for monolingual (MoIR) and cross-lingual information retrieval (CLIR) which relies on the induction of dense real-valued word vectors known as word embeddings (WE) from comparable data. To this end, we make several important contributions: (1) We present a novel word representation learning model called Bilingual Word Embeddings Skip-Gram (BWESG) which is the first model able to learn bilingual word embeddings solely on the basis of document-aligned comparable data; (2) We demonstrate a simple yet effective approach to building document embeddings from single word embeddings by utilizing models from compositional distributional semantics. BWESG induces a shared cross-lingual embedding vector space in which both words, queries, and documents may be presented as dense real-valued vectors; (3) We build novel ad-hoc MoIR and CLIR models which rely on the induced word and document embeddings and the shared cross-lingual embedding space; (4) Experiments for English and Dutch MoIR, as well as for English-to-Dutch and Dutch-to-English CLIR using benchmarking CLEF 2001-2003 collections and queries demonstrate the utility of our WE-based MoIR and CLIR models. The best results on the CLEF collections are obtained by the combination of the WE-based approach and a unigram language model. We also report on significant improvements in ad-hoc IR tasks of our WE-based framework over the state-of-the-art framework for learning text representations from comparable data based on latent Dirichlet allocation (LDA). Ivan Vulic, Marie-Francine Moens |
SIGIR | 1 |
| 2015 | Probabilistic topic modeling in multilingual settings: An overview of its methodology and applications
Ivan Vulic, Wim De Smet, Jie Tang 0001, Marie-Francine Moens |
Inf. Process. Manag. | 1 |
| 2014 | Probabilistic Models of Cross-Lingual Semantic Similarity in Context Based on Latent Cross-Lingual Concepts Induced from Comparable DataabstractWe propose the first probabilistic approach to modeling cross-lingual semantic sim-ilarity (CLSS) in context which requires only comparable data. The approach re-lies on an idea of projecting words and sets of words into a shared latent semantic space spanned by language-pair indepen-dent latent semantic concepts (e.g., cross-lingual topics obtained by a multilingual topic model). These latent cross-lingual concepts are induced from a comparable corpus without any additional lexical re-sources. Word meaning is represented as a probability distribution over the latent concepts, and a change in meaning is rep-resented as a change in the distribution over these latent concepts. We present new models that modulate the isolated out-of-context word representations with contex-tual knowledge. Results on the task of suggesting word translations in context for 3 language pairs reveal the utility of the proposed contextualized models of cross-lingual semantic similarity. 1 Ivan Vulic, Marie-Francine Moens |
EMNLP | 1 |
| 2014 | TermWise: A CAT-tool with Context-Sensitive Terminological Support
Kris Heylen, Stephen Bond, Dirk De Hertog, Ivan Vulic, Hendrik J. Kockaert |
LREC | 4 |
| 2014 | Learning to bridge colloquial and formal language applied to linking and search of E-Commerce dataabstractWe study the problem of linking information between different idiomatic usages of the same language, for example, colloquial and formal language. We propose a novel probabilistic topic model called multi-idiomatic LDA (MiLDA). Its modeling principles follow the intuition that certain words are shared between two idioms of the same language, while other words are non-shared, that is, idiom-specific. We demonstrate the ability of our model to learn relations between cross-idiomatic topics in a dataset containing product descriptions and reviews. We intrinsically evaluate our model by the perplexity measure. Following that, as an extrinsic evaluation, we present the utility of the new MiLDA topic model in a recently proposed IR task of linking Pinterest pins (given in colloquial English on the users' side) to online webshops (given in formal English on the retailers' side). We show that our multi-idiomatic model outperforms the standard monolingual LDA model and the pure bilingual LDA model both in terms of perplexity and MAP scores in the IR task. Ivan Vulic, Susana Zoghbi, Marie-Francine Moens |
SIGIR | 1 |
| 2014 | Multilingual probabilistic topic modeling and its applications in web mining and searchabstractMultilingual topic models are a fairly novel group of unsupervised, language-independent and generative machine learning models. This tutorial covers all key aspects of their probabilistic framework and demonstrates how to easily integrate these models into frameworks for cross-lingual and multilingual Web mining and search. Marie-Francine Moens, Ivan Vulic |
WSDM | 2 |
| 2013 | Monolingual and Cross-Lingual Probabilistic Topic Models and Their Applications in Information Retrieval
Marie-Francine Moens, Ivan Vulic |
ECIR | 2 |
| 2013 | A Unified Framework for Monolingual and Cross-Lingual Relevance Modeling Based on Probabilistic Topic Models
Ivan Vulic, Marie-Francine Moens |
ECIR | 1 |
| 2013 | A Study on Bootstrapping Bilingual Vector Spaces from Non-Parallel Data (and Nothing Else)abstractWe present a new language pair agnostic approach to inducing bilingual vector spaces from non-parallel data without any other resource in a bootstrapping fashion.The paper systematically introduces and describes all key elements of the bootstrapping procedure:(1) starting point or seed lexicon, (2) the confidence estimation and selection of new dimensions of the space, and (3) convergence.We test the quality of the induced bilingual vector spaces, and analyze the influence of the different components of the bootstrapping approach in the task of bilingual lexicon extraction (BLE) for two language pairs.Results reveal that, contrary to conclusions from prior work, the seeding of the bootstrapping process has a heavy impact on the quality of the learned lexicons.We also show that our approach outperforms the best performing fully corpus-based BLE methods on these test sets. Ivan Vulic, Marie-Francine Moens |
EMNLP | 1 |
| 2013 | Cross-Lingual Semantic Similarity of Words as the Similarity of Their Semantic Word Responses
Ivan Vulic, Marie-Francine Moens |
HLT-NAACL | 1 |
| 2013 | Cross-language information retrieval models based on latent topic models trained with document-aligned comparable corpora
Ivan Vulic, Wim De Smet, Marie-Francine Moens |
Inf. Retr. | 1 |
| 2012 | Sub-corpora Sampling with an Application to Bilingual Lexicon Extraction
Ivan Vulic, Marie-Francine Moens |
COLING | 1 |
| 2012 | Skip N-grams and Ranking Functions for Predicting Script Events
Bram Jans, Steven Bethard, Ivan Vulic, Marie-Francine Moens |
EACL | 3 |
| 2012 | Detecting Highly Confident Word Translations from Comparable Corpora without Any Prior Knowledge
Ivan Vulic, Marie-Francine Moens |
EACL | 1 |