VLDB 2026 Research / reviewers in the wild / expert
Tommaso Pasini
dblp:148/4516
· DBLP profile ↗
26ranked-venue papers
8as first author
10since 2021 · last 2022
0000-0002-0660-8389ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 8 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Visual Definition Modeling: Challenging Vision & Language Models to Define Words and Objects
Bianca Scarlini, Tommaso Pasini, Roberto Navigli |
AAAI | 2 |
| 2022 | FairLex: A Multilingual Benchmark for Evaluating Fairness in Legal Text ProcessingabstractIlias Chalkidis, Tommaso Pasini, Sheng Zhang, Letizia Tomada, Sebastian Schwemer, Anders Søgaard. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Ilias Chalkidis, Tommaso Pasini, Sheng Zhang 0022, Letizia Tomada, Sebastian Felix Schwemer, Anders Søgaard |
ACL (1) | 2 |
| 2022 | Reducing Disambiguation Biases in NMT by Leveraging Explicit Word Sense InformationabstractNiccolò Campolungo, Tommaso Pasini, Denis Emelin, Roberto Navigli. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Niccolò Campolungo, Tommaso Pasini, Denis Emelin, Roberto Navigli |
NAACL-HLT | 2 |
| 2021 | XL-WSD: An Extra-Large and Cross-Lingual Evaluation Framework for Word Sense DisambiguationabstractTransformer-based architectures brought a breeze of change to Word Sense Disambiguation (WSD), improving models' performances by a large margin. The fast development of new approaches has been further encouraged by a well-framed evaluation suite for English, which has allowed their performances to be kept track of and compared fairly. However, other languages have remained largely unexplored, as testing data are available for a few languages only and the evaluation setting is rather matted. In this paper, we untangle this situation by proposing XL-WSD, a cross-lingual evaluation benchmark for the WSD task featuring sense-annotated development and test sets in 18 languages from six different linguistic families, together with language-specific silver training data. We leverage XL-WSD datasets to conduct an extensive evaluation of neural and knowledge-based approaches, including the most recent multilingual language models. Results show that the zero-shot knowledge transfer across languages is a promising research direction within the WSD field, especially when considering low-resourced languages where large pre-trained multilingual models still perform poorly. We make the evaluation suite and the code for performing the experiments available at https://sapienzanlp.github.io/xl-wsd/. Tommaso Pasini, Alessandro Raganato, Roberto Navigli |
AAAI | 1 |
| 2021 | IR like a SIR: Sense-enhanced Information Retrieval for Multiple LanguagesabstractWith the advent of contextualized embeddings, attention towards neural ranking approaches for Information Retrieval increased considerably.However, two aspects have remained largely neglected: i) queries usually consist of few keywords only, which increases ambiguity and makes their contextualization harder, and ii) performing neural ranking on non-English documents is still cumbersome due to shortage of labeled datasets.In this paper we present SIR (Sense-enhanced Information Retrieval) to mitigate both problems by leveraging word sense information.At the core of our approach lies a novel multilingual query expansion mechanism based on Word Sense Disambiguation that provides sense definitions as additional semantic information for the query.Importantly, we use senses as a bridge across languages, thus allowing our model to perform considerably better than its supervised and unsupervised alternatives across French, German, Italian and Spanish languages on several CLEF benchmarks, while Rexhina Blloshmi, Tommaso Pasini, Niccolò Campolungo, Somnath Banerjee 0001, Roberto Navigli, Gabriella Pasi |
EMNLP (1) | 2 |
| 2021 | Exemplification Modeling: Can You Give Me an Example, Please?abstractRecently, generative approaches have been used effectively to provide definitions of words in their context. However, the opposite, i.e., generating a usage example given one or more words along with their definitions, has not yet been investigated. In this work, we introduce the novel task of Exemplification Modeling (ExMod), along with a sequence-to-sequence architecture and a training procedure for it. Starting from a set of (word, definition) pairs, our approach is capable of automatically generating high-quality sentences which express the requested semantics. As a result, we can drive the creation of sense-tagged data which cover the full range of meanings in any inventory of interest, and their interactions within sentences. Human annotators agree that the sentences generated are as fluent and semantically-coherent with the input definitions as the sentences in manually-annotated corpora. Indeed, when employed as training data for Word Sense Disambiguation, our examples enable the current state of the art to be outperformed, and higher results to be achieved than when using gold-standard datasets only. We release the pretrained model, the dataset and the software at https://github.com/SapienzaNLP/exmod. Edoardo Barba, Luigi Procopio, Caterina Lacerra, Tommaso Pasini, Roberto Navigli |
IJCAI | 4 |
| 2021 | Recent Trends in Word Sense Disambiguation: A SurveyabstractWord Sense Disambiguation (WSD) aims at making explicit the semantics of a word in context by identifying the most suitable meaning from a predefined sense inventory. Recent breakthroughs in representation learning have fueled intensive WSD research, resulting in considerable performance improvements, breaching the 80% glass ceiling set by the inter-annotator agreement. In this survey, we provide an extensive overview of current advances in WSD, describing the state of the art in terms of i) resources for the task, i.e., sense inventories and reference datasets for training and testing, as well as ii) automatic disambiguation approaches, detailing their peculiarities, strengths and weaknesses. Finally, we highlight the current limitations of the task itself, but also point out recent trends that could help expand the scope and applicability of WSD, setting up new promising directions for the future. Michele Bevilacqua, Tommaso Pasini, Alessandro Raganato, Roberto Navigli |
IJCAI | 2 |
| 2021 | ALaSca: an Automated approach for Large-Scale Lexical SubstitutionabstractThe lexical substitution task aims at finding suitable replacements for words in context. It has proved to be useful in several areas, such as word sense induction and text simplification, as well as in more practical applications such as writing-assistant tools. However, the paucity of annotated data has forced researchers to apply mainly unsupervised approaches, limiting the applicability of large pre-trained models and thus hampering the potential benefits of supervised approaches to the task. In this paper, we mitigate this issue by proposing ALaSca, a novel approach to automatically creating large-scale datasets for English lexical substitution. ALaSca allows examples to be produced for potentially any word in a language vocabulary and to cover most of the meanings it lists. Thanks to this, we can unleash the full potential of neural architectures and finetune them on the lexical substitution task. Indeed, when using our data, a transformer-based model performs substantially better than when using manually annotated data only. We release ALaSca at https://sapienzanlp.github.io/alasca/. Caterina Lacerra, Tommaso Pasini, Rocco Tripodi, Roberto Navigli |
IJCAI | 2 |
| 2021 | ESC: Redesigning WSD with Extractive Sense ComprehensionabstractWord Sense Disambiguation (WSD) is a historical NLP task aimed at linking words in contexts to discrete sense inventories and it is usually cast as a multi-label classification task.Recently, several neural approaches have employed sense definitions to better represent word meanings.Yet, these approaches do not observe the input sentence and the sense definition candidates all at once, thus potentially reducing the model performance and generalization power.We cope with this issue by reframing WSD as a span extraction problem -which we called Extractive Sense Comprehension (ESC) -and propose ESCHER, a transformer-based neural architecture for this new formulation.By means of an extensive array of experiments, we show that ESC unleashes the full potential of our model, leading it to outdo all of its competitors and to set a new state of the art on the English WSD task.In the few-shot scenario, ESCHER proves to exploit training data efficiently, attaining the same performance as its closest competitor while relying on almost three times fewer annotations.Furthermore, ESCHER can nimbly combine data annotated with senses from different lexical resources, achieving performances that were previously out of everyone's reach.The model along with data is available at https://github.com/ SapienzaNLP/esc. Edoardo Barba, Tommaso Pasini, Roberto Navigli |
NAACL-HLT | 2 |
| 2021 | Wikipedia Entities as Rendezvous across Languages: Grounding Multilingual Language Models by Predicting Wikipedia HyperlinksabstractIacer Calixto, Alessandro Raganato, Tommaso Pasini. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Iacer Calixto, Alessandro Raganato, Tommaso Pasini |
NAACL-HLT | 3 |
| 2020 | CSI: A Coarse Sense Inventory for 85% Word Sense DisambiguationabstractWord Sense Disambiguation (WSD) is the task of associating a word in context with one of its meanings. While many works in the past have focused on raising the state of the art, none has even come close to achieving an F-score in the 80% ballpark when using WordNet as its sense inventory. We contend that one of the main reasons for this failure is the excessively fine granularity of this inventory, resulting in senses that are hard to differentiate between, even for an experienced human annotator. In this paper we cope with this long-standing problem by introducing Coarse Sense Inventory (CSI), obtained by linking WordNet concepts to a new set of 45 labels. The results show that the coarse granularity of CSI leads a WSD model to achieve 85.9% F1, while maintaining a high expressive power. Our set of labels also exhibits ease of use in tagging and a descriptiveness that other coarse inventories lack, as demonstrated in two annotation tasks which we performed. Moreover, a few-shot evaluation proves that the class-based nature of CSI allows the model to generalise over unseen or under-represented words. Caterina Lacerra, Michele Bevilacqua, Tommaso Pasini, Roberto Navigli |
AAAI | 3 |
| 2020 | SensEmBERT: Context-Enhanced Sense Embeddings for Multilingual Word Sense DisambiguationabstractContextual representations of words derived by neural language models have proven to effectively encode the subtle distinctions that might occur between different meanings of the same word. However, these representations are not tied to a semantic network, hence they leave the word meanings implicit and thereby neglect the information that can be derived from the knowledge base itself. In this paper, we propose SensEmBERT, a knowledge-based approach that brings together the expressive power of language modelling and the vast amount of knowledge contained in a semantic network to produce high-quality latent semantic representations of word meanings in multiple languages. Our vectors lie in a space comparable with that of contextualized word embeddings, thus allowing a word occurrence to be easily linked to its meaning by applying a simple nearest neighbour approach.We show that, whilst not relying on manual semantic annotations, SensEmBERT is able to either achieve or surpass state-of-the-art results attained by most of the supervised neural approaches on the English Word Sense Disambiguation task. When scaling to other languages, our representations prove to be equally effective as their English counterpart and outperform the existing state of the art on all the Word Sense Disambiguation multilingual datasets. The embeddings are released in five different languages at http://sensembert.org. Bianca Scarlini, Tommaso Pasini, Roberto Navigli |
AAAI | 2 |
| 2020 | CluBERT: A Cluster-Based Approach for Learning Sense Distributions in Multiple LanguagesabstractKnowing the Most Frequent Sense (MFS) of a word has been proved to help Word Sense Disambiguation (WSD) models significantly.However, the scarcity of sense-annotated data makes it difficult to induce a reliable and highcoverage distribution of the meanings in a language vocabulary.To address this issue, in this paper we present CluBERT, an automatic and multilingual approach for inducing the distributions of word senses from a corpus of raw sentences.Our experiments show that Clu-BERT learns distributions over English senses that are of higher quality than those extracted by alternative approaches.When used to induce the MFS of a lemma, CluBERT attains state-of-the-art results on the English Word Sense Disambiguation tasks and helps to improve the disambiguation performance of two off-the-shelf WSD models.Moreover, our distributions also prove to be effective in other languages, beating all their alternatives for computing the MFS on the multilingual WSD tasks.We release our sense distributions in five different languages at https://github. com/SapienzaNLP/clubert. Tommaso Pasini, Federico Scozzafava, Bianca Scarlini |
ACL | 1 |
| 2020 | XL-WiC: A Multilingual Benchmark for Evaluating Semantic ContextualizationabstractThe ability to correctly model distinct meanings of a word is crucial for the effectiveness of semantic representation techniques.However, most existing evaluation benchmarks for assessing this criterion are tied to sense inventories (usually WordNet), restricting their usage to a small subset of knowledge-based representation techniques.The Word-in-Context dataset (WiC) addresses the dependence on sense inventories by reformulating the standard disambiguation task as a binary classification problem; but, it is limited to the English language.We put forward a large multilingual benchmark, XL-WiC, featuring gold standards in 12 new languages from varied language families and with different degrees of resource availability, opening room for evaluation scenarios such as zero-shot cross-lingual transfer.We perform a series of experiments to determine the reliability of the datasets and to set performance baselines for several recent contextualized multilingual models.Experimental results show that even when no tagged instances are available for a target language, models trained solely on the English data can attain competitive performance in the task of distinguishing different meanings of a word, even for distant languages.XL-WiC is available at https://pilehvar.github.io/xlwic/. Alessandro Raganato, Tommaso Pasini, José Camacho-Collados, Mohammad Taher Pilehvar |
EMNLP (1) | 2 |
| 2020 | With More Contexts Comes Better Performance: Contextualized Sense Embeddings for All-Round Word Sense DisambiguationabstractContextualized word embeddings have been employed effectively across several tasks in Natural Language Processing, as they have proved to carry useful semantic information.However, it is still hard to link them to structured sources of knowledge.In this paper we present ARES (context-AwaRe Embeddings of Senses), a semi-supervised approach to producing sense embeddings for the lexical meanings within a lexical knowledge base that lie in a space that is comparable to that of contextualized word vectors.ARES representations enable a simple 1-Nearest-Neighbour algorithm to outperform state-of-the-art models, not only in the English Word Sense Disambiguation task, but also in the multilingual one, whilst training on sense-annotated data in English only.We further assess the quality of our embeddings in the Word-in-Context task, where, when used as an external source of knowledge, they consistently improve the performance of a neural model, leading it to compete with other more complex architectures.ARES embeddings for all WordNet concepts and the automatically-extracted contexts used for creating the sense representations are freely available at http://sensembert.org/ares. Bianca Scarlini, Tommaso Pasini, Roberto Navigli |
EMNLP (1) | 2 |
| 2020 | MuLaN: Multilingual Label propagatioN for Word Sense DisambiguationabstractThe knowledge acquisition bottleneck strongly affects the creation of multilingual sense-annotated data, hence limiting the power of supervised systems when applied to multilingual Word Sense Disambiguation. In this paper, we propose a semi-supervised approach based upon a novel label propagation scheme, which, by jointly leveraging contextualized word embeddings and the multilingual information enclosed in a knowledge base, projects sense labels from a high-resource language, i.e., English, to lower-resourced ones. Backed by several experiments, we provide empirical evidence that our automatically created datasets are of a higher quality than those generated by other competitors and lead a supervised model to achieve state-of-the-art performances in all multilingual Word Sense Disambiguation tasks. We make our datasets available for research purposes at https://github.com/SapienzaNLP/mulan. Edoardo Barba, Luigi Procopio, Niccolò Campolungo, Tommaso Pasini, Roberto Navigli |
IJCAI | 4 |
| 2020 | The Knowledge Acquisition Bottleneck Problem in Multilingual Word Sense DisambiguationabstractWord Sense Disambiguation (WSD) is the task of identifying the meaning of a word in a given context. It lies at the base of Natural Language Processing as it provides semantic information for words. In the last decade, great strides have been made in this field and much effort has been devoted to mitigate the knowledge acquisition bottleneck problem, i.e., the problem of semantically annotating texts at a large scale and in different languages. This issue is ubiquitous in WSD as it hinders the creation of both multilingual knowledge bases and manually-curated training sets. In this work, we first introduce the reader to the task of WSD through a short historical digression and then take the stock of the advancements to alleviate the knowledge acquisition bottleneck problem. In that, we survey the literature on manual, semi-automatic and automatic approaches to create English and multilingual corpora tagged with sense annotations and present a clear overview over supervised models for WSD. Finally, we provide our view over the future directions that we foresee for the field. Tommaso Pasini |
IJCAI | 1 |
| 2020 | A Short Survey on Sense-Annotated CorporaabstractLarge sense-annotated datasets are increasingly necessary for training deep supervised systems in Word Sense Disambiguation. However, gathering high-quality sense-annotated data for as many instances as possible is a laborious and expensive task. This has led to the proliferation of automatic and semi-automatic methods for overcoming the so-called knowledge-acquisition bottleneck. In this short survey we present an overview of sense-annotated corpora, annotated either manually- or (semi)automatically, that are currently available for different languages and featuring distinct lexical resources as inventory of senses, i.e. WordNet, Wikipedia, BabelNet. Furthermore, we provide the reader with general statistics of each dataset and an analysis of their specific features. Tommaso Pasini, José Camacho-Collados |
LREC | 1 |
| 2020 | Sense-Annotated Corpora for Word Sense Disambiguation in Multiple Languages and DomainsabstractThe knowledge acquisition bottleneck problem dramatically hampers the creation of sense-annotated data for Word Sense Disambiguation (WSD). Sense-annotated data are scarce for English and almost absent for other languages. This limits the range of action of deep-learning approaches, which today are at the base of any NLP task and are hungry for data. We mitigate this issue and encourage further research in multilingual WSD by releasing to the NLP community five large datasets annotated with word-senses in five different languages, namely, English, French, Italian, German and Spanish, and 5 distinct datasets in English, each for a different semantic domain. We show that supervised WSD models trained on our data attain higher performance than when trained on other automatically-created corpora. We release all our data containing more than 15 million annotated instances in 5 different languages at http://trainomatic.org/onesec. Bianca Scarlini, Tommaso Pasini, Roberto Navigli |
LREC | 2 |
| 2020 | Train-O-Matic: Supervised Word Sense Disambiguation with no (manual) effortabstractWord Sense Disambiguation (WSD) is the task of associating the correct meaning with a word in a given context. WSD provides explicit semantic information that is beneficial to several downstream applications, such as question answering, semantic parsing and hypernym extraction. Unfortunately, WSD suffers from the well-known knowledge acquisition bottleneck problem: it is very expensive, in terms of both time and money, to acquire semantic annotations for a large number of sentences. To address this blocking issue we present Train-O-Matic, a knowledge-based and language-independent approach that is able to provide millions of training instances annotated automatically with word meanings. The approach is fully automatic, i.e., no human intervention is required, and the only type of human knowledge used is a task-independent WordNet-like resource. Moreover, as the sense distribution in the training set is pivotal to boosting the performance of WSD systems, we also present two unsupervised and language-independent methods that automatically induce a sense distribution when given a simple corpus of sentences. We show that, when the learned distributions are taken into account for generating the training sets, the performance of supervised methods is further enhanced. Experiments have proven that Train-O-Matic on its own, and also coupled with word sense distribution learning methods, lead a supervised system to achieve state-of-the-art performance consistently across gold standard datasets and languages. Importantly, we show how our sense distribution learning techniques aid Train-O-Matic to scale well over domains, without any extra human effort. To encourage future research, we release all the training sets in 5 different languages and the sense distributions for each domain of SemEval-13 and SemEval-15 at http://trainomatic.org. Tommaso Pasini, Roberto Navigli |
Artif. Intell. | 1 |
| 2019 | Just "OneSeC" for Producing Multilingual Sense-Annotated DataabstractThe well-known problem of knowledge acquisition is one of the biggest issues in Word Sense Disambiguation (WSD), where annotated data are still scarce in English and almost absent in other languages.In this paper we formulate the assumption of One Sense per Wikipedia Category and present OneSeC, a language-independent method for the automatic extraction of hundreds of thousands of sentences in which a target word is tagged with its meaning.Our automaticallygenerated data consistently lead a supervised WSD model to state-of-the-art performance when compared with other automatic and semi-automatic methods.Moreover, our approach outperforms its competitors on multilingual and domain-specific settings, where it beats the existing state of the art on all languages and most domains.All the training data are available for research purposes at http://trainomatic.org/onesec. Bianca Scarlini, Tommaso Pasini, Roberto Navigli |
ACL (1) | 2 |
| 2018 | Two Knowledge-based Methods for High-Performance Sense Distribution LearningabstractKnowing the correct distribution of senses within a corpus can potentially boost the performance of Word Sense Disambiguation (WSD) systems by many points. We present two fully automatic and language-independent methods for computing the distribution of senses given a raw corpus of sentences. Intrinsic and extrinsic evaluations show that our methods outperform the current state of the art in sense distribution learning and the strongest baselines for the most frequent sense in multiple languages and on domain-specific test sets. Our sense distributions are available at http://trainomatic.org. Tommaso Pasini, Roberto Navigli |
AAAI | 1 |
| 2018 | Huge Automatically Extracted Training-Sets for Multilingual Word SenseDisambiguation
Tommaso Pasini, Francesco Elia, Roberto Navigli |
LREC | 1 |
| 2017 | Train-O-Matic: Large-Scale Supervised Word Sense Disambiguation in Multiple Languages without Manual Training DataabstractAnnotating large numbers of sentences with senses is the heaviest requirement of current Word Sense Disambiguation.We present Train-O-Matic, a languageindependent method for generating millions of sense-annotated training instances for virtually all meanings of words in a language's vocabulary.The approach is fully automatic: no human intervention is required and the only type of human knowledge used is a WordNet-like resource.Train-O-Matic achieves consistently state-of-the-art performance across gold standard datasets and languages, while at the same time removing the burden of manual annotation.All the training data is available for research purposes at http://trainomatic.org. Tommaso Pasini, Roberto Navigli |
EMNLP | 1 |
| 2016 | MultiWiBi: The multilingual Wikipedia bitaxonomy project
Tiziano Flati, Daniele Vannella, Tommaso Pasini, Roberto Navigli |
Artif. Intell. | 3 |
| 2014 | Two Is Bigger (and Better) Than One: the Wikipedia Bitaxonomy ProjectabstractWe present WiBi, an approach to the automatic creation of a bitaxonomy for Wikipedia, that is, an integrated taxonomy of Wikipage pages and categories.We leverage the information available in either one of the taxonomies to reinforce the creation of the other taxonomy.Our experiments show higher quality and coverage than state-of-the-art resources like DBpedia, YAGO, MENTA, WikiNet and WikiTaxonomy.WiBi is available at http://wibitaxonomy.org. Tiziano Flati, Daniele Vannella, Tommaso Pasini, Roberto Navigli |
ACL (1) | 3 |