VLDB 2026 Research / reviewers in the wild / expert
Edoardo Barba
dblp:269/4565
· DBLP profile ↗
13ranked-venue papers
5as first author
12since 2021 · last 2024
0009-0004-4307-1498ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Maverick: Efficient and Accurate Coreference Resolution Defying Recent TrendsabstractLarge autoregressive generative models have emerged as the cornerstone for achieving the highest performance across several Natural Language Processing tasks.However, the urge to attain superior results has, at times, led to the premature replacement of carefully designed task-specific approaches without exhaustive experimentation.The Coreference Resolution task is no exception; all recent stateof-the-art solutions adopt large generative autoregressive models that outperform encoderbased discriminative systems.In this work, we challenge this recent trend by introducing Maverick, a carefully designed -yet simple -pipeline, which enables running a state-ofthe-art Coreference Resolution system within the constraints of an academic budget, outperforming models with up to 13 billion parameters with as few as 500 million parameters.Maverick achieves state-of-the-art performance on the CoNLL-2012 benchmark, training with up to 0.006x the memory resources and obtaining a 170x faster inference compared to previous state-of-the-art systems.We extensively validate the robustness of the Maverick framework with an array of diverse experiments, reporting improvements over prior systems in data-scarce, long-document, and out-of-domain settings.We release our code and models for research purposes at https: //github.com/SapienzaNLP/maverick-coref. Giuliano Martinelli 0001, Edoardo Barba, Roberto Navigli |
ACL (1) | 2 |
| 2024 | Guardians of the Machine Translation Meta-Evaluation: Sentinel Metrics Fall In!abstractStefano Perrella, Lorenzo Proietti, Alessandro Scirè, Edoardo Barba, Roberto Navigli. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Stefano Perrella, Lorenzo Proietti 0002, Alessandro Scirè, Edoardo Barba, Roberto Navigli |
ACL (1) | 4 |
| 2024 | Beyond Correlation: Interpretable Evaluation of Machine Translation MetricsabstractMachine Translation (MT) evaluation metrics assess translation quality automatically.Recently, researchers have employed MT metrics for various new use cases, such as data filtering and translation re-ranking.However, most MT metrics return assessments as scalar scores that are difficult to interpret, posing a challenge to making informed design choices.Moreover, MT metrics' capabilities have historically been evaluated using correlation with human judgment, which, despite its efficacy, falls short of providing intuitive insights into metric performance, especially in terms of new metric use cases.To address these issues, we introduce an interpretable evaluation framework for MT metrics.Within this framework, we evaluate metrics in two scenarios that serve as proxies for the data filtering and translation re-ranking use cases.Furthermore, by measuring the performance of MT metrics using Precision, Recall, and F -score, we offer clearer insights into their capabilities than correlation with human judgments.Finally, we raise concerns regarding the reliability of manually curated data following the Direct Assessments+Scalar Quality Metrics (DA+SQM) guidelines, reporting a notably low agreement with Multidimensional Quality Metrics (MQM) annotations. Stefano Perrella, Lorenzo Proietti 0002, Pere-Lluís Huguet Cabot, Edoardo Barba, Roberto Navigli |
EMNLP | 4 |
| 2024 | MOSAICo: a Multilingual Open-text Semantically Annotated Interlinked CorpusabstractSimone Conia, Edoardo Barba, Abelardo Carlos Martinez Lorenzo, Pere-Lluís Huguet Cabot, Riccardo Orlando, Luigi Procopio, Roberto Navigli. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Simone Conia, Edoardo Barba, Abelardo Carlos Martinez Lorenzo, Pere-Lluís Huguet Cabot, Riccardo Orlando, Luigi Procopio, Roberto Navigli |
NAACL-HLT | 2 |
| 2023 | Entity Disambiguation with Entity DefinitionsabstractLocal models have recently attained astounding performances in Entity Disambiguation (ED), with generative and extractive formulations being the most promising research directions.However, previous works have so far limited their studies to using, as the textual representation of each candidate, only its Wikipedia title.Although certainly effective, this strategy presents a few critical issues, especially when titles are not sufficiently informative or distinguishable from one another.In this paper, we address this limitation and investigate the extent to which more expressive textual representations can mitigate it.We evaluate our approach thoroughly against standard benchmarks in ED and find extractive formulations to be particularly well-suited to such representations.We report a new state of the art on 2 out of the 6 benchmarks we consider and strongly improve the generalization capability over unseen patterns.We release our code, data and model checkpoints at https: //github.com/SapienzaNLP/extend. Luigi Procopio, Simone Conia, Edoardo Barba, Roberto Navigli |
EACL | 3 |
| 2023 | LexicoMatic: Automatic Creation of Multilingual Lexical-Semantic DictionariesabstractFederico Martelli, Luigi Procopio, Edoardo Barba, Roberto Navigli. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Federico Martelli, Luigi Procopio, Edoardo Barba, Roberto Navigli |
IJCNLP (1) | 3 |
| 2022 | STEPS: Semantic Typing of Event Processes with a Sequence-to-Sequence ApproachabstractEnabling computers to comprehend the intent of human actions by processing language is one of the fundamental goals of Natural Language Understanding. An emerging task in this context is that of free-form event process typing, which aims at understanding the overall goal of a protagonist in terms of an action and an object, given a sequence of events. This task was initially treated as a learning-to-rank problem by exploiting the similarity between processes and action/object textual definitions. However, this approach appears to be overly complex, binds the output types to a fixed inventory for possible word definitions and, moreover, leaves space for further enhancements as regards performance. In this paper, we advance the field by reformulating the free-form event process typing task as a sequence generation problem and put forward STEPS, an end-to-end approach for producing user intent in terms of actions and objects only, dispensing with the need for their definitions. In addition to this, we eliminate several dataset constraints set by previous works, while at the same time significantly outperforming them. We release the data and software at https://github.com/SapienzaNLP/steps. Sveva Pepe, Edoardo Barba, Rexhina Blloshmi, Roberto Navigli |
AAAI | 2 |
| 2022 | ExtEnD: Extractive Entity DisambiguationabstractLocal models for Entity Disambiguation (ED) have today become extremely powerful, in most part thanks to the advent of large pretrained language models.However, despite their significant performance achievements, most of these approaches frame ED through classification formulations that have intrinsic limitations, both computationally and from a modeling perspective.In contrast with this trend, here we propose EXTEND, a novel local formulation for ED where we frame this task as a text extraction problem, and present two Transformer-based architectures that implement it.Based on experiments in and out of domain, and training over two different data regimes, we find our approach surpasses all its competitors in terms of both data efficiency and raw performance.EXTEND outperforms its alternatives by as few as 6 F 1 points on the more constrained of the two data regimes and, when moving to the other higher-resourced regime, sets a new state of the art on 4 out of 6 benchmarks under consideration, with average improvements of 0.7 F 1 points overall and 1.1 F 1 points out of domain.In addition, to gain better insights from our results, we also perform a fine-grained evaluation of our performances on different classes of label frequency, along with an ablation study of our architectural choices and an error analysis.We release our code and models for research purposes at https:// github.com/SapienzaNLP/extend. Edoardo Barba, Luigi Procopio, Roberto Navigli |
ACL (1) | 1 |
| 2021 | ConSeC: Word Sense Disambiguation as Continuous Sense ComprehensionabstractSupervised systems have nowadays become the standard recipe for Word Sense Disambiguation (WSD), with Transformer-based language models as their primary ingredient.However, while these systems have certainly attained unprecedented performances, virtually all of them operate under the constraining assumption that, given a context, each word can be disambiguated individually with no account of the other sense choices.To address this limitation and drop this assumption, we propose CONtinuous SEnse Comprehension (CONSEC), a novel approach to WSD: leveraging a recent re-framing of this task as a text extraction problem, we adapt it to our formulation and introduce a feedback loop strategy that allows the disambiguation of a target word to be conditioned not only on its context but also on the explicit senses assigned to nearby words.We evaluate CONSEC and examine how its components lead it to surpass all its competitors and set a new state of the art on English WSD.We also explore how CONSEC fares in the cross-lingual setting, focusing on 8 languages with various degrees of resource availability, and report significant improvements over prior systems.We release our code at https://github.com/ SapienzaNLP/consec. Edoardo Barba, Luigi Procopio, Roberto Navigli |
EMNLP (1) | 1 |
| 2021 | Exemplification Modeling: Can You Give Me an Example, Please?abstractRecently, generative approaches have been used effectively to provide definitions of words in their context. However, the opposite, i.e., generating a usage example given one or more words along with their definitions, has not yet been investigated. In this work, we introduce the novel task of Exemplification Modeling (ExMod), along with a sequence-to-sequence architecture and a training procedure for it. Starting from a set of (word, definition) pairs, our approach is capable of automatically generating high-quality sentences which express the requested semantics. As a result, we can drive the creation of sense-tagged data which cover the full range of meanings in any inventory of interest, and their interactions within sentences. Human annotators agree that the sentences generated are as fluent and semantically-coherent with the input definitions as the sentences in manually-annotated corpora. Indeed, when employed as training data for Word Sense Disambiguation, our examples enable the current state of the art to be outperformed, and higher results to be achieved than when using gold-standard datasets only. We release the pretrained model, the dataset and the software at https://github.com/SapienzaNLP/exmod. Edoardo Barba, Luigi Procopio, Caterina Lacerra, Tommaso Pasini, Roberto Navigli |
IJCAI | 1 |
| 2021 | MultiMirror: Neural Cross-lingual Word Alignment for Multilingual Word Sense DisambiguationabstractWord Sense Disambiguation (WSD), i.e., the task of assigning senses to words in context, has seen a surge of interest with the advent of neural models and a considerable increase in performance up to 80% F1 in English. However, when considering other languages, the availability of training data is limited, which hampers scaling WSD to many languages. To address this issue, we put forward MultiMirror, a sense projection approach for multilingual WSD based on a novel neural discriminative model for word alignment: given as input a pair of parallel sentences, our model -- trained with a low number of instances -- is capable of jointly aligning, at the same time, all source and target tokens with each other, surpassing its competitors across several language combinations. We demonstrate that projecting senses from English by leveraging the alignments produced by our model leads a simple mBERT-powered classifier to achieve a new state of the art on established WSD datasets in French, German, Italian, Spanish and Japanese. We release our software and all our datasets at https://github.com/SapienzaNLP/multimirror. Luigi Procopio, Edoardo Barba, Federico Martelli, Roberto Navigli |
IJCAI | 2 |
| 2021 | ESC: Redesigning WSD with Extractive Sense ComprehensionabstractWord Sense Disambiguation (WSD) is a historical NLP task aimed at linking words in contexts to discrete sense inventories and it is usually cast as a multi-label classification task.Recently, several neural approaches have employed sense definitions to better represent word meanings.Yet, these approaches do not observe the input sentence and the sense definition candidates all at once, thus potentially reducing the model performance and generalization power.We cope with this issue by reframing WSD as a span extraction problem -which we called Extractive Sense Comprehension (ESC) -and propose ESCHER, a transformer-based neural architecture for this new formulation.By means of an extensive array of experiments, we show that ESC unleashes the full potential of our model, leading it to outdo all of its competitors and to set a new state of the art on the English WSD task.In the few-shot scenario, ESCHER proves to exploit training data efficiently, attaining the same performance as its closest competitor while relying on almost three times fewer annotations.Furthermore, ESCHER can nimbly combine data annotated with senses from different lexical resources, achieving performances that were previously out of everyone's reach.The model along with data is available at https://github.com/ SapienzaNLP/esc. Edoardo Barba, Tommaso Pasini, Roberto Navigli |
NAACL-HLT | 1 |
| 2020 | MuLaN: Multilingual Label propagatioN for Word Sense DisambiguationabstractThe knowledge acquisition bottleneck strongly affects the creation of multilingual sense-annotated data, hence limiting the power of supervised systems when applied to multilingual Word Sense Disambiguation. In this paper, we propose a semi-supervised approach based upon a novel label propagation scheme, which, by jointly leveraging contextualized word embeddings and the multilingual information enclosed in a knowledge base, projects sense labels from a high-resource language, i.e., English, to lower-resourced ones. Backed by several experiments, we provide empirical evidence that our automatically created datasets are of a higher quality than those generated by other competitors and lead a supervised model to achieve state-of-the-art performances in all multilingual Word Sense Disambiguation tasks. We make our datasets available for research purposes at https://github.com/SapienzaNLP/mulan. Edoardo Barba, Luigi Procopio, Niccolò Campolungo, Tommaso Pasini, Roberto Navigli |
IJCAI | 1 |