Simone Conia

dblp:254/8205 · DBLP profile ↗
← Back
23ranked-venue papers
9as first author
20since 2021 · last 2026
0000-0002-6238-7816ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 9 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question Answering
abstract
Francesco Maria Molfese, Luca Moroni, Ciro Porcaro, Simone Conia, Roberto Navigli. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Francesco Molfese 0001, Luca Moroni, Ciro Porcaro, Simone Conia, Roberto Navigli
ACL (1)4
2025 Rewind and Render: Towards Factually Accurate Text-to-Video Generation with Distilled Knowledge Retrieval
abstract
Text-to-Video (T2V) models, despite recent advancements, struggle with factual accuracy, especially for knowledge-dense content. We introduce FACT-V (Factual Accuracy in Content Translation to Video), a system integrating multi-source knowledge retrieval into T2V pipelines. FACT-V offers two key benefits: i) improved factual accuracy of generated videos through dynamically retrieved information, and ii) increased interpretability by providing users with the augmented prompt information. A preliminary evaluation demonstrates the potential of knowledge-augmented approaches in improving the accuracy and reliability of T2V systems, particularly for entity-specific or time-sensitive prompts.
Arjun Chandra, Yunyao Li 0001, Simone Conia
AAAI5
2025 Do Large Language Models have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMs
abstract
Current Large Language Models (LLMs) are predominantly designed with English as the primary language, and even the few that are multilingual tend to exhibit strong English-centric biases.Much like speakers who might produce awkward expressions when learning a second language, LLMs often generate unnatural outputs in non-English languages, reflecting English-centric patterns in both vocabulary and grammar.Despite the importance of this issue, the naturalness of multilingual LLM outputs has received limited attention.In this paper, we address this gap by introducing novel automatic corpus-level metrics to assess the lexical and syntactic naturalness of LLM outputs in a multilingual context.Using our new metrics, we evaluate state-of-the-art LLMs on a curated benchmark in French and Chinese 1 , revealing a tendency towards English-influenced patterns.To mitigate this issue, we also propose a simple and effective alignment method to improve the naturalness of an LLM in a target language and domain, achieving consistent improvements in naturalness without compromising the performance on general-purpose benchmarks.Our work highlights the importance of developing multilingual metrics, resources and methods for the new wave of multilingual LLMs. * Work done during internship at Apple.
Yanzhu Guo, Simone Conia, Zelin Zhou, Saloni Potdar, Henry Xiao
ACL (1)2
2025 KG-TRICK: Unifying Textual and Relational Information Completion of Knowledge for Multilingual Knowledge Graphs
abstract
Multilingual knowledge graphs (KGs) provide high-quality relational and textual information for various NLP applications, but they are often incomplete, especially in non-English languages. Previous research has shown that combining information from KGs in different languages aids either Knowledge Graph Completion (KGC), the task of predicting missing relations between entities, or Knowledge Graph Enhancement (KGE), the task of predicting missing textual information for entities. Although previous efforts have considered KGC and KGE as independent tasks, we hypothesize that they are interdependent and mutually beneficial. To this end, we introduce KG-TRICK, a novel sequence-to-sequence framework that unifies the tasks of textual and relational information completion for multilingual KGs. KG-TRICK demonstrates that: i) it is possible to unify the tasks of KGC and KGE into a single framework, and ii) combining textual information from multiple languages is beneficial to improve the completeness of a KG. As part of our contributions, we also introduce WikiKGE10++, the largest manually-curated benchmark for textual information completion of KGs, which features over 25,000 entities across 10 diverse languages.
Zelin Zhou, Simone Conia, Shenglei Huang, Umar Farooq Minhas, Saloni Potdar, Henry Xiao, Yunyao Li 0001
COLING2
2025 Multi-LMentry: Can Multilingual LLMs Solve Elementary Tasks Across Languages?
abstract
Luca Moroni, Javier Aula-Blasco, Simone Conia, Irene Baucells, Naiara Perez, Silvia Paniagua Suárez, Anna Sallés, Malte Ostendorff, Júlia Falcão, Guijin Son, Aitor Gonzalez-Agirre, Roberto Navigli, Marta Villegas. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Luca Moroni, Javier Aula-Blasco, Simone Conia, Irene Baucells de la Peña, Naiara Pérez, Silvia Paniagua Suárez, Anna Salles, Malte Ostendorff, Júlia Falcão, Guijin Son, Aitor Gonzalez-Agirre, Roberto Navigli, Marta Villegas
EMNLP3
2024 Enhancing Machine Translation Experiences with Multilingual Knowledge Graphs
abstract
Translating entity names, especially when a literal translation is not correct, poses a significant challenge. Although Machine Translation (MT) systems have achieved impressive results, they still struggle to translate cultural nuances and language-specific context. In this work, we show that the integration of multilingual knowledge graphs into MT systems can address this problem and bring two significant benefits: i) improving the translation of utterances that contain entities by leveraging their human-curated aliases from a multilingual knowledge graph, and, ii) increasing the interpretability of the translation process by providing the user with information from the knowledge graph.
Simone Conia, Umar Farooq Minhas, Yunyao Li 0001
AAAI1
2024 CroCoAlign: A Cross-Lingual, Context-Aware and Fully-Neural Sentence Alignment System for Long Texts
abstract
Francesco Maria Molfese, Andrei Stefan Bejgu, Simone Tedeschi, Simone Conia, Roberto Navigli. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Francesco Molfese 0001, Andrei Stefan Bejgu, Simone Tedeschi, Simone Conia, Roberto Navigli
EACL (1)4
2024 ZEBRA: Zero-Shot Example-Based Retrieval Augmentation for Commonsense Question Answering
abstract
Current Large Language Models (LLMs) have shown strong reasoning capabilities in commonsense question answering benchmarks, but the process underlying their success remains largely opaque.As a consequence, recent approaches have equipped LLMs with mechanisms for knowledge retrieval, reasoning and introspection, not only to improve their capabilities but also to enhance the interpretability of their outputs.However, these methods require additional training, hand-crafted templates or human-written explanations.To address these issues, we introduce ZEBRA, a zero-shot question answering framework that combines retrieval, case-based reasoning and introspection and dispenses with the need for additional training of the LLM.Given an input question, ZEBRA retrieves relevant questionknowledge pairs from a knowledge base and generates new knowledge by reasoning over the relationships in these pairs.This generated knowledge is then used to answer the input question, improving the model's performance and interpretability.We evaluate our approach across 8 well-established commonsense reasoning benchmarks, demonstrating that ZEBRA consistently outperforms strong LLMs and previous knowledge integration approaches, achieving an average accuracy improvement of up to 4.5 points.
Francesco Molfese 0001, Simone Conia, Riccardo Orlando, Roberto Navigli
EMNLP2
2024 Towards Cross-Cultural Machine Translation with Retrieval-Augmented Generation from Multilingual Knowledge Graphs
abstract
Translating text that contains entity names is a challenging task, as cultural-related references can vary significantly across languages.These variations may also be caused by transcreation, an adaptation process that entails more than transliteration and word-for-word translation.In this paper, we address the problem of cross-cultural translation on two fronts: (i) we introduce XC-Translate, the first large-scale, manually-created benchmark for machine translation that focuses on text that contains potentially culturally-nuanced entity names, and (ii) we propose KG-MT, a novel end-to-end method to integrate information from a multilingual knowledge graph into a neural machine translation model by leveraging a dense retrieval mechanism.Our experiments and analyses show that current machine translation systems and large language models still struggle to translate texts containing entity names, whereas KG-MT outperforms state-of-the-art approaches by a large margin, obtaining a 129% and 62% relative improvement compared to NLLB-200 and GPT-4, respectively.
Simone Conia, Umar Farooq Minhas, Saloni Potdar, Yunyao Li 0001
EMNLP1
2024 MOSAICo: a Multilingual Open-text Semantically Annotated Interlinked Corpus
abstract
Simone Conia, Edoardo Barba, Abelardo Carlos Martinez Lorenzo, Pere-Lluís Huguet Cabot, Riccardo Orlando, Luigi Procopio, Roberto Navigli. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Simone Conia, Edoardo Barba, Abelardo Carlos Martinez Lorenzo, Pere-Lluís Huguet Cabot, Riccardo Orlando, Luigi Procopio, Roberto Navigli
NAACL-HLT1
2023 Entity Disambiguation with Entity Definitions
abstract
Local models have recently attained astounding performances in Entity Disambiguation (ED), with generative and extractive formulations being the most promising research directions.However, previous works have so far limited their studies to using, as the textual representation of each candidate, only its Wikipedia title.Although certainly effective, this strategy presents a few critical issues, especially when titles are not sufficiently informative or distinguishable from one another.In this paper, we address this limitation and investigate the extent to which more expressive textual representations can mitigate it.We evaluate our approach thoroughly against standard benchmarks in ED and find extractive formulations to be particularly well-suited to such representations.We report a new state of the art on 2 out of the 6 benchmarks we consider and strongly improve the generalization capability over unseen patterns.We release our code, data and model checkpoints at https: //github.com/SapienzaNLP/extend.
Luigi Procopio, Simone Conia, Edoardo Barba, Roberto Navigli
EACL2
2023 Increasing Coverage and Precision of Textual Information in Multilingual Knowledge Graphs
abstract
Recent work in Natural Language Processing and Computer Vision has been using textual information -e.g., entity names and descriptions -available in knowledge graphs to ground neural models to high-quality structured data.However, when it comes to non-English languages, the quantity and quality of textual information are comparatively scarce.To address this issue, we introduce the novel task of automatic Knowledge Graph Enhancement (KGE) and perform a thorough investigation on bridging the gap in both the quantity and quality of textual information between English and non-English languages.More specifically, we: i) bring to light the problem of increasing multilingual coverage and precision of entity names and descriptions in Wikidata; ii) demonstrate that state-of-the-art methods, namely, Machine Translation (MT), Web Search (WS), and Large Language Models (LLMs), struggle with this task; iii) present M-NTA, a novel unsupervised approach that combines MT, WS, and LLMs to generate high-quality textual information; and, iv) study the impact of increasing multilingual coverage and precision of non-English textual information in Entity Linking, Knowledge Graph Completion, and Question Answering.As part of our effort towards better multilingual knowledge graphs, we also introduce WikiKGE-10, the first human-curated benchmark to evaluate KGE approaches in 10 languages across 7 language families.
Simone Conia, Umar Farooq Minhas, Ihab F. Ilyas, Yunyao Li 0001
EMNLP1
2022 SRL4E - Semantic Role Labeling for Emotions: A Unified Evaluation Framework
abstract
In the field of sentiment analysis, several studies have highlighted that a single sentence may express multiple, sometimes contrasting, sentiments and emotions, each with its own experiencer, target and/or cause.To this end, over the past few years researchers have started to collect and annotate data manually, in order to investigate the capabilities of automatic systems not only to distinguish between emotions, but also to capture their semantic constituents.However, currently available gold datasets are heterogeneous in size, domain, format, splits, emotion categories and role labels, making comparisons across different works difficult and hampering progress in the area.In this paper, we tackle this issue and present a unified evaluation framework focused on Semantic Role Labeling for Emotions (SRL4E), in which we unify several datasets tagged with emotions and semantic roles by using a common labeling scheme.We use SRL4E as a benchmark to evaluate how modern pretrained language models perform and analyze where we currently stand in this task, hoping to provide the tools to facilitate studies in this complex area.
Cesare Campagnano, Simone Conia, Roberto Navigli
ACL (1)2
2022 Probing for Predicate Argument Structures in Pretrained Language Models
abstract
Thanks to the effectiveness and wide availability of modern pretrained language models (PLMs), recently proposed approaches have achieved remarkable results in dependencyand span-based, multilingual and cross-lingual Semantic Role Labeling (SRL).These results have prompted researchers to investigate the inner workings of modern PLMs with the aim of understanding how, where, and to what extent they encode information about SRL.In this paper, we follow this line of research and probe for predicate argument structures in PLMs.Our study shows that PLMs do encode semantic structures directly into the contextualized representation of a predicate, and also provides insights into the correlation between predicate senses and their structures, the degree of transferability between nominal and verbal structures, and how such structures are encoded across languages.Finally, we look at the practical implications of such insights and demonstrate the benefits of embedding predicate argument structure information into an SRL model.
Simone Conia, Roberto Navigli
ACL (1)1
2022 Nibbling at the Hard Core of Word Sense Disambiguation
abstract
With state-of-the-art systems having finally attained estimated human performance, Word Sense Disambiguation (WSD) has now joined the array of Natural Language Processing tasks that have seemingly been solved, thanks to the vast amounts of knowledge encoded into Transformer-based pre-trained language models.And yet, if we look below the surface of raw figures, it is easy to realize that current approaches still make trivial mistakes that a human would never make.In this work, we provide evidence showing why the F1 score metric should not simply be taken at face value and present an exhaustive analysis of the errors that seven of the most representative state-of-the-art systems for English all-words WSD make on traditional evaluation benchmarks.In addition, we produce and release a collection of test sets featuring (a) an amended version of the standard evaluation benchmark that fixes its lexical and semantic inaccuracies, (b) 42D, a challenge set devised to assess the resilience of systems with respect to least frequent word senses and senses not seen at training time, and (c) hardEN, a challenge set made up solely of instances which none of the investigated state-of-the-art systems can solve.We make all of the test sets and model predictions available to the research community at https://github.com/ SapienzaNLP/wsd-hard-benchmark.
Marco Maru, Simone Conia, Michele Bevilacqua, Roberto Navigli
ACL (1)2
2022 Universal Semantic Annotator: the First Unified API for WSD, SRL and Semantic Parsing
abstract
In this paper, we present the Universal Semantic Annotator (USeA), which offers the first unified API for high-quality automatic annotations of texts in 100 languages through state-of-the-art systems for Word Sense Disambiguation, Semantic Role Labeling and Semantic Parsing. Together, such annotations can be used to provide users with rich and diverse semantic information, help second-language learners, and allow researchers to integrate explicit semantic knowledge into downstream tasks and real-world applications.
Riccardo Orlando, Simone Conia, Stefano Faralli 0001, Roberto Navigli
LREC2
2021 Framing Word Sense Disambiguation as a Multi-Label Problem for Model-Agnostic Knowledge Integration
abstract
Recent studies treat Word Sense Disambiguation (WSD) as a single-label classification problem in which one is asked to choose only the best-fitting sense for a target word, given its context.However, gold data labelled by expert annotators suggest that maximizing the probability of a single sense may not be the most suitable training objective for WSD, especially if the sense inventory of choice is finegrained.In this paper, we approach WSD as a multi-label classification problem in which multiple senses can be assigned to each target word.Not only does our simple method bear a closer resemblance to how human annotators disambiguate text, but it can also be extended seamlessly to exploit structured knowledge from semantic networks to achieve stateof-the-art results in English all-words WSD.
Simone Conia, Roberto Navigli
EACL1
2021 Generating Senses and RoLes: An End-to-End Model for Dependency- and Span-based Semantic Role Labeling
abstract
Despite the recent great success of the sequence-to-sequence paradigm in Natural Language Processing, the majority of current studies in Semantic Role Labeling (SRL) still frame the problem as a sequence labeling task. In this paper we go against the flow and propose GSRL (Generating Senses and RoLes), the first sequence-to-sequence model for end-to-end SRL. Our approach benefits from recently-proposed decoder-side pretraining techniques to generate both sense and role labels for all the predicates in an input sentence at once, in an end-to-end fashion. Evaluated on standard gold benchmarks, GSRL achieves state-of-the-art results in both dependency- and span-based English SRL, proving empirically that our simple generation-based model can learn to produce complex predicate-argument structures. Finally, we propose a framework for evaluating the robustness of an SRL model in a variety of synthetic low-resource scenarios which can aid human annotators in the creation of better, more diverse, and more challenging gold datasets. We release GSRL at github.com/SapienzaNLP/gsrl.
Rexhina Blloshmi, Simone Conia, Rocco Tripodi, Roberto Navigli
IJCAI2
2021 Ten Years of BabelNet: A Survey
abstract
The intelligent manipulation of symbolic knowledge has been a long-sought goal of AI. However, when it comes to Natural Language Processing (NLP), symbols have to be mapped to words and phrases, which are not only ambiguous but also language-specific: multilinguality is indeed a desirable property for NLP systems, and one which enables the generalization of tasks where multiple languages need to be dealt with, without translating text. In this paper we survey BabelNet, a popular wide-coverage lexical-semantic knowledge resource obtained by merging heterogeneous sources into a unified semantic network that helps to scale tasks and applications to hundreds of languages. Over its ten years of existence, thanks to its promise to interconnect languages and resources in structured form, BabelNet has been employed in countless ways and directions. We first introduce the BabelNet model, its components and statistics, and then overview its successful use in a wide range of tasks in NLP as well as in other fields of AI.
Roberto Navigli, Michele Bevilacqua, Simone Conia, Dario Montagnini, Francesco Cecconi
IJCAI3
2021 Unifying Cross-Lingual Semantic Role Labeling with Heterogeneous Linguistic Resources
abstract
While cross-lingual techniques are finding increasing success in a wide range of Natural Language Processing tasks, their application to Semantic Role Labeling (SRL) has been strongly limited by the fact that each language adopts its own linguistic formalism, from Prop-Bank for English to AnCora for Spanish and PDT-Vallex for Czech, inter alia.In this work, we address this issue and present a unified model to perform cross-lingual SRL over heterogeneous linguistic resources.Our model implicitly learns a high-quality mapping for different formalisms across diverse languages without resorting to word alignment and/or translation techniques.We find that, not only is our cross-lingual system competitive with the current state of the art but that it is also robust to low-data scenarios.Most interestingly, our unified model is able to annotate a sentence in a single forward pass with all the inventories it was trained with, providing a tool for the analysis and comparison of linguistic theories across different languages.We release our code and model at https://github.com/ SapienzaNLP/unify-srl.
Simone Conia, Andrea Bacciu, Roberto Navigli
NAACL-HLT1
2020 Bridging the Gap in Multilingual Semantic Role Labeling: a Language-Agnostic Approach
abstract
Recent research indicates that taking advantage of complex syntactic features leads to favorable results in Semantic Role Labeling.Nonetheless, an analysis of the latest state-of-the-art multilingual systems reveals the difficulty of bridging the wide gap in performance between highresource (e.g., English) and low-resource (e.g., German) settings.To overcome this issue, we propose a fully language-agnostic model that does away with morphological and syntactic features to achieve robustness across languages.Our approach outperforms the state of the art in all the languages of the CoNLL-2009 benchmark dataset, especially whenever a scarce amount of training data is available.Our objective is not to reject approaches that rely on syntax, rather to set a strong and consistent language-independent baseline for future innovations in Semantic Role Labeling.We release our model code and checkpoints at https
Simone Conia, Roberto Navigli
COLING1
2020 Conception: Multilingually-Enhanced, Human-Readable Concept Vector Representations
abstract
To date, the most successful word, word sense, and concept modelling techniques have used large corpora and knowledge resources to produce dense vector representations that capture semantic similarities in a relatively low-dimensional space.Most current approaches, however, suffer from a monolingual bias, with their strength depending on the amount of data available across languages.In this paper we address this issue and propose Conception, a novel technique for building language-independent vector representations of concepts which places multilinguality at its core while retaining explicit relationships between concepts.Our approach results in highcoverage representations that outperform the state of the art in multilingual and cross-lingual Semantic Word Similarity and Word Sense Disambiguation, proving particularly robust on lowresource languages.Conception -its software and the complete set of representations -is available at https://github.
Simone Conia, Roberto Navigli
COLING1
2019 VerbAtlas: a Novel Large-Scale Verbal Semantic Resource and Its Application to Semantic Role Labeling
abstract
Andrea Di Fabio, Simone Conia, Roberto Navigli. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Andrea Di Fabio, Simone Conia, Roberto Navigli
EMNLP/IJCNLP (1)2