Davide Buscaldi

dblp:34/4842 · DBLP profile ↗
← Back
20ranked-venue papers in the field
6as first author
8since 2021 · last 2026
0000-0003-1112-3789ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 7 (2 first)Information Retrieval & Web Search · 6 (2 first)Database Systems & Data Management · 5 (2 first)Data Mining & Knowledge Discovery · 2
YearPublicationVenuePosition
2026 Leveraging knowledge graphs and LLMs for content-based reviewer assignment
abstract
Abstract The growing volume of academic submissions in recent years highlighted the need for scalable and accurate reviewer assignment systems, able to go beyond techniques based on manual processes and basic keyword matching. We propose a novel pipeline that integrates Knowledge Graphs (KGs) and Large Language Models (LLMs) to automate and enhance the reviewer assignment process. Our method extracts meaningful representations of papers and reviewer expertise using Open Information Extraction, the Computer Science Ontology classifier, and GLiNER to build KGs from research content. LLMs are employed to generate targeted keywords through prompt-based synthesis, refining both paper and reviewer profiles. The assignment relies on a hybrid similarity metric combining Cosine and Jaccard similarities to capture both lexical and semantic alignment. We evaluate the pipeline using standard metrics such as Mean Reciprocal Rank, Mean Average Precision, and Precision at K, on a dataset in the Computer Science domain, demonstrating its effectiveness in aligning submissions with appropriate reviewers. This approach offers a scalable and adaptive solution to the complexities of modern peer review.
Farid Bagheri, Davide Buscaldi, Diego Reforgiato Recupero
J. Intell. Inf. Syst.2
2025 Curvature constrained MPNNs: Improving message passing with local structural properties
abstract
International audience
Hugo Attali, Davide Buscaldi, Nathalie Pernelle
Data Knowl. Eng.2
2024 Workshop on Deep Learning and Large Language Models for Knowledge Graphs (DL4KG)
abstract
The use of Knowledge Graphs (KGs) which constitute large networks of real-world entities and their interrelationships, has grown rapidly. A substantial body of research has emerged, exploring the integration of deep learning (DL) and large language models (LLMs) with KGs. This workshop aims to bring together leading researchers in the field to discuss and foster collaborations on the intersection of KG and DL/LLMs.
Mehwish Alam, Davide Buscaldi, Michael Cochez, Genet Asefa Gesese, Francesco Osborne, Diego Reforgiato Recupero
KDD2
2024 Citation prediction by leveraging transformers and natural language processing heuristics
abstract
In scientific papers, it is common practice to cite other articles to substantiate claims, provide evidence for factual assertions, reference limitations, and research gaps, and fulfill various other purposes. When authors include a citation in a given sentence, there are two considerations they need to take into account: (i) where in the sentence to place the citation and (ii) which citation to choose to support the underlying claim. In this paper, we focus on the first task as it allows multiple potential approaches that rely on the researcher’s individual style and the specific norms and conventions of the relevant scientific community. We propose two automatic methodologies that leverage transformers architecture for either solving a Mask-Filling problem or a Named Entity Recognition problem. On top of the results of the proposed methodologies, we apply ad-hoc Natural Language Processing heuristics to further improve their outcome. We also introduce s2orc-9K, an open dataset for fine-tuning models on this task. A formal evaluation demonstrates that the generative approach significantly outperforms five alternative methods when fine-tuned on the novel dataset. Furthermore, this model’s results show no statistically significant deviation from the outputs of three senior researchers.
Davide Buscaldi, Danilo Dessì, Enrico Motta, Marco Murgia, Francesco Osborne, Diego Reforgiato Recupero
Inf. Process. Manag.1
2023 Detecting Artificially Generated Academic Text: The Importance of Mimicking Human Utilization of Large Language Models
Vijini Liyanage, Davide Buscaldi
NLDB2
2022 Transformer-Based Models for the Automatic Indexing of Scientific Documents in French
José-Ángel González, Davide Buscaldi, Emilio Sanchis Arnal, Lluís F. Hurtado
NLDB2
2022 CS-KG: A Large-Scale Knowledge Graph of Research Entities and Claims in Computer Science
Danilo Dessì, Francesco Osborne, Diego Reforgiato Recupero, Davide Buscaldi, Enrico Motta
ISWC4
2021 Sherloc: a knowledge-driven algorithm for geolocating microblog messages at sub-city level
abstract
Many solutions for coarse geolocating of users at the time they post a message exist. However, for many important applications, like traffic monitoring and event detection, finer geolocation at the level of city neighborhoods, i.e., at a sub-city level, is needed. Data-driven approaches often do not guarantee good accuracy and efficiency due to the higher number of sub-city level positions to be estimated and the low availability of balanced and large training sets. We claim that external information sources overcome limitations of data-driven approaches in achieving good accuracy for sub-city level geolocation and we present a knowledge-driven approach achieving good results once the reference area of a message is known. Our algorithm, called Sherloc, exploits toponyms in the message, extracts their semantic from a geographic gazetteer, and embeds them into a metric space that captures the semantic distance among them. We identify the semantically closest toponyms to a message and then cluster them with respect to their spatial locations. Sherloc requires no prior training, it can infer the location at sub-city level with high accuracy, and it is not limited to geolocating on a fixed spatial grid.
Laura Di Rocco, Federico Dassereto, Michela Bertolotto, Davide Buscaldi, Barbara Catania, Giovanna Guerrini
Int. J. Geogr. Inf. Sci.4
2020 AI-KG: An Automatically Generated Knowledge Graph of Artificial Intelligence
abstract
Scientific knowledge has been traditionally disseminated and preserved through research articles published in journals, conference proceedings, and online archives. However, this article-centric paradigm has been often criticized for not allowing to automatically process, categorize, and reason on this knowledge. An alternative vision is to generate a semantically rich and interlinked description of the content of research publications. In this paper, we present the Artificial Intelligence Knowledge Graph (AI-KG), a large-scale automatically generated knowledge graph that describes 820K research entities. AI-KG includes about 14M RDF triples and 1.2M reified statements extracted from 333K research publications in the field of AI, and describes 5 types of entities (tasks, methods, metrics, materials, others) linked by 27 relations. AI-KG has been designed to support a variety of intelligent services for analyzing and making sense of research dynamics, supporting researchers in their daily job, and helping to inform decision-making in funding bodies and research policymakers. AI-KG has been generated by applying an automatic pipeline that extracts entities and relationships using three tools: DyGIE++, Stanford CoreNLP, and the CSO Classifier. It then integrates and filters the resulting triples using a combination of deep learning and semantic technologies in order to produce a high-quality knowledge graph. This pipeline was evaluated on a manually crafted gold standard, yielding competitive results. AI-KG is available under CC BY 4.0 and can be downloaded as a dump or queried via a SPARQL endpoint.
Danilo Dessì, Francesco Osborne, Diego Reforgiato Recupero, Davide Buscaldi, Enrico Motta, Harald Sack
ISWC (2)4
2016 Event-Based Recognition of Lived Experiences in User Reviews
Ehab Hassan, Davide Buscaldi, Aldo Gangemi
EKAW2
2016 Unsupervised Relation Extraction in Specialized Corpora Using Sequence Mining
Kata Gábor, Haïfa Zargayouna, Isabelle Tellier, Davide Buscaldi, Thierry Charnois
IDA4
2013 Using the Semantics of Texts for Information Retrieval: A Concept- and Domain Relation-Based Approach
Davide Buscaldi, Marie-Noëlle Bessagnet, Albert Royer, Christian Sallaberry
ADBIS (2)1
2013 Effects of Ontology Pitfalls on Ontology-based Information Retrieval Systems
abstract
Nowadays, a growing number of information retrieval systems make use of ontologies to improve the access to textual information, especially in domain-specific scenarios, where the knowledge provided by ontologies represents a key factor. Such kinds of retrieval systems are often referred to as ontology-based or semantic information retrieval systems. The quality of ontologies plays an important role in such systems in the sense that modelling errors in the ontologies may deteriorate the quality of the results obtained by these systems. In this paper we provide a comprehensive analysis of how ontology pitfalls have an influence on these kinds of systems. This study allows us to have a more complete understanding of the role of ontology quality in the information retrieval field. Our survey shows that pitfalls may act as an indicator not only of possible problems in ontology design, but also of OWL features overseen by system developers.
Davide Buscaldi, Mari Carmen Suárez-Figueroa
KEOD1
2013 A semi-automatic approach for building ontologies from acollection of structured web documents
abstract
Many collections of structured documents are available on the web. The collection generally describes the characteristics of entities from a single type, where each page describes one entity. These documents are adequate knowledge sources for building ontologies. As they benefit from a strong and shared layout, they contain less well written text than plain text files but their architecture is very meaningful. Classical linguistic-based methods for identifying concepts and relations are no longer appropriate for analyzing them.The approach we propose in this paper exploits various properties of such documents, combining layout/formatting analysis and linguistic analysis, and using semantic annotation.
Mouna Kamel, Nathalie Aussenac-Gilles, Davide Buscaldi, Catherine Comparot
K-CAP3
2012 From humor recognition to irony detection: The figurative language of social media
Antonio Reyes, Paolo Rosso, Davide Buscaldi
Data Knowl. Eng.3
2011 A pretopological framework for the automatic construction of lexical-semantic structures from texts
abstract
We present in this paper a new approach for the automatic generation of lexical structures from texts. This tedious task is based on the strong hypothesis that simple statistical observations on textual usages can provide pieces of semantics about the lexicon. Using such "naive" observations only, we propose a (pre)-topological framework to formalize and combine various hypothesis on textual data usages and then to derive a structure similar to usual lexical knowledge basis such as WordNet. In addition we also consider the evaluation problem for obtained lexical structures ; a multi-level evaluation strategy is proposed that measures the fitting between a given reference structure and automatically generated structures on different point of views : intrinsic/structural and application-based points of view. The evaluation strategy is then used to quantify the contribution of the new structuring approach with respect to the corresponding solution proposed by (Sanderson et al. 2000) on two case studies that differs on the domain and the size of the lexicon.
Guillaume Cleuziou, Davide Buscaldi, Vincent Levorato, Gaël Dias
CIKM2
2010 Answering questions with an n-gram based passage retrieval engine
Davide Buscaldi, Paolo Rosso, José Manuel Gómez Soriano, Emilio Sanchis Arnal
J. Intell. Inf. Syst.1
2009 The Impact of Semantic and Morphosyntactic Ambiguity on Automatic Humour Recognition
Antonio Reyes, Davide Buscaldi, Paolo Rosso
NLDB2
2009 Toponym ambiguity in geographical information retrieval
abstract
The objectives of this research work is to study the effects of toponym (place name) ambiguity in the Geographical Information Retrieval (GIR) task. Our experience with GIR systems shows that toponym ambiguity may be an important factor in the inability of these systems to take advantage from geographical knowledge. Previous studies over ambiguity and Information Retrieval (IR) suggested that disambiguation may be useful in some specific IR scenario. We suppose that GIR may constitute such a scenario. This preliminary study was carried out over the WordNet based, manually disambiguated collection developed for the CLIR-WSD task, using the GeoCLEF collection of 100 geographically related topics. The employed GIR system was based on the GeoWorSE system that participated in GeoCLEF 2008. The experiments were carried out considering the manual disambiguation and comparing this result with those obtained by randomly disambiguating the document collection and those obtained by using always the most common referent. The obtained results show no significant difference in the overall results, although the work gave an insight into some errors that are produced by toponym ambiguity and how they may affect the results. These preliminary results also suggest that WordNet is not a suitable resource for the planned research.
Davide Buscaldi
SIGIR1
2008 A conceptual density-based approach for the disambiguation of toponyms
abstract
Nowadays, a huge quantity of information is stored in digital format. A great portion of this information is constituted by textual and unstructured documents, where geographical references are usually given by means of place names. A common problem with textual information retrieval is represented by polysemous words, that is, words can have more than one sense. This problem is present also in the geographical domain: place names may refer to different locations in the world. In this paper we investigate the use of our word sense disambiguation technique in the geographical domain, with the aim of resolving ambiguous place names. Our technique is based on WordNet conceptual density. Due to the lack of a reference corpus tagged with WordNet senses, we carried out the experiments over a set of 1,210 place names extracted from the SemCor corpus that we named GeoSemCor and made publicly available. We compared our method with the most‐frequent baseline and the enhanced‐Lesk method, which previously has not been tested in large contexts. The results show that a better precision can be achieved by using a small context (phrase level), whereas a greater coverage can be obtained by using large contexts (document level). The proposed method should be tested with other corpora, due to the fact that our experiments evidenced the excessive bias towards the most‐frequent sense of the GeoSemCor.
Davide Buscaldi, Paolo Rosso
Int. J. Geogr. Inf. Sci.1