Davide Buscaldi

dblp:34/4842 · DBLP profile ↗
← Back
42ranked-venue papers
12as first author
18since 2021 · last 2026
0000-0003-1112-3789ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 5 first-author · 12 since 2021Databases, data management, data science and information retrieval · 20 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Guess or Recall? Training CNNs to Classify and Localize Memorization in LLMs
abstract
Verbatim memorization in Large Language Models (LLMs) is a multifaceted phenomenon involving distinct underlying mechanisms. We introduce a novel method to analyze the different forms of memorization described by the existing taxonomy. Specifically, we train Convolutional Neural Networks (CNNs) on the attention weights of the LLM and evaluate the alignment between this taxonomy and the attention weights involved in decoding. We find that the existing taxonomy performs poorly and fails to reflect distinct mechanisms within the attention blocks. We propose a new taxonomy that maximizes alignment with the attention weights, consisting of three categories: memorized samples that are guessed using language modeling abilities, memorized samples that are recalled due to high duplication in the training set, and non-memorized samples. Our results reveal that few-shot verbatim memorization does not correspond to a distinct attention mechanism. We also show that a significant proportion of extractable samples are in fact guessed by the model and should therefore be studied separately. Finally, we develop a custom visual interpretability technique to localize the regions of the attention weights involved in each form of memorization.
Jérémie Dentan, Davide Buscaldi, Sonia Haddad-Vanier
AAAI2
2026 MUCH: A Multilingual Claim Hallucination Benchmark
abstract
Claim-level Uncertainty Quantification (UQ) is a promising approach to mitigate the lack of reliability in Large Language Models (LLMs). We introduce MUCH, the first claim-level UQ benchmark designed for fair and reproducible evaluation of future methods under realistic conditions. It includes 4,873 samples across four European languages (English, French, Spanish, and German) and four instruction-tuned open-weight LLMs. Unlike prior claim-level benchmarks, we release 24 generation logits per token, facilitating the development of future white-box methods without re-generating data. Moreover, in contrast to previous benchmarks that rely on manual or LLM-based segmentation, we propose a new deterministic algorithm capable of segmenting claims using as little as 0.2% of the LLM generation time. This makes our segmentation approach suitable for real-time monitoring of LLM outputs, ensuring that MUCH evaluates UQ methods under realistic deployment constraints. Finally, our evaluations show that current methods still have substantial room for improvement in both performance and efficiency.
Jérémie Dentan, Alexi Canesse, Davide Buscaldi, Aymen Shabou, Sonia Haddad-Vanier
LREC3
2026 Leveraging knowledge graphs and LLMs for content-based reviewer assignment
abstract
Abstract The growing volume of academic submissions in recent years highlighted the need for scalable and accurate reviewer assignment systems, able to go beyond techniques based on manual processes and basic keyword matching. We propose a novel pipeline that integrates Knowledge Graphs (KGs) and Large Language Models (LLMs) to automate and enhance the reviewer assignment process. Our method extracts meaningful representations of papers and reviewer expertise using Open Information Extraction, the Computer Science Ontology classifier, and GLiNER to build KGs from research content. LLMs are employed to generate targeted keywords through prompt-based synthesis, refining both paper and reviewer profiles. The assignment relies on a hybrid similarity metric combining Cosine and Jaccard similarities to capture both lexical and semantic alignment. We evaluate the pipeline using standard metrics such as Mean Reciprocal Rank, Mean Average Precision, and Precision at K, on a dataset in the Computer Science domain, demonstrating its effectiveness in aligning submissions with appropriate reviewers. This approach offers a scalable and adaptive solution to the complexities of modern peer review.
Farid Bagheri, Davide Buscaldi, Diego Reforgiato Recupero
J. Intell. Inf. Syst.2
2025 PACT: Pruning and Clustering-Based Token Reduction for Faster Visual Language Models
abstract
Visual Language Models require substantial computational resources for inference due to the additional input tokens needed to represent visual information. However, these visual tokens often contain redundant and unimportant information, resulting in an unnecessarily high number of tokens. To address this, we introduce PACT, a method that reduces inference time and memory usage by pruning irrelevant tokens and merging visually redundant ones at an early layer of the language model. Our approach uses a novel importance metric to identify unimportant tokens without relying on attention scores, making it compatible with FlashAttention. We also propose a novel clustering algorithm, called Distance Bounded Density Peak Clustering, which efficiently clusters visual tokens while constraining the distances between elements within a cluster by a predefined threshold. We demonstrate the effectiveness of PACT through extensive experiments.
Mohamed Dhouib, Davide Buscaldi, Sonia Haddad-Vanier, Aymen Shabou
CVPR2
2025 Predicting Memorization Within Large Language Models Fine-Tuned for Classification
abstract
Large Language Models have received significant attention due to their abilities to solve a wide range of complex tasks. However these models memorize a significant proportion of their training data, posing a serious threat when disclosed at inference time. To mitigate this unintended memorization, it is crucial to understand what elements are memorized and why. This area of research is largely unexplored, with most existing works providing a posteriori explanations. To address this gap, we propose a new approach to detect memorized samples a priori in LLMs fine-tuned for classification tasks. This method is effective from the early stages of training and readily adaptable to other classification settings, such as training vision models from scratch. Our method is supported by new theoretical results, and requires a low computational budget. We achieve strong empirical results, paving the way for the systematic identification and protection of vulnerable samples before they are memorized.
Jérémie Dentan, Davide Buscaldi, Aymen Shabou, Sonia Haddad-Vanier
ECAI2
2025 Curvature constrained MPNNs: Improving message passing with local structural properties
abstract
International audience
Hugo Attali, Davide Buscaldi, Nathalie Pernelle
Data Knowl. Eng.2
2025 Research hypothesis generation over scientific knowledge graphs
abstract
Generating research hypotheses is a crucial step in scientific investigation that involves the creation of precise, verifiable, and logically valid statements that can be empirically examined. Therefore, many efforts have been made to automate or assist this process through the use of various Artificial Intelligence solutions. However, most existing methods are tailored to very specific domains, particularly within the biomedical field. There have been recent attempts to formalize hypothesis generation as a link prediction task over knowledge graphs. This solution is potentially domain-independent and applicable across diverse disciplines. Nevertheless, current approaches for link prediction, which typically rely on embedding models or path-based methods, have shown limited success in accurately predicting new hypotheses. To address these limitations, this paper introduces ResearchLink, an innovative and domain-independent methodology for hypothesis generation over knowledge graphs. ResearchLink combines path-based features and knowledge graph embeddings with text embeddings, capturing the semantic context of entities within a given corpus, and integrates additional information from bibliometric databases to improve research collaboration predictions. To conduct a rigorous evaluation of ResearchLink, we constructed CSKG-600, a new dataset for hypothesis generation, consisting of 600 statements that were manually labelled by domain experts. ResearchLink achieved outstanding performance (78.7% P@20), significantly outperforming alternative approaches such as TransH (71.8%), TransD (71.7%), and RotatE (70.7%).
Agustín Borrego, Danilo Dessì, Daniel Ayala Hernández, Inma Hernández, Francesco Osborne, Diego Reforgiato Recupero, Davide Buscaldi, David Ruiz 0001, Enrico Motta
Knowl. Based Syst.7
2024 Delaunay Graph: Addressing Over-Squashing and Over-Smoothing Using Delaunay Triangulation
abstract
GNNs rely on the exchange of messages to distribute information along the edges of the graph. This approach makes the efficiency of architectures highly dependent on the specific structure of the input graph. Certain graph topologies lead to inefficient information propagation, resulting in a phenomenon known as over-squashing. While the majority of existing methods address over-squashing by rewiring the input graph, our novel approach involves constructing a graph directly from features using Delaunay Triangulation. We posit that the topological properties of the resulting graph prove advantageous for mitigate oversmoothing and over-squashing. Our extensive experimentation demonstrates that our method consistently outperforms established graph rewiring methods.
Hugo Attali, Davide Buscaldi, Nathalie Pernelle
ICML2
2024 Workshop on Deep Learning and Large Language Models for Knowledge Graphs (DL4KG)
abstract
The use of Knowledge Graphs (KGs) which constitute large networks of real-world entities and their interrelationships, has grown rapidly. A substantial body of research has emerged, exploring the integration of deep learning (DL) and large language models (LLMs) with KGs. This workshop aims to bring together leading researchers in the field to discuss and foster collaborations on the intersection of KG and DL/LLMs.
Mehwish Alam, Davide Buscaldi, Michael Cochez, Genet Asefa Gesese, Francesco Osborne, Diego Reforgiato Recupero
KDD2
2024 Citation prediction by leveraging transformers and natural language processing heuristics
abstract
In scientific papers, it is common practice to cite other articles to substantiate claims, provide evidence for factual assertions, reference limitations, and research gaps, and fulfill various other purposes. When authors include a citation in a given sentence, there are two considerations they need to take into account: (i) where in the sentence to place the citation and (ii) which citation to choose to support the underlying claim. In this paper, we focus on the first task as it allows multiple potential approaches that rely on the researcher’s individual style and the specific norms and conventions of the relevant scientific community. We propose two automatic methodologies that leverage transformers architecture for either solving a Mask-Filling problem or a Named Entity Recognition problem. On top of the results of the proposed methodologies, we apply ad-hoc Natural Language Processing heuristics to further improve their outcome. We also introduce s2orc-9K, an open dataset for fine-tuning models on this task. A formal evaluation demonstrates that the generative approach significantly outperforms five alternative methods when fine-tuned on the novel dataset. Furthermore, this model’s results show no statistically significant deviation from the outputs of three senior researchers.
Davide Buscaldi, Danilo Dessì, Enrico Motta, Marco Murgia, Francesco Osborne, Diego Reforgiato Recupero
Inf. Process. Manag.1
2023 Detecting Artificially Generated Academic Text: The Importance of Mimicking Human Utilization of Large Language Models
Vijini Liyanage, Davide Buscaldi
NLDB2
2022 A Benchmark Corpus for the Detection of Automatically Generated Text in Academic Publications
abstract
Automatic text generation based on neural language models has achieved performance levels that make the generated text almost indistinguishable from those written by humans. Despite the value that text generation can have in various applications, it can also be employed for malicious tasks. The diffusion of such practices represent a threat to the quality of academic publishing. To address these problems, we propose in this paper two datasets comprised of artificially generated research content: a completely synthetic dataset and a partial text substitution dataset. In the first case, the content is completely generated by the GPT-2 model after a short prompt extracted from original papers. The partial or hybrid dataset is created by replacing several sentences of abstracts with sentences that are generated by the Arxiv-NLP model. We evaluate the quality of the datasets comparing the generated texts to aligned original texts using fluency metrics such as BLEU and ROUGE. The more natural the artificial texts seem, the more difficult they are to detect and the better is the benchmark. We also evaluate the difficulty of the task of distinguishing original from generated text by using state-of-the-art classification models.
Vijini Liyanage, Davide Buscaldi, Adeline Nazarenko
LREC2
2022 Transformer-Based Models for the Automatic Indexing of Scientific Documents in French
José-Ángel González, Davide Buscaldi, Emilio Sanchis Arnal, Lluís F. Hurtado
NLDB2
2022 CS-KG: A Large-Scale Knowledge Graph of Research Entities and Claims in Computer Science
Danilo Dessì, Francesco Osborne, Diego Reforgiato Recupero, Davide Buscaldi, Enrico Motta
ISWC4
2022 Special issue on senti-mental health: Future generation sentiment analysis systems
Davide Buscaldi, Mauro Dragoni, Flavius Frasincar, Diego Reforgiato Recupero
Future Gener. Comput. Syst.1
2022 SCICERO: A deep learning and NLP approach for generating scientific knowledge graphs in the computer science domain
Danilo Dessì, Francesco Osborne, Diego Reforgiato Recupero, Davide Buscaldi, Enrico Motta
Knowl. Based Syst.4
2021 Generating knowledge graphs by employing Natural Language Processing and Machine Learning techniques within the scholarly domain
Danilo Dessì, Francesco Osborne, Diego Reforgiato Recupero, Davide Buscaldi, Enrico Motta
Future Gener. Comput. Syst.4
2021 Sherloc: a knowledge-driven algorithm for geolocating microblog messages at sub-city level
abstract
Many solutions for coarse geolocating of users at the time they post a message exist. However, for many important applications, like traffic monitoring and event detection, finer geolocation at the level of city neighborhoods, i.e., at a sub-city level, is needed. Data-driven approaches often do not guarantee good accuracy and efficiency due to the higher number of sub-city level positions to be estimated and the low availability of balanced and large training sets. We claim that external information sources overcome limitations of data-driven approaches in achieving good accuracy for sub-city level geolocation and we present a knowledge-driven approach achieving good results once the reference area of a message is known. Our algorithm, called Sherloc, exploits toponyms in the message, extracts their semantic from a geographic gazetteer, and embeds them into a metric space that captures the semantic distance among them. We identify the semantically closest toponyms to a message and then cluster them with respect to their spatial locations. Sherloc requires no prior training, it can infer the location at sub-city level with high accuracy, and it is not limited to geolocating on a fixed spatial grid.
Laura Di Rocco, Federico Dassereto, Michela Bertolotto, Davide Buscaldi, Barbara Catania, Giovanna Guerrini
Int. J. Geogr. Inf. Sci.4
2020 AI-KG: An Automatically Generated Knowledge Graph of Artificial Intelligence
abstract
Scientific knowledge has been traditionally disseminated and preserved through research articles published in journals, conference proceedings, and online archives. However, this article-centric paradigm has been often criticized for not allowing to automatically process, categorize, and reason on this knowledge. An alternative vision is to generate a semantically rich and interlinked description of the content of research publications. In this paper, we present the Artificial Intelligence Knowledge Graph (AI-KG), a large-scale automatically generated knowledge graph that describes 820K research entities. AI-KG includes about 14M RDF triples and 1.2M reified statements extracted from 333K research publications in the field of AI, and describes 5 types of entities (tasks, methods, metrics, materials, others) linked by 27 relations. AI-KG has been designed to support a variety of intelligent services for analyzing and making sense of research dynamics, supporting researchers in their daily job, and helping to inform decision-making in funding bodies and research policymakers. AI-KG has been generated by applying an automatic pipeline that extracts entities and relationships using three tools: DyGIE++, Stanford CoreNLP, and the CSO Classifier. It then integrates and filters the resulting triples using a combination of deep learning and semantic technologies in order to produce a high-quality knowledge graph. This pipeline was evaluated on a manually crafted gold standard, yielding competitive results. AI-KG is available under CC BY 4.0 and can be downloaded as a dump or queried via a SPARQL endpoint.
Danilo Dessì, Francesco Osborne, Diego Reforgiato Recupero, Davide Buscaldi, Enrico Motta, Harald Sack
ISWC (2)4
2017 Exploring Vector Spaces for Semantic Relations
abstract
Word embeddings are used with success for a variety of tasks involving lexical semantic similarities between individual words.Using unsupervised methods and just cosine similarity, encouraging results were obtained for analogical similarities.In this paper, we explore the potential of pre-trained word embeddings to identify generic types of semantic relations in an unsupervised experiment.We propose a new relational similarity measure based on the combination of word2vec's CBOW input and output vectors which outperforms alternative vector representations, when used for unsupervised clustering on SemEval 2010 Relation Classification data.
Kata Gábor, Haïfa Zargayouna, Isabelle Tellier, Davide Buscaldi, Thierry Charnois
EMNLP4
2016 Event-Based Recognition of Lived Experiences in User Reviews
Ehab Hassan, Davide Buscaldi, Aldo Gangemi
EKAW2
2016 Unsupervised Relation Extraction in Specialized Corpora Using Sequence Mining
Kata Gábor, Haïfa Zargayouna, Isabelle Tellier, Davide Buscaldi, Thierry Charnois
IDA4
2016 Semantic Annotation of the ACL Anthology Corpus for the Automatic Analysis of Scientific Literature
Kata Gábor, Haïfa Zargayouna, Davide Buscaldi, Isabelle Tellier, Thierry Charnois
LREC3
2015 Correlating Open Rating Systems and Event Extraction from Text
Ehab Hassan, Davide Buscaldi, Aldo Gangemi
ICONIP (4)2
2013 Using the Semantics of Texts for Information Retrieval: A Concept- and Domain Relation-Based Approach
Davide Buscaldi, Marie-Noëlle Bessagnet, Albert Royer, Christian Sallaberry
ADBIS (2)1
2013 Effects of Ontology Pitfalls on Ontology-based Information Retrieval Systems
abstract
Nowadays, a growing number of information retrieval systems make use of ontologies to improve the access to textual information, especially in domain-specific scenarios, where the knowledge provided by ontologies represents a key factor. Such kinds of retrieval systems are often referred to as ontology-based or semantic information retrieval systems. The quality of ontologies plays an important role in such systems in the sense that modelling errors in the ontologies may deteriorate the quality of the results obtained by these systems. In this paper we provide a comprehensive analysis of how ontology pitfalls have an influence on these kinds of systems. This study allows us to have a more complete understanding of the role of ontology quality in the information retrieval field. Our survey shows that pitfalls may act as an indicator not only of possible problems in ontology design, but also of OWL features overseen by system developers.
Davide Buscaldi, Mari Carmen Suárez-Figueroa
KEOD1
2013 A semi-automatic approach for building ontologies from acollection of structured web documents
abstract
Many collections of structured documents are available on the web. The collection generally describes the characteristics of entities from a single type, where each page describes one entity. These documents are adequate knowledge sources for building ontologies. As they benefit from a strong and shared layout, they contain less well written text than plain text files but their architecture is very meaningful. Classical linguistic-based methods for identifying concepts and relations are no longer appropriate for analyzing them.The approach we propose in this paper exploits various properties of such documents, combining layout/formatting analysis and linguistic analysis, and using semantic annotation.
Mouna Kamel, Nathalie Aussenac-Gilles, Davide Buscaldi, Catherine Comparot
K-CAP3
2012 From humor recognition to irony detection: The figurative language of social media
Antonio Reyes, Paolo Rosso, Davide Buscaldi
Data Knowl. Eng.3
2011 A pretopological framework for the automatic construction of lexical-semantic structures from texts
abstract
We present in this paper a new approach for the automatic generation of lexical structures from texts. This tedious task is based on the strong hypothesis that simple statistical observations on textual usages can provide pieces of semantics about the lexicon. Using such "naive" observations only, we propose a (pre)-topological framework to formalize and combine various hypothesis on textual data usages and then to derive a structure similar to usual lexical knowledge basis such as WordNet. In addition we also consider the evaluation problem for obtained lexical structures ; a multi-level evaluation strategy is proposed that measures the fitting between a given reference structure and automatically generated structures on different point of views : intrinsic/structural and application-based points of view. The evaluation strategy is then used to quantify the contribution of the new structuring approach with respect to the corresponding solution proposed by (Sanderson et al. 2000) on two case studies that differs on the domain and the size of the lexicon.
Guillaume Cleuziou, Davide Buscaldi, Vincent Levorato, Gaël Dias
CIKM2
2010 Evaluation Protocol and Tools for Question-Answering on Speech Transcripts
Nicolas Moreau, Olivier Hamon, Djamel Mostefa, Sophie Rosset, Olivier Galibert, Lori Lamel, Jordi Turmo, Pere Comas, Paolo Rosso, Davide Buscaldi, Khalid Choukri
LREC10
2010 Answering questions with an n-gram based passage retrieval engine
Davide Buscaldi, Paolo Rosso, José Manuel Gómez Soriano, Emilio Sanchis Arnal
J. Intell. Inf. Syst.1
2009 The Impact of Semantic and Morphosyntactic Ambiguity on Automatic Humour Recognition
Antonio Reyes, Davide Buscaldi, Paolo Rosso
NLDB2
2009 Toponym ambiguity in geographical information retrieval
abstract
The objectives of this research work is to study the effects of toponym (place name) ambiguity in the Geographical Information Retrieval (GIR) task. Our experience with GIR systems shows that toponym ambiguity may be an important factor in the inability of these systems to take advantage from geographical knowledge. Previous studies over ambiguity and Information Retrieval (IR) suggested that disambiguation may be useful in some specific IR scenario. We suppose that GIR may constitute such a scenario. This preliminary study was carried out over the WordNet based, manually disambiguated collection developed for the CLIR-WSD task, using the GeoCLEF collection of 100 geographically related topics. The employed GIR system was based on the GeoWorSE system that participated in GeoCLEF 2008. The experiments were carried out considering the manual disambiguation and comparing this result with those obtained by randomly disambiguating the document collection and those obtained by using always the most common referent. The obtained results show no significant difference in the overall results, although the work gave an insight into some errors that are produced by toponym ambiguity and how they may affect the results. These preliminary results also suggest that WordNet is not a suitable resource for the planned research.
Davide Buscaldi
SIGIR1
2008 Geo-WordNet: Automatic Georeferencing of WordNet
Davide Buscaldi, Paolo Rosso
LREC1
2008 A conceptual density-based approach for the disambiguation of toponyms
abstract
Nowadays, a huge quantity of information is stored in digital format. A great portion of this information is constituted by textual and unstructured documents, where geographical references are usually given by means of place names. A common problem with textual information retrieval is represented by polysemous words, that is, words can have more than one sense. This problem is present also in the geographical domain: place names may refer to different locations in the world. In this paper we investigate the use of our word sense disambiguation technique in the geographical domain, with the aim of resolving ambiguous place names. Our technique is based on WordNet conceptual density. Due to the lack of a reference corpus tagged with WordNet senses, we carried out the experiments over a set of 1,210 place names extracted from the SemCor corpus that we named GeoSemCor and made publicly available. We compared our method with the most‐frequent baseline and the enhanced‐Lesk method, which previously has not been tested in large contexts. The results show that a better precision can be achieved by using a small context (phrase level), whereas a greater coverage can be obtained by using large contexts (document level). The proposed method should be tested with other corpora, due to the fact that our experiments evidenced the excessive bias towards the most‐frequent sense of the GeoSemCor.
Davide Buscaldi, Paolo Rosso
Int. J. Geogr. Inf. Sci.1
2006 Sense Cluster Based Categorization and Clustering of Abstracts
Davide Buscaldi, Paolo Rosso, Mikhail Alexandrov, Alfons Juan-Císcar
CICLing1
2006 Verb Sense Disambiguation Using Support Vector Machines: Impact of WordNet-Extracted Features
Davide Buscaldi, Paolo Rosso, Ferran Plà, Encarna Segarra, Emilio Sanchis Arnal
CICLing1
2006 Mining Knowledge fromWikipedia for the Question Answering task
Davide Buscaldi, Paolo Rosso
LREC1
2006 Spoken QA Based on a Passage Retrieval Engine
abstract
This paper presents a Passage Retrieval-based approach for the development of spoken Question Answering systems. Question Answering can be seen as a particular aspect of Information Retrieval where the user's information need are not satisfied by means of a document, but by a portion of text. Currently, the performance of typical Question Answering systems is rather poor, if compared with other Information Retrieval tasks, however, it is predictable that these system will be used in combination with some speech interface when high precision will be achieved. The most important evaluation competitions for Question Answering, such as TREC and CLEF, still do not have a spoken Question Answering track, therefore this work presents a novel approach to this task in order to study the influence of recognition errors over Question Answering systems.
Emilio Sanchis Arnal, Davide Buscaldi, Sergio Grau, Lluís F. Hurtado, David Griol
SLT2
2005 Context Expansion with Global Keywords for a Conceptual Density-Based WSD
Davide Buscaldi, Paolo Rosso, Manuel Montes-y-Gómez
CICLing1
2005 Two Web-Based Approaches for Noun Sense Disambiguation
Paolo Rosso, Manuel Montes-y-Gómez, Davide Buscaldi, Aarón Pancardo-Rodríguez, Luis Villaseñor-Pineda
CICLing3
2003 Automatic Noun Sense Disambiguation
Paolo Rosso, Francesco Masulli, Davide Buscaldi, Ferran Plà, Antonio Molina
CICLing3