Vincenzo Della Mea

dblp:m/VincenzoDellaMea · DBLP profile ↗
← Back
19ranked-venue papers
7as first author
6since 2021 · last 2025
0000-0002-0144-3802ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 5 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Computer networks · 2 · 1 first-author
YearPublicationVenuePosition
2025 On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
abstract
Large Language Models (LLMs) effectiveness is usually evaluated by means of benchmarks such as MMLU, ARC-C, or HellaSwag, where questions are presented in their original wording, thus in a fixed, standardized format. However, real-world applications involve linguistic variability, requiring models to maintain their effectiveness across diverse rewordings of the same question or query. In this study, we systematically assess the robustness of LLMs to paraphrased benchmark questions and investigate whether benchmark-based evaluations provide a reliable measure of model capabilities. We systematically generate various paraphrases of all the questions across six different common benchmarks, and measure the resulting variations in effectiveness of 34 state-of-the-art LLMs, of different size and effectiveness. Our findings reveal that while LLM rankings remain relatively stable across paraphrased inputs, absolute effectiveness scores change, and decline significantly. This suggests that LLMs struggle with linguistic variability, raising concerns about their generalization abilities and evaluation methodologies. Furthermore, the observed performance drop challenges the reliability of benchmark-based evaluations, indicating that high benchmark scores may not fully capture a model’s robustness to real-world input variations. We discuss the implications of these findings for LLM evaluation methodologies, emphasizing the need for robustness-aware benchmarks that better reflect practical deployment scenarios.
Riccardo Lunardi, Vincenzo Della Mea, Stefano Mizzaro, Kevin Roitero
ECAI2
2025 Leveraging LLMs for Energy Forecasting: The AcegasApsAmga Case Study
Kevin Roitero, Andrea Zancola, Vincenzo Della Mea, Stefano Mizzaro
ECIR (5)3
2025 PILs of Knowledge: A Synthetic Benchmark for Evaluating Question Answering Systems in Healthcare
abstract
Patient Information Leaflets (PILs) provide essential information about medication usage, side effects, precautions, and interactions, making them a valuable resource for Question Answering (QA) systems in healthcare. However, no dedicated benchmark currently exists to evaluate QA systems specifically on PILs, limiting progress in this domain. To address this gap, we introduce a fact-supported synthetic benchmark composed of multiple-choice questions and answers generated from real PILs. We construct the benchmark using a fully automated pipeline that leverages multiple Large Language Models (LLMs) to generate diverse, realistic, and contextually relevant question-answer pairs. The benchmark is publicly released as a standardized evaluation framework for assessing the ability of LLMs to process and reason over PIL content. To validate its effectiveness, we conduct an initial evaluation with state-of-the-art LLMs, showing that the benchmark presents a realistic and challenging task, making it a valuable resource for advancing QA research in the healthcare domain.
Riccardo Lunardi, Michael Soprano, Paolo Coppola 0001, Vincenzo Della Mea, Stefano Mizzaro, Kevin Roitero
SIGIR4
2025 SNOMED CT entity linking challenge
abstract
OBJECTIVE: This paper presents the results from a competition challenging participants to develop entity linking models using a subset of annotated MIMIC-IV-Note data and the SNOMED CT Terminology. MATERIALS AND METHODS: As a basis for this work, a large set of 74 808 annotations was curated across 272 discharge notes spanning 6624 unique clinical concepts. Submissions were evaluated using the mean Intersection-over-Union metric, evaluated at the character level with the 3 best performing solutions awarded a cash prize. RESULTS: The winning solutions employed contrasting approaches: a dictionary-based method, an encoder-based method, and a decoder-based method. DISCUSSION: Our analysis reveals that concept frequency in training data significantly impacts model performance, with rare concepts proving particularly challenging. High concept entropy and annotation ambiguity were also associated with decreased performance. CONCLUSION: Findings from this work suggest that future projects should focus on improving entity linking for rare concepts and developing methods to better leverage contextual information when training examples are scarce.
Rory Davidson, Will Hardman, Guy Amit, Yonatan Bilu, Vincenzo Della Mea, Aleksandr Galaida, Irena Girshovitz, Mikhail Kulyabin, Mihai Horia Popescu, Kevin Roitero, Gleb Sokolov, Chen Yanover
J. Am. Medical Informatics Assoc.5
2024 Generative AI for Energy: Multi-Horizon Power Consumption Forecasting using Large Language Models
abstract
We leverage generative NLP-based models, specifically Transformer-Based models, for multi-horizon univariate and multivariate power consumption forecasting. We apply our approach to various datasets, focusing on short-term (1 day) and long-term (1 week) forecasts. We test several lag configurations with and without additional contextual information and achieve promising results. We evaluate the forecasts' effectiveness using a range of metrics, and aggregate the results on a monthly basis for a comprehensive understanding of the performance throughout the year.
Kevin Roitero, Gianluca D'Abrosca, Andrea Zancola, Vincenzo Della Mea, Stefano Mizzaro
CIKM4
2023 Can the crowd judge truthfulness? A longitudinal study on recent misinformation about COVID-19
abstract
Recently, the misinformation problem has been addressed with a crowdsourcing-based approach: to assess the truthfulness of a statement, instead of relying on a few experts, a crowd of non-expert is exploited. We study whether crowdsourcing is an effective and reliable method to assess truthfulness during a pandemic, targeting statements related to COVID-19, thus addressing (mis)information that is both related to a sensitive and personal issue and very recent as compared to when the judgment is done. In our experiments, crowd workers are asked to assess the truthfulness of statements, and to provide evidence for the assessments. Besides showing that the crowd is able to accurately judge the truthfulness of the statements, we report results on workers' behavior, agreement among workers, effect of aggregation functions, of scales transformations, and of workers background and bias. We perform a longitudinal study by re-launching the task multiple times with both novice and experienced workers, deriving important insights on how the behavior and quality change over time. Our results show that workers are able to detect and objectively categorize online (mis)information related to COVID-19; both crowdsourced and expert judgments can be transformed and aggregated to improve quality; worker background and other signals (e.g., source of information, behavior) impact the quality of the data. The longitudinal study demonstrates that the time-span has a major effect on the quality of the judgments, for both novice and experienced workers. Finally, we provide an extensive failure analysis of the statements misjudged by the crowd-workers.
Kevin Roitero, Michael Soprano, Beatrice Portelli, Massimiliano De Luise, Damiano Spina, Vincenzo Della Mea, Giuseppe Serra 0001, Stefano Mizzaro, Gianluca Demartini
Pers. Ubiquitous Comput.6
2020 Toward a Harmonized WHO Family of International Classifications Content Model
Samson W. Tu, Csongor Nyulas, Tania Tudorache, Mark A. Musen, Andrea Martinuzzi, Coen H. van Gool, Vincenzo Della Mea, Christopher G. Chute, Lucilla Frattura, Nicholas R. Hardiker, Huib ten Napel, Richard Madden, Ann-Helene Almborg, Jeewani Anupama Ginige, Catherine Sykes, Can Çelik, Robert Jakob
AMIA7
2020 The COVID-19 Infodemic: Can the Crowd Judge Recent Misinformation Objectively?
abstract
Misinformation is an ever increasing problem that is difficult to solve for the research community and has a negative impact on the society at large. Very recently, the problem has been addressed with a crowdsourcing-based approach to scale up labeling efforts: to assess the truthfulness of a statement, instead of relying on a few experts, a crowd of (non-expert) judges is exploited. We follow the same approach to study whether crowdsourcing is an effective and reliable method to assess statements truthfulness during a pandemic. We specifically target statements related to the COVID-19 health emergency, that is still ongoing at the time of the study and has arguably caused an increase of the amount of misinformation that is spreading online (a phenomenon for which the term "infodemic" has been used). By doing so, we are able to address (mis)information that is both related to a sensitive and personal issue like health and very recent as compared to when the judgment is done: two issues that have not been analyzed in related work.\n\nIn our experiment, crowd workers are asked to assess the truthfulness of statements, as well as to provide evidence for the assessments as a URL and a text justification. Besides showing that the crowd is able to accurately judge the truthfulness of the statements, we also report results on many different aspects, including: agreement among workers, the effect of different aggregation functions, of scales transformations, and of workers background / bias. We also analyze workers behavior, in terms of queries submitted, URLs found / selected, text justifications, and other behavioral data like clicks and mouse actions collected by means of an ad hoc logger.
Kevin Roitero, Michael Soprano, Beatrice Portelli, Damiano Spina, Vincenzo Della Mea, Giuseppe Serra 0001, Stefano Mizzaro, Gianluca Demartini
CIKM5
2020 Marker controlled superpixel nuclei segmentation and automatic counting on immunohistochemistry staining images
abstract
MOTIVATION: For the diagnosis of cancer, manually counting nuclei on massive histopathological images is tedious and the counting results might vary due to the subjective nature of the operation. RESULTS: This paper presents a new segmentation and counting method for nuclei, which can automatically provide nucleus counting results. This method segments nuclei with detected nuclei seed markers through a modified simple one-pass superpixel segmentation method. Rather than using a single pixel as a seed, we created a superseed for each nucleus to involve more information for improved segmentation results. Nucleus pixels are extracted by a newly proposed fusing method to reduce stain variations and preserve nucleus contour information. By evaluating segmentation results, the proposed method was compared to five existing methods on a dataset with 52 immunohistochemically (IHC) stained images. Our proposed method produced the highest mean F1-score of 0.668. By evaluating the counting results, another dataset with more than 30 000 IHC stained nuclei in 88 images were prepared. The correlation between automatically generated nucleus counting results and manual nucleus counting results was up to R2 = 0.901 (P < 0.001). By evaluating segmentation results of proposed method-based tool, we tested on a 2018 Data Science Bowl (DSB) competition dataset, three users obtained DSB score of 0.331 ± 0.006. AVAILABILITY AND IMPLEMENTATION: The proposed method has been implemented as a plugin tool in ImageJ and the source code can be freely downloaded. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jie Shu, Jingxin Liu 0005, Yongmei Zhang, Hao Fu 0001, Mohammad Ilyas, Giuseppe Faraci, Vincenzo Della Mea, Guoping Qiu
Bioinform.7
2019 Towards the Development of a Web Support System for Improving Accuracy in Coding Discharge Diagnosis
abstract
The practice of coding morbidity data using international standard diagnostic classifications has become increasingly important and recognized as a difficult and time-consuming task. Clinical coders/physicians assign codes to each patient episode based on their interpretation of the available case notes or the electronic patient record systems. Therefore, accurate coding depends on the legibility of the case notes and on the coders` understanding of medical terminology. Several studies have indicated poor reproducibility of clinical coding. To support physicians, in the coding of diagnoses and procedures using ICD 9th revision, Clinical Modifications (ICD-9-CM), and in identifying the main condition to be filled in Hospital Discharge Records, the SISCO.web service is proposed. The service has the aim of improving the accuracy of coding by using the combination of NLP algorithms, controlled vocabularies mapped to ICD-9-CM, as well as coding-rules and a decision trees for the identification of the main condition.
Elena Cardillo, Lucilla Frattura, Salvatore Ciambrini, Claudio Eccher, Elia Nardo, Carlo Zavaroni, Vincenzo Della Mea
ISCC7
2015 Mobile crowdsourcing: four experiments on platforms and tasks
Vincenzo Della Mea, Eddy Maddalena, Stefano Mizzaro
Distributed Parallel Databases1
2008 AI on the Move: Exploiting AI Techniques for Context Inference on Mobile Devices
abstract
Context aware computing is a computational paradigm that has faced a rapid growth in the last few years, especially in the field of mobile devices. One of the promises of context-awareness in this field is the possibility of automatically adapting the functioning mode of mobile devices to the environment and the current situation the user is in, with the aim of improving both their efficiency (using the scarce resources in a more efficient way) and effectiveness (providing better services to the user). We propose a novel approach for providing a basic infrastructure for context-aware applications on mobile devices, in which AI techniques (namely a principled combination of rule-based systems, Bayesian networks, and ontologies) are applied to context inference. The aim is to devise a general inferential framework to easier the development of context-aware applications by integrating the information coming from physical and logical sensors (e.g., position, agenda) and reasoning about this information in order to infer new and more abstract contexts. In previous contextaware applications, most researches focused almost exclusively on time and/or location and other few data, while the same contexts inference was limited to preconceived values. Our approach differs from previous works since we do not focus on particular contextual values, but rather we have developed an architecture where managed contexts can be easily replaced by new contexts, depending on the different needs. Moreover, the inferential infrastructure we designed is able to work in a more general way and can be easily adapted to different models of applications distribution. We show some concrete examples of applications built upon the inferential infrastructure and we discuss its strengths and limitations.
Adolfo Bulfoni, Paolo Coppola 0001, Vincenzo Della Mea, Luca Di Gaspero, Danny Mischis, Stefano Mizzaro, Ivan Scagnetto, Luca Vassena
ECAI3
2006 Experiments on Average Distance Measure
Vincenzo Della Mea, Gianluca Demartini, Luca Di Gaspero, Stefano Mizzaro
ECIR1
2005 Neuroimaging Services on the Net
abstract
The present paper reports a research project aimed at developing a cluster of clinical Institutions to share data and expertise in the fields of computational neuroanatomy. In particular, our first results concern a network system to transfer magnetic resonance (MR) images, explore and use Voxel-Based Morphometry (VBM)-based reports in four remote clinical settings. The cluster is composed of a central unit for the recording and processing of MR images, and peripheral federated units. The former is LENITEM -Lab of Epidemiology, Neuroimaging and TEleMedicine- Brescia, Italy; the latter are clinical units usually treating patients with neurodegenerative diseases. The cluster is meant to be grounded on a network infrastructure, providing services to the units: the first service to be designed and realised is VBM. The peripheral (client) units transfer MR images to the central one; here they are processed and analysed; a VBM-based report is made available, within a few hours, to the client unit. Neuroimaging techniques give rise to network services provided on demand.
Vito Roberto, Alessandro Zappia, Cristina Testa, Vincenzo Della Mea, Giovanni B. Frisoni
CBMS4
2004 Measuring retrieval effectiveness: A new proposal and a first experimental validation
abstract
Abstract Most common effectiveness measures for information retrieval systems are based on the assumptions of binary relevance (either a document is relevant to a given query or it is not) and binary retrieval (either a document is retrieved or it is not). In this article, these assumptions are questioned, and a new measure named ADM (average distance measure) is proposed, discussed from a conceptual point of view, and experimentally validated on Text Retrieval Conference (TREC) data. Both conceptual analysis and experimental evidence demonstrate ADM's adequacy in measuring the effectiveness of information retrieval systems. Some potential problems about precision and recall are also highlighted and discussed.
Vincenzo Della Mea, Stefano Mizzaro
J. Assoc. Inf. Sci. Technol.1
2001 Visualization Issues in Telepathology: The Role of the Internet Imaging Protocol
abstract
Image visualization may become a difficult task in telemedicine, especially when large data sets are to be delivered through the Internet and selectively analyzed by the physician. A new protocol - the Internet Imaging Protocol, IIP - has been recently proposed by a consortium of companies, with the aim of distributing efficiently multiresolution images on the WWW, for both visualization and printing purposes. The proposed protocol is indeed sufficiently general to be considered for a wide range of applications. The present paper analyses it from the point of view of telepathology, i.e., the transmission of medical images coming from microscopes. We argue that pathology images can be delivered by means of the IIP, with advantages on the visualization interface, which closely reproduces the direct use of a microscope. A sample application is presented, aimed at delivering pathology cases for second opinion diagnosis as well as for continuing medical education.
Vincenzo Della Mea, Vito Roberto, Carlo Alberto Beltrami
IV1
2001 Agents acting and moving in healthcare scenario - a paradigm for telemedical collaboration
abstract
This communications describes a novel approach to the analysis and development of telemedicine systems, based on the multiagent paradigm. An agent is an autonomous, social, reactive, and proactive entity, sometimes also mobile. Since telemedicine is grounded on communication and sharing of resources, agents are suitable for its analysis and implementation, and we adopted them for developing a prototype telemedical agent.
Vincenzo Della Mea
IEEE Trans. Inf. Technol. Biomed.1
1996 HTML Generation and Semantic Markup for Telepathology
Vincenzo Della Mea, Carlo Alberto Beltrami, Vito Roberto, Davide Brunato
Comput. Networks1
1995 A Graph-Based Approach to the Structural Analysis of Proliferative Breast Lesions
Vincenzo Della Mea, Nicoletta Finato, Carlo Alberto Beltrami
AIME1