VLDB 2026 Research / reviewers in the wild / expert
Eugenio Lomurno
dblp:297/1123
· DBLP profile ↗
12ranked-venue papers
6as first author
12since 2021 · last 2025
0000-0003-4007-3207ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FPBoost: Fully Parametric Gradient Boosting for Survival AnalysisabstractSurvival analysis is a statistical framework for modeling time-to-event data. It plays a pivotal role in medicine, reliability engineering, and social science research, where understanding event dynamics even with few data samples is critical. Recent advancements in machine learning, particularly those employing neural networks and decision trees, have introduced sophisticated algorithms for survival modeling. However, many of these methods rely on restrictive assumptions about the underlying event-time distribution, such as proportional hazard, time discretization, or accelerated failure time. In this study, we propose FPBoost, a survival model that combines a weighted sum of fully parametric hazard functions with gradient boosting. Distribution parameters are estimated with decision trees trained by maximizing the full survival likelihood. We show how FPBoost is a universal approximator of hazard functions, offering full event-time modeling flexibility while maintaining interpretability through the use of well-established parametric distributions. We evaluate concordance and calibration of FPBoost across multiple benchmark datasets, showcasing its robustness and versatility as a new tool for survival estimation. Alberto Archetti, Eugenio Lomurno, Diego Piccinotti, Matteo Matteucci |
ECAI | 2 |
| 2025 | Your image generator is your new private datasetabstractGenerative diffusion models have emerged as powerful tools to synthetically produce training data, offering potential solutions to data scarcity and reducing labelling costs for downstream supervised deep learning applications. However, existing approaches for synthetic dataset generation face significant limitations: previous methods like Knowledge Recycling rely on label-conditioned generation with models trained from scratch, limiting flexibility and requiring extensive computational resources, while simple class-based conditioning fails to capture the semantic diversity and intra-class variations found in real datasets. Additionally, effectively leveraging text-conditioned image generation for building classifier training sets requires addressing key issues: constructing informative textual prompts, adapting generative models to specific domains, and ensuring robust performance. This paper proposes the Text-Conditioned Knowledge Recycling (TCKR) pipeline to tackle these challenges. TCKR combines dynamic image captioning, parameter-efficient diffusion model fine-tuning, and Generative Knowledge Distillation techniques to create synthetic datasets tailored for image classification. The pipeline is rigorously evaluated on ten diverse image classification benchmarks. The results demonstrate that models trained solely on TCKR-generated data achieve classification accuracies on par with (and in several cases exceeding) models trained on real images. Furthermore, the evaluation reveals that these synthetic-data-trained models exhibit substantially enhanced privacy characteristics: their vulnerability to Membership Inference Attacks is significantly reduced, with the membership inference AUC lowered by 5.49 points on average compared to using real training data, demonstrating a substantial improvement in the performance-privacy trade-off. These findings indicate that high-fidelity synthetic data can effectively replace real data for training classifiers, yielding strong performance whilst simultaneously providing improved privacy protection as a valuable emergent property. The code and trained models are available in the accompanying open-source repository . • Introduces TCKR: text-conditioned diffusion + LoRA + distillation for synthetic data. • Proposes dynamic captioning (BLIP-2) to craft instance-specific prompts for images. • Synthetic-trained classifiers match or surpass real-trained ones on 10 benchmarks. • Synthetic training lowers MIA AUC by up to 8.14 vs models trained on real data. • Increasing synthetic size lifts accuracy but raises MIA risk. Nicolò Francesco Resmini, Eugenio Lomurno, Cristian Sbrolli, Matteo Matteucci |
Image Vis. Comput. | 2 |
| 2025 | Synthetic image learning: Preserving performance and preventing Membership Inference AttacksabstractGenerative artificial intelligence has transformed the generation of synthetic data, providing innovative solutions to challenges like data scarcity and privacy, which are particularly critical in fields such as medicine. However, the effective use of this synthetic data to train high-performance models remains a significant challenge. This paper addresses this issue by introducing Knowledge Recycling (KR), a pipeline designed to optimise the generation and use of synthetic data for training downstream classifiers. At the heart of this pipeline is Generative Knowledge Distillation, the proposed technique that significantly improves the quality and usefulness of the information provided to classifiers through a synthetic dataset regeneration and soft labelling mechanism. The KR pipeline has been tested on a variety of datasets, with a focus on six highly heterogeneous medical image datasets, ranging from retinal images to organ scans. The results show a significant reduction in the performance gap between models trained on real and synthetic data, with models based on synthetic data outperforming those trained on real data in some cases. Furthermore, the resulting models show almost complete immunity to Membership Inference Attacks, manifesting privacy properties missing in models trained with conventional techniques. • Synthetic data can create more private and high-performance image classifiers. • Larger synthetic datasets improve classifier test accuracy. • The Knowledge Recycling pipeline excels on small medical datasets. • Tuning dataset size, generator std and recycling rate enhances classifier performance. • Knowledge Recycling makes models highly resistant to Membership Inference Attacks. Eugenio Lomurno, Matteo Matteucci |
Pattern Recognit. Lett. | 1 |
| 2025 | Federated Knowledge Recycling: Privacy-preserving synthetic data sharingabstractFederated learning has emerged as a paradigm for collaborative learning, enabling the development of robust models without the need to centralise sensitive data. However, conventional federated learning techniques have privacy and security vulnerabilities due to the exposure of models, parameters or updates, which can be exploited as an attack surface. This paper presents Federated Knowledge Recycling (FedKR), a cross-silo federated learning approach that uses locally generated synthetic data to facilitate collaboration between institutions. FedKR combines advanced data generation techniques with a dynamic aggregation process to provide greater security against privacy attacks than existing methods, significantly reducing the attack surface. Experimental results on generic and medical datasets show that FedKR achieves competitive performance, with an average improvement in accuracy of 4.24% compared to training models from local data, demonstrating particular effectiveness in data scarcity scenarios. • FedKR is a synthetic data-based federated learning technique for enhanced privacy. • Synthetic data is exchanged and aggregated locally instead of models or metadata. • Robust against membership inference, model inversion and gradient leakage attacks. • Mitigates various privacy attacks whilst maintaining competitive performance. • Particularly effective in scenarios with limited data availability. Eugenio Lomurno, Matteo Matteucci |
Pattern Recognit. Lett. | 1 |
| 2025 | Age Group Discrimination via Free Handwriting IndicatorsabstractAgeing is associated with cognitive and functional decline, which can hamper daily activities and independent living. Chronic diseases may intensify this process. Early detection of unhealthy decline is key but hindered by similarity to normal ageing. This study presents an approach for early screening of healthy ageing, using an instrumented ink pen to ecologically assess handwriting performance in three age groups: 40-59, 60-69 and 70+ years old. Raw handwriting data from 60 healthy subjects were used to extract fourteen indicators related to gesture and tremor. The indicators were then used to discriminate between subjects of different age groups in three binary classification tasks, using machine learning algorithms. This approach produced remarkable results, particularly in identifying subjects at the very beginning of the ageing process (Group 2) from elderly subjects (Group 3), achieving an accuracy of 97.5%, an F1 score of 97.44% and a ROC-AUC of 95%. Analysis of the Shapley values revealed age-dependent sensitivity of handwriting and tremor-related indicators. The proposed method represents a promising solution for early detection of abnormal signs of ageing, designed for remote, non-invasive, unsupervised home monitoring to improve the care of older adults. Eugenio Lomurno, Simone Toffoli, Davide Di Febbo, Matteo Matteucci, Francesca Lunardini, Simona Ferrante |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | Harnessing the Computing Continuum Across Personalized Healthcare, Maintenance and Inspection, and Farming 4.0abstractThe AI-SPRINT project, launched in 2021 and funded by the European Commission, focuses on the development and implementation of AI applications across the computing continuum. This continuum ensures the coherent integration of computational resources and services from centralized data centers to edge devices, facilitating efficient and adaptive computation and application delivery. AI-SPRINT has achieved significant scientific advances, including streamlined processes, improved efficiency, and the ability to operate in real time, as evidenced by three practical use cases. This paper provides an in-depth examination of these applications – Personalized Healthcare, Maintenance and Inspection, and Farming 4.0 – highlighting their practical implementation and the objectives achieved with the integration of AI-SPRINT technologies. We analyze how the proposed toolchain effectively addresses a range of challenges and refines processes, discussing its relevance and impact in multiple domains. After a comprehensive overview of the main AI-SPRINT tools used in these scenarios, the paper summarizes of the findings and key lessons learned. Fatemeh Baghdadi, Davide Cirillo, Daniele Lezzi, Francesc Lordan, Fernando Vázquez, Eugenio Lomurno, Alberto Archetti, Danilo Ardagna, Matteo Matteucci |
CLOSER | 6 |
| 2024 | Stable Diffusion Dataset Generation for Downstream Classification TasksabstractRecent advances in generative artificial intelligence have enabled the creation of high-quality synthetic data that closely mimics real-world data.This paper explores the adaptation of the Stable Diffusion 2.0 model for generating synthetic datasets, using Transfer Learning, Fine-Tuning and generation parameter optimisation techniques to improve the utility of the dataset for downstream classification tasks.We present a class-conditional version of the model that exploits a Class-Encoder and optimisation of key generation parameters.Our methodology led to synthetic datasets that, in a third of cases, produced models that outperformed those trained on real datasets. * This paper is supported by the FAIR (Future Artificial Intelligence Research) project, funded by the NextGenerationEU program within the PNRR-PE-AI scheme (M4C2, investment 1.3, line on Artificial Intelligence).1 The authors have contributed in equal measure. Eugenio Lomurno, Matteo D'Oria, Matteo Matteucci |
ESANN | 1 |
| 2024 | An Efficient Neural Architecture Search Model for Medical Image ClassificationabstractAccurate classification of medical images is essential for modern diagnostics.Deep learning advancements led clinicians to increasingly use sophisticated models to make faster and more accurate decisions, sometimes replacing human judgment.However, model development is costly and repetitive.Neural Architecture Search (NAS) provides solutions by automating the design of deep learning architectures.This paper presents ZO-DARTS+, a differentiable NAS algorithm that improves search efficiency through a novel method of generating sparse probabilities by bilevel optimization.Experiments on five public medical datasets show that ZO-DARTS+ matches the accuracy of state-of-the-art solutions while reducing search times by up to three times. Lunchen Xie, Eugenio Lomurno, Matteo Gambella, Danilo Ardagna, Manuel Roveri, Matteo Matteucci, Qingjiang Shi |
ESANN | 2 |
| 2023 | Discriminative Adversarial Privacy: Balancing Accuracy and Membership Privacy in Neural Networks
Eugenio Lomurno, Alberto Archetti, Francesca Ausonio, Matteo Matteucci |
BMVC | 1 |
| 2023 | Bridging the Gap: Enhancing the Utility of Synthetic Data via Post-Processing Techniques
Eugenio Lomurno, Andrea Lampis, Matteo Matteucci |
BMVC | 1 |
| 2023 | Enhancing Once-For-All: A Study on Parallel Blocks, Skip Connections and Early ExitsabstractThe use of Neural Architecture Search (NAS) techniques to automate the design of neural networks has become increasingly popular in recent years. The proliferation of devices with different hardware characteristics using such neural networks, as well as the need to reduce the power consumption for their search, has led to the realisation of Once-For-All (OFA), an algorithm engineered for low CO2 consumption and characterised by the ability to generate easily adaptable models through a single learning process. In order to improve this paradigm and develop high-performance yet eco-friendly NAS techniques, this paper presents OFAv2, the extension of OFA aimed at improving its performance while maintaining the same ecological advantage. The algorithm is improved from an architectural point of view by including early exits, parallel blocks and dense skip connections. The training process is extended by two new steps called Elastic Level and Elastic Exit. A new Knowledge Distillation technique is presented to handle multi-output networks, and finally a new strategy for dynamic teacher network selection is proposed. These modifications allow OFAv2 to improve its accuracy performance on the Tiny ImageNet dataset by up to 12.07% compared to the original version of OFA, while maintaining the algorithm flexibility and advantages. Simone Sarti, Eugenio Lomurno, Andrea Falanti, Matteo Matteucci |
IJCNN | 2 |
| 2022 | POPNASv2: An Efficient Multi-Objective Neural Architecture Search TechniqueabstractAutomating the research for the best neural network model is a task that has gained more and more relevance in the last few years. In this context, Neural Architecture Search (NAS) represents the most effective technique whose results rival the state of the art hand-crafted architectures. However, this approach requires a lot of computational capabilities as well as research time, which make prohibitive its usage in many real-world scenarios. With its sequential model-based optimization strategy, Progressive Neural Architecture Search (PNAS) represents a possible step forward to face this resources issue. Despite the quality of the found network architectures, this technique is still limited in research time. A significant step in this direction has been done by Pareto-Optimal Progressive Neural Architecture Search (POPNAS), which expand PNAS with a time predictor to enable a trade-off between search time and accuracy, considering a multi-objective optimization problem. This paper proposes a new version of the Pareto-Optimal Progressive Neural Architecture Search, called POPNASv2. Our approach enhances its first version and improves its performance. We expanded the search space by adding new operators and improved the quality of both predictors to build more accurate Pareto fronts. Moreover, we introduced cell equivalence checks and enriched the search strategy with an adaptive greedy exploration step. Our efforts allow POPNASv2 to achieve PNAS-like performance with an average 4x factor search time speed-up. Code: https://doi.org/10.5281/zenodo.6574040 Andrea Falanti, Eugenio Lomurno, Stefano Samele, Danilo Ardagna, Matteo Matteucci |
IJCNN | 2 |