EDBT 2026 Demo / reviewers in the wild / expert
Niccolò Marini
dblp:275/9056
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0002-5273-5741ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Patch-Based Reconstruction and Multimodal Residual Learning for Generalized Deepfake DetectionabstractDeepfake detectors often achieve near-perfect accuracy in-domain but fail under dataset or forgery shifts, especially when only a few manipulated samples are available for training. We propose a two-stage framework for generalized and data-efficient deepfake detection based on reconstruction residuals. In Stage I, we learn a real-face prior with PM-VAE, a masked patch reconstructor that augments a Masked Autoencoder with a lightweight variational bottleneck to regularize the patch latent space and reduce memorization. In Stage II, the generator is frozen and used to produce forensic evidence from partially observed inputs via block-wise masking on the patch grid, the resulting inpainted reconstructions yield residual cues that are stable and localized. We then train a multi-branch Transformer to fuse (i) RGB context, (ii) spatial residuals, and (iii) wavelet-domain residuals that explicitly capture high-frequency inconsistencies missed by spatial errors alone. Extensive experiments on FaceForensics++ dataset under cross-forgery protocols, on external benchmarks (Celeb-DF, DFD, DFDC) for cross-dataset evaluation, and on synthetic generation (StyleGAN family and Stable Diffusion) show improved robustness over reconstruction baselines, with consistent gains in low-data regimes down to a handful of fake frames per video. Niccolò Marini, Andrea Ciamarra, Roberto Caldelli, Stefano Berretti |
IH&MMSec | 1 |
| 2026 | Mitigating hallucinations in synthesized clinical texts to improve multimodal deep learning for dermatologyabstract• An investigation into the effects of pairing synthesized clinical notes with image data to train a multimodal AI algorithm, using dermatology as example problem domain. • Leveraging metadata information to drive clinical note synthesis reduces hallucinations in Large Language Model (LLM) outputs. • Clinical notes generated by different LLMs using metadata lead to similar performance on downstream tasks when paired with real dermatology images. • Combination of multimodal data improves generalization performance on external datasets. Despite recent advancements in the development of foundation models and multimodal (MM) architectures in dermatology, their translation to clinical practice remains limited by the scarcity of large-scale multimodal (MM) datasets, as most publicly available resources are small, unimodal, and lack expressive clinical text. This paper investigates strategies to synthesize and exploit clinical notes paired with dermatological images to effectively train a MM architecture, focusing on solutions to limit the inherently hallucinated contents introduced by Large Language Models (LLMs) and to identify conditions under which synthetic clinical notes can be reliably leveraged. The paper proposes a MM architecture trained on real dermatological images paired with LLM-synthesized clinical notes. We systematically evaluate different note generation strategies, including metadata-guided prompting, alignment of image representations with specific keywords, sentence-level filtering of clinical notes, network architectural designs. Experiments involve 16,000 image-note couples collected from six public datasets for model training and over 37,000 images from fifteen public datasets as external data for generalization assessment. Performance is assessed on cross-modal retrieval and zero-shot learning tasks to quantify robustness and generalization. Results show that metadata inclusion into the prompts reduces the hallucinations within LLM outputs, providing more reliable notes. The resulting MM model trained with these notes show superior performance on multiple downstream tasks. Synthesized clinical notes can be paired with real dermatology images under specific conditions, providing a valuable resource to develop foundation models that can help reduce the dermatologists’ workload. Niccolò Marini, Zhaohui Liang, Sivaramakrishnan Rajaraman, Zhiyun Xue, Sameer K. Antani |
J. Biomed. Informatics | 1 |
| 2025 | The Hidden Threat of Hallucinations in Binary Chest X-Ray Pneumonia ClassificationabstractHallucination in deep learning (DL) classification, where DL models yield confidently erroneous predictions remains a pressing concern. This study investigates whether binary classifiers are truly learning disease-specific features when distinguishing overlapping radiological presentations among pneumonia subtypes on chest X-ray (CXR) images. Specifically, we evaluate if uncertainty measure is a valuable tool in classifying signs of different pathogen-specific subtypes of pneumonia. We evaluated two binary classifiers to classify bacterial pneumonia and viral pneumonia, respectively, from normal CXRs. A third classifier explored the ability to distinguish bacterial from viral pneumonia presentation to highlight our concern regarding the observed hallucinations in the former cases. Our comprehensive analysis computes the Matthews Correlation Coefficient and prediction entropy metrics on a pediatric CXR dataset and reveals that the normal/bacterial and normal/viral classifiers consistently and confidently misclassify the unseen pneumonia subtype to their respective disease class. These findings expose a critical limitation concerning the tendency of binary classifiers to hallucinate by relying on general pneumonia indicators rather than pathogen-specific patterns, thereby challenging their utility in clinical workflows. Sivaramakrishnan Rajaraman, Zhaohui Liang, Niccolò Marini, Zhiyun Xue, Sameer K. Antani |
CBMS | 3 |
| 2025 | FRED: The Florence RGB-Event Drone DatasetabstractSmall, fast, and lightweight drones present significant challenges for traditional RGB cameras due to their limitations in capturing fast-moving objects, especially under challenging lighting conditions. Event cameras offer an ideal solution, providing high temporal definition and dynamic range, yet existing benchmarks often lack fine temporal resolution or drone-specific motion patterns, hindering progress in these areas. This paper introduces the Florence RGB-Event Drone dataset (FRED), a novel multimodal dataset specifically designed for drone detection, tracking, and trajectory forecasting, combining RGB video and event streams. FRED features more than 7 hours of densely annotated drone trajectories, using 5 different drone models and including challenging scenarios such as rain and adverse lighting conditions. We provide detailed evaluation protocols and standard metrics for each task, facilitating reproducible benchmarking. The authors hope FRED will advance research in high-speed drone perception and multimodal spatiotemporal understanding. Gabriele Magrini, Niccolò Marini, Federico Becattini, Lorenzo Berlincioni, Niccolò Biondi, Pietro Pala, Alberto Del Bimbo |
ACM Multimedia | 2 |
| 2024 | A systematic comparison of deep learning methods for Gleason grading and scoringabstractProstate cancer is the second most frequent cancer in men worldwide after lung cancer. Its diagnosis is based on the identification of the Gleason score that evaluates the abnormality of cells in glands through the analysis of the different Gleason patterns within tissue samples. The recent advancements in computational pathology, a domain aiming at developing algorithms to automatically analyze digitized histopathology images, lead to a large variety and availability of datasets and algorithms for Gleason grading and scoring. However, there is no clear consensus on which methods are best suited for each problem in relation to the characteristics of data and labels. This paper provides a systematic comparison on nine datasets with state-of-the-art training approaches for deep neural networks (including fully-supervised learning, weakly-supervised learning, semi-supervised learning, Additive-MIL, Attention-Based MIL, Dual-Stream MIL, TransMIL and CLAM) applied to Gleason grading and scoring tasks. The nine datasets are collected from pathology institutes and openly accessible repositories. The results show that the best methods for Gleason grading and Gleason scoring tasks are fully supervised learning and CLAM, respectively, guiding researchers to the best practice to adopt depending on the task to solve and the labels that are available. Juan Pedro Dominguez-Morales, Lourdes Duran-Lopez, Niccolò Marini, Saturnino Vicente Diaz, Alejandro Linares-Barranco, Manfredo Atzori, Henning Müller |
Medical Image Anal. | 3 |
| 2024 | Multimodal representations of biomedical knowledge from limited training whole slide images and reports using deep learningabstractThe increasing availability of biomedical data creates valuable resources for developing new deep learning algorithms to support experts, especially in domains where collecting large volumes of annotated data is not trivial. Biomedical data include several modalities containing complementary information, such as medical images and reports: images are often large and encode low-level information, while reports include a summarized high-level description of the findings identified within data and often only concerning a small part of the image. However, only a few methods allow to effectively link the visual content of images with the textual content of reports, preventing medical specialists from properly benefitting from the recent opportunities offered by deep learning models. This paper introduces a multimodal architecture creating a robust biomedical data representation encoding fine-grained text representations within image embeddings. The architecture aims to tackle data scarcity (combining supervised and self-supervised learning) and to create multimodal biomedical ontologies. The architecture is trained on over 6,000 colon whole slide Images (WSI), paired with the corresponding report, collected from two digital pathology workflows. The evaluation of the multimodal architecture involves three tasks: WSI classification (on data from pathology workflow and from public repositories), multimodal data retrieval, and linking between textual and visual concepts. Noticeably, the latter two tasks are available by architectural design without further training, showing that the multimodal architecture that can be adopted as a backbone to solve peculiar tasks. The multimodal data representation outperforms the unimodal one on the classification of colon WSIs and allows to halve the data needed to reach accurate performance, reducing the computational power required and thus the carbon footprint. The combination of images and reports exploiting self-supervised algorithms allows to mine databases without needing new annotations provided by experts, extracting new information. In particular, the multimodal visual ontology, linking semantic concepts to images, may pave the way to advancements in medicine and biomedical analysis domains, not limited to histopathology. Niccolò Marini, Stefano Marchesin 0001, Marek Wodzinski, Alessandro Caputo, Damian Podareanu, Bryan Cardenas Guevara, Svetla Boytcheva, Simona Vatrano, Filippo Fraggetta, Francesco Ciompi, Gianmaria Silvello, Henning Müller, Manfredo Atzori |
Medical Image Anal. | 1 |
| 2024 | The ACROBAT 2022 challenge: Automatic registration of breast cancer tissueabstractThe alignment of tissue between histopathological whole-slide-images (WSI) is crucial for research and clinical applications. Advances in computing, deep learning, and availability of large WSI datasets have revolutionised WSI analysis. Therefore, the current state-of-the-art in WSI registration is unclear. To address this, we conducted the ACROBAT challenge, based on the largest WSI registration dataset to date, including 4,212 WSIs from 1,152 breast cancer patients. The challenge objective was to align WSIs of tissue that was stained with routine diagnostic immunohistochemistry to its H&E-stained counterpart. We compare the performance of eight WSI registration algorithms, including an investigation of the impact of different WSI properties and clinical covariates. We find that conceptually distinct WSI registration methods can lead to highly accurate registration performances and identify covariates that impact performances across methods. These results provide a comparison of the performance of current WSI registration methods and guide researchers in selecting and developing methods. Philippe Weitz, Masi Valkonen, Leslie Solorzano, Circe Carr, Kimmo Kartasalo, Constance Boissin, Sonja Koivukoski, Aino Kuusela, Dusan Rasic, Yanbo Feng, Sandra Kristiane Sinius Pouplier, Kajsa Ledesma Eriksson, Stephanie Robertson, Christian Marzahl, Chandler Gatenbee, Alexander R. A. Anderson, Marek Wodzinski, Artur Jurgas, Niccolò Marini, Manfredo Atzori, Henning Müller, Daniel Budelmann, Nick Weiss, Stefan Heldmann, Johannes Lotz 0002, Jelmer M. Wolterink, Bruno De Santi, Abhijeet Patil, Amit Sethi, Satoshi Kondo, Satoshi Kasai, Kousuke Hirasawa, Mahtab Farrokh, Neeraj Kumar 0002, Russell Greiner, Leena Latonen, Anne-Vibeke Laenkholm, Johan Hartman, Pekka Ruusuvuori, Mattias Rantalainen |
Medical Image Anal. | 20 |
| 2021 | Semi-supervised training of deep convolutional neural networks with heterogeneous data and few local annotations: An experiment on prostate histopathology image classificationabstractConvolutional neural networks (CNNs) are state-of-the-art computer vision techniques for various tasks, particularly for image classification. However, there are domains where the training of classification models that generalize on several datasets is still an open challenge because of the highly heterogeneous data and the lack of large datasets with local annotations of the regions of interest, such as histopathology image analysis. Histopathology concerns the microscopic analysis of tissue specimens processed in glass slides to identify diseases such as cancer. Digital pathology concerns the acquisition, management and automatic analysis of digitized histopathology images that are large, having in the order of 100′0002 pixels per image. Digital histopathology images are highly heterogeneous due to the variability of the image acquisition procedures. Creating locally labeled regions (required for the training) is time-consuming and often expensive in the medical field, as physicians usually have to annotate the data. Despite the advances in deep learning, leveraging strongly and weakly annotated datasets to train classification models is still an unsolved problem, mainly when data are very heterogeneous. Large amounts of data are needed to create models that generalize well. This paper presents a novel approach to train CNNs that generalize to heterogeneous datasets originating from various sources and without local annotations. The data analysis pipeline targets Gleason grading on prostate images and includes two models in sequence, following a teacher/student training paradigm. The teacher model (a high-capacity neural network) automatically annotates a set of pseudo-labeled patches used to train the student model (a smaller network). The two models are trained with two different teacher/student approaches: semi-supervised learning and semi-weekly supervised learning. For each of the two approaches, three student training variants are presented. The baseline is provided by training the student model only with the strongly annotated data. Classification performance is evaluated on the student model at the patch level (using the local annotations of the Tissue Micro-Arrays Zurich dataset) and at the global level (using the TCGA-PRAD, The Cancer Genome Atlas-PRostate ADenocarcinoma, whole slide image Gleason score). The teacher/student paradigm allows the models to better generalize on both datasets, despite the inter-dataset heterogeneity and the small number of local annotations used. The classification performance is improved both at the patch-level (up to κ=0.6127±0.0133 from κ=0.5667±0.0285), at the TMA core-level (Gleason score) (up to κ=0.7645±0.0231 from κ=0.7186±0.0306) and at the WSI-level (Gleason score) (up to κ=0.4529±0.0512 from κ=0.2293±0.1350). The results show that with the teacher/student paradigm, it is possible to train models that generalize on datasets from entirely different sources, despite the inter-dataset heterogeneity and the lack of large datasets with local annotations. Niccolò Marini, Juan Sebastian Otálora Montenegro, Henning Müller, Manfredo Atzori |
Medical Image Anal. | 1 |