EDBT 2026 Demo / reviewers in the wild / expert
Andrea Montibeller
dblp:325/5161
· DBLP profile ↗
9ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-0794-312XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Backbone is All You Need: Assessing Vulnerabilities of Frozen Foundation Models in Synthetic Image ForensicsabstractAs AI-generated synthetic images become increasingly realistic, Vision Transformers (ViTs) have emerged as a cornerstone of modern deepfake detection. However, the prevailing reliance on frozen, pre-trained backbones introduces a subtle yet critical vulnerability. In this work, we present the Surrogate Iterative Adversarial Attack (SIAA), a gray-box attack that exploits knowledge of the detector’s ViT backbone alone and operates entirely within the target detector’s feature space to craft highly effective adversarial examples. Through our experiments, involving multiple ViT-based detectors and diverse gray-box scenarios, including few-shot learning, complete training misalignment and attack transferability tests, we demonstrate that this vulnerability consistently yields high attack success rates, often approaching white-box performance. By doing so, we reveal that backbone knowledge alone is sufficient to undermine detector reliability, highlighting the urgent need for more resilient defenses in adversarial multimedia forensics. Chiara Musso, Joy Battocchio, Andrea Montibeller, Giulia Boato |
IH&MMSec | 3 |
| 2026 | AINPAINT: A comprehensive dataset and dual branch architecture for practical video inpainting localizationabstractThe rapid evolution of generative artificial intelligence has made video inpainting and object removal highly realistic, posing a severe threat to multimedia integrity. While various forensic detectors have been proposed, they predominantly rely on high frequency noise or specific artefact signatures that are easily destroyed by real world degradations like H.264 and HEVC compression, and AI based post processing. To address this critical gap, we introduce AINPAINT, a large scale forensic dataset containing over 25,000 video sequences manipulated with nine diverse generative techniques, explicitly including variants subjected to temporal smoothing and heavy compression. On top of AINPAINT, we propose two complementary architectures for video inpainting localization built upon a LoRA adapted DINOv2 backbone. The first method extracts rich semantic spatial features, while the second augments these features with temporal motion anomalies derived from dense optical flow. Beyond merely establishing new performance baselines, our ablation provides a functional decision guide for the forensics community, clarifying when spatial features alone are preferable and when motion anomalies provide a measurable gain in the presence of post-processing, H.264 and HEVC compression and data-shifts. The dataset and code implementation are available at: Andrea Montibeller, Giulia Boato, Luisa Verdoliva |
Comput. Vis. Image Underst. | 1 |
| 2025 | Advance Fake Video Detection via Vision TransformersabstractRecent advancements in AI-based multimedia generation have enabled the creation of hyper-realistic images and videos, raising concerns about their potential use in spreading misinformation. The widespread accessibility of generative techniques, which allow for the production of fake multimedia from prompts or existing media, along with their continuous refinement, underscores the urgent need for highly accurate and generalizable AI-generated media detection methods, underlined also by new regulations like the European Digital AI Act. In this paper, we draw inspiration from Vision Transformer (ViT)-based fake image detection and extend this idea to video. We propose an original framework that effectively integrates ViT embeddings over time to enhance detection performance. Our method shows promising accuracy, generalization, and few-shot learning capabilities across a new, large and diverse dataset of videos generated using five open source generative techniques from the state-of-the-art, as well as a separate dataset containing videos produced by proprietary generative methods. Joy Battocchio, Stefano Dell'Anna, Andrea Montibeller, Giulia Boato |
IH&MMSec | 3 |
| 2025 | WILD: a new in-the-Wild Image Linkage Dataset for synthetic image attributionabstractSynthetic image source attribution is an open challenge, with an increasing number of image generators being released yearly. The complexity and the sheer number of available generative techniques, as well as the scarcity of high-quality open source datasets of diverse nature for this task, make training and benchmarking synthetic image source attribution models very challenging. WILD1is a new in-the-Wild Image Linkage Dataset designed to provide a powerful training and benchmarking tool for synthetic image attribution models. The dataset is built out of a closed set of 10 popular commercial generators, which constitutes the training base of attribution models, and an open set of 10 additional generators, simulating a real-world in-the-wild scenario. Each generator is represented by 1,000 images, for a total of 10,000 images in the closed set and 10,000 images in the open set. Half of the images are post-processed with a wide range of operators. WILD allows benchmarking attribution models in a wide range of tasks, including closed and open set identification and verification, and robust attribution with respect to post-processing and adversarial attacks. Models trained on WILD are expected to benefit from the challenging scenario represented by the dataset itself. Moreover, an assessment of seven baseline methodologies on closed and open set attribution is presented, including robustness tests with respect to post-processing. Pietro Bongini, Sara Mandelli, Andrea Montibeller, Mirko Casu, Orazio Pontorno, Claudio Vittorio Ragaglia, Luca Zanchetta, Mattia Aquilina, Taiba Majid Wani, Luca Guarnera, Benedetta Tondi, Giulia Boato, Paolo Bestagini, Irene Amerini, Francesco G. B. De Natale, Sebastiano Battiato, Mauro Barni |
IJCNN | 3 |
| 2025 | TrueFake: A Real World Case Dataset of Last Generation Fake Images also Shared on Social NetworksabstractAI-generated synthetic media are increasingly used in real-world scenarios, often with the purpose of spreading misinformation and propaganda through social media platforms, where compression and other processing can degrade fake detection cues. Currently, many forensic tools fail to account for these in-the-wild challenges. In this work, we introduce TrueFake, a large-scale benchmarking dataset of 600,000 images including top notch generative techniques and sharing via three different social networks. This dataset allows for rigorous evaluation of state-of-the-art fake image detectors under very realistic and challenging conditions. Through extensive experimentation, we analyze how social media sharing impacts detection performance, and identify current most effective detection and training strategies. Our findings highlight the need for evaluating forensic models in conditions that mirror real-world use. Stefano Dell'Anna, Andrea Montibeller, Giulia Boato |
IJCNN | 2 |
| 2024 | Shedding Light on some Leaks in PRNU-based Source AttributionabstractForensic image source attribution aims at deciding whether a query image was taken by a specific camera. While various algorithms leveraging forensic traces have been proposed, the most effective techniques rely on Photo Response Non-Uniformity (PRNU), a pattern introduced by camera sensors during the image acquisition process. In recent years, advances in image acquisition and processing technologies in modern devices have been found to impact the performance of PRNU, seemingly challenging its uniqueness. In this paper, we build upon recent discoveries of leaks in PRNU uniqueness, focusing on the dataset recently published by Iuliani et al. which has been instrumental in identifying numerous issues related to source attribution. Specifically, we analyze the effects in terms of false positive of visible watermarks applied to Xiaomi Mi 9 images, and reveal artifacts in the magnitude of the Discrete Fourier Transform of Samsung A50 images, indicative of the absence of non-unique artifacts. Furthermore, we demonstrate how several false positive cases are attributed to mislabeled devices. Finally, we show that a number of false negatives from the dataset are traceable to radially corrected images, and to images processed by third-party software that had not been previously noticed. Andrea Montibeller, Roy Alia Asiku, Fernando Pérez-González, Giulia Boato |
IH&MMSec | 1 |
| 2024 | An Adaptive Method for Camera Attribution Under Complex Radial Distortion CorrectionsabstractRadial distortion correction, applied by in-camera or out-camera software/firmware alters the supporting grid of the image so as to hamper PRNU-based camera attribution. Existing solutions to deal with this problem try to invert/estimate the correction using radial transformations parameterized with few variables in order to restrain the computational load; however, with ever more prevalent complex distortion corrections their performance is unsatisfactory. In this paper we propose an adaptive algorithm that by dividing the image into concentric annuli is able to deal with sophisticated corrections like those applied out-camera by third party software like Adobe Lightroom, Photoshop, Gimp and PT-Lens. We also introduce a statistic called cumulative peak of correlation energy (CPCE) that allows for an efficient early stopping strategy. Experiments on a large dataset of in-camera and out-camera radially corrected images and on a in-the-wild dataset of images from smartphones show that our solution improves the state of the art in terms of both accuracy and computational cost. Andrea Montibeller, Fernando Pérez-González |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2023 | Exploiting PRNU and Linear Patterns in Forensic Camera Attribution under Complex Lens Distortion CorrectionabstractMore complex and ever more common lens distortion correction post-processing is seriously hampering state-of-the-art camera attribution techniques. In this paper, we show that the two main existing techniques, namely PRNU (Photo Response Non Uniformity)-based and linear-pattern-based, can be successfully combined to improve performance. Moreover, we introduce a novel method that is able to correctly invert adaptive distortion correction transformations by successively maximizing the peak-to-correlation energy (PCE) and the linear-pattern energy for much more reliable camera attribution. A novel validation procedure to quickly discard mismatched test images is also proposed. Finally, we show how great reductions in running time can be achieved by using a GPU for interpolation, resampling, and PCE computation. The code is available at https://github.com/AMontiB/PSLR. Andrea Montibeller, Fernando Pérez-González |
ICASSP | 1 |
| 2022 | Gpu-Accelerated Sift-Aided Source Identification of Stabilized VideosabstractVideo stabilization is an in-camera processing commonly applied by modern acquisition devices. While significantly improving the visual quality of the resulting videos, it has been shown that such operation typically hinders the forensic analysis of video signals. In fact, the correct identification of the acquisition source usually based on Photo Response non-Uniformity (PRNU) is subject to the estimation of the transformation applied to each frame in the stabilization phase. A number of techniques have been proposed for dealing with this problem, which however typically suffer from a high computational burden due to the grid search in the space of inversion parameters. Our work attempts to alleviate these short-comings by exploiting the parallelization capabilities of Graphics Processing Units (GPUs), typically used for deep learning applications, in the framework of stabilised frames inversion. Moreover, we propose to exploit SIFT features to estimate the camera momentum and identify less stabilized temporal segments, thus enabling a more accurate identification analysis, and to efficiently initialize the frame-wise parameter search of consecutive frames. Experiments on a consolidated benchmark dataset confirm the effectiveness of the proposed approach in reducing the required computational time and improving the source identification accuracy. The code is available at https://github.com/AMontiB/GPU-PRNU-SIFT. Andrea Montibeller, Cecilia Pasquini, Giulia Boato, Stefano Dell'Anna, Fernando Pérez-González |
ICIP | 1 |