Luca Guarnera

dblp:206/9884 · DBLP profile ↗
← Back
15ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0001-8315-351XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Proto-LeakNet: Towards signal-leak aware attribution in synthetic human face imagery
abstract
The growing sophistication of synthetic image and deepfake generation models has turned source attribution and authenticity verification into a critical challenge for modern computer vision systems. Recent studies suggest that diffusion pipelines unintentionally imprint persistent statistical traces, known as signal-leaks, within their outputs, particularly in latent representations. Building on this observation, we propose Proto-LeakNet, a signal-leak-aware and interpretable attribution framework that integrates Closed-set classification with a density-based Open-set evaluation on the learned embeddings, enabling analysis of unseen generators without retraining. Acting in the latent domain of diffusion models, our method re-simulates partial forward diffusion to expose residual generator-specific cues. A temporal attention encoder aggregates multi-step latent features, while a feature-weighted prototype head structures the embedding space and enables transparent attribution. Trained solely on closed data and achieving a Macro AUC of 98.13%, Proto-LeakNet learns a latent geometry that remains robust under post-processing, surpassing state-of-the-art methods, and achieves strong separability both between real images and known generators, and between known and unseen ones. The codebase will be available after acceptance.
Claudio Giusti, Luca Guarnera, Sebastiano Battiato
Comput. Vis. Image Underst.2
2026 Fraud is not just rarity: A causal prototype attention approach to realistic synthetic oversampling
abstract
Detecting fraudulent credit card transactions remains a significant challenge, due to the extreme class imbalance in real-world data and the often subtle patterns that separate fraud from legitimate activity. Existing research commonly attempts to address this by generating synthetic samples for the minority class using approaches such as GANs, VAEs (Variational Autoencoders), or hybrid generative models. However, these techniques, particularly when applied only to minority-class data, tend to result in overconfident classifiers and poor latent cluster separation, ultimately limiting real-world detection performance. In this study, we propose the Causal Prototype Attention Classifier (CPAC), an interpretable architecture that promotes class-aware clustering and improved latent space structure through prototype-based attention mechanisms and we couple it with the encoder of a Variational Autoencoder–Generative Adversarial Network (VAE-GAN) in order to achieve improved latent cluster separation moving beyond post-hoc sample augmentation. We compared CPAC-augmented models to traditional oversamplers, such as SMOTE, as well as to state-of-the-art generative models, both with and without CPAC-based latent classifiers. Our results show that classifier-guided latent shaping with CPAC delivers superior performance, achieving an F1-score of 93.74% and recall of 92.85%, along with improved latent cluster separation. Further ablation studies and visualizations provide deeper insight into the benefits and limitations of classifier-driven representation learning for fraud detection. The codebase for this work will be available at final submission.
Claudio Giusti, Luca Guarnera, Mirko Casu, Sebastiano Battiato
Knowl. Based Syst.2
2025 WILD: a new in-the-Wild Image Linkage Dataset for synthetic image attribution
abstract
Synthetic image source attribution is an open challenge, with an increasing number of image generators being released yearly. The complexity and the sheer number of available generative techniques, as well as the scarcity of high-quality open source datasets of diverse nature for this task, make training and benchmarking synthetic image source attribution models very challenging. WILD1is a new in-the-Wild Image Linkage Dataset designed to provide a powerful training and benchmarking tool for synthetic image attribution models. The dataset is built out of a closed set of 10 popular commercial generators, which constitutes the training base of attribution models, and an open set of 10 additional generators, simulating a real-world in-the-wild scenario. Each generator is represented by 1,000 images, for a total of 10,000 images in the closed set and 10,000 images in the open set. Half of the images are post-processed with a wide range of operators. WILD allows benchmarking attribution models in a wide range of tasks, including closed and open set identification and verification, and robust attribution with respect to post-processing and adversarial attacks. Models trained on WILD are expected to benefit from the challenging scenario represented by the dataset itself. Moreover, an assessment of seven baseline methodologies on closed and open set attribution is presented, including robustness tests with respect to post-processing.
Pietro Bongini, Sara Mandelli, Andrea Montibeller, Mirko Casu, Orazio Pontorno, Claudio Vittorio Ragaglia, Luca Zanchetta, Mattia Aquilina, Taiba Majid Wani, Luca Guarnera, Benedetta Tondi, Giulia Boato, Paolo Bestagini, Irene Amerini, Francesco G. B. De Natale, Sebastiano Battiato, Mauro Barni
IJCNN10
2025 End-to-end Audio Deepfake Detection from RAW Waveforms: a RawNet-Based Approach with Cross-Dataset Evaluation
abstract
Audio deepfakes represent a growing threat to digital security and trust, leveraging advanced generative models to produce synthetic speech that closely mimics real human voices. Detecting such manipulations is especially challenging under open-world conditions, where spoofing methods encountered during testing may differ from those seen during training. In this work, we propose an end-to-end deep learning framework for audio deepfake detection that operates directly on raw waveforms. Our model, RawNetLite, is a lightweight convolutional-recurrent architecture designed to capture both spectral and temporal features without handcrafted preprocessing. To enhance robustness, we introduce a training strategy that combines data from multiple domains and adopts Focal Loss to emphasize difficult or ambiguous samples. We further demonstrate that incorporating codec-based manipulations and applying waveform-level audio augmentations (e.g., pitch shifting, noise, and time stretching) leads to significant generalization improvements under realistic acoustic conditions. The proposed model achieves over 99.7% F1 and 0.25% EER on in-domain data (FakeOrReal), and up to 83.4% F1 with 16.4% EER on a challenging out-of-distribution test set (AVSpoof2021 + CodecFake). These findings highlight the importance of diverse training data, tailored objective functions and audio augmentations in building resilient and generalizable audio forgery detectors. Code and pretrained models are available at https://iplab.dmi.unict.it/mfs/Deepfakes/PaperRawNet2025/.
Andrea Di Pierno, Luca Guarnera, Dario Allegra, Sebastiano Battiato
IJCNN2
2025 Adversarial Attacks on Deepfake Detectors: A Challenge in the Era of AI-Generated Media (AADD-2025)
abstract
The rapid proliferation of AI-generated media, particularly hyper-realistic deepfakes, has underscored the critical need for robust detection systems to mitigate risks such as misinformation and identity theft. However, state-of-the-art deepfake detectors remain vulnerable to adversarial attacks-subtle perturbations designed to evade classification. To address this gap, we organized the Adversarial Attacks on Deepfake Detectors (AADD-2025) challenge, a competitive evaluation aimed at advancing methodologies to expose and strengthen weaknesses in deepfake detection models. The challenge tasked participants with generating adversarial examples capable of evading four diverse classifiers (including ResNet, DenseNet, and two blind models) while preserving structural similarity to original deepfakes. A dataset comprising 16 subsets of high- and low-quality deepfake images generated by GAN-based and diffusion models (e.g., StableDiffusion, StyleGAN3) was provided. Participants were evaluated using a weighted combination of Structural Similarity Index (SSIM) and attack success rates across all classifiers. Thirteen teams proposed innovative solutions leveraging techniques such as latent-space manipulation, ensemble gradient optimization, surrogate modeling, and frequency-domain perturbation. Top-performing approaches, including MR-CAS (1st place), Safe AI (2nd place), and RoMa (3rd place), achieved high SSIM scores (0.74-0.93) while successfully misleading classifiers. Notably, MR-CAS's latent diffusion model inversion strategy and Safe AI's consensus-orthogonal gradient weighting framework demonstrated superior transferability across architectures, including Vision Transformers. The challenge revealed critical insights: latent-space attacks outperformed pixel-level methods, ensemble-based strategies enhanced cross-model robustness, and adversarial perturbations optimized for both CNNs and transformers proved most effective. However, gaps persist in generalizing attacks across heterogeneous models and maintaining perceptual fidelity, highlighting the urgency of developing adaptive defenses and hybrid detection mechanisms. By fostering collaboration and innovation, AADD-2025 provides a benchmark for evaluating adversarial robustness in deepfake detection and underscores the need for resilient systems in the era of AI-generated media.
Sebastiano Battiato, Mirko Casu, Francesco Guarnera, Luca Guarnera, Giovanni Puglisi, Orazio Pontorno, Claudio Vittorio Ragaglia, Zahid Akhtar
ACM Multimedia4
2025 (DFF '25) 1st Deepfake Forensics Workshop: Detection, Attribution, Recognition, and Adversarial Challenges in the Era of AI-Generated Media
abstract
The proliferation of generative models, particularly Generative Adversarial Networks (GANs) and Diffusion Models, has reshaped multimedia content creation. Alongside creative and commercial opportunities, they have introduced unprecedented risks through the production of highly realistic synthetic content, or deepfakes. These artifacts challenge visual and auditory trust, with major implications for media, security, politics, and law. This workshop provides a forum to examine deepfake technology from forensic, technical, legal, and social perspectives. It will bring together experts to advance robust and explainable detection methods, define benchmarking practices, and address ethical and regulatory frameworks. Topics include detection and attribution, adversarial countermeasures, multimodal analysis, model traceability, legal admissibility of synthetic content, as well as real-world deployment challenges and dataset creation. Further information about the workshop is available at https://iplab.dmi.unict.it/mfs/acm-dff-ws-2025/
Sebastiano Battiato, Mirko Casu, Francesco Guarnera, Luca Guarnera, Giovanni Puglisi, Orazio Pontorno, Claudio Vittorio Ragaglia, Zahid Akhtar
ACM Multimedia4
2025 DeepFeatureX-SN: Generalization of deepfake detection via contrastive learning
abstract
Abstract The rapid advancement of generative artificial intelligence, particularly in the domains of Generative Adversarial Networks (GANs) and Diffusion Models (DMs), has led to the creation of increasingly sophisticated deepfakes. These synthetic images pose significant challenges for detection systems and present growing concerns in the realm of Cybersecurity. The potential misuse of deepfakes for disinformation, fraud, and identity theft underscores the critical need for robust detection methods. This paper introduces DeepFeatureX-SN (‘Deep Features eXtractors based Siamese Network’), an innovative deep learning model designed to address the complex task of not only distinguishing between real and synthetic images but also identifying the specific employed generative technique (GAN or DM). Our approach makes use of a tripartite structure of specialized base models, each trained using Siamese networks and contrastive learning techniques, to extract discriminative features unique to real, GAN-generated, and DM-generated images. These features are then combined through a CNN-based classifier for final categorization. Extensive experiments demonstrate the model’s superior performance, with a detection accuracy of 97.29%, strong generalization to unseen generative architectures (achieving an average accuracy of 67.40%, which surpasses most existing approaches by over 10%) and robustness against various image manipulations, all of which are crucial for real-world Cybersecurity applications. DeepFeatureX-SN achieves state-of-the-art results across multiple datasets, showing particular strength in detecting images from novel GAN and DM implementations. Furthermore, a comprehensive ablation study validates the effectiveness of each component in our proposed architecture. This research contributes significantly to the field, offering a more nuanced and accurate approach to identifying and categorizing synthetic images. The results obtained in the different configurations in the generalization tests demonstrate the good capabilities of the model, outperforming methods found in the literature. Codes and models are available at https://iplab.dmi.unict.it/mfs/Deepfakes/DeepFeatureX-SN/ .
Orazio Pontorno, Luca Guarnera, Sebastiano Battiato
Multim. Tools Appl.2
2025 Benchmarking computer vision architectures for cloud detection from lidar ceilometer backscatter data
abstract
Abstract Cloud detection is fundamental for accurate weather monitoring, often achieved through remote sensing technology, such as satellite imagery or radar. This study explores the use of lidar ceilometer backscatter data, a rich but noisy source of atmospheric information, to enhance cloud detection. Leveraging data acquired from a Lufft CHM 15k ceilometer over three months near Mount Etna, Italy, we gathered a novel dataset comprising time-height plots derived from backscatter profiles. The Weather Research and Forecasting (WRF) model was used for ground-truth data labeling, ensuring reliable model validation. We benchmarked state-of-the-art deep learning architectures, including CNN-based models (e.g., ResNet50, VGG16, InceptionV3, EfficientNet) and the Vision Transformer (ViT), on our collected dataset. Among these, ResNet50 achieved the highest accuracy ( $$89.57\%$$ 89.57 % ), closely followed by ViT ( $$89.36\%$$ 89.36 % ), showcasing the efficacy of residual learning and transformer-based approaches in extracting complex patterns from atmospheric data. Our results highlight the potential of lidar-based systems for accurate cloud detection, complementing other remote sensing technologies. Our work contributes to the field by introducing a publicly available dataset and providing comprehensive benchmarking results that establish a baseline for future research. This study also opens avenues for broader applications of ceilometer data, such as the detection of pollutants and other atmospheric phenomena. Our dataset is publicly available at https://zenodo.org/records/10616434 .
Alessio Barbaro Chisari, Luca Guarnera, Alessandro Ortis, Wladimiro Carlo Patatu, Sebastiano Battiato, Mario Valerio Giuffrida
Vis. Comput.2
2024 Innovative Methods for Non-Destructive Inspection of Handwritten Documents
abstract
Handwritten document analysis is an area of forensic science, with the goal of establishing authorship of documents through examination of inherent characteristics. Law enforcement agencies use standard protocols based on manual processing of handwritten documents. This method is time-consuming, is often subjective in its evaluation, and is not replicable. To overcome these limitations, in this paper we present a framework capable of extracting and analyzing intrinsic measures of manuscript documents related to text line heights, space between words, and character sizes using image processing and deep learning techniques. The final feature vector for each document involved consists of the mean (η) and standard deviation (σ) for every type of measure collected. By quantifying the Euclidean distance between the feature vectors of the documents to be compared, authorship can be discerned. Our study pioneered the comparison between traditionally handwritten documents and those produced with digital tools (e.g., tablets). Experimental results demonstrate the ability of our method to objectively determine authorship in different writing media, outperforming the state of the art.
Eleonora Breci, Luca Guarnera, Sebastiano Battiato
ICASSP2
2024 On the Cloud Detection from Backscattered Images Generated from a Lidar-Based Ceilometer: Current State and Opportunities
abstract
Accurate weather monitoring depends significantly on cloud detection, a crucial process achievable through remote sensing tools such as satellite imagery and radar or through the analysis of data obtained from ceilometers. A ceilometer is a lidar-based device allowing to analyse the atmosphere and detect the presence of particles within clouds. The data retrieved from ceilometers involve analysis of the backscatter of the lidar signal returning to the surface. Given the inherent noise in this data, we leverage deep learning models to detect the presence of clouds in the data. To label the data, we take advantage of a Weather Research & Forecasting (WRF) model, which provided us with ground-truth used for validation purposes. We performed a comparative analysis with current state-of-the-art deep learning architectures on this specialist domain. This comparative analysis shows that the best model is ResNet 50, but also a transformer-based model, such as ViT, achieves great results. These preliminary results pave the scenario for future works aimed at detecting other particles composing the atmosphere, such as polluting agents that can be detected from the ceilometer backscatter data.
Alessio Barbaro Chisari, Alessandro Ortis, Luca Guarnera, Wladimiro Carlo Patatu, Rosaria Ausilia Giandolfo, Emanuele Spampinato, Sebastiano Battiato, Mario Valerio Giuffrida
ICIP3
2024 On the Exploitation of DCT-Traces in the Generative-AI Domain
abstract
Deepfakes represent one of the toughest challenges in the world of Cybersecurity and Digital Forensics, especially considering the high-quality results obtained with recent generative AI-based solutions. Almost all generative models leave unique traces in synthetic data that, if analyzed and identified in detail, can be exploited to improve the generalization limitations of existing deepfake detectors. In this paper we analyzed deepfake images in the frequency domain generated by both GAN and Diffusion Model engines, examining in detail the underlying statistical distribution of Discrete Cosine Transform (DCT) coefficients. Recognizing that not all coefficients contribute equally to image detection, we hypothesize the existence of a unique “discriminative fingerprint”, embedded in specific combinations of coefficients. To identify them, Machine Learning classifiers were trained on various combinations of coefficients. In addition, the Explainable AI (XAI) LIME algorithm was used to search for intrinsic discriminative combinations of coefficients. Finally, we performed a robustness test to analyze the persistence of traces by applying JPEG compression. The experimental results reveal the existence of traces left by the generative models that are more discriminative and persistent at JPEG attacks. Code and dataset are available at github/opontorno/dcts_analysis_deepfakes.
Orazio Pontorno, Luca Guarnera, Sebastiano Battiato
ICIP2
2024 DeepFeatureX Net: Deep Features eXtractors Based Network for Discriminating Synthetic from Real Images
Orazio Pontorno, Luca Guarnera, Sebastiano Battiato
ICPR (21)2
2024 Mastering Deepfake Detection: A Cutting-edge Approach to Distinguish GAN and Diffusion-model Images
abstract
Detecting and recognizing deepfakes is a pressing issue in the digital age. In this study, we first collected a dataset of pristine images and fake ones properly generated by nine different Generative Adversarial Network (GAN) architectures and four Diffusion Models (DM). The dataset contained a total of 83,000 images, with equal distribution between the real and deepfake data. Then, to address different deepfake detection and recognition tasks, we proposed a hierarchical multi-level approach. At the first level, we classified real images from AI-generated ones. At the second level, we distinguished between images generated by GANs and DMs. At the third level (composed of two additional sub-levels), we recognized the specific GAN and DM architectures used to generate the synthetic data. Experimental results demonstrated that our approach achieved more than 97% classification accuracy, outperforming existing state-of-the-art methods. The models obtained in the different levels turn out to be robust to various attacks such as JPEG compression (with different quality factor values) and resize (and others), demonstrating that the framework can be used and applied in real-world contexts (such as the analysis of multimedia data shared in the various social platforms) for support even in forensic investigations to counter the illicit use of these powerful and modern generative models. We are able to identify the specific GAN and DM architecture used to generate the image, which is critical in tracking down the source of the deepfake. Our hierarchical multi-level approach to deepfake detection and recognition shows promising results in identifying deepfakes allowing focus on underlying task by improving (about 2% on the average) standard multiclass flat detection systems. The proposed method has the potential to enhance the performance of deepfake detection systems, aid in the fight against the spread of fake images, and safeguard the authenticity of digital media.
Luca Guarnera, Oliver Giudice, Sebastiano Battiato
ACM Trans. Multim. Comput. Commun. Appl.1
2023 Assessing forensic ballistics three-dimensionally through graphical reconstruction and immersive VR observation
abstract
Abstract A crime scene can provide valuable evidence critical to explain reason and modality of the occurred crime, and it can also lead to the arrest of criminals. The type of evidence collected by crime scene investigators or by law enforcement may accordingly effective involved cases. Bullets and cartridge cases examination is of paramount importance in forensic science because they may contain traces of microscopic striations, impressions and markings, which are unique and reproducible as “ballistic fingerprints”. The analysis of bullets and cartridge cases is a complicated and challenging process, typically based on optical comparison, leading to the identification of the employed firearm. New methods have recently been proposed for more accurate comparisons, which rely on three-dimensionally reconstructed data. This paper aims at further advancing recent methods by introducing a novel immersive technique for ballistics comparison by means of Virtual Reality. Users can three-dimensionally examine the cartridge cases shapes through intuitive natural gestures, from any vantage viewpoint (including internal iper-magnified views), while having at their disposal sets of visual aids which could not be easily implemented in desktop-based applications. A user study was conducted to assess viability and performance of our solution, which involved fourteen individuals acquainted with the standard procedures used by law enforcement agencies. Results clearly indicated that our approach lead to faster adaptation of users to the UI/UX and more accurate and explainable ballistics examination results.
Luca Guarnera, Oliver Giudice, Salvatore Livatino, Antonino Barbaro Paratore, Angelo Salici, Sebastiano Battiato
Multim. Tools Appl.1
2019 Siamese Ballistics Neural Network
abstract
Firearm identification is crucial in many investigative scenario. The crime scene often contains traces left by firearms in terms of bullets and cartridges. Traces analysis is a fundamental step in the Forensics Ballistics Analysis Process to identify which firearm fired a specific cartridge. In this paper we present a fully automated technique to compare cartridges represented as a set of 3D point-clouds. The overall approach is based on Siamese Neural Network learning paradigm that we use to build a suitable embedding space where the 3D point-cloud of the cartridges are compared. The proposed approach has been assessed by considering the NBTRD dataset. Obtained results support the exploitation of the proposed technique in ballistic analysis.
Oliver Giudice, Luca Guarnera, Antonino Barbaro Paratore, Giovanni Maria Farinella, Sebastiano Battiato
ICIP2