Andrea Ciamarra

dblp:320/4746 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0001-7053-0830ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Improving Generalization in AI-Generated Facial Image Detection via Explainable Recovery
abstract
Existing deepfake detectors achieve near-perfect accuracy when trained and tested on the same generation method, often without requiring complicated detection pipelines. However, these detectors struggle to generalize well to unseen fake images, because the artifacts on which they rely to distinguish real from fake content are not consistently present or distinctive across varying data distributions. In this study, we focus on understanding where and why detectors fail to generalize in cross-dataset scenarios, leveraging Explainable AI (XAI) methods to identify the specific failure points via a thoroughly feature-level inspection. Based on this analysis, we propose a straightforward recovery method designed to restore detection ability without significantly compromising overall performance; such an approach can be easily integrated into existing detection pipelines as a plug-and-play solution. We use a simple CNN-based synthetic image detector in our experiments for making an understanding of generalization issues and exploring the recovery process. Our approach addresses the generalization gap observed between in-dataset and cross-dataset contexts. Experiments conducted on various generative methods and implementations demonstrate the effectiveness of the proposed recovery strategy.
Giulia Ciacci, Alesssandra Spinaci, Niccolò Biondi, Andrea Ciamarra, Roberto Caldelli
IH&MMSec4
2026 Patch-Based Reconstruction and Multimodal Residual Learning for Generalized Deepfake Detection
abstract
Deepfake detectors often achieve near-perfect accuracy in-domain but fail under dataset or forgery shifts, especially when only a few manipulated samples are available for training. We propose a two-stage framework for generalized and data-efficient deepfake detection based on reconstruction residuals. In Stage I, we learn a real-face prior with PM-VAE, a masked patch reconstructor that augments a Masked Autoencoder with a lightweight variational bottleneck to regularize the patch latent space and reduce memorization. In Stage II, the generator is frozen and used to produce forensic evidence from partially observed inputs via block-wise masking on the patch grid, the resulting inpainted reconstructions yield residual cues that are stable and localized. We then train a multi-branch Transformer to fuse (i) RGB context, (ii) spatial residuals, and (iii) wavelet-domain residuals that explicitly capture high-frequency inconsistencies missed by spatial errors alone. Extensive experiments on FaceForensics++ dataset under cross-forgery protocols, on external benchmarks (Celeb-DF, DFD, DFDC) for cross-dataset evaluation, and on synthetic generation (StyleGAN family and Stable Diffusion) show improved robustness over reconstruction baselines, with consistent gains in low-data regimes down to a handful of fake frames per video.
Niccolò Marini, Andrea Ciamarra, Roberto Caldelli, Stefano Berretti
IH&MMSec2
2026 Revealing GAN-generated faces through local camera surface frame analysis
abstract
The ability of AI to generate highly realistic, fully synthetic images, particularly of human faces, is rapidly advancing, making it increasingly difficult to distinguish between real and artificially generated content. This growing realism highlights the urgent need for reliable methods to detect subtle inconsistencies introduced during the image generation process. A fundamental distinction between authentic and deepfake content lies in the absence, for the latter, of an acquisition process by a real camera. As a result, the intricate relationships among scene elements, such as lighting, reflectance, and spatial positioning, are not captured from the physical world but are artificially reconstructed. Motivated by this observation, we propose the use of local camera surface frames as a feature to encode such environment-specific attributes. Our experimental results demonstrate that this representation not only achieves high detection accuracy but also exhibits strong and robust generalisation capabilities across different GAN-based generative models.
Andrea Ciamarra, Roberto Caldelli, Alberto Del Bimbo
J. Inf. Secur.1
2025 Text-Oriented Image Query Representation for Zero-Shot Composed Image Retrieval
abstract
Zero-Shot Composed Image Retrieval (ZS-CIR) is the task of retrieving a target image based on a query that combines a reference image with a textual description specifying desired modifications in a zero-shot setting. Existing ZS-CIR models typically fuse visual and textual modalities into a single query representation, but often struggle to capture the fine-grained distinctions essential for accurate retrieval. In this paper, we present TEOZCIR, a transformer-based model that introduces a balanced semantic fusion module and an enhancement mechanism to more effectively integrate multimodal information. The model is built around two core components: the Text-Aware Query Combiner (TAQC) and the Query Enhancer Network (QENet). These components operate in tandem: TAQC dynamically adjusts the semantic contributions of the visual context based on the input text, generating a balanced query representation. This representation is then further refined by QENet, which enhances the fused features to better align with the target image. Throughout the entire process, the model maintains a lightweight architecture with significantly fewer trainable parameters compared to conventional training-based methods. Experiments carried out on three benchmark datasets CIRR, Fashion IQ, and CIRCO to demonstrate that TEOZCIR significantly improves ZS-CIR performance, setting a new bench-mark for multimodal retrieval.
Pavan K. Rachabathuni, Andrea Ciamarra, Roberto Caldelli, Marco Bertini 0001
CBMI2
2024 FLODCAST: Flow and depth forecasting via multimodal recurrent architectures
abstract
Forecasting motion and spatial positions of objects is of fundamental importance, especially in safety-critical settings such as autonomous driving. In this work, we address the issue by forecasting two different modalities that carry complementary information, namely optical flow and depth. To this end we propose FLODCAST a flow and depth forecasting model that leverages a multitask recurrent architecture, trained to jointly forecast both modalities at once. We stress the importance of training using flows and depth maps together, demonstrating that both tasks improve when the model is informed of the other modality. We train the proposed model to also perform predictions for several timesteps in the future. This provides better supervision and leads to more precise predictions, retaining the capability of the model to yield outputs autoregressively for any future time horizon. We test our model on the challenging Cityscapes dataset, obtaining state of the art results for both flow and depth forecasting. Thanks to the high quality of the generated flows, we also report benefits on the downstream task of segmentation forecasting, injecting our predictions in a flow-based mask-warping framework.
Andrea Ciamarra, Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo
Pattern Recognit.1