Roberto Caldelli

dblp:52/5048 · DBLP profile ↗
← Back
45ranked-venue papers
12as first author
15since 2021 · last 2026
0000-0003-3471-1196ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 5 first-author · 7 since 2021Security and privacy · 13 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorComputer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Improving Generalization in AI-Generated Facial Image Detection via Explainable Recovery
abstract
Existing deepfake detectors achieve near-perfect accuracy when trained and tested on the same generation method, often without requiring complicated detection pipelines. However, these detectors struggle to generalize well to unseen fake images, because the artifacts on which they rely to distinguish real from fake content are not consistently present or distinctive across varying data distributions. In this study, we focus on understanding where and why detectors fail to generalize in cross-dataset scenarios, leveraging Explainable AI (XAI) methods to identify the specific failure points via a thoroughly feature-level inspection. Based on this analysis, we propose a straightforward recovery method designed to restore detection ability without significantly compromising overall performance; such an approach can be easily integrated into existing detection pipelines as a plug-and-play solution. We use a simple CNN-based synthetic image detector in our experiments for making an understanding of generalization issues and exploring the recovery process. Our approach addresses the generalization gap observed between in-dataset and cross-dataset contexts. Experiments conducted on various generative methods and implementations demonstrate the effectiveness of the proposed recovery strategy.
Giulia Ciacci, Alesssandra Spinaci, Niccolò Biondi, Andrea Ciamarra, Roberto Caldelli
IH&MMSec5
2026 Patch-Based Reconstruction and Multimodal Residual Learning for Generalized Deepfake Detection
abstract
Deepfake detectors often achieve near-perfect accuracy in-domain but fail under dataset or forgery shifts, especially when only a few manipulated samples are available for training. We propose a two-stage framework for generalized and data-efficient deepfake detection based on reconstruction residuals. In Stage I, we learn a real-face prior with PM-VAE, a masked patch reconstructor that augments a Masked Autoencoder with a lightweight variational bottleneck to regularize the patch latent space and reduce memorization. In Stage II, the generator is frozen and used to produce forensic evidence from partially observed inputs via block-wise masking on the patch grid, the resulting inpainted reconstructions yield residual cues that are stable and localized. We then train a multi-branch Transformer to fuse (i) RGB context, (ii) spatial residuals, and (iii) wavelet-domain residuals that explicitly capture high-frequency inconsistencies missed by spatial errors alone. Extensive experiments on FaceForensics++ dataset under cross-forgery protocols, on external benchmarks (Celeb-DF, DFD, DFDC) for cross-dataset evaluation, and on synthetic generation (StyleGAN family and Stable Diffusion) show improved robustness over reconstruction baselines, with consistent gains in low-data regimes down to a handful of fake frames per video.
Niccolò Marini, Andrea Ciamarra, Roberto Caldelli, Stefano Berretti
IH&MMSec3
2026 The 5th ACM International Workshop on Multimedia AI against Disinformation (MAD'26)
abstract
Verifying the authenticity of media has become an increasingly challenging task. Rapid advances in AI-generated content, spanning modalities like text, images, video, audio have significantly blurred the line between genuine and synthetic information. Nowadays, powerful foundation models can easily be leveraged to create, amplify and disseminate information at scale, enabling disinformation campaigns, defamation, or impersonation. This results in the erosion of trust in online information, which poses a great threat to society. The MAD’26 workshop seeks to address this problem by bringing together researchers and practitioners from diverse disciplines, united by the goal of combating disinformation through AI-driven approaches. Now in its fifth edition, the workshop aims to cultivate a collaborative environment that encourages the exchange of ideas, methodologies, and practical experiences. The workshop focuses on key research directions, including the detection of AI-generated and manipulated content, the analysis of disinformation propagation, and the examination of its broader societal impact.
Dan-Cristian Stanciu, Symeon Papadopoulos, Giorgos Kordopatis-Zilos, Bogdan Ionescu, Adrian Popescu 0001, Roberto Caldelli, Milica Gerhardt, Vera Schmitt
ICMR6
2026 Synthetic frequency patterns injection for data-agnostic deepfake detection
abstract
Deepfake detectors are typically trained on large sets of pristine and generated images, resulting in limited generalization capacity; they excel at identifying deepfakes created through methods encountered during training but struggle with those generated by unknown techniques.This paper introduces a learning approach aimed at significantly enhancing the generalization capabilities of deepfake detectors. Our method takes inspiration from the unique "fingerprints" that image generation processes consistently introduce into the frequency domain. These fingerprints manifest as structured and distinctly recognizable frequency patterns. We propose to train detectors using only pristine images injecting in part of them crafted frequency patterns, simulating the effects of various deepfake generation techniques without being specific to any. These synthetic patterns are based on generic shapes, grids, or auras.We evaluated our approach using diverse architectures across 25 different generation methods. The models trained with our approach were able to perform state-of-the-art deepfake detection, demonstrating also superior generalization capabilities in comparison with previous methods. Indeed, they are untied to any specific generation technique and can effectively identify deepfakes regardless of how they were made.The code to use the proposed approach and reproduce the presented experiments is available at: https://github.com/davide-coccomini/Deepfake-Detection-without-Deepfakes-Generalization-via-Synthetic-Frequency-Patterns-Injection
Davide Coccomini, Roberto Caldelli, Claudio Gennaro, Giuseppe Fiameni, Giuseppe Amato 0001, Fabrizio Falchi
Comput. Vis. Image Underst.2
2026 Revealing GAN-generated faces through local camera surface frame analysis
abstract
The ability of AI to generate highly realistic, fully synthetic images, particularly of human faces, is rapidly advancing, making it increasingly difficult to distinguish between real and artificially generated content. This growing realism highlights the urgent need for reliable methods to detect subtle inconsistencies introduced during the image generation process. A fundamental distinction between authentic and deepfake content lies in the absence, for the latter, of an acquisition process by a real camera. As a result, the intricate relationships among scene elements, such as lighting, reflectance, and spatial positioning, are not captured from the physical world but are artificially reconstructed. Motivated by this observation, we propose the use of local camera surface frames as a feature to encode such environment-specific attributes. Our experimental results demonstrate that this representation not only achieves high detection accuracy but also exhibits strong and robust generalisation capabilities across different GAN-based generative models.
Andrea Ciamarra, Roberto Caldelli, Alberto Del Bimbo
J. Inf. Secur.2
2025 Text-Oriented Image Query Representation for Zero-Shot Composed Image Retrieval
abstract
Zero-Shot Composed Image Retrieval (ZS-CIR) is the task of retrieving a target image based on a query that combines a reference image with a textual description specifying desired modifications in a zero-shot setting. Existing ZS-CIR models typically fuse visual and textual modalities into a single query representation, but often struggle to capture the fine-grained distinctions essential for accurate retrieval. In this paper, we present TEOZCIR, a transformer-based model that introduces a balanced semantic fusion module and an enhancement mechanism to more effectively integrate multimodal information. The model is built around two core components: the Text-Aware Query Combiner (TAQC) and the Query Enhancer Network (QENet). These components operate in tandem: TAQC dynamically adjusts the semantic contributions of the visual context based on the input text, generating a balanced query representation. This representation is then further refined by QENet, which enhances the fused features to better align with the target image. Throughout the entire process, the model maintains a lightweight architecture with significantly fewer trainable parameters compared to conventional training-based methods. Experiments carried out on three benchmark datasets CIRR, Fashion IQ, and CIRCO to demonstrate that TEOZCIR significantly improves ZS-CIR performance, setting a new bench-mark for multimodal retrieval.
Pavan K. Rachabathuni, Andrea Ciamarra, Roberto Caldelli, Marco Bertini 0001
CBMI3
2025 MAD'25: 4th ACM International Workshop on Multimedia AI against Disinformation
abstract
2148
Dan-Cristian Stanciu, Bogdan Ionescu, Symeon Papadopoulos, Giorgos Kordopatis-Zilos, Adrian Popescu 0001, Roberto Caldelli, Milica Gerhardt, Vera Schmitt
ICMR6
2025 High-fidelity reversible data hiding using novel comprehensive rhombus predictor
Rajeev Kumar 0007, Roberto Caldelli, Koksheik Wong, Aruna Malik, Ki-Hyun Jung
Multim. Tools Appl.2
2024 MAD '24 Workshop: Multimedia AI against Disinformation
abstract
1339
Cristian Lucian Stanciu, Bogdan Ionescu, Luca Cuccovillo, Symeon Papadopoulos, Giorgos Kordopatis-Zilos, Adrian Popescu 0001, Roberto Caldelli
ICMR7
2024 MINTIME: Multi-Identity Size-Invariant Video Deepfake Detection
abstract
In this paper, we present MINTIME, a video deepfake detection method that effectively captures spatial and temporal inconsistencies in videos that depict multiple individuals and varying face sizes. Unlike previous approaches that either employ simplistic a-posteriori aggregation schemes, i.e., averaging or max operations, or only focus on the largest face in the video, our proposed method learns to accurately detect spatio-temporal inconsistencies across multiple identities in a video through a Spatio-Temporal Transformer combined with a Convolutional Neural Network backbone. This is achieved through an Identity-aware Attention mechanism that applies a masking operation on the face sequence to process each identity independently, which enables effective video-level aggregation. Furthermore, our system incorporates two novel embedding schemes: (i) the Temporal Coherent Positional Embedding, which encodes the temporal information of the face sequences of each identity, and (ii) the Size Embedding, which captures the relative sizes of the faces to the video frames. MINTIME achieves state-of-the-art performance on the ForgeryNet dataset, with a remarkable improvement of up to 14% AUC in videos containing multiple people. Moreover, it demonstrates very robust generalization capabilities in cross-forgery and cross-dataset settings. The code is publicly available at: https://github.com/davide-coccomini/MINTIME-Multi-Ident ity-size-iNvariant-TIMEsformer-for-Video-Deepfake-Detection.
Davide Coccomini, Giorgos Kordopatis-Zilos, Giuseppe Amato 0001, Roberto Caldelli, Fabrizio Falchi, Symeon Papadopoulos, Claudio Gennaro
IEEE Trans. Inf. Forensics Secur.4
2023 On helping users in writing network slice intents through NLP and User Profiling
abstract
Intent-based Networking (IBN) has emerged as an innovative approach to automate the provisioning of network services while abstracting the details of the underlying infrastructure and simplifying the interaction between the users and the network. In this paper, we present an intent-based framework that allows for the deployment of SDN-based and QoS-aware network slices. The main objective of the work is to describe the role of artificial intelligence techniques such as Natural Language Processing (NLP) and user profiling in helping non-expert users easily interact with the IBN system and express their desired operational goals. Such innovative solutions offer a customized support to the users to improve their Quality of Experience (QoE) while increasing the automation in the network configuration process.
Roberto Caldelli, Piero Castoldi, Molka Gharbaoui, Barbara Martini, M. Matarazzo, Filippo Sciarrone
NetSoft1
2022 Optical Network Authentication through Rayleigh Backscattering Fingerprints of the Composing Fibers
abstract
We propose a new method for network authentication, identification, and secure communication, using the optical physical unclonable function Challenge-Response (PUF-CRPs) database protocol. We investigated the database performance generated with the proposed protocol by identifying 150 networks ID. Different methods and strategies for generating and exerting PUF-based CRP databases are proposed and numerically validated. The proposed method may find applications for identification and security networks with different architectures.
Pantea Nadimi Goki, Thomas Teferi Mulugeta, Nicola Sambo, Roberto Caldelli, Luca Potì
GLOBECOM4
2022 Tuning Neural ODE Networks to Increase Adversarial Robustness in Image Forensics
abstract
Although deep-learning-based solutions are pervading different application sectors, many doubts have arisen about their reliability and, above all, their security against threats that can mislead their decision mechanisms. In this work, we considered a particular kind of deep neural network, the Neural Ordinary Differential Equations (N-ODE) networks, which have shown intrinsic robustness against adversarial samples by properly tuning their tolerance parameter at test time. Their behaviour has never been investigated in image forensics tasks such as distinguishing between an original and an altered image. Following this direction, we demonstrate how tuning the tolerance parameter during the prediction phase can control and increase N-ODE’s robustness versus adversarial attacks. We performed experiments on basic image transformations used to generate tampered data, providing encouraging results in terms of adversarial rejection and preservation of the correct classification of pristine images.
Roberto Caldelli, Fabio Carrara, Fabrizio Falchi
ICIP1
2021 Optical Flow based CNN for detection of unlearnt deepfake manipulations
abstract
A new phenomenon named Deepfakes constitutes a serious threat in video manipulation. AI-based technologies have provided easy-to-use methods to create extremely realistic videos. On the side of multimedia forensics, being able to individuate this kind of fake contents becomes ever more crucial. In this work, a new forensic technique able to detect fake and original video sequences is proposed; it is based on the use of CNNs trained to distinguish possible motion dissimilarities in the temporal structure of a video sequence by exploiting optical flow fields. The results obtained highlight comparable performances with the state-of-the-art methods which, in general, only resort to single video frames. Furthermore, the proposed optical flow based detection scheme also provides a superior robustness in the more realistic cross-forgery operative scenario and can even be combined with frame-based approaches to improve their global effectiveness.
Roberto Caldelli, Leonardo Galteri, Irene Amerini, Alberto Del Bimbo
Pattern Recognit. Lett.1
2021 Local Moment Driven PVO Based Reversible Data Hiding
abstract
Pixel-value-ordering (PVO) is one of the most widely used reversible data hiding (RDH) framework which efficiently utilizes smooth pixels of the cover image to provide high-fidelity stego-image but with limited embedding capacity. This letter proposes an RDH scheme based on local moment driven pixel value ordering (LM-PVO) which further enhances the effectiveness of smooth pixel's utilization by dividing the fixed-size blocks into two groups. Pixels of each group are sub-divided into two sub-groups based on the local moment of the block so that correlation among the pixels of each sub-group is enhanced. Thus doing, the pixels of each sub-group are grouped based on their intensity values instead of their position as in the conventional PVO-based schemes; this enables information hider to embed a higher amount of secret data while also enhancing the stego-image quality. Experimental results also validate the superiority of the proposed scheme over the existing PVO-based RDH schemes.
Neeraj Kumar 0001, Rajeev Kumar 0007, Roberto Caldelli
IEEE Signal Process. Lett.3
2020 Exploiting Prediction Error Inconsistencies through LSTM-based Classifiers to Detect Deepfake Videos
abstract
The ability of artificial intelligence techniques to build synthesized brand new videos or to alter the facial expression of already existing ones has been efficiently demonstrated in the literature. The identification of such new threat generally known as Deepfake, but consisting of different techniques, is fundamental in multimedia forensics. In fact this kind of manipulated information could undermine and easily distort the public opinion on a certain person or about a specific event. Thus, in this paper, a new technique able to distinguish synthetic generated portrait videos from natural ones is introduced by exploiting inconsistencies due to the prediction error in the re-encoding phase. In particular, features based on inter-frame prediction error have been investigated jointly with a Long Short-Term Memory (LSTM) model network able to learn the temporal correlation among consecutive frames. Preliminary results have demonstrated that such sequence-based approach, used to distinguish between original and manipulated videos, highlights promising performances.
Irene Amerini, Roberto Caldelli
IH&MMSec2
2019 Tracking Multiple Image Sharing on Social Networks
abstract
Social Networks (SN) and Instant Messaging Apps (IMA) are more and more engaging people in their personal relations taking possession of an important part of their daily life. Huge amounts of multimedia contents, mainly photos, are poured and successively shared on these networks so quickly that is not possible to follow their paths. This last issue surely grants anonymity and impunity thus it consequently makes easier to commit crimes such as reputation attack and cyberbullying. In fact, contents published within a restricted group of friends on an IMA can be rapidly delivered and viewed on a SN by acquaintances and then by strangers without any sort of tracking. In a forensic scenario (e.g., during an investigation), succeeding in understanding this flow could be strategic, thus allowing to reveal all the intermediate steps a certain content has followed. This work aims at tracking multiple sharing on social networks, by extracting specific traces left by each SN within the image file, due to the process each of them applies, to perform a multi-class classification. Innovative strategies, based on deep learning, are proposed and satisfactory results are achieved in recovering till triple up-downloads.
Quoc-Tin Phan, Giulia Boato, Roberto Caldelli, Irene Amerini
ICASSP3
2019 Exploiting CNN Layer Activations to Improve Adversarial Image Classification
abstract
Neural networks are now used in many sectors of our daily life thanks to efficient solutions such instruments provide for diverse tasks. Leaving to artificial intelligence the chance to make choices on behalf of humans inevitably exposes these tools to be fraudulently attacked. In fact, adversarial examples, intentionally crafted to fool a neural network, can dangerously induce a misclassification though appearing innocuous for a human observer. On such a basis, this paper focuses on the problem of image classification and proposes an analysis to better insight what happens inside a convolutional neural network (CNN) when it evaluates an adversarial example. In particular, the activations of the internal network layers have been analyzed and exploited to design possible countermeasures to reduce CNN vulnerability. Experimental results confirm that layer activations can be adopted to detect adversarial inputs.
Roberto Caldelli, Rudy Becarelli, Fabio Carrara, Fabrizio Falchi, Giuseppe Amato 0001
ICIP1
2019 Adversarial image detection in deep neural networks
Fabio Carrara, Fabrizio Falchi, Roberto Caldelli, Giuseppe Amato 0001, Rudy Becarelli
Multim. Tools Appl.3
2019 Special issue on Deep Learning in Image and Video Forensics
Roberto Caldelli, Marc Chaumont, Chang-Tsun Li, Irene Amerini
Signal Process. Image Commun.1
2017 Media trustworthiness verification and event assessment through an integrated framework: a case-study
Irene Amerini, Rudy Becarelli, Francesco Brancati, Roberto Caldelli, Gabriele Giunta, Massimiliano Leone Itria
Multim. Tools Appl.4
2017 Dealing with video source identification in social networks
Irene Amerini, Roberto Caldelli, Andrea Del Mastio, Andrea Di Fuccia, Cristiano Molinari, Anna Paola Rizzo
Signal Process. Image Commun.2
2017 Smartphone Fingerprinting Combining Features of On-Board Sensors
abstract
Many everyday activities involve the exchange of confidential information through the use of a smartphone in mobility, i.e., sending on e-mail, checking bank account, buying on-line, accessing cloud platforms, and health monitoring. This demonstrates how security issues related to these operations are a major challenge in our society and in particular in the cyber-security domain. This paper focuses on the use of the smartphone intrinsic and physical characteristics as a mean to build a smartphone fingerprint to enable devices identification. The basic idea proposed in this paper is to investigate how to generate a specific fingerprint that allows to distinctively and reliably characterize each smartphone. In particular, the accelerometer, the gyroscope, the magnetometer, and the audio system (microphone-speaker) are taken into account to build up a composite fingerprint based on a set of their distinctive features. Many experiments have been carried out by analyzing different classification methods, diverse features combination configurations, and operative scenarios. Satisfactory results have been obtained showing that the combination of such sensors improves smartphone distinctiveness.
Irene Amerini, Rudy Becarelli, Roberto Caldelli, Alessio Melani, Moreno Niccolai
IEEE Trans. Inf. Forensics Secur.3
2017 Image Origin Classification Based on Social Network Provenance
abstract
Recognizing information about the origin of a digital image has been individuated as a crucial task to be tackled by the image forensic scientific community. Understanding something on the previous history of an image could be strategic to address any successive assessment to be made on it: knowing the kind of device used for acquisition or, better, the model of the camera could focus investigations in a specific direction. Sometimes just revealing that a determined post-processing, such as an interpolation or a filtering, has been performed on an image could be of fundamental importance to go back to its provenance. This paper locates in such a context and proposes an innovative method to inquire if an image derives from a social network and, in particular, try to distinguish from, which one has been downloaded. The technique is based on the assumption that each social network applies a peculiar and mostly unknown manipulation that, however, leaves some distinctive traces on the image; such traces can be extracted to feature every platform. By resorting at trained classifiers, the presented methodology is satisfactorily able to discern different social network origins. Experimental results carried out on diverse image datasets and in various operative conditions witness that such a distinction is possible. In addition, the proposed method is also able to go back to the original JPEG quality factor the image had before being uploaded on a social network.
Roberto Caldelli, Rudy Becarelli, Irene Amerini
IEEE Trans. Inf. Forensics Secur.1
2015 Acquisition source identification through a blind image classification
abstract
Image forensics, besides understanding if a digital image has been forged, often aims at determining information about image origin. In particular, it could be worthy to individuate which is the kind of source (digital camera, scanner or computer graphics software) that has generated a certain photo. Such an issue has already been studied in literature, but the problem of doing that in a blind manner has not been faced so far. It is easy to understand that in many application scenarios information at disposal is usually very limited; this is the case when, given a set of L images, the authors want to establish if they belong to K different classes of acquisition sources, without having any previous knowledge about the number of specific types of generation processes. The proposed system is able, in an unsupervised and fast manner, to blindly classify a group of photos without neither any initial information about their membership nor by resorting at a trained classifier. Experimental results have been carried out to verify actual performances of the proposed methodology and a comparative analysis with two SVM‐based clustering techniques has been performed too.
Irene Amerini, Rudy Becarelli, B. Bertini, Roberto Caldelli
IET Image Process.4
2014 Exploiting perceptual quality issues in countering SIFT-based Forensic methods
abstract
Scale Invariant Feature Transform (SIFT) has been widely employed in several image application domains, including Image Forensics (e.g. detection of copy-move forgery or near duplicates). Recently, a number of methods allowing to remove SIFT keypoints from an original image have been devised studying the problem of SIFT security against malicious procedures. Such techniques are quite effective in producing an attacked image with very few (or no) keypoints, but at the expense of an image distortion. Final perceptual quality has been taken in account very roughly so far. In this paper, effectiveness of the attacking methods is evaluated also from the side of perceptual image quality; a new version of a SIFT keypoint removal method, based on a perceptual metric, is presented and an extended series of perceptive experiments is reported.
Irene Amerini, Federica Battisti, Roberto Caldelli, Marco Carli, Andrea Costanzo
ICASSP3
2014 Blind image clustering based on the Normalized Cuts criterion for camera identification
Irene Amerini, Roberto Caldelli, Pierluigi Crescenzi, Andrea Del Mastio, Andrea Marino 0001
Signal Process. Image Commun.2
2014 Forensic Analysis of SIFT Keypoint Removal and Injection
abstract
Attacks capable of removing SIFT keypoints from images have been recently devised with the intention of compromising the correct functioning of SIFT-based copy-move forgery detection. To tackle with these attacks, we propose three novel forensic detectors for the identification of images whose SIFT keypoints have been globally or locally removed. The detectors look for inconsistencies like the absence or anomalous distribution of keypoints within textured image regions. We first validate the methods on state-of-the-art keypoint removal techniques, then we further assess their robustness by devising a counter-forensic attack injecting fake SIFT keypoints in the attempt to cover the traces of removal. We apply the detectors to a practical image forensic scenario of SIFT-based copy-move forgery detection, assuming the presence of a counterfeiter who resorts to keypoint removal and injection to create copy-move forgeries that successfully elude SIFT-based detectors but are in turn exposed by the newly proposed tools.
Andrea Costanzo, Irene Amerini, Roberto Caldelli, Mauro Barni
IEEE Trans. Inf. Forensics Secur.3
2013 SIFT keypoint removal and injection for countering matching-based image forensics
abstract
Scale Invariant Feature Transform (SIFT) has been widely employed in several image application domains, including Image Forensics (e.g. detection of copy-move forgery or near duplicates). Until now, the research community has focused on studying the robustness of SIFT against legitimate image processing, but rarely concerned itself with the problem of SIFT security against malicious procedures. Recently, a number of methods allowing to remove SIFT keypoints from an original image have been devised. Although quite effective, such methods produce an attacked image with very few (or no) keypoints, thus leaving cues that can be easily exploited by a forensic analyst to reveal the occurred manipulation. In this paper, we explore the topic of reintroducing fake SIFT keypoints into a previously cleaned image in order to address the main weakness of the existing removal attacks. In particular, we evaluate the fitness of locally adaptive contrast enhancement methods to the task of injecting new keypoints. The results we obtained are encouraging: (i) it is possible to effectively introduce new keypoints whose descriptors do not match with those of the original image, thus concealing the removal forgery; (ii) the perceptual quality of the image following the removal and injection attacks is comparable to the one of the original image.
Irene Amerini, Mauro Barni, Roberto Caldelli, Andrea Costanzo
IH&MMSec3
2013 Removal and injection of keypoints for SIFT-based copy-move counter-forensics
abstract
Abstract Recent studies exposed the weaknesses of scale-invariant feature transform (SIFT)-based analysis by removing keypoints without significantly deteriorating the visual quality of the counterfeited image. As a consequence, an attacker can leverage on such weaknesses to impair or directly bypass with alarming efficacy some applications that rely on SIFT. In this paper, we further investigate this topic by addressing the dual problem of keypoint removal, i.e., the injection of fake SIFT keypoints in an image whose authentic keypoints have been previously deleted. Our interest stemmed from the consideration that an image with too few keypoints is per se a clue of counterfeit, which can be used by the forensic analyst to reveal the removal attack. Therefore, we analyse five injection tools reducing the perceptibility of keypoint removal and compare them experimentally. The results are encouraging and show that injection is feasible without causing a successive detection at SIFT matching level. To demonstrate the practical effectiveness of our procedure, we apply the best performing tool to create a forensically undetectable copy-move forgery, whereby traces of keypoint removal are hidden by means of keypoint injection.
Irene Amerini, Mauro Barni, Roberto Caldelli, Andrea Costanzo
EURASIP J. Inf. Secur.3
2013 Copy-move forgery detection and localization by means of robust clustering with J-Linkage
Irene Amerini, Lamberto Ballan, Roberto Caldelli, Alberto Del Bimbo, Luca Del Tongo, Giuseppe Serra 0001
Signal Process. Image Commun.3
2011 A SIFT-Based Forensic Method for Copy-Move Attack Detection and Transformation Recovery
abstract
One of the principal problems in image forensics is determining if a particular image is authentic or not. This can be a crucial task when images are used as basic evidence to influence judgment like, for example, in a court of law. To carry out such forensic analysis, various technological instruments have been developed in the literature. In this paper, the problem of detecting if an image has been forged is investigated; in particular, attention has been paid to the case in which an area of an image is copied and then pasted onto another zone to create a duplication or to cancel something that was awkward. Generally, to adapt the image patch to the new context a geometric transformation is needed. To detect such modifications, a novel methodology based on scale invariant features transform (SIFT) is proposed. Such a method allows us to both understand if a copy-move attack has occurred and, furthermore, to recover the geometric transformation used to perform cloning. Extensive experimental results are presented to confirm that the technique is able to precisely individuate the altered area and, in addition, to estimate the geometric transformation parameters with high reliability. The method also deals with multiple cloning.
Irene Amerini, Lamberto Ballan, Roberto Caldelli, Alberto Del Bimbo, Giuseppe Serra 0001
IEEE Trans. Inf. Forensics Secur.3
2010 Geometric tampering estimation by means of a SIFT-based forensic analysis
abstract
In many application scenarios digital images play a basic role and often it is important to assess if their content is realistic or has been manipulated to mislead watcher's opinion. Image forensics tools provide answers to similar questions. This paper, in particular, focuses on the problem of detecting if a feigned image has been created by cloning an area of the image onto another zone to make a duplication or to cancel something awkward. The proposed method is based on SIFT features and allows both to understand which are the image points involved in the counterfeit attack and, furthermore, to recover the parameters of the geometric transformation. Experimental results are provided to witness the powerfulness of the proposed technique.
Irene Amerini, Lamberto Ballan, Roberto Caldelli, Alberto Del Bimbo, Giuseppe Serra 0001
ICASSP3
2010 Reversible Watermarking Techniques: An Overview and a Classification
Roberto Caldelli, Francesco Filippini, Rudy Becarelli
EURASIP J. Inf. Secur.1
2010 A DVB-MHP web browser to pursue convergence between Digital Terrestrial Television and Internet
Irene Amerini, Giovanni Ballocca, Rudy Becarelli, Roberto Borri, Roberto Caldelli, Francesco Filippini
Multim. Tools Appl.5
2009 Integration between Digital Terrestrial Television and Internet by Means of a DVB-MHP Web Browser
Irene Amerini, Roberto Caldelli, Rudy Becarelli, Francesco Filippini, Giovanni Ballocca, Roberto Borri
WEBIST2
2006 Joint near-lossless compression and watermarking of still images for authentication and tamper localization
Roberto Caldelli, Francesco Filippini, Mauro Barni
Signal Process. Image Commun.1
2005 Effectiveness of ST-DM Watermarking Against Intra-video Collusion
Roberto Caldelli, Alessandro Piva, Mauro Barni, Andrea Carboni
IWDW1
2004 Joint near-lossless watermarking and compression for the authentication of remote sensing images
abstract
In this paper we present a new watermarking algorithm for joint near-lossless compression and authentication of remote sensing images. The adopted compression algorithm is the standard JPEG-LS algorithm. Our methodology has been designed by integrating into the standard JPEG-LS compression algorithm, by means of a stripe approach, a known authentication technique derived from Fridrich. This procedure points out two advantages: firstly, the produced bit-stream is perfectly compliant with the JPEG-LS standard, secondly, when the image has been decoded, it is always authenticated because information has been embedded in the reconstructed values. Near-lossless coding does not harm authentication procedure and robustness against different attacks is preserved
Roberto Caldelli, Giovanni Macaluso, Mauro Barni, Enrico Magli
IGARSS1
2004 Data hiding for error concealment in H.264/AVC
abstract
Recently, data hiding has been proposed to improve the performance of error concealment algorithms. In this paper, a new data hiding-based error concealment algorithm is proposed, that allows the increase of video quality in H.264/AVC wireless video transmission and real-time applications. Data hiding is used for carrying to the decoder the values of some inner pixels to be used to reconstruct lost macro blocks into intra frames through a bi-linear interpolation process.
Alessandro Piva, Roberto Caldelli, Francesco Filippini
MMSP2
2002 Metadata hiding tightly binding information to content
Roberto Caldelli, Franco Bartolini, Vito Cappellini
Dublin Core Conference1
2001 Direct estimate of motion parameters by means of Markov random fields
abstract
Motion estimation in image sequences is undoubtedly one of the most studied problems because for many applications, going from video coding to pattern recognition, motion estimation is a fundamental tool. A new methodology which, by minimizing a specific potential function, determines for each image pixel its motion parameter set is presented. The approach is based on MRFs (Markov random fields) acting on a first-order neighborhood for each selected point and on a simple motion model that accounts for rotations and translations. Experimental results on synthetic and real world sequences have demonstrated the good performance of the adopted technique and moreover a quantitative and qualitative comparison with another well-known approach has confirmed the goodness of the proposed algorithm.
Franco Bartolini, Roberto Caldelli, Vittorio Romagnoli
ICIP (2)2
2000 Geometric-Invariant Robust Watermarking through Constellation Matching in the Frequency Domain
abstract
So far digital watermarking has been indicated as the most feasible answer for multimedia copyright protection issues, though many problems, especially regarding aspects of robustness against geometrical attacks, have not been completely and adequately solved yet. Robustness against geometric manipulations has been dealt with by inserting, together with the watermark, a synchronization template to be used later in the detection phase, to determine if a geometric distortion occurred and invert it before looking for the mark. A novel technique is presented, which, by exploiting the theory of geometric invariants, inserts a watermark intrinsically resistant to this sort of manipulations, thus avoiding the need of a synchronization pattern. Preliminary experimental results proving the goodness of the methodology are discussed along with some implementation problems due to the computational complexity.
Roberto Caldelli, Mauro Barni, Franco Bartolini, Alessandro Piva
ICIP1
2000 A DWT-Based Object Watermarking System for MPEG-4 Video Streams
abstract
The MPEG-4 standard is revealing very attractive for a large set of applications. In some of them a copy protection system allowing to control the distribution of multimedia data is required. A new technology useful for copyright protection is watermarking: a digital code (watermark), indicating the copyright owner, is directly embedded into the video signal. The possibility of the MPEG-4 standard to directly access objects within a video sequence introduces a constraint to the watermarking process: even if a video object is transferred from a sequence to another, the copyright data of the single object has to be correctly detected. Another requirement is that, in order to be robust against format conversions, the watermark has to be inserted before compression. The method proposed in this paper satisfies the previous requirements by relying on an image watermarking algorithm which embeds the code in the discrete wavelet transform of each frame.
Alessandro Piva, Roberto Caldelli, Alessia De Rosa
ICIP2
1999 Regularization of optic flow estimates by means of weighted vector median filtering
abstract
Vector median filtering has been recently proposed as an effective method to refine estimated velocity fields. Here, the use of a weighted vector median filtering is suggested to improve the regularization of the optic flow field across motion boundaries. Information about the confidence of the estimated pixel velocities is exploited for the choice of the filter weights. Experimental results, on both synthetic and real-world sequences, show the effectiveness of the proposed procedure.
Luciano Alparone, Mauro Barni, Franco Bartolini, Roberto Caldelli
IEEE Trans. Image Process.4