Irene Amerini

dblp:46/2805 · DBLP profile ↗
← Back
51ranked-venue papers
17as first author
33since 2021 · last 2026
0000-0002-6461-1391ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 9 first-author · 9 since 2021Artificial intelligence and machine learning · 16 · 2 first-author · 16 since 2021Security and privacy · 11 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A comprehensive evaluation framework and efficient regularization strategy for adversarial robustness of deepfake detectors under closed and open set scenarios
Najeh Alhalawani, Lorenzo Cirillo, Irene Amerini
Comput. Vis. Image Underst.3
2026 Geometric algebra transformer with spectral-spatial-volumetric feature extraction for hyperspectral image classification
abstract
• Weight-free multivector embedding of spatial, spectral, and volumetric data. • Geometric Algebra Transformer enabling relational reasoning in hyperspectral data. • Compact and interpretable geometric representation with reduced model complexity. • State-of-the-art accuracy with fewer parameters than baseline models. • Opens Geometric Algebra Transformer to multimodal fusion and broader sensing tasks. Current approaches to hyperspectral imaging classification have achieved important results, mostly depending on convolution-based preprocessing that adds complexity to the approach. At the same time, these methods often do not directly integrate all the data’s spatial, spectral, and volumetric dependencies at once. This article introduces Hyperspectral Geometric Algebra Transformer (H-GATr), an approach that addresses these limitations by using a Geometric Algebra (GA) architecture to integrate hyperspectral information into multivector representations, without additional trainable parameters. Thus, exploiting the intrinsic geometric structure of the data is possible to obtain an efficient integration of all hyperspectral sample dependencies. Experiments carried out on various benchmark datasets demonstrate that H-GATr is capable of achieving performance comparable to and even superior to current models, offering a compact solution for remote sensing applications.
Jacopo Fabi, Claudio Schiavella, Irene Amerini
Pattern Recognit. Lett.3
2026 Multi-Scale Self-Supervised Learning for Efficient Audio Deepfake Detection
abstract
Neural vocoders enable highly realistic synthetic speech that challenges multimedia authentication; however, existing detection approaches suffer from limited robustness to unseen synthesis methods and inadequate deployment readiness. We propose MASD (Multi-scale Artifact-aware Self-supervised Deepfake detector), combining multi-scale SSL with handcrafted features. MASD decomposes spectrograms into three frequency bands, processed through an encoder pretrained using masked reconstruction, contrastive predictive coding, and adversarial vocoder classification. Features fuse with phase coherence, spectral flux, and high-frequency energy through cross-attention, classified by temperature-scaled SVM. Evaluation on ASVspoof 2019 LA demonstrates state-of-the-art performance. Ablation studies confirm adversarial augmentation as the primary driver of robustness, improving EER from 1.52% to 0.39%, while zero-shot cross-dataset evaluation validates generalization effectiveness, establishing MASD as a practical solution.
Taiba Majid Wani, Irene Amerini
IEEE Signal Process. Lett.2
2025 R3ST: A Synthetic 3D Dataset with Realistic Trajectories
Simone Teglia, Claudia Melis Tonti, Francesco Pro, Leonardo Russo, Andrea Alfarano, Matteo Pentassuglia, Irene Amerini
CAIP (2)7
2025 Enabling Smart Urban Mobility with Edge AI
abstract
Road accidents are a major cause of death in urban areas. Cooperative Intelligent Transport Systems may mitigate or prevent road accidents by providing drivers with relevant and timely warnings, thanks to V2X communications. However, this requires fast detection of dangers, performed at the edge, to minimize transmission latencies and ensure the timeliness of the warnings. The PMDI project extends the STEP platform for intelligent mobility to support real-time operations and ETSI message generation, and provides integration with mobile and embedded systems.
Andrea Tuscano, Paolo Giuseppetti, Pietro Amato, Alessandro Solinas, Federico Saluz, Andrea Tessieri, Mario Pedol, Manuel Pernigotto, Massimo Fioravanti, Giovanni Agosta, William Fornaciari, Paolo Maffezzoni, Fabio Salice, Irene Amerini, Francesco Pro, Paolo Satta, Giovanni Trovini
DSD15
2025 Training-Free Deepfake Detection: An Identity-Guided Approach Using Facial Embeddings
abstract
Artificially generated images that exploit personal identity to undermine individual dignity present a significant challenge. Conventional detection methods, which rely on classifiers trained to recognize artifacts specific to known generators, often fail when confronted with forgeries from unknown sources. This paper introduces a training-free framework for deepfake detection that leverages pre-trained models and facial embeddings. Our approach, using models exclusively trained on authentic images, retrieves a set of genuine reference images corresponding to the target image’s identity and applies specialized operations on the resulting facial embeddings to assess authenticity. The technique achieves detection accuracies approaching 90%, underscoring its potential as a robust countermeasure against identity-based deepfake generation.
Mattia Aquilina, Taiba Majid Wani, Irene Amerini
IJCNN3
2025 WILD: a new in-the-Wild Image Linkage Dataset for synthetic image attribution
abstract
Synthetic image source attribution is an open challenge, with an increasing number of image generators being released yearly. The complexity and the sheer number of available generative techniques, as well as the scarcity of high-quality open source datasets of diverse nature for this task, make training and benchmarking synthetic image source attribution models very challenging. WILD1is a new in-the-Wild Image Linkage Dataset designed to provide a powerful training and benchmarking tool for synthetic image attribution models. The dataset is built out of a closed set of 10 popular commercial generators, which constitutes the training base of attribution models, and an open set of 10 additional generators, simulating a real-world in-the-wild scenario. Each generator is represented by 1,000 images, for a total of 10,000 images in the closed set and 10,000 images in the open set. Half of the images are post-processed with a wide range of operators. WILD allows benchmarking attribution models in a wide range of tasks, including closed and open set identification and verification, and robust attribution with respect to post-processing and adversarial attacks. Models trained on WILD are expected to benefit from the challenging scenario represented by the dataset itself. Moreover, an assessment of seven baseline methodologies on closed and open set attribution is presented, including robustness tests with respect to post-processing.
Pietro Bongini, Sara Mandelli, Andrea Montibeller, Mirko Casu, Orazio Pontorno, Claudio Vittorio Ragaglia, Luca Zanchetta, Mattia Aquilina, Taiba Majid Wani, Luca Guarnera, Benedetta Tondi, Giulia Boato, Paolo Bestagini, Irene Amerini, Francesco G. B. De Natale, Sebastiano Battiato, Mauro Barni
IJCNN14
2025 Shedding Light on Depth: Explainability Assessment in Monocular Depth Estimation
abstract
Explainable artificial intelligence is increasingly employed to understand the decision-making process of deep learning models and create trustworthiness in their adoption. However, the explainability of Monocular Depth Estimation (MDE) remains largely unexplored despite its wide deployment in real-world applications. In this work, we study how to analyze MDE networks to map the input image to the predicted depth map. More in detail, we investigate well-established feature attribution methods, Saliency Maps, Integrated Gradients, and Attention Roll-out on different computationally complex models for MDE: METER, a lightweight network, and PixelFormer, a deep network. We assess the quality of the generated visual explanations by selectively perturbing the most relevant and irrelevant pixels, as identified by the explainability methods, and analyzing the impact of these perturbations on the model’s output. Moreover, since existing evaluation metrics can have some limitations in measuring the validity of visual explanations for MDE, we additionally introduce the Attribution Fidelity. This metric evaluates the reliability of the feature attributions by assessing their consistency with the predicted depth map. Experimental results demonstrate that Saliency Maps and Integrated Gradients have good performance in highlighting the most important input features for MDE lightweight and deep models, respectively. Furthermore, we show that Attribution Fidelity effectively identifies whether an explainability method fails to produce reliable visual maps, even in scenarios where conventional metrics might suggest satisfactory results.
Lorenzo Cirillo, Claudio Schiavella, Lorenzo Papa, Paolo Russo 0001, Irene Amerini
IJCNN5
2025 STLight: A Fully Convolutional Approach for Efficient Predictive Learning by Spatio-Temporal Joint Processing
abstract
Spatio-Temporal predictive Learning is a self-supervised learning paradigm that enables models to identify spatial and temporal patterns by predicting future frames based on past frames. Traditional methods, which use recurrent neural networks to capture temporal patterns, have proven their effectiveness but come with high system complexity and computational demand. Convolutions could offer a more efficient alternative but are limited by their characteristic of treating all previous frames equally, resulting in poor temporal characterization, and by their local receptive field, limiting the capacity to capture distant correlations among frames. In this paper, we propose STLight, a novel method for spatiotemporal learning that relies solely on channel-wise and depth-wise convolutions as learnable layers. STLight overcomes the limitations of traditional convolutional approaches by rearranging spatial and temporal dimensions together, using a single convolution to mix both types of features into a comprehensive spatiotemporal patch representation. This representation is then processed in a purely convolutional framework, capable of focusing simultaneously on the interaction among near and distant patches, and subsequently allowing for efficient reconstruction of the predicted frames. Our architecture achieves state-of-the-art performance on STL benchmarks across different datasets and settings, while significantly improving computational efficiency in terms of parameters and computational FLOPs. The code is publicly available11https://github.com/AlfaranoAndrea/STLight/.
Andrea Alfarano, Alberto Alfarano, Linda Friso, Andrea Bacciu, Irene Amerini, Fabrizio Silvestri
WACV5
2025 Explainability-driven adversarial robustness assessment for generalized deepfake detectors
abstract
The capabilities of generative models to produce high-quality fake images require deepfake detectors to be accurate and have strong generalization performance. Moreover, the explainability and adversarial robustness of deepfake detectors are critical to apply such models in real-world scenarios. In this paper, we propose a framework that leverages explainability to assess the adversarial robustness of deepfake detectors. Specifically, we apply feature attribution methods to identify image regions where the model is focusing to make its prediction. Then we use the generated heatmaps to perform an explainability-driven attack, perturbing the most relevant and irrelevant regions with gradient-based adversarial techniques. We feed the model with the resulting adversarial images and measure the accuracy drop and the attack success rate. We tested our methodology on state-of-the-art models with strong generalization abilities, providing a comprehensive and explainability-driven evaluation of their robustness. Experimental results show the explainability analysis serves as a tool to reveal vulnerabilities of generalized deepfake detectors to adversarial attacks.
Lorenzo Cirillo, Andrea Gervasio, Irene Amerini
EURASIP J. Inf. Secur.3
2025 ASTDT: an Interpretable Adaptive Spectro-Temporal Diffusion Transformer for audio deepfake detection
abstract
Advances in audio synthesis techniques have led to the creation of highly realistic audio deepfakes, posing growing threats to digital integrity and public trust. These synthetic manipulations mimic natural speech with high fidelity, making detection increasingly challenging and fueling the spread of misinformation, identity fraud, and voice-based attacks. To address these concerns, this study proposes the Adaptive Spectro-Temporal Diffusion Transformer (ASTDT), a novel detection framework that tackles key challenges in generalization, interpretability, and adaptability across diverse audio generation techniques. ASTDT integrates a score-based diffusion model to augment training spectrograms with realistic deepfake variations, improving generalization to unseen text-to-speech and voice conversion attacks. An adaptive spectro-temporal feature extraction mechanism partitions audio into interpretable frequency and temporal segments, while a dual-modal attention fusion module jointly processes magnitude and phase features. These fused features are processed by a transformer encoder with diffusion-aware attention, enabling effective modeling of long-range temporal dependencies. To enhance transparency, ASTDT includes an interpretability module that combines quantitative feature attributions and spatial heatmaps to explain model predictions. Experimental results across four benchmark datasets demonstrate the effectiveness of ASTDT, with the model achieving the lowest equal error rate of 1.20% on the ASVspoof 2019 dataset.
Taiba Maijd Wani, Syed Asif Ahmad Qadri, Arselan Ashraf, Irene Amerini
EURASIP J. Inf. Secur.4
2025 Dynamic knowledge condensation with audio-selective transformer for audio deepfake detection
abstract
Abstract The rapid evolution of audio deepfakes has raised significant challenges for the security and reliability of voice-driven systems. While recent detection frameworks achieve high accuracy in controlled environments, their performance often degrades under real-world conditions involving codec compression, signal preprocessing, or domain shifts. To address these challenges, we propose dynamic knowledge condensation with audio-selective transformer (DK-CAST), a novel tri-stream knowledge distillation framework designed for robust audio deepfake detection. DK-CAST employs a high-capacity XLS-R teacher trained on clean speech to supervise a compact student model operating on degraded and preprocessed audio. The student employs a custom audio-selective transformer with dual-stream encoding, dynamic fusion, and phoneme-gated attention to emphasize linguistically relevant cues. Knowledge is transferred via multi-level supervision, including logits, embeddings, and phoneme posteriors, and modulated through a codec-aware loss weighting scheme. To enhance generalization, DK-CAST also includes a compression-agnostic embedding alignment module based on MMD and Center Loss. Evaluations on ASVspoof 2019-LA and ASVspoof 2021-DF demonstrate state-of-the-art performance, achieving EERs of 0.38 and 2.18%, respectively. Furthermore, DK-CAST maintains strong performance under codec degradation, achieving an EER of 3.01% on ASVspoof 2021-DF when tested under MP3 compression.
Taiba Maijd Wani, Irene Amerini
Discov. Comput.2
2025 Unsupervised pedestrian intention estimation through deep neural embeddings and spatio-temporal graph convolutional networks
abstract
Abstract A deep understanding of pedestrian intention and crossing behaviors is crucial in applications like pedestrian attribute recognition and autonomous driving. While vehicles need to predict the movements of pedestrians accurately for safety, the recognition and re-identification systems rely on behavioral cues that help them enhance identity tracking and attribute analysis. Traditional trajectory-based methods for pedestrian intention estimation evaluate the future positions of pedestrians based on their past movements but may fail to capture their true intentions. A more effective approach will anticipate actions by analyzing underlying intent, improving the precision of pedestrian recognition and the motion prediction. Current research on estimating pedestrian intentions primarily depends on supervised learning methods. In contrast, this work introduces an unsupervised learning approach to learn intention representations. This method is based on the idea that similar intentions lead to comparable behaviors among pedestrians, and, therefore, they can be clustered. To achieve this, this paper introduces UnPIE, an unsupervised method for predicting pedestrian intentions. It utilizes Spatio-Temporal Graph Convolutional Networks to encode intentions from videos and map them into a D-dimensional latent space. The training phase incorporates Instance Recognition to increase separation between embeddings from different classes and Local Aggregation to form soft clusters of related embeddings. A supervised non-parametric classifier is used to evaluate the performance of the method. The results demonstrate that UnPIE has comparable performance with respect to supervised approaches and even surpasses them, achieving a higher Precision by about 7% on the Pedestrian Intention Estimation dataset.
Simone Scaccia, Francesco Pro, Irene Amerini
Pattern Anal. Appl.3
2024 Robust CLIP-Based Detector for Exposing Diffusion Model-Generated Images
abstract
Diffusion models (DMs) have revolutionized image generation, producing high-quality images with applications spanning various fields. However, their ability to create hyper-realistic images poses significant challenges in distinguishing between real and synthetic content, raising concerns about digital authenticity and potential misuse in creating deepfakes. This work introduces a robust detection framework that integrates image and text features extracted by CLIP model with a Multilayer Perceptron (MLP) classifier. We propose a novel loss that can improve the detector’s robustness and handle imbalanced datasets. Additionally, we flatten the loss landscape during the model training to improve the detector’s generalization capabilities. The effectiveness of our method, which outperforms traditional detection techniques, is demonstrated through extensive experiments, underscoring its potential to set a new state-of-the-art approach in DM-generated image detection. The code is available at https://github.com/Purdue-M2/RobustDM_Generated_Image_Detection.
Santosh, Irene Amerini, Xin Wang 0045, Shu Hu 0001
AVSS3
2024 Audio Deepfake Detection: A Continual Approach with Feature Distillation and Dynamic Class Rebalancing
Taiba Majid Wani, Irene Amerini
ICPR (21)2
2024 Learning from Unlabelled data with Transformers: Domain Adaptation for Semantic Segmentation of High Resolution Aerial Images
abstract
Data from satellites or aerial vehicles are most of the times unlabelled. Annotating such data accurately is difficult, requires expertise, and is costly in terms of time. Even if Earth Observation (EO) data were correctly labelled, labels might change over time. Learning from unlabelled data within a semi-supervised learning framework for segmentation of aerial images is challenging. In this paper, we develop a new model for semantic segmentation of unlabelled images, the Non-annotated Earth Observation Semantic Segmentation (NEOS) model. NEOS performs domain adaptation as the target domain does not have ground truth masks. The distribution inconsistencies between the target and source domains are due to differences in acquisition scenes, environment conditions, sensors, and times. Our model aligns the learned representations of the different domains to make them coincide. The evaluation results show that it is successful and outperforms other models for semantic segmentation of unlabelled data.
Nikolaos Dionelis, Francesco Pro, Luca Maiano, Irene Amerini, Bertrand Le Saux
IGARSS4
2024 A Semantic Segmentation-Guided Approach for Ground-to-Aerial Image Matching
abstract
Nowadays the accurate geo-localization of ground-view images has an important role across domains as diverse as journalism, forensics analysis, transports, and Earth Observation. This work addresses the problem of matching a query ground-view image with the corresponding satellite image without GPS data. This is done by comparing the features from a ground-view image and a satellite one, innovatively leveraging the corresponding latter’s segmentation mask through a three-stream Siamese-like network. The proposed method, Semantic Align Net (SAN), focuses on limited Field-of-View (FoV) and ground panorama images (images with a FoV of 360°). The novelty lies in the fusion of satellite images in combination with their segmentation masks, aimed at ensuring that the model can extract useful features and focus on the significant parts of the images. This work shows how SAN through semantic analysis of images improves the performance on the unlabelled CVUSA dataset for all the tested FoVs.
Francesco Pro, Nikolaos Dionelis, Luca Maiano, Bertrand Le Saux, Irene Amerini
IGARSS5
2024 Detecting Audio Deepfakes: Integrating CNN and BiLSTM with Multi-Feature Concatenation
abstract
Audio deepfake detection is emerging as a crucial field in digital media, as distinguishing real audio from deepfakes becomes increasingly challenging due to the advancement of deepfake technologies. These methods threaten information authenticity and pose serious security risks. Addressing this challenge, we propose a novel architecture that combines Convolutional Neural Networks (CNN) and Bidirectional Long Short-Term Memory (BiLSTM) for effective deepfake audio detection. Our approach is distinguished by the feature concatenation of a comprehensive set of acoustic features: Mel Frequency Cepstral Coefficients (MFCC), Mel spectrograms, Constant Q Cepstral Coefficients (CQCC), and Constant-Q Transform (CQT) vectors. In the proposed architecture, features processed by a CNN are concatenated into two multi-dimensional features for comprehensive analysis, then analyzed by a BiLSTM network to capture temporal dynamics and contextual dependencies in audio data. This synergistic method ensures an understanding of both spatial and sequential audio characteristics. We validate our model on the ASVSpoof 2019 and FoR datasets, using accuracy and Equal Error Rate (EER) metrics for the evaluation.
Taiba Majid Wani, Syed Asif Ahmad Qadri, Danilo Comminiello, Irene Amerini
IH&MMSec4
2024 Estimating optical flow: A comprehensive review of the state of the art
abstract
Optical flow estimation is a crucial task in computer vision that provides low-level motion information. Despite recent advances, real-world applications still present significant challenges. This survey provides an overview of optical flow techniques and their application. For a comprehensive review, this survey covers both classical frameworks and the latest AI-based techniques. In doing so, we highlight the limitations of current benchmarks and metrics, underscoring the need for more representative datasets and comprehensive evaluation methods. The survey also highlights the importance of integrating industry knowledge and adopting training practices optimized for deep learning-based models. By addressing these issues, future research can aid the development of robust and efficient optical flow methods that can effectively address real-world scenarios. • Investigating integration of traditional techniques in modern models. • Surveying key challenges of optical flow in real-world applications. • Offering the most comprehensive survey of datasets for optical flow. • Presenting a complete overview of Classical and Modern Optical Flow methods. • Highlighting crucial open questions, paving way for future research.
Andrea Alfarano, Luca Maiano, Lorenzo Papa, Irene Amerini
Comput. Vis. Image Underst.4
2024 Continuous fake media detection: Adapting deepfake detectors to new generative techniques
abstract
Generative techniques continue to evolve at an impressively high rate, driven by the hype about these technologies. This rapid advancement severely limits the application of deepfake detectors, which, despite numerous efforts by the scientific community, struggle to achieve sufficiently robust performance against the ever-changing content. To address these limitations, in this paper, we propose an analysis of two continuous learning techniques on a Short and a Long sequence of fake media. Both sequences include a complex and heterogeneous range of deepfakes (generated images and videos) from GANs, computer graphics techniques, and unknown sources. Our experiments show that continual learning could be important in mitigating the need for generalizability. In fact, we show that, although with some limitations, continual learning methods help to maintain good performance across the entire training sequence. For these techniques to work in a sufficiently robust way, however, it is necessary that the tasks in the sequence share similarities. In fact, according to our experiments, the order and similarity of the tasks can affect the performance of the models over time. To address this problem, we show that it is possible to group tasks based on their similarity. This small measure allows for a significant improvement even in longer sequences. This result suggests that continual techniques can be combined with the most promising detection methods, allowing them to catch up with the latest generative techniques. In addition to this, we propose an overview of how this learning approach can be integrated into a deepfake detection pipeline for continuous integration and continuous deployment (CI/CD). This allows you to keep track of different funds, such as social networks, new generative tools, or third-party datasets, and through the integration of continuous learning, allows constant maintenance of the detectors.
Francesco Tassone, Luca Maiano, Irene Amerini
Comput. Vis. Image Underst.3
2024 D-Fence layer: an ensemble framework for comprehensive deepfake detection
Asha S 0001, P. Vinod 0001, Irene Amerini, Varun G. Menon
Multim. Tools Appl.3
2024 A Survey on Efficient Vision Transformers: Algorithms, Techniques, and Performance Benchmarking
abstract
Vision Transformer (ViT) architectures are becoming increasingly popular and widely employed to tackle computer vision applications. Their main feature is the capacity to extract global information through the self-attention mechanism, outperforming earlier convolutional neural networks. However, ViT deployment and performance have grown steadily with their size, number of trainable parameters, and operations. Furthermore, self-attention's computational and memory cost quadratically increases with the image resolution. Generally speaking, it is challenging to employ these architectures in real-world applications due to many hardware and environmental restrictions, such as processing and computational capabilities. Therefore, this survey investigates the most efficient methodologies to ensure sub-optimal estimation performances. More in detail, four efficient categories will be analyzed: compact architecture, pruning, knowledge distillation, and quantization strategies. Moreover, a new metric called Efficient Error Rate has been introduced in order to normalize and compare models' features that affect hardware devices at inference time, such as the number of parameters, bits, FLOPs, and model size. Summarizing, this paper first mathematically defines the strategies used to make Vision Transformer efficient, describes and discusses state-of-the-art methodologies, and analyzes their performances over different application scenarios. Toward the end of this paper, we also discuss open challenges and promising research directions.
Lorenzo Papa, Paolo Russo 0001, Irene Amerini, Luping Zhou
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Editorial for pattern recognition letters special issue on Advances in Disinformation Detection and Media Forensics
Irene Amerini, Victor Sanchez, Luca Maiano
Pattern Recognit. Lett.1
2024 A guided-based approach for deepfake detection: RGB-depth integration via features fusion
abstract
Deep fake technology paves the way for a new generation of super realistic artificial content. While this opens the door to extraordinary new applications, the malicious use of deepfakes allows for far more realistic disinformation attacks than ever before. In this paper, we start from the intuition that generating fake content introduces possible inconsistencies in the depth of the generated images. This extra information provides valuable spatial and semantic cues that can reveal inconsistencies facial generative methods introduce. To test this idea, we evaluate different strategies for integrating depth information into an RGB detector and we propose an attention mechanism that makes it possible to integrate information from depth effectively. In addition to being more accurate than an RGB model, our Masked Depthfake Network method is +3.2% more robust against common adversarial attacks on average than a typical RGB detector. Furthermore, we show how this technique allows the model to learn more discriminative features than RGB alone.
Giorgio Leporoni, Luca Maiano, Lorenzo Papa, Irene Amerini
Pattern Recognit. Lett.4
2024 D4D: An RGBD Diffusion Model to Boost Monocular Depth Estimation
abstract
Ground-truth RGBD data are fundamental for a wide range of computer vision applications; however, those labeled samples are difficult to collect and time-consuming to produce. A common solution to overcome this lack of data is to employ graphic engines to produce synthetic proxies; however, those data do not often reflect real-world images, resulting in poor performance of the trained models at the inference step. In this paper we propose a novel training pipeline that incorporates Diffusion4D (D4D), a customized 4-channels diffusion model able to generate realistic RGBD samples. We show the effectiveness of the developed solution in improving the performances of deep learning models on the monocular depth estimation task, where the correspondence between RGB and depth map is crucial to achieving accurate measurements. Our supervised training pipeline, enriched by the generated samples, outperforms synthetic and original data performances achieving an RMSE reduction of (8.2%, 11.9%) and (8.1%, 6.1%) respectively on the indoor NYU Depth v2 and the outdoor KITTI dataset.
Lorenzo Papa, Paolo Russo 0001, Irene Amerini
IEEE Trans. Circuits Syst. Video Technol.3
2023 Convolutional neural networks combined with feature selection for radio-frequency fingerprinting
abstract
Abstract Radio‐frequency fingerprinting is a technique for the authentication and identification of wireless devices using their intrinsic physical features and an analysis of the digitized signal collected during transmission. The technique is based on the fact that the unique physical features of the devices generate discriminating features in the transmitted signal, which can then be analyzed using signal‐processing and machine‐learning algorithms. Deep learning and more specifically convolutional neural networks (CNNs) have been successfully applied to the problem of radio‐frequency fingerprinting using a spectral domain representation of the signal. A potential problem is the large size of the data to be processed, because this size impacts on the processing time during the application of the CNN. We propose an approach to addressing this problem, based on dimensionality reduction using feature‐selection algorithms before the spectrum domain representation is given as an input to the CNN. The approach is applied to two public data sets of radio‐frequency devices using different feature‐selection algorithms for different values of the signal‐to‐noise ratio. The results show that the approach is able to achieve not only a shorter processing time; it also provides a superior classification performance in comparison to the direct application of CNNs.
Gianmarco Baldini, Irene Amerini, Franc Dimc, Fausto Bonavitacola
Comput. Intell.2
2023 A deep-learning-based antifraud system for car-insurance claims
abstract
The annual cost of vehicle insurance fraud is estimated to exceed 40 billion dollars. This is an enormous amount considering the number of new vehicles insured yearly. In terms of higher premiums, it implies that insurance fraud incurs an additional annual cost to each U.S. family of $400 to $700, on average. Many frauds can be attributed to previously reported damages, which are submitted a second time to the insurance company. In these cases, it does not suffice to check the customer’s history to identify them: Damaged car panels can be removed from one vehicle and reassembled on another to make an insurance claim on the second car. To deal with these fraud attempts, in this paper, we propose an end-to-end solution to support the special investigation unit of the insurance companies in their antifraud investigations. For each claim, we organize the images sent to the insurance company and analyze them to extract basic vehicle information. Subsequently, we use these images to identify any damage to the bodywork, and, finally, we verify that the damage has not already been processed in previous claims. To the best of our knowledge, this is the first published work that deals with this problem through a pipeline that covers the entire claim management process. To validate our proposal, we compare our solution with other state-of-the-art models for estimating image similarity. Our results show that our solution is, on average, superior by 15.27% in terms of mean average precision (mAP). In addition, we report the challenges faced in scaling such a system in a production environment. This aspect is often ignored, but applying these solutions to industrial settings is of fundamental importance. Our proposed end-to-end system can reduce by up to 18% the number of false positives produced by the damage reidentification system. We show that encapsulating several specialized components and merging their intermediate results leads to a 72% reduction in possible alerts. Finally, to support the discussion and comparison of explanations for this new task, we introduce a new dataset as a benchmark for damage reidentification.
Luca Maiano, Antonio Montuschi, Marta Caserio, Egon Ferri, Federico Kieffer, Chiara Germanò, Lorenzo Baiocco, Lorenzo Ricciardi Celsi, Irene Amerini, Aris Anagnostopoulos
Expert Syst. Appl.9
2023 METER: A Mobile Vision Transformer Architecture for Monocular Depth Estimation
abstract
Depth estimation is a fundamental knowledge for autonomous systems that need to assess their own state and perceive the surrounding environment. Deep learning algorithms for depth estimation have gained significant interest in recent years, owing to the potential benefits of this methodology in overcoming the limitations of active depth sensing systems. Moreover, due to the low cost and size of monocular cameras, researchers have focused their attention on monocular depth estimation (MDE), which consists in estimating a dense depth map from a single RGB video frame. State of the art MDE models typically rely on vision transformers (ViT) architectures that are highly deep and complex, making them unsuitable for fast inference on devices with hardware constraints. Purposely, in this paper, we address the problem of exploiting ViT in MDE on embedded devices. Those systems are usually characterized by limited memory capabilities and low-power CPU/GPU. We propose METER, a novel lightweight vision transformer architecture capable of achieving state of the art estimations and low latency inference performances on the considered embedded hardwares: NVIDIA Jetson TX1 and NVIDIA Jetson Nano. We provide a solution consisting of three alternative configurations of METER, a novel loss function to balance pixel estimation and reconstruction of image details, and a new data augmentation strategy to improve the overall final predictions. The proposed method outperforms previous lightweight works over the two benchmark datasets: the indoor NYU Depth v2 and the outdoor KITTI.
Lorenzo Papa, Paolo Russo 0001, Irene Amerini
IEEE Trans. Circuits Syst. Video Technol.3
2022 PRaNA: PRNU-based Technique to Tell Real and Deepfake Videos Apart
abstract
Videos are a powerful source of communication adopted in several contexts and used for both benign and malicious purposes (e.g., education vs. reputation damage). Nowadays, realistic video manipulation strategies like deepfake generators constitute a severe threat to our society in term of misinformation. While the wide range of the current research focuses on deepfake detection as a binary task, the identification of a real video among a pool of deepfakes sharing the same origin is not widely investigated. While the pool task might be more rare in real-life compared to the binary one, outcomes that derives from these analyses might let us better understand deepfake behaviours, benefiting binary deep fake detection as well. In this paper, we address the less investigated scenario by investigating the role of Photo Response Non-Uniformity (PRNU) in deepfake detection. Our analysis, in agreement with prior studies, shows that PRNU can be a valuable source to identify deepfake videos. In particular, we found that unique PRNU characteristics exist to distinguish real videos from their deepfake versions: real video autocorrelations tend to be lower compared to their deepfakes versions. Motivated by this, we propose PRaNA, a training-free strategy that leverages PRNU autocorrelation. Our results on three well-known datasets confirm our algorithm's robustness and transferability, with accuracy up to 66% when considering one real video in a pool of four deepfakes using the real video as a source, and up to 80% when only one deepfake is considered. Our work aims to open different strategies to counter deepfake diffusion.
Irene Amerini, Mauro Conti, Pietro Giacomazzi, Luca Pajola
IJCNN1
2022 Online Distributed Denial of Service (DDoS) intrusion detection based on adaptive sliding window and morphological fractal dimension
abstract
Distributed Denial of Service (DDOS) attacks are important threats to network services and applications. Studies in literature have proposed various approaches including Intrusion Detection Systems (IDS) based on the application of machine learning and deep learning, but their computational cost can be significant. For this reason, other studies have proposed efficient IDS algorithms based on the online real-time analysis of the network traffic with a sliding window and entropy or other statistical measures. This paper proposes an online algorithm based on a sliding window with the novel application of the Morphological Fractal Dimension (MFD) to this problem. The results presented in this study show that the application of MFD to the recent CICIDS2017 public data set can obtain a significant improvement in the detection of the DDoS attack in comparison to entropy based approaches. In addition, this paper proposes a novel algorithm for the automatic definition of the sliding window size. This paper reports the impact of the different hyper-parameters, including the parameters present in the definition of MFD and the evaluation of the distance measures, where the Chebyschev distance provides the optimal detection accuracy. The results show a detection accuracy of 99.30%, which performs better than similar approaches on the same data set.
Gianmarco Baldini, Irene Amerini
Comput. Networks2
2021 Learning Double-Compression Video Fingerprints Left From Social-Media Platforms
abstract
Social media and messaging apps have become major communication platforms. Multimedia contents promote improved user engagement and have thus become a very important communication tool. However, fake news and manipulated content can easily go viral, so, being able to verify the source of videos and images as well as to distinguish between native and downloaded content becomes essential. Most of the work performed so far on social media provenance has concentrated on images; in this paper, we propose a CNN architecture that analyzes video content to trace videos back to their social network of origin. The experiments demonstrate that stating platform provenance is possible for videos as well as images with very good accuracy.
Irene Amerini, Aris Anagnostopoulos, Luca Maiano, Lorenzo Ricciardi Celsi
ICASSP1
2021 Media forensics on social media platforms: a survey
abstract
Abstract The dependability of visual information on the web and the authenticity of digital media appearing virally in social media platforms has been raising unprecedented concerns. As a result, in the last years the multimedia forensics research community pursued the ambition to scale the forensic analysis to real-world web-based open systems. This survey aims at describing the work done so far on the analysis of shared data, covering three main aspects: forensics techniques performing source identification and integrity verification on media uploaded on social networks, platform provenance analysis allowing to identify sharing platforms, and multimedia verification algorithms assessing the credibility of media objects in relation to its associated textual information. The achieved results are highlighted together with current open issues and research challenges to be addressed in order to advance the field in the next future.
Cecilia Pasquini, Irene Amerini, Giulia Boato
EURASIP J. Inf. Secur.2
2021 Optical Flow based CNN for detection of unlearnt deepfake manipulations
abstract
A new phenomenon named Deepfakes constitutes a serious threat in video manipulation. AI-based technologies have provided easy-to-use methods to create extremely realistic videos. On the side of multimedia forensics, being able to individuate this kind of fake contents becomes ever more crucial. In this work, a new forensic technique able to detect fake and original video sequences is proposed; it is based on the use of CNNs trained to distinguish possible motion dissimilarities in the temporal structure of a video sequence by exploiting optical flow fields. The results obtained highlight comparable performances with the state-of-the-art methods which, in general, only resort to single video frames. Furthermore, the proposed optical flow based detection scheme also provides a superior robustness in the more realistic cross-forgery operative scenario and can even be combined with frame-based approaches to improve their global effectiveness.
Roberto Caldelli, Leonardo Galteri, Irene Amerini, Alberto Del Bimbo
Pattern Recognit. Lett.3
2020 Exploiting Prediction Error Inconsistencies through LSTM-based Classifiers to Detect Deepfake Videos
abstract
The ability of artificial intelligence techniques to build synthesized brand new videos or to alter the facial expression of already existing ones has been efficiently demonstrated in the literature. The identification of such new threat generally known as Deepfake, but consisting of different techniques, is fundamental in multimedia forensics. In fact this kind of manipulated information could undermine and easily distort the public opinion on a certain person or about a specific event. Thus, in this paper, a new technique able to distinguish synthetic generated portrait videos from natural ones is introduced by exploiting inconsistencies due to the prediction error in the re-encoding phase. In particular, features based on inter-frame prediction error have been investigated jointly with a Long Short-Term Memory (LSTM) model network able to learn the temporal correlation among consecutive frames. Preliminary results have demonstrated that such sequence-based approach, used to distinguish between original and manipulated videos, highlights promising performances.
Irene Amerini, Roberto Caldelli
IH&MMSec1
2019 Tracking Multiple Image Sharing on Social Networks
abstract
Social Networks (SN) and Instant Messaging Apps (IMA) are more and more engaging people in their personal relations taking possession of an important part of their daily life. Huge amounts of multimedia contents, mainly photos, are poured and successively shared on these networks so quickly that is not possible to follow their paths. This last issue surely grants anonymity and impunity thus it consequently makes easier to commit crimes such as reputation attack and cyberbullying. In fact, contents published within a restricted group of friends on an IMA can be rapidly delivered and viewed on a SN by acquaintances and then by strangers without any sort of tracking. In a forensic scenario (e.g., during an investigation), succeeding in understanding this flow could be strategic, thus allowing to reveal all the intermediate steps a certain content has followed. This work aims at tracking multiple sharing on social networks, by extracting specific traces left by each SN within the image file, due to the process each of them applies, to perform a multi-class classification. Innovative strategies, based on deep learning, are proposed and satisfactory results are achieved in recovering till triple up-downloads.
Quoc-Tin Phan, Giulia Boato, Roberto Caldelli, Irene Amerini
ICASSP4
2019 Special issue on Deep Learning in Image and Video Forensics
Roberto Caldelli, Marc Chaumont, Chang-Tsun Li, Irene Amerini
Signal Process. Image Commun.4
2017 Media trustworthiness verification and event assessment through an integrated framework: a case-study
Irene Amerini, Rudy Becarelli, Francesco Brancati, Roberto Caldelli, Gabriele Giunta, Massimiliano Leone Itria
Multim. Tools Appl.1
2017 Dealing with video source identification in social networks
Irene Amerini, Roberto Caldelli, Andrea Del Mastio, Andrea Di Fuccia, Cristiano Molinari, Anna Paola Rizzo
Signal Process. Image Commun.1
2017 Smartphone Fingerprinting Combining Features of On-Board Sensors
abstract
Many everyday activities involve the exchange of confidential information through the use of a smartphone in mobility, i.e., sending on e-mail, checking bank account, buying on-line, accessing cloud platforms, and health monitoring. This demonstrates how security issues related to these operations are a major challenge in our society and in particular in the cyber-security domain. This paper focuses on the use of the smartphone intrinsic and physical characteristics as a mean to build a smartphone fingerprint to enable devices identification. The basic idea proposed in this paper is to investigate how to generate a specific fingerprint that allows to distinctively and reliably characterize each smartphone. In particular, the accelerometer, the gyroscope, the magnetometer, and the audio system (microphone-speaker) are taken into account to build up a composite fingerprint based on a set of their distinctive features. Many experiments have been carried out by analyzing different classification methods, diverse features combination configurations, and operative scenarios. Satisfactory results have been obtained showing that the combination of such sensors improves smartphone distinctiveness.
Irene Amerini, Rudy Becarelli, Roberto Caldelli, Alessio Melani, Moreno Niccolai
IEEE Trans. Inf. Forensics Secur.1
2017 Image Origin Classification Based on Social Network Provenance
abstract
Recognizing information about the origin of a digital image has been individuated as a crucial task to be tackled by the image forensic scientific community. Understanding something on the previous history of an image could be strategic to address any successive assessment to be made on it: knowing the kind of device used for acquisition or, better, the model of the camera could focus investigations in a specific direction. Sometimes just revealing that a determined post-processing, such as an interpolation or a filtering, has been performed on an image could be of fundamental importance to go back to its provenance. This paper locates in such a context and proposes an innovative method to inquire if an image derives from a social network and, in particular, try to distinguish from, which one has been downloaded. The technique is based on the assumption that each social network applies a peculiar and mostly unknown manipulation that, however, leaves some distinctive traces on the image; such traces can be extracted to feature every platform. By resorting at trained classifiers, the presented methodology is satisfactorily able to discern different social network origins. Experimental results carried out on diverse image datasets and in various operative conditions witness that such a distinction is possible. In addition, the proposed method is also able to go back to the original JPEG quality factor the image had before being uploaded on a social network.
Roberto Caldelli, Rudy Becarelli, Irene Amerini
IEEE Trans. Inf. Forensics Secur.3
2015 Acquisition source identification through a blind image classification
abstract
Image forensics, besides understanding if a digital image has been forged, often aims at determining information about image origin. In particular, it could be worthy to individuate which is the kind of source (digital camera, scanner or computer graphics software) that has generated a certain photo. Such an issue has already been studied in literature, but the problem of doing that in a blind manner has not been faced so far. It is easy to understand that in many application scenarios information at disposal is usually very limited; this is the case when, given a set of L images, the authors want to establish if they belong to K different classes of acquisition sources, without having any previous knowledge about the number of specific types of generation processes. The proposed system is able, in an unsupervised and fast manner, to blindly classify a group of photos without neither any initial information about their membership nor by resorting at a trained classifier. Experimental results have been carried out to verify actual performances of the proposed methodology and a comparative analysis with two SVM‐based clustering techniques has been performed too.
Irene Amerini, Rudy Becarelli, B. Bertini, Roberto Caldelli
IET Image Process.1
2014 Exploiting perceptual quality issues in countering SIFT-based Forensic methods
abstract
Scale Invariant Feature Transform (SIFT) has been widely employed in several image application domains, including Image Forensics (e.g. detection of copy-move forgery or near duplicates). Recently, a number of methods allowing to remove SIFT keypoints from an original image have been devised studying the problem of SIFT security against malicious procedures. Such techniques are quite effective in producing an attacked image with very few (or no) keypoints, but at the expense of an image distortion. Final perceptual quality has been taken in account very roughly so far. In this paper, effectiveness of the attacking methods is evaluated also from the side of perceptual image quality; a new version of a SIFT keypoint removal method, based on a perceptual metric, is presented and an extended series of perceptive experiments is reported.
Irene Amerini, Federica Battisti, Roberto Caldelli, Marco Carli, Andrea Costanzo
ICASSP1
2014 Blind image clustering based on the Normalized Cuts criterion for camera identification
Irene Amerini, Roberto Caldelli, Pierluigi Crescenzi, Andrea Del Mastio, Andrea Marino 0001
Signal Process. Image Commun.1
2014 Forensic Analysis of SIFT Keypoint Removal and Injection
abstract
Attacks capable of removing SIFT keypoints from images have been recently devised with the intention of compromising the correct functioning of SIFT-based copy-move forgery detection. To tackle with these attacks, we propose three novel forensic detectors for the identification of images whose SIFT keypoints have been globally or locally removed. The detectors look for inconsistencies like the absence or anomalous distribution of keypoints within textured image regions. We first validate the methods on state-of-the-art keypoint removal techniques, then we further assess their robustness by devising a counter-forensic attack injecting fake SIFT keypoints in the attempt to cover the traces of removal. We apply the detectors to a practical image forensic scenario of SIFT-based copy-move forgery detection, assuming the presence of a counterfeiter who resorts to keypoint removal and injection to create copy-move forgeries that successfully elude SIFT-based detectors but are in turn exposed by the newly proposed tools.
Andrea Costanzo, Irene Amerini, Roberto Caldelli, Mauro Barni
IEEE Trans. Inf. Forensics Secur.2
2013 SIFT keypoint removal and injection for countering matching-based image forensics
abstract
Scale Invariant Feature Transform (SIFT) has been widely employed in several image application domains, including Image Forensics (e.g. detection of copy-move forgery or near duplicates). Until now, the research community has focused on studying the robustness of SIFT against legitimate image processing, but rarely concerned itself with the problem of SIFT security against malicious procedures. Recently, a number of methods allowing to remove SIFT keypoints from an original image have been devised. Although quite effective, such methods produce an attacked image with very few (or no) keypoints, thus leaving cues that can be easily exploited by a forensic analyst to reveal the occurred manipulation. In this paper, we explore the topic of reintroducing fake SIFT keypoints into a previously cleaned image in order to address the main weakness of the existing removal attacks. In particular, we evaluate the fitness of locally adaptive contrast enhancement methods to the task of injecting new keypoints. The results we obtained are encouraging: (i) it is possible to effectively introduce new keypoints whose descriptors do not match with those of the original image, thus concealing the removal forgery; (ii) the perceptual quality of the image following the removal and injection attacks is comparable to the one of the original image.
Irene Amerini, Mauro Barni, Roberto Caldelli, Andrea Costanzo
IH&MMSec1
2013 Removal and injection of keypoints for SIFT-based copy-move counter-forensics
abstract
Abstract Recent studies exposed the weaknesses of scale-invariant feature transform (SIFT)-based analysis by removing keypoints without significantly deteriorating the visual quality of the counterfeited image. As a consequence, an attacker can leverage on such weaknesses to impair or directly bypass with alarming efficacy some applications that rely on SIFT. In this paper, we further investigate this topic by addressing the dual problem of keypoint removal, i.e., the injection of fake SIFT keypoints in an image whose authentic keypoints have been previously deleted. Our interest stemmed from the consideration that an image with too few keypoints is per se a clue of counterfeit, which can be used by the forensic analyst to reveal the removal attack. Therefore, we analyse five injection tools reducing the perceptibility of keypoint removal and compare them experimentally. The results are encouraging and show that injection is feasible without causing a successive detection at SIFT matching level. To demonstrate the practical effectiveness of our procedure, we apply the best performing tool to create a forensically undetectable copy-move forgery, whereby traces of keypoint removal are hidden by means of keypoint injection.
Irene Amerini, Mauro Barni, Roberto Caldelli, Andrea Costanzo
EURASIP J. Inf. Secur.1
2013 Copy-move forgery detection and localization by means of robust clustering with J-Linkage
Irene Amerini, Lamberto Ballan, Roberto Caldelli, Alberto Del Bimbo, Luca Del Tongo, Giuseppe Serra 0001
Signal Process. Image Commun.1
2011 A SIFT-Based Forensic Method for Copy-Move Attack Detection and Transformation Recovery
abstract
One of the principal problems in image forensics is determining if a particular image is authentic or not. This can be a crucial task when images are used as basic evidence to influence judgment like, for example, in a court of law. To carry out such forensic analysis, various technological instruments have been developed in the literature. In this paper, the problem of detecting if an image has been forged is investigated; in particular, attention has been paid to the case in which an area of an image is copied and then pasted onto another zone to create a duplication or to cancel something that was awkward. Generally, to adapt the image patch to the new context a geometric transformation is needed. To detect such modifications, a novel methodology based on scale invariant features transform (SIFT) is proposed. Such a method allows us to both understand if a copy-move attack has occurred and, furthermore, to recover the geometric transformation used to perform cloning. Extensive experimental results are presented to confirm that the technique is able to precisely individuate the altered area and, in addition, to estimate the geometric transformation parameters with high reliability. The method also deals with multiple cloning.
Irene Amerini, Lamberto Ballan, Roberto Caldelli, Alberto Del Bimbo, Giuseppe Serra 0001
IEEE Trans. Inf. Forensics Secur.1
2010 Geometric tampering estimation by means of a SIFT-based forensic analysis
abstract
In many application scenarios digital images play a basic role and often it is important to assess if their content is realistic or has been manipulated to mislead watcher's opinion. Image forensics tools provide answers to similar questions. This paper, in particular, focuses on the problem of detecting if a feigned image has been created by cloning an area of the image onto another zone to make a duplication or to cancel something awkward. The proposed method is based on SIFT features and allows both to understand which are the image points involved in the counterfeit attack and, furthermore, to recover the parameters of the geometric transformation. Experimental results are provided to witness the powerfulness of the proposed technique.
Irene Amerini, Lamberto Ballan, Roberto Caldelli, Alberto Del Bimbo, Giuseppe Serra 0001
ICASSP1
2010 A DVB-MHP web browser to pursue convergence between Digital Terrestrial Television and Internet
Irene Amerini, Giovanni Ballocca, Rudy Becarelli, Roberto Borri, Roberto Caldelli, Francesco Filippini
Multim. Tools Appl.1
2009 Integration between Digital Terrestrial Television and Internet by Means of a DVB-MHP Web Browser
Irene Amerini, Roberto Caldelli, Rudy Becarelli, Francesco Filippini, Giovanni Ballocca, Roberto Borri
WEBIST1