Paolo Bestagini

dblp:53/11030 · DBLP profile ↗
← Back
85ranked-venue papers
7as first author
45since 2021 · last 2026
0000-0003-0406-0222ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 57 · 6 first-author · 23 since 2021Security and privacy · 12 · 9 since 2021Artificial intelligence and machine learning · 11 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Lightweight On-Device Anti-Spoofing Detection using Ternary Neural Networks
abstract
In recent years, the security risks posed by speech deepfakes have increased significantly, as synthetic speech can convincingly impersonate target identities. Although numerous deepfake detectors have been proposed, many rely on deep learning architectures that remain computationally demanding for deployment on resource-constrained devices such as smartphones. In this work, we investigate ternary quantization as a strategy to enable efficient on-device speech deepfake detection. We apply a dynamic ternary quantization scheme to a LCNN architecture, obtaining a model with 98.92% sparsity. The resulting network achieves a 98.8% reduction in multiply–accumulate operations and a 55.5% decrease in storage size compared to a lightweight state-of-the-art baseline. Experimental results show minimal degradation relative to the full-precision counterpart, competitive in-domain performance, and improved cross-dataset generalization, while maintaining robustness under codec compression and reverberation distortions. These findings indicate that drastic quantization can drastically reduce computational and memory requirements without substantially compromising detection capability, paving the way for practical real-time speech deepfake detection on edge devices.
Matteo Benzo, Davide Salvi, Paolo Bestagini, Stefano Tubaro
IH&MMSec3
2026 Forensic Similarity for Speech Deepfakes
abstract
In this paper, we introduce the concept of forensic similarity in the speech deepfake detection domain, which aims to determine whether two audio segments share the same underlying forensic traces. Our approach is inspired by prior work in the image domain. To transfer this idea to the audio domain, we propose a two-stage deep learning framework consisting of a Siamese-based feature extractor and a core decision module, referred to as the similarity network. The system goal to assess whether two speech samples originate from the same source by comparing their forensic characteristics. In practice, the model maps pairs of audio segments to a similarity score indicating whether they contain identical or different forensic traces. We evaluate the proposed method on the emerging task of source verification, demonstrating its ability to determine whether two speech samples were generated by the same model. In addition, we explore its applicability to audio splicing detection as a complementary use case. Experimental results show that the proposed approach generalizes well to previously unseen forensic traces, highlighting its robustness, flexibility, and practical relevance for digital audio forensics.
Viola Negroni, Davide Salvi, Daniele Ugo Leonzio, Paolo Bestagini, Stefano Tubaro
IH&MMSec4
2026 Interpretable detection of singing voice manipulations using audio-language models
abstract
Singing voice manipulations have become increasingly common in modern music production. While such techniques can serve as creative tools to enhance artists’ expressive possibilities, they can also raise concerns about content authenticity and media integrity. To counter potential misuse of vocal manipulation tools, recent research has developed detection systems. However, these are typically limited to binary classification, indicating only whether a vocal track has been altered, without providing any further interpretable information. In this work, we address this limitation and propose a novel framework that combines forensic audio analysis with natural language generation to both detect and describe modifications in singing voice signals. Building on recent advances in audio-language models, we construct a dataset of manipulated and synthetic vocals annotated with detailed textual annotations, which we use to train and evaluate our framework. Our approach identifies and characterises a wide range of vocal transformations, including pitch correction, pitch shifting, time stretching, and singing voice deepfake generation. Experimental results show that the proposed method not only surpasses existing baselines in classification accuracy but also provides substantially greater interpretability, as it provides explanations of the outputs in natural language, making them understandable to non-experts. This makes the system particularly relevant for music production, media forensics, and copyright verification, offering a transparent and descriptive account of vocal alterations.
Mahyar Gohari, Davide Salvi, Paolo Bestagini, Nicola Adami
Comput. Vis. Image Underst.3
2026 Splicing detection and localization for speech deepfakes using audio novelty
Davide Salvi, Francesco Castelli, Viola Negroni, Paolo Bestagini, Stefano Tubaro
Comput. Vis. Image Underst.4
2025 DiffSSD: A Diffusion-Based Dataset For Speech Forensics
abstract
Diffusion-based speech generators are ubiquitous. These methods can generate very high quality synthetic speech and several recent incidents report their malicious use. To counter such misuse, synthetic speech detectors have been developed. Many of these detectors are trained on datasets which do not include diffusion-based synthesizers. In this paper, we demonstrate that existing detectors trained on one such dataset, ASVspoof2019, do not perform well in detecting synthetic speech from recent diffusion-based synthesizers. We propose the Diffusion-Based Synthetic Speech Dataset (DiffSSD), a dataset consisting of about 200 hours of labeled speech, including synthetic speech generated by 8 diffusion-based open-source and 2 commercial generators. We also examine the performance of existing synthetic speech detectors on DiffSSD in both closed-set and open-set scenarios. The results highlight the importance of this dataset in detecting synthetic speech generated from recent open-source and commercial speech generators.
Kratika Bhagtani, Amit Kumar Singh Yadav, Paolo Bestagini, Edward J. Delp
ICASSP3
2025 Audio Features Investigation for Singing Voice Deepfake Detection
abstract
The audio forensics field has recently faced a new challenge: singing voice deepfake detection. Current approaches to tackle this problem have borrowed methods initially developed for the more established task of speech deepfake detection, often simply retraining these systems on singing voice data. However, effective speech detection techniques may not necessarily perform well on singing voice, and there has been limited research on identifying the factors that can improve detection specifically in the singing domain. This paper investigates the effectiveness of various audio representations and features for discriminating real and synthetically generated singing voice signals. We evaluate two Convolutional Neural Network (CNN)-based detection systems using a wide range of audio representations, including handcrafted, learning-based, and pre-trained features. Through a systematic analysis, we aim to understand the key factors that can improve the performance of deepfake detection methods for singing voices. Additionally, we investigate the differences between singing voice and speech detection, highlighting the implications of the feature sets considered. Our results offer valuable insights and guidance for developing more advanced and effective singing voice deepfake detection systems in the future.
Mahyar Gohari, Davide Salvi, Paolo Bestagini, Nicola Adami
ICASSP3
2025 Leveraging Mixture of Experts for Improved Speech Deepfake Detection
abstract
Speech deepfakes pose a significant threat to personal security and content authenticity. Several detectors have been proposed in the literature, and one of the primary challenges these systems have to face is the generalization over unseen data to identify fake signals across a wide range of datasets. In this paper, we introduce a novel approach for enhancing speech deepfake detection performance using a Mixture of Experts architecture. The Mixture of Experts framework is well-suited for the speech deepfake detection task due to its ability to specialize in different input types and handle data variability efficiently. This approach offers superior generalization and adaptability to unseen data compared to traditional single models or ensemble methods. Additionally, its modular structure supports scalable updates, making it more flexible in managing the evolving complexity of deepfake techniques while maintaining high detection accuracy. We propose an efficient, lightweight gating mechanism to dynamically assign expert weights for each input, optimizing detection performance. Experimental results across multiple datasets demonstrate the effectiveness and potential of our proposed approach.
Viola Negroni, Davide Salvi, Alessandro Ilic Mezza, Paolo Bestagini, Stefano Tubaro
ICASSP4
2025 Freeze and Learn: Continual Learning with Selective Freezing for Speech Deepfake Detection
abstract
In speech deepfake detection, one of the critical aspects is developing detectors able to generalize on unseen data and distinguish fake signals across different datasets. Common approaches to this challenge involve incorporating diverse data into the training process or fine-tuning models on unseen datasets. However, these solutions can be computationally demanding and may lead to the loss of knowledge acquired from previously learned data. Continual learning techniques offer a potential solution to this problem, allowing the models to learn from unseen data without losing what they have already learned. Still, the optimal way to apply these algorithms for speech deepfake detection remains unclear, and we do not know which is the best way to apply these algorithms to the developed models. In this paper we address this aspect and investigate whether, when retraining a speech deepfake detector, it is more effective to apply continual learning across the entire model or to update only some of its layers while freezing others. Our findings, validated across multiple models, indicate that the most effective approach among the analyzed ones is to update only the weights of the initial layers, which are responsible for processing the input features of the detector.
Davide Salvi, Viola Negroni, Luca Bondi, Paolo Bestagini, Stefano Tubaro
ICASSP4
2025 WILD: a new in-the-Wild Image Linkage Dataset for synthetic image attribution
abstract
Synthetic image source attribution is an open challenge, with an increasing number of image generators being released yearly. The complexity and the sheer number of available generative techniques, as well as the scarcity of high-quality open source datasets of diverse nature for this task, make training and benchmarking synthetic image source attribution models very challenging. WILD1is a new in-the-Wild Image Linkage Dataset designed to provide a powerful training and benchmarking tool for synthetic image attribution models. The dataset is built out of a closed set of 10 popular commercial generators, which constitutes the training base of attribution models, and an open set of 10 additional generators, simulating a real-world in-the-wild scenario. Each generator is represented by 1,000 images, for a total of 10,000 images in the closed set and 10,000 images in the open set. Half of the images are post-processed with a wide range of operators. WILD allows benchmarking attribution models in a wide range of tasks, including closed and open set identification and verification, and robust attribution with respect to post-processing and adversarial attacks. Models trained on WILD are expected to benefit from the challenging scenario represented by the dataset itself. Moreover, an assessment of seven baseline methodologies on closed and open set attribution is presented, including robustness tests with respect to post-processing.
Pietro Bongini, Sara Mandelli, Andrea Montibeller, Mirko Casu, Orazio Pontorno, Claudio Vittorio Ragaglia, Luca Zanchetta, Mattia Aquilina, Taiba Majid Wani, Luca Guarnera, Benedetta Tondi, Giulia Boato, Paolo Bestagini, Irene Amerini, Francesco G. B. De Natale, Sebastiano Battiato, Mauro Barni
IJCNN13
2025 Source Verification for Speech Deepfakes
abstract
With the proliferation of speech deepfake generators, it becomes crucial not only to assess the authenticity of synthetic audio but also to trace its origin. While source attribution models attempt to address this challenge, they often struggle in open-set conditions against unseen generators. In this paper, we introduce the source verification task, which, inspired by speaker verification, determines whether a test track was produced using the same model as a set of reference signals. Our approach leverages embeddings from a classifier trained for source attribution, computing distance scores between tracks to assess whether they originate from the same source. We evaluate multiple models across diverse scenarios, analyzing the impact of speaker diversity, language mismatch, and post-processing operations. This work provides the first exploration of source verification, highlighting its potential and vulnerabilities, and offers insights for real-world forensic applications.
Viola Negroni, Davide Salvi, Paolo Bestagini, Stefano Tubaro
INTERSPEECH3
2025 Enhanced Water Leak Detection with Convolutional Neural Networks and One-Class Support Vector Machine
Daniele Ugo Leonzio, Paolo Bestagini, Marco Marcon, Stefano Tubaro
Networking2
2025 Hiding Local Manipulations on SAR Images: A Counter-Forensic Attack
abstract
The vast accessibility of Synthetic Aperture Radar (SAR) images through online portals has propelled the research across various fields. This widespread use and easy availability have unfortunately made SAR data susceptible to malicious alterations, such as local editing applied to the images for inserting or covering the presence of sensitive targets. To contrast malicious manipulations, in the last years the forensic community has begun to dig into the SAR manipulation issue, proposing detectors that effectively localize the tampering traces in amplitude images. Nonetheless, in this paper we demonstrate that an expert practitioner can exploit the complex nature of SAR data to obscure any signs of manipulation within a locally altered amplitude image. We refer to this approach as a counter-forensic attack. To achieve the concealment of manipulation traces, the attacker can simulate a re-acquisition of the manipulated scene by the SAR system that initially generated the pristine image. In doing so, the attacker can obscure any evidence of manipulation, making it appear as if the image was legitimately produced by the system. This attack has unique features that make it both highly generalizable and relatively easy to apply. First, it is a black-box attack, meaning it is not designed to deceive a specific forensic detector. Furthermore, it does not require a training phase and is not based on adversarial operations. We assess the effectiveness of the proposed counter-forensic approach across diverse scenarios, examining various manipulation operations. The obtained results indicate that our devised attack successfully eliminates traces of manipulation, deceiving even the most advanced forensic detectors.
Sara Mandelli, Edoardo Daniele Cannas, Paolo Bestagini, Stefano Tebaldini, Stefano Tubaro
IEEE Trans. Image Process.3
2024 A One-Class Approach to Detect Super-Resolution Satellite Imagery with Spectral Features
abstract
Satellite imagery has a vital role in many applications and several techniques exist to enhance its quality. An example is given by image Super Resolution (SR), which aims at increasing the pixel resolution to recover lost high-frequency details. Due to their importance, satellite images are also a target for malicious manipulations. In such a context, knowing if an image has been super-resolved (SRV) is crucial for guaranteeing the correct usage of this kind of data. In this paper, we propose a pipeline for detecting satellite images SRV through State-Of-The-Art (SOTA) techniques based on Convolutional Neural Networks. Our solution employs an anomaly detection algorithm, i.e., a one-class Support Vector Machine trained on Fourier spectral features of native high-resolution images. Even though we never process SRV samples during training, the experiments show that we can reject images generated through different SOTA techniques, and our solution is robust against image compression.
Edoardo Daniele Cannas, P. Beaus, Paolo Bestagini, F. Marques, Stefano Tubaro
ICASSP3
2024 Water Leak Detection via Domain Adaptation
abstract
Outdated infrastructure contributes to significant water wastage, where leaks can represent as much as 30% of urban water supply losses. Rapid and precise leak detection is therefore crucial for economic and environmental reasons. Data-driven methods have emerged as promising solutions to detect water leaks due to their accurate performance. However, they encounter obstacles like limited labeled datasets and adapting to various situations. To address these challenges, we explore Semi Supervised Learning (SSL) and Transfer Learning (TL) techniques in the context of water leak detection. We propose to address the problem of leak detection in case of limited labeled data, using a Convolutional Neural Network (CNN) trained on a laboratory-scale network, then adapted to work on real data. To do so, we compare three different domain adaptation techniques that leverage only a small amount of data from the new domain. Our results show that Self-Tuning techniques proves better than the others for this task, even with limited data.
Daniele Ugo Leonzio, Paolo Bestagini, Marco Marcon, Gian Paolo Quarta, Stefano Tubaro
ICASSP2
2024 Mdrt: Multi-Domain Synthetic Speech Localization
abstract
With recent advancements in generating synthetic speech, tools to generate high-quality synthetic speech impersonating any human speaker are easily available. Several incidents report misuse of high-quality synthetic speech for spreading misinformation and for large-scale financial frauds. Many methods have been proposed for detecting synthetic speech; however, there is limited work on localizing the synthetic segments within the speech signal. In this work, our goal is to localize the synthetic speech segments in a partially synthetic speech signal. Most existing methods for synthetic speech localization obtain features from either the time domain waveform or the spectrogram representation of the speech signal. In this work, we propose Multi-Domain ResNet Transformer (MDRT) that obtains multi-domain features from both the time domain and the spectrogram representation of a speech signal to localize synthetic speech segments. MDRT uses transformer neural networks to obtain multi-domain features and processes them using a ResNet-style neural network. We use the PartialSpoof dataset to examine the performance of MDRT on localizing synthetic speech segments of varying duration. Our results show that MDRT performs better than several existing synthetic speech localization methods.
Amit Kumar Singh Yadav, Kratika Bhagtani, Sriram Baireddy, Paolo Bestagini, Stefano Tubaro, Edward J. Delp
ICASSP4
2024 Back to the Future: GNN-Based No2 Forecasting Via Future Covariates
abstract
Due to the latest environmental concerns in keeping at bay contaminants emissions in urban areas, air pollution forecasting has been rising the forefront of all researchers around the world. When predicting pollutant concentrations, it is common to include the effects of environmental factors that influence these concentrations within an extended period, like traffic, meteorological conditions and geographical information. Most of the existing approaches exploit this information as past covariates, i.e., past exogenous variables that affected the pollutant but were not affected by it. In this paper, we present a novel forecasting methodology to predict NO2 concentration via both past and future covariates. Future covariates are represented by weather forecasts and future calendar events, which are already known at prediction time. In particular, we deal with air quality observations in a city-wide network of ground monitoring stations, modeling the data structure and estimating the predictions with a Spatiotemporal Graph Neural Network (STGNN). We propose a conditioning block that embeds past and future covariates into the current observations. After extracting meaningful spatiotemporal representations, these are fused together and projected into the forecasting horizon to generate the final prediction. To the best of our knowledge, it is the first time that future covariates are included in time series predictions in a structured way. Remarkably, we find that conditioning on future weather information has a greater impact than considering past traffic conditions. We release our code implementation at https://github.com/polimi-ispl/MAGCRN.
Antonio Giganti, Sara Mandelli, Paolo Bestagini, Umberto Giuriato, Alessandro D'Ausilio, Marco Marcon, Stefano Tubaro
IGARSS3
2024 Are Recent Deepfake Speech Generators Detectable?
abstract
Deep learning methods can generate high-quality synthetic speech which is perceptually indistinguishable from real human speech. Synthetic speech can be maliciously used for fraud. Synthetic speech detection methods have been proposed which perform well on ASVspoof2019 and ASVspoof2021 Datasets. These datasets consist of synthetic speech from conventional neural network speech generators. Recently, many voice cloning methods have been proposed which use diffusion models and generative adversarial networks for high-quality speech synthesis. In this work, we present a new synthetic speech dataset containing 25,000 synthetic speech signals for 11 distinct speakers, with a total duration of 52 hours. We have developed this dataset using 5 recent diffusion model-based synthetic speech generators. These generators can clone a speaker's voice from text using only a few minutes of their real speech. We evaluate 6 of the best synthetic speech detectors that work well on the ASVspoof2019 Dataset on this new dataset, and demonstrate their performance using Equal Error Rate (EER).
Kratika Bhagtani, Amit Kumar Singh Yadav, Paolo Bestagini, Edward J. Delp
IH&MMSec3
2024 Investigating Translation Invariance and Shiftability in CNNs for Robust Multimedia Forensics: A JPEG Case Study
Edoardo Daniele Cannas, Sara Mandelli, Paolo Bestagini, Stefano Tubaro
IH&MMSec3
2024 Self-Supervised Seismic Swell Noise Suppression From Noisy Seismic Data
abstract
Seismic swell noise, often observed in marine seismic data, is characterized by high amplitude and low frequencies. This noise significantly hides useful signals, underscoring the importance of attenuating it in the processing pipeline for marine seismic data. To date, most deep learning methods proposed for swell noise suppression have focused on supervised paradigms, which require a large dataset of paired noisy and clean data to gain insights into signal features and swell noise statistics at the training phase. In field data processing, however, obtaining generalizable training data often proves to be challenging. The inherent complexity and variability of the real world often make it difficult to acquire unbiased, noise-free datasets in the context of marine seismic data processing. To address this problem, we present a self-supervised deep learning method for suppressing swell noise even in the absence of access to clean seismic training data. First, we propose a strategy for the synthetic swell noise generation based on the amplitude, frequency, and coherence features of noisy traces. We implement the noisy-as-clean (NAC) strategy, wherein either the original noisy seismic data or its reorganized variant acts as the network’s target. Simultaneously, the network receives the observed noisy seismic data combined with simulated swell noise as its input. By leveraging these “noisy-noisy” pairs, we train a DnCNN network. Experimental evaluations conducted on both synthetic and field data demonstrate that, through the integration of swell noise simulation and the NAC strategy, the trained network consistently achieves superior denoising performance.
Weiwei Xu 0004, Vincenzo Lipari, Paolo Bestagini, Stefano Tubaro
IEEE Trans. Geosci. Remote. Sens.3
2023 Water Leak Detection and Localization Using Convolutional Autoencoders
abstract
Water is a valuable resource that has to be handled appropriately. However, a significant volume of water is wasted annually due to leaks in Water Distribution Networks (WDNs). This emphasizes the necessity for reliable and effective leak detection and localization systems. Several types of solutions have been proposed during the last few years. Among these solutions, data-driven ones are gaining more traction due to their impressive performance. In this paper, we propose a new method for leak detection and localization. The method is based on water pressure measurements acquired at a series of nodes of a WDN. Our technique is a fully data-driven solution that makes only use of the knowledge of the WDN topology, and a series of pressure data acquisitions obtained in absence of leaks. The proposed solution is based on an autoencoder trained on no-leak data, so that leaks are detected as anomalies. The results achieved on the LeakDB dataset demonstrate that the proposed solution outperforms recent methods for leak detection and localization.
Daniele Ugo Leonzio, Paolo Bestagini, Marco Marcon, Gian Paolo Quarta, Stefano Tubaro
ICASSP2
2023 Reliability Estimation for Synthetic Speech Detection
abstract
Recent advances in speech synthesis and counterfeit audio generation have pushed the multimedia forensics community to develop speech deepfake detection techniques to avoid threats and unpleasant situations. Although synthetic speech detectors show excellent performance in controlled conditions, they are not always reliable in open set cases, when evaluated on data that are very different from those seen during training. This can lead to misleading scores and poorly indicative results in real-world scenarios. In this paper, we propose a method for estimating the reliability of a prediction performed by a speech deepfake detector. This enables us to perform the detection only on the most relevant portions of a signal, i.e., the time windows on which we obtain more reliable scores. This increases the final accuracy of the developed systems. As some audio fragments may not contain enough traces for the task at hand and negatively affect the system output, a reliability estimator allows us to discard them and focus only on the most pertinent data. The proposed method proves to positively impact the performance of the considered detector and shows excellent generalization capabilities on unseen datasets.
Davide Salvi, Paolo Bestagini, Stefano Tubaro
ICASSP2
2023 ASSD: Synthetic Speech Detection in the AAC Compressed Domain
abstract
Synthetic human speech signals have become very easy to generate given modern text-to-speech methods. When these signals are shared on social media they are often compressed using the Advanced Audio Coding (AAC) standard. Our goal is to study if a small set of coding metadata contained in the AAC compressed bit stream is sufficient to detect synthetic speech. This would avoid decompressing of the speech signals before analysis. We call our proposed method AAC Synthetic Speech Detection (ASSD). ASSD extracts information from the AAC compressed bit stream without decompressing the speech signal. ASSD analyzes the information using a transformer neural network. In our experiments, we compressed the ASVspoof2019 dataset according to the AAC standard using different data rates. We compared the performance of ASSD to a time domain based and a spectrogram based synthetic speech detection methods. We evaluated ASSD on approximately 71k compressed speech signals. The results show that our proposed method typically only requires 1000 bits per speech block/frame from the AAC compressed bit stream to detect synthetic speech. This is much lower than other reported methods. Our method also had a 9.7 percentage points higher detection accuracy compared to existing methods.
Amit Kumar Singh Yadav, Ziyue Xiang, Emily R. Bartusiak, Paolo Bestagini, Stefano Tubaro, Edward J. Delp
ICASSP4
2023 Super-Resolution of BVOC Maps by Adapting Deep Learning Methods
abstract
Biogenic Volatile Organic Compounds (BVOCs) play a critical role in biosphere-atmosphere interactions, being a key factor in the physical and chemical properties of the atmosphere and climate. Acquiring large and fine-grained BVOC emission maps is expensive and time-consuming, so most available BVOC data are obtained on a loose and sparse sampling grid or on small regions. However, high-resolution BVOC data are desirable in many applications, such as air quality, atmospheric chemistry, and climate monitoring. In this work, we investigate the possibility of enhancing BVOC acquisitions, further explaining the relationships between the environment and these compounds. We do so by comparing the performances of several state-of-the-art neural networks proposed for image Super-Resolution (SR), adapting them to overcome the challenges posed by the large dynamic range of the emission and reduce the impact of outliers in the prediction. Moreover, we also consider realistic scenarios, considering both temporal and geographical constraints. Finally, we present possible future developments regarding SR generalization, considering the scale-invariance property and super-resolving emissions from unseen compounds.
Antonio Giganti, Sara Mandelli, Paolo Bestagini, Marco Marcon, Stefano Tubaro
ICIP3
2023 It Wasn't Me: Irregular Identity in Deepfake Videos
abstract
With the rapid development in media generation technologies, the creation of DeepFake videos is within everyone’s reach. As the widespread diffusion of DeepFakes can lead to severe consequences (e.g., defamation, fake news spreading, etc.), detecting DeepFakes is becoming a crucial task within the forensic community. However, most of the existing DeepFake detectors suffer from two issues: i) they are hardly explainable as they build upon black-box data-driven techniques rather than interpretable features; ii) they are often tailored to low-level texture features, failing to generalize on low-quality DeepFake videos. In this work we propose a video DeepFake detector that aims at solving these issues. The proposed detector relies on the fact that most DeepFake generators work on a frame-by-frame basis, thus breaking the temporal consistency of facial features across frames. In particular, we noticed that facial identity features tend to be less stable in time on DeepFake videos than original ones. We therefore propose a framework trained on time series of facial identity features. The use of high-level semantic features makes the detector interpretable and robust against low-quality DeepFake videos. Extensive experiments show that our method achieves outstanding performance on low-quality DeepFake video and obtains promising results on unseen dataset evaluation. The code is available at https://github.com/HongguLiu/Identity-Inconsistency-DeepFake-Detection
Honggu Liu, Paolo Bestagini, Wenbo Zhou 0004, Stefano Tubaro, Weiming Zhang 0001, Nenghai Yu
ICIP2
2023 DSVAE: Disentangled Representation Learning for Synthetic Speech Detection
abstract
Tools to generate high quality synthetic speech that is perceptually indistinguishable from speech recorded from hu-man speakers are easily available. Many incidents report misuse of synthetic speech for spreading misinformation and committing financial fraud. Several approaches have been proposed for detecting synthetic speech. Many of these approaches use deep learning methods without providing reasoning for the decisions they make. This limits the explainability of these approaches. In this paper, we use disentangled representation learning for developing a synthetic speech detector. We propose Disentangled Spectrogram Variational Auto Encoder (DSVAE) which is a two stage trained variational autoencoder that processes spectrograms of speech to generate features that disentangle synthetic and bona fide speech. We evaluated DSVAE using the ASVspoof2019 dataset. Our experimental results show high accuracy (> 98%) on detecting synthetic speech from 6 known and 10 unknown speech synthesizers. Further, the visualization of disentangled features obtained from DSVAE provides rea-soning behind the working principle of DSVAE, improving its explainability. DSVAE performs well compared to several existing methods. Additionally, DSVAE works in practical scenarios such as detecting synthetic speech uploaded on social platforms and against simple attacks such as removing silence regions.
Amit Kumar Singh Yadav, Kratika Bhagtani, Ziyue Xiang, Paolo Bestagini, Stefano Tubaro, Edward J. Delp
ICMLA4
2023 PS3DT: Synthetic Speech Detection Using Patched Spectrogram Transformer
abstract
Many deep learning synthetic speech generation tools are readily available. The use of synthetic speech has caused financial fraud, impersonation of people, and misinformation to spread. For this reason forensic methods that can detect synthetic speech have been proposed. Existing methods often overfit on one dataset and their performance reduces substantially in practical scenarios such as detecting synthetic speech shared on social platforms. In this paper we propose, Patched Spectrogram Synthetic Speech Detection Transformer (PS3DT), a synthetic speech detector that converts a time domain speech signal to a mel-spectrogram and processes it in patches using a trans-former neural network. We evaluate the detection performance of PS3DT on ASVspoof2019 dataset. Our experiments show that PS3DT performs well on ASVspoof2019 dataset compared to other approaches using spectrogram for synthetic speech detection. We also investigate generalization performance of PS3DT on In-the-Wild dataset. PS3DT generalizes well than several existing methods on detecting synthetic speech from an out-of-distribution dataset. We also evaluate robustness of PS3DT to detect telephone quality synthetic speech and synthetic speech shared on social platforms (compressed speech). PS3DT is robust to compression and can detect telephone quality synthetic speech better than several existing methods.
Amit Kumar Singh Yadav, Ziyue Xiang, Kratika Bhagtani, Paolo Bestagini, Stefano Tubaro, Edward J. Delp
ICMLA4
2023 Robust Water Leak Detection and Localization with Graph Signal Processing
abstract
Water is a resource that has to be managed properly. Nevertheless, a sizable amount of water is lost each year because of leaks in Water Distribution Networks (WDNs). The need for trustworthy and efficient leak detection and localization systems is therefore an urgent necessity. For this reason, different solutions have been put out in recent years. Due to their outstanding performance, data-driven methods are among those that are gaining the most popularity. However, the performance of data-driven approaches depend on the coherence between data on which they are trained and data on which they are tested. For example, if the acquired test data look corrupted and incoherent with training ones due to sensor failure, the performance of the overall system may be severely hindered. In this work we present a resilient water leak detection and localization algorithm. It is based on two main steps: the first step analyzes acquired data to possibly recover corrupted ones by means of graph interpolation; the second step finds leaks exploiting an autoencoder-based anomaly detector proposed in the literature. The results show that the suggested approach for signal recovery by means of graph interpolation enables the detector to work in situations in which it would otherwise fail. In doing so, we address a problem that has so far received little attention in the literature: potential sensor failures when acquiring data.
Daniele Ugo Leonzio, Paolo Bestagini, Marco Marcon, Gian Paolo Quarta, Stefano Tubaro
IECON2
2023 Super-Resolution of Bvoc Emission Maps Via Domain Adaptation
abstract
Enhancing the resolution of Biogenic Volatile Organic Compound (BVOC) emission maps is a critical task in remote sensing. Recently, some Super-Resolution (SR) methods based on Deep Learning (DL) have been proposed, leveraging data from numerical simulations for their training process. However, when dealing with data derived from satellite observations, the reconstruction is particularly challenging due to the scarcity of measurements to train SR algorithms with. In our work, we aim at super-resolving low resolution emission maps derived from satellite observations by leveraging the information of emission maps obtained through numerical simulations. To do this, we combine a SR method based on DL with Domain Adaptation (DA) techniques, harmonizing the different aggregation strategies and spatial information used in simulated and observed domains to ensure compatibility. We investigate the effectiveness of DA strategies at different stages by systematically varying the number of simulated and observed emissions used, exploring the implications of data scarcity on the adaptation strategies. To the best of our knowledge, there are no prior investigations of DA in satellite-derived BVOC maps enhancement. Our work represents a first step toward the development of robust strategies for the reconstruction of observed BVOC emissions.
Antonio Giganti, Sara Mandelli, Paolo Bestagini, Marco Marcon, Stefano Tubaro
IGARSS3
2023 Synthesized Speech Attribution Using The Patchout Spectrogram Attribution Transformer
abstract
The malicious use of synthetic speech has increased with the recent availability of speech generation tools. It is important to determine whether a speech signal is authentic (spoken by a human) or is synthesized and to determine the generation method used to create it. Identifying the synthesis method is known as synthetic speech attribution. In this paper, we propose the use of a transformer deep learning method that analyzes mel-spectrograms for synthetic speech attribution. Our method known as Patchout Spectrogram Attribution Transformer (PSAT) can distinguish new, unseen speech generation methods from those seen during training. PSAT demonstrates high performance in attributing synthetic speech signals. Evaluation on the DARPA SemaFor Audio Attribution Dataset and the ASVSpoof2019 Dataset shows that our method achieves more than 95% accuracy in synthetic speech attribution and performs better than existing deep learning approaches.
Kratika Bhagtani, Emily R. Bartusiak, Amit Kumar Singh Yadav, Paolo Bestagini, Edward J. Delp
IH&MMSec4
2023 Extracting Efficient Spectrograms From MP3 Compressed Speech Signals for Synthetic Speech Detection
abstract
Many speech signals are compressed with MP3 to reduce the data rate. In many synthetic speech detection methods the spectrogram of the speech signal is used. This usually requires the speech signal to be fully decompressed. We show that the design of MP3 compression allows one to approximate the spectrogram of the MP3 compressed speech efficiently without fully decoding the compressed speech. We denote the spectograms obtained using our proposed approach by Efficient Spectrograms (E-Specs). E-Spec can reduce the complexity of spectrogram computation by ~77.60 percentage points (p.p.) and save ~37.87 p.p. of MP3 decoding time. E-Spec bypasses the reconstruction artifacts introduced by the MP3 synthesis filterbank, which makes it useful in speech forensics tasks. We tested E-Spec in the synthetic speech detection, where a detector is asked to determine whether a speech signal is synthesized or recorded from a human. We examined 4 different neural network architectures to evaluate the performance of E-Spec compared to speech features extracted from the fully decoded speech signal. E-Spec achieved the best synthetic speech detection performance for 3 architectures; it also achieved the best overall detection performance across architectures. The computation of E-Spec is an approximation to Short Time Fourier Transform (STFT). E-Spec can be extended to other audio compression methods.
Ziyue Xiang, Amit Kumar Singh Yadav, Stefano Tubaro, Paolo Bestagini, Edward J. Delp
IH&MMSec4
2023 BiFPro: A Bidirectional Facial-data Protection Framework against DeepFake
abstract
The rapid progress of the DeepFake technique has caused severe privacy problems. Thus protecting facial data against DeepFake becomes an urgent requirement. Face protection can be regarded as a bidirectional process: Face-out-detection (FOD) and Face-in-forensics (FIF). For FOD, the detectability should be satisfied when using the protected face to replace other faces. For FIF, traceability should be guaranteed when the protected face is replaced by others. For this, we propose a Bidirectional Facial-data Protection Framework (BiFPro) to protect face data comprehensively. This framework is composed of three main parts: Watermarking embedding, Face-out-detection (FOD) and Face-in-forensics (FIF). For the FOD case, we ensure the vulnerability of the original face by embedding fragile watermarking. Once the protected facial image is used to replace other faces, the watermarking information will be corrupted in the synthesized face images which can be used to detect the authenticity of the protected facial images. As for the FIF case, we guarantee the traceability of the protected face image by embedding robust watermarking, with which the fake faces can be traced with the reserved watermarking even after the face is swapped. Experimental results demonstrate that our proposed BiFPro could generate the watermarking which is fragile to FOD and at the same time robust to FIF with an average watermark extraction success rate reaching more than 95% when defending against the four advanced DeepFake techniques. Finally, we hope this work can encourage more initiative countermeasures against DeepFake.
Honggu Liu, Wenbo Zhou 0004, Han Fang 0004, Paolo Bestagini, Weiming Zhang 0001, Yuefeng Chen, Stefano Tubaro, Nenghai Yu, Yuan He 0011, Hui Xue 0001
ACM Multimedia5
2023 Audio Splicing Detection and Localization Based on Acquisition Device Traces
abstract
In recent years, the multimedia forensic community has put a great effort in developing solutions to assess the integrity and authenticity of multimedia objects, focusing especially on manipulations applied by means of advanced deep learning techniques. However, in addition to complex forgeries as the deepfakes, very simple yet effective manipulation techniques not involving any use of state-of-the-art editing tools still exist and prove dangerous. This is the case of audio splicing for speech signals, i.e., to concatenate and combine multiple speech segments obtained from different recordings of a person in order to cast a new fake speech. Indeed, by simply adding a few words to an existing speech we can completely alter its meaning. In this work, we address the overlooked problem of detection and localization of audio splicing from different models of acquisition devices. Our goal is to determine whether an audio track under analysis is pristine, or it has been manipulated by splicing one or multiple segments obtained from different device models. Moreover, if a recording is detected as spliced, we identify where the modification has been introduced in the temporal dimension. The proposed method is based on a Convolutional Neural Network (CNN) that extracts model-specific features from the audio recording. After extracting the features, we determine whether there has been a manipulation through a clustering algorithm. Finally, we identify the point where the modification has been introduced through a distance-measuring technique. The proposed method allows to detect and localize multiple splicing points within a recording.
Daniele Ugo Leonzio, Luca Cuccovillo, Paolo Bestagini, Marco Marcon, Patrick Aichroth, Stefano Tubaro
IEEE Trans. Inf. Forensics Secur.3
2022 Panchromatic Imagery Copy-Paste Localization Through Data-Driven Sensor Attribution
abstract
Overhead images can be obtained using different acquisition and processing techniques, and they are becoming more and more popular. As with common photographs, they can be forged and manipulated by malicious users. However, not all image forensics methods tailored to normal photos can be successfully applied out of the box to overhead images. In this paper we consider the problem of localizing copy-paste forgeries on panchromatic images acquired with different satellites. We leverage a set of Convolutional Neural Networks (CNNs) that extract traces of the acquisition satellite directly from image patches. We then determine whether an image region appears to have been acquired with a different satellite than the rest of the picture. Results show that the proposed technique outperforms more sophisticated image forensics tools tailoring common photographs.
Edoardo Daniele Cannas, János Horváth, Sriram Baireddy, Paolo Bestagini, Edward J. Delp, Stefano Tubaro
ICASSP4
2022 Deepfake Speech Detection Through Emotion Recognition: A Semantic Approach
abstract
In recent years, audio and video deepfake technology has advanced relentlessly, severely impacting people’s reputation and reliability. Several factors have facilitated the growing deepfake threat. On the one hand, the hyper-connected society of social and mass media enables the spread of multimedia content worldwide in real-time, facilitating the dissemination of counterfeit material. On the other hand, neural network-based techniques have made deepfakes easier to produce and difficult to detect, showing that the analysis of low-level features is no longer sufficient for the task. This situation makes it crucial to design systems that allow detecting deepfakes at both video and audio levels. In this paper, we propose a new audio spoofing detection system leveraging emotional features. The rationale behind the proposed method is that audio deepfake techniques cannot correctly synthesize natural emotional behavior. Therefore, we feed our deepfake detector with high-level features obtained from a state-of-the-art Speech Emotion Recognition (SER) system. As the used descriptors capture semantic audio information, the proposed system proves robust in cross-dataset scenarios outperforming the considered baseline on multiple datasets.
Emanuele Conti, Davide Salvi, Clara Borrelli, Brian C. Hosler, Paolo Bestagini, Fabio Antonacci, Augusto Sarti, Matthew C. Stamm, Stefano Tubaro
ICASSP5
2022 A Data-Driven Approach for Acoustic Parameter Similarity Estimation of Speech Recording
abstract
Speech audio acquisitions exhibit different quality and reverberation properties depending on the recording setup and environment. For this reason, it is expected that speech analysis systems that work correctly on certain audio recordings may fail on others acquired in different acoustic contexts. Therefore, to be able to tell whether a track under analysis shares the same acoustic characteristics of a reference one may be useful to understand if it can be successfully processed by a given speech analysis system. Alternatively, in a forensic scenario, an estimate of acoustic parameter similarity between two tracks can be used to verify whether the recordings have been likely acquired in the same environment or not. In this work, we propose two methods to estimate acoustic parameter similarity between a speech recording under analysis and a reference one. The first method relies on the estimation of channel-based acoustic indicators that are then compared to extract a similarity measure. The second method directly learns a parameter similarity measure through siamese neural networks.
Mattia Papa, Clara Borrelli, Paolo Bestagini, Fabio Antonacci, Augusto Sarti, Stefano Tubaro
ICASSP3
2022 Forensic Analysis and Localization of Multiply Compressed MP3 Audio Using Transformers
abstract
Audio signals are often stored and transmitted in compressed formats. Among the many available audio compression schemes, MPEG-1 Audio Layer III (MP3) is very popular and widely used. Since MP3 is lossy it leaves characteristic traces in the compressed audio which can be used forensically to expose the past history of an audio file. In this paper, we consider the scenario of audio signal manipulation done by temporal splicing of compressed and uncompressed audio signals. We propose a method to find the temporal location of the splices based on transformer networks. Our method identifies which temporal portions of a audio signal have undergone single or multiple compression at the temporal frame level, which is the smallest temporal unit of MP3 compression. We tested our method on a dataset of 486,743 MP3 audio clips. Our method achieved higher performance and demonstrated robustness with respect to different MP3 data when compared with existing methods.
Ziyue Xiang, Paolo Bestagini, Stefano Tubaro, Edward J. Delp
ICASSP2
2022 Detecting Gan-Generated Images by Orthogonal Training of Multiple CNNs
abstract
In the last few years, we have witnessed the rise of a series of deep learning methods to generate synthetic images that look extremely realistic. These techniques prove useful in the movie industry and for artistic purposes. However, they also prove dangerous if used to spread fake news or to generate fake online accounts. For this reason, detecting if an image is an actual photograph or has been synthetically generated is becoming an urgent necessity. This paper proposes a detector of synthetic images based on an ensemble of Convolutional Neural Networks (CNNs). We consider the problem of detecting images generated with techniques not available at training time. This is a common scenario, given that new image generators are published more and more frequently. To solve this issue, we leverage two main ideas: (i) CNNs should provide "orthogonal" results to better contribute to the ensemble; (ii) the original-image class is better defined than the synthetic-image one, thus it should be better trusted at testing time. Experiments show that pursuing these two ideas improves the detector accuracy on NVIDIA's newly generated StyleGAN3 images, never used in training.
Sara Mandelli, Nicolò Bonettini, Paolo Bestagini, Stefano Tubaro
ICIP3
2022 DIPPAS: a deep image prior PRNU anonymization scheme
abstract
Abstract Source device identification is an important topic in image forensics since it allows to trace back the origin of an image. Its forensics counterpart is source device anonymization, that is, to mask any trace on the image that can be useful for identifying the source device. A typical trace exploited for source device identification is the photo response non-uniformity (PRNU), a noise pattern left by the device on the acquired images. In this paper, we devise a methodology for suppressing such a trace from natural images without a significant impact on image quality. Expressly, we turn PRNU anonymization into the combination of a global optimization problem in a deep image prior (DIP) framework followed by local post-processing operations. In a nutshell, a convolutional neural network (CNN) acts as a generator and iteratively returns several images with attenuated PRNU traces. By exploiting straightforward local post-processing and assembly on these images, we produce a final image that is anonymized with respect to the source PRNU, still maintaining high visual quality. With respect to widely adopted deep learning paradigms, the used CNN is not trained on a set of input-target pairs of images. Instead, it is optimized to reconstruct output images from the original image under analysis itself. This makes the approach particularly suitable in scenarios where large heterogeneous databases are analyzed. Moreover, it prevents any problem due to the lack of generalization. Through numerical examples on publicly available datasets, we prove our methodology to be effective compared to state-of-the-art techniques.
Francesco Picetti, Sara Mandelli, Paolo Bestagini, Vincenzo Lipari, Stefano Tubaro
EURASIP J. Inf. Secur.3
2022 Deep Prior-Based Unsupervised Reconstruction of Irregularly Sampled Seismic Data
abstract
Irregularity and coarse spatial sampling of seismic data strongly affect the performances of processing and imaging algorithms. Therefore, interpolation is a usual preprocessing step in most of the processing workflows. In this work, we propose a seismic data interpolation method based on the deep prior paradigm: anad hocconvolutional neural network is used as a prior to solve the interpolation inverse problem, avoiding any costly and prone-to-overfitting training stage. In particular, the proposed method leverages a multiresolution U-Net with 3-D convolution kernels exploiting correlations in cubes of seismic data, at different scales in all directions. Numerical examples on different corrupted synthetic and field data sets show the effectiveness and promising features of the proposed approach.
Fantong Kong, Francesco Picetti, Vincenzo Lipari, Paolo Bestagini, Xiaoming Tang, Stefano Tubaro
IEEE Geosci. Remote. Sens. Lett.4
2022 Intelligent Seismic Deblending Through Deep Preconditioner
abstract
Seismic deblending is an ill-posed inverse problem that involves counteracting the effect of a blending matrix derived from the shots position and firing time. In this letter, we propose a seismic deblending method based on so-called deep preconditioners. A convolutional Autoencoder (AE) is first trained in a patch-wise fashion to learn an effective sparse representation of the common receiver gathers (CRGs) we aim to reconstruct. Then, the decoder branch of the trained AE is used as a nonlinear preconditioner for the deblending problem. Particularly, to avoid the explicit creation of a training dataset, we suggest to use the common shot gathers (CSGs) of the blended dataset itself to train the AE network, as they are not affected by incoherent blending noise. Numerical examples on synthetic and field datasets demonstrate the effectiveness of the proposed method in comparison with significantly comparable techniques: a dictionary-learning based deblending method; an end-to-end deblending convolution neutral network (CNN).
Weiwei Xu 0004, Vincenzo Lipari, Paolo Bestagini, Matteo Ravasi, Stefano Tubaro
IEEE Geosci. Remote. Sens. Lett.3
2021 Open-Set Source Attribution for Panchromatic Satellite Imagery
abstract
In the last few years, several companies started offering the possibility of buying different kinds of overhead images acquired by satellites orbiting around the planet. This market is interesting for several customers, from those who simply fancy a shot of their house from space, to those aiming to acquire strategic information on portions of land. Due to the sensitive nature of this data, which can be maliciously altered by anyone, the forensic community has started investigating methodologies to verify overhead imagery authenticity and integrity. Within this context, in this paper we investigate the possibility of using Convolutional Neural Networks (CNNs) to attribute a panchromatic satellite image to the satellite used to acquire it. In our investigation we tackle both closed-set and, adapting Deep Ensemble (DE) and Monte Carlo Dropout (MCD) techniques, open-set image attribution problems.
Edoardo Daniele Cannas, Sriram Baireddy, Emily R. Bartusiak, Sri Yarlagadda, Daniel Mas Montserrat, Paolo Bestagini, Stefano Tubaro, Edward J. Delp
ICIP6
2021 Anti-Aliasing Add-On For Deep Prior Seismic Data Interpolation
abstract
Data interpolation is a fundamental step in any seismic processing workflow. Among machine learning techniques recently proposed to solve data interpolation as an inverse problem, Deep Prior paradigm aims at employing a convolutional neural network to capture priors on the data in order to regularize the inversion. However, this technique lacks of reconstruction precision when interpolating highly decimated data due to the presence of aliasing. In this work, we propose to improve Deep Prior inversion by adding a directional Laplacian as regularization term to the problem. This regularizer drives the optimization towards solutions that honor the slopes estimated from the interpolated data low frequencies. We provide some numerical examples to showcase the methodology devised in this manuscript, showing that our results are less prone to aliasing also in presence of noisy and corrupted data.
Francesco Picetti, Vincenzo Lipari, Paolo Bestagini, Stefano Tubaro
ICIP3
2021 Synthetic speech detection through short-term and long-term prediction traces
abstract
Abstract Several methods for synthetic audio speech generation have been developed in the literature through the years. With the great technological advances brought by deep learning, many novel synthetic speech techniques achieving incredible realistic results have been recently proposed. As these methods generate convincing fake human voices, they can be used in a malicious way to negatively impact on today’s society (e.g., people impersonation, fake news spreading, opinion formation). For this reason, the ability of detecting whether a speech recording is synthetic or pristine is becoming an urgent necessity. In this work, we develop a synthetic speech detector. This takes as input an audio recording, extracts a series of hand-crafted features motivated by the speech-processing literature, and classify them in either closed-set or open-set. The proposed detector is validated on a publicly available dataset consisting of 17 synthetic speech generation algorithms ranging from old fashioned vocoders to modern deep learning solutions. Results show that the proposed method outperforms recently proposed detectors in the forensics literature.
Clara Borrelli, Paolo Bestagini, Fabio Antonacci, Augusto Sarti, Stefano Tubaro
EURASIP J. Inf. Secur.2
2021 Reconstructing Speech From CNN Embeddings
abstract
The complete understanding of the decision-making process of Convolutional Neural Networks (CNNs) is far from being fully reached. Many researchers proposed techniques to interpret what a network actually “learns” from data. Nevertheless many questions still remain unanswered. In this work we study one aspect of this problem by reconstructing speech from the intermediate embeddings computed by a CNNs. Specifically, we consider a pre-trained network that acts as a feature extractor from speech audio. We investigate the possibility of inverting these features, reconstructing the input signals in a black-box scenario, and quantitatively measure the reconstruction quality by measuring the word-error-rate of an off-the-shelf ASR model. Experiments performed using two different CNN architectures trained for six different classification tasks, show that it is possible to reconstruct time-domain speech signals that preserve the semantic content, whenever the embeddings are extracted before the fully connected layers.
Luca Comanducci, Paolo Bestagini, Marco Tagliasacchi, Augusto Sarti, Stefano Tubaro
IEEE Signal Process. Lett.2
2021 Landmine Detection Using Autoencoders on Multipolarization GPR Volumetric Data
abstract
Buried landmines and unexploded remnants of war are a constant threat for the population of many countries that have been hit by wars in the past years. The huge amount of casualties has been a strong motivation for the research community toward the development of safe and robust techniques designed for landmine clearance. Nonetheless, being able to detect and localize buried landmines with high precision in an automatic fashion is still considered a challenging task due to the many different boundary conditions that characterize this problem (e.g., several kinds of objects to detect, different soils and meteorological conditions, etc.). In this article, we propose a novel technique for buried object detection tailored to unexploded landmine discovery. The proposed solution exploits a specific kind of convolutional neural network (CNN) known as autoencoder to analyze volumetric data acquired with ground penetrating radar (GPR) using different polarizations. This method works in an anomaly detection framework, indeed we only train the autoencoder on GPR data acquired on landmine-free areas. The system then recognizes landmines as objects that are dissimilar to the soil used during the training step. Experiments conducted on real data show that the proposed technique requires little training and no ad hoc data preprocessing to achieve accuracy higher than 93% on challenging data sets.
Paolo Bestagini, Federico Lombardi, Maurizio Lualdi, Francesco Picetti, Stefano Tubaro
IEEE Trans. Geosci. Remote. Sens.1
2020 Phylogenetic Minimum Spanning Tree Reconstruction Using Autoencoders
abstract
The history of a shared and re-posted multimedia content can be reconstructed by analyzing the mutual relations between all of its near-duplicate copies and solving a minimum spanning tree (MST) problem, as shown by multimedia phylogeny research field, Unfortunately, MST estimation strategies are severely impaired by the noise affecting dissimilarity measures between pairs of near-duplicate contents, For this reason, researchers have recently been investigating robust dissimilarity metrics.This paper proposes a matrix denoising solution that both mitigates dissimilarity noise and reconstruct the desired phylogenetic tree at the same time, The proposed strategy is a first attempt to estimate a MST via a denoising autoencoder that returns an approximation of the adjacency matrix corresponding to the underlying tree, Experimental results prove that the proposed solution outperforms the previous approaches and easily adapts to different analysis scenarios.
Riccardo Castelletto, Simone Milani, Paolo Bestagini
ICASSP3
2020 Multimodal Violence Detection in Videos
abstract
Effective tools for detection of violence are highly demanded, specially when dealing with video streams. Such tools have a wide range of applications, from forensics and law enforcement to parental control over the ever increasing amount of videos available online. Prior studies showed that deep learning has great potential in detecting violence, but focuses on detecting violence in general, or only specific cases of violent behavior. While the concept of violence is broad and highly subjective, simpler concepts such as fights, explosions, and gunshots, convey the idea of violence while being more objective. Even though different concepts relate to this same broader idea of violence, they differ widely in relation to whether or not they convey the idea of movement, the presence of a specific object, or even if they generate distinctive sounds. In this study, we propose to analyze different concepts related to violence and how to better describe these concepts exploring visual and auditory cues in order to reach a robust method to detect violence.
Bruno Peixoto, Bahram Lavi, Paolo Bestagini, Zanoni Dias, Anderson Rocha 0001
ICASSP3
2020 A Modified Fourier-Mellin Approach For Source Device Identification On Stabilized Videos
abstract
To decide whether a digital video has been captured by a given device, multimedia forensic tools usually exploit characteristic noise traces left by the camera sensor on the acquired frames. This analysis requires that the noise pattern characterizing the camera and the noise pattern extracted from video frames under analysis are geometrically aligned. However, in many practical scenarios this does not occur, thus a re-alignment or synchronization has to be performed. Current solutions often require time consuming search of the realignment transformation parameters. In this paper, we propose to overcome this limitation by searching scaling and rotation parameters in the frequency domain. The proposed algorithm tested on real videos from a well-known state-of-the-art dataset shows promising results.
Sara Mandelli, Fabrizio Argenti, Paolo Bestagini, Massimo Iuliani, Alessandro Piva, Stefano Tubaro
ICIP3
2020 On the use of Benford's law to detect GAN-generated images
abstract
The advent of Generative Adversarial Network (GAN) architectures has given anyone the ability of generating incredibly realistic synthetic imagery. The malicious diffusion of GAN-generated images may lead to serious social and political consequences (e.g., fake news spreading, opinion formation, etc.). It is therefore important to regulate the widespread distribution of synthetic imagery by developing solutions able to detect them. In this paper, we study the possibility of using Benford's law to discriminate GAN-generated images from natural photographs. Benford's law describes the distribution of the most significant digit for quantized Discrete Cosine Transform (DCT) coefficients. Extending and generalizing this property, we show that it is possible to extract a compact feature vector from an image. This feature vector can be fed to an extremely simple classifier for GAN-generated image detection purpose.
Nicolò Bonettini, Paolo Bestagini, Simone Milani, Stefano Tubaro
ICPR2
2020 Video Face Manipulation Detection Through Ensemble of CNNs
abstract
In the last few years, several techniques for facial manipulation in videos have been successfully developed and made available to the masses (i.e., FaceSwap, deepfake, etc.). These methods enable anyone to easily edit faces in video sequences with incredibly realistic results and a very little effort. Despite the usefulness of these tools in many fields, if used maliciously, they can have a significantly bad impact on society (e.g., fake news spreading, cyber bullying through fake revenge porn). The ability of objectively detecting whether a face has been manipulated in a video sequence is then a task of utmost importance. In this paper, we tackle the problem of face manipulation detection in video sequences targeting modern facial manipulation techniques. In particular, we study the ensembling of different trained Convolutional Neural Network (CNN) models. In the proposed solution, different models are obtained starting from a base network (i.e., EfficientNetB4) making use of two different concepts: (i) attention layers; (ii) siamese training. We show that combining these networks leads to promising face manipulation detection results on two publicly available datasets with more than 119000 videos.
Nicolò Bonettini, Edoardo Daniele Cannas, Sara Mandelli, Luca Bondi, Paolo Bestagini, Stefano Tubaro
ICPR5
2020 CNN-Based Fast Source Device Identification
abstract
Source identification is an important topic in image forensics, since it allows to trace back the origin of an image. This represents a precious information to claim intellectual property but also to reveal the authors of illicit materials. In this letter we address the problem of device identification based on sensor noise and propose a fast and accurate solution using convolutional neural networks (CNNs). Specifically, we propose a 2-channel-based CNN that learns a way of comparing camera fingerprint and image noise at patch level. The proposed solution turns out to be much faster than the conventional approach and to ensure an increased accuracy. This makes the approach particularly suitable in scenarios where large databases of images are analyzed, like over social networks. In this vein, since images uploaded on social media usually undergo at least two compression stages, we include investigations on double JPEG compressed images, always reporting higher accuracy than standard approaches.
Sara Mandelli, Davide Cozzolino, Paolo Bestagini, Luisa Verdoliva, Stefano Tubaro
IEEE Signal Process. Lett.3
2020 Source Localization Using Distributed Microphones in Reverberant Environments Based on Deep Learning and Ray Space Transform
abstract
In this article we present a methodology for source localization in reverberant environments from Generalized Cross Correlations (GCCs) computed between spatially distributed individual microphones. Reverberation tends to negatively affect localization based on Time Differences of Arrival (TDOAs), which become inaccurate due to the presence of spurious peaks in the GCC. We therefore adopt a data-driven approach based on a convolutional neural network, which, using the GCCs as input, estimates the source location in two steps. It first computes the Ray Space Transform (RST) from multiple arrays. The RST is a convenient representation of the acoustic rays impinging on the array in a parametric space, called Ray Space. Rays produced by a source are visualized in the RST as patterns, whose position is uniquely related to the source location. The second step consists of estimating the source location through a nonlinear fitting, which estimates the coordinates that best approximate the RST pattern obtained through the first step. It is worth noting that training can be accomplished on simulated data only, thus relaxing the need of actually deploying microphone arrays in the acoustic scene. The localization accuracy of the proposed techniques is similar to the one of SRP-PHAT, however our method demonstrates an increased robustness regarding different distributed microphones configurations. Moreover, the use of the RST as an intermediate representation makes it possible for the network to generalize to data unseen during training.
Luca Comanducci, Federico Borra, Paolo Bestagini, Fabio Antonacci, Stefano Tubaro, Augusto Sarti
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Facing Device Attribution Problem for Stabilized Video Sequences
abstract
A problem deeply investigated by multimedia forensics researchers is that of detecting which device has been used to capture a video. This enables us to trace down the owner of a video sequence, which proves extremely helpful to solve copyright infringement cases as well as to fight distribution of illicit material (e.g., child exploitation clips and terroristic threats). Currently, the most promising methods to tackle this task exploit unique noise traces left by camera sensors on acquired images. However, given the recent advancements in motion stabilization of video content, robustness of sensor pattern noise-based techniques is strongly hindered. Indeed, video stabilization introduces geometric transformations to video frames, thus making camera fingerprint estimation problematic with classical approaches. In this paper, we deal with the challenging problem of attributing stabilized videos to their recording device. Specifically, we propose: 1) a strategy to extract the characteristic fingerprint of a device, starting from either a set of images or stabilized video sequences and 2) a strategy to match a stabilized video sequence with a given fingerprint. The proposed methodology is tested on videos coming from a set of different smartphones, taken from the modern publicly available Vision Dataset. The conducted experiments also provide an interesting insight on the effect of modern smartphones video stabilization algorithms on specific video frames.
Sara Mandelli, Paolo Bestagini, Luisa Verdoliva, Stefano Tubaro
IEEE Trans. Inf. Forensics Secur.2
2019 "Hello? Who Am I Talking to?" A Shallow CNN Approach for Human vs. Bot Speech Classification
abstract
Automatic speech generation algorithms, enhanced by deep learning techniques, enable an increasingly seamless and immediate machine-to-human interaction. As a result, the latest generation of phone-calling bots sounds more convincingly human than previous generations. The application of this technology has a strong social impact in terms of privacy issues (e.g., in customer-care services), fraudulent actions (e.g., social hacking) and erosion of trust (e.g., generation of fake conversation). For these reasons, it is crucial to identify the nature of a speaker, as either a human or a bot. In this paper, we propose a speech classification algorithm based on Convolutional Neural Networks (CNNs), which enables the automatic classification of human vs non-human speakers from the analysis of short audio excerpts. We evaluate the effectiveness of the proposed solution by exploiting a real human speech database populated with audio recordings from various sources, and automatically generated speeches using state-of-the-art text-to-speech generators based on deep learning (e.g., Google WaveNet).
Alessandro Lieto, Daniele Moro, Francesco Devoti, Claudia Parera, Vincenzo Lipari, Paolo Bestagini, Stefano Tubaro
ICASSP6
2019 Shadow Removal Detection and Localization for Forensics Analysis
abstract
The recent advancements in image processing and computer vision allow realistic photo manipulations. In order to avoid the distribution of fake imagery, the image forensics community is working towards the development of image authenticity verification tools. Methods based on shadow analysis are particularly reliable since they are part of the physical integrity of the scene, thus detecting forgeries is possible whenever inconsistencies are found (e.g., shadows not coherent with the light direction). An attacker can easily delete inconsistent shadows and replace them with correctly cast shadows in order to fool forensics detectors based on physical analysis. In this paper, we propose a method to detect shadow removal done with state-of-the-art tools. The proposed method is based on a conditional generative adversarial network (cGAN) specifically trained for shadow removal detection.
Sri Yarlagadda, David Guera, Daniel Mas Montserrat, Fengqing Zhu 0001, Edward J. Delp, Paolo Bestagini, Stefano Tubaro
ICASSP6
2019 Image Anonymization Detection with Deep Handcrafted Features
abstract
In recent years, the number of images shared online has continuously grown. The forensics community has kept the pace by developing techniques to both reliably extract information from these images, but also to remove it. In particular, the latest developments in image anonymization methods exposes an attack vector when used by skilled ill-intentioned image producers that may want to elude prosecution. We present an approach to detect whether or not an image has undergone a laundering process, i.e., it has been tampered with so that its unique characterizing features have been changed to avoid detection. We focus on the photo response non uniformity (PRNU) noise unique to every imaging sensor, and we consider that an image has been "laundered" when we detect the absence of PRNU from an image. We propose a per image preprocessing pipeline that generates information-rich features later used as input of fine-tuned convolutional neural networks (CNNs). We study the performance of the proposed approach using various CNN architectures and blind anonymization techniques and show its effectiveness under several training and testing scenarios. Our results also show that CNN models trained with the proposed feature are capable of generalizing over unseen devices and are robust against non-geometric transformations.
Nicolò Bonettini, David Guera, Luca Bondi, Paolo Bestagini, Edward J. Delp, Stefano Tubaro
ICIP4
2019 A Prnu-Based Method to Expose Video Device Compositions in Open-Set Setups
abstract
As the diffusion of altered video sequences over social networks and the internet can have severe consequences (e.g., fake news spreading, false accusations, etc.), the forensic research community has started actively working toward the development of methodologies tailored to assess the authenticity and integrity of video sequences. In this work, we focus on the problem of spotting video compilations composed by temporally concatenating sequences acquired with different devices. To solve this problem, we leverage trace characteristics of each video recording device left on each video at acquisition time. Specifically, we propose a method leveraging Photo-Response Non-Uniformity (PRNU) traces along with a binary classifier in order to understand which frames of a video have been acquired with the same device used to record the first few frames. Results show that the proposed solution outperforms baselines based on standard PRNU correlation and thresholding tests. Experiments have been carried out in an open-set scenario and show promising results.
Pedro Ribeiro Mendes Júnior, Luca Bondi, Paolo Bestagini, Anderson Rocha 0001, Stefano Tubaro
ICIP3
2019 Detection and Synchronization of Video Sequences for Event Reconstruction
abstract
With an ever-growing amount of unexpected menaces in crowded places such as terrorist attacks, it is paramount to develop techniques to aid investigators reconstructing all details about an event of interest. To extract reliable information about the event, all kinds of available clues must be jointly exploited. As a matter of fact, today's sources of information are plenty and varied, as important events affecting many people are typically documented by different sources. Both witnesses' smartphones and security cameras can provide valuable information coming from multiple viewpoints and time instants - "the eyes of the crowd". In this paper, we focus on the specific problem of automatically detecting and temporally synchronizing videos depicting the same event of interest. Videos can be either near-duplicates (i.e., edited copies of the same original source) or sequences shot by different users from different vantage points. The proposed method relies upon a video fingerprinting technique capable of describing how video semantic content evolves in time. The solution does not assume a priori information about cameras location, and it only exploits visual cues, not relying on audio channels.
Giuliano Pinheiro, Marcos V. M. Cirne, Paolo Bestagini, Stefano Tubaro, Anderson Rocha 0001
ICIP3
2019 Multimedia Forensics
abstract
With the availability of powerful and easy-to-use media editing tools, falsifying images and videos has become widespread in the last few years. Coupled with ubiquitous social networks, this allows for the viral dissemination of fake news. This raises huge concerns on multimedia security. This scenario became even worse with the advent of deep learning. New, sophisticated methods have been proposed to accomplish manipulations that were previously unthinkable (e.g., deepfake). This tutorial will present the most reliable methods for detection of manipulated images and for source identification. These are important tools nowadays to carry out fact checking and authorship verification. Hence, this is a timely and relevant research topic in the multimedia security research community.
Luisa Verdoliva, Paolo Bestagini
ACM Multimedia2
2019 Improving PRNU Compression Through Preprocessing, Quantization, and Coding
abstract
In the last decade, the extremely rapid proliferation of digital devices capable of acquiring and sharing images over the Web has significantly increased the amount of digital images publicly accessible by everyone with Internet access. Despite the obvious benefits of such technological improvements, it is becoming mandatory to verify the origin and trustfulness of such shared pictures. Photo response non-uniformity (PRNU) is the reference signal for forensic investigators when it comes to verifying or identifying which camera device shot a picture under analysis. In spite of this, PRNU is almost a white-shaped noise, thus being very difficult to compress for storage or large scale search purposes, which are frequent investigation scenarios. To overcome the issue, the forensic community has developed a series of compression algorithms. Lately, Gaussian random projections have proved to achieve state-of-the-art performance. In this paper, we propose two additional steps that help improving even more Gaussian random projections compression rate: 1) a decimation preprocessing step tailored at attenuating frequency components in which PRNU traces are already suppressed in JPEG compressed images and 2) a dead-zone quantizer (rather than the commonly used binary one) that enables an entropy coding scheme to save bitrate when storing PRNU fingerprints or sending residuals over a communication channel. Reported results show the effectiveness of proposed improvements, both under controlled JPEG compression and in a real case scenario.
Luca Bondi, Paolo Bestagini, Fernando Pérez-González, Stefano Tubaro
IEEE Trans. Inf. Forensics Secur.2
2018 Multiple Jpeg Compression Detection Through Task-Driven Non-Negative Matrix Factorization
abstract
Due to the increasingly unbridled practice of sharing visual content on the web, tracing back past history of uploaded images is getting far from being an easy task. Nonetheless, forensic analysts might be interested in probing digital history of content published on the web to assess its authenticity. In this vein, a possible indicator of image integrity is the number of JPEG compressions a picture underwent. As a matter of fact, JPEG compression is typically operated first at image inception time directly on the acquisition device. Then, it is customary re-applied every time an image is manipulated or shared through social media. For this reason, the more the applied JPEG compressions, the more the likelihood that an image underwent some editing. In this work, we propose an algorithm to detect multiple JPEG compressions, specifically up to four coding cycles. This approach leverages the Task-driven Non-negative Matrix Factorization (TNMF) model, fed with histograms of the Discrete Cosine Transform (DCT) of the image under analysis. Experimental results show the effectiveness of the method if compared with the state-of-the-art, confirming this strategy as a viable solution for detecting multiple JPEG compressions.
Sara Mandelli, Nicolò Bonettini, Paolo Bestagini, Vincenzo Lipari, Stefano Tubaro
ICASSP3
2018 Video Codec Forensics Based on Convolutional Neural Networks
abstract
The recent development of multimedia has made video editing accessible to everyone. Unfortunately, forensic analysis tools capable of detecting traces left by video processing operations in a blind fashion are still at their beginnings. One of the reasons is that videos are customary stored and distributed in a compressed format, and codec-related traces tends to mask previous processing operations. In this paper, we propose to capture video codec traces through convolutional neural networks (CNNs) and exploit them as an asset. Specifically, we train two CNN s to extract information about the used video codec and coding quality, respectively. Building upon these CNN s, we propose a system to detect and localize temporal splicing for video sequences generated from the concatenation of different video segments, which are characterized by inconsistent coding schemes and/or parameters (e.g., video compilations from different sources or broadcasting channels). The proposed solution is validated using videos at different resolutions (i.e., CIF, 4CIF, PAL and 720p) encoded with four common codecs (i.e., MPEG2, MPEG4, H264 and H265) at different qualities (i.e., different constant and variable bitrates, as well as constant quantization parameters).
Sebastiano Verde, Luca Bondi, Paolo Bestagini, Simone Milani, Giancarlo Calvagno, Stefano Tubaro
ICIP3
2018 Reliability Map Estimation for CNN-Based Camera Model Attribution
abstract
Among the image forensic issues investigated in the last few years, great attention has been devoted to blind camera model attribution. This refers to the problem of detecting which camera model has been used to acquire an image by only exploiting pixel information. Solving this problem has great impact on image integrity assessment as well as on authenticity verification. Recent advancements that use convolutional neural networks (CNNs) in the media forensic field have enabled camera model attribution methods to work well even on small image patches. These improvements are also important for determining forgery localization. Some patches of an image may not contain enough information related to the camera model (e.g., saturated patches). In this paper, we propose a CNN-based solution to estimate the camera model attribution reliability of a given image patch. We show that we can estimate a reliabilitymap indicating which portions of the image contain reliable camera traces. Testing using a well known dataset confirms that by using this information, it is possible to increase small patch camera model attribution accuracy by more than 8% on a single patch.
David Guera, Fengqing Zhu 0001, Sri Yarlagadda, Stefano Tubaro, Paolo Bestagini, Edward J. Delp
WACV5
2017 Spotting the difference: Context retrieval and analysis for improved forgery detection and localization
abstract
As image tampering becomes ever more sophisticated and commonplace, the need for image forensics algorithms that can accurately and quickly detect forgeries grows. In this paper, we revisit the ideas of image querying and retrieval to provide clues to better localize forgeries. We propose a method to perform large-scale image forensics on the order of one million images using the help of an image search algorithm and database to gather contextual clues as to where tampering may have taken place. In this vein, we introduce five new strongly invariant image comparison methods and test their effectiveness under heavy noise, rotation, and color space changes. Lastly, we show the effectiveness of these methods compared to passive image forensics using Nimble [1], a new, state-of-the-art dataset from the National Institute of Standards and Technology (NIST).
Joel Brogan, Paolo Bestagini, Aparna Bharati, Allan Pinto, Daniel Moreira, Kevin W. Bowyer, Patrick J. Flynn, Anderson Rocha 0001, Walter J. Scheirer
ICIP2
2017 Inpainting-Based camera anonymization
abstract
Over the years, the forensic community has developed a series of very accurate camera attribution algorithms enabling to detect which device has been used to acquire an image with outstanding results. Many of these methods are based on photo response non uniformity (PRNU) that allows tracing back a picture to the camera used to shoot it. However, when privacy is required, it would be desirable to anonymize photos, unlinking them from their specific device. This paper investigates a new and alternative approach to image anonymization task. The proposed method leverages image inpainting described as an inverse regularized problem, and does not need any priors about the PRNU to remove. Specifically, we show how PRNU pattern can be strongly attenuated by reconstructing each pixel of an image from its neighbors, only slightly affecting visual quality. Results confirm this approach as a viable alternative solution for image anonymization.
Sara Mandelli, Luca Bondi, Silvia Lameri, Vincenzo Lipari, Paolo Bestagini, Stefano Tubaro
ICIP5
2017 Aligned and non-aligned double JPEG detection using convolutional neural networks
Mauro Barni, Luca Bondi, Nicolò Bonettini, Paolo Bestagini, Andrea Costanzo, Marco Maggini, Benedetta Tondi, Stefano Tubaro
J. Vis. Commun. Image Represent.4
2017 First Steps Toward Camera Model Identification With Convolutional Neural Networks
abstract
Detecting the camera model used to shoot a picture enables to solve a wide series of forensic problems, from copyright infringement to ownership attribution. For this reason, the forensic community has developed a set of camera model identification algorithms that exploit characteristic traces left on acquired images by the processing pipelines specific of each camera model. In this letter, we investigate a novel approach to solve camera model identification problem. Specifically, we propose a data-driven algorithm based on convolutional neural networks, which learns features characterizing each camera model directly from the acquired pictures. Results on a well-known dataset of 18 camera models show that: 1) the proposed method outperforms up-to-date state-of-the-art algorithms on classification of 64 × 64 color image patches; 2) features learned by the proposed network generalize to camera models never used for training.
Luca Bondi, Luca Baroffio, David Guera, Paolo Bestagini, Edward J. Delp, Stefano Tubaro
IEEE Signal Process. Lett.4
2017 Data-Driven Feature Characterization Techniques for Laser Printer Attribution
abstract
Laser printer attribution is an increasing problem with several applications, such as pointing out the ownership of crime proofs and authentication of printed documents. However, as commonly proposed methods for this task are based on custom-tailored features, they are limited by modeling assumptions about printing artifacts. In this paper, we explore solutions able to learn discriminant-printing patterns directly from the available data during an investigation, without any further feature engineering, proposing the first approach based on deep learning to laser printer attribution. This allows us to avoid any prior assumption about printing artifacts that characterize each printer, thus highlighting almost invisible and difficult printer footprints generated during the printing process. The proposed approach merges, in a synergistic fashion, convolutional neural networks (CNNs) applied on multiple representations of multiple data. Multiple representations, generated through different pre-processing operations, enable the use of the small and lightweight CNNs whilst the use of multiple data enable the use of aggregation procedures to better determine the provenance of a document. Experimental results show that the proposed method is robust to noisy data and outperforms existing counterparts in the literature for this problem.
Anselmo Ferreira, Luca Bondi, Luca Baroffio, Paolo Bestagini, Jiwu Huang, Jefersson A. dos Santos, Stefano Tubaro, Anderson Rocha 0001
IEEE Trans. Inf. Forensics Secur.4
2016 Image phylogeny tree reconstruction based on region selection
abstract
Nowadays, everyone can download, edit and republish any picture on the web, thus contributing to the diffusion of near-duplicate (ND) images. In order to gain an interesting insight on the way NDs are distributed online, recent works have focused on the reconstruction of the image phylogeny tree (IPT), i.e., an acyclic graph describing the genealogical relationship between ND image pairs. IPT reconstruction methods typically leverage the possibility of reconstructing one image from another one only if they are in parent-child relationship. However, as estimating the possible parent-child transformation is computationally expensive, usually a limited set of global editing operations is considered (i.e., compression, geometric and colour transformations applied to the whole image). However, in a real-world scenario it is customary to edit images also using local operations (e.g., logo insertion, object removal, splicing, etc.), which hinder the possibility of correctly estimating the parent-child relationship. In this paper, we propose an algorithm for IPT reconstruction that deals with the presence of local editing operations.
Paolo Bestagini, Marco Tagliasacchi, Stefano Tubaro
ICASSP1
2016 Phylogenetic analysis of near-duplicate images using processing age metrics
abstract
Recent researches on image forensics have led to the design of algorithms to study the phylogenetic relationship between near-duplicate (ND) images. The proposed solutions aim at reconstructing the image phylogeny tree (IPT), and they have immediate applications in security, law and copyright enforcement, and news tracking services. Anyway, the effectiveness of such strategies strictly depends on the accuracy in characterizing image similarities. In this paper, we show that it is possible to take into account additional information to better reconstruct the IPT. More specifically, we propose a set of features that blindly model the processing age of an image, i.e., how much an image has been edited in its lifetime. By exploiting these features, it is possible to improve the performance of IPT reconstruction by increasing the accuracy and reducing the computational complexity.
Simone Milani, Marco Fontana, Paolo Bestagini, Stefano Tubaro
ICASSP3
2016 Codec and GOP Identification in Double Compressed Videos
abstract
Video content is routinely acquired and distributed in a digital compressed format. In many cases, the same video content is encoded multiple times. This is the typical scenario that arises when a video, originally encoded directly by the acquisition device, is then re-encoded, either after an editing operation, or when uploaded to a sharing website. The analysis of the bitstream reveals details of the last compression step (i.e., the codec adopted and the corresponding encoding parameters), while masking the previous compression history. Therefore, in this paper, we consider a processing chain of two coding steps, and we propose a method that exploits coding-based footprints to identify both the codec and the size of the group of pictures (GOPs) used in the first coding step. This sort of analysis is useful in video forensics, when the analyst is interested in determining the characteristics of the originating source device, and in video quality assessment, since quality is determined by the whole compression history. The proposed method relies on the fact that lossy coding is an (almost) idempotent operation. That is, re-encoding a video sequence with the same codec and coding parameters produces a sequence that is similar to the former. As a consequence, if the second codec in the chain does not significantly alter the sequence, it is possible to analyze this sort of similarity to identify the first codec and the adopted GOP size. The method was extensively validated on a very large data set of video sequences generated by encoding content with a diversity of codecs (MPEG-2, MPEG-4, H.264/AVC, and DIRAC) and different encoding parameters. In addition, a proof of concept showing that the proposed method can also be used on videos downloaded from YouTube is reported.
Paolo Bestagini, Simone Milani, Marco Tagliasacchi, Stefano Tubaro
IEEE Trans. Image Process.1
2015 Phylogeny reconstruction for misaligned and compressed video sequences
abstract
In the last few years, the amount of videos distributed online has dramatically increased due to the popularity of media sharing platforms (e.g., YouTube, Vimeo, etc.). However, distributed videos are often edited copies of original content, typically referred to as near duplicates. In this paper, we face the problem of reconstructing a video phylogeny tree, i.e., given a set of near-duplicate videos, we want to reconstruct the relationships between every pair of videos to detect which one generated the others and trace back their evolution history. Solving this problem is of paramount importance when the first published video within a set is sought, e.g., to solve copyright infringement cases or to pinpoint criminal impersonation online. The technique we propose exploits the same rationale of previous works in the field of image and video phylogeny. However, we embed in the commonly used pipeline of operations the possibility of dealing with temporally misaligned and encoded video sequences, thus making the method applicable to user-generated videos shared on online platforms. Results computed on a wide dataset of video sequences highlight the importance of taking care of both coding and misalignment in the reconstruction pipeline.
Filipe de Oliveira Costa, Silvia Lameri, Paolo Bestagini, Zanoni Dias, Anderson Rocha 0001, Marco Tagliasacchi, Stefano Tubaro
ICIP3
2015 Near-duplicate detection and alignment for multi-view videos
abstract
The increasing popularity of video sharing platforms (e.g., YouTube, Vimeo, etc.) has determined the widespread diffusion of near-duplicate videos, i.e., sequences obtained applying different editing operations to the same original clip. However, it is also possible to come across sequences referring to the same specific event shot from different viewpoints. This is a very common situation that arises when analyzing user-generated content acquired with mobile devices. Therefore, for some applications, it can be useful to extend the concept of near-duplicates considering also all the videos (and their edited versions) referring to the same event even if shot from different viewpoints. In this paper we consider such challenging scenario. More specifically, we focus on the problem of multi-view near-duplicate video detection and temporal alignment. In doing so, we show the limitations of a state-of-the-art algorithm based on robust hashing, and propose a processing pipeline that allows to deal also with sequences taken from significantly different viewpoints.
Andrea Melloni, Silvia Lameri, Paolo Bestagini, Marco Tagliasacchi, Stefano Tubaro
ICIP3
2015 A Robust and Low-Complexity Source Localization Algorithm for Asynchronous Distributed Microphone Networks
abstract
In this paper, we propose a robust and low-complexity acoustic source localization technique based on time differences of arrival (TDOA), which addresses the scenario of distributed sensor networks in 3D environments. Network nodes are assumed to be unsynchronized, i.e., TDOAs between microphones belonging to different nodes are not available. We begin with showing how to select feasible TDOAs for each sensor node, exploiting both geometrical considerations and a characterization of the overall generalized cross correlation (GCC) shape. We then show how to localize sources in the space-range reference frame, where TDOA measurements have a clear geometrical interpretation that can be fruitfully used in the scenario of unsynchronized sensors. In this framework, in fact, the source corresponds to the apex of a hypercone passing through points described by the sole microphone positions and TDOA measurements. The localization problem is therefore approached as a hypercone fitting problem. Finally, in order to improve the robustness of the estimate, we include an outlier detection procedure based on the evaluation of the hypercone fitting residuals. A refinement of source location estimate is then performed ignoring the contributions coming from outlier measurements. A set of simulations shows the performance of individual blocks of the system, with particular focus on the effect of TDOA selection on source localization and refinement steps. Experiments on real data validate the localization algorithm in an everyday scenario, proving that good accuracy can be obtained while saving computational cost in comparison with state-of-the-art techniques.
Antonio Canclini, Paolo Bestagini, Fabio Antonacci, Marco Compagnoni, Augusto Sarti, Stefano Tubaro
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 Demosaicing strategy identification via eigenalgorithms
abstract
The identification of the camera that has acquired a specific image can be performed via several device-related footprints. Among these, it is possible to look for the traces left by the adopted color demosaicing strategy, which varies according to the camera model and vendor. The paper presents an identification strategy that re-processes the analyzed image with a set of distinctive CFA interpolation algorithms (eigenalgorithms) and, according to the correlation of the output with the original image, builds a set of features that permits identifying the algorithm. The proposed solution performs well with respect to other state-of-the-art solutions also when the analyzed image is severely compressed.
Simone Milani, Paolo Bestagini, Marco Tagliasacchi, Stefano Tubaro
ICASSP2
2014 Antiforensic synthesis of motion vectors using template algorithms
abstract
The identification of the video camera employed to acquire a video sequence is made possible by a large set of different footprints. Since video signals are always available in a compressed format, some of the most significant traces can be related to the coding tools of the implemented video codec (e.g., rate-distortion optimization, motion estimation strategy, etc.). As a matter of fact, an effective antiforensic attack, which aims at fooling the tools that identify the acquisition device, must appropriately alter these footprints. In the paper, we present an antiforensic strategy that targets a video camera detector which is based on the identification of the motion estimation strategy used by the video coder. The proposed approach synthesizes a set of motion vectors that approximate those that would have been generated by the algorithm to be mimicked. This method proves to be effective in attacking the detector while preserving the coding efficiency.
Simone Milani, Paolo Bestagini, Marco Tagliasacchi, Stefano Tubaro
ICASSP2
2014 Audio tampering detection using multimodal features
abstract
The authenticity verification of a User Generated Audio-Video content relative to a real event can be a very critical task especially when the content is shared on the Internet. Audio-Video files need to be checked in order to verify the origin of the content and the absence of alterations that could have changed their semantic content. The paper presents a multimodal approach for audio tampering detection that analyzes both the audio component and the video component of a recorded video file. The proposed solution estimates the volumetric characteristics of the environment where the multimedia content has been captured both from the video and audio signals. Then, the approach checks the consistency of the environment characteristics estimated from the audio signal with respect to those estimated from video files. The proposed solution proves to be useful in identifying video fakes and bootlegs, although it proves to be useful for the localization of added audio effects in a movie or radio track.
Simone Milani, Pier Francesco Piazza, Paolo Bestagini, Stefano Tubaro
ICASSP3
2014 Who is my parent? Reconstructing video sequences from partially matching shots
abstract
Nowadays, a significant fraction of the available video content is created by reusing already existing online videos. In these cases, the source video is seldom reused as is. Conversely, it is typically time clipped to extract only a subset of the original frames, and other transformations are commonly applied (e.g., cropping, logo insertion, etc.). In this paper, we analyze a pool of videos related to the same event or topic. We propose a method that aims at automatically reconstructing the content of the original source videos, i.e., the parent sequences, by splicing together sets of near-duplicate shots seemingly extracted from the same parent sequence. The result of the analysis shows how content is reused, thus revealing the intent of content creators, and enables us to reconstruct a parent sequence also when it is no longer available online. In doing so, we make use of a robust-hash algorithm that allows us to detect whether groups of frames are near-duplicates. Based on that, we developed an algorithm to automatically find near-duplicate matchings between multiple parts of multiple sequences. All the near-duplicate parts are finally temporally aligned to reconstruct the parent sequence. The proposed method is validated with both synthetic and real world datasets downloaded from YouTube.
Silvia Lameri, Paolo Bestagini, Andrea Melloni, Simone Milani, Anderson Rocha 0001, Marco Tagliasacchi, Stefano Tubaro
ICIP2
2013 Detection of temporal interpolation in video sequences
abstract
Nowadays, considering the availability of relatively cheap devices and powerful editing software, video tampering is a relatively easy task. Video sequences can be tampered with by performing, e.g., temporal splicing. However, if the sequences spliced together do not share the same frame rate, they have to be temporally interpolated beforehand. This operation is often made using motion compensated interpolators, which allow to minimize visual artifacts. In this paper we propose a detector of this kind of interpolation. Moreover, the detector is capable of identifying the interpolation factor used, allowing an analyst to uncover the original frame rate of a sequence. This method relies on the analysis of the correlation introduced by the filter adopted by the interpolator. Results show that detection is successful, provided that the number of observed interpolated frames is large enough. Moreover, tests on compressed sequences obtained from television broadcasts validate the method in a real world scenario.
Paolo Bestagini, S. Battaglia, Simone Milani, Marco Tagliasacchi, Stefano Tubaro
ICASSP1
2013 Video recapture detection based on ghosting artifact analysis
abstract
Video forensics is becoming a popular field of research and an increasing number of forensic techniques have been proposed in the last few years. However, a simple yet effective method to fool many detectors consists in recapturing a video sequence with a camcorder. For this reason being able to detect video recapture is a topic of interest for a forensic analyst. In this paper, we first characterize the video recapture model, focusing on the common scenario of a sequence recaptured from a LCD monitor using a digital camcorder, then we propose a recapture detector for this case. The detector is based on the analysis of a characteristic ghosting artifact left by the recapture process. The presented algorithm is finally validated by means of tests on original and recaptured sequences. These tests prove that the algorithm achieves high accuracy results.
Paolo Bestagini, Marco Visentini Scarzanella, Marco Tagliasacchi, Pier Luigi Dragotti, Stefano Tubaro
ICIP1
2013 Local tampering detection in video sequences
abstract
Video sequences are often believed to provide stronger forensic evidence than still images, e.g., when used in lawsuits. However, a wide set of powerful and easy-to-use video authoring tools is today available to anyone. Therefore, it is possible for an attacker to maliciously forge a video sequence, e.g., by removing or inserting an object in a scene. These forms of manipulation can be performed with different techniques. For example, a portion of the original video may be replaced by either a still image repeated in time or, in more complex cases, by a video sequence. Moreover, the attacker might use as source data either a spatio-temporal region of the same video, or a region taken from an external sequence. In this paper we present the analysis of the footprints left when tampering with a video sequence, and propose a detection algorithm that allows a forensic analyst to reveal video forgeries and localize them in the spatio-temporal domain. With respect to the state-of-the-art, the proposed method is completely unsupervised and proves to be robust to compression. The algorithm is validated against a dataset of forged videos available online.
Paolo Bestagini, Simone Milani, Marco Tagliasacchi, Stefano Tubaro
MMSP1
2012 Video codec identification
abstract
Video content is routinely acquired and distributed in digital format. Therefore, it is customary to have the content encoded multiple times. In this paper we consider a processing chain of two coding steps and we propose a method that aims at identifying the type of codec used in the first step, by analyzing its coding-based footprints. The method relies on the fact that lossy coding is an almost idempotent operation, i.e., re-encoding the reconstructed sequence with the same codec and coding parameters produces a sequence that is highly correlated with the input one. As a consequence, it is possible to analyze this sort of correlation to identify the first codec provided that the second codec does not introduce severe quality degradation. The proposed solution finds several applications in the field of multi-media forensics, e.g. to identify the device that generated the original video stream or detect collages of different sequences.
Paolo Bestagini, Simone Milani, Marco Tagliasacchi, Stefano Tubaro
ICASSP1
2012 Multiple compression detection for video sequences
abstract
Nowadays, thanks to the increasingly availability of powerful processors and user friendly applications, the editing of video sequences is becoming more and more frequent. Moreover, after each editing step, any video object is almost always encoded in order to store it using a less amount of memory. For this reason, inferring the number of compression steps that have been applied to such a multimedia object is an important clue in order to assess its authenticity. In this paper we propose a method to recover the number of compression steps applied to a video sequence. In order to accomplish this goal, we make use of a classifier based on multiple Support Vector Machines (SVM) exploiting the Benford's law. Indeed, the feature vectors used to train and test the SVM are based on the statistics of the most significant digit of quantized transform coefficients. The proposed method is tested with a generic hybrid video encoder combining motion-compensation and block coding. Results show that this method is able to discriminate up to three compression stages with high accuracy.
Simone Milani, Paolo Bestagini, Marco Tagliasacchi, Stefano Tubaro
MMSP2
2012 Localization of Acoustic Sources Through the Fitting of Propagation Cones Using Multiple Independent Arrays
abstract
In this paper, we propose a novel acoustic source localization method that accommodates the general scenario of multiple independent microphone arrays. The method is based on a 3-D parameter space defined by the 2-D spatial location of a source and the range difference extracted from the time difference of arrival (TDOA). In this space, the set of points that correspond to a given range lie on a circle that expand as the range increases, forming a cone whose apex is the actual location of the source. In this parameter space, the lack of synchronization between arrays results in the fact that clusters of data associated to individual arrays are free to shift along the range axis. The cone constraint, in fact, enables the realignment of such clusters while positioning the cone vertex (source location), thus resulting in a joint data re-synchronization and source localization. We also propose a novel and general analysis methodology for swiftly assessing the localization error as a function of the TDOA uncertainties, which is remarkably accurate for small localization bias. With the aid of this method, simulations and experiments on real data, we show that the cone-fitting process offers excellent localization accuracy in the scenario of multiple unsynchronized arrays, as well as in simpler single-array scenarios, also in comparison with state-of-the-art techniques. We also show that the proposed method offers the desired flexibility for adapting to arbitrary geometries of microphone clusters.
Marco Compagnoni, Paolo Bestagini, Fabio Antonacci, Augusto Sarti, Stefano Tubaro
IEEE Trans. Speech Audio Process.2
2010 Geometric calibration of distributed microphone arrays from acoustic source correspondences
abstract
This paper proposes a method that solves the problem of geometric calibration of microphone arrays. We consider a distributed system, in which each array is controlled by separate acquisition devices that do not share a common synchronization clock. Given a set of probing sources, e.g. loudspeakers, each array computes an estimate of the source locations using a conventional TDOA-based algorithm. These observations are fused together by the proposed method, in order to estimate the position and pose of one array with respect to the other. Unlike previous approaches, we explicitly consider the anisotropic distribution of localization errors. As such, the proposed method is able to address the problem of geometric calibration when the probing sources are located both in the near- and far-field of the microphone arrays. Experimental results demonstrate that the improvement in terms of calibration accuracy with respect to state-of-the-art algorithms can be substantial, especially in the far-field.
S. Daniele Valente, Marco Tagliasacchi, Fabio Antonacci, Paolo Bestagini, Augusto Sarti, Stefano Tubaro
MMSP4