Daniele Baracchi

dblp:198/0951 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0002-7364-1955ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Security and privacy · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 One for All: Synthesis-Free Fingerprint Learning for Attribution of In-the-Wild Synthetic Images
abstract
Attributing synthetic images to their source generative models is critical for digital forensics and security. While most existing attribution methods can distinguish images produced by known models and reject those from unknown ones, they are unable to verify whether a given image was produced by a specific, previously unseen model. To address this limitation, we formulate an open-set verification problem: determining whether a given image was generated by a specific model. Our key insight is that synthetic images from different models show consistent, content-independent fingerprints in their amplitude spectrum. Based on this insight, we design a dynamic fingerprint simulator capable of simulating over 1.6 trillion generative model architectures. We further train an extractor to capture model-specific fingerprint representations with supervised contrastive learning, enabling accurate attribution of synthetic images, even from previously unseen models. Our method does not rely on any synthetic images, instead, it is trained solely on real images. On DMDetection and AIGCBenchmark, which comprises dozens of state-of-the-art and in-the-wild generative models, our method improves the attribution performance (AUC) of the prior method from random level to 94.05% and 83.05%, respectively. On GenImage and OSMA datasets, we obtain 85.08%, and 88.48% OSCR, outperforming the SOTA methods by 4.30% and 9.37% under the same settings.
Jianwei Fei, Yunshu Dai, Peipeng Yu, Zhihua Xia, Dasara Shullani, Daniele Baracchi, Alessandro Piva
AAAI6
2026 Beyond the Brush++: a flexible pipeline for fully automated generation of realistic inpainted images
abstract
Abstract Partially manipulated images pose a growing threat to the reliability of online content. The rapid spread of diffusion-based inpainting tools has made the creation of such manipulations increasingly easy to perform. As a result, the multimedia forensics community is disadvantaged compared to the attackers, as developing effective localization techniques often requires the creation of large datasets, a resource-intensive process due to the necessary human effort. In this paper, we present Beyond the Brush+ + (BtB++), a fully automated pipeline for generating large-scale datasets of realistic inpainted images. Our experiments demonstrate that BtB++ is both flexible and easily integrates different models and configurations, offering the adaptability required to address evolving models and application scenarios. Moreover, an automatic filtering mechanism ensures quality control by discarding low-quality generated images. To provide an initial assessment of the proposed filtering strategy, we also conducted a small-scale human evaluation, studying the alignment between human perceptual judgments and the automatic metrics used for filtering.
Giulia Bertazzini, Chiara Albisani, Daniele Baracchi, Dasara Shullani, Alessandro Piva
J. Inf. Secur.3
2025 ForensiCam-215K: A Large Scale Image and Video Dataset for Forensic Analysis
abstract
Determining the origin of a digital image or video, namely device source identification, is widely used in courtroom evidence and copyright protection. Currently, device source identification primarily focuses on images captured using single camera with default settings. However, with the advancement of imaging technology, there is a large number of smartphones equipped with multiple cameras and various shooting modes for acquiring images, which may pose a significant challenge to device source identification. Therefore, to assess the performance of image source identification algorithm for modern smartphones and promote further research, it is crucial to build a dataset of image and video captured by modern smartphones. In this paper, we present a large-scale image and video dataset for forensic analysis, ForensiCam-215K. The dataset includes over 215K media contents captured by 130 modern smartphones of 10 major brands. We used the latest equipment to capture images from the main, wide-angle, and telephoto cameras in six different shooting modes, and the media were collected under a strictly controlled procedure to reduce the bias caused by differences in the acquisition process between different devices. Additionally, we used the Photo Response Non-Uniformity (PRNU) method to perform device source identification tests on the dataset. The results indicate that device source identification is a challenging task especially for images and videos captured by smartphones with multiple cameras and various shooting modes. The dataset will be released as open-source and freely available for use by the multimedia forensics research community at https://github.com/dswdsw21072/ForensiCam-215K.
Suwen Du, Pengpeng Yang 0001, Daniele Baracchi, Jinglian Jin, Dasara Shullani, Alessandro Piva
ICASSP3
2025 Deepfake audio detection with spectral features and ResNeXt-based architecture
abstract
The increasing prevalence of deepfake audio technologies and their potential for malicious use in fields such as politics and media has raised significant concerns regarding the ability to distinguish fake from authentic audio recordings. This study proposes a robust technique for detecting synthetic audio by leveraging three spectral features: Linear Frequency Cepstral Coefficients (LFCC), Mel Frequency Cepstral Coefficients (MFCC), and Constant Q Cepstral Coefficients (CQCC). These features are processed using an enhanced ResNeXt architecture to improve classification accuracy between genuine and spoofed audio. Additionally, a Multi-Layer Perceptron (MLP)-based fusion technique is employed to further boost the model’s performance. Extensive experiments were conducted using three datasets: the ASVspoof 2019 Logical Access (LA) dataset—featuring text-to-speech (TTS) and voice conversion attacks—the ASVspoof 2019 Physical Access (PA) dataset—including replay attacks—and the ASVspoof 2021 LA, PA and DF datasets. The proposed approach has demonstrated superior performance compared to state-of-the-art methods across all three datasets, particularly in detecting fake audio generated by text-to-speech (TTS) attacks. Its overall performance is summarized as follows: the system achieved an Equal Error Rate (EER) of 1.05% and a minimum tandem Detection Cost Function (min-tDCF) of 0.028 on the ASVspoof 2019 Logical Access (LA) dataset, and an EER of 1.14% and min-tDCF of 0.03 on the ASVspoof 2019 Physical Access(PA) dataset, demonstrating its robustness in detecting various types of audio spoofing attacks. Finally, on the ASVspoof 2021 LA dataset the method achieved an EER of 7.44% and min-tDCF of 0.35.
Gul Tahaoglu, Daniele Baracchi, Dasara Shullani, Massimo Iuliani, Alessandro Piva
Knowl. Based Syst.2
2025 Self-Supervised SAR Despeckling Using Deep Image Prior
abstract
Speckle noise produces a strong degradation in SAR images, characterized by a multiplicative model. Its removal is an important step of any processing chain exploiting such data. To perform this task, several model-based despeckling methods were proposed in the past years as well as, more recently, deep learning approaches. However, most of the latter ones need to be trained on a large number of pairs of noisy and clean images that, in the case of SAR images, can only be produced with the aid of synthetic noise. In this paper, we propose a self-supervised learning method based on the use of Deep Image Prior, which is extended to deal with speckle noise. The major advantage of the proposed approach lies in its ability to perform denoising without requiring any reference clean image during training. A new loss function is introduced in order to reproduce a multiplicative noise having statistics close to those of a typical speckle noise and composed also by a guidance term derived from model-based denoisers. Experimental results are presented to show the effectiveness of the proposed method and compare its performance with other reference despeckling algorithms. • A self-supervised learning method for SAR despeckling, denoted as S3DIP, is proposed. • The Deep Image Prior (DIP) approach is extended introducing a learnable noise matrix. • A histogram loss based on the knowledge of the speckle noise statistics is designed. • Existing model-based denoisers are exploited through the use of a guidance loss. • S3DIP outperforms both DIP and existing model-based denoisers on the intended task.
Chiara Albisani, Daniele Baracchi, Alessandro Piva, Fabrizio Argenti
Pattern Recognit. Lett.2
2024 A Codec-Based Approach for Video Life-Cycle Characterization in Social Networks
abstract
Over the past decade, the proliferation of social networks introduced new challenges in the multimedia forensic field, such as the identification of the originating platform. Significant strides have been made in the characterization of digital images, exploiting features related to the media container and content. Within the realm of videos, several efforts have been directed towards analyzing the container aspect. However, the utilization of content-based features remains limited due to the intricate nature of video encoding. In this paper, we introduce an approach to identify the source social network of a digital video by leveraging codec-based features. For the purpose, we designed a method to extract and efficiently organize detailed information from H.264/AVC-encoded videos based on a bespoke version of the video decoder tool JM. We show how the proposed method can significantly improve the process of determining the source social network, even when confronted with container-based laundering operations, surpassing existing state-of-the-art results.
Giulia Bertazzini, Daniele Baracchi, Dasara Shullani, Massimo Iuliani, Alessandro Piva
ICASSP2
2024 Structure Matters: Analyzing Videos Via Graph Neural Networks for Social Media Platform Attribution
abstract
Detecting the origin of a digital video within a social network is a critical task that aids law enforcement and intelligence agencies in identifying the creators of misleading visual content. In this research, we introduce an innovative method for identifying the original social network of a video, even when the video has been altered through actions like group of frames removal and file container reconstruction. The proposed method takes advantage of the video encoding’s temporal uniformity, leveraging motion vectors to characterize the specific features associated to various social media platforms. Each video is represented by a graph where nodes correspond to macroblocks. These macroblocks are interconnected by following the inter-prediction rules outlined in the H.264/AVC codec standard. Such a structure can be then classified using a graph neural network to predict the platform on which the video has been shared. Experimental results demonstrate that this approach outperforms both codec- and content-based approaches, underscoring the effectiveness of a structural approach in attributing the social media platform from which videos originated.
Andrea Gemelli, Dasara Shullani, Daniele Baracchi, Simone Marinai, Alessandro Piva
ICASSP3
2024 CoFFEE: a codec-based forensic feature extraction and evaluation software for H.264 videos
abstract
Abstract The forensic analysis of digital videos is becoming increasingly relevant to deal with forensic cases, propaganda, and fake news. The research community has developed numerous forensic tools to address various challenges, such as integrity verification, manipulation detection, and source characterization. Each tool exploits characteristic traces to reconstruct the video life-cycle. Among these traces, a significant source of information is provided by the specific way in which the video has been encoded. While several tools are available to analyze codec-related information for images, a similar approach has been overlooked for videos, since video codecs are extremely complex and involve the analysis of a huge amount of data. In this paper, we present a new tool designed for extracting and parsing a plethora of video compression information from H.264 encoded files, including macroblocks structure, prediction residuals, and motion vectors. We demonstrate how the extracted features can be effectively exploited to address various forensic tasks, such as social network identification, source characterization, and double compression detection. We provide a detailed description of the developed software, which is released free of charge to enable its use by the research community to create new tools for forensic analysis of video files.
Giulia Bertazzini, Daniele Baracchi, Dasara Shullani, Massimo Iuliani, Alessandro Piva
EURASIP J. Inf. Secur.2
2024 Uncovering the authorship: Linking media content to social user profiles
abstract
The extensive spread of fake news on social networks is carried out by a diverse range of users, encompassing private individuals, newspapers, and organizations. With widely accessible image and video editing tools, malicious users can easily create manipulated media. They can then distribute this content through multiple fake profiles, aiming to maximize its social impact. To tackle this problem effectively, it is crucial to possess the ability to analyze shared media to identify the originators of fake news. To this end, multimedia forensics research has advanced tools that examine traces in media, revealing valuable insights into its origins. While combining these tools has proven to be highly efficient in creating profiles of image and video creators, it is important to note that most of these tools are not specifically designed to function effectively in the complex environment of content exchange on social networks. In this paper, we introduce the problem of establishing associations between images and their source profiles as a means to tackle the spread of disinformation on social platforms. To this end, we assembled SocialNews, an extensive image dataset comprising more than 12,000 images sourced from 21 user profiles across Facebook, Instagram, and Twitter, and we propose three increasingly realistic and challenging experimental scenarios. We present two simple yet effective techniques as benchmarks, one based on statistical analysis of Discrete Cosine Transform (DCT) coefficients and one employing a neural network model based on ResNet, and we compare their performance against the state of the art. Experimental results show that the proposed approaches exhibit superior performance in accurately classifying the originating user profiles.
Daniele Baracchi, Dasara Shullani, Massimo Iuliani, Damiano Giani, Alessandro Piva
Pattern Recognit. Lett.1
2024 Continual learning for adaptive social network identification
abstract
The popularity of social networks as primary mediums for sharing visual content has made it crucial for forensic experts to identify the original platform of multimedia content. Various methods address this challenge, but the constant emergence of new platforms and updates to existing ones often render forensic tools ineffective shortly after release. This necessitates the regular updating of methods and models, which can be particularly cumbersome for techniques based on neural networks which cannot quickly adapt to new classes without sacrificing performance on previously learned ones – a phenomenon known as catastrophic forgetting. Recently, researchers aimed at mitigating this problem via a family of techniques known as continual learning. In this paper we study the applicability of continual learning techniques to the social network identification task by evaluating two relevant forensic scenarios: Incremental Social Platform Classification, for handling newly introduced social media platforms, and Incremental Social Version Classification, for addressing updated versions of a set of existing social networks. We perform an extensive experimental evaluation of a variety of continual learning approaches applied to these two scenarios. Experimental results demonstrate that, although Continual Social Network Identification remains a difficult problem, catastrophic forgetting can be significantly mitigated in both scenarios by retaining only a fraction of the image patches from past task training samples or by employing previous tasks prototypes.
Simone Magistri, Daniele Baracchi, Dasara Shullani, Andrew D. Bagdanov, Alessandro Piva
Pattern Recognit. Lett.2
2022 Social Network Identification of Laundered Videos Based on DCT Coefficient Analysis
abstract
Identifying the originating social network of a digital video is considered a relevant task to support law enforcement agencies and intelligence services in tracing producers of deceptive visual contents. Recent advances in video forensics highlighted how the structure of video containers can be extremely effective in determining the social network of provenance. However, current studies do not consider that a malicious user could easily launder the traces of the social network by rebuilding the container without transcoding. In this letter, we propose a method to identify a video’s originating social network, even when the video container structure is completely unreliable. The proposed method exploits the statistics of DCT coefficients to characterize the different social media encoding properties. With this work, we also built and made available over 1000 videos of different provenance (native, manipulated, exchanged through social networks) to aid the forensic community further researching this topic.
Dasara Shullani, Daniele Baracchi, Massimo Iuliani, Alessandro Piva
IEEE Signal Process. Lett.2
2020 Facing Image Source Attribution on iPhone X
Daniele Baracchi, Massimo Iuliani, Andrea G. Nencini, Alessandro Piva
IWDW1
2019 Towards Learned Color Representations for Image Splicing Detection
abstract
The detection of images that are spliced from multiple sources is one important goal of image forensics. Several methods have been proposed for this task, but particularly since the rise of social media, it is an ongoing challenge to devise forensic approaches that are highly robust to common processing operations such as strong JPEG recompression and downsampling.In this work, we make a first step towards a novel type of cue for image splicing, which is based on the color formation of an image. We make the assumption that the color formation is a joint result of the camera hardware, the software settings, and the depicted scene, and as such can be used to locate spliced patches that originally stem from different images. To this end, we train a two-stage classifier on the full set of colors from a Macbeth color chart, and compare two patches for their color consistency. Our preliminary results on a challenging dataset on downsampled data of identical scenes indicate that the color distribution can be a useful forensic tool that is highly resistant to JPEG compression.
Benjamin Hadwiger, Daniele Baracchi, Alessandro Piva, Christian Riess
ICASSP2