VLDB 2026 Research / reviewers in the wild / expert
Massimo Iuliani
dblp:173/8262
· DBLP profile ↗
16ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-5501-4667ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 3 since 2021Security and privacy · 5 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Explainable-by-Design Audio Deepfake Detection via Wiener-Hopf Linear PredictionabstractThe rapid advancement of synthetic speech generation methods has made audio deepfake detection a critical challenge in multimedia forensics. While recent approaches achieve high detection accuracy, they typically rely on black-box architectures that offer limited interpretability and high computational complexity. In this paper, we propose an explainable-by-design audio deepfake detection framework based on Wiener-Hopf linear prediction, processed by a lightweight 2D Convolutional Neural Network (CNN). This design enables a direct and transparent connection between classification outcomes and the acoustic properties of the signal. Experimental results on benchmark datasets demonstrate competitive detection performance while maintaining significantly lower computational complexity compared to state-of-the-art solutions. The interpretability analysis using Grad-CAM reveals that the classifier focuses on low-order predictor coefficients and on silence and transition regions, suggesting that the Wiener-Hopf predictor captures reverberation characteristics and subtle statistical inconsistencies in synthetic speech. Finally, robustness experiments show that fine-tuning effectively recovers detection performance under common post-processing degradations, including additive noise, MP3 compression, and telephone filtering. Mattia Tamiazzo, Simone Milani, Massimo Iuliani, Marco Fontani |
IH&MMSec | 3 |
| 2025 | Deepfake audio detection with spectral features and ResNeXt-based architectureabstractThe increasing prevalence of deepfake audio technologies and their potential for malicious use in fields such as politics and media has raised significant concerns regarding the ability to distinguish fake from authentic audio recordings. This study proposes a robust technique for detecting synthetic audio by leveraging three spectral features: Linear Frequency Cepstral Coefficients (LFCC), Mel Frequency Cepstral Coefficients (MFCC), and Constant Q Cepstral Coefficients (CQCC). These features are processed using an enhanced ResNeXt architecture to improve classification accuracy between genuine and spoofed audio. Additionally, a Multi-Layer Perceptron (MLP)-based fusion technique is employed to further boost the model’s performance. Extensive experiments were conducted using three datasets: the ASVspoof 2019 Logical Access (LA) dataset—featuring text-to-speech (TTS) and voice conversion attacks—the ASVspoof 2019 Physical Access (PA) dataset—including replay attacks—and the ASVspoof 2021 LA, PA and DF datasets. The proposed approach has demonstrated superior performance compared to state-of-the-art methods across all three datasets, particularly in detecting fake audio generated by text-to-speech (TTS) attacks. Its overall performance is summarized as follows: the system achieved an Equal Error Rate (EER) of 1.05% and a minimum tandem Detection Cost Function (min-tDCF) of 0.028 on the ASVspoof 2019 Logical Access (LA) dataset, and an EER of 1.14% and min-tDCF of 0.03 on the ASVspoof 2019 Physical Access(PA) dataset, demonstrating its robustness in detecting various types of audio spoofing attacks. Finally, on the ASVspoof 2021 LA dataset the method achieved an EER of 7.44% and min-tDCF of 0.35. Gul Tahaoglu, Daniele Baracchi, Dasara Shullani, Massimo Iuliani, Alessandro Piva |
Knowl. Based Syst. | 4 |
| 2024 | A Codec-Based Approach for Video Life-Cycle Characterization in Social NetworksabstractOver the past decade, the proliferation of social networks introduced new challenges in the multimedia forensic field, such as the identification of the originating platform. Significant strides have been made in the characterization of digital images, exploiting features related to the media container and content. Within the realm of videos, several efforts have been directed towards analyzing the container aspect. However, the utilization of content-based features remains limited due to the intricate nature of video encoding. In this paper, we introduce an approach to identify the source social network of a digital video by leveraging codec-based features. For the purpose, we designed a method to extract and efficiently organize detailed information from H.264/AVC-encoded videos based on a bespoke version of the video decoder tool JM. We show how the proposed method can significantly improve the process of determining the source social network, even when confronted with container-based laundering operations, surpassing existing state-of-the-art results. Giulia Bertazzini, Daniele Baracchi, Dasara Shullani, Massimo Iuliani, Alessandro Piva |
ICASSP | 4 |
| 2024 | CoFFEE: a codec-based forensic feature extraction and evaluation software for H.264 videosabstractAbstract The forensic analysis of digital videos is becoming increasingly relevant to deal with forensic cases, propaganda, and fake news. The research community has developed numerous forensic tools to address various challenges, such as integrity verification, manipulation detection, and source characterization. Each tool exploits characteristic traces to reconstruct the video life-cycle. Among these traces, a significant source of information is provided by the specific way in which the video has been encoded. While several tools are available to analyze codec-related information for images, a similar approach has been overlooked for videos, since video codecs are extremely complex and involve the analysis of a huge amount of data. In this paper, we present a new tool designed for extracting and parsing a plethora of video compression information from H.264 encoded files, including macroblocks structure, prediction residuals, and motion vectors. We demonstrate how the extracted features can be effectively exploited to address various forensic tasks, such as social network identification, source characterization, and double compression detection. We provide a detailed description of the developed software, which is released free of charge to enable its use by the research community to create new tools for forensic analysis of video files. Giulia Bertazzini, Daniele Baracchi, Dasara Shullani, Massimo Iuliani, Alessandro Piva |
EURASIP J. Inf. Secur. | 4 |
| 2024 | Uncovering the authorship: Linking media content to social user profilesabstractThe extensive spread of fake news on social networks is carried out by a diverse range of users, encompassing private individuals, newspapers, and organizations. With widely accessible image and video editing tools, malicious users can easily create manipulated media. They can then distribute this content through multiple fake profiles, aiming to maximize its social impact. To tackle this problem effectively, it is crucial to possess the ability to analyze shared media to identify the originators of fake news. To this end, multimedia forensics research has advanced tools that examine traces in media, revealing valuable insights into its origins. While combining these tools has proven to be highly efficient in creating profiles of image and video creators, it is important to note that most of these tools are not specifically designed to function effectively in the complex environment of content exchange on social networks. In this paper, we introduce the problem of establishing associations between images and their source profiles as a means to tackle the spread of disinformation on social platforms. To this end, we assembled SocialNews, an extensive image dataset comprising more than 12,000 images sourced from 21 user profiles across Facebook, Instagram, and Twitter, and we propose three increasingly realistic and challenging experimental scenarios. We present two simple yet effective techniques as benchmarks, one based on statistical analysis of Discrete Cosine Transform (DCT) coefficients and one employing a neural network model based on ResNet, and we compare their performance against the state of the art. Experimental results show that the proposed approaches exhibit superior performance in accurately classifying the originating user profiles. Daniele Baracchi, Dasara Shullani, Massimo Iuliani, Damiano Giani, Alessandro Piva |
Pattern Recognit. Lett. | 3 |
| 2022 | Social Network Identification of Laundered Videos Based on DCT Coefficient AnalysisabstractIdentifying the originating social network of a digital video is considered a relevant task to support law enforcement agencies and intelligence services in tracing producers of deceptive visual contents. Recent advances in video forensics highlighted how the structure of video containers can be extremely effective in determining the social network of provenance. However, current studies do not consider that a malicious user could easily launder the traces of the social network by rebuilding the container without transcoding. In this letter, we propose a method to identify a video’s originating social network, even when the video container structure is completely unreliable. The proposed method exploits the statistics of DCT coefficients to characterize the different social media encoding properties. With this work, we also built and made available over 1000 videos of different provenance (native, manipulated, exchanged through social networks) to aid the forensic community further researching this topic. Dasara Shullani, Daniele Baracchi, Massimo Iuliani, Alessandro Piva |
IEEE Signal Process. Lett. | 3 |
| 2021 | Checking PRNU Usability on Modern DevicesabstractThe image source identification task is mainly addressed by exploiting the unique traces of the sensor pattern noise, that ensure a negligible false alarm rate when comparing patterns extracted from different devices, even of the same brand or model. However, most recent smartphones are equipped with proprietary in-camera processing that can possibly expose unexpected correlated patterns within images belonging to different sensors.In this paper, we first highlight that wrong source attribution can happen on smartphones belonging to the same brand when images are acquired both in default and in bokeh mode. While the bokeh mode is proved to introduce a correlated pattern due to the specific in-camera post-processing, we also show that natural images also expose such issue, even when a reference from flat images is available. Furthermore, different camera models expose different correlation patterns since they are reasonably related to developers’ choices. Then, we propose a general strategy that allows the forensic practitioner to determine whether a questioned device may suffer from these correlated patterns, thus avoiding the risk of false image attribution. Chiara Albisani, Massimo Iuliani, Alessandro Piva |
ICASSP | 2 |
| 2020 | A Modified Fourier-Mellin Approach For Source Device Identification On Stabilized VideosabstractTo decide whether a digital video has been captured by a given device, multimedia forensic tools usually exploit characteristic noise traces left by the camera sensor on the acquired frames. This analysis requires that the noise pattern characterizing the camera and the noise pattern extracted from video frames under analysis are geometrically aligned. However, in many practical scenarios this does not occur, thus a re-alignment or synchronization has to be performed. Current solutions often require time consuming search of the realignment transformation parameters. In this paper, we propose to overcome this limitation by searching scaling and rotation parameters in the frequency domain. The proposed algorithm tested on real videos from a well-known state-of-the-art dataset shows promising results. Sara Mandelli, Fabrizio Argenti, Paolo Bestagini, Massimo Iuliani, Alessandro Piva, Stefano Tubaro |
ICIP | 4 |
| 2020 | Facing Image Source Attribution on iPhone X
Daniele Baracchi, Massimo Iuliani, Andrea G. Nencini, Alessandro Piva |
IWDW | 2 |
| 2020 | A vision-based fully automated approach to robust image cropping detection
Marco Fanfani, Massimo Iuliani, Fabio Bellavia, Carlo Colombo, Alessandro Piva |
Signal Process. Image Commun. | 2 |
| 2019 | Prnu Pattern Alignment for Images and Videos Based on Scene ContentabstractThis paper proposes a novel approach for registering the PRNU pattern between different camera acquisition modes by relying on the imaged scene content. First, images are aligned by establishing correspondences between local descriptors: The result can then optionally be refined by maximizing the PRNU correlation. Comparative evaluations show that this approach outperforms those based on brute-force and particle swarm optimization in terms of reliability, accuracy and speed. The proposed scene-based approach for PRNU pattern alignment is suitable for video source identification in multimedia forensics applications. Fabio Bellavia, Massimo Iuliani, Marco Fanfani, Carlo Colombo, Alessandro Piva |
ICIP | 2 |
| 2019 | FISH: Face intensity-shape histogram representation for automatic face splicing detection
Marco Fanfani, Fabio Bellavia, Massimo Iuliani, Alessandro Piva, Carlo Colombo |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | A Video Forensic Framework for the Unsupervised Analysis of MP4-Like File ContainerabstractVideo forensics keeps developing new technologies to verify the authenticity and the integrity of digital videos. While most of the existing methods rely on the analysis of the video data stream, recently, a new line of research was introduced to investigate video life cycle based on the analysis of the video container. Anyway, existing contributions in this field are based on manual comparison of video container structure and content, which is time demanding and error-prone. In this paper, we introduce a method for unsupervised analysis of video file containers, and present two main forensic applications of such method: the first one deals with video integrity verification, based on the dissimilarity between a reference and a query file container; the second one focuses on the identification and classification of the source device brand, based on the analysis of containers structure and content. Noticeably, the latter application relies on the likelihood-ratio framework, which is more and more approved by the forensic community as the appropriate way to exhibit findings in court. We tested and proved the effectiveness of both applications on a dataset composed by 578 videos taken with modern smartphones from major brands and models. The proposed approaches are proved to be valuable also for requiring an extremely small computational cost as opposed to all available techniques based on the video stream analysis or manual inspection of file containers. Massimo Iuliani, Dasara Shullani, Marco Fontani, Saverio Meucci, Alessandro Piva |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2017 | Image forgery detection confronts image composition
Victor Schetinger, Massimo Iuliani, Alessandro Piva, Manuel Menezes de Oliveira Neto |
Comput. Graph. | 2 |
| 2017 | VISION: a video and image dataset for source identificationabstractForensic research community keeps proposing new techniques to analyze digital images and videos. However, the performance of proposed tools are usually tested on data that are far from reality in terms of resolution, source device, and processing history. Remarkably, in the latest years, portable devices became the preferred means to capture images and videos, and contents are commonly shared through social media platforms (SMPs, for example, Facebook, YouTube, etc.). These facts pose new challenges to the forensic community: for example, most modern cameras feature digital stabilization, that is proved to severely hinder the performance of video source identification technologies; moreover, the strong re-compression enforced by SMPs during upload threatens the reliability of multimedia forensic tools. On the other hand, portable devices capture both images and videos with the same sensor, opening new forensic opportunities. The goal of this paper is to propose the VISION dataset as a contribution to the development of multimedia forensics. The VISION dataset is currently composed by 34,427 images and 1914 videos, both in the native format and in their social version (Facebook, YouTube, and WhatsApp are considered), from 35 portable devices of 11 major brands. VISION can be exploited as benchmark for the exhaustive evaluation of several image and video forensic tools. Dasara Shullani, Marco Fontani, Massimo Iuliani, Omar Al Shaya, Alessandro Piva |
EURASIP J. Inf. Secur. | 3 |
| 2017 | Reliability assessment of principal point estimates for forensic applications
Massimo Iuliani, Marco Fanfani, Carlo Colombo, Alessandro Piva |
J. Vis. Commun. Image Represent. | 1 |