Luca Cuccovillo

dblp:137/2363 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0001-5559-6508ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Towards Explainable Person-of-Interest-based Audio Synthesis Detection
abstract
Generalization and explainability are two key challenges in synthetic audio detection. Effective detectors should not only reliably classify unseen data from unknown synthesis algorithms, but also provide insight into their decision-making process and explain why a given input was classified as real or fake. To promote generalization we use the Person-of-Interest approach, which allows us to detect synthetic audio using a model trained only on real data, provided that some pristine audio of the putative speaker is provided. To support explainability, we instead use an encoder-decoder backbone such that the bottleneck features ensure syntactic and semantic fidelity to the input, as well as enable reliable decisions. Experiments show that our approach outperforms both state-of-the-art models based on supervised learning and methods based on speaker verification.
Alessandro Pianese, Luca Cuccovillo, Giovanni Poggi, Thomas Le Roux, Patrick Aichroth
IJCNN2
2024 Audio Transformer for Synthetic Speech Detection via Formant Magnitude and Phase Analysis
abstract
This paper introduces a novel multi-task transformer for synthetic speech detection. The network encodes magnitude and phase of the input speech with a feature bottleneck, used to autoencode the input magnitude, to predict the trajectory of the fundamental frequency (f0), and to discern if the input speech is synthetic or natural. The approach achieves state-of-the-art performance on the ASVspoof 2019 LA dataset while still retaining interpretability, with an AUC score of 0.910.
Luca Cuccovillo, Milica Gerhardt, Patrick Aichroth
ICASSP1
2024 MAD '24 Workshop: Multimedia AI against Disinformation
abstract
1339
Cristian Lucian Stanciu, Bogdan Ionescu, Luca Cuccovillo, Symeon Papadopoulos, Giorgos Kordopatis-Zilos, Adrian Popescu 0001, Roberto Caldelli
ICMR3
2023 MAD '23 Workshop: Multimedia AI against Disinformation
abstract
With recent advancements in synthetic media manipulation and generation, verifying multimedia content posted online has become increasingly difficult. Additionally, the malicious exploitation of AI technologies by actors to disseminate disinformation on social media, and more generally the Web, at an alarming pace poses significant threats to society and democracy. Therefore, the development of AI-powered tools that facilitate media verification is urgently needed. The MAD ’23 workshop aims to bring together individuals working on the wider topic of detecting disinformation in multimedia to exchange their experiences and discuss innovative ideas, attracting people with varying backgrounds and expertise. The research areas of interest include identifying manipulated and synthetic content in multimedia, as well as examining the dissemination of disinformation and its impact on society. The multimedia aspect is very important since content most often contains a mix of modalities and their joint analysis can boost the performance of verification methods.
Luca Cuccovillo, Bogdan Ionescu, Giorgos Kordopatis-Zilos, Symeon Papadopoulos, Adrian Popescu 0001
ICMR1
2023 Audio Splicing Detection and Localization Based on Acquisition Device Traces
abstract
In recent years, the multimedia forensic community has put a great effort in developing solutions to assess the integrity and authenticity of multimedia objects, focusing especially on manipulations applied by means of advanced deep learning techniques. However, in addition to complex forgeries as the deepfakes, very simple yet effective manipulation techniques not involving any use of state-of-the-art editing tools still exist and prove dangerous. This is the case of audio splicing for speech signals, i.e., to concatenate and combine multiple speech segments obtained from different recordings of a person in order to cast a new fake speech. Indeed, by simply adding a few words to an existing speech we can completely alter its meaning. In this work, we address the overlooked problem of detection and localization of audio splicing from different models of acquisition devices. Our goal is to determine whether an audio track under analysis is pristine, or it has been manipulated by splicing one or multiple segments obtained from different device models. Moreover, if a recording is detected as spliced, we identify where the modification has been introduced in the temporal dimension. The proposed method is based on a Convolutional Neural Network (CNN) that extracts model-specific features from the audio recording. After extracting the features, we determine whether there has been a manipulation through a clustering algorithm. Finally, we identify the point where the modification has been introduced through a distance-measuring technique. The proposed method allows to detect and localize multiple splicing points within a recording.
Daniele Ugo Leonzio, Luca Cuccovillo, Paolo Bestagini, Marco Marcon, Patrick Aichroth, Stefano Tubaro
IEEE Trans. Inf. Forensics Secur.2
2022 MAD '22 Workshop: Multimedia AI against Disinformation
abstract
The verification of multimedia content posted online becomes increasingly challenging due to recent advancements in synthetic media manipulation and generation. Moreover, malicious actors can easily exploit AI technologies to spread disinformation across social media at a rapid pace, which poses very high risks for society and democracy. There is, therefore, an urgent need for AI-powered tools that facilitate the media verification process. The objective of the MAD '22 workshop is to bring together those who work on the broader topic of disinformation detection in multimedia in order to share their experiences and discuss their novel ideas, reaching out to people with different backgrounds and expertise. The research domains of interest vary from the detection of manipulated and synthetic content in multimedia to the analysis of the spread of disinformation and its impact on society. The MAD '22 workshop proceedings are available at: https://dl.acm.org/citation.cfm?id=3512732.
Bogdan Ionescu, Giorgos Kordopatis-Zilos, Adrian Popescu 0001, Luca Cuccovillo, Symeon Papadopoulos
ICMR4
2021 Detection and localization of partial audio matches in various application scenarios
abstract
Abstract In this paper, we describe various application scenarios for archive management, broadcast/stream analysis, media search and media forensics which require the detection and accurate localization of unknown partial audio matches within items and datasets. We explain why they cannot be addressed with state-of-the-art matching approaches based on fingerprinting, and propose a new partial matching algorithm which can satisfy the relevant requirements. We propose two distinct requirement sets and hence two variants / settings for our proposed approach: One focusing on lower time granularity and hence lower computational complexity, to be able to deal with large datasets, and one focusing on fine-grain analysis for small datasets and individual items. Both variants are tested using distinct evaluation sets and methodologies and compared with a popular audio matching algorithm, thereby demonstrating that the proposed algorithm achieves convincing performance for the relevant application scenarios beyond the current state-of-the-art.
Milica Gerhardt, Patrick Aichroth, Luca Cuccovillo
Multim. Tools Appl.3
2018 Detection and Localization of Partial Audio Matches
abstract
Within recent years, several applications have emerged which require detection and accurate localization of unknown partial audio matches within a dataset. This requirement cannot be adequately addressed with state-of-the-art matching approaches based on fingerprinting. We propose a new approach that supports partial matching detection and localization within a dataset, and we evaluate it against a popular audio matching algorithm, showing that it performs significantly better for the given problem domain.
Milica Gerhardt, Patrick Aichroth, Luca Cuccovillo
CBMI3
2017 Phylogeny analysis for MP3 and AAC coding transformations
abstract
The following paper presents our work on audio phylogeny with a focus on two application scenarios: audiovisual (A/V) archives and tampering detection. Starting from a set of near-duplicate audio files, our goal is to determine the processing history for the set, and detect the transformations that have been applied on each linked pair of nodes. Our approach targets AAC and MP3 encoding operations and is addressing both music and speech material.
Milica Gerhardt, Luca Cuccovillo, Patrick Aichroth
ICME2
2016 Open-set microphone classification via blind channel analysis
abstract
In this paper, we present a new algorithm for open-set microphone classification, which is based on a pre-existing blind channel estimation approach. The proposed method achieves a Rand index above 93% for AAC, MP3 and PCM-encoded recordings from eight different mobile devices.
Luca Cuccovillo, Patrick Aichroth
ICASSP1
2016 AAC encoding detection and bitrate estimation using a convolutional neural network
abstract
In this paper, we propose a new method for AAC encoding detection and bitrate estimation from PCM material. The algorithm is based on a Convolutional Neural Network that can distinguish between eight different bitrates. It achieves an average accuracy of 94.65% by analysis of only 116.10 ms of content.
Daniel Seichter, Luca Cuccovillo, Patrick Aichroth
ICASSP2
2013 Audio tampering detection via microphone classification
abstract
In this paper, we present a new approach for audio tampering detection based on microphone classification. The underlying algorithm is based on a blind channel estimation, specifically designed for recordings from mobile devices. It is applied to detect a specific type of tampering, i.e., to detect whether footprints from more than one microphone exist within a given content item. As will be shown, the proposed method achieves an accuracy above 95% for AAC, MP3 and PCM-encoded recordings.
Luca Cuccovillo, Sebastian Mann, Marco Tagliasacchi, Patrick Aichroth
MMSP1