EDBT 2026 Demo / reviewers in the wild / expert
Davide Salvi
dblp:09/1965
· DBLP profile ↗
12ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0002-5163-3364ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2Security and privacy · 2 · 2 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Lightweight On-Device Anti-Spoofing Detection using Ternary Neural NetworksabstractIn recent years, the security risks posed by speech deepfakes have increased significantly, as synthetic speech can convincingly impersonate target identities. Although numerous deepfake detectors have been proposed, many rely on deep learning architectures that remain computationally demanding for deployment on resource-constrained devices such as smartphones. In this work, we investigate ternary quantization as a strategy to enable efficient on-device speech deepfake detection. We apply a dynamic ternary quantization scheme to a LCNN architecture, obtaining a model with 98.92% sparsity. The resulting network achieves a 98.8% reduction in multiply–accumulate operations and a 55.5% decrease in storage size compared to a lightweight state-of-the-art baseline. Experimental results show minimal degradation relative to the full-precision counterpart, competitive in-domain performance, and improved cross-dataset generalization, while maintaining robustness under codec compression and reverberation distortions. These findings indicate that drastic quantization can drastically reduce computational and memory requirements without substantially compromising detection capability, paving the way for practical real-time speech deepfake detection on edge devices. Matteo Benzo, Davide Salvi, Paolo Bestagini, Stefano Tubaro |
IH&MMSec | 2 |
| 2026 | Forensic Similarity for Speech DeepfakesabstractIn this paper, we introduce the concept of forensic similarity in the speech deepfake detection domain, which aims to determine whether two audio segments share the same underlying forensic traces. Our approach is inspired by prior work in the image domain. To transfer this idea to the audio domain, we propose a two-stage deep learning framework consisting of a Siamese-based feature extractor and a core decision module, referred to as the similarity network. The system goal to assess whether two speech samples originate from the same source by comparing their forensic characteristics. In practice, the model maps pairs of audio segments to a similarity score indicating whether they contain identical or different forensic traces. We evaluate the proposed method on the emerging task of source verification, demonstrating its ability to determine whether two speech samples were generated by the same model. In addition, we explore its applicability to audio splicing detection as a complementary use case. Experimental results show that the proposed approach generalizes well to previously unseen forensic traces, highlighting its robustness, flexibility, and practical relevance for digital audio forensics. Viola Negroni, Davide Salvi, Daniele Ugo Leonzio, Paolo Bestagini, Stefano Tubaro |
IH&MMSec | 2 |
| 2026 | Interpretable detection of singing voice manipulations using audio-language modelsabstractSinging voice manipulations have become increasingly common in modern music production. While such techniques can serve as creative tools to enhance artists’ expressive possibilities, they can also raise concerns about content authenticity and media integrity. To counter potential misuse of vocal manipulation tools, recent research has developed detection systems. However, these are typically limited to binary classification, indicating only whether a vocal track has been altered, without providing any further interpretable information. In this work, we address this limitation and propose a novel framework that combines forensic audio analysis with natural language generation to both detect and describe modifications in singing voice signals. Building on recent advances in audio-language models, we construct a dataset of manipulated and synthetic vocals annotated with detailed textual annotations, which we use to train and evaluate our framework. Our approach identifies and characterises a wide range of vocal transformations, including pitch correction, pitch shifting, time stretching, and singing voice deepfake generation. Experimental results show that the proposed method not only surpasses existing baselines in classification accuracy but also provides substantially greater interpretability, as it provides explanations of the outputs in natural language, making them understandable to non-experts. This makes the system particularly relevant for music production, media forensics, and copyright verification, offering a transparent and descriptive account of vocal alterations. Mahyar Gohari, Davide Salvi, Paolo Bestagini, Nicola Adami |
Comput. Vis. Image Underst. | 2 |
| 2026 | Splicing detection and localization for speech deepfakes using audio novelty
Davide Salvi, Francesco Castelli, Viola Negroni, Paolo Bestagini, Stefano Tubaro |
Comput. Vis. Image Underst. | 1 |
| 2025 | Audio Features Investigation for Singing Voice Deepfake DetectionabstractThe audio forensics field has recently faced a new challenge: singing voice deepfake detection. Current approaches to tackle this problem have borrowed methods initially developed for the more established task of speech deepfake detection, often simply retraining these systems on singing voice data. However, effective speech detection techniques may not necessarily perform well on singing voice, and there has been limited research on identifying the factors that can improve detection specifically in the singing domain. This paper investigates the effectiveness of various audio representations and features for discriminating real and synthetically generated singing voice signals. We evaluate two Convolutional Neural Network (CNN)-based detection systems using a wide range of audio representations, including handcrafted, learning-based, and pre-trained features. Through a systematic analysis, we aim to understand the key factors that can improve the performance of deepfake detection methods for singing voices. Additionally, we investigate the differences between singing voice and speech detection, highlighting the implications of the feature sets considered. Our results offer valuable insights and guidance for developing more advanced and effective singing voice deepfake detection systems in the future. Mahyar Gohari, Davide Salvi, Paolo Bestagini, Nicola Adami |
ICASSP | 2 |
| 2025 | Leveraging Mixture of Experts for Improved Speech Deepfake DetectionabstractSpeech deepfakes pose a significant threat to personal security and content authenticity. Several detectors have been proposed in the literature, and one of the primary challenges these systems have to face is the generalization over unseen data to identify fake signals across a wide range of datasets. In this paper, we introduce a novel approach for enhancing speech deepfake detection performance using a Mixture of Experts architecture. The Mixture of Experts framework is well-suited for the speech deepfake detection task due to its ability to specialize in different input types and handle data variability efficiently. This approach offers superior generalization and adaptability to unseen data compared to traditional single models or ensemble methods. Additionally, its modular structure supports scalable updates, making it more flexible in managing the evolving complexity of deepfake techniques while maintaining high detection accuracy. We propose an efficient, lightweight gating mechanism to dynamically assign expert weights for each input, optimizing detection performance. Experimental results across multiple datasets demonstrate the effectiveness and potential of our proposed approach. Viola Negroni, Davide Salvi, Alessandro Ilic Mezza, Paolo Bestagini, Stefano Tubaro |
ICASSP | 2 |
| 2025 | Freeze and Learn: Continual Learning with Selective Freezing for Speech Deepfake DetectionabstractIn speech deepfake detection, one of the critical aspects is developing detectors able to generalize on unseen data and distinguish fake signals across different datasets. Common approaches to this challenge involve incorporating diverse data into the training process or fine-tuning models on unseen datasets. However, these solutions can be computationally demanding and may lead to the loss of knowledge acquired from previously learned data. Continual learning techniques offer a potential solution to this problem, allowing the models to learn from unseen data without losing what they have already learned. Still, the optimal way to apply these algorithms for speech deepfake detection remains unclear, and we do not know which is the best way to apply these algorithms to the developed models. In this paper we address this aspect and investigate whether, when retraining a speech deepfake detector, it is more effective to apply continual learning across the entire model or to update only some of its layers while freezing others. Our findings, validated across multiple models, indicate that the most effective approach among the analyzed ones is to update only the weights of the initial layers, which are responsible for processing the input features of the detector. Davide Salvi, Viola Negroni, Luca Bondi, Paolo Bestagini, Stefano Tubaro |
ICASSP | 1 |
| 2025 | Source Verification for Speech DeepfakesabstractWith the proliferation of speech deepfake generators, it becomes crucial not only to assess the authenticity of synthetic audio but also to trace its origin. While source attribution models attempt to address this challenge, they often struggle in open-set conditions against unseen generators. In this paper, we introduce the source verification task, which, inspired by speaker verification, determines whether a test track was produced using the same model as a set of reference signals. Our approach leverages embeddings from a classifier trained for source attribution, computing distance scores between tracks to assess whether they originate from the same source. We evaluate multiple models across diverse scenarios, analyzing the impact of speaker diversity, language mismatch, and post-processing operations. This work provides the first exploration of source verification, highlighting its potential and vulnerabilities, and offers insights for real-world forensic applications. Viola Negroni, Davide Salvi, Paolo Bestagini, Stefano Tubaro |
INTERSPEECH | 2 |
| 2023 | Reliability Estimation for Synthetic Speech DetectionabstractRecent advances in speech synthesis and counterfeit audio generation have pushed the multimedia forensics community to develop speech deepfake detection techniques to avoid threats and unpleasant situations. Although synthetic speech detectors show excellent performance in controlled conditions, they are not always reliable in open set cases, when evaluated on data that are very different from those seen during training. This can lead to misleading scores and poorly indicative results in real-world scenarios. In this paper, we propose a method for estimating the reliability of a prediction performed by a speech deepfake detector. This enables us to perform the detection only on the most relevant portions of a signal, i.e., the time windows on which we obtain more reliable scores. This increases the final accuracy of the developed systems. As some audio fragments may not contain enough traces for the task at hand and negatively affect the system output, a reliability estimator allows us to discard them and focus only on the most pertinent data. The proposed method proves to positively impact the performance of the considered detector and shows excellent generalization capabilities on unseen datasets. Davide Salvi, Paolo Bestagini, Stefano Tubaro |
ICASSP | 1 |
| 2022 | Deepfake Speech Detection Through Emotion Recognition: A Semantic ApproachabstractIn recent years, audio and video deepfake technology has advanced relentlessly, severely impacting people’s reputation and reliability. Several factors have facilitated the growing deepfake threat. On the one hand, the hyper-connected society of social and mass media enables the spread of multimedia content worldwide in real-time, facilitating the dissemination of counterfeit material. On the other hand, neural network-based techniques have made deepfakes easier to produce and difficult to detect, showing that the analysis of low-level features is no longer sufficient for the task. This situation makes it crucial to design systems that allow detecting deepfakes at both video and audio levels. In this paper, we propose a new audio spoofing detection system leveraging emotional features. The rationale behind the proposed method is that audio deepfake techniques cannot correctly synthesize natural emotional behavior. Therefore, we feed our deepfake detector with high-level features obtained from a state-of-the-art Speech Emotion Recognition (SER) system. As the used descriptors capture semantic audio information, the proposed system proves robust in cross-dataset scenarios outperforming the considered baseline on multiple datasets. Emanuele Conti, Davide Salvi, Clara Borrelli, Brian C. Hosler, Paolo Bestagini, Fabio Antonacci, Augusto Sarti, Matthew C. Stamm, Stefano Tubaro |
ICASSP | 2 |
| 2001 | A Fault-Tolerance Scheme for a MIN-Based Multi-Sensor System
Monica Alderighi, Fabio Casini, Sergio D'Angelo, Davide Salvi, Giacomo R. Sechi |
FCCM | 4 |
| 1998 | Novel Technique for Testing FPGAsabstractThis paper presents a novel technique for testing Field Programmable Gate Arrays (FPGAs), suitable for use in the case of frequent FPGA reuse and rapid dynamic modifiability of the implemented function. Cecilia Metra, Michel Renovell, Giovanni A. Mojoli, Jean-Michel Portal, Sandro Pastore, Joan Figueras, Yervant Zorian, Davide Salvi, Giacomo R. Sechi |
DATE | 8 |