Luca Bondi

dblp:166/2798 · DBLP profile ↗
← Back
24ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0003-3974-7542ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Security and privacy · 2 · 1 first-authorComputer networks · 1 · 1 first-author
YearPublicationVenuePosition
2025 Freeze and Learn: Continual Learning with Selective Freezing for Speech Deepfake Detection
abstract
In speech deepfake detection, one of the critical aspects is developing detectors able to generalize on unseen data and distinguish fake signals across different datasets. Common approaches to this challenge involve incorporating diverse data into the training process or fine-tuning models on unseen datasets. However, these solutions can be computationally demanding and may lead to the loss of knowledge acquired from previously learned data. Continual learning techniques offer a potential solution to this problem, allowing the models to learn from unseen data without losing what they have already learned. Still, the optimal way to apply these algorithms for speech deepfake detection remains unclear, and we do not know which is the best way to apply these algorithms to the developed models. In this paper we address this aspect and investigate whether, when retraining a speech deepfake detector, it is more effective to apply continual learning across the entire model or to update only some of its layers while freezing others. Our findings, validated across multiple models, indicate that the most effective approach among the analyzed ones is to update only the weights of the initial layers, which are responsible for processing the input features of the detector.
Davide Salvi, Viola Negroni, Luca Bondi, Paolo Bestagini, Stefano Tubaro
ICASSP3
2025 Towards Few-Shot Training-Free Anomaly Sound Detection
Ho-Hsiang Wu, Abinaya Kumar, Luca Bondi, Shabnam Ghaffarzadegan, Juan Pablo Bello
INTERSPEECH4
2024 Can Synthetic Data Boost the Training of Deep Acoustic Vehicle Counting Networks?
abstract
In the design of traffic monitoring solutions for optimizing the urban mobility infrastructure, acoustic vehicle counting models have received attention due to their cost effectiveness and energy efficiency. Although deep learning has proven effective for visual traffic monitoring, its use has not been thoroughly investigated in the audio domain, likely due to real-world data scarcity. In this work, we propose a novel approach to acoustic vehicle counting by developing: i) a traffic noise simulation framework to synthesize realistic vehicle pass-by events; ii) a strategy to mix synthetic and real data to train a deep-learning model for traffic counting. The proposed system is capable of simultaneously counting cars and commercial vehicles driving on a two-lane road, and identifying their direction of travel under moderate traffic density conditions. With only 24 hours of labeled real-world traffic noise, we are able to improve counting accuracy on real-world data from 63% to 88% for cars and from 86% to 94% for commercial vehicles.
Stefano Damiano, Luca Bondi, Shabnam Ghaffarzadegan, Andre Guntoro, Toon van Waterschoot
ICASSP2
2024 Multi-Modal Continual Pre-Training For Audio Encoders
abstract
Several approaches have been proposed to pre-train an audio encoder to learn fundamental audio knowledge. These training frameworks range from supervised learning to self-supervised learning with a contrastive objective under multi-modal supervision. However, these approaches are constrained to a single pretext task, preventing their adaptability to multi-modal interactions beyond the modalities provided in training data. Continual learning (CL), in the meantime, allows machine learning systems to incrementally learn a new task while preserving the previously acquired knowledge, making the system more knowledgeable over time. The existing CL approaches are limited to learning downstream tasks such as classification. In this work, we propose to combine CL methods with several audio encoder pre-training methods. The audio encoders, when pre-trained continually over a sequence of multi-modal tasks, namely audiovisual and audio-text, exhibit improved performance across various downstream tasks compared to their non-continual learning counterparts, due to knowledge accumulation. The audio encoders are also capable of performing cross-modal tasks of all learned modalities.
Gyuhak Kim, Ho-Hsiang Wu, Luca Bondi, Bing Liu 0001
ICASSP3
2024 CLAP4Emo: ChatGPT-Assisted Speech Emotion Retrieval with Natural Language Supervision
abstract
Speech emotion retrieval is an important technique for large-scale and high-quality data collection. Conventional approach using ensemble of classification models might limit the retrieved emotion diversity and/or underperform in out-of-domain acoustic conditions. Natural language is diverse and agnostic to specific acoustic concepts, embedding a huge potential for developing language-based speech emotion retrieval system. In this paper we introduce CLAP4Emo, a novel framework to retrieve emotional speech via natural language prompts based on contrastive language-audio pretraining. To compensate for the absence of training captions in existing public datasets, we propose a systematic framework that applies ChatGPT to generate emotion captions. The experimental results demonstrate that our method can effectively improve the retrieved sample diversity while maintaining high precision across five benchmark datasets. By leveraging large language models, we establish a connection between audio and language for emotion description, culminating in an intuitive and interactive retrieval system. We release the generated emotion captions at: https://github.com/boschresearch/soundsee-emo-caps
Shabnam Ghaffarzadegan, Luca Bondi, Abinaya Kumar, Samarjit Das, Ho-Hsiang Wu
ICASSP3
2024 Learning Audio Concepts from Counterfactual Natural Language
abstract
Conventional audio classification relied on predefined classes, lacking the ability to learn from free-form text. Recent methods unlock learning joint audio-text embeddings from raw audio-text pairs describing audio in natural language. Despite recent advancements, there is little exploration of systematic methods to train models for recognizing sound events and sources in alternative scenarios, such as distinguishing fireworks from gunshots at outdoor events in similar situations. This study introduces causal reasoning and counterfactual analysis in the audio domain. We use counterfactual instances and include them in our model across different aspects. Our model considers acoustic characteristics and sound source information from human-annotated reference texts. To validate the effectiveness of our model, we conducted pre-training utilizing multiple audio captioning datasets. We then evaluate with several common downstream tasks, demonstrating the merits of the proposed method as one of the first works leveraging counterfactual information in audio domain. Specifically, the top-1 accuracy in open-ended language-based audio retrieval task increased by more than 43%.
Ali Vosoughi, Luca Bondi, Ho-Hsiang Wu, Chenliang Xu
ICASSP2
2024 Sound of Traffic: A Dataset for Acoustic Traffic Identification and Counting
Shabnam Ghaffarzadegan, Luca Bondi, Abinaya Kumar, Ho-Hsiang Wu, Hans-Georg Horst, Samarjit Das
INTERSPEECH2
2023 Active Learning for Abnormal Lung Sound Data Curation and Detection in Asthma
Shabnam Ghaffarzadegan, Luca Bondi, Ho-Hsiang Wu, Sirajum Munir, Kelly J. Shields, Samarjit Das, Joseph Aracri
INTERSPEECH2
2023 Background Domain Switch: A Novel Data Augmentation Technique for Robust Sound Event Detection
Luca Bondi, Shabnam Ghaffarzadegan
INTERSPEECH2
2022 Acoustic Imaging Aboard The International Space Station (ISS): Challenges and Preliminary Results
abstract
Design and execution of high fidelity acoustic sensing in complex environments poses a number of practical challenges, from accurately measuring the geometry of the setup and estimating the channel response, to time synchronization amongst the sources and receivers. When acoustic experiments are performed on-board the International Space Stations (ISS), the number of constraints and obstacles vastly increases, due to the combination of a highly unpredictable acoustic environment, and restricted availability of crew time. In this paper, we present our preliminary results with a first-of-a-kind acoustic imaging experiment performed aboard the ISS, highlighting the difference between simulations, laboratory measurements, and in-space experiments. We hope that these experiments and results will help the research community in realizing high performance acoustic imaging capabilities in complex environments.
Luca Bondi, Gabriel Chuang, Christopher Ick, Adarsh Dave, Charles Shelton, Brian Coltin, Trey Smith, Samarjit Das
ICASSP1
2022 Urban Sound & Sight: Dataset And Benchmark For Audio-Visual Urban Scene Understanding
abstract
Automatic audio-visual urban traffic understanding is a growing area of research with many potential applications of value to industry, academia, and the public sector. Yet, the lack of well-curated resources for training and evaluating models to research in this area hinders their development. To address this we present a curated audio-visual dataset, Urban Sound & Sight (Urbansas), developed for investigating the detection and localization of sounding vehicles in the wild. Urbansas consists of 12 hours of unlabeled data along with 3 hours of manually annotated data, including bounding boxes with classes and unique id of vehicles, and strong audio labels featuring vehicle types and indicating off-screen sounds. We discuss the challenges presented by the dataset and how to use its annotations for the localization of vehicles in the wild through audio models.
Magdalena Fuentes, Bea Steers, Pablo Zinemanas, Martín Rocamora, Luca Bondi, Julia Wilkins, Qianyi Shi, Yao Hou, Samarjit Das, Xavier Serra, Juan Pablo Bello
ICASSP5
2022 Learning to Adapt to Domain Shifts with Few-shot Samples in Anomalous Sound Detection
abstract
Anomaly detection has many important applications, such as monitoring industrial equipment. Despite recent advances in anomaly detection with deep-learning methods, it is unclear how existing solutions would perform under out-of-distribution scenarios, e.g., due to shifts in machine load or environmental noise. Grounded in the application of machine health monitoring, we propose a framework that adapts to new conditions with few-shot samples. Building upon prior work, we adopt a classification-based approach for anomaly detection and show its equivalence to mixture density estimation of the normal samples. We incorporate an episodic training procedure to match the few-shot setting during inference. We define multiple auxiliary classification tasks based on meta-information and leverage gradient-based meta-learning to improve generalization to different shifts. We evaluate our proposed method on a recently-released dataset of audio measurements from different machine types. It improved upon two baselines by around 10% and is on par with best-performing model reported on the dataset.
Bingqing Chen, Luca Bondi, Samarjit Das
ICPR2
2020 Video Face Manipulation Detection Through Ensemble of CNNs
abstract
In the last few years, several techniques for facial manipulation in videos have been successfully developed and made available to the masses (i.e., FaceSwap, deepfake, etc.). These methods enable anyone to easily edit faces in video sequences with incredibly realistic results and a very little effort. Despite the usefulness of these tools in many fields, if used maliciously, they can have a significantly bad impact on society (e.g., fake news spreading, cyber bullying through fake revenge porn). The ability of objectively detecting whether a face has been manipulated in a video sequence is then a task of utmost importance. In this paper, we tackle the problem of face manipulation detection in video sequences targeting modern facial manipulation techniques. In particular, we study the ensembling of different trained Convolutional Neural Network (CNN) models. In the proposed solution, different models are obtained starting from a base network (i.e., EfficientNetB4) making use of two different concepts: (i) attention layers; (ii) siamese training. We show that combining these networks leads to promising face manipulation detection results on two publicly available datasets with more than 119000 videos.
Nicolò Bonettini, Edoardo Daniele Cannas, Sara Mandelli, Luca Bondi, Paolo Bestagini, Stefano Tubaro
ICPR4
2019 Image Anonymization Detection with Deep Handcrafted Features
abstract
In recent years, the number of images shared online has continuously grown. The forensics community has kept the pace by developing techniques to both reliably extract information from these images, but also to remove it. In particular, the latest developments in image anonymization methods exposes an attack vector when used by skilled ill-intentioned image producers that may want to elude prosecution. We present an approach to detect whether or not an image has undergone a laundering process, i.e., it has been tampered with so that its unique characterizing features have been changed to avoid detection. We focus on the photo response non uniformity (PRNU) noise unique to every imaging sensor, and we consider that an image has been "laundered" when we detect the absence of PRNU from an image. We propose a per image preprocessing pipeline that generates information-rich features later used as input of fine-tuned convolutional neural networks (CNNs). We study the performance of the proposed approach using various CNN architectures and blind anonymization techniques and show its effectiveness under several training and testing scenarios. Our results also show that CNN models trained with the proposed feature are capable of generalizing over unseen devices and are robust against non-geometric transformations.
Nicolò Bonettini, David Guera, Luca Bondi, Paolo Bestagini, Edward J. Delp, Stefano Tubaro
ICIP3
2019 A Prnu-Based Method to Expose Video Device Compositions in Open-Set Setups
abstract
As the diffusion of altered video sequences over social networks and the internet can have severe consequences (e.g., fake news spreading, false accusations, etc.), the forensic research community has started actively working toward the development of methodologies tailored to assess the authenticity and integrity of video sequences. In this work, we focus on the problem of spotting video compilations composed by temporally concatenating sequences acquired with different devices. To solve this problem, we leverage trace characteristics of each video recording device left on each video at acquisition time. Specifically, we propose a method leveraging Photo-Response Non-Uniformity (PRNU) traces along with a binary classifier in order to understand which frames of a video have been acquired with the same device used to record the first few frames. Results show that the proposed solution outperforms baselines based on standard PRNU correlation and thresholding tests. Experiments have been carried out in an open-set scenario and show promising results.
Pedro Ribeiro Mendes Júnior, Luca Bondi, Paolo Bestagini, Anderson Rocha 0001, Stefano Tubaro
ICIP2
2019 Improving PRNU Compression Through Preprocessing, Quantization, and Coding
abstract
In the last decade, the extremely rapid proliferation of digital devices capable of acquiring and sharing images over the Web has significantly increased the amount of digital images publicly accessible by everyone with Internet access. Despite the obvious benefits of such technological improvements, it is becoming mandatory to verify the origin and trustfulness of such shared pictures. Photo response non-uniformity (PRNU) is the reference signal for forensic investigators when it comes to verifying or identifying which camera device shot a picture under analysis. In spite of this, PRNU is almost a white-shaped noise, thus being very difficult to compress for storage or large scale search purposes, which are frequent investigation scenarios. To overcome the issue, the forensic community has developed a series of compression algorithms. Lately, Gaussian random projections have proved to achieve state-of-the-art performance. In this paper, we propose two additional steps that help improving even more Gaussian random projections compression rate: 1) a decimation preprocessing step tailored at attenuating frequency components in which PRNU traces are already suppressed in JPEG compressed images and 2) a dead-zone quantizer (rather than the commonly used binary one) that enables an entropy coding scheme to save bitrate when storing PRNU fingerprints or sending residuals over a communication channel. Reported results show the effectiveness of proposed improvements, both under controlled JPEG compression and in a real case scenario.
Luca Bondi, Paolo Bestagini, Fernando Pérez-González, Stefano Tubaro
IEEE Trans. Inf. Forensics Secur.1
2018 Video Codec Forensics Based on Convolutional Neural Networks
abstract
The recent development of multimedia has made video editing accessible to everyone. Unfortunately, forensic analysis tools capable of detecting traces left by video processing operations in a blind fashion are still at their beginnings. One of the reasons is that videos are customary stored and distributed in a compressed format, and codec-related traces tends to mask previous processing operations. In this paper, we propose to capture video codec traces through convolutional neural networks (CNNs) and exploit them as an asset. Specifically, we train two CNN s to extract information about the used video codec and coding quality, respectively. Building upon these CNN s, we propose a system to detect and localize temporal splicing for video sequences generated from the concatenation of different video segments, which are characterized by inconsistent coding schemes and/or parameters (e.g., video compilations from different sources or broadcasting channels). The proposed solution is validated using videos at different resolutions (i.e., CIF, 4CIF, PAL and 720p) encoded with four common codecs (i.e., MPEG2, MPEG4, H264 and H265) at different qualities (i.e., different constant and variable bitrates, as well as constant quantization parameters).
Sebastiano Verde, Luca Bondi, Paolo Bestagini, Simone Milani, Giancarlo Calvagno, Stefano Tubaro
ICIP2
2017 Inpainting-Based camera anonymization
abstract
Over the years, the forensic community has developed a series of very accurate camera attribution algorithms enabling to detect which device has been used to acquire an image with outstanding results. Many of these methods are based on photo response non uniformity (PRNU) that allows tracing back a picture to the camera used to shoot it. However, when privacy is required, it would be desirable to anonymize photos, unlinking them from their specific device. This paper investigates a new and alternative approach to image anonymization task. The proposed method leverages image inpainting described as an inverse regularized problem, and does not need any priors about the PRNU to remove. Specifically, we show how PRNU pattern can be strongly attenuated by reconstructing each pixel of an image from its neighbors, only slightly affecting visual quality. Results confirm this approach as a viable alternative solution for image anonymization.
Sara Mandelli, Luca Bondi, Silvia Lameri, Vincenzo Lipari, Paolo Bestagini, Stefano Tubaro
ICIP2
2017 Aligned and non-aligned double JPEG detection using convolutional neural networks
Mauro Barni, Luca Bondi, Nicolò Bonettini, Paolo Bestagini, Andrea Costanzo, Marco Maggini, Benedetta Tondi, Stefano Tubaro
J. Vis. Commun. Image Represent.2
2017 First Steps Toward Camera Model Identification With Convolutional Neural Networks
abstract
Detecting the camera model used to shoot a picture enables to solve a wide series of forensic problems, from copyright infringement to ownership attribution. For this reason, the forensic community has developed a set of camera model identification algorithms that exploit characteristic traces left on acquired images by the processing pipelines specific of each camera model. In this letter, we investigate a novel approach to solve camera model identification problem. Specifically, we propose a data-driven algorithm based on convolutional neural networks, which learns features characterizing each camera model directly from the acquired pictures. Results on a well-known dataset of 18 camera models show that: 1) the proposed method outperforms up-to-date state-of-the-art algorithms on classification of 64 × 64 color image patches; 2) features learned by the proposed network generalize to camera models never used for training.
Luca Bondi, Luca Baroffio, David Guera, Paolo Bestagini, Edward J. Delp, Stefano Tubaro
IEEE Signal Process. Lett.1
2017 Data-Driven Feature Characterization Techniques for Laser Printer Attribution
abstract
Laser printer attribution is an increasing problem with several applications, such as pointing out the ownership of crime proofs and authentication of printed documents. However, as commonly proposed methods for this task are based on custom-tailored features, they are limited by modeling assumptions about printing artifacts. In this paper, we explore solutions able to learn discriminant-printing patterns directly from the available data during an investigation, without any further feature engineering, proposing the first approach based on deep learning to laser printer attribution. This allows us to avoid any prior assumption about printing artifacts that characterize each printer, thus highlighting almost invisible and difficult printer footprints generated during the printing process. The proposed approach merges, in a synergistic fashion, convolutional neural networks (CNNs) applied on multiple representations of multiple data. Multiple representations, generated through different pre-processing operations, enable the use of the small and lightweight CNNs whilst the use of multiple data enable the use of aggregation procedures to better determine the provenance of a document. Experimental results show that the proposed method is robust to noisy data and outperforms existing counterparts in the literature for this problem.
Anselmo Ferreira, Luca Bondi, Luca Baroffio, Paolo Bestagini, Jiwu Huang, Jefersson A. dos Santos, Stefano Tubaro, Anderson Rocha 0001
IEEE Trans. Inf. Forensics Secur.2
2016 EZ-VSN: An Open-Source and Flexible Framework for Visual Sensor Networks
abstract
We present a complete, open-source framework for rapid experimentation of visual sensor network (VSN) solutions. From the software point of view, we base our architecture on open-source and widely known C++ libraries to provide the basic image processing and networking primitives. The resulting system can be leveraged to create different types of VSNs, characterized by the presence of multiple cameras, relays and cooperator nodes, and can be run on any Linux-based hardware platform, such as the BeagleBone Black. To demonstrate the flexibility of the proposed framework, we describe two different application scenarios typical of VSNs, namely object recognition and parking monitoring. The framework is then used to evaluate the benefits of two complementary paradigms for networked visual analysis recently discussed in the literature. In the traditional compress-then-analyze (CTA) paradigm, compressed images are transmitted from camera nodes to a central controller, where they are analyzed. In the novel analyze-then-compress (ATC) paradigm, camera nodes extract and compress local features from the acquired images. Such features are transmitted to the central controller and used to perform visual analysis. We show that the ATC paradigm outperforms CTA from the consumed energy point of view, at the same target analysis accuracy in both the application scenarios.
Luca Bondi, Luca Baroffio, Matteo Cesana, Alessandro Redondi, Marco Tagliasacchi
IEEE Internet Things J.1
2016 Rate-energy-accuracy optimization of convolutional architectures for face recognition
Luca Bondi, Luca Baroffio, Matteo Cesana, Marco Tagliasacchi, Giovani Chiachia, Anderson Rocha 0001
J. Vis. Commun. Image Represent.1
2016 Deep Convolutional Neural Networks for pedestrian detection
Denis Tomè, Federico Monti, Luca Baroffio, Luca Bondi, Marco Tagliasacchi, Stefano Tubaro
Signal Process. Image Commun.4