Stefano Tubaro

dblp:52/3820 · DBLP profile ↗
← Back
220ranked-venue papers
2as first author
40since 2021 · last 2026
0000-0002-1990-9869ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 177 · 2 first-author · 22 since 2021Artificial intelligence and machine learning · 23 · 4 since 2021Security and privacy · 11 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Computer networks · 4 · 1 since 2021Systems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Lightweight On-Device Anti-Spoofing Detection using Ternary Neural Networks
abstract
In recent years, the security risks posed by speech deepfakes have increased significantly, as synthetic speech can convincingly impersonate target identities. Although numerous deepfake detectors have been proposed, many rely on deep learning architectures that remain computationally demanding for deployment on resource-constrained devices such as smartphones. In this work, we investigate ternary quantization as a strategy to enable efficient on-device speech deepfake detection. We apply a dynamic ternary quantization scheme to a LCNN architecture, obtaining a model with 98.92% sparsity. The resulting network achieves a 98.8% reduction in multiply–accumulate operations and a 55.5% decrease in storage size compared to a lightweight state-of-the-art baseline. Experimental results show minimal degradation relative to the full-precision counterpart, competitive in-domain performance, and improved cross-dataset generalization, while maintaining robustness under codec compression and reverberation distortions. These findings indicate that drastic quantization can drastically reduce computational and memory requirements without substantially compromising detection capability, paving the way for practical real-time speech deepfake detection on edge devices.
Matteo Benzo, Davide Salvi, Paolo Bestagini, Stefano Tubaro
IH&MMSec4
2026 Forensic Similarity for Speech Deepfakes
abstract
In this paper, we introduce the concept of forensic similarity in the speech deepfake detection domain, which aims to determine whether two audio segments share the same underlying forensic traces. Our approach is inspired by prior work in the image domain. To transfer this idea to the audio domain, we propose a two-stage deep learning framework consisting of a Siamese-based feature extractor and a core decision module, referred to as the similarity network. The system goal to assess whether two speech samples originate from the same source by comparing their forensic characteristics. In practice, the model maps pairs of audio segments to a similarity score indicating whether they contain identical or different forensic traces. We evaluate the proposed method on the emerging task of source verification, demonstrating its ability to determine whether two speech samples were generated by the same model. In addition, we explore its applicability to audio splicing detection as a complementary use case. Experimental results show that the proposed approach generalizes well to previously unseen forensic traces, highlighting its robustness, flexibility, and practical relevance for digital audio forensics.
Viola Negroni, Davide Salvi, Daniele Ugo Leonzio, Paolo Bestagini, Stefano Tubaro
IH&MMSec5
2026 Splicing detection and localization for speech deepfakes using audio novelty
Davide Salvi, Francesco Castelli, Viola Negroni, Paolo Bestagini, Stefano Tubaro
Comput. Vis. Image Underst.5
2025 Low-Power Hierarchical Network: Pervasive Eye-Tracking on Smart Eyewear
abstract
Pervasive eye-tracking technology for eyewear devices represents a major advancement in wearable computing, enabling intuitive interaction and improving accessibility. However, the low-power constraints of these devices present a significant challenge in balancing accuracy with limited computational capacity. This study focuses on developing and evaluating algorithms for a low-power wearable infrared eye-tracking system conceived to work 24/7. The system includes a custom-built prototype that integrates infrared LEDs and photodiodes, strategically positioned on smart eyewear to estimate gaze direction. A humanoid robot, Ami Desktop, was utilized to create a controlled and robust dataset. Two deep learning architectures were investigated: a Multi-Layer Perceptron (MLP) and a tailored Hierarchical Neural Network (HNN). Variants of these models incorporating dimensionality reduction techniques were implemented to optimize performance and efficiency for lowpower microcontrollers. The results demonstrate the superior accuracy and reasonable computational demands of the HNN models, highlighting their potential for continuous, real-time and portable eye-tracking applications.
Carlo Pezzoli, Emanuele Santoro, Marco Paracchini, Giulio Marano, Daniele Bani, Luca Francesco Raduzzi, Daniele M. Crafa, Marco Carminati, Luca Merigo, Tommaso Ongarello, Marco Marcon, Stefano Tubaro
ETRA12
2025 Leveraging Mixture of Experts for Improved Speech Deepfake Detection
abstract
Speech deepfakes pose a significant threat to personal security and content authenticity. Several detectors have been proposed in the literature, and one of the primary challenges these systems have to face is the generalization over unseen data to identify fake signals across a wide range of datasets. In this paper, we introduce a novel approach for enhancing speech deepfake detection performance using a Mixture of Experts architecture. The Mixture of Experts framework is well-suited for the speech deepfake detection task due to its ability to specialize in different input types and handle data variability efficiently. This approach offers superior generalization and adaptability to unseen data compared to traditional single models or ensemble methods. Additionally, its modular structure supports scalable updates, making it more flexible in managing the evolving complexity of deepfake techniques while maintaining high detection accuracy. We propose an efficient, lightweight gating mechanism to dynamically assign expert weights for each input, optimizing detection performance. Experimental results across multiple datasets demonstrate the effectiveness and potential of our proposed approach.
Viola Negroni, Davide Salvi, Alessandro Ilic Mezza, Paolo Bestagini, Stefano Tubaro
ICASSP5
2025 Freeze and Learn: Continual Learning with Selective Freezing for Speech Deepfake Detection
abstract
In speech deepfake detection, one of the critical aspects is developing detectors able to generalize on unseen data and distinguish fake signals across different datasets. Common approaches to this challenge involve incorporating diverse data into the training process or fine-tuning models on unseen datasets. However, these solutions can be computationally demanding and may lead to the loss of knowledge acquired from previously learned data. Continual learning techniques offer a potential solution to this problem, allowing the models to learn from unseen data without losing what they have already learned. Still, the optimal way to apply these algorithms for speech deepfake detection remains unclear, and we do not know which is the best way to apply these algorithms to the developed models. In this paper we address this aspect and investigate whether, when retraining a speech deepfake detector, it is more effective to apply continual learning across the entire model or to update only some of its layers while freezing others. Our findings, validated across multiple models, indicate that the most effective approach among the analyzed ones is to update only the weights of the initial layers, which are responsible for processing the input features of the detector.
Davide Salvi, Viola Negroni, Luca Bondi, Paolo Bestagini, Stefano Tubaro
ICASSP5
2025 Source Verification for Speech Deepfakes
abstract
With the proliferation of speech deepfake generators, it becomes crucial not only to assess the authenticity of synthetic audio but also to trace its origin. While source attribution models attempt to address this challenge, they often struggle in open-set conditions against unseen generators. In this paper, we introduce the source verification task, which, inspired by speaker verification, determines whether a test track was produced using the same model as a set of reference signals. Our approach leverages embeddings from a classifier trained for source attribution, computing distance scores between tracks to assess whether they originate from the same source. We evaluate multiple models across diverse scenarios, analyzing the impact of speaker diversity, language mismatch, and post-processing operations. This work provides the first exploration of source verification, highlighting its potential and vulnerabilities, and offers insights for real-world forensic applications.
Viola Negroni, Davide Salvi, Paolo Bestagini, Stefano Tubaro
INTERSPEECH4
2025 Enhanced Water Leak Detection with Convolutional Neural Networks and One-Class Support Vector Machine
Daniele Ugo Leonzio, Paolo Bestagini, Marco Marcon, Stefano Tubaro
Networking4
2025 Hiding Local Manipulations on SAR Images: A Counter-Forensic Attack
abstract
The vast accessibility of Synthetic Aperture Radar (SAR) images through online portals has propelled the research across various fields. This widespread use and easy availability have unfortunately made SAR data susceptible to malicious alterations, such as local editing applied to the images for inserting or covering the presence of sensitive targets. To contrast malicious manipulations, in the last years the forensic community has begun to dig into the SAR manipulation issue, proposing detectors that effectively localize the tampering traces in amplitude images. Nonetheless, in this paper we demonstrate that an expert practitioner can exploit the complex nature of SAR data to obscure any signs of manipulation within a locally altered amplitude image. We refer to this approach as a counter-forensic attack. To achieve the concealment of manipulation traces, the attacker can simulate a re-acquisition of the manipulated scene by the SAR system that initially generated the pristine image. In doing so, the attacker can obscure any evidence of manipulation, making it appear as if the image was legitimately produced by the system. This attack has unique features that make it both highly generalizable and relatively easy to apply. First, it is a black-box attack, meaning it is not designed to deceive a specific forensic detector. Furthermore, it does not require a training phase and is not based on adversarial operations. We assess the effectiveness of the proposed counter-forensic approach across diverse scenarios, examining various manipulation operations. The obtained results indicate that our devised attack successfully eliminates traces of manipulation, deceiving even the most advanced forensic detectors.
Sara Mandelli, Edoardo Daniele Cannas, Paolo Bestagini, Stefano Tebaldini, Stefano Tubaro
IEEE Trans. Image Process.5
2024 A One-Class Approach to Detect Super-Resolution Satellite Imagery with Spectral Features
abstract
Satellite imagery has a vital role in many applications and several techniques exist to enhance its quality. An example is given by image Super Resolution (SR), which aims at increasing the pixel resolution to recover lost high-frequency details. Due to their importance, satellite images are also a target for malicious manipulations. In such a context, knowing if an image has been super-resolved (SRV) is crucial for guaranteeing the correct usage of this kind of data. In this paper, we propose a pipeline for detecting satellite images SRV through State-Of-The-Art (SOTA) techniques based on Convolutional Neural Networks. Our solution employs an anomaly detection algorithm, i.e., a one-class Support Vector Machine trained on Fourier spectral features of native high-resolution images. Even though we never process SRV samples during training, the experiments show that we can reject images generated through different SOTA techniques, and our solution is robust against image compression.
Edoardo Daniele Cannas, P. Beaus, Paolo Bestagini, F. Marques, Stefano Tubaro
ICASSP5
2024 Water Leak Detection via Domain Adaptation
abstract
Outdated infrastructure contributes to significant water wastage, where leaks can represent as much as 30% of urban water supply losses. Rapid and precise leak detection is therefore crucial for economic and environmental reasons. Data-driven methods have emerged as promising solutions to detect water leaks due to their accurate performance. However, they encounter obstacles like limited labeled datasets and adapting to various situations. To address these challenges, we explore Semi Supervised Learning (SSL) and Transfer Learning (TL) techniques in the context of water leak detection. We propose to address the problem of leak detection in case of limited labeled data, using a Convolutional Neural Network (CNN) trained on a laboratory-scale network, then adapted to work on real data. To do so, we compare three different domain adaptation techniques that leverage only a small amount of data from the new domain. Our results show that Self-Tuning techniques proves better than the others for this task, even with limited data.
Daniele Ugo Leonzio, Paolo Bestagini, Marco Marcon, Gian Paolo Quarta, Stefano Tubaro
ICASSP5
2024 Mdrt: Multi-Domain Synthetic Speech Localization
abstract
With recent advancements in generating synthetic speech, tools to generate high-quality synthetic speech impersonating any human speaker are easily available. Several incidents report misuse of high-quality synthetic speech for spreading misinformation and for large-scale financial frauds. Many methods have been proposed for detecting synthetic speech; however, there is limited work on localizing the synthetic segments within the speech signal. In this work, our goal is to localize the synthetic speech segments in a partially synthetic speech signal. Most existing methods for synthetic speech localization obtain features from either the time domain waveform or the spectrogram representation of the speech signal. In this work, we propose Multi-Domain ResNet Transformer (MDRT) that obtains multi-domain features from both the time domain and the spectrogram representation of a speech signal to localize synthetic speech segments. MDRT uses transformer neural networks to obtain multi-domain features and processes them using a ResNet-style neural network. We use the PartialSpoof dataset to examine the performance of MDRT on localizing synthetic speech segments of varying duration. Our results show that MDRT performs better than several existing synthetic speech localization methods.
Amit Kumar Singh Yadav, Kratika Bhagtani, Sriram Baireddy, Paolo Bestagini, Stefano Tubaro, Edward J. Delp
ICASSP5
2024 Back to the Future: GNN-Based No2 Forecasting Via Future Covariates
abstract
Due to the latest environmental concerns in keeping at bay contaminants emissions in urban areas, air pollution forecasting has been rising the forefront of all researchers around the world. When predicting pollutant concentrations, it is common to include the effects of environmental factors that influence these concentrations within an extended period, like traffic, meteorological conditions and geographical information. Most of the existing approaches exploit this information as past covariates, i.e., past exogenous variables that affected the pollutant but were not affected by it. In this paper, we present a novel forecasting methodology to predict NO2 concentration via both past and future covariates. Future covariates are represented by weather forecasts and future calendar events, which are already known at prediction time. In particular, we deal with air quality observations in a city-wide network of ground monitoring stations, modeling the data structure and estimating the predictions with a Spatiotemporal Graph Neural Network (STGNN). We propose a conditioning block that embeds past and future covariates into the current observations. After extracting meaningful spatiotemporal representations, these are fused together and projected into the forecasting horizon to generate the final prediction. To the best of our knowledge, it is the first time that future covariates are included in time series predictions in a structured way. Remarkably, we find that conditioning on future weather information has a greater impact than considering past traffic conditions. We release our code implementation at https://github.com/polimi-ispl/MAGCRN.
Antonio Giganti, Sara Mandelli, Paolo Bestagini, Umberto Giuriato, Alessandro D'Ausilio, Marco Marcon, Stefano Tubaro
IGARSS7
2024 Investigating Translation Invariance and Shiftability in CNNs for Robust Multimedia Forensics: A JPEG Case Study
Edoardo Daniele Cannas, Sara Mandelli, Paolo Bestagini, Stefano Tubaro
IH&MMSec4
2024 Self-Supervised Seismic Swell Noise Suppression From Noisy Seismic Data
abstract
Seismic swell noise, often observed in marine seismic data, is characterized by high amplitude and low frequencies. This noise significantly hides useful signals, underscoring the importance of attenuating it in the processing pipeline for marine seismic data. To date, most deep learning methods proposed for swell noise suppression have focused on supervised paradigms, which require a large dataset of paired noisy and clean data to gain insights into signal features and swell noise statistics at the training phase. In field data processing, however, obtaining generalizable training data often proves to be challenging. The inherent complexity and variability of the real world often make it difficult to acquire unbiased, noise-free datasets in the context of marine seismic data processing. To address this problem, we present a self-supervised deep learning method for suppressing swell noise even in the absence of access to clean seismic training data. First, we propose a strategy for the synthetic swell noise generation based on the amplitude, frequency, and coherence features of noisy traces. We implement the noisy-as-clean (NAC) strategy, wherein either the original noisy seismic data or its reorganized variant acts as the network’s target. Simultaneously, the network receives the observed noisy seismic data combined with simulated swell noise as its input. By leveraging these “noisy-noisy” pairs, we train a DnCNN network. Experimental evaluations conducted on both synthetic and field data demonstrate that, through the integration of swell noise simulation and the NAC strategy, the trained network consistently achieves superior denoising performance.
Weiwei Xu 0004, Vincenzo Lipari, Paolo Bestagini, Stefano Tubaro
IEEE Trans. Geosci. Remote. Sens.5
2023 Water Leak Detection and Localization Using Convolutional Autoencoders
abstract
Water is a valuable resource that has to be handled appropriately. However, a significant volume of water is wasted annually due to leaks in Water Distribution Networks (WDNs). This emphasizes the necessity for reliable and effective leak detection and localization systems. Several types of solutions have been proposed during the last few years. Among these solutions, data-driven ones are gaining more traction due to their impressive performance. In this paper, we propose a new method for leak detection and localization. The method is based on water pressure measurements acquired at a series of nodes of a WDN. Our technique is a fully data-driven solution that makes only use of the knowledge of the WDN topology, and a series of pressure data acquisitions obtained in absence of leaks. The proposed solution is based on an autoencoder trained on no-leak data, so that leaks are detected as anomalies. The results achieved on the LeakDB dataset demonstrate that the proposed solution outperforms recent methods for leak detection and localization.
Daniele Ugo Leonzio, Paolo Bestagini, Marco Marcon, Gian Paolo Quarta, Stefano Tubaro
ICASSP5
2023 Reliability Estimation for Synthetic Speech Detection
abstract
Recent advances in speech synthesis and counterfeit audio generation have pushed the multimedia forensics community to develop speech deepfake detection techniques to avoid threats and unpleasant situations. Although synthetic speech detectors show excellent performance in controlled conditions, they are not always reliable in open set cases, when evaluated on data that are very different from those seen during training. This can lead to misleading scores and poorly indicative results in real-world scenarios. In this paper, we propose a method for estimating the reliability of a prediction performed by a speech deepfake detector. This enables us to perform the detection only on the most relevant portions of a signal, i.e., the time windows on which we obtain more reliable scores. This increases the final accuracy of the developed systems. As some audio fragments may not contain enough traces for the task at hand and negatively affect the system output, a reliability estimator allows us to discard them and focus only on the most pertinent data. The proposed method proves to positively impact the performance of the considered detector and shows excellent generalization capabilities on unseen datasets.
Davide Salvi, Paolo Bestagini, Stefano Tubaro
ICASSP3
2023 ASSD: Synthetic Speech Detection in the AAC Compressed Domain
abstract
Synthetic human speech signals have become very easy to generate given modern text-to-speech methods. When these signals are shared on social media they are often compressed using the Advanced Audio Coding (AAC) standard. Our goal is to study if a small set of coding metadata contained in the AAC compressed bit stream is sufficient to detect synthetic speech. This would avoid decompressing of the speech signals before analysis. We call our proposed method AAC Synthetic Speech Detection (ASSD). ASSD extracts information from the AAC compressed bit stream without decompressing the speech signal. ASSD analyzes the information using a transformer neural network. In our experiments, we compressed the ASVspoof2019 dataset according to the AAC standard using different data rates. We compared the performance of ASSD to a time domain based and a spectrogram based synthetic speech detection methods. We evaluated ASSD on approximately 71k compressed speech signals. The results show that our proposed method typically only requires 1000 bits per speech block/frame from the AAC compressed bit stream to detect synthetic speech. This is much lower than other reported methods. Our method also had a 9.7 percentage points higher detection accuracy compared to existing methods.
Amit Kumar Singh Yadav, Ziyue Xiang, Emily R. Bartusiak, Paolo Bestagini, Stefano Tubaro, Edward J. Delp
ICASSP5
2023 Super-Resolution of BVOC Maps by Adapting Deep Learning Methods
abstract
Biogenic Volatile Organic Compounds (BVOCs) play a critical role in biosphere-atmosphere interactions, being a key factor in the physical and chemical properties of the atmosphere and climate. Acquiring large and fine-grained BVOC emission maps is expensive and time-consuming, so most available BVOC data are obtained on a loose and sparse sampling grid or on small regions. However, high-resolution BVOC data are desirable in many applications, such as air quality, atmospheric chemistry, and climate monitoring. In this work, we investigate the possibility of enhancing BVOC acquisitions, further explaining the relationships between the environment and these compounds. We do so by comparing the performances of several state-of-the-art neural networks proposed for image Super-Resolution (SR), adapting them to overcome the challenges posed by the large dynamic range of the emission and reduce the impact of outliers in the prediction. Moreover, we also consider realistic scenarios, considering both temporal and geographical constraints. Finally, we present possible future developments regarding SR generalization, considering the scale-invariance property and super-resolving emissions from unseen compounds.
Antonio Giganti, Sara Mandelli, Paolo Bestagini, Marco Marcon, Stefano Tubaro
ICIP5
2023 It Wasn't Me: Irregular Identity in Deepfake Videos
abstract
With the rapid development in media generation technologies, the creation of DeepFake videos is within everyone’s reach. As the widespread diffusion of DeepFakes can lead to severe consequences (e.g., defamation, fake news spreading, etc.), detecting DeepFakes is becoming a crucial task within the forensic community. However, most of the existing DeepFake detectors suffer from two issues: i) they are hardly explainable as they build upon black-box data-driven techniques rather than interpretable features; ii) they are often tailored to low-level texture features, failing to generalize on low-quality DeepFake videos. In this work we propose a video DeepFake detector that aims at solving these issues. The proposed detector relies on the fact that most DeepFake generators work on a frame-by-frame basis, thus breaking the temporal consistency of facial features across frames. In particular, we noticed that facial identity features tend to be less stable in time on DeepFake videos than original ones. We therefore propose a framework trained on time series of facial identity features. The use of high-level semantic features makes the detector interpretable and robust against low-quality DeepFake videos. Extensive experiments show that our method achieves outstanding performance on low-quality DeepFake video and obtains promising results on unseen dataset evaluation. The code is available at https://github.com/HongguLiu/Identity-Inconsistency-DeepFake-Detection
Honggu Liu, Paolo Bestagini, Wenbo Zhou 0004, Stefano Tubaro, Weiming Zhang 0001, Nenghai Yu
ICIP5
2023 DSVAE: Disentangled Representation Learning for Synthetic Speech Detection
abstract
Tools to generate high quality synthetic speech that is perceptually indistinguishable from speech recorded from hu-man speakers are easily available. Many incidents report misuse of synthetic speech for spreading misinformation and committing financial fraud. Several approaches have been proposed for detecting synthetic speech. Many of these approaches use deep learning methods without providing reasoning for the decisions they make. This limits the explainability of these approaches. In this paper, we use disentangled representation learning for developing a synthetic speech detector. We propose Disentangled Spectrogram Variational Auto Encoder (DSVAE) which is a two stage trained variational autoencoder that processes spectrograms of speech to generate features that disentangle synthetic and bona fide speech. We evaluated DSVAE using the ASVspoof2019 dataset. Our experimental results show high accuracy (> 98%) on detecting synthetic speech from 6 known and 10 unknown speech synthesizers. Further, the visualization of disentangled features obtained from DSVAE provides rea-soning behind the working principle of DSVAE, improving its explainability. DSVAE performs well compared to several existing methods. Additionally, DSVAE works in practical scenarios such as detecting synthetic speech uploaded on social platforms and against simple attacks such as removing silence regions.
Amit Kumar Singh Yadav, Kratika Bhagtani, Ziyue Xiang, Paolo Bestagini, Stefano Tubaro, Edward J. Delp
ICMLA5
2023 PS3DT: Synthetic Speech Detection Using Patched Spectrogram Transformer
abstract
Many deep learning synthetic speech generation tools are readily available. The use of synthetic speech has caused financial fraud, impersonation of people, and misinformation to spread. For this reason forensic methods that can detect synthetic speech have been proposed. Existing methods often overfit on one dataset and their performance reduces substantially in practical scenarios such as detecting synthetic speech shared on social platforms. In this paper we propose, Patched Spectrogram Synthetic Speech Detection Transformer (PS3DT), a synthetic speech detector that converts a time domain speech signal to a mel-spectrogram and processes it in patches using a trans-former neural network. We evaluate the detection performance of PS3DT on ASVspoof2019 dataset. Our experiments show that PS3DT performs well on ASVspoof2019 dataset compared to other approaches using spectrogram for synthetic speech detection. We also investigate generalization performance of PS3DT on In-the-Wild dataset. PS3DT generalizes well than several existing methods on detecting synthetic speech from an out-of-distribution dataset. We also evaluate robustness of PS3DT to detect telephone quality synthetic speech and synthetic speech shared on social platforms (compressed speech). PS3DT is robust to compression and can detect telephone quality synthetic speech better than several existing methods.
Amit Kumar Singh Yadav, Ziyue Xiang, Kratika Bhagtani, Paolo Bestagini, Stefano Tubaro, Edward J. Delp
ICMLA5
2023 Robust Water Leak Detection and Localization with Graph Signal Processing
abstract
Water is a resource that has to be managed properly. Nevertheless, a sizable amount of water is lost each year because of leaks in Water Distribution Networks (WDNs). The need for trustworthy and efficient leak detection and localization systems is therefore an urgent necessity. For this reason, different solutions have been put out in recent years. Due to their outstanding performance, data-driven methods are among those that are gaining the most popularity. However, the performance of data-driven approaches depend on the coherence between data on which they are trained and data on which they are tested. For example, if the acquired test data look corrupted and incoherent with training ones due to sensor failure, the performance of the overall system may be severely hindered. In this work we present a resilient water leak detection and localization algorithm. It is based on two main steps: the first step analyzes acquired data to possibly recover corrupted ones by means of graph interpolation; the second step finds leaks exploiting an autoencoder-based anomaly detector proposed in the literature. The results show that the suggested approach for signal recovery by means of graph interpolation enables the detector to work in situations in which it would otherwise fail. In doing so, we address a problem that has so far received little attention in the literature: potential sensor failures when acquiring data.
Daniele Ugo Leonzio, Paolo Bestagini, Marco Marcon, Gian Paolo Quarta, Stefano Tubaro
IECON5
2023 Super-Resolution of Bvoc Emission Maps Via Domain Adaptation
abstract
Enhancing the resolution of Biogenic Volatile Organic Compound (BVOC) emission maps is a critical task in remote sensing. Recently, some Super-Resolution (SR) methods based on Deep Learning (DL) have been proposed, leveraging data from numerical simulations for their training process. However, when dealing with data derived from satellite observations, the reconstruction is particularly challenging due to the scarcity of measurements to train SR algorithms with. In our work, we aim at super-resolving low resolution emission maps derived from satellite observations by leveraging the information of emission maps obtained through numerical simulations. To do this, we combine a SR method based on DL with Domain Adaptation (DA) techniques, harmonizing the different aggregation strategies and spatial information used in simulated and observed domains to ensure compatibility. We investigate the effectiveness of DA strategies at different stages by systematically varying the number of simulated and observed emissions used, exploring the implications of data scarcity on the adaptation strategies. To the best of our knowledge, there are no prior investigations of DA in satellite-derived BVOC maps enhancement. Our work represents a first step toward the development of robust strategies for the reconstruction of observed BVOC emissions.
Antonio Giganti, Sara Mandelli, Paolo Bestagini, Marco Marcon, Stefano Tubaro
IGARSS5
2023 Extracting Efficient Spectrograms From MP3 Compressed Speech Signals for Synthetic Speech Detection
abstract
Many speech signals are compressed with MP3 to reduce the data rate. In many synthetic speech detection methods the spectrogram of the speech signal is used. This usually requires the speech signal to be fully decompressed. We show that the design of MP3 compression allows one to approximate the spectrogram of the MP3 compressed speech efficiently without fully decoding the compressed speech. We denote the spectograms obtained using our proposed approach by Efficient Spectrograms (E-Specs). E-Spec can reduce the complexity of spectrogram computation by ~77.60 percentage points (p.p.) and save ~37.87 p.p. of MP3 decoding time. E-Spec bypasses the reconstruction artifacts introduced by the MP3 synthesis filterbank, which makes it useful in speech forensics tasks. We tested E-Spec in the synthetic speech detection, where a detector is asked to determine whether a speech signal is synthesized or recorded from a human. We examined 4 different neural network architectures to evaluate the performance of E-Spec compared to speech features extracted from the fully decoded speech signal. E-Spec achieved the best synthetic speech detection performance for 3 architectures; it also achieved the best overall detection performance across architectures. The computation of E-Spec is an approximation to Short Time Fourier Transform (STFT). E-Spec can be extended to other audio compression methods.
Ziyue Xiang, Amit Kumar Singh Yadav, Stefano Tubaro, Paolo Bestagini, Edward J. Delp
IH&MMSec3
2023 BiFPro: A Bidirectional Facial-data Protection Framework against DeepFake
abstract
The rapid progress of the DeepFake technique has caused severe privacy problems. Thus protecting facial data against DeepFake becomes an urgent requirement. Face protection can be regarded as a bidirectional process: Face-out-detection (FOD) and Face-in-forensics (FIF). For FOD, the detectability should be satisfied when using the protected face to replace other faces. For FIF, traceability should be guaranteed when the protected face is replaced by others. For this, we propose a Bidirectional Facial-data Protection Framework (BiFPro) to protect face data comprehensively. This framework is composed of three main parts: Watermarking embedding, Face-out-detection (FOD) and Face-in-forensics (FIF). For the FOD case, we ensure the vulnerability of the original face by embedding fragile watermarking. Once the protected facial image is used to replace other faces, the watermarking information will be corrupted in the synthesized face images which can be used to detect the authenticity of the protected facial images. As for the FIF case, we guarantee the traceability of the protected face image by embedding robust watermarking, with which the fake faces can be traced with the reserved watermarking even after the face is swapped. Experimental results demonstrate that our proposed BiFPro could generate the watermarking which is fragile to FOD and at the same time robust to FIF with an average watermark extraction success rate reaching more than 95% when defending against the four advanced DeepFake techniques. Finally, we hope this work can encourage more initiative countermeasures against DeepFake.
Honggu Liu, Wenbo Zhou 0004, Han Fang 0004, Paolo Bestagini, Weiming Zhang 0001, Yuefeng Chen, Stefano Tubaro, Nenghai Yu, Yuan He 0011, Hui Xue 0001
ACM Multimedia8
2023 Audio Splicing Detection and Localization Based on Acquisition Device Traces
abstract
In recent years, the multimedia forensic community has put a great effort in developing solutions to assess the integrity and authenticity of multimedia objects, focusing especially on manipulations applied by means of advanced deep learning techniques. However, in addition to complex forgeries as the deepfakes, very simple yet effective manipulation techniques not involving any use of state-of-the-art editing tools still exist and prove dangerous. This is the case of audio splicing for speech signals, i.e., to concatenate and combine multiple speech segments obtained from different recordings of a person in order to cast a new fake speech. Indeed, by simply adding a few words to an existing speech we can completely alter its meaning. In this work, we address the overlooked problem of detection and localization of audio splicing from different models of acquisition devices. Our goal is to determine whether an audio track under analysis is pristine, or it has been manipulated by splicing one or multiple segments obtained from different device models. Moreover, if a recording is detected as spliced, we identify where the modification has been introduced in the temporal dimension. The proposed method is based on a Convolutional Neural Network (CNN) that extracts model-specific features from the audio recording. After extracting the features, we determine whether there has been a manipulation through a clustering algorithm. Finally, we identify the point where the modification has been introduced through a distance-measuring technique. The proposed method allows to detect and localize multiple splicing points within a recording.
Daniele Ugo Leonzio, Luca Cuccovillo, Paolo Bestagini, Marco Marcon, Patrick Aichroth, Stefano Tubaro
IEEE Trans. Inf. Forensics Secur.6
2022 Panchromatic Imagery Copy-Paste Localization Through Data-Driven Sensor Attribution
abstract
Overhead images can be obtained using different acquisition and processing techniques, and they are becoming more and more popular. As with common photographs, they can be forged and manipulated by malicious users. However, not all image forensics methods tailored to normal photos can be successfully applied out of the box to overhead images. In this paper we consider the problem of localizing copy-paste forgeries on panchromatic images acquired with different satellites. We leverage a set of Convolutional Neural Networks (CNNs) that extract traces of the acquisition satellite directly from image patches. We then determine whether an image region appears to have been acquired with a different satellite than the rest of the picture. Results show that the proposed technique outperforms more sophisticated image forensics tools tailoring common photographs.
Edoardo Daniele Cannas, János Horváth, Sriram Baireddy, Paolo Bestagini, Edward J. Delp, Stefano Tubaro
ICASSP6
2022 Deepfake Speech Detection Through Emotion Recognition: A Semantic Approach
abstract
In recent years, audio and video deepfake technology has advanced relentlessly, severely impacting people’s reputation and reliability. Several factors have facilitated the growing deepfake threat. On the one hand, the hyper-connected society of social and mass media enables the spread of multimedia content worldwide in real-time, facilitating the dissemination of counterfeit material. On the other hand, neural network-based techniques have made deepfakes easier to produce and difficult to detect, showing that the analysis of low-level features is no longer sufficient for the task. This situation makes it crucial to design systems that allow detecting deepfakes at both video and audio levels. In this paper, we propose a new audio spoofing detection system leveraging emotional features. The rationale behind the proposed method is that audio deepfake techniques cannot correctly synthesize natural emotional behavior. Therefore, we feed our deepfake detector with high-level features obtained from a state-of-the-art Speech Emotion Recognition (SER) system. As the used descriptors capture semantic audio information, the proposed system proves robust in cross-dataset scenarios outperforming the considered baseline on multiple datasets.
Emanuele Conti, Davide Salvi, Clara Borrelli, Brian C. Hosler, Paolo Bestagini, Fabio Antonacci, Augusto Sarti, Matthew C. Stamm, Stefano Tubaro
ICASSP9
2022 A Data-Driven Approach for Acoustic Parameter Similarity Estimation of Speech Recording
abstract
Speech audio acquisitions exhibit different quality and reverberation properties depending on the recording setup and environment. For this reason, it is expected that speech analysis systems that work correctly on certain audio recordings may fail on others acquired in different acoustic contexts. Therefore, to be able to tell whether a track under analysis shares the same acoustic characteristics of a reference one may be useful to understand if it can be successfully processed by a given speech analysis system. Alternatively, in a forensic scenario, an estimate of acoustic parameter similarity between two tracks can be used to verify whether the recordings have been likely acquired in the same environment or not. In this work, we propose two methods to estimate acoustic parameter similarity between a speech recording under analysis and a reference one. The first method relies on the estimation of channel-based acoustic indicators that are then compared to extract a similarity measure. The second method directly learns a parameter similarity measure through siamese neural networks.
Mattia Papa, Clara Borrelli, Paolo Bestagini, Fabio Antonacci, Augusto Sarti, Stefano Tubaro
ICASSP6
2022 Forensic Analysis and Localization of Multiply Compressed MP3 Audio Using Transformers
abstract
Audio signals are often stored and transmitted in compressed formats. Among the many available audio compression schemes, MPEG-1 Audio Layer III (MP3) is very popular and widely used. Since MP3 is lossy it leaves characteristic traces in the compressed audio which can be used forensically to expose the past history of an audio file. In this paper, we consider the scenario of audio signal manipulation done by temporal splicing of compressed and uncompressed audio signals. We propose a method to find the temporal location of the splices based on transformer networks. Our method identifies which temporal portions of a audio signal have undergone single or multiple compression at the temporal frame level, which is the smallest temporal unit of MP3 compression. We tested our method on a dataset of 486,743 MP3 audio clips. Our method achieved higher performance and demonstrated robustness with respect to different MP3 data when compared with existing methods.
Ziyue Xiang, Paolo Bestagini, Stefano Tubaro, Edward J. Delp
ICASSP3
2022 Detecting Gan-Generated Images by Orthogonal Training of Multiple CNNs
abstract
In the last few years, we have witnessed the rise of a series of deep learning methods to generate synthetic images that look extremely realistic. These techniques prove useful in the movie industry and for artistic purposes. However, they also prove dangerous if used to spread fake news or to generate fake online accounts. For this reason, detecting if an image is an actual photograph or has been synthetically generated is becoming an urgent necessity. This paper proposes a detector of synthetic images based on an ensemble of Convolutional Neural Networks (CNNs). We consider the problem of detecting images generated with techniques not available at training time. This is a common scenario, given that new image generators are published more and more frequently. To solve this issue, we leverage two main ideas: (i) CNNs should provide "orthogonal" results to better contribute to the ensemble; (ii) the original-image class is better defined than the synthetic-image one, thus it should be better trusted at testing time. Experiments show that pursuing these two ideas improves the detector accuracy on NVIDIA's newly generated StyleGAN3 images, never used in training.
Sara Mandelli, Nicolò Bonettini, Paolo Bestagini, Stefano Tubaro
ICIP4
2022 DIPPAS: a deep image prior PRNU anonymization scheme
abstract
Abstract Source device identification is an important topic in image forensics since it allows to trace back the origin of an image. Its forensics counterpart is source device anonymization, that is, to mask any trace on the image that can be useful for identifying the source device. A typical trace exploited for source device identification is the photo response non-uniformity (PRNU), a noise pattern left by the device on the acquired images. In this paper, we devise a methodology for suppressing such a trace from natural images without a significant impact on image quality. Expressly, we turn PRNU anonymization into the combination of a global optimization problem in a deep image prior (DIP) framework followed by local post-processing operations. In a nutshell, a convolutional neural network (CNN) acts as a generator and iteratively returns several images with attenuated PRNU traces. By exploiting straightforward local post-processing and assembly on these images, we produce a final image that is anonymized with respect to the source PRNU, still maintaining high visual quality. With respect to widely adopted deep learning paradigms, the used CNN is not trained on a set of input-target pairs of images. Instead, it is optimized to reconstruct output images from the original image under analysis itself. This makes the approach particularly suitable in scenarios where large heterogeneous databases are analyzed. Moreover, it prevents any problem due to the lack of generalization. Through numerical examples on publicly available datasets, we prove our methodology to be effective compared to state-of-the-art techniques.
Francesco Picetti, Sara Mandelli, Paolo Bestagini, Vincenzo Lipari, Stefano Tubaro
EURASIP J. Inf. Secur.5
2022 Deep Prior-Based Unsupervised Reconstruction of Irregularly Sampled Seismic Data
abstract
Irregularity and coarse spatial sampling of seismic data strongly affect the performances of processing and imaging algorithms. Therefore, interpolation is a usual preprocessing step in most of the processing workflows. In this work, we propose a seismic data interpolation method based on the deep prior paradigm: anad hocconvolutional neural network is used as a prior to solve the interpolation inverse problem, avoiding any costly and prone-to-overfitting training stage. In particular, the proposed method leverages a multiresolution U-Net with 3-D convolution kernels exploiting correlations in cubes of seismic data, at different scales in all directions. Numerical examples on different corrupted synthetic and field data sets show the effectiveness and promising features of the proposed approach.
Fantong Kong, Francesco Picetti, Vincenzo Lipari, Paolo Bestagini, Xiaoming Tang, Stefano Tubaro
IEEE Geosci. Remote. Sens. Lett.6
2022 Intelligent Seismic Deblending Through Deep Preconditioner
abstract
Seismic deblending is an ill-posed inverse problem that involves counteracting the effect of a blending matrix derived from the shots position and firing time. In this letter, we propose a seismic deblending method based on so-called deep preconditioners. A convolutional Autoencoder (AE) is first trained in a patch-wise fashion to learn an effective sparse representation of the common receiver gathers (CRGs) we aim to reconstruct. Then, the decoder branch of the trained AE is used as a nonlinear preconditioner for the deblending problem. Particularly, to avoid the explicit creation of a training dataset, we suggest to use the common shot gathers (CSGs) of the blended dataset itself to train the AE network, as they are not affected by incoherent blending noise. Numerical examples on synthetic and field datasets demonstrate the effectiveness of the proposed method in comparison with significantly comparable techniques: a dictionary-learning based deblending method; an end-to-end deblending convolution neutral network (CNN).
Weiwei Xu 0004, Vincenzo Lipari, Paolo Bestagini, Matteo Ravasi, Stefano Tubaro
IEEE Geosci. Remote. Sens. Lett.6
2021 Open-Set Source Attribution for Panchromatic Satellite Imagery
abstract
In the last few years, several companies started offering the possibility of buying different kinds of overhead images acquired by satellites orbiting around the planet. This market is interesting for several customers, from those who simply fancy a shot of their house from space, to those aiming to acquire strategic information on portions of land. Due to the sensitive nature of this data, which can be maliciously altered by anyone, the forensic community has started investigating methodologies to verify overhead imagery authenticity and integrity. Within this context, in this paper we investigate the possibility of using Convolutional Neural Networks (CNNs) to attribute a panchromatic satellite image to the satellite used to acquire it. In our investigation we tackle both closed-set and, adapting Deep Ensemble (DE) and Monte Carlo Dropout (MCD) techniques, open-set image attribution problems.
Edoardo Daniele Cannas, Sriram Baireddy, Emily R. Bartusiak, Sri Yarlagadda, Daniel Mas Montserrat, Paolo Bestagini, Stefano Tubaro, Edward J. Delp
ICIP7
2021 Anti-Aliasing Add-On For Deep Prior Seismic Data Interpolation
abstract
Data interpolation is a fundamental step in any seismic processing workflow. Among machine learning techniques recently proposed to solve data interpolation as an inverse problem, Deep Prior paradigm aims at employing a convolutional neural network to capture priors on the data in order to regularize the inversion. However, this technique lacks of reconstruction precision when interpolating highly decimated data due to the presence of aliasing. In this work, we propose to improve Deep Prior inversion by adding a directional Laplacian as regularization term to the problem. This regularizer drives the optimization towards solutions that honor the slopes estimated from the interpolated data low frequencies. We provide some numerical examples to showcase the methodology devised in this manuscript, showing that our results are less prone to aliasing also in presence of noisy and corrupted data.
Francesco Picetti, Vincenzo Lipari, Paolo Bestagini, Stefano Tubaro
ICIP4
2021 Synthetic speech detection through short-term and long-term prediction traces
abstract
Abstract Several methods for synthetic audio speech generation have been developed in the literature through the years. With the great technological advances brought by deep learning, many novel synthetic speech techniques achieving incredible realistic results have been recently proposed. As these methods generate convincing fake human voices, they can be used in a malicious way to negatively impact on today’s society (e.g., people impersonation, fake news spreading, opinion formation). For this reason, the ability of detecting whether a speech recording is synthetic or pristine is becoming an urgent necessity. In this work, we develop a synthetic speech detector. This takes as input an audio recording, extracts a series of hand-crafted features motivated by the speech-processing literature, and classify them in either closed-set or open-set. The proposed detector is validated on a publicly available dataset consisting of 17 synthetic speech generation algorithms ranging from old fashioned vocoders to modern deep learning solutions. Results show that the proposed method outperforms recently proposed detectors in the forensics literature.
Clara Borrelli, Paolo Bestagini, Fabio Antonacci, Augusto Sarti, Stefano Tubaro
EURASIP J. Inf. Secur.5
2021 Reconstructing Speech From CNN Embeddings
abstract
The complete understanding of the decision-making process of Convolutional Neural Networks (CNNs) is far from being fully reached. Many researchers proposed techniques to interpret what a network actually “learns” from data. Nevertheless many questions still remain unanswered. In this work we study one aspect of this problem by reconstructing speech from the intermediate embeddings computed by a CNNs. Specifically, we consider a pre-trained network that acts as a feature extractor from speech audio. We investigate the possibility of inverting these features, reconstructing the input signals in a black-box scenario, and quantitatively measure the reconstruction quality by measuring the word-error-rate of an off-the-shelf ASR model. Experiments performed using two different CNN architectures trained for six different classification tasks, show that it is possible to reconstruct time-domain speech signals that preserve the semantic content, whenever the embeddings are extracted before the fully connected layers.
Luca Comanducci, Paolo Bestagini, Marco Tagliasacchi, Augusto Sarti, Stefano Tubaro
IEEE Signal Process. Lett.5
2021 Landmine Detection Using Autoencoders on Multipolarization GPR Volumetric Data
abstract
Buried landmines and unexploded remnants of war are a constant threat for the population of many countries that have been hit by wars in the past years. The huge amount of casualties has been a strong motivation for the research community toward the development of safe and robust techniques designed for landmine clearance. Nonetheless, being able to detect and localize buried landmines with high precision in an automatic fashion is still considered a challenging task due to the many different boundary conditions that characterize this problem (e.g., several kinds of objects to detect, different soils and meteorological conditions, etc.). In this article, we propose a novel technique for buried object detection tailored to unexploded landmine discovery. The proposed solution exploits a specific kind of convolutional neural network (CNN) known as autoencoder to analyze volumetric data acquired with ground penetrating radar (GPR) using different polarizations. This method works in an anomaly detection framework, indeed we only train the autoencoder on GPR data acquired on landmine-free areas. The system then recognizes landmines as objects that are dissimilar to the soil used during the training step. Experiments conducted on real data show that the proposed technique requires little training and no ad hoc data preprocessing to achieve accuracy higher than 93% on challenging data sets.
Paolo Bestagini, Federico Lombardi, Maurizio Lualdi, Francesco Picetti, Stefano Tubaro
IEEE Trans. Geosci. Remote. Sens.5
2020 A Modified Fourier-Mellin Approach For Source Device Identification On Stabilized Videos
abstract
To decide whether a digital video has been captured by a given device, multimedia forensic tools usually exploit characteristic noise traces left by the camera sensor on the acquired frames. This analysis requires that the noise pattern characterizing the camera and the noise pattern extracted from video frames under analysis are geometrically aligned. However, in many practical scenarios this does not occur, thus a re-alignment or synchronization has to be performed. Current solutions often require time consuming search of the realignment transformation parameters. In this paper, we propose to overcome this limitation by searching scaling and rotation parameters in the frequency domain. The proposed algorithm tested on real videos from a well-known state-of-the-art dataset shows promising results.
Sara Mandelli, Fabrizio Argenti, Paolo Bestagini, Massimo Iuliani, Alessandro Piva, Stefano Tubaro
ICIP6
2020 On the use of Benford's law to detect GAN-generated images
abstract
The advent of Generative Adversarial Network (GAN) architectures has given anyone the ability of generating incredibly realistic synthetic imagery. The malicious diffusion of GAN-generated images may lead to serious social and political consequences (e.g., fake news spreading, opinion formation, etc.). It is therefore important to regulate the widespread distribution of synthetic imagery by developing solutions able to detect them. In this paper, we study the possibility of using Benford's law to discriminate GAN-generated images from natural photographs. Benford's law describes the distribution of the most significant digit for quantized Discrete Cosine Transform (DCT) coefficients. Extending and generalizing this property, we show that it is possible to extract a compact feature vector from an image. This feature vector can be fed to an extremely simple classifier for GAN-generated image detection purpose.
Nicolò Bonettini, Paolo Bestagini, Simone Milani, Stefano Tubaro
ICPR4
2020 Video Face Manipulation Detection Through Ensemble of CNNs
abstract
In the last few years, several techniques for facial manipulation in videos have been successfully developed and made available to the masses (i.e., FaceSwap, deepfake, etc.). These methods enable anyone to easily edit faces in video sequences with incredibly realistic results and a very little effort. Despite the usefulness of these tools in many fields, if used maliciously, they can have a significantly bad impact on society (e.g., fake news spreading, cyber bullying through fake revenge porn). The ability of objectively detecting whether a face has been manipulated in a video sequence is then a task of utmost importance. In this paper, we tackle the problem of face manipulation detection in video sequences targeting modern facial manipulation techniques. In particular, we study the ensembling of different trained Convolutional Neural Network (CNN) models. In the proposed solution, different models are obtained starting from a base network (i.e., EfficientNetB4) making use of two different concepts: (i) attention layers; (ii) siamese training. We show that combining these networks leads to promising face manipulation detection results on two publicly available datasets with more than 119000 videos.
Nicolò Bonettini, Edoardo Daniele Cannas, Sara Mandelli, Luca Bondi, Paolo Bestagini, Stefano Tubaro
ICPR6
2020 Deep skin detection on low resolution grayscale images
Marco Paracchini, Marco Marcon, Federica A. Villa, Stefano Tubaro
Pattern Recognit. Lett.4
2020 CNN-Based Fast Source Device Identification
abstract
Source identification is an important topic in image forensics, since it allows to trace back the origin of an image. This represents a precious information to claim intellectual property but also to reveal the authors of illicit materials. In this letter we address the problem of device identification based on sensor noise and propose a fast and accurate solution using convolutional neural networks (CNNs). Specifically, we propose a 2-channel-based CNN that learns a way of comparing camera fingerprint and image noise at patch level. The proposed solution turns out to be much faster than the conventional approach and to ensure an increased accuracy. This makes the approach particularly suitable in scenarios where large databases of images are analyzed, like over social networks. In this vein, since images uploaded on social media usually undergo at least two compression stages, we include investigations on double JPEG compressed images, always reporting higher accuracy than standard approaches.
Sara Mandelli, Davide Cozzolino, Paolo Bestagini, Luisa Verdoliva, Stefano Tubaro
IEEE Signal Process. Lett.5
2020 A Methodology for the Robust Estimation of the Radiation Pattern of Acoustic Sources
abstract
We propose a novel methodology for estimating the radiation pattern of acoustic sources, which is general enough as to be suitable for a wide variety of sources without the need of anechoic conditions of operation. Multiple plenacoustic cameras (which can be thought of as arrays of acoustic cameras) scan the source while keeping reflections and interferers at bay through deconvolution and windowing of the measured response. In the case of a moving source (e.g. a musical instrument while it is being played), the plenacoustic cameras are also used for tracking the position of the source. As for its orientation, we propose practical solutions for tracking that as well, whenever such information is not known in advance. Two experiments are conducted in order to validate the proposed solution. The former focuses on a commercial loudspeaker cabinet, whose radiation pattern is known in advance and can be used as groundtruth. The latter concerns violins, which exhibit an extremely rich and hard to predict acoustic behavior, due to their inherent structural and constructional complexity. Our method allows us to capture the radiation pattern of the instrument while it is being played, thus returning data corresponding to the natural timbre of the instrument, including the unavoidable acoustic shadow of the violinist's head. Experimental results confirm a relevant improvement in accuracy and robustness afforded by the adoption of dynamic plenacoustic solutions with respect to state-of-the-art techniques.
Antonio Canclini, Fabio Antonacci, Stefano Tubaro, Augusto Sarti
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Source Localization Using Distributed Microphones in Reverberant Environments Based on Deep Learning and Ray Space Transform
abstract
In this article we present a methodology for source localization in reverberant environments from Generalized Cross Correlations (GCCs) computed between spatially distributed individual microphones. Reverberation tends to negatively affect localization based on Time Differences of Arrival (TDOAs), which become inaccurate due to the presence of spurious peaks in the GCC. We therefore adopt a data-driven approach based on a convolutional neural network, which, using the GCCs as input, estimates the source location in two steps. It first computes the Ray Space Transform (RST) from multiple arrays. The RST is a convenient representation of the acoustic rays impinging on the array in a parametric space, called Ray Space. Rays produced by a source are visualized in the RST as patterns, whose position is uniquely related to the source location. The second step consists of estimating the source location through a nonlinear fitting, which estimates the coordinates that best approximate the RST pattern obtained through the first step. It is worth noting that training can be accomplished on simulated data only, thus relaxing the need of actually deploying microphone arrays in the acoustic scene. The localization accuracy of the proposed techniques is similar to the one of SRP-PHAT, however our method demonstrates an increased robustness regarding different distributed microphones configurations. Moreover, the use of the RST as an intermediate representation makes it possible for the network to generalize to data unseen during training.
Luca Comanducci, Federico Borra, Paolo Bestagini, Fabio Antonacci, Stefano Tubaro, Augusto Sarti
IEEE ACM Trans. Audio Speech Lang. Process.5
2020 A Parametric Approach to Virtual Miking for Sources of Arbitrary Directivity
abstract
In this article we propose a methodology for the reconstruction of sound fields in arbitrary locations based on the signals acquired by a spatial distribution of compact microphone arrays (virtual miking). The proposed method is suitable for operating in reverberant environments, thanks to a two-stage analysis process, the former of which aims at separating the direct and the diffuse components of the sound field. The method that we propose is inherently parametric, as the sources of the acoustic scene are characterized by parameters describing location and directivity (spherical harmonics expansion), which are extracted from the exterior model of the direct component of the sound field. Once the parameters of the sources are extracted, the direct sound field at an arbitrary location is reconstructed. The diffuse component is reconstructed from the joint knowledge of the diffuse component at the locations of the distributed microphone arrays, under the assumption of isotropic behavior. Results show that the proposed technique is able to analyze the sound field and reconstruct the parameters of the sources that are active in the scene. In addition, the synthesis of the signals at the virtual microphone locations turns out to accurately match (in terms of spatial cues) the actual sound field, as measured by a microphone places in the desired location.
Mirco Pezzoli, Federico Borra, Fabio Antonacci, Stefano Tubaro, Augusto Sarti
IEEE ACM Trans. Audio Speech Lang. Process.4
2020 Facing Device Attribution Problem for Stabilized Video Sequences
abstract
A problem deeply investigated by multimedia forensics researchers is that of detecting which device has been used to capture a video. This enables us to trace down the owner of a video sequence, which proves extremely helpful to solve copyright infringement cases as well as to fight distribution of illicit material (e.g., child exploitation clips and terroristic threats). Currently, the most promising methods to tackle this task exploit unique noise traces left by camera sensors on acquired images. However, given the recent advancements in motion stabilization of video content, robustness of sensor pattern noise-based techniques is strongly hindered. Indeed, video stabilization introduces geometric transformations to video frames, thus making camera fingerprint estimation problematic with classical approaches. In this paper, we deal with the challenging problem of attributing stabilized videos to their recording device. Specifically, we propose: 1) a strategy to extract the characteristic fingerprint of a device, starting from either a set of images or stabilized video sequences and 2) a strategy to match a stabilized video sequence with a given fingerprint. The proposed methodology is tested on videos coming from a set of different smartphones, taken from the modern publicly available Vision Dataset. The conducted experiments also provide an interesting insight on the effect of modern smartphones video stabilization algorithms on specific video frames.
Sara Mandelli, Paolo Bestagini, Luisa Verdoliva, Stefano Tubaro
IEEE Trans. Inf. Forensics Secur.4
2019 "Hello? Who Am I Talking to?" A Shallow CNN Approach for Human vs. Bot Speech Classification
abstract
Automatic speech generation algorithms, enhanced by deep learning techniques, enable an increasingly seamless and immediate machine-to-human interaction. As a result, the latest generation of phone-calling bots sounds more convincingly human than previous generations. The application of this technology has a strong social impact in terms of privacy issues (e.g., in customer-care services), fraudulent actions (e.g., social hacking) and erosion of trust (e.g., generation of fake conversation). For these reasons, it is crucial to identify the nature of a speaker, as either a human or a bot. In this paper, we propose a speech classification algorithm based on Convolutional Neural Networks (CNNs), which enables the automatic classification of human vs non-human speakers from the analysis of short audio excerpts. We evaluate the effectiveness of the proposed solution by exploiting a real human speech database populated with audio recordings from various sources, and automatically generated speeches using state-of-the-art text-to-speech generators based on deep learning (e.g., Google WaveNet).
Alessandro Lieto, Daniele Moro, Francesco Devoti, Claudia Parera, Vincenzo Lipari, Paolo Bestagini, Stefano Tubaro
ICASSP7
2019 Shadow Removal Detection and Localization for Forensics Analysis
abstract
The recent advancements in image processing and computer vision allow realistic photo manipulations. In order to avoid the distribution of fake imagery, the image forensics community is working towards the development of image authenticity verification tools. Methods based on shadow analysis are particularly reliable since they are part of the physical integrity of the scene, thus detecting forgeries is possible whenever inconsistencies are found (e.g., shadows not coherent with the light direction). An attacker can easily delete inconsistent shadows and replace them with correctly cast shadows in order to fool forensics detectors based on physical analysis. In this paper, we propose a method to detect shadow removal done with state-of-the-art tools. The proposed method is based on a conditional generative adversarial network (cGAN) specifically trained for shadow removal detection.
Sri Yarlagadda, David Guera, Daniel Mas Montserrat, Fengqing Zhu 0001, Edward J. Delp, Paolo Bestagini, Stefano Tubaro
ICASSP7
2019 Image Anonymization Detection with Deep Handcrafted Features
abstract
In recent years, the number of images shared online has continuously grown. The forensics community has kept the pace by developing techniques to both reliably extract information from these images, but also to remove it. In particular, the latest developments in image anonymization methods exposes an attack vector when used by skilled ill-intentioned image producers that may want to elude prosecution. We present an approach to detect whether or not an image has undergone a laundering process, i.e., it has been tampered with so that its unique characterizing features have been changed to avoid detection. We focus on the photo response non uniformity (PRNU) noise unique to every imaging sensor, and we consider that an image has been "laundered" when we detect the absence of PRNU from an image. We propose a per image preprocessing pipeline that generates information-rich features later used as input of fine-tuned convolutional neural networks (CNNs). We study the performance of the proposed approach using various CNN architectures and blind anonymization techniques and show its effectiveness under several training and testing scenarios. Our results also show that CNN models trained with the proposed feature are capable of generalizing over unseen devices and are robust against non-geometric transformations.
Nicolò Bonettini, David Guera, Luca Bondi, Paolo Bestagini, Edward J. Delp, Stefano Tubaro
ICIP6
2019 A Prnu-Based Method to Expose Video Device Compositions in Open-Set Setups
abstract
As the diffusion of altered video sequences over social networks and the internet can have severe consequences (e.g., fake news spreading, false accusations, etc.), the forensic research community has started actively working toward the development of methodologies tailored to assess the authenticity and integrity of video sequences. In this work, we focus on the problem of spotting video compilations composed by temporally concatenating sequences acquired with different devices. To solve this problem, we leverage trace characteristics of each video recording device left on each video at acquisition time. Specifically, we propose a method leveraging Photo-Response Non-Uniformity (PRNU) traces along with a binary classifier in order to understand which frames of a video have been acquired with the same device used to record the first few frames. Results show that the proposed solution outperforms baselines based on standard PRNU correlation and thresholding tests. Experiments have been carried out in an open-set scenario and show promising results.
Pedro Ribeiro Mendes Júnior, Luca Bondi, Paolo Bestagini, Anderson Rocha 0001, Stefano Tubaro
ICIP5
2019 Detection and Synchronization of Video Sequences for Event Reconstruction
abstract
With an ever-growing amount of unexpected menaces in crowded places such as terrorist attacks, it is paramount to develop techniques to aid investigators reconstructing all details about an event of interest. To extract reliable information about the event, all kinds of available clues must be jointly exploited. As a matter of fact, today's sources of information are plenty and varied, as important events affecting many people are typically documented by different sources. Both witnesses' smartphones and security cameras can provide valuable information coming from multiple viewpoints and time instants - "the eyes of the crowd". In this paper, we focus on the specific problem of automatically detecting and temporally synchronizing videos depicting the same event of interest. Videos can be either near-duplicates (i.e., edited copies of the same original source) or sequences shot by different users from different vantage points. The proposed method relies upon a video fingerprinting technique capable of describing how video semantic content evolves in time. The solution does not assume a priori information about cameras location, and it only exploits visual cues, not relying on audio channels.
Giuliano Pinheiro, Marcos V. M. Cirne, Paolo Bestagini, Stefano Tubaro, Anderson Rocha 0001
ICIP4
2019 View-synthesis from uncalibrated cameras and parallel planes
Antonio Canclini, Francesco Malapelle, Marco Marcon, Stefano Tubaro, Andrea Fusiello
Signal Process. Image Commun.4
2019 Improving PRNU Compression Through Preprocessing, Quantization, and Coding
abstract
In the last decade, the extremely rapid proliferation of digital devices capable of acquiring and sharing images over the Web has significantly increased the amount of digital images publicly accessible by everyone with Internet access. Despite the obvious benefits of such technological improvements, it is becoming mandatory to verify the origin and trustfulness of such shared pictures. Photo response non-uniformity (PRNU) is the reference signal for forensic investigators when it comes to verifying or identifying which camera device shot a picture under analysis. In spite of this, PRNU is almost a white-shaped noise, thus being very difficult to compress for storage or large scale search purposes, which are frequent investigation scenarios. To overcome the issue, the forensic community has developed a series of compression algorithms. Lately, Gaussian random projections have proved to achieve state-of-the-art performance. In this paper, we propose two additional steps that help improving even more Gaussian random projections compression rate: 1) a decimation preprocessing step tailored at attenuating frequency components in which PRNU traces are already suppressed in JPEG compressed images and 2) a dead-zone quantizer (rather than the commonly used binary one) that enables an entropy coding scheme to save bitrate when storing PRNU fingerprints or sending residuals over a communication channel. Reported results show the effectiveness of proposed improvements, both under controlled JPEG compression and in a real case scenario.
Luca Bondi, Paolo Bestagini, Fernando Pérez-González, Stefano Tubaro
IEEE Trans. Inf. Forensics Secur.4
2018 Multiple Jpeg Compression Detection Through Task-Driven Non-Negative Matrix Factorization
abstract
Due to the increasingly unbridled practice of sharing visual content on the web, tracing back past history of uploaded images is getting far from being an easy task. Nonetheless, forensic analysts might be interested in probing digital history of content published on the web to assess its authenticity. In this vein, a possible indicator of image integrity is the number of JPEG compressions a picture underwent. As a matter of fact, JPEG compression is typically operated first at image inception time directly on the acquisition device. Then, it is customary re-applied every time an image is manipulated or shared through social media. For this reason, the more the applied JPEG compressions, the more the likelihood that an image underwent some editing. In this work, we propose an algorithm to detect multiple JPEG compressions, specifically up to four coding cycles. This approach leverages the Task-driven Non-negative Matrix Factorization (TNMF) model, fed with histograms of the Discrete Cosine Transform (DCT) of the image under analysis. Experimental results show the effectiveness of the method if compared with the state-of-the-art, confirming this strategy as a viable solution for detecting multiple JPEG compressions.
Sara Mandelli, Nicolò Bonettini, Paolo Bestagini, Vincenzo Lipari, Stefano Tubaro
ICASSP5
2018 Estimation of the Sound Field at Arbitrary Positions in Distributed Microphone Networks Based on Distributed Ray Space Transform
abstract
In this paper we propose a parametric sound field reconstruction approach. In particular, the technique is based on the estimation of three parameters for each acoustic source (source position, radiation pattern and source signal) given the signals acquired by few arbitrarily placed microphone arrays. This allows us to synthesize the signal of a virtual microphone placed in any point of the acoustic scene.
Mirco Pezzoli, Federico Borra, Fabio Antonacci, Augusto Sarti, Stefano Tubaro
ICASSP5
2018 Video Codec Forensics Based on Convolutional Neural Networks
abstract
The recent development of multimedia has made video editing accessible to everyone. Unfortunately, forensic analysis tools capable of detecting traces left by video processing operations in a blind fashion are still at their beginnings. One of the reasons is that videos are customary stored and distributed in a compressed format, and codec-related traces tends to mask previous processing operations. In this paper, we propose to capture video codec traces through convolutional neural networks (CNNs) and exploit them as an asset. Specifically, we train two CNN s to extract information about the used video codec and coding quality, respectively. Building upon these CNN s, we propose a system to detect and localize temporal splicing for video sequences generated from the concatenation of different video segments, which are characterized by inconsistent coding schemes and/or parameters (e.g., video compilations from different sources or broadcasting channels). The proposed solution is validated using videos at different resolutions (i.e., CIF, 4CIF, PAL and 720p) encoded with four common codecs (i.e., MPEG2, MPEG4, H264 and H265) at different qualities (i.e., different constant and variable bitrates, as well as constant quantization parameters).
Sebastiano Verde, Luca Bondi, Paolo Bestagini, Simone Milani, Giancarlo Calvagno, Stefano Tubaro
ICIP6
2018 Reliability Map Estimation for CNN-Based Camera Model Attribution
abstract
Among the image forensic issues investigated in the last few years, great attention has been devoted to blind camera model attribution. This refers to the problem of detecting which camera model has been used to acquire an image by only exploiting pixel information. Solving this problem has great impact on image integrity assessment as well as on authenticity verification. Recent advancements that use convolutional neural networks (CNNs) in the media forensic field have enabled camera model attribution methods to work well even on small image patches. These improvements are also important for determining forgery localization. Some patches of an image may not contain enough information related to the camera model (e.g., saturated patches). In this paper, we propose a CNN-based solution to estimate the camera model attribution reliability of a given image patch. We show that we can estimate a reliabilitymap indicating which portions of the image contain reliable camera traces. Testing using a well known dataset confirms that by using this information, it is possible to increase small patch camera model attribution accuracy by more than 8% on a single patch.
David Guera, Fengqing Zhu 0001, Sri Yarlagadda, Stefano Tubaro, Paolo Bestagini, Edward J. Delp
WACV4
2017 Inpainting-Based camera anonymization
abstract
Over the years, the forensic community has developed a series of very accurate camera attribution algorithms enabling to detect which device has been used to acquire an image with outstanding results. Many of these methods are based on photo response non uniformity (PRNU) that allows tracing back a picture to the camera used to shoot it. However, when privacy is required, it would be desirable to anonymize photos, unlinking them from their specific device. This paper investigates a new and alternative approach to image anonymization task. The proposed method leverages image inpainting described as an inverse regularized problem, and does not need any priors about the PRNU to remove. Specifically, we show how PRNU pattern can be strongly attenuated by reconstructing each pixel of an image from its neighbors, only slightly affecting visual quality. Results confirm this approach as a viable alternative solution for image anonymization.
Sara Mandelli, Luca Bondi, Silvia Lameri, Vincenzo Lipari, Paolo Bestagini, Stefano Tubaro
ICIP6
2017 Multicamera rig calibration by double-sided thick checkerboard
abstract
A multi‐camera rig calibration algorithm based on a double sided planar target is proposed. Due to their inherently simple realisation, low cost and accuracy, planar calibration targets came out as one of the most largely adopted calibration tools both for intrinsic and extrinsic camera parameters. However, concerning the estimation of extrinsic parameters, one of the major drawbacks of these targets is their requirement for distinct target visibility from both cameras. This prevents many configurations from being adopted where, e.g. two cameras are facing each other. An inexpensive solution could be based on printing/pasting a planar pattern on both target sides, however, the relative misalignment between the patterns on the two sides and the target thickness could be unknown. The authors propose a solution where double‐sided target displacement error is estimated together with the extrinsic parameters allowing the reuse of all the available planar calibration tools in less constrained configurations. To assess their approach the authors tested the system in two scenarios, one using two professional 4K cameras and one using two smartphones.
Marco Marcon, Augusto Sarti, Stefano Tubaro
IET Comput. Vis.3
2017 Aligned and non-aligned double JPEG detection using convolutional neural networks
Mauro Barni, Luca Bondi, Nicolò Bonettini, Paolo Bestagini, Andrea Costanzo, Marco Maggini, Benedetta Tondi, Stefano Tubaro
J. Vis. Commun. Image Represent.8
2017 First Steps Toward Camera Model Identification With Convolutional Neural Networks
abstract
Detecting the camera model used to shoot a picture enables to solve a wide series of forensic problems, from copyright infringement to ownership attribution. For this reason, the forensic community has developed a set of camera model identification algorithms that exploit characteristic traces left on acquired images by the processing pipelines specific of each camera model. In this letter, we investigate a novel approach to solve camera model identification problem. Specifically, we propose a data-driven algorithm based on convolutional neural networks, which learns features characterizing each camera model directly from the acquired pictures. Results on a well-known dataset of 18 camera models show that: 1) the proposed method outperforms up-to-date state-of-the-art algorithms on classification of 64 × 64 color image patches; 2) features learned by the proposed network generalize to camera models never used for training.
Luca Bondi, Luca Baroffio, David Guera, Paolo Bestagini, Edward J. Delp, Stefano Tubaro
IEEE Signal Process. Lett.6
2017 Data-Driven Feature Characterization Techniques for Laser Printer Attribution
abstract
Laser printer attribution is an increasing problem with several applications, such as pointing out the ownership of crime proofs and authentication of printed documents. However, as commonly proposed methods for this task are based on custom-tailored features, they are limited by modeling assumptions about printing artifacts. In this paper, we explore solutions able to learn discriminant-printing patterns directly from the available data during an investigation, without any further feature engineering, proposing the first approach based on deep learning to laser printer attribution. This allows us to avoid any prior assumption about printing artifacts that characterize each printer, thus highlighting almost invisible and difficult printer footprints generated during the printing process. The proposed approach merges, in a synergistic fashion, convolutional neural networks (CNNs) applied on multiple representations of multiple data. Multiple representations, generated through different pre-processing operations, enable the use of the small and lightweight CNNs whilst the use of multiple data enable the use of aggregation procedures to better determine the provenance of a document. Experimental results show that the proposed method is robust to noisy data and outperforms existing counterparts in the literature for this problem.
Anselmo Ferreira, Luca Bondi, Luca Baroffio, Paolo Bestagini, Jiwu Huang, Jefersson A. dos Santos, Stefano Tubaro, Anderson Rocha 0001
IEEE Trans. Inf. Forensics Secur.7
2017 Distributed 3D Source Localization from 2D DOA Measurements Using Multiple Linear Arrays
abstract
This manuscript addresses the problem of 3D source localization from direction of arrivals (DOAs) in wireless acoustic sensor networks. In this context, multiple sensors measure the DOA of the source, and a central node combines the measurements to yield the source location estimate. Traditional approaches require 3D DOA measurements; that is, each sensor estimates the azimuth and elevation of the source by means of a microphone array, typically in a planar or spherical configuration. The proposed methodology aims at reducing the hardware and computational costs by combining measurements related to 2D DOAs estimated from linear arrays arbitrarily displaced in the 3D space. Each sensor measures the DOA in the plane containing the array and the source. Measurements are then translated into an equivalent planar geometry, in which a set of coplanar equivalent arrays observe the source preserving the original DOAs. This formulation is exploited to define a cost function, whose minimization leads to the source location estimation. An extensive simulation campaign validates the proposed approach and compares its accuracy with state-of-the-art methodologies.
Antonio Canclini, Fabio Antonacci, Augusto Sarti, Stefano Tubaro
Wirel. Commun. Mob. Comput.4
2016 Fast keypoint detection in video sequences
abstract
Several computer vision tasks exploit a succinct representation of the visual content in the form of sets of local features. Given an input image, feature extraction algorithms identify keypoints and assign to each of them a descriptor, based on the characteristics of the surrounding visual content. Several tasks might require local features to be extracted from a video sequence, on a frame-by-frame basis. Although temporal downsampling has been proven to be an effective solution for mobile augmented reality and visual search, high temporal resolution is a key requirement for time-critical applications such as object tracking, event recognition, pedestrian detection, surveillance. In recent years, more and more computationally efficient visual feature detectors and descriptors have been proposed. Nonetheless, such approaches are tailored to still images. In this paper we propose a fast keypoint detection algorithm for video sequences, that exploits the temporal coherence of the sequence of keypoints. According to the proposed method, each frame is preprocessed so as to identify the parts of the input frame for which keypoint detection and description need to be performed. Our experiments show that it is possible to achieve a reduction in computational time of up to 40%, without significantly affecting the task accuracy.
Luca Baroffio, Matteo Cesana, Alessandro Redondi, Marco Tagliasacchi, Stefano Tubaro
ICASSP5
2016 Image phylogeny tree reconstruction based on region selection
abstract
Nowadays, everyone can download, edit and republish any picture on the web, thus contributing to the diffusion of near-duplicate (ND) images. In order to gain an interesting insight on the way NDs are distributed online, recent works have focused on the reconstruction of the image phylogeny tree (IPT), i.e., an acyclic graph describing the genealogical relationship between ND image pairs. IPT reconstruction methods typically leverage the possibility of reconstructing one image from another one only if they are in parent-child relationship. However, as estimating the possible parent-child transformation is computationally expensive, usually a limited set of global editing operations is considered (i.e., compression, geometric and colour transformations applied to the whole image). However, in a real-world scenario it is customary to edit images also using local operations (e.g., logo insertion, object removal, splicing, etc.), which hinder the possibility of correctly estimating the parent-child relationship. In this paper, we propose an algorithm for IPT reconstruction that deals with the presence of local editing operations.
Paolo Bestagini, Marco Tagliasacchi, Stefano Tubaro
ICASSP3
2016 A linear operator for the computation of soundfield maps
abstract
In the process of soundfield imaging, as defined in the literature, a microphone array is subdivided into overlapping sub-arrays and soundfield images are obtained by juxtaposition of spatial spectra computed from individual subarray data. In this paper we show that the whole process can be conveniently seen as a linear transformation applied to array data. This linear transformation embeds a nonlinear mapping to cast the directional information in a more convenient domain: the ray space. We show by simulations that the proposed formulation is suitable for fast implementation of the soundfield imaging operation, and, more specifically, for the localization of acoustic sources.
Lucio Bianchi, V. Baldini Anastasio, Dejan Markovic, Fabio Antonacci, Augusto Sarti, Stefano Tubaro
ICASSP6
2016 A low-cost solution to 3D pinna modeling for HRTF prediction
abstract
We propose an infrared (IR) stereo-vision system for estimating the 3D model of the pinna, based on low-cost devices. A commercial IR calibrated stereo camera is used in conjunction with a structured IR light projector, to acquire highly textured snapshots of the pinna. A point cloud is computed for each snapshot by triangulating the stereo correspondences detected in the acquired IR images. A complete 3D model is computed by aligning and merging the point clouds, and then creating a polygonal mesh surface. The nominal accuracy of the proposed system turns to be about 1 mm, which enables an accurate prediction of the Head Related Transfer Function (HRTF) through numerical acoustic simulation.
Luca Bonacina, Antonio Canclini, Fabio Antonacci, Marco Marcon, Augusto Sarti, Stefano Tubaro
ICASSP6
2016 Phylogenetic analysis of near-duplicate images using processing age metrics
abstract
Recent researches on image forensics have led to the design of algorithms to study the phylogenetic relationship between near-duplicate (ND) images. The proposed solutions aim at reconstructing the image phylogeny tree (IPT), and they have immediate applications in security, law and copyright enforcement, and news tracking services. Anyway, the effectiveness of such strategies strictly depends on the accuracy in characterizing image similarities. In this paper, we show that it is possible to take into account additional information to better reconstruct the IPT. More specifically, we propose a set of features that blindly model the processing age of an image, i.e., how much an image has been edited in its lifetime. By exploiting these features, it is possible to improve the performance of IPT reconstruction by increasing the accuracy and reducing the computational complexity.
Simone Milani, Marco Fontana, Paolo Bestagini, Stefano Tubaro
ICASSP4
2016 Toothbrush motion analysis to help children learn proper tooth brushing
Marco Marcon, Augusto Sarti, Stefano Tubaro
Comput. Vis. Image Underst.3
2016 Deep Convolutional Neural Networks for pedestrian detection
Denis Tomè, Federico Monti, Luca Baroffio, Luca Bondi, Marco Tagliasacchi, Stefano Tubaro
Signal Process. Image Commun.6
2016 Extraction of Acoustic Sources Through the Processing of Sound Field Maps in the Ray Space
abstract
Our goal is to develop a model-based approach to acoustic source extraction from microphone array data, which is suitable for both near-field and far-field sources. A signal representation based on plane-wave (PW) decomposition is suitable for acoustic sources in the far field as the resulting spectrum turns out to be impulsive. When the source approaches the array, however, the curvature of the wavefront causes the spectrum of the PW components to depart from impulsive behavior, thus making source extraction harder to attain. In this paper, we adopt a sound field representation based on the local estimation of the plenacoustic function along the array line. This approach consists of dividing the array into subarrays, and applying the PW analysis on individual subarrays. This has the immediate result of extending the range of validity of the far-field hypothesis, as a source that enters the near-field range of the extended array is still in the far-field range of the subarrays. PW analysis on subarrays allows us to construct the so-called sound field map in a domain of acoustic visibility called ray space. The extraction of the desired source is accomplished through spatial filtering of the sound field map. The design of the spatial filter relies on a linear minimum mean square error criterion defined on the sound field map. The effectiveness of the proposed methodology is proven through an extensive simulation campaign as well as real experiments.
Dejan Markovic, Fabio Antonacci, Lucio Bianchi, Stefano Tubaro, Augusto Sarti
IEEE ACM Trans. Audio Speech Lang. Process.4
2016 Codec and GOP Identification in Double Compressed Videos
abstract
Video content is routinely acquired and distributed in a digital compressed format. In many cases, the same video content is encoded multiple times. This is the typical scenario that arises when a video, originally encoded directly by the acquisition device, is then re-encoded, either after an editing operation, or when uploaded to a sharing website. The analysis of the bitstream reveals details of the last compression step (i.e., the codec adopted and the corresponding encoding parameters), while masking the previous compression history. Therefore, in this paper, we consider a processing chain of two coding steps, and we propose a method that exploits coding-based footprints to identify both the codec and the size of the group of pictures (GOPs) used in the first coding step. This sort of analysis is useful in video forensics, when the analyst is interested in determining the characteristics of the originating source device, and in video quality assessment, since quality is determined by the whole compression history. The proposed method relies on the fact that lossy coding is an (almost) idempotent operation. That is, re-encoding a video sequence with the same codec and coding parameters produces a sequence that is similar to the former. As a consequence, if the second codec in the chain does not significantly alter the sequence, it is possible to analyze this sort of similarity to identify the first codec and the adopted GOP size. The method was extensively validated on a very large data set of video sequences generated by encoding content with a diversity of codecs (MPEG-2, MPEG-4, H.264/AVC, and DIRAC) and different encoding parameters. In addition, a proof of concept showing that the proposed method can also be used on videos downloaded from YouTube is reported.
Paolo Bestagini, Simone Milani, Marco Tagliasacchi, Stefano Tubaro
IEEE Trans. Image Process.4
2016 Identification of Transform Coding Chains
abstract
Transform coding is routinely used for lossy compression of discrete sources with memory. The input signal is divided into N-dimensional vectors, which are transformed by means of a linear mapping. Then, transform coefficients are quantized and entropy coded. In this paper, we consider the problem of identifying the transform matrix as well as the quantization step sizes. First, we study the case in which the only available information is a set of P transform decoded vectors. We formulate the problem in terms of finding the lattice with the largest determinant that contains all observed vectors. We propose an algorithm that is able to find the optimal solution and we formally study its convergence properties. Three potential realms of application are considered as example scenarios for the proposed theory: 1) parameter retrieval in the presence of a chain of two transform coders; 2) image tampering identification; and 3) parameter estimation for predictive coders. We show that, despite their differences, all three scenarios can be tackled by applying the same fundamental methodology. Experiments on both the synthetic data and the real images validate the proposed approach.
Marco Tagliasacchi, Marco Visentini Scarzanella, Pier Luigi Dragotti, Stefano Tubaro
IEEE Trans. Image Process.4
2016 3D Beam Tracing Based on Visibility Lookup for Interactive Acoustic Modeling
abstract
We present a method for accelerating the computation of specular reflections in complex 3D enclosures, based on acoustic beam tracing. Our method constructs the beam tree on the fly through an iterative lookup process of a precomputed data structure that collects the information on the exact mutual visibility among all reflectors in the environment (region-to-region visibility). This information is encoded in the form of visibility regions that are conveniently represented in the space of acoustic rays using the Plücker coordinates. During the beam tracing phase, the visibility of the environment from the source position (the beam tree) is evaluated by traversing the precomputed visibility data structure and testing the presence of beams inside the visibility regions. The Plücker parameterization simplifies this procedure and reduces its computational burden, as it turns out to be an iterative intersection of linear subspaces. Similarly, during the path determination phase, acoustic paths are found by testing their presence within the nodes of the beam tree data structure. The simulations show that, with an average computation time per beam in the order of a dozen of microseconds, the proposed method can compute a large number of beams at rates suitable for interactive applications with moving sources and receivers.
Dejan Markovic, Fabio Antonacci, Augusto Sarti, Stefano Tubaro
IEEE Trans. Vis. Comput. Graph.4
2015 A Dimensional Contextual Semantic Model for music description and retrieval
abstract
Several paradigms for high-level music descriptions have been proposed to develop effective system for browsing and retrieving musical content in large repositories. Such paradigms are based on either categorical or dimensional models. The interest in dimensional models has recently grown a great deal, as they define a semantic relation between concepts through graded descriptions. One problem that affects semantic descriptions is the ambiguity that often arises from using the same descriptor in different contexts. In order to overcome this difficulty, it is important to model and address polysemy, which is the property of words to take on different meanings depending on the use-context. In this paper we propose a Dimensional Contextual Semantic Model for defining semantic relations among descriptors in a context-aware fashion. This model is here used for developing a semantic music search engine. In order to evaluate the effectiveness of our model, we compare this engine with two systems that are based on different description models.
Michele Buccoli, Massimiliano Zanoni, Augusto Sarti, Stefano Tubaro
ICASSP5
2015 Hybrid coding of visual content and local image features
abstract
Distributed visual analysis applications, such as mobile visual search or Visual Sensor Networks (VSNs) require the transmission of visual content on a bandwidth-limited network, from a peripheral node to a processing unit. Traditionally, a “Compress-Then-Analyze” approach has been pursued, in which sensing nodes acquire and encode the pixel-level representation of the visual content, that is subsequently transmitted to a sink node in order to be processed. This approach might not represent the most effective solution, since several analysis applications leverage a compact representation of the content, thus resulting in an inefficient usage of network resources. Furthermore, coding artifacts might significantly impact the accuracy of the visual task at hand. To tackle such limitations, an orthogonal approach named “Analyze-Then-Compress” has been proposed [1]. According to such a paradigm, sensing nodes are responsible for the extraction of visual features, that are encoded and transmitted to a sink node for further processing. In spite of improved task efficiency, such paradigm implies the central processing node not being able to reconstruct a pixel-level representation of the visual content. In this paper we propose an effective compromise between the two paradigms, namely “Hybrid-Analyze-Then-Compress” (HATC) that aims at jointly encoding visual content and local image features. Furthermore, we show how a target tradeoff between image quality and task accuracy might be achieved by accurately allocating the bitrate to either visual content or local features.
Luca Baroffio, Matteo Cesana, Alessandro Redondi, Marco Tagliasacchi, Stefano Tubaro
ICIP5
2015 Phylogeny reconstruction for misaligned and compressed video sequences
abstract
In the last few years, the amount of videos distributed online has dramatically increased due to the popularity of media sharing platforms (e.g., YouTube, Vimeo, etc.). However, distributed videos are often edited copies of original content, typically referred to as near duplicates. In this paper, we face the problem of reconstructing a video phylogeny tree, i.e., given a set of near-duplicate videos, we want to reconstruct the relationships between every pair of videos to detect which one generated the others and trace back their evolution history. Solving this problem is of paramount importance when the first published video within a set is sought, e.g., to solve copyright infringement cases or to pinpoint criminal impersonation online. The technique we propose exploits the same rationale of previous works in the field of image and video phylogeny. However, we embed in the commonly used pipeline of operations the possibility of dealing with temporally misaligned and encoded video sequences, thus making the method applicable to user-generated videos shared on online platforms. Results computed on a wide dataset of video sequences highlight the importance of taking care of both coding and misalignment in the reconstruction pipeline.
Filipe de Oliveira Costa, Silvia Lameri, Paolo Bestagini, Zanoni Dias, Anderson Rocha 0001, Marco Tagliasacchi, Stefano Tubaro
ICIP7
2015 Piecewise distortion correction for fisheye lenses
abstract
Lens distortion is a well-known problem for camera calibration. In particular in applications where a large amount of low-quality acquisition devices is adopted, like, e.g. embedded systems or Cyber-Physical Systems (CPS), a complete re-sectioning and undistortion in different operating conditions (e.g. autofocus, zooming) could not be feasible and accurate. Usually undistortion is obtained building a proper invertible geometrical distortion model with a specific number of parameters, but, unfortunately with the actually available low cost wide-angle and ultra wide-angle lenses a simple mathematical model for a global closed form solution can be ineffective in practical cases in particular in the peripheral image regions. In order to account for these problems we present a novel local correction approach based on the planarity and orthogonality constrains for a planar target (a checkerboard) where a Look-Up Table (LUT) is built to provide the minimum displacement for every pixel in the target region minimizing interpolation and residual distortion. The proposed method can also be considered a preliminary step in order to recognize image elements before global undistortion e.g. for features matching in images stitching.
Marco Marcon, Augusto Sarti, Stefano Tubaro
ICIP3
2015 Near-duplicate detection and alignment for multi-view videos
abstract
The increasing popularity of video sharing platforms (e.g., YouTube, Vimeo, etc.) has determined the widespread diffusion of near-duplicate videos, i.e., sequences obtained applying different editing operations to the same original clip. However, it is also possible to come across sequences referring to the same specific event shot from different viewpoints. This is a very common situation that arises when analyzing user-generated content acquired with mobile devices. Therefore, for some applications, it can be useful to extend the concept of near-duplicates considering also all the videos (and their edited versions) referring to the same event even if shot from different viewpoints. In this paper we consider such challenging scenario. More specifically, we focus on the problem of multi-view near-duplicate video detection and temporal alignment. In doing so, we show the limitations of a state-of-the-art algorithm based on robust hashing, and propose a processing pipeline that allows to deal also with sequences taken from significantly different viewpoints.
Andrea Melloni, Silvia Lameri, Paolo Bestagini, Marco Tagliasacchi, Stefano Tubaro
ICIP5
2015 A Robust and Low-Complexity Source Localization Algorithm for Asynchronous Distributed Microphone Networks
abstract
In this paper, we propose a robust and low-complexity acoustic source localization technique based on time differences of arrival (TDOA), which addresses the scenario of distributed sensor networks in 3D environments. Network nodes are assumed to be unsynchronized, i.e., TDOAs between microphones belonging to different nodes are not available. We begin with showing how to select feasible TDOAs for each sensor node, exploiting both geometrical considerations and a characterization of the overall generalized cross correlation (GCC) shape. We then show how to localize sources in the space-range reference frame, where TDOA measurements have a clear geometrical interpretation that can be fruitfully used in the scenario of unsynchronized sensors. In this framework, in fact, the source corresponds to the apex of a hypercone passing through points described by the sole microphone positions and TDOA measurements. The localization problem is therefore approached as a hypercone fitting problem. Finally, in order to improve the robustness of the estimate, we include an outlier detection procedure based on the evaluation of the hypercone fitting residuals. A refinement of source location estimate is then performed ignoring the contributions coming from outlier measurements. A set of simulations shows the performance of individual blocks of the system, with particular focus on the effect of TDOA selection on source localization and refinement steps. Experiments on real data validate the localization algorithm in an everyday scenario, proving that good accuracy can be obtained while saving computational cost in comparison with state-of-the-art techniques.
Antonio Canclini, Paolo Bestagini, Fabio Antonacci, Marco Compagnoni, Augusto Sarti, Stefano Tubaro
IEEE ACM Trans. Audio Speech Lang. Process.6
2015 Multiview Soundfield Imaging in the Projective Ray Space
abstract
A soundfield image is a data structure that efficiently encodes and represents the wave field as captured by a microphone array. Its representation is based on the directional plenacoustic function, which is defined as the radiance of the acoustic paths (rays) that cross the segment that the array lies upon. The soundfield image can be processed “as is” to develop a variety of applications. In its original formulation, the soundfield image is based on a Euclidean parameterization that can accommodate a limited range of rays and is suitable for managing a single array only. In this paper, we generalize this methodology to the case of multiple microphone arrays deployed in space. The use of multiple arrays allows us to capture truly global information on the sound field but requires us to rethink the ray space, and adopt a global representation of the acoustic rays based on projective geometry. After introducing the new parameterization, we present two examples of applications: the estimation of the mutual poses of two or more arrays (self-calibration); and the localization of multiple acoustic sources. The effectiveness of these applications is proven through simulations as well as real data experiments.
Dejan Markovic, Fabio Antonacci, Augusto Sarti, Stefano Tubaro
IEEE ACM Trans. Audio Speech Lang. Process.4
2015 Coding Local and Global Binary Visual Features Extracted From Video Sequences
abstract
Binary local features represent an effective alternative to real-valued descriptors, leading to comparable results for many visual analysis tasks while being characterized by significantly lower computational complexity and memory requirements. When dealing with large collections, a more compact representation based on global features is often preferred, which can be obtained from local features by means of, e.g., the bag-of-visual word model. Several applications, including, for example, visual sensor networks and mobile augmented reality, require visual features to be transmitted over a bandwidth-limited network, thus calling for coding techniques that aim at reducing the required bit budget while attaining a target level of efficiency. In this paper, we investigate a coding scheme tailored to both local and global binary features, which aims at exploiting both spatial and temporal redundancy by means of intra- and inter-frame coding. In this respect, the proposed coding scheme can conveniently be adopted to support the analyze-then-compress (ATC) paradigm. That is, visual features are extracted from the acquired content, encoded at remote nodes, and finally transmitted to a central controller that performs the visual analysis. This is in contrast with the traditional approach, in which visual content is acquired at a node, compressed and then sent to a central unit for further processing, according to the compress-then-analyze (CTA) paradigm. In this paper, we experimentally compare the ATC and the CTA by means of rate-efficiency curves in the context of two different visual analysis tasks: 1) homography estimation and 2) content-based retrieval. Our results show that the novel ATC paradigm based on the proposed coding primitives can be competitive with the CTA, especially in bandwidth limited scenarios.
Luca Baroffio, Antonio Canclini, Matteo Cesana, Alessandro Redondi, Marco Tagliasacchi, Stefano Tubaro
IEEE Trans. Image Process.6
2014 Robust beamforming under uncertainties in the loudspeakers directivity pattern
abstract
In this paper we propose a robust beamforming technique which takes into account uncertainties and variations in the radiation pattern of the loudspeakers. The proposed technique is based on the solution of a robust least-square problem in which the propagation matrix is to some extent unknown. Both simulations and experimental results prove the validity of the proposed methodology in terms of directivity index and white noise gain.
Lucio Bianchi, R. Magalotti, Fabio Antonacci, Augusto Sarti, Stefano Tubaro
ICASSP5
2014 Demosaicing strategy identification via eigenalgorithms
abstract
The identification of the camera that has acquired a specific image can be performed via several device-related footprints. Among these, it is possible to look for the traces left by the adopted color demosaicing strategy, which varies according to the camera model and vendor. The paper presents an identification strategy that re-processes the analyzed image with a set of distinctive CFA interpolation algorithms (eigenalgorithms) and, according to the correlation of the output with the original image, builds a set of features that permits identifying the algorithm. The proposed solution performs well with respect to other state-of-the-art solutions also when the analyzed image is severely compressed.
Simone Milani, Paolo Bestagini, Marco Tagliasacchi, Stefano Tubaro
ICASSP4
2014 Antiforensic synthesis of motion vectors using template algorithms
abstract
The identification of the video camera employed to acquire a video sequence is made possible by a large set of different footprints. Since video signals are always available in a compressed format, some of the most significant traces can be related to the coding tools of the implemented video codec (e.g., rate-distortion optimization, motion estimation strategy, etc.). As a matter of fact, an effective antiforensic attack, which aims at fooling the tools that identify the acquisition device, must appropriately alter these footprints. In the paper, we present an antiforensic strategy that targets a video camera detector which is based on the identification of the motion estimation strategy used by the video coder. The proposed approach synthesizes a set of motion vectors that approximate those that would have been generated by the algorithm to be mimicked. This method proves to be effective in attacking the detector while preserving the coding efficiency.
Simone Milani, Paolo Bestagini, Marco Tagliasacchi, Stefano Tubaro
ICASSP4
2014 Audio tampering detection using multimodal features
abstract
The authenticity verification of a User Generated Audio-Video content relative to a real event can be a very critical task especially when the content is shared on the Internet. Audio-Video files need to be checked in order to verify the origin of the content and the absence of alterations that could have changed their semantic content. The paper presents a multimodal approach for audio tampering detection that analyzes both the audio component and the video component of a recorded video file. The proposed solution estimates the volumetric characteristics of the environment where the multimedia content has been captured both from the video and audio signals. Then, the approach checks the consistency of the environment characteristics estimated from the audio signal with respect to those estimated from video files. The proposed solution proves to be useful in identifying video fakes and bootlegs, although it proves to be useful for the localization of added audio effects in a movie or radio track.
Simone Milani, Pier Francesco Piazza, Paolo Bestagini, Stefano Tubaro
ICASSP4
2014 Who is my parent? Reconstructing video sequences from partially matching shots
abstract
Nowadays, a significant fraction of the available video content is created by reusing already existing online videos. In these cases, the source video is seldom reused as is. Conversely, it is typically time clipped to extract only a subset of the original frames, and other transformations are commonly applied (e.g., cropping, logo insertion, etc.). In this paper, we analyze a pool of videos related to the same event or topic. We propose a method that aims at automatically reconstructing the content of the original source videos, i.e., the parent sequences, by splicing together sets of near-duplicate shots seemingly extracted from the same parent sequence. The result of the analysis shows how content is reused, thus revealing the intent of content creators, and enables us to reconstruct a parent sequence also when it is no longer available online. In doing so, we make use of a robust-hash algorithm that allows us to detect whether groups of frames are near-duplicates. Based on that, we developed an algorithm to automatically find near-duplicate matchings between multiple parts of multiple sequences. All the near-duplicate parts are finally temporally aligned to reconstruct the parent sequence. The proposed method is validated with both synthetic and real world datasets downloaded from YouTube.
Silvia Lameri, Paolo Bestagini, Andrea Melloni, Simone Milani, Anderson Rocha 0001, Marco Tagliasacchi, Stefano Tubaro
ICIP7
2014 Detectability-quality trade-off in JPEG counter-forensics
abstract
Removing JPEG quantization footprints from an image inevitably introduces artifacts and traces in the spatial domain. Recently, several robust methods have been proposed to detect footprints of counter-forensics and recover the image's compression history. In this paper we investigate the limitations of these detectors, by proposing an improved counter-forensic attack which adds a postprocessing denoising step besides dithering. We consider both a general-purpose denoising algorithm and one targeted to JPEG images. In the latter case, we show that this approach can successfully reduce the accuracy of detectors in the literature to that of a random decision. As a second contribution, we study the trade-off between the detectability of counter-forensics and quality of the tampered image, and show that the loss of quality is not sufficient for the analyst to use available no-reference quality assessment tools as an indicator of an attack.
Giuseppe Valenzise, Marco Tagliasacchi, Stefano Tubaro
ICIP3
2014 Coding Visual Features Extracted From Video Sequences
abstract
Visual features are successfully exploited in several applications (e.g., visual search, object recognition and tracking, etc.) due to their ability to efficiently represent image content. Several visual analysis tasks require features to be transmitted over a bandwidth-limited network, thus calling for coding techniques to reduce the required bit budget, while attaining a target level of efficiency. In this paper, we propose, for the first time, a coding architecture designed for local features (e.g., SIFT, SURF) extracted from video sequences. To achieve high coding efficiency, we exploit both spatial and temporal redundancy by means of intraframe and interframe coding modes. In addition, we propose a coding mode decision based on rate-distortion optimization. The proposed coding scheme can be conveniently adopted to implement the analyze-then-compress (ATC) paradigm in the context of visual sensor networks. That is, sets of visual features are extracted from video frames, encoded at remote nodes, and finally transmitted to a central controller that performs visual analysis. This is in contrast to the traditional compress-then-analyze (CTA) paradigm, in which video sequences acquired at a node are compressed and then sent to a central unit for further processing. In this paper, we compare these coding paradigms using metrics that are routinely adopted to evaluate the suitability of visual features in the context of content-based retrieval, object recognition, and tracking. Experimental results demonstrate that, thanks to the significant coding gains achieved by the proposed coding scheme, ATC outperforms CTA with respect to all evaluation metrics.
Luca Baroffio, Matteo Cesana, Alessandro Redondi, Marco Tagliasacchi, Stefano Tubaro
IEEE Trans. Image Process.5
2013 Detection of temporal interpolation in video sequences
abstract
Nowadays, considering the availability of relatively cheap devices and powerful editing software, video tampering is a relatively easy task. Video sequences can be tampered with by performing, e.g., temporal splicing. However, if the sequences spliced together do not share the same frame rate, they have to be temporally interpolated beforehand. This operation is often made using motion compensated interpolators, which allow to minimize visual artifacts. In this paper we propose a detector of this kind of interpolation. Moreover, the detector is capable of identifying the interpolation factor used, allowing an analyst to uncover the original frame rate of a sequence. This method relies on the analysis of the correlation introduced by the filter adopted by the interpolator. Results show that detection is successful, provided that the number of observed interpolated frames is large enough. Moreover, tests on compressed sequences obtained from television broadcasts validate the method in a real world scenario.
Paolo Bestagini, S. Battaglia, Simone Milani, Marco Tagliasacchi, Stefano Tubaro
ICASSP5
2013 Localization of virtual acoustic sources based on the Hough transform for sound field rendering applications
abstract
In this paper we propose a methodology for the localization of virtual acoustic sources for sound field rendering applications. After the reconstruction of the sound field in the listening area by means of circular harmonic decomposition, the virtual source location is found through the Hough transform. We prove the accuracy of the proposed methodology by comparing the source locations estimates with those of a subjective test campaign.
Lucio Bianchi, Fabio Antonacci, Antonio Canclini, Augusto Sarti, Stefano Tubaro
ICASSP5
2013 Improving action classification with volumetric data using 3D morphological operators
abstract
This work deals with the definition of a framework for interpreting, modeling and classifying sequences of body movements into a pre-defined vocabulary of actions. Starting from sequences of volumetric reconstructions of the actor pose in each frame, we split action recognition into three separated tasks. The first task is the representation of the four-dimensional patterns reconstructed from each sequence, the second task is the extraction of motion descriptors, and the third task is the classification into action classes. In particular, we extract the curve skeleton from the reconstructed volumes in order to underly the actor movements and to reduce the system dependence from the actor gender and the body shape. The proposed method increases the action recognition rate.
Eliana Frigerio, Marco Marcon, Stefano Tubaro
ICASSP3
2013 Antiforensics attacks to Benford's law for the detection of double compressed images
abstract
Researchers have been recently challenging the robustness of forensic algorithms by designing antiforensic strategies that try to fool them. In this paper, we propose an antiforensic strategy that targets double image compression detectors based on Benford's law (or first digit law). The proposed approach is able to modify the first digit statistics of the considered data (a double compressed image) to fool single/double compression detectors based on Benford's law. In this way, the proposed strategy tries to mimick the effects of a single compression with limited additional distortion. The presented algorithm performs better than previous state-of-the-art antiforensic strategies and can be easily extended to other fraud detection methods.
Simone Milani, Marco Tagliasacchi, Stefano Tubaro
ICASSP3
2013 Transform coder identification
abstract
The widespread popularity of transform coding has made it central to a wide range of methods in forensics, quality assessment and digital restoration. However, most approaches require prior knowledge of the transform coding parameters. In this paper, we consider the challenging problem of identifying the transform matrix as well as the quantization step sizes of a transform coder, given a set of P non-overlapping N-dimensional vectors observed as its output. We formulate the problem in terms of finding the lattice with the largest determinant that contains all observed vectors and we propose an algorithm that is able to find the optimal solution. Our experimental analysis shows that the probability of success of the algorithm quickly approaches 1 for small values of (P - N). The complexity of the proposed algorithm grows linearly with the dimensionality N.
Marco Tagliasacchi, Marco Visentini Scarzanella, Pier Luigi Dragotti, Stefano Tubaro
ICASSP4
2013 Coding video sequences of visual features
abstract
Visual features provide a convenient representation of the image content, which is exploited in several applications, e.g., visual search, object tracking, etc. In several cases, visual features need to be transmitted over a bandwidth-limited network, thus calling for coding techniques to reduce the required rate, while attaining a target efficiency for the task at hand. Although the literature has recently addressed the problem of coding local features extracted from still images, in this paper we propose, for the first time, a coding architecture designed for local features extracted from video content. We exploit both spatial and temporal redundancy by means of intra-frame and inter-frame coding modes. In addition, we propose a coding mode decision based on rate-distortion optimization. Experimental results demonstrate that, in the case of SIFT descriptors, exploiting temporal redundancy leads to substantial gains in terms of coding efficiency.
Luca Baroffio, Matteo Cesana, Alessandro Redondi, Stefano Tubaro, Marco Tagliasacchi
ICIP4
2013 Video recapture detection based on ghosting artifact analysis
abstract
Video forensics is becoming a popular field of research and an increasing number of forensic techniques have been proposed in the last few years. However, a simple yet effective method to fool many detectors consists in recapturing a video sequence with a camcorder. For this reason being able to detect video recapture is a topic of interest for a forensic analyst. In this paper, we first characterize the video recapture model, focusing on the common scenario of a sequence recaptured from a LCD monitor using a digital camcorder, then we propose a recapture detector for this case. The detector is based on the analysis of a characteristic ghosting artifact left by the recapture process. The presented algorithm is finally validated by means of tests on original and recaptured sequences. These tests prove that the algorithm achieves high accuracy results.
Paolo Bestagini, Marco Visentini Scarzanella, Marco Tagliasacchi, Pier Luigi Dragotti, Stefano Tubaro
ICIP5
2013 No-reference quality metric for depth maps
abstract
The performances of several 3D imaging/video applications (going from 3DTV to video surveillance) benefit from the estimation or acquisition of accurate and high quality depth maps. However, the characteristics of depth information is strongly affected by the procedure employed in its acquisition or estimation (e.g., stereo evaluation, ToF cameras, structured light sensors, etc.), and the very definition of “quality” for a depth map is still under investigation. In this paper we proposed an unsupervised quality metric for depth information in Depth Image Based Rendering signals that predicts the accuracy in synthesizing 3D models and lateral views by using the considered depth information. The metric has been tested on depth maps generate with different algorithms and sensors. Moreover, experimental results show how it is possible to progressively improve the performance of 3D modelization by controlling the device/algorithm with this metric.
Simone Milani, Daniele Ferrario, Stefano Tubaro
ICIP3
2013 Identification of the motion estimation strategy using eigenalgorithms
abstract
The identification of the device, or device model, that was used to acquire a video sequence is a very challenging task, since it has to rely on subtle traces left by the processing steps applied to the raw acquired data. Previous works have tried to address this problem leveraging the traces left by the imaging sensor. However, in the case of video, lossy coding is often quite aggressive, thus making these methods impractical. In this work, we reverse the analysis strategy and exploit the traces left by lossy coding as telltale for the adopted acquisition device. Specifically, we aim at detecting the implementation of the video codec by identifying the adopted motion estimation algorithm. Indeed, motion estimation is not defined in video coding standards and, as such, it represents one of the non-normative tools that can be customized in the design of the encoder. The key tenet consists in studying the correlation between the motion vectors obtained from the decoded bitstream, and those computed using a set of known and diverse motion estimation algorithms, called eigenalgorithms. In our work, we generalize a method recently appeared in the literature, which assumes that the motion estimation algorithm used is necessarily one of those available during the analysis. Experimental results show that the approach is able to successfully identify the motion estimation algorithm in most cases.
Simone Milani, Marco Tagliasacchi, Stefano Tubaro
ICIP3
2013 Transform coder identification with double quantized data
abstract
The analysis of chains of double transform coders has been recently addressed in the image forensic literature, especially for the case of double JPEG compression. In that case, the transform is assumed to be known a priori (e.g., 2D-DCT), whereas the quantization steps of the first coder need to be determined. In this work, we generalize the analysis to the challenging case in which nothing is known about the first coder, but that the transform is orthonormal. Given a set of vectors observed as output of a chain of two transform coders, we identify both the transform and the quantizer of the first. The key idea is to denoise the observed vectors exploiting the constraints imposed by the first quantizer and then apply our previously proposed method, which successfully performs transform identification in the case of noiseless observations. Experiments on real images validate the proposed approach.
Marco Tagliasacchi, Marco Visentini Scarzanella, Pier Luigi Dragotti, Stefano Tubaro
ICIP4
2013 Local tampering detection in video sequences
abstract
Video sequences are often believed to provide stronger forensic evidence than still images, e.g., when used in lawsuits. However, a wide set of powerful and easy-to-use video authoring tools is today available to anyone. Therefore, it is possible for an attacker to maliciously forge a video sequence, e.g., by removing or inserting an object in a scene. These forms of manipulation can be performed with different techniques. For example, a portion of the original video may be replaced by either a still image repeated in time or, in more complex cases, by a video sequence. Moreover, the attacker might use as source data either a spatio-temporal region of the same video, or a region taken from an external sequence. In this paper we present the analysis of the footprints left when tampering with a video sequence, and propose a detection algorithm that allows a forensic analyst to reveal video forgeries and localize them in the spatio-temporal domain. With respect to the state-of-the-art, the proposed method is completely unsupervised and proves to be robust to compression. The algorithm is validated against a dataset of forged videos available online.
Paolo Bestagini, Simone Milani, Marco Tagliasacchi, Stefano Tubaro
MMSP4
2013 Rendering of directional sources through loudspeaker arrays based on plane wave decomposition
abstract
In this paper we present a technique for the rendering of directional sources by means of loudspeaker arrays. The proposed methodology is based on a decomposition of the sound field in terms of plane waves. Within this framework the directivity of the source is naturally included in the rendering problem, therefore accommodating the directivity into the picture becomes much simpler. For this purpose, the loudspeaker array is subdivided into overlapping sub-arrays, each generating a plane wave component. The individual plane waves are then weighed by the desired directivity pattern. Simulations and experimental results show that the proposed technique is able to reproduce the sound field of directional sources with an improved accuracy with respect to existing techniques.
Lucio Bianchi, Fabio Antonacci, Augusto Sarti, Stefano Tubaro
MMSP4
2013 A music search engine based on semantic text-based query
abstract
Search and retrieval of songs from a large music repository usually relies on added meta-information (e.g., title, artist or musical genre); or on specific descriptors (e.g. mood); or on categorical music descriptors; none of which can specify the desired intensity. In this work, we propose an early example of semantic text-based music search engine. The semantic description takes into account emotional and non-emotional musical aspects. The method also includes a query-by-similarity search approach performed using semantic cues. We model both concepts and musical content in dimensional spaces that are suitable for carrying intensity information on the descriptors. We process the semantic query with a Natural Language parser to capture only the relevant words and qualifiers. We rely on Bayesian Decision theory to model concepts and songs as probability distributions. The resulted ranked list of songs are produced through a posterior probability model. A prototype of the system has been proposed to 53 subjects for evaluation, with good ratings on performance, usefulness and potential.
Michele Buccoli, Massimiliano Zanoni, Augusto Sarti, Stefano Tubaro
MMSP4
2013 A phylogenetic analysis of near-duplicate audio tracks
abstract
We present a content-based system for the analysis of near-duplicate audio tracks. The objective is to infer the structure of modifications, represented as trees, underneath a pool of near-duplicates, possibly specifying the operations the tracks have gone through. A pilot study was carried out for a set of plausible processing operators, including trim, fade and perceptual audio coding, generating near-duplicates from an original audio track. The proposed method measures the similarity between pairs of near-duplicates and reconstructs a tree representing the causal dependencies in the analyzed pool. Experimental results demonstrate that the structure of the tree can be successfully recovered, also in the challenging case in which some pieces of information are missing, i.e., when only a subset of near-duplicate tracks is available.
Matteo Nucci, Marco Tagliasacchi, Stefano Tubaro
MMSP3
2013 Acoustic Source Localization With Distributed Asynchronous Microphone Networks
abstract
We propose a method for localizing an acoustic source with distributed microphone networks. Time Differences of Arrival (TDOAs) of signals pertaining the same sensor are estimated through Generalized Cross-Correlation. After a TDOA filtering stage that discards measurements that are potentially unreliable, source localization is performed by minimizing a fourth-order polynomial that combines hyperbolic constraints from multiple sensors. The algorithm turns to exhibit a significantly lower computational cost compared with state-of-the-art techniques, while retaining an excellent localization accuracy in fairly reverberant conditions.
Antonio Canclini, Fabio Antonacci, Augusto Sarti, Stefano Tubaro
IEEE Trans. Speech Audio Process.4
2013 Soundfield Imaging in the Ray Space
abstract
In this work we propose a general approach to acoustic scene analysis based on a novel data structure (ray-space image) that encodes the directional plenacoustic function over a line segment (Observation Window, OW). We define and describe a system for acquiring a ray-space image using a microphone array and refer to it as ray-space (or “soundfield”) camera. The method consists of acquiring the pseudo-spectra corresponding to a grid of sampling points over the OW, and remapping them onto the ray space, which parameterizes acoustic paths crossing the OW. The resulting ray-space image displays the information gathered by the sensors in such a way that the elements of the acoustic scene (sources and reflectors) will be easy to discern, recognize and extract. The key advantage of this method is that ray-space images, irrespective of the application, are generated by a common (and highly parallelizable) processing layer, and can be processed using methods coming from the extensive literature of pattern analysis. After defining the ideal ray-space image in terms of the directional plenacoustic function, we show how to acquire it using a microphone array. We also discuss resolution and aliasing issues and show two simple examples of applications of ray-space imaging.
Dejan Markovic, Fabio Antonacci, Augusto Sarti, Stefano Tubaro
IEEE ACM Trans. Audio Speech Lang. Process.4
2013 Revealing the Traces of JPEG Compression Anti-Forensics
abstract
Due to the lossy nature of transform coding, JPEG introduces characteristic traces in the compressed images. A forensic analyst might reveal these traces by analyzing the histogram of discrete cosine transform (DCT) coefficients and exploit them to identify local tampering, copy-move forgery, etc. At the same time, it has been recently shown that a knowledgeable adversary can possibly conceal the traces of JPEG compression, by adding a dithering noise signal in the DCT domain, in order to restore the histogram of the original image. In this paper, we study the processing chain that arises in the case of JPEG compression anti-forensics. We take the perspective of the forensic analyst, and we show how it is possible to counter the aforementioned anti-forensic method revealing the traces of JPEG compression, regardless of the quantization matrix being used. Tests on a large image dataset demonstrated that the proposed detector was able to achieve an average accuracy equal to 93%, rising above 99% when excluding the case of nearly lossless JPEG compression.
Giuseppe Valenzise, Marco Tagliasacchi, Stefano Tubaro
IEEE Trans. Inf. Forensics Secur.3
2012 Video codec identification
abstract
Video content is routinely acquired and distributed in digital format. Therefore, it is customary to have the content encoded multiple times. In this paper we consider a processing chain of two coding steps and we propose a method that aims at identifying the type of codec used in the first step, by analyzing its coding-based footprints. The method relies on the fact that lossy coding is an almost idempotent operation, i.e., re-encoding the reconstructed sequence with the same codec and coding parameters produces a sequence that is highly correlated with the input one. As a consequence, it is possible to analyze this sort of correlation to identify the first codec provided that the second codec does not introduce severe quality degradation. The proposed solution finds several applications in the field of multi-media forensics, e.g. to identify the device that generated the original video stream or detect collages of different sequences.
Paolo Bestagini, Simone Milani, Marco Tagliasacchi, Stefano Tubaro
ICASSP5
2012 Discriminating multiple JPEG compression using first digit features
abstract
The analysis of double-compressed images is a problem largely studied by the multimedia forensics community, as it might be exploited, e.g., for tampering localization or source device identification. In many practical scenarios, e.g. photos uploaded on blogs, on-line albums, and photo sharing Web sites, images might be compressed several times. However, the identification of the number of compression stages applied to an image remains an open issue. This paper proposes a forensic method based on the analysis of the distribution of the first significant digits of DCT coefficients, which is modeled according to Benford's law. The method relies on a set of Support Vector Machine (SVM) classifiers and allows us to accurately identify the number of compression stages applied to an image. Up to four consecutive compression stages were considered in the experimental validation. The proposed approach extends and outperforms the previously published methods aimed at detecting double JPEG compression.
Simone Milani, Marco Tagliasacchi, Stefano Tubaro
ICASSP3
2012 Correction method for nonideal iris recognition
abstract
The use of iris as biometric trait has emerged as one of the most preferred method because of its uniqueness, lifetime stability and regular shape. Moreover it shows public acceptance and new user-friendly capture devices are developed and used in a broadened range of applications. Currently, iris recognition systems work well with frontal iris images from cooperative users. Nonideal iris images are still a challenge for iris recognition and can significantly affect the accuracy of iris recognition systems. In this paper, we propose a method to correct off-angle iris image. Taking into account the eye morphology and the reflectance properties of the external transparent layers, we can evaluate the distorting effect that is present in the acquired image. The correction algorithm proposed includes a first modeling phase of the human eye, a segmentation of the acquired image, and a simulation phase where the acquisition geometry is reproduced and the distortions are evaluated. Finally we obtain an image which does not contain the distorting effects due to jumps in the refractive index. We show how this correction process reduce the intra-class variations for off-angle iris images.
Eliana Frigerio, Marco Marcon, Augusto Sarti, Stefano Tubaro
ICIP4
2012 Multiple compression detection for video sequences
abstract
Nowadays, thanks to the increasingly availability of powerful processors and user friendly applications, the editing of video sequences is becoming more and more frequent. Moreover, after each editing step, any video object is almost always encoded in order to store it using a less amount of memory. For this reason, inferring the number of compression steps that have been applied to such a multimedia object is an important clue in order to assess its authenticity. In this paper we propose a method to recover the number of compression steps applied to a video sequence. In order to accomplish this goal, we make use of a classifier based on multiple Support Vector Machines (SVM) exploiting the Benford's law. Indeed, the feature vectors used to train and test the SVM are based on the statistics of the most significant digit of quantized transform coefficients. The proposed method is tested with a generic hybrid video encoder combining motion-compensation and block coding. Results show that this method is able to discriminate up to three compression stages with high accuracy.
Simone Milani, Paolo Bestagini, Marco Tagliasacchi, Stefano Tubaro
MMSP4
2012 3D wide baseline correspondences using depth-maps
Marco Marcon, Eliana Frigerio, Augusto Sarti, Stefano Tubaro
Signal Process. Image Commun.4
2012 Inference of Room Geometry From Acoustic Impulse Responses
abstract
Acoustic scene reconstruction is a process that aims to infer characteristics of the environment from acoustic measurements. We investigate the problem of locating planar reflectors in rooms, such as walls and furniture, from signals obtained using distributed microphones. Specifically, localization of multiple two- dimensional (2-D) reflectors is achieved by estimation of the time of arrival (TOA) of reflected signals by analysis of acoustic impulse responses (AIRs). The estimated TOAs are converted into elliptical constraints about the location of the line reflector, which is then localized by combining multiple constraints. When multiple walls are present in the acoustic scene, an ambiguity problem arises, which we show can be addressed using the Hough transform. Additionally, the Hough transform significantly improves the robustness of the estimation for noisy measurements. The proposed approach is evaluated using simulated rooms under a variety of different controlled conditions where the floor and ceiling are perfectly absorbing. Results using AIRs measured in a real environment are also given. Additionally, results showing the robustness to additive noise in the TOA information are presented, with particular reference to the improvement achieved through the use of the Hough transform.
Fabio Antonacci, Jason Filos, Mark R. P. Thomas, Emanuël A. P. Habets, Augusto Sarti, Patrick A. Naylor, Stefano Tubaro
IEEE Trans. Speech Audio Process.7
2012 Localization of Acoustic Sources Through the Fitting of Propagation Cones Using Multiple Independent Arrays
abstract
In this paper, we propose a novel acoustic source localization method that accommodates the general scenario of multiple independent microphone arrays. The method is based on a 3-D parameter space defined by the 2-D spatial location of a source and the range difference extracted from the time difference of arrival (TDOA). In this space, the set of points that correspond to a given range lie on a circle that expand as the range increases, forming a cone whose apex is the actual location of the source. In this parameter space, the lack of synchronization between arrays results in the fact that clusters of data associated to individual arrays are free to shift along the range axis. The cone constraint, in fact, enables the realignment of such clusters while positioning the cone vertex (source location), thus resulting in a joint data re-synchronization and source localization. We also propose a novel and general analysis methodology for swiftly assessing the localization error as a function of the TDOA uncertainties, which is remarkably accurate for small localization bias. With the aid of this method, simulations and experiments on real data, we show that the cone-fitting process offers excellent localization accuracy in the scenario of multiple unsynchronized arrays, as well as in simpler single-array scenarios, also in comparison with state-of-the-art techniques. We also show that the proposed method offers the desired flexibility for adapting to arbitrary geometries of microphone clusters.
Marco Compagnoni, Paolo Bestagini, Fabio Antonacci, Augusto Sarti, Stefano Tubaro
IEEE Trans. Speech Audio Process.5
2012 No-Reference Pixel Video Quality Monitoring of Channel-Induced Distortion
abstract
Video transmitted over an error-prone network may be received at the decoder with degradations due to packet losses. No-reference quality monitoring algorithms are the most practical way to measure the quality of the received video, since they do not impose any change with respect to the network architecture. Conventionally, these methods assume the availability of the corrupted bitstream. In some situations this is not possible, e.g., because the bitstream is encrypted or processed by third-party decoders, and only the decoded pixel values can be used. The major issue in this scenario is the lack of knowledge about which regions of the video have been actually lost, which is a fundamental ingredient for estimating channel-induced distortion. In this paper, we propose a maximum a posteriori estimation of the pattern of lost macroblocks, which assumes the knowledge of the decoded pixels only. This information can be used as input to a no-reference quality monitoring system, which produces an accurate estimate of the mean-square-error (MSE) distortion introduced by channel errors. The results of the proposed method are well correlated with the MSE distortion computed in full-reference mode, with a linear correlation coefficient equal to 0.9 at frame level and 0.98 at sequence level.
Giuseppe Valenzise, Stefano Magni, Marco Tagliasacchi, Stefano Tubaro
IEEE Trans. Circuits Syst. Video Technol.4
2011 A methodology for evaluating the accuracy of wave field rendering techniques
abstract
In this paper we propose a methodology for assessing the accuracy of techniques of wave field rendering through loudspeaker arrays. In order to measure the rendered wave field we adopt a solution based on a circular harmonic analysis of the sound field captured by a virtual microphone array. As a result of this analysis stage, we are able to compare the target, the theoretical and the measured wave fields, which may differ due to the non-ideality in the loudspeaker array or in the environment that generates some spurious reverberations. Moreover, in order to quantify the error between target, theoretical and measured wave fields, we define some evaluation metrics, based on RMSE and modal analysis of the acquired wave fields. We show some experimental results on real data.
Antonio Canclini, Paolo Annibale, Fabio Antonacci, Augusto Sarti, Rudolf Rabenstein, Stefano Tubaro
ICASSP6
2011 From direction of arrival estimates to localization of planar reflectors in a two dimensional geometry
abstract
In this paper we propose a novel technique to localize planar obstacles through the measurement of the Direction of Arrival by a microphone array. The measurement of the Direction of Arrival of the reflected path is turned into a quadratic constraint where the unknowns are the line parameters of the reflector. A cost function that combines multiple constraints is then derived. A parametric description of the obstacle is found by minimization of the cost function. Some simulations and experimental results show the feasibility of the proposed method.
Antonio Canclini, Paolo Annibale, Fabio Antonacci, Augusto Sarti, Rudolf Rabenstein, Stefano Tubaro
ICASSP6
2011 The cost of JPEG compression anti-forensics
abstract
The statistical footprint left by JPEG compression can be a valuable source of information for the forensic analyst. Recently, it has been shown that a suitable anti-forensic method can be used to destroy these traces, by properly adding a noise-like signal to the quantized DCT coefficients. In this paper we analyze the cost of this technique in terms of introduced distortion and loss of image quality. We characterize the dependency of the distortion on the image statistics in the DCT domain and on the quantization step used in JPEG compression. We also evaluate the loss of quality as measured by means of a perceptual metric, showing that a perceptually-optimized version of the anti-forensic method fails to completely conceal the forgery. Our conclusion is that removing the traces of the JPEG compression history could be much more challenging than it might appear, as anti-forensic methods are bound to leave characteristic traces.
Giuseppe Valenzise, Marco Tagliasacchi, Stefano Tubaro
ICASSP3
2011 Countering JPEG anti-forensics
abstract
JPEG coding leaves characteristic footprints that can be leveraged to reveal doctored images, e.g. providing the evidence for local tampering, copy-move forgery, etc. Recently, it has been shown that a knowledgeable attacker might attempt to remove such footprints by adding a suitable anti-forensic dithering signal to the image in the DCT domain. Such noise-like signal restores the distribution of the DCT coefficients of the original picture, at the cost of affecting image quality. In this paper we show that it is possible to detect this kind of attack by measuring the noisiness of images obtained by re-compressing the forged image at different quality factors. When tested on a large set of images, our method was able to correctly detect forged images in 97% of the cases. In addition, the original quality factor could be accurately estimated.
Giuseppe Valenzise, Vitaliano Nobile, Marco Tagliasacchi, Stefano Tubaro
ICIP4
2010 Geometric reconstruction of the environment from its response to multiple acoustic emissions
abstract
In this paper we propose a method for reconstructing the 2D geometry of the surrounding environment based on the signals acquired by a fixed microphone, when a series of acoustic stimula are produced in different positions in space. After estimating the Times Of Arrival (TOAs) of the reflective paths, we turn each TOA into a projective geometric constraint that can be used for determining the locations of the reflectors. The result consists of a collection of planar surfaces that correspond to the reflectors' locations. In this paper we present the whole processing chain and prove its effectiveness through experimental results.
Fabio Antonacci, Augusto Sarti, Stefano Tubaro
ICASSP3
2010 A H.264/AVC video database for the evaluation of quality metrics
abstract
This paper describes a publicly available database of subjective scores, relative to quality assessment of 156 video streams encoded with H.264/AVC and corrupted by simulating packet losses over an error-prone network. The data has been collected in controlled test environments at the premises of two academic institutions. A detailed statistical analysis of subjective results has been performed, showing high consistency of the collected scores. In addition to subjective scores, we have made available to the research community both the uncompressed files and the H.264/AVC bitstreams of each video sequence, in order to provide a common database of benchmark data to test and compare the performance of full-reference, reduced-reference and no-reference video quality assessment algorithms.
Francesca De Simone, Marco Tagliasacchi, Matteo Naccari, Stefano Tubaro, Stefano Ebrahimi
ICASSP4
2010 Visibility-based beam tracing for soundfield rendering
abstract
In this paper we present a visibility-based beam tracing solution for the simulation of the acoustics of environment that makes use of a projective geometry representation. More specifically, projective geometry turns out to be useful for the pre-computation of the visibility among all the reflectors in the environment. The simulation engine has a straightforward application in the rendering of the acoustics of virtual environments using loudspeaker arrays. More specifically, the acoustic wavefield is conceived as a superposition of acoustic beams, whose parameters (i.e. origin, orientation and aperture) are computed using the fast beam tracing methodology presented here. This information is processed by the rendering engine to compute spatial filters to be applied to the loudspeakers within the array. Simulative results show that an accurate simulation of the acoustic wavefield can be obtained using this approach.
Dejan Markovic, Antonio Canclini, Fabio Antonacci, Augusto Sarti, Stefano Tubaro
MMSP5
2010 Geometric calibration of distributed microphone arrays from acoustic source correspondences
abstract
This paper proposes a method that solves the problem of geometric calibration of microphone arrays. We consider a distributed system, in which each array is controlled by separate acquisition devices that do not share a common synchronization clock. Given a set of probing sources, e.g. loudspeakers, each array computes an estimate of the source locations using a conventional TDOA-based algorithm. These observations are fused together by the proposed method, in order to estimate the position and pose of one array with respect to the other. Unlike previous approaches, we explicitly consider the anisotropic distribution of localization errors. As such, the proposed method is able to address the problem of geometric calibration when the probing sources are located both in the near- and far-field of the microphone arrays. Experimental results demonstrate that the improvement in terms of calibration accuracy with respect to state-of-the-art algorithms can be substantial, especially in the far-field.
S. Daniele Valente, Marco Tagliasacchi, Fabio Antonacci, Paolo Bestagini, Augusto Sarti, Stefano Tubaro
MMSP6
2010 A reduced-reference structural similarity approximation for videos corrupted by channel errors
Marco Tagliasacchi, Giuseppe Valenzise, Matteo Naccari, Stefano Tubaro
Multim. Tools Appl.4
2010 Joint Compressive Video Coding and Analysis
abstract
Traditionally, video acquisition, coding and analysis have been designed and optimized as independent tasks. This has a negative impact in terms of consumed resources, as most of the raw information captured by conventional acquisition devices is discarded in the coding phase, while the analysis step only requires a few descriptors of salient video characteristics. Recent compressive sensing literature has partially broken this paradigm by proposing to integrate sensing and coding in a unified architecture composed by a light encoder and a more complex decoder, which exploits sparsity of the underlying signal for efficient recovery. However, a clear understanding of how to embed video analysis in this scheme is still missing. In this paper, we propose a joint compressive video coding and analysis scheme and, as a specific application example, we consider the problem of object tracking in video sequences. We show that, weaving together compressive sensing and the information computed by the analysis module, the bit-rate required to perform reconstruction and tracking of the foreground objects can be considerably reduced, with respect to a conventional disjoint approach that postpones the analysis after the video signal is recovered in the pixel domain. These findings suggest that considerable gains in performance can be potentially obtained in video analysis applications, provided that a joint analysis-aware design of acquisition, coding and signal recovery is carried out.
M. Cossalter, Giuseppe Valenzise, Marco Tagliasacchi, Stefano Tubaro
IEEE Trans. Multim.4
2009 A reduced-reference video structural similarity metric based on no-reference estimation of channel-induced distortion
abstract
The reduced-reference (RR) approximation of a full-reference (FR) video quality assessment method is a convenient way to build evaluation metrics which are both intrinsically well correlated with human judgments and feasible to implement in a network scenario, without the need to explore the perceptual significance of new video features through mean opinion score tests. In this paper, we propose a RR approximation of the video structural similarity index (VSSIM), a FR metric which is known to be well descriptive of the video quality perceived by users. We focus on the visual degradation produced by channel transmission errors: first, at the encoder, a small set of salient structural video features is assembled and transmitted through the RR channel to the end-user; then, at the decoder the feature vector is combined with a fine-granularity, no-reference estimate of the channel-induced distortion to produce the VSSIM approximation. By uniformly quantizing the feature vector and compressing it using a context-adaptive, variable length encoder, we show that good correlation coefficients with ground-truth VSSIM (rho = 0.85) may be achieved spending, respectively, less than 12 and 27 kbps for a video sequence with CIF or SD resolution.
Andrea Albonico, Giuseppe Valenzise, Matteo Naccari, Marco Tagliasacchi, Stefano Tubaro
ICASSP5
2009 Musical audio semantic segmentation exploiting analysis of prominent spectral energy peaks and multi-feature refinement
abstract
In this paper we present a novel hierarchical and scalable three-stage algorithm to effectively perform musical audio semantic segmentation. In the first stage, the energy spectrum of the entire audio track is analyzed to find significant energy textures that may characterize different semantic segments; in the second and third stages, tonal and timbric features are used to refine the segmentation by moving or deleting segment boundaries. Experimental results on a set of 58 songs show that our algorithm is able to attain good semantic segmentation just after the first step, with a precision of 64% and a recall of 96%. After second step the precision increases to 79%; the best precision result is obtained after the third step, where a value of 85% is reached. In this step the minimum average recall value of 92% is obtained.
P. Romano, Giorgio Prandi, Augusto Sarti, Stefano Tubaro
ICASSP4
2009 Subjective evaluation of a NO-reference video quality Monitoring algorithm for H.264/AVC video over a noisy channel
abstract
In this paper we evaluate NORM, a no-reference video quality monitoring algorithm we proposed in a previous work, for the prediction of the subjective quality of H.264/AVC video transmitted over a noisy packet-switched network. NORM produces an estimate of the mean square error distortion at the macroblock level between the noiseless and noisy sequence, without having access to the former. The output of NORM can be readily converted into a no-reference estimate of the PSNR at the sequence level. We carried out an extensive subjective evaluation campaign on CIF and 4CIF resolution sequences encoded with H.264/AVC and transmitted over a channel that drops packets at different packet loss rates, to obtain the differential mean opinion scores. Our results show that the estimated PSNR achieves good correlation with the subjective scores, very close to the ones achieved by the PSNR computed in full-reference mode, i.e. as if the noiseless sequence would be available at the decoder.
Matteo Naccari, Marco Tagliasacchi, Stefano Tubaro
ICIP3
2009 A compressive-sensing based watermarking scheme for sparse image tampering identification
abstract
In this paper we describe a robust watermarking scheme for image tampering identification and localization. A compact representation of the image is first produced by assembling a feature vector consisting of pseudo-random projections of the decimated image. Then, the quantized projections are encoded to form a hash, which is robustly embedded as a watermark in the image. By recovering the watermark the random projections are obtained, and then used to estimate the distortion of the received image. If tampering is sufficiently sparse or compressible in some basis description, a map of the introduced modification is recovered. The system relies on compressive sensing and distributed source coding principles to reduce the size of the hash of a 1024 × 1024 image, to about 4,000 bits. With this hash length, tampering sparse up to 20% and with a tampering energy around a PSNR of 15 dB can be successfully localized.
Giuseppe Valenzise, Marco Tagliasacchi, Stefano Tubaro, Giacomo Cancelli, Mauro Barni
ICIP3
2009 Geometric and radiometric modeling of 3D scenes
abstract
Modeling of 3D scenes is a hot topic in computer vision from more that thirty years, and probably its history is longer than a century considering also photogrammetry. In the recent years the rapid technological improvements that characterized the acquisition devices (photo-cameras, video-cameras, ..), illumination devices (lasers, structured light sources) and computational units allowed the application of 3D shape estimation methods, based on image analysis techniques, in a wide set of applications. Furthermore real-time 3D analysis is becoming a common tool in virtual and augmented reality contexts. Aim of this presentation is a rapid description of recent major advances on geometric and radiometric modeling of 3D scenes based on image analysis.
Marco Marcon, Augusto Sarti, Stefano Tubaro
ICME3
2009 Special issue on scalable coded media beyond compression
G. Charith K. Abhayaratne, Ebroul Izquierdo, Marta Mrak, Stefano Tubaro
Signal Process. Image Commun.4
2009 Hash-Based Identification of Sparse Image Tampering
abstract
In the last decade, the increased possibility to produce, edit, and disseminate multimedia contents has not been adequately balanced by similar advances in protecting these contents from unauthorized diffusion of forged copies. When the goal is to detect whether or not a digital content has been tampered with in order to alter its semantics, the use of multimedia hashes turns out to be an effective solution to offer proof of legitimacy and to possibly identify the introduced tampering. We propose an image hashing algorithm based on compressive sensing principles, which solves both the authentication and the tampering identification problems. The original content producer generates a hash using a small bit budget by quantizing a limited number of random projections of the authentic image. The content user receives the (possibly altered) image and uses the hash to estimate the mean square error distortion between the original and the received image. In addition, if the introduced tampering is sparse in some orthonormal basis or redundant dictionary, an approximation is given in the pixel domain. We emphasize that the hash is universal, e.g., the same hash signature can be used to detect and identify different types of tampering. At the cost of additional complexity at the decoder, the proposed algorithm is robust to moderate content-preserving transformations including cropping, scaling, and rotation. In addition, in order to keep the size of the hash small, hash encoding/decoding takes advantage of distributed source codes.
Marco Tagliasacchi, Giuseppe Valenzise, Stefano Tubaro
IEEE Trans. Image Process.3
2009 No-Reference Video Quality Monitoring for H.264/AVC Coded Video
abstract
When video is transmitted over a packet-switched network, the sequence reconstructed at the receiver side might suffer from impairments introduced by packet losses, which can only be partially healed by the action of error concealment techniques. In this context we propose NORM (NO-Reference video quality Monitoring), an algorithm to assess the quality degradation of H.264/AVC video affected by channel errors. NORM works at the receiver side where both the original and the uncorrupted video content is unavailable. We explicitly account for distortion introduced by spatial and temporal error concealment together with the effect of temporal motion-compensation. NORM provides an estimate of the mean square error distortion at the macroblock level, showing good linear correlation (correlation coefficient greater than 0.80) with the distortion computed in full-reference mode. In addition, the estimate at the macroblock level can be successfully exploited by forward quality monitoring systems that compute quality objective metrics to predict mean opinion score (MOS) values. As a proof of concept, we feed the output of NORM to a reduced-reference quality monitoring system that computes an estimate of the structural similarity metric (SSIM) score, which is known to be well correlated with perceptual quality.
Matteo Naccari, Marco Tagliasacchi, Stefano Tubaro
IEEE Trans. Multim.3
2008 Minimum variance multiplexing of multimedia objects
abstract
This paper addresses the problem of simultaneous transmission of multiple multimedia objects (such as images or video sequences) over a bandwidth-limited channel. The trivial strategy of partitioning in equal parts the available rate among the bitstreams is suboptimal, when the multimedia objects have different coding complexities. Exploiting object diversity allows us to allocate the bandwidth according to some optimality criteria, e.g. minimizing the average total distortion or minimizing the variance between the distortions of each object. By describing the rate-distortion characteristics of each multimedia object in terms of a simple exponential model, we provide a closed form solution for both the minimum average and the minimum variance problems. In addition, if we consider the statistical distribution of the rate-distortion model parameters, we can show that the minimum variance solution can effectively reduce the quality fluctuations among the objects, with an overall coding efficiency loss, w.r.t. the minimum average solution, of only 0.5dB on average. Some experiments, carried out on different H.264/AVC video sequences, validate our theoretical results.
Giuseppe Valenzise, Marco Tagliasacchi, Stefano Tubaro
ICASSP3
2008 No-reference modeling of the channel induced distortion at the decoder for H.264/AVC video coding
abstract
This paper proposes a model for estimating, at the decoder side, the distortion induced by the transmission over an error-prone channel, when error-free reconstructed frames are not available as a reference. The proposed estimation model considers explicitly the temporal error concealment algorithm adopted at the decoder. In the evaluation of the induced distortion, we model the effects of the absence of motion vectors and prediction residuals in the decoding process. In addition, we take into account the error propagation along successive frames. Experimental results conducted over real video sequences coded with the state-of-art H.264/AVC video coding standard validate the proposed model. In fact, the distortion estimated when no reference is available is strongly correlated both at the frame and group of pictures level with the actual distortion. This technique represents an effective no-reference video quality monitoring tool that can be embedded in any H.264/AVC compliant decoder.
Matteo Naccari, Marco Tagliasacchi, Fernando Pereira 0001, Stefano Tubaro
ICIP4
2008 Localization of sparse image tampering via random projections
abstract
Hashes can be used to provide authentication of multimedia contents. In the case of images, a hash can be used to detect whether the data has been modified in an illegitimate way. When the authentication check fails, it might be useful to localize the tampering in the spatial domain. This paper proposes an algorithm based on compressive sensing principles, which solves both the authentication and the localization problems. The encoder produces a hash using a small bit budget by quantizing a limited number of random projections of the authentic image. The decoder uses the hash to estimate the distortion between the original and the received image. In addition, if the attack is sparse, it can be also localized. In order to keep the size of the hash small, encoding/decoding takes advantage of distributed source codes. This paper also investigates experimentally the tradeoff between the rate allocated to the hash and the performance achieved in terms of tampering localization.
Marco Tagliasacchi, Giuseppe Valenzise, Stefano Tubaro
ICIP3
2008 Uncalibrated view synthesis from Relative Affine Structure based on planes parallelism
abstract
This paper focuses on the generation of physically valid views from two or more uncalibrated images acquired by standard cameras. The problem is faced without trying to yield a three dimensional reconstruction of the imaged scene, which would be unfeasible without the exact knowledge of the positions of the cameras in the Euclidean frame where the scene is to be described. Instead, starting from the previous works of Shashua and Navab on relative affine structure (1996) and the article of Fusiello on views synthesis from uncalibrated views (2007) we propose a novel approach that does not require the presence of a plane at infinity to define the homography between two views but merely the parallelism between couples of planes. This allows our approach to be applied to numerous scenes where two parallel planes can be defined (indoor scenes, straight streets and avenues). Experiments with synthetic images illustrate the approach.
Stefano Tebaldini, Marco Marcon, Augusto Sarti, Stefano Tubaro
ICIP4
2008 Reduced-reference estimation of channel-induced video distortion using distributed source coding
abstract
Channel-induced distortion estimation is an important aspect in the delivery of video contents over IP networks: the QoS requirements of both content providers and content users conflict with the intrinsic best-effort nature of packet-switched networks, which may introduce annoying artifacts in the received streams due to channel errors or jitter. In this paper we propose a Reduced-Reference video quality assessment method, based on objective quality metrics, which enables distortion estimation at the macroblock level. The content provider transmits a small feature vector for each frame, starting from random projections computed for each macroblock. In order to reduce the bit rate of the transmitted feature vector, we encode it using Distributed Source Coding (DSC) tools. The content user decodes the feature vector using the received sequence as side information. Additionally, the end-user may take advantage of some prior information about the support of the errors in the frame in such a way that the required bit length of the transmitted feature vector is further reduced. In our experiments, using 4 random projections, the use of DSC enables a bit saving of 70% w.r.t. scalar quantization and transmission of the original feature vector; when also the a priori error map is available at the decoder, the average length of the transmitted partial reference can be further reduced by another 5% of average.
Giuseppe Valenzise, Matteo Naccari, Marco Tagliasacchi, Stefano Tubaro
ACM Multimedia4
2008 Fast PDE approach to surface reconstruction from large cloud of points
Marco Marcon, Luca Piccarreta, Augusto Sarti, Stefano Tubaro
Comput. Vis. Image Underst.4
2008 3D Motion from structures of points, lines and planes
Andrea Dell'Acqua, Augusto Sarti, Stefano Tubaro
Image Vis. Comput.3
2008 Rate allocation for robust video streaming based on distributed video coding
Riccardo Bernardini, Matteo Naccari, Roberto Rinaldo, Marco Tagliasacchi, Stefano Tubaro, Pamela Zontone
Signal Process. Image Commun.5
2008 Fast Tracing of Acoustic Beams and Paths Through Visibility Lookup
abstract
The beam tracing method can be used for the fast tracing of a large number of acoustic paths through a direct lookup of a special tree-like data structure (beam tree) that describes the iterated visibility information from one specific position. This structure describes the branching of bundles of rays (beams) as they encounter reflectors in their paths. For this reason, beam tracing is suitable for real-time acoustic rendering even when the receiver is moving. In this paper, we propose a novel technique that enables the fast tracing of a large number of acoustic beams through the iterative lookup of a special data structure that describes the global visibility between reflectors. The method enables the immediate generation of the beam tree corresponding to an arbitrary source location, which can then be used for path tracing through direct lookup. In practice, this technique generalizes the traditional beam-tracing method as it makes it suitable for real-time acoustic rendering not just when the receiver is moving but also when the source is moving. The method enables real-time modeling of acoustic propagation and real-time auralization in complex 2-D and 2-Dtimes1-D environments (e.g., vertical walls limited by horizontal floor and ceiling), which makes it suitable for applications of real-time virtual acoustics, immersive gaming, and advanced acoustic rendering. Some experimental results show the effectiveness of fast beam tracing with respect to the state of the art in acoustic beam tracing.
Fabio Antonacci, Marco Foco, Augusto Sarti, Stefano Tubaro
IEEE Trans. Speech Audio Process.4
2008 Minimum Variance Optimal Rate Allocation for Multiplexed H.264/AVC Bitstreams
abstract
Consider the problem of transmitting multiple video streams to fulfill a constant bandwidth constraint. The available bit budget needs to be distributed across the sequences in order to meet some optimality criteria. For example, one might want to minimize the average distortion or, alternatively, minimize the distortion variance, in order to keep almost constant quality among the encoded sequences. By working in the rho-domain, we propose a low-delay rate allocation scheme that, at each time instant, provides a closed form solution for either the aforementioned problems. We show that minimizing the distortion variance instead of the average distortion leads, for each of the multiplexed sequences, to a coding penalty less than 0.5 dB, in terms of average PSNR. In addition, our analysis provides an explicit relationship between model parameters and this loss. In order to smooth the distortion also along time, we accommodate a shared encoder buffer to compensate for rate fluctuations. Although the proposed scheme is general, and it can be adopted for any video and image coding standard, we provide experimental evidence by transcoding bitstreams encoded using the state-of-the-art H.264/AVC standard. The results of our simulations reveal that is it possible to achieve distortion smoothing both in time and across the sequences, without sacrificing coding efficiency.
Marco Tagliasacchi, Giuseppe Valenzise, Stefano Tubaro
IEEE Trans. Image Process.3
2007 Tracking of two acoustic sources in reverberant environments using a particle swarm optimizer
abstract
In this paper we consider the problem of tracking multiple acoustic sources in reverberant environments. The solution that we propose is based on the combination of two techniques. A blind source separation (BSS) method known as TRINICON [5] is applied to the signals acquired by the microphone arrays. The TRINICON de-mixing filters are used to obtain the Time Differences of Arrival (TDOAs), which are related to the source location through a nonlinear function. A particle filter is then applied in order to localize the sources. Particles move according to a swarm-like dynamics, which significatively reduces the number of particles involved with respect to traditional particle filter. We discuss results for the case of two sources and four microphone pairs. In addition, we propose a method, based on detecting source inactivity, which overcomes the ambiguities that intrinsically arise when only two microphone pairs are used. Experimental results demonstrate that the average localization error on a variety of pseudo-random trajectories is around 40 cm when the T60reverberation time is 0.6s.
Fabio Antonacci, Davide Riva, Augusto Sarti, Marco Tagliasacchi, Stefano Tubaro
AVSS5
2007 Hash-Based Motion Modeling in Wyner-Ziv Video Coding
abstract
Generally, distributed video coding (DVC) schemes perform motion estimation at the decoder side, without the current frame being available. In order to generate the side-information reliably, one solution consists in allocating a limited bit budget to send a hash of the current frame. At the decoder, this auxiliary hash is used to perform motion estimation. This paper studies the accuracy of hash-based motion estimation and compares it to conventional encoder-side motion estimation. We show that, at low rates, the very limited bit-budget of the hash does not ensure a reliable motion estimation, while at medium to high rates the motion accuracy is comparable with the finite precision used to represent motion vectors. Then, we derive the rate-distortion characteristic, which combines the cost of encoding the hash and the prediction residuals after decoder-side motion compensation. We show that, at high rates, hash-based motion modeling can virtually achieve the same coding efficiency as motion-compensated predictive coding. Instead, at medium-to-low rates we observe a significant coding loss. Experimental results on real video sequences validate the results of the proposed model.
Marco Tagliasacchi, Stefano Tubaro
ICASSP (1)2
2007 Analysis of Coding Efficiency of Motion-Compensated Interpolation at the Decoder in Distributed Video Coding
abstract
This paper analyzes the coding efficiency of distributed video coding (DVC) schemes that perform motion-compensated interpolation at the decoder. The decoder has access only to the key frames when generating the side information for intermediate frames. This fact introduces a displacement estimation error that depends on several factors: 1) the overall motion complexity; 2) the temporal coherence of the motion field; 3) the temporal distance between successive key frames. Adopting a state-space model and a Kalman filtering framework, we obtain an estimate of the displacement error variance. This is used to determine the rate-distortion function of the overall coding scheme, that takes into account both intra-coded key frames and DVC-coded frames. The proposed model shows that motion-compensated interpolation is unable to achieve the coding efficiency of conventional motion-compensated predictive coding.
Marco Tagliasacchi, Laura Frigerio, Stefano Tubaro
ICIP (3)3
2007 Symmetric Distributed Coding of Stereo Video Sequences
abstract
In this paper we present a novel video coding scheme to compress stereo video sequences. We consider a wireless sensor network scenario, where the sensing nodes cannot communicate with each other and are characterized by limited computational complexity. The joint decoder exploits both the temporal and inter-view correlation to generate the side information. To this end, we propose a fusion algorithm that adaptively selects either the temporal or the inter-view side information on a pixel-by-pixel basis. In addition, the coding algorithm is symmetric with respect to the two cameras. We also propose a practical stopping criterion for turbo decoding that determines when decoding is successful. Experimental results on stereo video sequences show that a coding efficiency gain up to 4dB can be obtained by the proposed scheme at high bit-rates.
Marco Tagliasacchi, Giorgio Prandi, Stefano Tubaro
ICIP (2)3
2007 Rate-Distortion Analysis of Motion-Compensated Interpolation at the Decoder in Distributed Video Coding
abstract
This letter analyzes the coding efficiency of distributed video coding (DVC) schemes that perform motion-compensated interpolation at the decoder. The decoder has access only to the key frames, when generating the side information for intermediate frames. Therefore, the true motion field necessary for this operation is not directly available, and the motion vectors must be estimated at the decoder side, thus introducing displacement estimation errors. The accuracy of the motion-compensated interpolation at the decoder depends on several factors: 1, the overall motion complexity; 2, the temporal coherence of the motion field; and 3, the temporal distance between successive key frames. Adopting a state-space model and a Kalman filtering framework, we obtain an estimate of the displacement error variance. This is used to determine the rate-distortion function of the overall coding scheme, that takes into account both intra-coded key frames and DVC-coded frames. The proposed model shows that motion-compensated interpolation is unable to achieve the coding efficiency of conventional motion-compensated predictive coding. In addition, the model provides a good estimate of the group of pictures size that optimizes the coding efficiency. Experimental results on real video sequences validate the results of the proposed model.
Marco Tagliasacchi, Laura Frigerio, Stefano Tubaro
IEEE Signal Process. Lett.3
2006 A Proposal to Suppress the Training Stage in a Coset-Based Distributed Video Codec
abstract
Distributed video coding (DVC) is a coding paradigm that gives the decoder the task to exploit the source statistics to achieve efficient compression. Many approaches to the DVC problem have recently appeared in the literature, including the PRISM codec. Instead of encoding the deterministic quantized prediction error residual, PRISM partitions the quantization lattice into cosets and sends the index of the coset each quantized coefficient belongs to. Estimating the number of cosets is of crucial importance to achieve good coding efficiency. In PRISM, this is determined during an offline training phase. The present work aims at being a starting point for the suppression of the training stage of PRISM at the cost of sending the number of cosets for each DCT coefficient. The statistics of the number of cosets are analyzed to figure out the maximum compression efficiency achievable by entropy coding. Furthermore the paper discusses some techniques that might be used to lower the amount of transmitted bits. Based on these results, directions for future works are proposed
Xavier Artigas, Marco Tagliasacchi, Stefano Tubaro
ICASSP (2)4
2006 Improved Bit Allocation in an Error-Resilient Scheme Based on Distributed Source Coding
abstract
In this work we propose an error-resilient scheme that allows enhancing the robustness of a video stream. Based on distributed source coding (DSC) principles, an auxiliary stream is sent in parallel to the main stream as a redundant representation of the sequence that is used to correct errors at the decoder, thus reducing the impact of drift. In order to perform an optimal bit allocation in the auxiliary stream, the encoder needs to compute a reliable estimate of the expected video distortion observed at the decoder side due to channel loss. This paper proposes an algorithm to calculate the expected distortion of decoded DCT-coefficients (dubbed EDDD) and its application to the bit allocation problem in a DSC based auxiliary stream
Marco Fumagalli, Marco Tagliasacchi, Stefano Tubaro
ICASSP (2)3
2006 3-D Body Posture Tracking For Human Action Template Matching
abstract
In this paper we present a novel approach to 3-D human action classification based on the analysis of volumetric data obtained form the joint processing of video sequences acquired by a multiple-camera system. The use of volumetric data makes the system very robust and avoids problems related the typical human body self-occlusions and motion ambiguities, very common in an independent camera-by-camera analysis. A shape descriptor of a human body is obtained in order to capture only posture-dependent characteristics and its outputs at each time instant are collected together in action feature matrices. The use of dynamic time warping approach for action template matching accounts for possible temporal nonlinear distortions among different instances of the same gesture and allows gesture classification
Massimiliano Pierobon, Marco Marcon, Augusto Sarti, Stefano Tubaro
ICASSP (2)4
2006 Mixed 2D-3D Information for Pose Estimation and Face Recognition
abstract
Face recognition based on 3D techniques is a promising approach since it takes advantage of the additional information provided by depth which makes the whole approach more robust against illumination and pose variations. However, these 3D approaches require the cooperation of the person to acquire accurate 3D data; thus, they are not appropriated for some applications such as video surveillance or restricted area access points where only a 2D face image is disposable. In this paper, a novel approach is presented which takes advantage of 3D data in the training stage but only requires 2D data in the recognition stage. The proposed method can be used for both pose estimation and face recognition. Moreover, the estimation of the pose can be used as side information to improve the performance of the face recognition stage. Experiments have been carried out on the public UPC face database which is composed of a total of 621 face images of several persons taken from different views and illuminations
Antonio Rama, Francesc Tarres, Davide Onofrio, Stefano Tubaro
ICASSP (2)4
2006 Intra Mode Decision Based on Spatio-Temporal Cues in Pixel Domain Wyner-ZIV Video Coding
abstract
Distributed source coding principles have been recently applied to video coding in order to achieve a flexible distribution of the complexity burden between the encoder and the decoder. In this paper we elaborate on a pixel based Wyner-Ziv video codec that shifts all the complexity of the motion estimation phase to the decoder, thus achieving light encoding. We observe that the correlation noise statistics describing the relationship between the frame to be encoded and the side information available at the decoder is not spatially stationary. For this reason we introduce a mode decision scheme either at the encoder or at the decoder in such a way that when the estimated correlation is weak we opt for intra coding on a block-by-block basis. Both spatial and temporal criteria are used to determine whether a block is better intra coded or not
Marco Tagliasacchi, Alan Trapanese, Stefano Tubaro, João Ascenso, Catarina Brites, Fernando Pereira 0001
ICASSP (2)3
2006 A Robust Method for the Estimation of Reliable Wide Baseline Correspondences
abstract
In this paper we present a complete method to retrieve reliable correspondences among wide baseline images, that is images of the same scene/object acquired from very different viewpoints. We propose a solution based on matching of affine co-variant features, composed by the following four steps: interest region detection, normalization, description and matching. In our method we implemented improved versions of some techniques recently introduced in the literature: the MSER detector (maximally stable extremal regions) and SIFT and RIFT descriptors (scale/rotation invariant feature transform). After a general introduction to the wide baseline problems and a summary of the recent state-of-the-art solutions, we illustrate the proposed method detailing the added improvements, then we present some experimental results obtained on wide baseline images.
Francesco Colletto, Marco Marcon, Augusto Sarti, Stefano Tubaro
ICIP4
2006 P2CA: How Much Face Information is Needed?
abstract
Multimodal 2D+3D face biometrics commonly report that performance improves relative to that of a single modality. Complete 2D and 3D data can be available during training because they are acquired in a controlled scenario. However, in the evaluation scenario, only partial 2D and 3D data can be acquired and hence available for recognition. In this paper we present experimental results that determine how partial data contribute to the task of recognition using partial principal component analysis (P2CA) algorithm in a multimodal scheme. From our results it seems that discrimination power on individuals is ascribed to different regions of the face if we consider 2D or 3D data.
Davide Onofrio, Antonio Rama, Francesc Tarres, Stefano Tubaro
ICIP4
2006 On the Modeling of Motion in Wyner-Ziv Video Coding
abstract
In the past few years, a number of practical video coding schemes following distributed source coding principles have emerged. One of the main goals of distributed video coding (DVC) is to enable a flexible distribution of the computational complexity between the encoder and the decoder, while approaching the coding efficiency of conventional closed-loop motion-compensated predictive codecs. In this paper we perform a rate-distortion analysis of a well-known Wyner-Ziv architecture, while focusing our attention on the impact of the motion modeling that is used for generating the side information at the decoder. Our analysis is structured according to a Kalman filtering problem and it allows us to compare three different scenarios: motion estimation at the encoder; motion interpolation at the decoder; and motion extrapolation at the decoder.
Marco Tagliasacchi, Stefano Tubaro, Augusto Sarti
ICIP2
2006 Exploiting Spatial Redundancy in Pixel Domain Wyner-Ziv Video Coding
abstract
Distributed video coding is a recent paradigm that enables a flexible distribution of the computational complexity between the encoder and the decoder building on top of distributed source coding principles. In this paper we focus on the scenario where most of the complexity is shifted to the decoder, thus achieving light encoding. We elaborate on a well known pixel based Wyner-Ziv architecture and we improve its coding efficiency by exploiting both spatial and temporal correlation at the decoder side, without the need of performing any transform at the encoder. In order to generate the side information, the decoder adaptively chooses spatial or temporal information, based on the local estimate of the correlation noise. Simulations on test sequences demonstrate that a coding gain of up to +1.8 dB can be obtained with respect to the case that generates the side information by motion interpolation only.
Marco Tagliasacchi, Alan Trapanese, Stefano Tubaro, João Ascenso, Catarina Brites, Fernando Pereira 0001
ICIP3
2006 Motion Estimation by Quadtree Pruning and Merging
abstract
In this paper we propose a rate-distortion optimized motion estimation algorithm that is built upon a quadtree structure. Each node of the quadtree represents a block in the current frame together with its motion vector, and the block size decreases from the root to the leaves. In the first step, the quadtree is pruned according to a rate-distortion criterion in order to obtain blocks of variable sizes. A further rate rebate can be achieved by merging those leaf nodes of the quadtree that can be efficiently represented by the same motion vector. The proposed merging scheme provides a reduction of up to 50% of the rate spent for the motion model with respect to the case that performs pruning only
Marco Tagliasacchi, Mauro Sarchi, Stefano Tubaro
ICME3
2006 Video Quality Assessment from the Perspective of a Network Service Provider
abstract
In this paper we consider the problem of estimating the quality of a decoded video that is transmitted over an unreliable packet-switched network. The goal is to allow the encoder to calculate the expected decoder distortion without any direct comparison between the original and the decoded sequence. This work is based on ROPE algorithm for the video quality estimate. ROPE calculates, at the encoder side, the expected decoded video distortion: therefore it is reasonably performing in estimating the average obtained distortion. The main question this paper would like to answer is how accurate ROPE distortion estimate is in the case of a single realization, i.e., in the case of distortion estimate for a single user
Marco Fumagalli, Rosa Lancini, Stefano Tubaro
MMSP3
2006 Robust wireless video multicast based on a distributed source coding approach
Marco Tagliasacchi, Abhik Majumdar, Kannan Ramchandran, Stefano Tubaro
Signal Process.4
2006 A sequence-based error-concealment algorithm for an unbalanced multiple description video coding system
Marco Fumagalli, Rosa Lancini, Stefano Tubaro
Signal Process. Image Commun.3
2005 Colored visual tags: a robust approach for augmented reality
abstract
This paper presents a robust method for fast visual tags reading, suitable for augmented reality (AR) environments. Tag detection is based on well known tools of image-processing, but their combination, together with the use of colored markers, allows a robust recognition even with low-cost CMOS or CCD cameras and in poorly illuminated environments. In particular the color mix and the structure of the tag are quite unusual in common environments and can be easily detected with color filtering and geometric analysis. The proposed tag carries binary information encoded in its structure: in the presented implementation a 32-bit code with 12 parity bits is encoded in the tag but extensions to longer codes can be easily devised.
Andrea Dell'Acqua, Marco Ferrari 0001, Marco Marcon, Augusto Sarti, Stefano Tubaro
AVSS5
2005 Clustering of human actions using invariant body shape descriptor and dynamic time warping
abstract
We propose a human action clustering method based on a 3D representation of the body in terms of volumetric coordinates. Features representing body postures are extracted directly from 3D data, making the system inherently insensitive to viewpoint dependence, motion ambiguities and self-occlusions. An invariant shape descriptor of human body is obtained in order to capture only posture-dependent characteristics, despite possible differences in translation, orientation, scale and body size. Frame-by-frame descriptions, generated from a gesture sequence, are collected together in matrices. Clustering of action matrices is eventually performed, and through a dynamic time warping (while computing the distance metric), we gain independence from possible temporal nonlinear distortions among different instances of the same gesture.
Massimiliano Pierobon, Marco Marcon, Augusto Sarti, Stefano Tubaro
AVSS4
2005 Efficient source localization and tracking in reverberant environments using microphone arrays
abstract
In this paper, we propose an algorithm for acoustic source localization and tracking that is suitable for reverberant environments. The approach that we propose is based on the iterative identification of the FIR channels that link source and microphones through an LMS method (multi-channel LMS), but we propose additional solutions that significantly improve this method in terms of computational efficiency and localization reliability, without affecting its convergence properties. This is achieved through a modified block-wise implementation of the approach combined with Kalman filtering. We also show the results of extensive comparative testing using novel performance parameters for the assessment of localization reliability.
Fabio Antonacci, Davide Lonoce, Marco Motta, Augusto Sarti, Stefano Tubaro
ICASSP (4)5
2005 3D object modeling with a voxelset carving approach
abstract
In the past few years several systems for object reconstruction based on the analysis of 2D images have been proposed. In order for such systems to be of practical use, the 3D data extraction process is expected to be fast and reliable. In this paper we propose a general approach for the reconstruction of complete 3D objects based on a mesh fusion algorithm. Every surface patch is obtained as a depth map using an algorithm based on graph cuts theory. Each depth map is then triangulated before using it in a fusion algorithm based on a voxel-set carving approach. The result of the process is a closed mesh representing the object surface with sub-voxel resolution.
Giovanni Dainese, Marco Marcon, Augusto Sarti, Stefano Tubaro
ICIP (1)4
2005 Combining MCTF with distributed source coding
abstract
Motion compensated temporal filtering (MCTF) has proved to be an efficient coding tool in the design of open-loop scalable video codecs. In this paper we propose a MCTF video coding scheme based on lifting where the prediction step is implemented using PRISM (power efficient, robust, high compression syndrome-based multimedia coding), a video coding framework built on distributed source coding principles. We study the effect of integrating the update step at the encoder or at the decoder side. We show that the latter approach allows improving the quality of the side information exploited during decoding. We present the analytical results obtained by modeling the video signal along the motion trajectories as an AR(1) process showing that the update step at the decoder allows to half the contribution of the quantization noise. We also include experimental results with real video data that demonstrate the potential of this approach when the video sequences are coded at low bitrates.
Marco Tagliasacchi, Stefano Tubaro, Augusto Sarti
ICIP (1)2
2005 A Model Based Energy Minimization Method for 3D Face Reconstruction
abstract
In the paper we present a model based method for generating 3D face models from multiple images by means of an energy minimization algorithm. The energy function takes into account of: i) how well the luminance profiles are transferred (through the 3D model) from one image to the others; ii) how smooth are the reconstructed surfaces; iii) how it is congruent with an adapted face template mode (Candide model). It is important to notice that with the proposed method for each considered face two different 3D models are reconstructed: one with high resolution and one with low resolution obtained a reshaped Candide model. The first model can be used, for example, in 3D face analysis/recognition systems, while the second can be more useful for searching a specific face model in a large database
Davide Onofrio, Stefano Tubaro
ICME2
2005 Using partial information for face recognition and pose estimation
abstract
The main achievement of this work is the development of a new face recognition approach called partial principal component analysis (P/sup 2/CA), which exploits the novel concept of using only partial information for the recognition stage. This approach uses 3D data in the training stage but it permits to use either 2D or 3D data in the recognition stage, making the whole system more flexible. Preliminary experiments carried out on a multi-view face database composed of 18 individuals have shown robustness against big pose variations obtaining higher recognition rates than the conventional PCA method. Moreover, the P/sup 2/CA method can estimate the pose of the face under different illuminations with accuracy of the 96.15% when classifying the face images in 0/spl deg/, /spl plusmn/30/spl deg/, /spl plusmn/45/spl deg/, /spl plusmn/60/spl deg/ and /spl plusmn/90/spl deg/ views.
Antonio Rama, Francesc Tarres, Davide Onofrio, Stefano Tubaro
ICME4
2005 Multiple description video coding for scalable and robust transmission over IP
abstract
In this paper, we address the problem of video transmission over unreliable networks, such as the Internet, where packet losses occur. The most recent literature indicates multiple description (MD) as a promising coding approach to handle this issue. Moreover, it has been shown also how important the use of motion compensation prediction is in an MD-coding scheme. This paper proposes two architectures for multiple description video coding, both of them are based on the motion compensation prediction loop. The common characteristic of the two architectures is the use of a polyphase down-sampling technique to create the MDs and to introduce cross-redundancy among the descriptions. The first scheme, that we call drift-compensation multiple description video coder (DC-MDVC) appears very robust when used in an error-prone environment, but it can provide only two descriptions. The second architecture, called independent flow multiple description video coder (IF-MDVC), generates multiple sets of data before the motion compensation loop; in this case, there are no severe limitations in the selection of the number of descriptions used by the coder.
Nicola Franchi, Marco Fumagalli, Rosa Lancini, Stefano Tubaro
IEEE Trans. Circuits Syst. Video Technol.4
2004 Action modeling with volumetric data
abstract
In this paper we propose and test an action recognition algorithm in which the images of the scene captured by a significant number of cameras are first used to generate a volumetric representation of a moving human body in terms of voxsets by means of volumetric intersection. The recognition stage is then performed directly on 3D data, allowing the system to avoid critical problems like viewpoint dependence and motion trajectory variability. Suitable features are extracted from the voxset representing the body and fed to a classical hidden Markov model to produce a finite-state description of the motion.
Fabio Cuzzolin, Augusto Sarti, Stefano Tubaro
ICIP3
2004 Scalable coding of variable size blocks motion vectors
abstract
In this paper we discuss an algorithm that is able to provide a scalable (multiresolution) representation of the motion field information. It has been recently demonstrated that, for an open loop wavelet video coder, it is possible to use at the decoder side a scaled version of the original motion information and the residual coefficients computed with the full resolution/quality motion field without incurring into drift. We propose a fully scalable wavelet based video coder that performs motion estimation-compensation in the wavelet domain. In particular the coding scheme is specifically designed for variable size block matching algorithms. In this scenario motion vectors are distributed across an irregular lattice according to a quadtree structure. The developed system allows a scalable representation of the motion field and a flexible allocation of the bit budget between motion and residual data. The simulations that we have carried out show, at low bit-rates, a significant gain of the proposed approach with respect to the case in which the motion information is coded lossless.
Davide Maestroni, Augusto Sarti, Marco Tagliasacchi, Stefano Tubaro
ICIP4
2004 Fast in-band motion estimation with variable size block matching
abstract
In this paper we propose a fast motion estimation technique that works in the wavelet domain. The computational cost of the algorithm turns out to be proportional to the linear size of the search window instead of its area. We complete our proposal with a variable-size block-matching scheme in the wavelet domain. We integrated the motion estimation algorithm in a fully-scalable wavelet in-band prediction coder inspired by the IB-MCTF (in-band motion compensation temporal filtering) proposed in J.C Ye et al., (2003). Comparative tests prove that our coder provides the same quality level than IB-MCTF at a very reduced computational cost. Moreover, although our method turns out to match the performance of MCTF-EZBC (P. Chen, 2003) in terms of PSNR, it clearly outperforms it in terms of perceptual quality, as it is completely free from blocking artifacts.
Davide Maestroni, Augusto Sarti, Marco Tagliasacchi, Stefano Tubaro
ICIP4
2004 Area matching based on belief propagation with applications to face modeling
Davide Onofrio, Augusto Sarti, Stefano Tubaro
ICIP3
2004 Accurate and fast audio-realistic rendering of sounds in virtual environments
abstract
In this paper we propose a novel method for real-time auralization of sounds in complex environments using visibility diagrams. The method accounts for both specular and diffracted reflections with receivers and sources that are free to move. Our solution concentrates in a pre-processing phase all the operations that can be conducted without knowledge of either source or receiver locations. In fact we pre-compute a set of visibility diagrams and diffracted beam trees. Once the source location is specified, we can determine the reflective beam trees through a simple lookup process on the visibility diagrams. The additional knowledge of the receiver location allows us to immediately generate all reflective and diffractive paths that link source and receiver.
Fabio Antonacci, Marco Foco, Augusto Sarti, Stefano Tubaro
MMSP4
2004 Invariant action classification with volumetric data
abstract
We propose an action recognition algorithm in which the image sequences capturing a moving human body produced by a significant number of cameras are first used to generate a volumetric representation of the body by means of volumetric intersection. Classification is then performed directly on 3D data, making the system inherently insensitive to viewpoint dependence and motion trajectory variability. Suitable features are extracted from the voxset approximating the body, and fed to a hidden Markov model to produce a finite-state description of the motion. The Kullback-Leibler distance is finally used to classify new sequences.
Fabio Cuzzolin, Augusto Sarti, Stefano Tubaro
MMSP3
2004 Detection of linear objects in GPR data
Andrea Dell'Acqua, Augusto Sarti, Stefano Tubaro, Luigi Zanzi
Signal Process.3
2003 Three-view camera calibration using geometric algebra
abstract
In a former work of ours C. Defferara et al. (2002), we proposed a new way to express and interpret the epipolar constraint using Geometric Algebra, and we derived from it a novel and efficient 2-view camera calibration technique. In this paper we extend this GA approach to the 3-view case. After expressing the trifocal constraint in terms of bivectors and trivectors, we provide an alternative geometric interpretation of the coefficients of the trifocal tensor. On the basis of that, we propose a novel solution for the simultaneous determination of the focal lengths of the cameras and the rigid motion between three views.
Andrea Dell'Acqua, Augusto Sarti, Stefano Tubaro
ICIP (1)3
2003 A space domain approach for multiple description video coding
abstract
In this paper we address the problem of video transmission over unreliable networks using multiple description (MD) coding approach. Recent literature on MD techniques has shown the success of using motion compensation prediction scheme. Even though these MD systems have proved to be robust to error environment, they are forced to use only two descriptions. For this reason this paper proposes two architectures of MD video coders. The first, called DC-MDVC, is based on the classical MD video codec architecture and the principle contribution is the use of a new MD algorithm based on a spatial domain polyphase down-sampling technique. The second architecture, called IF-MDVC, proposes multi-level scalability generating a flexible number of descriptions. The results of both architectures are interesting and the second architecture, the subject of the current research, appears to perform impressively.
Nicola Franchi, Marco Fumagalli, Rosa Lancini, Stefano Tubaro
ICIP (3)4
2003 Video indexing application based on watermarking using turbocode and side information
Agnese Buccoliero, Rosa Lancini, Francesco Mapelli, Stefano Tubaro
VCIP4
2002 Robust real-time intrusion detection with fuzzy classification
abstract
We propose a novel system for indoor video surveillance. It is able to detect and track moving objects even in the presence of significant variations of scene illumination. After a preliminary analysis and clustering of temporal changes in the video sequence, the algorithm performs a classification based on fuzzy logic, aimed at identifying moving regions that really correspond to unexpected objects in the scene. The proposed approach tends to discard shadows, reflections and luminance profile changes due to illumination variations. One key feature of our system is its modest computation complexity, which allows it to operate in real-time on a common PC platform. The system has been tested on a wide variety of situations, proving its effectiveness and robustness.
Giovanni Milanesi, Augusto Sarti, Stefano Tubaro
ICIP (3)3
2002 New perspectives on camera calibration using geometric algebra
abstract
We propose a new approach to the camera self-calibration problem, based on geometric algebra. After an introduction on the adopted Clifford algebra framework, we provide new insight on the epipolar constraint as defined in terms of bivectors. On the basis of that, we propose a novel solution for the simultaneous determination of the focal lengths of the cameras and the rigid motion between views.
Augusto Sarti, Claudio Defferara, Fabio Negroni, Stefano Tubaro
ICIP (2)4
2002 A robust video watermarking technique for compression and transcoding processing
abstract
We present a video watermarking technique proposed as an indexing technique for video sequences. The watermarking approach is based on an algorithm working in the spatial video domain, in order to reduce the computational cost, while the invisibility is obtained by masks computed in space and time domains. Finally, to improve robustness we work with different error correcting codes and we test their performances. The proposed scheme has been implemented and tested. Our tests are mainly focused on compression and transcoding attacks that are the most important operations that the video suffers in a journalistic report production. The robustness is increased by using an error correction code. The obtained results show that this is a good approach.
Rosa Lancini, Francesco Mapelli, Stefano Tubaro
ICME (1)3
2002 Image-based surface modeling: a multi-resolution approach
Augusto Sarti, Stefano Tubaro
Signal Process.2
2002 Detection and characterisation of planar fractures using a 3D Hough transform
Augusto Sarti, Stefano Tubaro
Signal Process.2
2001 Image-based object modeling: a multiresolution level-set approach
abstract
We propose an image-based 3D modeling method based on a multi-resolution evolution of a level-set of a volumetric function, steered by the texture mismatch between views. A key feature of this approach is that it operates in an adaptive multi-resolution fashion, which boosts up the computational efficiency.
A. Colosimo, Augusto Sarti, Stefano Tubaro
ICIP (2)3
2001 A multiresolution level-set approach to surface fusion
abstract
We propose a novel solution to the problem of combining a number of partial reconstructions through a process of 3D "patchworking". Our method is able to seamlessly "sew" the surface overlaps together, and to reasonably "mend" all the holes that remain after surface assembly, which usually correspond to the non-visible portions of the object surface. The approach is based on the temporal evolution of the zero level set volumetric function, driven by surface curvature and distance from data and works entirely in a multi-resolution fashion.
Augusto Sarti, Stefano Tubaro
ICIP (2)2
2001 Audiowatermarking Based Technologies for Automatic Identification of Musical Pieces in Audiotracks
abstract
In this paper we propose a fast and robust watermarking scheme for digital audio signals that allows automatic identification of musical pieces transmitted in TV broadcasting programs. Usually watermark is used as a way for hiding information on digital media. The watermarked information may be used to allow copyright protection or user and media identification. In our application the watermark must be, obviously, imperceptible to the users to mantain the quality of the source, should be robust to standard TV and radio editing and its processing has to require a very low complexity. This last item is essential to allow a software real-time implementation of the insertion and detection of watermarks using only a minimum amount of the computation power of a modern PC. Simulation results obtained on a long set of subjective test prove the quality and the robustness of the proposed approach.
Giuseppe Caccia, Rosa Lancini, Francesco Mapelli, Stefano Tubaro
ICME4
2001 Watermarking for musical pieces indexing used in automatic cue sheet generation systems
abstract
We propose a fast and robust watermarking scheme for digital audio signals that allows automatic identification of musical pieces transmitted in TV broadcast programs. Usually the watermark is used as a way for hiding information on digital media. The watermarked information may allow copyright protection or user identification and media indexing. In our application the watermark must be, obviously, imperceptible to the users to maintain the quality of the source, should be robust to standard TV and radio editing and its processing has to require a very low complexity. This last item is essential to allow a software real-time implementation of the insertion and detection of watermark using only a minimum amount of the computation power of a modern PC. In the proposed algorithm we divide the audio signal in frames, on each frame a masking function is computed based on the human auditory system (HAS), and the watermark is embedded using a spread spectrum technique. Simulation results obtained on a long set of subjective tests prove the quality and the robustness of the proposed approach.
Giuseppe Caccia, Rosa Lancini, M. Ludovico, Francesco Mapelli, Stefano Tubaro
MMSP5
2000 A Blind & Readable Watermarking Technique for Color Images
abstract
The paper presents a wavelet domain watermarking technique for color images. The characteristics of this transform domain are well suited for masking consideration, due to its good localization in both space and frequency. The different color sensitivity can be exploited in order to increase the watermarking power, and therefore its reliability. Thanks to the multi-resolution nature of the discrete wavelet transform (DWT) the watermark detection can be obtained iteratively. The technique has been successfully evaluated against attacks such as JPEG and JPEG-2000 compression, filtering, cropping, dithering and digital to analog and analog to digital conversions.
Marcello Caramma, Rosa Lancini, Francesco Mapelli, Stefano Tubaro
ICIP4
2000 Multi-Resolution Corner Detection
abstract
In this paper we propose a novel technique for wavelet-based multi-resolution corner detection. The method is based on a search of the zero crossings of the Laplacian at different scales along the line that describes the trajectory of the maxima as the scale varies. The proposed techniques allow us to achieve sub-pixel accuracy and provide us with useful scale information on the detected features.
Federico Pedersini, Elena Pozzoli, Augusto Sarti, Stefano Tubaro
ICIP4
2000 Multi-Resolution Area Matching
abstract
We present a general and robust approach to the problem of close-range partial 3D reconstruction of objects from multi-resolution texture matching. The method is based on the progressive refinement of a parametric surface, which is described using an increasing number of radial functions.
Federico Pedersini, Augusto Sarti, Stefano Tubaro
ICIP3
2000 Multicamera motion estimation for high-accuracy 3D reconstruction
Federico Pedersini, Pasquale Pigazzini, Augusto Sarti, Stefano Tubaro
Signal Process.4
2000 Visible surface reconstruction with accurate localization of object boundaries
abstract
A common limitation of many techniques for 3-D reconstruction from multiple perspective views is the poor quality of the results near the object boundaries. The interpolation process applied to "unstructured" 3-D data ("clouds" of non-connected 3-D points) plays a crucial role in the global quality of the 3-D reconstruction. We present a method for interpolating unstructured 3-D data, which is able to perform a segmentation of such data into different data sets that correspond to different objects. The algorithm is also able to perform an accurate localization of the boundaries of the objects. The method is based on an iterative optimization algorithm. As a first step, a set of surfaces and boundary curves are generated for the various objects. Then, the edges of the original images are used for refining such boundaries as best as possible. Experimental results with real data are presented for proving the effectiveness of the proposed algorithm.
Federico Pedersini, Augusto Sarti, Stefano Tubaro
IEEE Trans. Circuits Syst. Video Technol.3
1999 Estimation of Radiometric Parameters for a Realistic Rendering of 3D Models
abstract
In order to obtain realistic results in the rendering of a 3D model, we need accurate information on both structure (shape) and radiometric properties (reflectivity) of its surfaces. In this paper, we approach the problem of correctly estimating the radiometric characteristics of an imaged surface through the analysis of several of its views. The images are acquired with one or more CCD cameras that move around the object, and the 3D model is assumed available (it can be constructed from the available images). The adopted reflectivity model is non-Lambertian as it takes into account both the diffuse lobe and the specular lobe. The method implements a robust technique which is able to estimate the reflectivity parameters of the object's surfaces even in the presence of modeling imperfections and/or when the relative camera-object position and orientation are only approximately known.
Federico Pedersini, Luca Piccarreta, Augusto Sarti, Stefano Tubaro
ICIP (4)4
1999 Accurate and simple geometric calibration of multi-camera systems
Federico Pedersini, Augusto Sarti, Stefano Tubaro
Signal Process.3
1998 A Moving Object Identification Algorithm for Image Sequence Interpolation
abstract
An effective motion compensated image interpolation technique is presented. The proposed algorithm has been developed for the interpolation of missing frames in an image sequence and it is based on two principal elements. The first is a region-based motion estimation and representation technique, which combines, for each known image, an approximate initial motion field estimate and an initial image over-segmentation (obtained using both luminance and chrominance information) to produce a very accurate affine-model regularized motion field. The proposed algorithm uses a robust identification technique for the estimation of the motion parameters associated to each region. The second important element is the use of an interpolation technique, that starting from the available motion information defines, for each point of the image to be interpolated a suitable reconstruction strategy taking into account the different situations that can appear (for example a moving object can occlude either a stationary background or an another object and so on). The principal aim of the proposed algorithm is the reconstruction of the missing frames in a reasonable way, without introducing significant artifacts and assuring a pleasant representation of object displacement.
Rosa Lancini, M. Ripamonti, P. Vicari, Marcello Caramma, Stefano Tubaro
ICIP (2)5
1998 Accurate Feature Detection and Matching for the Tracking of Calibration Parameters in Multi-Camera Acquisition Systems
abstract
The 3D reconstruction's quality of multiple-camera acquisition systems is strongly influenced by the accuracy of the camera calibration procedure. The acquisition of long sequences is, in fact, very sensitive to mechanical shocks, vibrations and thermal changes on cameras and supports, as they could result in a significant drift of the camera parameters. In this paper we propose a technique which is able to keep track of the camera parameters and, whenever possible, to correct them accordingly. This technique does not need any a-priori knowledge or test objects to be placed in the scene, but exploits features that are already present in the scene itself. In fact it performs an accurate detection, matching and back-projection of luminance corners and spots in the scene space. Experimental results on real sequences are reported in order to prove the ability of the proposed technique to detect a change in the calibration and to re-calibrate the camera setup with an accuracy that depends on the number of available feature points.
Federico Pedersini, Augusto Sarti, Stefano Tubaro
ICIP (2)3
1998 Combined Surface Interpolation and Object Segmentation for Automatic 3-D Scene Reconstruction
abstract
A common limitation of many techniques for 3D reconstruction from multiple perspective views is the poor quality of the results near the object boundaries. The interpolation process applied to "unstructured" 3D data ("clouds" of non-connected 3D points) plays a crucial role in the global quality of the 3D reconstruction. We present a method for interpolating unstructured 3D data, which is able to perform a segmentation of such data into different data sets that correspond to different objects. The algorithm is also able to perform an accurate localization of the boundaries of the objects. The method is based on an iterative optimization algorithm. As a first step, a set of surfaces and boundary curves are generated for the various objects. Then, the edges of the original images are used for refining such boundaries as best as possible. Experimental results with real data are presented for proving the effectiveness of the proposed algorithm.
Federico Pedersini, Augusto Sarti, Stefano Tubaro
ICIP (2)3
1998 3D area matching with arbitrary multiview geometry
Federico Pedersini, Pasquale Pigazzini, Augusto Sarti, Stefano Tubaro
Signal Process. Image Commun.4
1998 Improving the performance of edge localization techniques through error compensation
Federico Pedersini, Augusto Sarti, Stefano Tubaro
Signal Process. Image Commun.3
1997 Egomotion Estimation of a Multicamera System Through Line Correspondence
abstract
In this paper we propose a method for estimating the egomotion of a calibrated multi-camera system from an analysis of the luminance edges. The method works entirely in the 3D space as all edges of each one set of views are previously localized, matched and back-projected onto the object space. In fact, it searches for the rigid motion that best merges the sets of 3D contours extracted from each one of the multi-views. The method uses both straight and curved 3D contours.
Federico Pedersini, Augusto Sarti, Stefano Tubaro
ICIP (2)3
1997 Robust Area Matching
abstract
We propose a general and robust solution to the problem of close-range 3D reconstruction of objects from stereo correspondence of luminance profiles. The method does not require a particular camera geometry, and can be implemented with an arbitrary number of CCD cameras. Its robustness can be mainly attributed to the physicality of the matching process, which is performed in the 3D space, while taking both geometric and radiometric distortions into account. Extensive tests have been performed over a variety of real scenes, using a calibrated trinocular camera system.
Federico Pedersini, Augusto Sarti, Stefano Tubaro
ICIP (2)3
1997 Estimation and Compensation of Subpixel Edge Localization Error
abstract
We propose and analyze a method for improving the performance of subpixel edge localization (EL) techniques through compensation of the systematic portion of the localization error. The method is based on the estimation of the EL characteristic through statistical analysis of a test image and is independent of the EL technique in use.
Federico Pedersini, Augusto Sarti, Stefano Tubaro
IEEE Trans. Pattern Anal. Mach. Intell.3
1997 Joint automatic design of prefilter and decimation grid for the reduction of spectral redundancy in 2D digital signals
Federico Pedersini, Augusto Sarti, Stefano Tubaro
Signal Process.3
1997 An image coding scheme based on image projections and geometrical VQ
Rosa Lancini, Stefano Tubaro
Signal Process. Image Commun.2
1996 Combined motion and edge analysis for a layer-based representation of image sequences
abstract
We propose a method for the motion-congruent segmentation of image sequences based on both motion field and luminance information. In order to do so, the affine motion models are determined by analyzing the motion field through a clustering procedure, while their regions of validity are determined by an MRF-based region estimator. This last block performs a pixel-wise re-assignment of a limited number of affine models (those determined via clustering).
Federico Pedersini, Augusto Sarti, Stefano Tubaro
ICIP (1)3
1996 3D motion estimation of a trinocular system for a full-3D object reconstruction
abstract
Motion estimation of a calibrated multi-ocular acquisition system can be performed either on two-dimensional data, by applying a rigidity constraint to a set of matched points on monocular views, or on three-dimensional data, by determining the rigid motion that best overlaps homologous 3D points obtained through stereo matching and back-projection. We propose, analyze and compare these two possible solutions, and present a low-cost high-accuracy full-3D reconstruction system based on multiple trinocular views at standard TV resolution.
Federico Pedersini, Augusto Sarti, Stefano Tubaro
ICIP (2)3
1995 Synthesis of virtual views using non-Lambertian reflectivity models and stereo matching
abstract
A technique for the synthesis of virtual views of a 3D scene, starting from images taken by a calibrated multicamera system, is proposed and tested. Surface interpolation is performed over a set of 3D edges, computed with stereometric algorithms, and additional curvature-tuning points, scattered where the reflectivity model is sufficiently reliable. The 3D coordinates of the tuning points are computed by minimizing the MSE between the available real views and the corresponding synthetic views. Synthesis is finally carried out by reprojecting on the new image plane the estimated object surface over which texture-mapping of the reflection-corrected luminance function has been performed. Texture correction, which uses an estimate of a non-Lambertian reflectivity model, is done in such a way to simulate the migration of reflections due to the change of viewpoint. The technique has been tested on real images, producing realistic synthesized views.
Federico Pedersini, Augusto Sarti, Stefano Tubaro
ICIP3
1995 Multistage motion estimation for image interpolation
Pierangelo Migliorati, Stefano Tubaro
Signal Process. Image Commun.2
1995 Adaptive vector quantization for picture coding using neural networks
abstract
The paper presents applications of neural network algorithms to the design of an adaptive vector quantizer. Vector quantization has been applied to the problem of displaying natural images with a reduced set of colors (colormap) and to the interframe coding of image sequences. The first step was to test classical Linde Buzo Gray (LGB), self organizing feature maps (SOFM) and Competitive Learning (CL) algorithms for the codebook design. The best results for the reconstructed quality image and the computational time are obtained using a CL algorithm with a new initialization strategy that solves the problem of underutilized nodes. An adaptive vector quantization algorithm is proposed and tested in a motion compensated image coder. The results of the simulations are very promising. In fact the coder performance, compared with that using a fixed VQ, is considerably improved and the subjective quality of the coded images is much better than that obtained using standard vector quantization, especially when rapid motion is present in the scene.>
Rosa Lancini, Stefano Tubaro
IEEE Trans. Commun.2
1994 Image coding with overlapped projection and pyramid vector quantization
abstract
A new effective still image coding scheme based on overlapped projections and pyramid vector quantization is presented. In standard image coding the spatial redundancy of the images is reduced subdividing the incoming data into small blocks and applying to these a linear transformation (DCT) to obtain highly independent coefficients, that are scalar quantized. This block by block coding takes advantage of the spatial correlation but it induces visually annoying blocking artifacts (especially at high compression ratios). The authors resent a different coding scheme where the image is subdivided into circular partially overlapped regions and then the projections along predefined directions are determined and vector quantized. Simulations show that the projection technique allows the detection of the most important information locally present on the image while the use of circular overlapped regions reduces blocking artifacts.>
Rosa Lancini, Emanuele Marconetti, Stefano Tubaro
ICASSP (5)3
1994 Progressive Image Transmission Based on Image Projections
abstract
We present a progressive image transmission coder, where geometrical vector quantization is applied to overlapped image projections. Progressive image transmission is widely used in many applications, because it offers successive improved reconstructions of an image. It is possible to have a rough version of information available in a very short time and, then, the viewer can decide whether further transmission is necessary or not. In the proposed coding scheme the still image is subdivided into circular partially overlapped regions and, then, some projections along predefined directions are determined and vector quantized by using geometrical VQ. The utilization of overlapped circular regions exploits the spatial redundancy present on the image, as the block subdivision does, and at the same time, it avoids the problem of blocking artifacts. The projection technique allows the detection of the most important information locally present on the image.>
Rosa Lancini, Stefano Tubaro
ICIP (3)2
1993 Semantic segmentation applied to image interpolation in the case of camera panning and zooming
Pierangelo Migliorati, Federico Pedersini, L. Sorcinelli, Stefano Tubaro
ICASSP (5)4
1992 Neural network approach for adaptive vector quantization of images
abstract
The problem of adaptive vector quantization (VQ) for image sequence coding is addressed. The goal of this work is to overcome the limits that reduce the possibility of a hardware implementation of the schemes already presented in the literature. The limits principally concern the complexity of the implementation in real time of the classical Linde-Buzo-Gray (LBG) algorithm, which is necessary to introduce some adaptive capabilities in a VQ. Neural network methods, because of their fast codebook design, seem an interesting alternative to solve this problem. The authors have used an unsupervised neural network approach to introduce adaptivity in a vector quantizer by using a codebook replenishment method. The proposed adaptive VQ algorithm has been tested in a motion compensated interframe image coding scheme. The results of the simulations are very promising. A considerable rise of the coder performance with respect to the use of fixed VQ has been obtained.>
Rosa Lancini, F. Perego, Stefano Tubaro
ICASSP3
1991 A two layers video coding scheme for ATM networks
Stefano Tubaro
Signal Process. Image Commun.1
1990 A hybrid image coder with vector quantizer
Stefano Tubaro
Signal Process. Image Commun.1
1990 Motion compensated image interpolation
abstract
Field skipping is a variable technique for reducing drastically the bit rate necessary to transmit a television signal. All fields have to be reconstructed at the receiver end, but nearest-neighbor or linear interpolations give poor performances when significant scene activity is present; therefore, some motion compensation scheme is mandatory. Although any of the algorithms proposed for motion estimation can be used for motion compensated interpolation, in this paper only the pel-recursive algorithm is considered. Past experience is briefly summarized and, based on criticism, improvements to the basic algorithm are proposed that seem to lead to an estimated motion field that is accurate enough. Multiple recursions and appropriate selection rules are proposed to counteract streaking effects of the recursive algorithm when it crosses image boundaries. The estimation of a forward and a backward motion field is used to identify image areas where occlusion effects are present. Experimental results are presented that seem to indicate that the proposed interpolation scheme has good performance.>
Ciro Cafforio, Fabio Rocca, Stefano Tubaro
IEEE Trans. Commun.3
1986 An audio computer interface: A case study of structured electronic equipment design
Sergio Brofferio, Maurizio Piacentini, Stefano Tubaro
Microprocessing and Microprogramming3