Marco Tagliasacchi

dblp:10/166 · DBLP profile ↗
← Back
134ranked-venue papers
23as first author
18since 2021 · last 2025
0000-0002-7682-6795ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 100 · 20 first-author · 12 since 2021Artificial intelligence and machine learning · 17 · 2 first-author · 11 since 2021Databases, data management, data science and information retrieval · 15Computer networks · 7Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorSoftware engineering, systems software and programming languages · 2Security and privacy · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 MAD Speech: Measures of Acoustic Diversity of Speech
abstract
Matthieu Futeral, Andrea Agostinelli, Marco Tagliasacchi, Neil Zeghidour, Eugene Kharitonov. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Matthieu Futeral, Andrea Agostinelli, Marco Tagliasacchi, Neil Zeghidour, Eugene Kharitonov
NAACL (Long Papers)3
2024 Data Summarization via Bilevel Optimization
abstract
The increasing availability of massive data sets poses various challenges for machine learning. Prominent among these is learning models under hardware or human resource constraints. In such resource-constrained settings, a simple yet powerful approach is operating on small subsets of the data. Coresets are weighted subsets of the data that provide approximation guarantees for the optimization objective. However, existing coreset constructions are highly model-specific and are limited to simple models such as linear regression, logistic regression, and k-means. In this work, we propose a generic coreset construction framework that formulates the coreset selection as a cardinality-constrained bilevel optimization problem. In contrast to existing approaches, our framework does not require model-specific adaptations and applies to any twice differentiable model, including neural networks. We show the effectiveness of our framework for a wide range of models in various settings, including training non-convex models online and batch active learning.
Zalan Borsos, Mojmír Mutný, Marco Tagliasacchi, Andreas Krause 0001
J. Mach. Learn. Res.3
2023 LMCodec: A Low Bitrate Speech Codec with Causal Transformer Models
abstract
We introduce LMCodec, a causal neural speech codec that provides high quality audio at very low bitrates. The backbone of the system is a causal convolutional codec that encodes audio into a hierarchy of coarse-to-fine tokens using residual vector quantization. LMCodec trains a Transformer language model to predict the fine tokens from the coarse ones in a generative fashion, allowing for the transmission of fewer codes. A second Transformer predicts the uncertainty of the next codes given the past transmitted codes, and is used to perform conditional entropy coding. A MUSHRA subjective test was conducted and shows that the quality is comparable to reference codecs at higher bitrates. Example audio is available at https://mjenrungrot.github.io/chrome-media-audio-papers/publications/lmcodec.
Teerapat Jenrungrot, Michael Chinen, W. Bastiaan Kleijn, Jan Skoglund, Zalan Borsos, Neil Zeghidour, Marco Tagliasacchi
ICASSP7
2023 Disentangling Speech from Surroundings with Neural Embeddings
abstract
We present a method to separate speech signals from noisy environments in the embedding space of a neural audio codec. We introduce a new training procedure that allows our model to produce structured encodings of audio waveforms given by embedding vectors, where one part of the embedding vector represents the speech signal, and the rest represent the environment. We achieve this by partitioning the embeddings of different input waveforms and training the model to faithfully reconstruct audio from mixed partitions, thereby ensuring each partition encodes a separate audio attribute. As use cases, we demonstrate the separation of speech from background noise or from reverberation characteristics. Our method also allows for targeted adjustments of the audio output characteristics.
Ahmed Omran, Neil Zeghidour, Zalan Borsos, Félix de Chaumont Quitry, Malcolm Slaney, Marco Tagliasacchi
ICASSP6
2023 TokenSplit: Using Discrete Speech Representations for Direct, Refined, and Transcript-Conditioned Speech Separation and Recognition
Hakan Erdogan, Scott Wisdom, Xuankai Chang, Zalan Borsos, Marco Tagliasacchi, Neil Zeghidour, John R. Hershey
INTERSPEECH5
2023 Real time spectrogram inversion on mobile phone
Oleg Rybakov, Marco Tagliasacchi, Yunpeng Li 0008, Liyang Jiang, Fadi Biadsy
INTERSPEECH2
2023 Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision
abstract
Abstract We introduce SPEAR-TTS, a multi-speaker text-to-speech (TTS) system that can be trained with minimal supervision. By combining two types of discrete speech representations, we cast TTS as a composition of two sequence-to-sequence tasks: from text to high-level semantic tokens (akin to “reading”) and from semantic tokens to low-level acoustic tokens (“speaking”). Decoupling these two tasks enables training of the “speaking” module using abundant audio-only data, and unlocks the highly efficient combination of pretraining and backtranslation to reduce the need for parallel data when training the “reading” component. To control the speaker identity, we adopt example prompting, which allows SPEAR-TTS to generalize to unseen speakers using only a short sample of 3 seconds, without any explicit speaker representation or speaker labels. Our experiments demonstrate that SPEAR-TTS achieves a character error rate that is competitive with state-of-the-art methods using only 15 minutes of parallel data, while matching ground-truth speech in naturalness and acoustic quality.
Eugene Kharitonov, Damien Vincent, Zalan Borsos, Raphaël Marinier, Sertan Girgin, Olivier Pietquin, Matthew Sharifi, Marco Tagliasacchi, Neil Zeghidour
Trans. Assoc. Comput. Linguistics8
2023 AudioLM: A Language Modeling Approach to Audio Generation
abstract
We introduce AudioLM, a framework for high-quality audio generation with long-term consistency. AudioLM maps the input audio to a sequence of discrete tokens and casts audio generation as a language modeling task in this representation space. We show how existing audio tokenizers provide different trade-offs between reconstruction quality and long-term structure, and we propose a hybrid tokenization scheme to achieve both objectives. Namely, we leverage the discretized activations of a masked language model pre-trained on audio to capture long-term structure and the discrete codes produced by a neural audio codec to achieve high-quality synthesis. By training on large corpora of raw audio waveforms, AudioLM learns to generate natural and coherent continuations given short prompts. When trained on speech, and without any transcript or annotation, AudioLM generates syntactically and semantically plausible speech continuations while also maintaining speaker identity and prosody for unseen speakers. Furthermore, we demonstrate how our approach extends beyond speech by generating coherent piano music continuations, despite being trained without any symbolic representation of music.
Zalan Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matthew Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, Neil Zeghidour
IEEE ACM Trans. Audio Speech Lang. Process.10
2022 SpeechPainter: Text-conditioned Speech Inpainting
Zalan Borsos, Matthew Sharifi, Marco Tagliasacchi
INTERSPEECH3
2022 Text-Driven Separation of Arbitrary Sounds
Kevin Kilgour, Beat Gfeller, Aren Jansen, Scott Wisdom, Marco Tagliasacchi
INTERSPEECH6
2022 CycleGAN-based Unpaired Speech Dereverberation
abstract
Typically, neural network-based speech dereverberation models are trained on paired data, composed of a dry utterance and its corresponding reverberant utterance.The main limitation of this approach is that such models can only be trained on large amounts of data and a variety of room impulse responses when the data is synthetically reverberated, since acquiring real paired data is costly.In this paper we propose a CycleGAN-based approach that enables dereverberation models to be trained on unpaired data.We quantify the impact of using unpaired data by comparing the proposed unpaired model to a paired model with the same architecture and trained on the paired version of the same dataset.We show that the performance of the unpaired model is comparable to the performance of the paired model on two different datasets, according to objective evaluation metrics.Furthermore, we run two subjective evaluations and show that both models achieve comparable subjective quality on the AMI dataset, which was not seen during training.
Hannah Muckenhirn, Aleksandr Safin, Hakan Erdogan, Félix de Chaumont Quitry, Marco Tagliasacchi, Scott Wisdom, John R. Hershey
INTERSPEECH5
2022 SoundStream: An End-to-End Neural Audio Codec
abstract
We presentSoundStream, a novel neural audio codec that can efficiently compress speech, music and general audio at bitrates normally targeted by speech-tailored codecs.SoundStreamrelies on a model architecture composed by a fully convolutional encoder/decoder network and a residual vector quantizer, which are trained jointly end-to-end. Training leverages recent advances in text-to-speech and speech enhancement, which combine adversarial and reconstruction losses to allow the generation of high-quality audio content from quantized embeddings. By training with structured dropout applied to quantizer layers, a single model can operate across variable bitrates from 3 kbps to 18 kbps, with a negligible quality loss when compared with models trained at fixed bitrates. In addition, the model is amenable to a low latency implementation, which supports streamable inference and runs in real time on a smartphone CPU. In subjective evaluations using audio at 24 kHz sampling rate,SoundStreamat 3 kbps outperforms Opus at 12 kbps and approaches EVS at 9.6 kbps. Moreover, we are able to perform joint compression and enhancement either at the encoder or at the decoder side with no additional latency, which we demonstrate through background noise suppression for speech.
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, Marco Tagliasacchi
IEEE ACM Trans. Audio Speech Lang. Process.5
2021 Micaugment: One-Shot Microphone Style Transfer
abstract
A crucial aspect for the successful deployment of audio-based models "in-the-wild" is the robustness to the transformations introduced by heterogeneous acquisition conditions. In this work, we propose a method to perform one-shot microphone style transfer. Given only a few seconds of audio recorded by a target device, MicAugment identifies the transformations associated with the input acquisition pipeline and uses the learned transformations to synthesize audio as if it were recorded under the same conditions as the target audio. We show that our method can successfully apply the style transfer to real audio and that it significantly increases model robustness when used as data augmentation in the downstream tasks.
Zalan Borsos, Yunpeng Li 0008, Beat Gfeller, Marco Tagliasacchi
ICASSP4
2021 Semi-Supervised Batch Active Learning Via Bilevel Optimization
abstract
Active learning is an effective technique for reducing the labeling cost by improving data efficiency. In this work, we propose a novel batch acquisition strategy for active learning in the setting where the model training is performed in a semi-supervised manner. We formulate our approach as a data summarization problem via bilevel optimization, where the queried batch consists of the points that best summarize the unlabeled data pool. We show that our method is highly effective in keyword detection tasks in the regime when only few labeled samples are available.
Zalan Borsos, Marco Tagliasacchi, Andreas Krause 0001
ICASSP2
2021 One-Shot Conditional Audio Filtering of Arbitrary Sounds
abstract
We consider the problem of separating a particular sound source from a single-channel mixture, based on only a short sample of the target source (from the same recording). Using SoundFilter, a wave-to-wave neural network architecture, we can train a model without using any sound class labels. Using a conditioning encoder model which is learned jointly with the source separation network, the trained model can be "configured" to filter arbitrary sound sources, even ones that it has not seen during training. Evaluated on the FSD50k dataset, our model obtains an SI-SDR improvement of 9.6 dB for mixtures of two sounds. When trained on Librispeech, our model achieves an SI-SDR improvement of 14.0 dB when separating one voice from a mixture of two speakers. Moreover, we show that the representation learned by the conditioning encoder clusters acoustically similar sounds together in the embedding space, even though it is trained without using any labels.
Beat Gfeller, Dominik Roblek, Marco Tagliasacchi
ICASSP3
2021 Real-Time Speech Frequency Bandwidth Extension
abstract
In this paper we propose a lightweight model for frequency bandwidth extension of speech signals, increasing the sampling frequency from 8kHz to 16kHz while restoring the high frequency content to a level almost indistinguishable from the 16kHz ground truth. The model architecture is based on SEANet (Sound EnhAncement Network), a wave-to-wave fully convolutional model, which uses a combination of feature losses and adversarial losses to reconstruct an enhanced version of the input speech. In addition, we propose a variant of SEANet that can be deployed on-device in streaming mode, achieving an architectural latency of 16ms. When profiled on a single core of a mobile CPU, processing one 16ms frame takes only 1.5ms. The low latency makes it viable for bi-directional voice communication systems.
Yunpeng Li 0008, Marco Tagliasacchi, Oleg Rybakov, Victor Ungureanu, Dominik Roblek
ICASSP2
2021 LEAF: A Learnable Frontend for Audio Classification
Neil Zeghidour, Olivier Teboul, Félix de Chaumont Quitry, Marco Tagliasacchi
ICLR4
2021 Reconstructing Speech From CNN Embeddings
abstract
The complete understanding of the decision-making process of Convolutional Neural Networks (CNNs) is far from being fully reached. Many researchers proposed techniques to interpret what a network actually “learns” from data. Nevertheless many questions still remain unanswered. In this work we study one aspect of this problem by reconstructing speech from the intermediate embeddings computed by a CNNs. Specifically, we consider a pre-trained network that acts as a feature extractor from speech audio. We investigate the possibility of inverting these features, reconstructing the input signals in a black-box scenario, and quantitatively measure the reconstruction quality by measuring the word-error-rate of an off-the-shelf ASR model. Experiments performed using two different CNN architectures trained for six different classification tasks, show that it is possible to reconstruct time-domain speech signals that preserve the semantic content, whenever the embeddings are extracted before the fully connected layers.
Luca Comanducci, Paolo Bestagini, Marco Tagliasacchi, Augusto Sarti, Stefano Tubaro
IEEE Signal Process. Lett.3
2020 Pitch Estimation Via Self-Supervision
abstract
We present a method to estimate the fundamental frequency in monophonic audio, often referred to as pitch estimation. In contrast to existing methods, our neural network can be fully trained only on unlabeled data, using self-supervision. A tiny amount of labeled data is needed solely for mapping the network outputs to absolute pitch values. The key to this is the observation that if one creates two examples from one original audio clip by pitch shifting both, the difference between the correct outputs is known, without even knowing the actual pitch value in the original clip. Somewhat surprisingly, this idea combined with an auxiliary reconstruction loss allows training a pitch estimation model. Our results show that our pitch estimation method obtains an accuracy comparable to fully supervised models on monophonic audio, without the need for large labeled datasets. In addition, we are able to train a voicing detection output in the same model, again without using any labels.
Beat Gfeller, Christian Havnø Frank, Dominik Roblek, Matthew Sharifi, Marco Tagliasacchi, Mihajlo Velimirovic
ICASSP5
2020 Towards Learning a Universal Non-Semantic Representation of Speech
abstract
The ultimate goal of transfer learning is to reduce labeled data requirements by exploiting a pre-existing embedding model trained for different datasets or tasks. The visual and language communities have established benchmarks to compare embeddings, but the speech community has yet to do so. This paper proposes a benchmark for comparing speech representations on non-semantic tasks, and proposes a representation based on an unsupervised triplet-loss objective. The proposed representation outperforms other representations on the benchmark, and even exceeds state-of-the-art performance on a number of transfer learning tasks. The embedding is trained on a publicly available dataset, and it is tested on a variety of low-resource downstream tasks, including personalization tasks and medical domain. The benchmark, models, and evaluation code are publicly released.
Joel Shor, Aren Jansen, Ronnie Maor, Oran Lang, Omry Tuval, Félix de Chaumont Quitry, Marco Tagliasacchi, Ira Shavitt, Dotan Emanuel, Yinnon Haviv
INTERSPEECH7
2020 SEANet: A Multi-Modal Speech Enhancement Network
abstract
We explore the possibility of leveraging accelerometer data to perform speech enhancement in very noisy conditions. Although it is possible to only partially reconstruct user's speech from the accelerometer, the latter provides a strong conditioning signal that is not influenced from noise sources in the environment. Based on this observation, we feed a multi-modal input to SEANet (Sound EnhAncement Network), a wave-to-wave fully convolutional model, which adopts a combination of feature losses and adversarial losses to reconstruct an enhanced version of user's speech. We trained our model with data collected by sensors mounted on an earbud and synthetically corrupted by adding different kinds of noise sources to the audio signal. Our experimental results demonstrate that it is possible to achieve very high quality results, even in the case of interfering speech at the same level of loudness. A sample of the output produced by our model is available at https://google-research.github.io/seanet/multimodal/speech.
Marco Tagliasacchi, Yunpeng Li 0008, Karolis Misiunas, Dominik Roblek
INTERSPEECH1
2020 Pre-Training Audio Representations With Self-Supervision
abstract
We explore self-supervision as a way to learn general purpose audio representations. Specifically, we propose two self-supervised tasks: Audio2Vec, which aims at reconstructing a spectrogram slice from past and future slices and TemporalGap, which estimates the distance between two short audio segments extracted at random from the same audio clip. We evaluate how the representations learned via self-supervision transfer to different downstream tasks, either training a task-specific linear classifier on top of the pretrained embeddings, or fine-tuning a model end-to-end for each downstream task. Our results show that the representations learned with Audio2Vec transfer better than those learned by fully-supervised training on Audioset. In addition, by fine-tuning Audio2Vec representations it is possible to outperform fully-supervised models trained from scratch on each task, when limited data is available, thus improving label efficiency.
Marco Tagliasacchi, Beat Gfeller, Félix de Chaumont Quitry, Dominik Roblek
IEEE Signal Process. Lett.1
2020 Multi-Task Adapters for On-Device Audio Inference
abstract
The deployment of deep networks on mobile devices requires to efficiently use the scarce computational resources, expressed as either available memory or computing cost. When addressing multiple tasks simultaneously, it is extremely important to share resources across tasks, especially when they all consume the same input data, e.g., audio samples captured by the on-board microphones. In this paper we propose a multi-task model architecture that consists of a shared encoder and multiple task-specific adapters. During training, we learn the model parameters as well as the allocation of the task-specific additional resources across both tasks and layers. A global tuning parameter can be used to obtain different multi-task network configurations finding the desired trade-off between cost and the level of accuracy across tasks. Our results show that this solution significantly outperforms a multi-head model baseline. Interestingly, we observe that the optimal resource allocation depends on both the task intrinsic characteristics as well as on the targeted cost measure (e.g., memory or computing cost).
Marco Tagliasacchi, Félix de Chaumont Quitry, Dominik Roblek
IEEE Signal Process. Lett.1
2020 SPICE: Self-Supervised Pitch Estimation
abstract
We propose a model to estimate the fundamental frequency in monophonic audio, often referred to as pitch estimation. We acknowledge the fact that obtaining ground truth annotations at the required temporal and frequency resolution is a particularly daunting task. Therefore, we propose to adopt a self-supervised learning technique, which is able to estimate pitch without any form of supervision. The key observation is that pitch shift maps to a simple translation when the audio signal is analysed through the lens of the constant-Q transform (CQT). We design a self-supervised task by feeding two shifted slices of the CQT to the same convolutional encoder, and require that the difference in the outputs is proportional to the corresponding difference in pitch. In addition, we introduce a small model head on top of the encoder, which is able to determine the confidence of the pitch estimate, so as to distinguish between voiced and unvoiced audio. Our results show that the proposed method is able to estimate pitch at a level of accuracy comparable to fully supervised models, both on clean and noisy audio samples, although it does not require access to large labeled datasets.
Beat Gfeller, Christian Havnø Frank, Dominik Roblek, Matthew Sharifi, Marco Tagliasacchi, Mihajlo Velimirovic
IEEE ACM Trans. Audio Speech Lang. Process.5
2016 Fast keypoint detection in video sequences
abstract
Several computer vision tasks exploit a succinct representation of the visual content in the form of sets of local features. Given an input image, feature extraction algorithms identify keypoints and assign to each of them a descriptor, based on the characteristics of the surrounding visual content. Several tasks might require local features to be extracted from a video sequence, on a frame-by-frame basis. Although temporal downsampling has been proven to be an effective solution for mobile augmented reality and visual search, high temporal resolution is a key requirement for time-critical applications such as object tracking, event recognition, pedestrian detection, surveillance. In recent years, more and more computationally efficient visual feature detectors and descriptors have been proposed. Nonetheless, such approaches are tailored to still images. In this paper we propose a fast keypoint detection algorithm for video sequences, that exploits the temporal coherence of the sequence of keypoints. According to the proposed method, each frame is preprocessed so as to identify the parts of the input frame for which keypoint detection and description need to be performed. Our experiments show that it is possible to achieve a reduction in computational time of up to 40%, without significantly affecting the task accuracy.
Luca Baroffio, Matteo Cesana, Alessandro Redondi, Marco Tagliasacchi, Stefano Tubaro
ICASSP4
2016 Image phylogeny tree reconstruction based on region selection
abstract
Nowadays, everyone can download, edit and republish any picture on the web, thus contributing to the diffusion of near-duplicate (ND) images. In order to gain an interesting insight on the way NDs are distributed online, recent works have focused on the reconstruction of the image phylogeny tree (IPT), i.e., an acyclic graph describing the genealogical relationship between ND image pairs. IPT reconstruction methods typically leverage the possibility of reconstructing one image from another one only if they are in parent-child relationship. However, as estimating the possible parent-child transformation is computationally expensive, usually a limited set of global editing operations is considered (i.e., compression, geometric and colour transformations applied to the whole image). However, in a real-world scenario it is customary to edit images also using local operations (e.g., logo insertion, object removal, splicing, etc.), which hinder the possibility of correctly estimating the parent-child relationship. In this paper, we propose an algorithm for IPT reconstruction that deals with the presence of local editing operations.
Paolo Bestagini, Marco Tagliasacchi, Stefano Tubaro
ICASSP2
2016 Crowdsourcing for top-K query processing over uncertain data
abstract
Both social media and sensing infrastructures are producing an unprecedented mass of data characterized by their uncertain nature, due to either the noise inherent in sensors or the imprecision of human contributions. Therefore query processing over uncertain data has become an active research field. In the well-known class of applications commonly referred to as “top-K queries”, the objective is to find the best K objects matching the user's information need, formulated as a scoring function over the objects' attribute values. If both the data and the scoring function are deterministic, the best K objects can be univocally determined and totally ordered so as to produce a single ranked result set (as long as ties are broken by some deterministic rule). However, in application scenarios involving uncertain data and fuzzy information needs, this does not hold: when either the attribute values or the scoring function are nondeterministic, there may be no consensus on a single ordering, but rather a space of possible orderings. To determine the correct ordering, one needs to acquire additional information so as to reduce the amount of uncertainty associated with the queried data and consequently the number of orderings in such a space. An emerging trend in data processing is crowdsourcing, defined as the systematic engagement of humans in the resolution of tasks through online distributed work. Our approach combines human and automatic computation in order to solve complex problems: when data ambiguity can be resolved by human judgment, crowdsourcing becomes a viable tool for converging towards a unique or at least less uncertain query result. The goal of this paper is to define and compare task selection policies for uncertainty reduction via crowdsourcing, with emphasis on the case of top-K queries.
Eleonora Ciceri, Piero Fraternali, Davide Martinenghi, Marco Tagliasacchi
ICDE4
2016 Multi-view coding and routing of local features in Visual Sensor Networks
abstract
Visual Sensor Networks (VSNs) have been recently used for implementing automatic visual analysis tasks where local image features, instead of images, are compressed and transmitted to a central controller. Such features may also be compressed in a multi-view fashion, exploiting the redundancy between overlapping views. In this paper we analyze the problem of multi-view coding and routing of features in VSNs. We empirically analyze the relationship between the bitrate reduction obtained with a practical multi-view local features encoder and several geometry-based, image-based and feature-based predictors. The purpose of this analysis is to identify the most accurate, yet compact predictor of the achievable compression efficiency when jointly encoding correlated streams of local features. Then, we propose a robust optimization framework that exploits the aforementioned predictors. The proposed mathematical problem maximizes the amount of data extracted from the VSN by properly routing the streams of features, subject to capacity, interference and energy constraints, explicitly considering the uncertainty in the compression efficiency estimation. Extensive experiments on simulated VSNs show that multi-view coding maximizes the amount of data extracted from camera nodes, while the robust optimization approach provides significant improvement in uncertain scenarios compared to the optimal solution of a deterministic approach.
Alessandro Redondi, Luca Baroffio, Matteo Cesana, Marco Tagliasacchi
INFOCOM4
2016 EZ-VSN: An Open-Source and Flexible Framework for Visual Sensor Networks
abstract
We present a complete, open-source framework for rapid experimentation of visual sensor network (VSN) solutions. From the software point of view, we base our architecture on open-source and widely known C++ libraries to provide the basic image processing and networking primitives. The resulting system can be leveraged to create different types of VSNs, characterized by the presence of multiple cameras, relays and cooperator nodes, and can be run on any Linux-based hardware platform, such as the BeagleBone Black. To demonstrate the flexibility of the proposed framework, we describe two different application scenarios typical of VSNs, namely object recognition and parking monitoring. The framework is then used to evaluate the benefits of two complementary paradigms for networked visual analysis recently discussed in the literature. In the traditional compress-then-analyze (CTA) paradigm, compressed images are transmitted from camera nodes to a central controller, where they are analyzed. In the novel analyze-then-compress (ATC) paradigm, camera nodes extract and compress local features from the acquired images. Such features are transmitted to the central controller and used to perform visual analysis. We show that the ATC paradigm outperforms CTA from the consumed energy point of view, at the same target analysis accuracy in both the application scenarios.
Luca Bondi, Luca Baroffio, Matteo Cesana, Alessandro Redondi, Marco Tagliasacchi
IEEE Internet Things J.5
2016 Rate-energy-accuracy optimization of convolutional architectures for face recognition
Luca Bondi, Luca Baroffio, Matteo Cesana, Marco Tagliasacchi, Giovani Chiachia, Anderson Rocha 0001
J. Vis. Commun. Image Represent.4
2016 Deep Convolutional Neural Networks for pedestrian detection
Denis Tomè, Federico Monti, Luca Baroffio, Luca Bondi, Marco Tagliasacchi, Stefano Tubaro
Signal Process. Image Commun.5
2016 Codec and GOP Identification in Double Compressed Videos
abstract
Video content is routinely acquired and distributed in a digital compressed format. In many cases, the same video content is encoded multiple times. This is the typical scenario that arises when a video, originally encoded directly by the acquisition device, is then re-encoded, either after an editing operation, or when uploaded to a sharing website. The analysis of the bitstream reveals details of the last compression step (i.e., the codec adopted and the corresponding encoding parameters), while masking the previous compression history. Therefore, in this paper, we consider a processing chain of two coding steps, and we propose a method that exploits coding-based footprints to identify both the codec and the size of the group of pictures (GOPs) used in the first coding step. This sort of analysis is useful in video forensics, when the analyst is interested in determining the characteristics of the originating source device, and in video quality assessment, since quality is determined by the whole compression history. The proposed method relies on the fact that lossy coding is an (almost) idempotent operation. That is, re-encoding a video sequence with the same codec and coding parameters produces a sequence that is similar to the former. As a consequence, if the second codec in the chain does not significantly alter the sequence, it is possible to analyze this sort of similarity to identify the first codec and the adopted GOP size. The method was extensively validated on a very large data set of video sequences generated by encoding content with a diversity of codecs (MPEG-2, MPEG-4, H.264/AVC, and DIRAC) and different encoding parameters. In addition, a proof of concept showing that the proposed method can also be used on videos downloaded from YouTube is reported.
Paolo Bestagini, Simone Milani, Marco Tagliasacchi, Stefano Tubaro
IEEE Trans. Image Process.3
2016 Identification of Transform Coding Chains
abstract
Transform coding is routinely used for lossy compression of discrete sources with memory. The input signal is divided into N-dimensional vectors, which are transformed by means of a linear mapping. Then, transform coefficients are quantized and entropy coded. In this paper, we consider the problem of identifying the transform matrix as well as the quantization step sizes. First, we study the case in which the only available information is a set of P transform decoded vectors. We formulate the problem in terms of finding the lattice with the largest determinant that contains all observed vectors. We propose an algorithm that is able to find the optimal solution and we formally study its convergence properties. Three potential realms of application are considered as example scenarios for the proposed theory: 1) parameter retrieval in the presence of a chain of two transform coders; 2) image tampering identification; and 3) parameter estimation for predictive coders. We show that, despite their differences, all three scenarios can be tackled by applying the same fundamental methodology. Experiments on both the synthetic data and the real images validate the proposed approach.
Marco Tagliasacchi, Marco Visentini Scarzanella, Pier Luigi Dragotti, Stefano Tubaro
IEEE Trans. Image Process.1
2016 Crowdsourcing for Top-K Query Processing over Uncertain Data
abstract
Querying uncertain data has become a prominent application due to the proliferation of user-generated content from social media and of data streams from sensors. When data ambiguity cannot be reduced algorithmically, crowdsourcing proves a viable approach, which consists of posting tasks to humans and harnessing their judgment for improving the confidence about data values or relationships. This paper tackles the problem of processing top-K queries over uncertain data with the help of crowdsourcing for quickly converging to the realordering of relevant results. Several offline and online approaches for addressing questions to a crowd are defined and contrasted on both synthetic and real data sets, with the aim of minimizing the crowd interactions necessary to find the realordering of the result set.
Eleonora Ciceri, Piero Fraternali, Davide Martinenghi, Marco Tagliasacchi
IEEE Trans. Knowl. Data Eng.4
2016 Compress-then-Analyze versus Analyze-then-Compress: What Is Best in Visual Sensor Networks?
abstract
Visual sensor networks (VSNs) have attracted the interest of researchers worldwide in the last few years, and are expected to play a major role in the evolution of the Internet-of-Things (IoT). When used to perform visual analysis tasks, VSNs may be operated according to two different paradigms. In the traditional compress-then-analyze paradigm, images are acquired, compressed and transmitted for further analysis. Conversely, in the analyze-then-compress paradigm, image features are extracted by visual sensor nodes, encoded and then delivered to a remote destination where analysis is performed. The question this paper aims to answer is What is the best visual analysis paradigm in VSNs?To do this, first we empirically characterize the rate-energy-accuracy performance of the two aforementioned paradigms. Then, we leverage such models to formulate a resource allocation problem for VSNs. The problem optimally allocates the specific paradigm used by each camera node in the network and the related transmission source rate, with the objective of optimizing the accuracy of the visual analysis task and the VSN coverage. Experimental results over several VSNs instances demonstrate that there is no “winning” paradigm, but the best performance are obtained by allowing the coexistence of the two and by properly optimizing their utilization.
Alessandro Redondi, Luca Baroffio, Lucio Bianchi, Matteo Cesana, Marco Tagliasacchi
IEEE Trans. Mob. Comput.5
2016 Estimating Snow Cover From Publicly Available Images
abstract
In this paper, we study the problem of estimating snow cover in mountainous regions, that is, the spatial extent of the earth surface covered by snow. We argue that publicly available visual content, in the form of user-generated photographs and image feeds from outdoor webcams, can both be leveraged as additional measurement sources, complementing existing ground, satellite, and airborne sensor data. To this end, we describe two content acquisition and processing pipelines that are tailored to such sources, addressing the specific challenges posed by each of them, e.g., identifying the mountain peaks, filtering out images taken in bad weather conditions, and handling varying illumination conditions. The final outcome is summarized in a snow cover index, which indicates for a specific mountain and day of the year the fraction of visible area covered by snow, possibly at different elevations. We created a manually labeled dataset to assess the accuracy of the image snow covered area estimation, achieving 90.0% precision at 91.1% recall. In addition, we show that seasonal trends related to air temperature are captured by the snow cover index.
Roman Fedorov 0001, Alessandro Camerada, Piero Fraternali, Marco Tagliasacchi
IEEE Trans. Multim.4
2015 Distributed object recognition in Visual Sensor Networks
abstract
This work focuses on Visual Sensor Networks (VSNs) which perform visual analysis tasks such as object recognition. There, the goal is to find the image in a reference database which is the closest match to the image captured by camera sensor nodes. Recognition is performed by relying on visual features extracted from the acquired image, which are matched against a database of labeled features in order to find the closest image match. The matching functionalities are often implemented at a central controller outside the VSN. In contrast, we study the performance trade-offs involved in distributing the matching functionalities inside the VSN by letting sensor nodes performing parts of the matching process. We propose an optimization framework to optimally distribute the matching task to in-network sensor nodes with the goal of minimizing the overall completion time of the recognition task. The proposed optimization framework is then used to assess the performance of distributed matching, comparing it to a traditional, centralized approach in realistic VSN scenarios.
Stefano Paris, Alessandro Redondi, Matteo Cesana, Marco Tagliasacchi
ICC4
2015 Hybrid coding of visual content and local image features
abstract
Distributed visual analysis applications, such as mobile visual search or Visual Sensor Networks (VSNs) require the transmission of visual content on a bandwidth-limited network, from a peripheral node to a processing unit. Traditionally, a “Compress-Then-Analyze” approach has been pursued, in which sensing nodes acquire and encode the pixel-level representation of the visual content, that is subsequently transmitted to a sink node in order to be processed. This approach might not represent the most effective solution, since several analysis applications leverage a compact representation of the content, thus resulting in an inefficient usage of network resources. Furthermore, coding artifacts might significantly impact the accuracy of the visual task at hand. To tackle such limitations, an orthogonal approach named “Analyze-Then-Compress” has been proposed [1]. According to such a paradigm, sensing nodes are responsible for the extraction of visual features, that are encoded and transmitted to a sink node for further processing. In spite of improved task efficiency, such paradigm implies the central processing node not being able to reconstruct a pixel-level representation of the visual content. In this paper we propose an effective compromise between the two paradigms, namely “Hybrid-Analyze-Then-Compress” (HATC) that aims at jointly encoding visual content and local image features. Furthermore, we show how a target tradeoff between image quality and task accuracy might be achieved by accurately allocating the bitrate to either visual content or local features.
Luca Baroffio, Matteo Cesana, Alessandro Redondi, Marco Tagliasacchi, Stefano Tubaro
ICIP4
2015 Phylogeny reconstruction for misaligned and compressed video sequences
abstract
In the last few years, the amount of videos distributed online has dramatically increased due to the popularity of media sharing platforms (e.g., YouTube, Vimeo, etc.). However, distributed videos are often edited copies of original content, typically referred to as near duplicates. In this paper, we face the problem of reconstructing a video phylogeny tree, i.e., given a set of near-duplicate videos, we want to reconstruct the relationships between every pair of videos to detect which one generated the others and trace back their evolution history. Solving this problem is of paramount importance when the first published video within a set is sought, e.g., to solve copyright infringement cases or to pinpoint criminal impersonation online. The technique we propose exploits the same rationale of previous works in the field of image and video phylogeny. However, we embed in the commonly used pipeline of operations the possibility of dealing with temporally misaligned and encoded video sequences, thus making the method applicable to user-generated videos shared on online platforms. Results computed on a wide dataset of video sequences highlight the importance of taking care of both coding and misalignment in the reconstruction pipeline.
Filipe de Oliveira Costa, Silvia Lameri, Paolo Bestagini, Zanoni Dias, Anderson Rocha 0001, Marco Tagliasacchi, Stefano Tubaro
ICIP6
2015 Near-duplicate detection and alignment for multi-view videos
abstract
The increasing popularity of video sharing platforms (e.g., YouTube, Vimeo, etc.) has determined the widespread diffusion of near-duplicate videos, i.e., sequences obtained applying different editing operations to the same original clip. However, it is also possible to come across sequences referring to the same specific event shot from different viewpoints. This is a very common situation that arises when analyzing user-generated content acquired with mobile devices. Therefore, for some applications, it can be useful to extend the concept of near-duplicates considering also all the videos (and their edited versions) referring to the same event even if shot from different viewpoints. In this paper we consider such challenging scenario. More specifically, we focus on the problem of multi-view near-duplicate video detection and temporal alignment. In doing so, we show the limitations of a state-of-the-art algorithm based on robust hashing, and propose a processing pipeline that allows to deal also with sequences taken from significantly different viewpoints.
Andrea Melloni, Silvia Lameri, Paolo Bestagini, Marco Tagliasacchi, Stefano Tubaro
ICIP4
2015 A Mathematical Programming Approach to Task Offloading in Visual Sensor Networks
abstract
This work studies how visual analysis tasks based on feature extraction can be speeded up in the context of Visual Sensor Networks. The main catch is for the camera node to leverage the presence of neighboring sensor nodes and offload the task, thus parallelizing its execution. We propose two mathematical programming formulations for the optimal visual task offloading problem: the first one targets the minimization of the overall task completion time while enforcing energy consumption constraints onto the nodes; the second maximizes the overall sensor network lifetime subject to a temporal constraint on the task completion time. The aforementioned formulations are used to characterize the achievable speed-up and consequent energy consumption in representative visual sensor network topologies.
Alessandro Redondi, Matteo Cesana, Luca Baroffio, Marco Tagliasacchi
VTC Spring4
2015 Cooperative image analysis in visual sensor networks
Alessandro Redondi, Matteo Cesana, Marco Tagliasacchi, Ilario Filippini, György Dán, Viktoria Fodor
Ad Hoc Networks3
2015 Coding Local and Global Binary Visual Features Extracted From Video Sequences
abstract
Binary local features represent an effective alternative to real-valued descriptors, leading to comparable results for many visual analysis tasks while being characterized by significantly lower computational complexity and memory requirements. When dealing with large collections, a more compact representation based on global features is often preferred, which can be obtained from local features by means of, e.g., the bag-of-visual word model. Several applications, including, for example, visual sensor networks and mobile augmented reality, require visual features to be transmitted over a bandwidth-limited network, thus calling for coding techniques that aim at reducing the required bit budget while attaining a target level of efficiency. In this paper, we investigate a coding scheme tailored to both local and global binary features, which aims at exploiting both spatial and temporal redundancy by means of intra- and inter-frame coding. In this respect, the proposed coding scheme can conveniently be adopted to support the analyze-then-compress (ATC) paradigm. That is, visual features are extracted from the acquired content, encoded at remote nodes, and finally transmitted to a central controller that performs the visual analysis. This is in contrast with the traditional approach, in which visual content is acquired at a node, compressed and then sent to a central unit for further processing, according to the compress-then-analyze (CTA) paradigm. In this paper, we experimentally compare the ATC and the CTA by means of rate-efficiency curves in the context of two different visual analysis tasks: 1) homography estimation and 2) content-based retrieval. Our results show that the novel ATC paradigm based on the proposed coding primitives can be competitive with the CTA, especially in bandwidth limited scenarios.
Luca Baroffio, Antonio Canclini, Matteo Cesana, Alessandro Redondi, Marco Tagliasacchi, Stefano Tubaro
IEEE Trans. Image Process.5
2014 Energy Consumption of Visual Sensor Networks: Impact of Spatio-Temporal Coverage Based on Single-Hop Topologies
Alessandro Redondi, Dujdow Buranapanichkit, Matteo Cesana, Marco Tagliasacchi, Yiannis Andreopoulos
EWSN4
2014 Demosaicing strategy identification via eigenalgorithms
abstract
The identification of the camera that has acquired a specific image can be performed via several device-related footprints. Among these, it is possible to look for the traces left by the adopted color demosaicing strategy, which varies according to the camera model and vendor. The paper presents an identification strategy that re-processes the analyzed image with a set of distinctive CFA interpolation algorithms (eigenalgorithms) and, according to the correlation of the output with the original image, builds a set of features that permits identifying the algorithm. The proposed solution performs well with respect to other state-of-the-art solutions also when the analyzed image is severely compressed.
Simone Milani, Paolo Bestagini, Marco Tagliasacchi, Stefano Tubaro
ICASSP3
2014 Antiforensic synthesis of motion vectors using template algorithms
abstract
The identification of the video camera employed to acquire a video sequence is made possible by a large set of different footprints. Since video signals are always available in a compressed format, some of the most significant traces can be related to the coding tools of the implemented video codec (e.g., rate-distortion optimization, motion estimation strategy, etc.). As a matter of fact, an effective antiforensic attack, which aims at fooling the tools that identify the acquisition device, must appropriately alter these footprints. In the paper, we present an antiforensic strategy that targets a video camera detector which is based on the identification of the motion estimation strategy used by the video coder. The proposed approach synthesizes a set of motion vectors that approximate those that would have been generated by the algorithm to be mimicked. This method proves to be effective in attacking the detector while preserving the coding efficiency.
Simone Milani, Paolo Bestagini, Marco Tagliasacchi, Stefano Tubaro
ICASSP3
2014 Coding binary local features extracted from video sequences
abstract
Local features represent a powerful tool which is exploited in several applications such as visual search, object recognition and tracking, etc. In this context, binary descriptors provide an efficient alternative to real-valued descriptors, due to low computational complexity, limited memory footprint and fast matching algorithms. The descriptor consists of a binary vector, in which each bit is the result of a pairwise comparison between smoothed pixel intensities. In several cases, visual features need to be transmitted over a bandwidth-limited network. To this end, it is useful to compress the descriptor to reduce the required rate, while attaining a target accuracy for the task at hand. The past literature thoroughly addressed the problem of coding visual features extracted from still images and, only very recently, the problem of coding real-valued features (e.g., SIFT, SURF) extracted from video sequences. In this paper we propose a coding architecture specifically designed for binary local features extracted from video content. We exploit both spatial and temporal redundancy by means of intra-frame and inter-frame coding modes, showing that significant coding gains can be attained for a target level of accuracy of the visual analysis task.
Luca Baroffio, João Ascenso, Matteo Cesana, Alessandro Redondi, Marco Tagliasacchi
ICIP5
2014 Briskola: BRISK optimized for low-power ARM architectures
abstract
Local visual features are commonly adopted to accomplish analysis tasks such as object recognition/tracking and image retrieval. Recently, several visual features extraction algorithms tailored to low-power architectures have been proposed, in order to enable image analysis on energy-constrained devices such as smart-phones or Visual Sensor Networks (VSN). In this work, we dissect and analyze BRISK, a state-of-the-art low-power visual feature extractor, in order to evaluate the impact of its individual building blocks on the overall energy consumption. For each building block, we propose a solution to limit the energy consumption without affecting the overall analysis performance. The resulting BRISKOLA (BRISK Optimized for Low-power ARM architectures) feature extractor exhibits energy savings up to 30% with respect to the original implementation.
Luca Baroffio, Antonio Canclini, Matteo Cesana, Alessandro Redondi, Marco Tagliasacchi
ICIP5
2014 Enabling visual analysis in wireless sensor networks
abstract
This demo showcases some of the results obtained by the GreenEyes project, whose main objective is to enable visual analysis on resource-constrained multimedia sensor networks. The demo features a multi-hop visual sensor network operated by BeagleBones Linux computers with IEEE 802.15.4 communication capabilities, and capable of recognizing and tracking objects according to two different visual paradigms. In the traditional compress-then-analyze (CTA) paradigm, JPEG compressed images are transmitted through the network from a camera node to a central controller, where the analysis takes place. In the alternative analyze-then-compress (ATC) paradigm, the camera node extracts and compresses local binary visual features from the acquired images (either locally or in a distributed fashion) and transmits them to the central controller, where they are used to perform object recognition/tracking. We show that, in a bandwidth constrained scenario, the latter paradigm allows to reach better results in terms of application frame rates, still ensuring excellent analysis performance.
Luca Baroffio, Antonio Canclini, Matteo Cesana, Alessandro Redondi, Marco Tagliasacchi, György Dán, Emil Eriksson, Viktoria Fodor, João Ascenso, Pedro Monteiro
ICIP5
2014 Bamboo: A fast descriptor based on AsymMetric pairwise BOOsting
abstract
A robust hash, or content-based fingerprint, is a succinct representation of the perceptually most relevant parts of a multimedia object. A key requirement of fingerprinting is that elements with perceptually similar content should map to the same fingerprint, even if their bit-level representations are different. In this work we propose BAMBOO (Binary descriptor based on AsymMetric pairwise BOOsting), a binary local descriptor that exploits a combination of content-based fingerprinting techniques and computationally efficient filters (box filters, Haar-like features, etc.) applied to image patches. In particular, we define a possibly large set of filters and iteratively select the most discriminative ones resorting to an asymmetric pair-wise boosting technique. The output values of the filtering process are quantized to one bit, leading to a very compact binary descriptor. Results show that such descriptor leads to compelling results, significantly outperforming binary descriptors having comparable complexity (e.g., BRISK), and approaching the discriminative power of state-of-the-art descriptors which are significantly more complex (e.g., SIFT and BinBoost).
Luca Baroffio, Matteo Cesana, Alessandro Redondi, Marco Tagliasacchi
ICIP4
2014 Snow phenomena modeling through online public media
abstract
We propose a method for the environmental monitoring through the publicly available media User Generated Content (UGC). In particular we address the problem of the snow cover and level estimation by analyzing the social media data such as geotagged photographs and public webcams installed in mountain regions. The entire pipeline of the process is presented to the audience: from the data crawling and automatic relevance classification (does or does not the photograph contain a significant mountain profile) to the image content analysis and environmental models (identification of the snow covered area on the photograph). Each presented component is self-contained and can be inspected individually, the connections between the components however are strongly highlighted allowing the viewer to understand intuitively the entire pipeline structure.
Roman Fedorov 0001, Piero Fraternali, Marco Tagliasacchi
ICIP3
2014 Who is my parent? Reconstructing video sequences from partially matching shots
abstract
Nowadays, a significant fraction of the available video content is created by reusing already existing online videos. In these cases, the source video is seldom reused as is. Conversely, it is typically time clipped to extract only a subset of the original frames, and other transformations are commonly applied (e.g., cropping, logo insertion, etc.). In this paper, we analyze a pool of videos related to the same event or topic. We propose a method that aims at automatically reconstructing the content of the original source videos, i.e., the parent sequences, by splicing together sets of near-duplicate shots seemingly extracted from the same parent sequence. The result of the analysis shows how content is reused, thus revealing the intent of content creators, and enables us to reconstruct a parent sequence also when it is no longer available online. In doing so, we make use of a robust-hash algorithm that allows us to detect whether groups of frames are near-duplicates. Based on that, we developed an algorithm to automatically find near-duplicate matchings between multiple parts of multiple sequences. All the near-duplicate parts are finally temporally aligned to reconstruct the parent sequence. The proposed method is validated with both synthetic and real world datasets downloaded from YouTube.
Silvia Lameri, Paolo Bestagini, Andrea Melloni, Simone Milani, Anderson Rocha 0001, Marco Tagliasacchi, Stefano Tubaro
ICIP6
2014 Detectability-quality trade-off in JPEG counter-forensics
abstract
Removing JPEG quantization footprints from an image inevitably introduces artifacts and traces in the spatial domain. Recently, several robust methods have been proposed to detect footprints of counter-forensics and recover the image's compression history. In this paper we investigate the limitations of these detectors, by proposing an improved counter-forensic attack which adds a postprocessing denoising step besides dithering. We consider both a general-purpose denoising algorithm and one targeted to JPEG images. In the latter case, we show that this approach can successfully reduce the accuracy of detectors in the literature to that of a random decision. As a second contribution, we study the trade-off between the detectability of counter-forensics and quality of the tampered image, and show that the loss of quality is not sufficient for the analyst to use available no-reference quality assessment tools as an indicator of an attack.
Giuseppe Valenzise, Marco Tagliasacchi, Stefano Tubaro
ICIP2
2014 HistoGraph - A Visualization Tool for Collaborative Analysis of Networks from Historical Social Multimedia Collections
abstract
This paper describes the design and development of histoGraph, an interactive tool for explorative visualization and collaborative investigation of historical social networks from multimedia collections. Developed in an interdisciplinary collaboration of computer scientists, historians, HCI researchers and interface designers, the tool aims at supporting historians in the discovery and historical analysis of relationships between people, places and events. A special focus is on the identification and interactive visualization of social relations from historical photo collections through a combination of automatic analysis and expert-based crowd sourcing. The tool design bridges the gap between established network analysis and visualization techniques and traditional hermeneutic research methods in historical research. It integrates visual exploration with hybrid social graph construction, hypothesis formulation and the consultation of digitized primary sources. A formative evaluation of the current prototype, developed as a domain-specific application for historians in the field of European integration points to opportunities and critical factors in applying this approach to support and further current research practices in digital humanities.
Jasminko Novak, Isabel Micheel, Mark S. Melenhorst, Lars Wieneke, Marten Düring, Javier Garcia Moron, Chiara Pasini, Marco Tagliasacchi, Piero Fraternali
IV8
2014 Robust aggregation of GWAP tracks for local image annotation
abstract
The possibility of assigning labels to localized regions in an image enables flexible image retrieval paradigms. However, the process of automatically segmenting and tagging images is notoriously hard, due to the presence of occlusions, noise, challenging illumination conditions, background clutter, etc. For this reason, human computation has recently emerged as a viable alternative when computer vision algorithms fail to provide a satisfactory answer. For example, Games with a purpose (GWAP) represent a powerful crowdsourcing mechanism to collect implicit annotations from human players. In this paper we consider the problem of aggregating the gaming tracks collected by a GWAP we developed to solve challenging instances of image segmentation problems. In particular we consider the existence of malicious players, who might try to fool the rules of the game to achieve higher rewards. The proposed solution can automatically estimate the reliability of human players, thus identifying cheaters. This information is exploited to aggregate the gaming tracks, thus significantly improving the image segmentation result and the quality of local image annotations.
Carlo Bernaschina, Piero Fraternali, Luca Galli, Davide Martinenghi, Marco Tagliasacchi
ICMR5
2014 Energy Consumption of Visual Sensor Networks: Impact of Spatio-Temporal Coverage
abstract
Wireless visual sensor networks (VSNs) are expected to play a major role in future IEEE 802.15.4 personal area networks (PANs) under recently established collision-free medium access control (MAC) protocols, such as the IEEE 802.15.4e-2012 MAC. In such environments, the VSN energy consumption is affected by a number of camera sensors deployed (spatial coverage), as well as a number of captured video frames of which each node processes and transmits data (temporal coverage). In this paper we explore this aspect for uniformly formed VSNs, that is, networks comprising identical wireless visual sensor nodes connected to a collection node via a balanced cluster-tree topology, with each node producing independent identically distributed bitstream sizes after processing the video frames captured within each network activation interval. We derive analytic results for the energy-optimal spatio-temporal coverage parameters of such VSNs under a priori known bounds for the number of frames to process per sensor and the number of nodes to deploy within each tier of the VSN. Our results are parametric to the probability density function characterizing the bitstream size produced by each node and the energy consumption rates of the system of interest. Experimental results are derived from a deployment of TelosB motes and reveal that our analytic results are always within 7% of the energy consumption measurements for a wide range of settings. In addition, results obtained via motion JPEG encoding and feature extraction on a multimedia subsystem (BeagleBone Linux Computer) show that the optimal spatio-temporal settings derived by our framework allow for substantial reduction of energy consumption in comparison with ad hoc settings.
Alessandro Redondi, Dujdow Buranapanichkit, Matteo Cesana, Marco Tagliasacchi, Yiannis Andreopoulos
IEEE Trans. Circuits Syst. Video Technol.4
2014 Coding Visual Features Extracted From Video Sequences
abstract
Visual features are successfully exploited in several applications (e.g., visual search, object recognition and tracking, etc.) due to their ability to efficiently represent image content. Several visual analysis tasks require features to be transmitted over a bandwidth-limited network, thus calling for coding techniques to reduce the required bit budget, while attaining a target level of efficiency. In this paper, we propose, for the first time, a coding architecture designed for local features (e.g., SIFT, SURF) extracted from video sequences. To achieve high coding efficiency, we exploit both spatial and temporal redundancy by means of intraframe and interframe coding modes. In addition, we propose a coding mode decision based on rate-distortion optimization. The proposed coding scheme can be conveniently adopted to implement the analyze-then-compress (ATC) paradigm in the context of visual sensor networks. That is, sets of visual features are extracted from video frames, encoded at remote nodes, and finally transmitted to a central controller that performs visual analysis. This is in contrast to the traditional compress-then-analyze (CTA) paradigm, in which video sequences acquired at a node are compressed and then sent to a central unit for further processing. In this paper, we compare these coding paradigms using metrics that are routinely adopted to evaluate the suitability of visual features in the context of content-based retrieval, object recognition, and tracking. Experimental results demonstrate that, thanks to the significant coding gains achieved by the proposed coding scheme, ATC outperforms CTA with respect to all evaluation metrics.
Luca Baroffio, Matteo Cesana, Alessandro Redondi, Marco Tagliasacchi, Stefano Tubaro
IEEE Trans. Image Process.4
2013 Quantisation Invariants for Transform Parameter Estimation in Coding Chains
abstract
We examine the case of a signal going through a processing chain consisting of two transform coding stages, with the aim of recovering the unknown parameters of the first encoder. Through number theoretical considerations, we identify a lattice of quantisation invariant points, whose coordinates are not affected by the double quantisation and whose parameters are closely related to the unknown transform. The conditions for this lattice to exist are then discussed, and its uniqueness properties analysed. Finally, an algorithmic procedure to recover the invariants from a sparse set of points is shown together with numerical results.
Marco Visentini Scarzanella, Marco Tagliasacchi, Pier Luigi Dragotti
DCC2
2013 Detection of temporal interpolation in video sequences
abstract
Nowadays, considering the availability of relatively cheap devices and powerful editing software, video tampering is a relatively easy task. Video sequences can be tampered with by performing, e.g., temporal splicing. However, if the sequences spliced together do not share the same frame rate, they have to be temporally interpolated beforehand. This operation is often made using motion compensated interpolators, which allow to minimize visual artifacts. In this paper we propose a detector of this kind of interpolation. Moreover, the detector is capable of identifying the interpolation factor used, allowing an analyst to uncover the original frame rate of a sequence. This method relies on the analysis of the correlation introduced by the filter adopted by the interpolator. Results show that detection is successful, provided that the number of observed interpolated frames is large enough. Moreover, tests on compressed sequences obtained from television broadcasts validate the method in a real world scenario.
Paolo Bestagini, S. Battaglia, Simone Milani, Marco Tagliasacchi, Stefano Tubaro
ICASSP4
2013 Antiforensics attacks to Benford's law for the detection of double compressed images
abstract
Researchers have been recently challenging the robustness of forensic algorithms by designing antiforensic strategies that try to fool them. In this paper, we propose an antiforensic strategy that targets double image compression detectors based on Benford's law (or first digit law). The proposed approach is able to modify the first digit statistics of the considered data (a double compressed image) to fool single/double compression detectors based on Benford's law. In this way, the proposed strategy tries to mimick the effects of a single compression with limited additional distortion. The presented algorithm performs better than previous state-of-the-art antiforensic strategies and can be easily extended to other fraud detection methods.
Simone Milani, Marco Tagliasacchi, Stefano Tubaro
ICASSP2
2013 Transform coder identification
abstract
The widespread popularity of transform coding has made it central to a wide range of methods in forensics, quality assessment and digital restoration. However, most approaches require prior knowledge of the transform coding parameters. In this paper, we consider the challenging problem of identifying the transform matrix as well as the quantization step sizes of a transform coder, given a set of P non-overlapping N-dimensional vectors observed as its output. We formulate the problem in terms of finding the lattice with the largest determinant that contains all observed vectors and we propose an algorithm that is able to find the optimal solution. Our experimental analysis shows that the probability of success of the algorithm quickly approaches 1 for small values of (P - N). The complexity of the proposed algorithm grows linearly with the dimensionality N.
Marco Tagliasacchi, Marco Visentini Scarzanella, Pier Luigi Dragotti, Stefano Tubaro
ICASSP1
2013 Coding video sequences of visual features
abstract
Visual features provide a convenient representation of the image content, which is exploited in several applications, e.g., visual search, object tracking, etc. In several cases, visual features need to be transmitted over a bandwidth-limited network, thus calling for coding techniques to reduce the required rate, while attaining a target efficiency for the task at hand. Although the literature has recently addressed the problem of coding local features extracted from still images, in this paper we propose, for the first time, a coding architecture designed for local features extracted from video content. We exploit both spatial and temporal redundancy by means of intra-frame and inter-frame coding modes. In addition, we propose a coding mode decision based on rate-distortion optimization. Experimental results demonstrate that, in the case of SIFT descriptors, exploiting temporal redundancy leads to substantial gains in terms of coding efficiency.
Luca Baroffio, Matteo Cesana, Alessandro Redondi, Stefano Tubaro, Marco Tagliasacchi
ICIP5
2013 Video recapture detection based on ghosting artifact analysis
abstract
Video forensics is becoming a popular field of research and an increasing number of forensic techniques have been proposed in the last few years. However, a simple yet effective method to fool many detectors consists in recapturing a video sequence with a camcorder. For this reason being able to detect video recapture is a topic of interest for a forensic analyst. In this paper, we first characterize the video recapture model, focusing on the common scenario of a sequence recaptured from a LCD monitor using a digital camcorder, then we propose a recapture detector for this case. The detector is based on the analysis of a characteristic ghosting artifact left by the recapture process. The presented algorithm is finally validated by means of tests on original and recaptured sequences. These tests prove that the algorithm achieves high accuracy results.
Paolo Bestagini, Marco Visentini Scarzanella, Marco Tagliasacchi, Pier Luigi Dragotti, Stefano Tubaro
ICIP3
2013 Identification of the motion estimation strategy using eigenalgorithms
abstract
The identification of the device, or device model, that was used to acquire a video sequence is a very challenging task, since it has to rely on subtle traces left by the processing steps applied to the raw acquired data. Previous works have tried to address this problem leveraging the traces left by the imaging sensor. However, in the case of video, lossy coding is often quite aggressive, thus making these methods impractical. In this work, we reverse the analysis strategy and exploit the traces left by lossy coding as telltale for the adopted acquisition device. Specifically, we aim at detecting the implementation of the video codec by identifying the adopted motion estimation algorithm. Indeed, motion estimation is not defined in video coding standards and, as such, it represents one of the non-normative tools that can be customized in the design of the encoder. The key tenet consists in studying the correlation between the motion vectors obtained from the decoded bitstream, and those computed using a set of known and diverse motion estimation algorithms, called eigenalgorithms. In our work, we generalize a method recently appeared in the literature, which assumes that the motion estimation algorithm used is necessarily one of those available during the analysis. Experimental results show that the approach is able to successfully identify the motion estimation algorithm in most cases.
Simone Milani, Marco Tagliasacchi, Stefano Tubaro
ICIP2
2013 Rate-accuracy optimization of binary descriptors
abstract
Binary descriptors have recently emerged as low-complexity alternatives to state-of-the-art descriptors such as SIFT. The descriptor is represented by means of a binary string, in which each bit is the result of the pair-wise comparison of smoothed pixel values properly selected in a patch around each keypoint. Previous works have focused on the construction of the descriptor neglecting the opportunity of performing lossless compression. In this paper, we propose two contributions. First, design an entropy coding scheme that seeks the internal ordering of the descriptor that minimizes the number of bits necessary to represent it. Second, we compare different selection strategies that can be adopted to identify which pair-wise comparisons to use when building the descriptor. Unlike previous works, we evaluate the discriminative power of descriptors as a function of rate, in order to investigate the trade-offs in a bandwidth constrained scenario.
Alessandro Redondi, Luca Baroffio, João Ascenso, Matteo Cesana, Marco Tagliasacchi
ICIP5
2013 Transform coder identification with double quantized data
abstract
The analysis of chains of double transform coders has been recently addressed in the image forensic literature, especially for the case of double JPEG compression. In that case, the transform is assumed to be known a priori (e.g., 2D-DCT), whereas the quantization steps of the first coder need to be determined. In this work, we generalize the analysis to the challenging case in which nothing is known about the first coder, but that the transform is orthonormal. Given a set of vectors observed as output of a chain of two transform coders, we identify both the transform and the quantizer of the first. The key idea is to denoise the observed vectors exploiting the constraints imposed by the first quantizer and then apply our previously proposed method, which successfully performs transform identification in the case of noiseless observations. Experiments on real images validate the proposed approach.
Marco Tagliasacchi, Marco Visentini Scarzanella, Pier Luigi Dragotti, Stefano Tubaro
ICIP1
2013 Local tampering detection in video sequences
abstract
Video sequences are often believed to provide stronger forensic evidence than still images, e.g., when used in lawsuits. However, a wide set of powerful and easy-to-use video authoring tools is today available to anyone. Therefore, it is possible for an attacker to maliciously forge a video sequence, e.g., by removing or inserting an object in a scene. These forms of manipulation can be performed with different techniques. For example, a portion of the original video may be replaced by either a still image repeated in time or, in more complex cases, by a video sequence. Moreover, the attacker might use as source data either a spatio-temporal region of the same video, or a region taken from an external sequence. In this paper we present the analysis of the footprints left when tampering with a video sequence, and propose a detection algorithm that allows a forensic analyst to reveal video forgeries and localize them in the spatio-temporal domain. With respect to the state-of-the-art, the proposed method is completely unsupervised and proves to be robust to compression. The algorithm is validated against a dataset of forged videos available online.
Paolo Bestagini, Simone Milani, Marco Tagliasacchi, Stefano Tubaro
MMSP3
2013 Audio tampering detection via microphone classification
abstract
In this paper, we present a new approach for audio tampering detection based on microphone classification. The underlying algorithm is based on a blind channel estimation, specifically designed for recordings from mobile devices. It is applied to detect a specific type of tampering, i.e., to detect whether footprints from more than one microphone exist within a given content item. As will be shown, the proposed method achieves an accuracy above 95% for AAC, MP3 and PCM-encoded recordings.
Luca Cuccovillo, Sebastian Mann, Marco Tagliasacchi, Patrick Aichroth
MMSP3
2013 A phylogenetic analysis of near-duplicate audio tracks
abstract
We present a content-based system for the analysis of near-duplicate audio tracks. The objective is to infer the structure of modifications, represented as trees, underneath a pool of near-duplicates, possibly specifying the operations the tracks have gone through. A pilot study was carried out for a set of plausible processing operators, including trim, fade and perceptual audio coding, generating near-duplicates from an original audio track. The proposed method measures the similarity between pairs of near-duplicates and reconstructs a tree representing the causal dependencies in the analyzed pool. Experimental results demonstrate that the structure of the tree can be successfully recovered, also in the challenging case in which some pieces of information are missing, i.e., when only a subset of near-duplicate tracks is available.
Matteo Nucci, Marco Tagliasacchi, Stefano Tubaro
MMSP2
2013 Compress-then-analyze vs. analyze-then-compress: Two paradigms for image analysis in visual sensor networks
abstract
We compare two paradigms for image analysis in visual sensor networks (VSN). In the compress-then-analyze (CTA) paradigm, images acquired from camera nodes are compressed and sent to a central controller for further analysis. Conversely, in the analyze-then-compress (ATC) approach, camera nodes perform visual feature extraction and transmit a compressed version of these features to a central controller. We focus on state-of-the-art binary features which are particularly suitable for resource-constrained VSNs, and we show that the “winning” paradigm depends primarily on the network conditions. Indeed, while the ATC approach might be the only possible way to perform analysis at low available bitrates, the CTA approach reaches the best results when the available bandwidth enables the transmission of high-quality images.
Alessandro Redondi, Luca Baroffio, Matteo Cesana, Marco Tagliasacchi
MMSP4
2013 Comparison of two paradigms for image analysis in visual sensor networks
abstract
This interactive demo presents and compares two different paradigms for image analysis in visual sensor networks (VSN), using a testbed based on battery-operated Beagle-Bone platforms with sight and wireless communication capabilities.
Antonio Canclini, Luca Baroffio, Matteo Cesana, Alessandro Redondi, Marco Tagliasacchi
SenSys5
2013 An integrated system based on wireless sensor networks for patient monitoring, localization and tracking
Alessandro Redondi, Marco Chirico, Luca Borsani, Matteo Cesana, Marco Tagliasacchi
Ad Hoc Networks5
2013 Provisional reporting for rank joins
Adnan Abid, Marco Tagliasacchi
J. Intell. Inf. Syst.2
2013 Introduction to the Special Issue on "Recent advances on analysis and processing for distributed video systems"
Chia-Wen Lin, Weiyao Lin, Zhenzhong Chen 0001, Marco Tagliasacchi, Shantanu Rane
J. Vis. Commun. Image Represent.4
2013 Revealing the Traces of JPEG Compression Anti-Forensics
abstract
Due to the lossy nature of transform coding, JPEG introduces characteristic traces in the compressed images. A forensic analyst might reveal these traces by analyzing the histogram of discrete cosine transform (DCT) coefficients and exploit them to identify local tampering, copy-move forgery, etc. At the same time, it has been recently shown that a knowledgeable adversary can possibly conceal the traces of JPEG compression, by adding a dithering noise signal in the DCT domain, in order to restore the histogram of the original image. In this paper, we study the processing chain that arises in the case of JPEG compression anti-forensics. We take the perspective of the forensic analyst, and we show how it is possible to counter the aforementioned anti-forensic method revealing the traces of JPEG compression, regardless of the quantization matrix being used. Tests on a large image dataset demonstrated that the proposed detector was able to achieve an average accuracy equal to 93%, rising above 99% when excluding the case of nearly lossless JPEG compression.
Giuseppe Valenzise, Marco Tagliasacchi, Stefano Tubaro
IEEE Trans. Inf. Forensics Secur.2
2013 Top-k diversity queries over bounded regions
abstract
Top-k diversity queries over objects embedded in a low-dimensional vector space aim to retrieve the best k objects that are both relevant to given user's criteria and well distributed over a designated region. An interesting case is provided by spatial Web objects, which are produced in great quantity by location-based services that let users attach content to places and are found also in domains like trip planning, news analysis, and real estate. In this article we present a technique for addressing such queries that, unlike existing methods for diversified top- k queries, does not require accessing and scanning all relevant objects in order to find the best k results. Our Space Partitioning and Probing (SPP) algorithm works by progressively exploring the vector space, while keeping track of the already seen objects and of their relevance and position. The goal is to provide a good quality result set in terms of both relevance and diversity. We assess quality by using as a baseline the result set computed by MMR, one of the most popular diversification algorithms, while minimizing the number of accessed objects. In order to do so, SPP exploits score-based and distance-based access methods, which are available, for instance, in most geo-referenced Web data sources. Experiments with both synthetic and real data show that SPP produces results that are relevant and spatially well distributed, while significantly reducing the number of accessed objects and incurring a very low computational overhead.
Ilio Catallo, Eleonora Ciceri, Piero Fraternali, Davide Martinenghi, Marco Tagliasacchi
ACM Trans. Database Syst.5
2012 Video codec identification
abstract
Video content is routinely acquired and distributed in digital format. Therefore, it is customary to have the content encoded multiple times. In this paper we consider a processing chain of two coding steps and we propose a method that aims at identifying the type of codec used in the first step, by analyzing its coding-based footprints. The method relies on the fact that lossy coding is an almost idempotent operation, i.e., re-encoding the reconstructed sequence with the same codec and coding parameters produces a sequence that is highly correlated with the input one. As a consequence, it is possible to analyze this sort of correlation to identify the first codec provided that the second codec does not introduce severe quality degradation. The proposed solution finds several applications in the field of multi-media forensics, e.g. to identify the device that generated the original video stream or detect collages of different sequences.
Paolo Bestagini, Simone Milani, Marco Tagliasacchi, Stefano Tubaro
ICASSP4
2012 Discriminating multiple JPEG compression using first digit features
abstract
The analysis of double-compressed images is a problem largely studied by the multimedia forensics community, as it might be exploited, e.g., for tampering localization or source device identification. In many practical scenarios, e.g. photos uploaded on blogs, on-line albums, and photo sharing Web sites, images might be compressed several times. However, the identification of the number of compression stages applied to an image remains an open issue. This paper proposes a forensic method based on the analysis of the distribution of the first significant digits of DCT coefficients, which is modeled according to Benford's law. The method relies on a set of Support Vector Machine (SVM) classifiers and allows us to accurately identify the number of compression stages applied to an image. Up to four consecutive compression stages were considered in the experimental validation. The proposed approach extends and outperforms the previously published methods aimed at detecting double JPEG compression.
Simone Milani, Marco Tagliasacchi, Stefano Tubaro
ICASSP2
2012 Rate-accuracy optimization in visual wireless sensor networks
abstract
We consider the problem of allocating the resources in a wireless sensor network, which is designed to perform visual analysis (e.g. object recognition). We depart from the traditional compress-then-analyze paradigm, in which nodes sense, compress and transmit visual data to a sink node. Instead, we study the case in which nodes extract and lossy code local features from pixel-domain representations of the sensed visual scene. The formulation of the allocation problem entails maximizing the lifetime of the visual sensor network subject to a target accuracy of the analysis task, together with energy, bandwidth and routing constraints. To this end, we contribute with the definition of a rate-accuracy model, which plays the role of the traditional rate-distortion model commonly adopted in visual communication. The proposed model captures the impact of: i) the number of selected local features; ii) the number of bits used for quantizing local features; iii) the criterion used to select the subset of local features to be transmitted. We verify the correctness of the models on two widely adopted visual dataset and we demonstrate the network lifetime gain that can be achieved by an optimal allocation of the resources.
Alessandro Redondi, Matteo Cesana, Marco Tagliasacchi
ICIP3
2012 Diversification for Multi-domain Result Sets
Alessandro Bozzon, Marco Brambilla 0001, Piero Fraternali, Marco Tagliasacchi
ICWE4
2012 Multiple compression detection for video sequences
abstract
Nowadays, thanks to the increasingly availability of powerful processors and user friendly applications, the editing of video sequences is becoming more and more frequent. Moreover, after each editing step, any video object is almost always encoded in order to store it using a less amount of memory. For this reason, inferring the number of compression steps that have been applied to such a multimedia object is an important clue in order to assess its authenticity. In this paper we propose a method to recover the number of compression steps applied to a video sequence. In order to accomplish this goal, we make use of a classifier based on multiple Support Vector Machines (SVM) exploiting the Benford's law. Indeed, the feature vectors used to train and test the SVM are based on the statistics of the most significant digit of quantized transform coefficients. The proposed method is tested with a generic hybrid video encoder combining motion-compensation and block coding. Results show that this method is able to discriminate up to three compression stages with high accuracy.
Simone Milani, Paolo Bestagini, Marco Tagliasacchi, Stefano Tubaro
MMSP3
2012 Secure image databases through distributed source coding of SIFT descriptors
abstract
The adoption of distributed databases calls for storing data at two or more sites, in order to address application-specific requirements including, e.g., redundancy, data locality, and so on. When visual data (including images, videos and their corresponding descriptors) need to be stored, synchronization across different sites might require significant bandwidth resources. In this paper we explore the use of distributed source coding to encode local SIFT descriptors extracted from static images. The key tenet is to exploit, at the decoder side, the correlation between matching pairs of descriptors extracted, respectively, from the out-of-date and up-to-date image. Preliminary results show that a coding efficiency gain up to 1 bit/descriptor can be achieved in the case of ideal lossless coding. In the case of distributed source coding with LDPC codes, a practical average gain of 0.19 bit/descriptor is observed.
Athira Nambiar, Marco Tagliasacchi, Enrico Magli
MMSP2
2012 Low bitrate coding schemes for local image descriptors
abstract
Efficient coding of local image descriptors is of paramount importance when they need to be transmitted to a remote destination on bandwidth constrained networks. This is a case that arises, e.g., in mobile visual search and visual wireless sensor networks. In this work we consider SURF, a popular descriptor suitable for low-complexity devices, and we provide a comparative study of lossy coding schemes operating at low bitrate (e.g., less than 128 bits / descriptor). Our investigation covers schemes that address both intra- and inter-descriptor redundancy, including methods that have not been tested before in this context, e.g., sparse coding, lifting-based coding on trees, and hybrid intra and inter-descriptor coding. The experimental evaluation is carried out on two publicly available datasets, in terms of both rate-distortion and rate-accuracy, for the specific task of object recognition. Our results show that a rate saving of 15-30% can be achieved by exploiting intra-descriptor redundancy. On the other side, addressing inter-descriptor redundancy does not lead to substantial gains when applied alone, whereas it leads to marginal gains (up to 3%) when used in hybrid schemes jointly with intra-descriptor coding.
Alessandro Redondi, Matteo Cesana, Marco Tagliasacchi
MMSP3
2012 Top-k bounded diversification
abstract
This paper investigates diversity queries over objects embedded in a low-dimensional vector space. An interesting case is provided by spatial Web objects, which are produced in great quantity by location-based services that let users attach content to places, and arise also in trip planning, news analysis, and real estate scenarios. The targeted queries aim at retrieving the best set of objects relevant to given user criteria and well distributed over a region of interest. Such queries are a particular case of diversified top-k queries, for which existing methods are too costly, as they evaluate diversity by accessing and scanning all relevant objects, even if only a small subset is needed. We therefore introduce Space Partitioning and Probing (SPP), an algorithm that minimizes the number of accessed objects while finding exactly the same result as MMR, the most popular diversification algorithm. SPP belongs to a family of algorithms that rely only on score-based and distance-based access methods, which are available in most geo-referenced Web data sources, and do not require retrieving all the relevant objects. Experiments show that SPP significantly reduces the number of accessed objects while incurring a very low computational overhead.
Piero Fraternali, Davide Martinenghi, Marco Tagliasacchi
SIGMOD Conference3
2012 No-Reference Pixel Video Quality Monitoring of Channel-Induced Distortion
abstract
Video transmitted over an error-prone network may be received at the decoder with degradations due to packet losses. No-reference quality monitoring algorithms are the most practical way to measure the quality of the received video, since they do not impose any change with respect to the network architecture. Conventionally, these methods assume the availability of the corrupted bitstream. In some situations this is not possible, e.g., because the bitstream is encrypted or processed by third-party decoders, and only the decoded pixel values can be used. The major issue in this scenario is the lack of knowledge about which regions of the video have been actually lost, which is a fundamental ingredient for estimating channel-induced distortion. In this paper, we propose a maximum a posteriori estimation of the pattern of lost macroblocks, which assumes the knowledge of the decoded pixels only. This information can be used as input to a no-reference quality monitoring system, which produces an accurate estimate of the mean-square-error (MSE) distortion introduced by channel errors. The results of the proposed method are well correlated with the MSE distortion computed in full-reference mode, with a linear correlation coefficient equal to 0.9 at frame level and 0.98 at sequence level.
Giuseppe Valenzise, Stefano Magni, Marco Tagliasacchi, Stefano Tubaro
IEEE Trans. Circuits Syst. Video Technol.3
2012 Cost-Aware Rank Join with Random and Sorted Access
abstract
In this paper, we address the problem of joining ranked results produced by two or more services on the web. We consider services endowed with two kinds of access that are often available: 1) sorted access, which returns tuples sorted by score; 2) random access, which returns tuples matching a given join attribute value. Rank join operators combine objects of two or more relations and output the k combinations with the highest aggregate score. While the past literature has studied suitable bounding schemes for this setting, in this paper we focus on the definition of a pulling strategy, which determines the order of invocation of the joined services. We propose the Cost-Aware with Random and Sorted access (CARS) pulling strategy, which is derived at compile-time and is oblivious of the query-dependent score distributions. We cast CARS as the solution of an optimization problem based on a small set of parameters characterizing the joined services. We validate the proposed strategy with experiments on both real and synthetic data sets. We show that CARS outperforms prior proposals and that its overall access cost is always within a very short margin from that of an oracle-based optimal strategy. In addition, CARS is shown to be robust w.r.t. the uncertainty that may characterize the estimated parameters.
Davide Martinenghi, Marco Tagliasacchi
IEEE Trans. Knowl. Data Eng.2
2012 Proximity measures for rank join
abstract
We introduce the proximity rank join problem, where we are given a set of relations whose tuples are equipped with a score and a real-valued feature vector. Given a target feature vector, the goal is to return the K combinations of tuples with high scores that are as close as possible to the target and to each other, according to some notion of distance or dissimilarity. The setting closely resembles that of traditional rank join, but the geometry of the vector space plays a distinctive role in the computation of the overall score of a combination. Also, the input relations typically return their results either by distance from the target or by score. Because of these aspects, it turns out that traditional rank join algorithms, such as the well-known HRJN , have shortcomings in solving the proximity rank join problem, as they may read more input than needed. To overcome this weakness, we define a tight bound (used as a stopping criterion) that guarantees instance optimality, that is, an I/O cost is achieved that is always within a constant factor of optimal. The tight bound can also be used to drive an adaptive pulling strategy, deciding at each step which relation to access next. For practically relevant classes of problems, we show how to compute the tight bound efficiently. An extensive experimental study validates our results and demonstrates significant gains over existing solutions.
Davide Martinenghi, Marco Tagliasacchi
ACM Trans. Database Syst.2
2011 Diversification for multi-domain result sets
abstract
Multi-domain search answers to queries spanning multiple entities, like "Find an affordable house in a city with low criminality index, good schools and medical services", by producing ranked sets of entity combinations that maximize relevance, measured by a function expressing the user's preferences. Due to the combinatorial nature of results, good entity instances (e.g., inexpensive houses) tend to appear repeatedly in top-ranked combinations. To improve the quality of the result set, it is important to balance relevance (i.e., high values of the ranking function) with diversity, which promotes different, yet almost equally relevant, entities in the top-k combinations. This paper explores two different notions of diversity for multi-domain result sets, compares experimentally alternative algorithms for the trade-off between relevance and diversity, and performs a user study for evaluating the utility of diversification in multi-domain queries.
Alessandro Bozzon, Marco Brambilla 0001, Piero Fraternali, Marco Tagliasacchi
CIKM4
2011 The cost of JPEG compression anti-forensics
abstract
The statistical footprint left by JPEG compression can be a valuable source of information for the forensic analyst. Recently, it has been shown that a suitable anti-forensic method can be used to destroy these traces, by properly adding a noise-like signal to the quantized DCT coefficients. In this paper we analyze the cost of this technique in terms of introduced distortion and loss of image quality. We characterize the dependency of the distortion on the image statistics in the DCT domain and on the quantization step used in JPEG compression. We also evaluate the loss of quality as measured by means of a perceptual metric, showing that a perceptually-optimized version of the anti-forensic method fails to completely conceal the forgery. Our conclusion is that removing the traces of the JPEG compression history could be much more challenging than it might appear, as anti-forensic methods are bound to leave characteristic traces.
Giuseppe Valenzise, Marco Tagliasacchi, Stefano Tubaro
ICASSP2
2011 Countering JPEG anti-forensics
abstract
JPEG coding leaves characteristic footprints that can be leveraged to reveal doctored images, e.g. providing the evidence for local tampering, copy-move forgery, etc. Recently, it has been shown that a knowledgeable attacker might attempt to remove such footprints by adding a suitable anti-forensic dithering signal to the image in the DCT domain. Such noise-like signal restores the distribution of the DCT coefficients of the original picture, at the cost of affecting image quality. In this paper we show that it is possible to detect this kind of attack by measuring the noisiness of images obtained by re-compressing the forged image at different quality factors. When tested on a large set of images, our method was able to correctly detect forged images in 97% of the cases. In addition, the original quality factor could be accurately estimated.
Giuseppe Valenzise, Vitaliano Nobile, Marco Tagliasacchi, Stefano Tubaro
ICIP3
2011 Parallel Data Access for Multiway Rank Joins
Adnan Abid, Marco Tagliasacchi
ICWE2
2011 Semantically improved genome-wide prediction of Gene Ontology annotations
abstract
Genomic annotations describing structural and functional features of genes and gene products through controlled terminologies and ontologies are extremely valuable, especially for computational analyses aimed at inferring new biomedical knowledge, which rely on available annotations. Yet, they are incomplete, especially for recently studied genomes, and only some of available annotations represent highly reliable human curated information. In order to help and speedup the time-consuming curation process and improve available annotations, computational methods able to provide prioritized lists of predicted annotations are paramount. Starting from a previous work on automatic prediction of Gene Ontology annotations based on singular value decomposition (SVD) of gene-to-term annotation matrix, here we propose a novel prediction algorithm that incorporates gene clustering based on gene functional similarity computed on Gene Ontology annotations. We tested both prediction methods performing k-fold cross-validation on two organism genomes, Saccharomyces cerevisiae (SGD) and Drosophila melanogaster (FlyBase). Results demonstrate effectiveness of our approach.
Marco Masseroli, Marco Tagliasacchi, Davide Chicco
ISDA2
2011 Ranking with uncertain scoring functions: semantics and sensitivity measures
abstract
Ranking queries report the top-K results according to a user-defined scoring function. A widely used scoring function is the weighted summation of multiple scores. Often times, users cannot precisely specify the weights in such functions in order to produce the preferred order of results. Adopting uncertain/incomplete scoring functions (e.g., using weight ranges and partially-specified weight preferences) can better capture user's preferences in this scenario.
Mohamed A. Soliman, Ihab F. Ilyas, Davide Martinenghi, Marco Tagliasacchi
SIGMOD Conference4
2010 Prediction of Gene Ontology Annotations Based on Gene Functional Clustering
abstract
We propose an algorithm that predicts potentially missing Gene Ontology annotations, in order to speed up the time-consuming annotation curation process. The proposed method extends a previous work based on the singular value decomposition of the gene-term annotation matrix and incorporates gene clustering, based on gene functional similarity computed by means of the Gene Ontology annotations. We tested the prediction method by performing K-fold cross-validation on the genomes of two organisms, Saccharomyces cerevisiae (SGD) and Drosophila melanogaster (FlyBase).
Marco Tagliasacchi, Roberto Sarati, Marco Masseroli
BIBE1
2010 A H.264/AVC video database for the evaluation of quality metrics
abstract
This paper describes a publicly available database of subjective scores, relative to quality assessment of 156 video streams encoded with H.264/AVC and corrupted by simulating packet losses over an error-prone network. The data has been collected in controlled test environments at the premises of two academic institutions. A detailed statistical analysis of subjective results has been performed, showing high consistency of the collected scores. In addition to subjective scores, we have made available to the research community both the uncompressed files and the H.264/AVC bitstreams of each video sequence, in order to provide a common database of benchmark data to test and compare the performance of full-reference, reduced-reference and no-reference video quality assessment algorithms.
Francesca De Simone, Marco Tagliasacchi, Matteo Naccari, Stefano Tubaro, Stefano Ebrahimi
ICASSP2
2010 Geometric calibration of distributed microphone arrays from acoustic source correspondences
abstract
This paper proposes a method that solves the problem of geometric calibration of microphone arrays. We consider a distributed system, in which each array is controlled by separate acquisition devices that do not share a common synchronization clock. Given a set of probing sources, e.g. loudspeakers, each array computes an estimate of the source locations using a conventional TDOA-based algorithm. These observations are fused together by the proposed method, in order to estimate the position and pose of one array with respect to the other. Unlike previous approaches, we explicitly consider the anisotropic distribution of localization errors. As such, the proposed method is able to address the problem of geometric calibration when the probing sources are located both in the near- and far-field of the microphone arrays. Experimental results demonstrate that the improvement in terms of calibration accuracy with respect to state-of-the-art algorithms can be substantial, especially in the far-field.
S. Daniele Valente, Marco Tagliasacchi, Fabio Antonacci, Paolo Bestagini, Augusto Sarti, Stefano Tubaro
MMSP2
2010 A reduced-reference structural similarity approximation for videos corrupted by channel errors
Marco Tagliasacchi, Giuseppe Valenzise, Matteo Naccari, Stefano Tubaro
Multim. Tools Appl.1
2010 Proximity Rank Join
abstract
We introduce the proximity rank join problem, where we are given a set of relations whose tuples are equipped with a score and a real-valued feature vector. Given a target feature vector, the goal is to return the K combinations of tuples with high scores that are as close as possible to the target and to each other, according to some notion of distance. The setting closely resembles that of traditional rank join, but the geometry of the vector space plays a distinctive role in the computation of the overall score of a combination. Also, the input relations typically return their results either by distance from the target or by score. Because of these aspects, it turns out that traditional rank join algorithms, such as the well-known HRJN , have shortcomings in solving the proximity rank join problem, as they may read more input than needed. To overcome this weakness, we define a tight bound (used as a stopping criterion) that guarantees instance optimality, i.e., an I/O cost is achieved that is always within a constant factor of optimal. The tight bound can also be used to drive an adaptive pulling strategy, deciding at each step which relation to access next. For practically relevant classes of problems, we show how to compute the tight bound efficiently. An extensive experimental study validates our results and demonstrates significant gains over existing solutions.
Davide Martinenghi, Marco Tagliasacchi
Proc. VLDB Endow.2
2010 Joint Compressive Video Coding and Analysis
abstract
Traditionally, video acquisition, coding and analysis have been designed and optimized as independent tasks. This has a negative impact in terms of consumed resources, as most of the raw information captured by conventional acquisition devices is discarded in the coding phase, while the analysis step only requires a few descriptors of salient video characteristics. Recent compressive sensing literature has partially broken this paradigm by proposing to integrate sensing and coding in a unified architecture composed by a light encoder and a more complex decoder, which exploits sparsity of the underlying signal for efficient recovery. However, a clear understanding of how to embed video analysis in this scheme is still missing. In this paper, we propose a joint compressive video coding and analysis scheme and, as a specific application example, we consider the problem of object tracking in video sequences. We show that, weaving together compressive sensing and the information computed by the analysis module, the bit-rate required to perform reconstruction and tracking of the foreground objects can be considerably reduced, with respect to a conventional disjoint approach that postpones the analysis after the video signal is recovered in the pixel domain. These findings suggest that considerable gains in performance can be potentially obtained in video analysis applications, provided that a joint analysis-aware design of acquisition, coding and signal recovery is carried out.
M. Cossalter, Giuseppe Valenzise, Marco Tagliasacchi, Stefano Tubaro
IEEE Trans. Multim.3
2009 Privacy-Enabled Object Tracking in Video Sequences Using Compressive Sensing
abstract
In a typical video analysis framework, video sequences are decoded and reconstructed in the pixel domain before being processed for high level tasks such as classification or detection.Nevertheless, in some application scenarios, it might be of interest to complete these analysis tasks without disclosing sensitive data, e.g. the identity of people captured by surveillance cameras. In this paper we propose a new coding scheme suitable for video surveillance applications that allows tracking of video objects without the need to reconstruct the sequence,thus enabling privacy protection. By taking advantage of recent findings in the compressive sensing literature, we encode a video sequence with a limited number of pseudo-random projections of each frame. At the decoder, we exploit the sparsity that characterizes background subtracted images in order to recover the location of the foreground object. We also leverage the prior knowledge about the estimated location of the object, which is predicted by means of a particle filter, to improve the recovery of the foreground object location. The proposed framework enables privacy, in the sense it is impossible to reconstruct the original video content from the encoded random projections alone, as well as secrecy, since decoding is prevented if the seed used to generate the random projections is not available.
M. Cossalter, Marco Tagliasacchi, Giuseppe Valenzise
AVSS2
2009 Anomaly-free Prediction of Gene Ontology Annotations Using Bayesian Networks
abstract
Gene and protein structural and functional annotations expressed through controlled terminologies and ontologies are paramount especially for the aim of inferring new biomedical knowledge through computational analyses. However, the available annotations are incomplete, in particular for recently studied genomes, and only a few of them are highly reliable human curated information. To support and speed up the time-consuming curation process, prioritized lists of computationally predicted annotations are hence extremely useful. In this paper we leverage a previous work on the automatic prediction of gene ontology annotations based on the singular value decomposition (SVD) of the gene-to-term annotation matrix, and we propose a novel post-processing method that uses a Bayesian network to eliminate predictions of anomalous annotations. In fact, we observed that the predicted annotation profiles might suggest that a gene shall be annotated to a term, but not to one of its ancestors, thus violating the constraint imposed by the gene ontology. To this end, the proposed algorithm processes the annotation profiles predicted by a SVD based method, and produces a ranked list of computationally discovered candidate annotations which is consistent with the gene ontology.
Marco Tagliasacchi, Marco Masseroli
BIBE1
2009 A reduced-reference video structural similarity metric based on no-reference estimation of channel-induced distortion
abstract
The reduced-reference (RR) approximation of a full-reference (FR) video quality assessment method is a convenient way to build evaluation metrics which are both intrinsically well correlated with human judgments and feasible to implement in a network scenario, without the need to explore the perceptual significance of new video features through mean opinion score tests. In this paper, we propose a RR approximation of the video structural similarity index (VSSIM), a FR metric which is known to be well descriptive of the video quality perceived by users. We focus on the visual degradation produced by channel transmission errors: first, at the encoder, a small set of salient structural video features is assembled and transmitted through the RR channel to the end-user; then, at the decoder the feature vector is combined with a fine-granularity, no-reference estimate of the channel-induced distortion to produce the VSSIM approximation. By uniformly quantizing the feature vector and compressing it using a context-adaptive, variable length encoder, we show that good correlation coefficients with ground-truth VSSIM (rho = 0.85) may be achieved spending, respectively, less than 12 and 27 kbps for a video sequence with CIF or SD resolution.
Andrea Albonico, Giuseppe Valenzise, Matteo Naccari, Marco Tagliasacchi, Stefano Tubaro
ICASSP4
2009 Subjective evaluation of a NO-reference video quality Monitoring algorithm for H.264/AVC video over a noisy channel
abstract
In this paper we evaluate NORM, a no-reference video quality monitoring algorithm we proposed in a previous work, for the prediction of the subjective quality of H.264/AVC video transmitted over a noisy packet-switched network. NORM produces an estimate of the mean square error distortion at the macroblock level between the noiseless and noisy sequence, without having access to the former. The output of NORM can be readily converted into a no-reference estimate of the PSNR at the sequence level. We carried out an extensive subjective evaluation campaign on CIF and 4CIF resolution sequences encoded with H.264/AVC and transmitted over a channel that drops packets at different packet loss rates, to obtain the differential mean opinion scores. Our results show that the estimated PSNR achieves good correlation with the subjective scores, very close to the ones achieved by the PSNR computed in full-reference mode, i.e. as if the noiseless sequence would be available at the decoder.
Matteo Naccari, Marco Tagliasacchi, Stefano Tubaro
ICIP2
2009 A compressive-sensing based watermarking scheme for sparse image tampering identification
abstract
In this paper we describe a robust watermarking scheme for image tampering identification and localization. A compact representation of the image is first produced by assembling a feature vector consisting of pseudo-random projections of the decimated image. Then, the quantized projections are encoded to form a hash, which is robustly embedded as a watermark in the image. By recovering the watermark the random projections are obtained, and then used to estimate the distortion of the received image. If tampering is sufficiently sparse or compressible in some basis description, a map of the introduced modification is recovered. The system relies on compressive sensing and distributed source coding principles to reduce the size of the hash of a 1024 × 1024 image, to about 4,000 bits. With this hash length, tampering sparse up to 20% and with a tampering energy around a PSNR of 15 dB can be successfully localized.
Giuseppe Valenzise, Marco Tagliasacchi, Stefano Tubaro, Giacomo Cancelli, Mauro Barni
ICIP2
2009 Geometric calibration of distributed microphone arrays
abstract
Computational auditory scene analysis exploits signals acquired by means of microphone arrays. In some circumstances, more than one array is deployed in the same environment. In order to effectively fuse the information gathered by each array, the relative location and pose of the arrays needs to be obtained solving a problem of geometric inter-array calibration. We consider the case where the arrays do not share a synchronous clock, which impairs the use of time-difference of arrival measures across arrays. Conversely, each array produces an acoustic image, which describes the energy of acoustic signals received from different directions. We jointly consider acoustic images acquired by the different arrays and adapt computer vision techniques to solve the calibration problem, thus estimating the location and pose of microphone arrays sensing the same auditory scene. We evaluate the robustness of the calibration process in a simulated environment and we investigate the effect of the various system parameters, namely the number of probing signal locations, the resolution of the acoustic images, the non-ideal intra-array calibration.
Alessandro Redondi, Marco Tagliasacchi, Fabio Antonacci, Augusto Sarti
MMSP2
2009 Hash-Based Identification of Sparse Image Tampering
abstract
In the last decade, the increased possibility to produce, edit, and disseminate multimedia contents has not been adequately balanced by similar advances in protecting these contents from unauthorized diffusion of forged copies. When the goal is to detect whether or not a digital content has been tampered with in order to alter its semantics, the use of multimedia hashes turns out to be an effective solution to offer proof of legitimacy and to possibly identify the introduced tampering. We propose an image hashing algorithm based on compressive sensing principles, which solves both the authentication and the tampering identification problems. The original content producer generates a hash using a small bit budget by quantizing a limited number of random projections of the authentic image. The content user receives the (possibly altered) image and uses the hash to estimate the mean square error distortion between the original and the received image. In addition, if the introduced tampering is sparse in some orthonormal basis or redundant dictionary, an approximation is given in the pixel domain. We emphasize that the hash is universal, e.g., the same hash signature can be used to detect and identify different types of tampering. At the cost of additional complexity at the decoder, the proposed algorithm is robust to moderate content-preserving transformations including cropping, scaling, and rotation. In addition, in order to keep the size of the hash small, hash encoding/decoding takes advantage of distributed source codes.
Marco Tagliasacchi, Giuseppe Valenzise, Stefano Tubaro
IEEE Trans. Image Process.1
2009 No-Reference Video Quality Monitoring for H.264/AVC Coded Video
abstract
When video is transmitted over a packet-switched network, the sequence reconstructed at the receiver side might suffer from impairments introduced by packet losses, which can only be partially healed by the action of error concealment techniques. In this context we propose NORM (NO-Reference video quality Monitoring), an algorithm to assess the quality degradation of H.264/AVC video affected by channel errors. NORM works at the receiver side where both the original and the uncorrupted video content is unavailable. We explicitly account for distortion introduced by spatial and temporal error concealment together with the effect of temporal motion-compensation. NORM provides an estimate of the mean square error distortion at the macroblock level, showing good linear correlation (correlation coefficient greater than 0.80) with the distortion computed in full-reference mode. In addition, the estimate at the macroblock level can be successfully exploited by forward quality monitoring systems that compute quality objective metrics to predict mean opinion score (MOS) values. As a proof of concept, we feed the output of NORM to a reduced-reference quality monitoring system that computes an estimate of the structural similarity metric (SSIM) score, which is known to be well correlated with perceptual quality.
Matteo Naccari, Marco Tagliasacchi, Stefano Tubaro
IEEE Trans. Multim.2
2008 Resource constrained efficient acoustic source localization and tracking using a distributed network of microphones
abstract
In this paper we present an efficient method to perform acoustic source localization and tracking using a distributed network of microphones. In this scenario, there is a trade-off between the localization performance and the expense of resources: in fact, a minimization of the localization error would require to use as many sensors as possible; at the same time, as the number of microphones increases, the cost of the network inevitably tends to grow, while in practical applications only a limited amount of resources is available. Therefore, at each time instant only a subset of the sensors should be enabled in order to meet the cost constraints. We propose a heuristic method for the optimal selection of this subset of microphones, using as distortion metrics the Cramer-Rao lower bound (CRLB) and as cost function the total distance between the selected sensors. The heuristic approach has been compared to an optimal algorithm, which searches the best sensor configuration among the full set of microphones, while satisfying the cost constraint. The proposed heuristic algorithm yields similar performance w.r.t. the full-search procedure, but at a much less computational cost. We show that this method can be used effectively in an acoustic source tracking application.
Giuseppe Valenzise, Giorgio Prandi, Marco Tagliasacchi, Augusto Sarti
ICASSP3
2008 Minimum variance multiplexing of multimedia objects
abstract
This paper addresses the problem of simultaneous transmission of multiple multimedia objects (such as images or video sequences) over a bandwidth-limited channel. The trivial strategy of partitioning in equal parts the available rate among the bitstreams is suboptimal, when the multimedia objects have different coding complexities. Exploiting object diversity allows us to allocate the bandwidth according to some optimality criteria, e.g. minimizing the average total distortion or minimizing the variance between the distortions of each object. By describing the rate-distortion characteristics of each multimedia object in terms of a simple exponential model, we provide a closed form solution for both the minimum average and the minimum variance problems. In addition, if we consider the statistical distribution of the rate-distortion model parameters, we can show that the minimum variance solution can effectively reduce the quality fluctuations among the objects, with an overall coding efficiency loss, w.r.t. the minimum average solution, of only 0.5dB on average. Some experiments, carried out on different H.264/AVC video sequences, validate our theoretical results.
Giuseppe Valenzise, Marco Tagliasacchi, Stefano Tubaro
ICASSP2
2008 No-reference modeling of the channel induced distortion at the decoder for H.264/AVC video coding
abstract
This paper proposes a model for estimating, at the decoder side, the distortion induced by the transmission over an error-prone channel, when error-free reconstructed frames are not available as a reference. The proposed estimation model considers explicitly the temporal error concealment algorithm adopted at the decoder. In the evaluation of the induced distortion, we model the effects of the absence of motion vectors and prediction residuals in the decoding process. In addition, we take into account the error propagation along successive frames. Experimental results conducted over real video sequences coded with the state-of-art H.264/AVC video coding standard validate the proposed model. In fact, the distortion estimated when no reference is available is strongly correlated both at the frame and group of pictures level with the actual distortion. This technique represents an effective no-reference video quality monitoring tool that can be embedded in any H.264/AVC compliant decoder.
Matteo Naccari, Marco Tagliasacchi, Fernando Pereira 0001, Stefano Tubaro
ICIP2
2008 Localization of sparse image tampering via random projections
abstract
Hashes can be used to provide authentication of multimedia contents. In the case of images, a hash can be used to detect whether the data has been modified in an illegitimate way. When the authentication check fails, it might be useful to localize the tampering in the spatial domain. This paper proposes an algorithm based on compressive sensing principles, which solves both the authentication and the localization problems. The encoder produces a hash using a small bit budget by quantizing a limited number of random projections of the authentic image. The decoder uses the hash to estimate the distortion between the original and the received image. In addition, if the attack is sparse, it can be also localized. In order to keep the size of the hash small, encoding/decoding takes advantage of distributed source codes. This paper also investigates experimentally the tradeoff between the rate allocated to the hash and the performance achieved in terms of tampering localization.
Marco Tagliasacchi, Giuseppe Valenzise, Stefano Tubaro
ICIP1
2008 Wyner-Ziv video coding: A review of the early architectures and further developments
abstract
In 2002, the video coding community faced the emergence of a new video coding paradigm, the so-called Wyner-Ziv video coding, which was represented by two early solutions designed by the Stanford University and the University of California, Berkeley research teams. This paper intends to briefly review, and compare these two early Wyner-Ziv video coding solutions, notably from the functional point of view. Moreover, this paper reviews some important developments of the Stanford Wyner-Ziv coding architecture, which has become the most popular in the literature.
Fernando Pereira 0001, Catarina Brites, João Ascenso, Marco Tagliasacchi
ICME4
2008 Reduced-reference estimation of channel-induced video distortion using distributed source coding
abstract
Channel-induced distortion estimation is an important aspect in the delivery of video contents over IP networks: the QoS requirements of both content providers and content users conflict with the intrinsic best-effort nature of packet-switched networks, which may introduce annoying artifacts in the received streams due to channel errors or jitter. In this paper we propose a Reduced-Reference video quality assessment method, based on objective quality metrics, which enables distortion estimation at the macroblock level. The content provider transmits a small feature vector for each frame, starting from random projections computed for each macroblock. In order to reduce the bit rate of the transmitted feature vector, we encode it using Distributed Source Coding (DSC) tools. The content user decodes the feature vector using the received sequence as side information. Additionally, the end-user may take advantage of some prior information about the support of the errors in the frame in such a way that the required bit length of the transmitted feature vector is further reduced. In our experiments, using 4 random projections, the use of DSC enables a bit saving of 70% w.r.t. scalar quantization and transmission of the original feature vector; when also the a priori error map is available at the decoder, the average length of the transmitted partial reference can be further reduced by another 5% of average.
Giuseppe Valenzise, Matteo Naccari, Marco Tagliasacchi, Stefano Tubaro
ACM Multimedia3
2008 Rate allocation for robust video streaming based on distributed video coding
Riccardo Bernardini, Matteo Naccari, Roberto Rinaldo, Marco Tagliasacchi, Stefano Tubaro, Pamela Zontone
Signal Process. Image Commun.4
2008 Minimum Variance Optimal Rate Allocation for Multiplexed H.264/AVC Bitstreams
abstract
Consider the problem of transmitting multiple video streams to fulfill a constant bandwidth constraint. The available bit budget needs to be distributed across the sequences in order to meet some optimality criteria. For example, one might want to minimize the average distortion or, alternatively, minimize the distortion variance, in order to keep almost constant quality among the encoded sequences. By working in the rho-domain, we propose a low-delay rate allocation scheme that, at each time instant, provides a closed form solution for either the aforementioned problems. We show that minimizing the distortion variance instead of the average distortion leads, for each of the multiplexed sequences, to a coding penalty less than 0.5 dB, in terms of average PSNR. In addition, our analysis provides an explicit relationship between model parameters and this loss. In order to smooth the distortion also along time, we accommodate a shared encoder buffer to compensate for rate fluctuations. Although the proposed scheme is general, and it can be adopted for any video and image coding standard, we provide experimental evidence by transcoding bitstreams encoded using the state-of-the-art H.264/AVC standard. The results of our simulations reveal that is it possible to achieve distortion smoothing both in time and across the sequences, without sacrificing coding efficiency.
Marco Tagliasacchi, Giuseppe Valenzise, Stefano Tubaro
IEEE Trans. Image Process.1
2007 Tracking of two acoustic sources in reverberant environments using a particle swarm optimizer
abstract
In this paper we consider the problem of tracking multiple acoustic sources in reverberant environments. The solution that we propose is based on the combination of two techniques. A blind source separation (BSS) method known as TRINICON [5] is applied to the signals acquired by the microphone arrays. The TRINICON de-mixing filters are used to obtain the Time Differences of Arrival (TDOAs), which are related to the source location through a nonlinear function. A particle filter is then applied in order to localize the sources. Particles move according to a swarm-like dynamics, which significatively reduces the number of particles involved with respect to traditional particle filter. We discuss results for the case of two sources and four microphone pairs. In addition, we propose a method, based on detecting source inactivity, which overcomes the ambiguities that intrinsically arise when only two microphone pairs are used. Experimental results demonstrate that the average localization error on a variety of pseudo-random trajectories is around 40 cm when the T60reverberation time is 0.6s.
Fabio Antonacci, Davide Riva, Augusto Sarti, Marco Tagliasacchi, Stefano Tubaro
AVSS4
2007 Scream and gunshot detection and localization for audio-surveillance systems
abstract
This paper describes an audio-based video surveillance system which automatically detects anomalous audio events in a public square, such as screams or gunshots, and localizes the position of the acoustic source, in such a way that a video-camera is steered consequently. The system employs two parallel GMM classifiers for discriminating screams from noise and gunshots from noise, respectively. Each classifier is trained using different features, chosen from a set of both conventional and innovative audio features. The location of the acoustic source which has produced the sound event is estimated by computing the time difference of arrivals of the signal at a microphone array and using linear-correction least square localization algorithm. Experimental results show that our system can detect events with a precision of 93% at a false rejection rate of 5% when the SNR is 10dB, while the source direction can be estimated with a precision of one degree. A real-time implementation of the system is going to be installed in a public square of Milan.
Giuseppe Valenzise, Luigi Gerosa, Marco Tagliasacchi, Fabio Antonacci, Augusto Sarti
AVSS3
2007 Hash-Based Motion Modeling in Wyner-Ziv Video Coding
abstract
Generally, distributed video coding (DVC) schemes perform motion estimation at the decoder side, without the current frame being available. In order to generate the side-information reliably, one solution consists in allocating a limited bit budget to send a hash of the current frame. At the decoder, this auxiliary hash is used to perform motion estimation. This paper studies the accuracy of hash-based motion estimation and compares it to conventional encoder-side motion estimation. We show that, at low rates, the very limited bit-budget of the hash does not ensure a reliable motion estimation, while at medium to high rates the motion accuracy is comparable with the finite precision used to represent motion vectors. Then, we derive the rate-distortion characteristic, which combines the cost of encoding the hash and the prediction residuals after decoder-side motion compensation. We show that, at high rates, hash-based motion modeling can virtually achieve the same coding efficiency as motion-compensated predictive coding. Instead, at medium-to-low rates we observe a significant coding loss. Experimental results on real video sequences validate the results of the proposed model.
Marco Tagliasacchi, Stefano Tubaro
ICASSP (1)1
2007 Analysis of Coding Efficiency of Motion-Compensated Interpolation at the Decoder in Distributed Video Coding
abstract
This paper analyzes the coding efficiency of distributed video coding (DVC) schemes that perform motion-compensated interpolation at the decoder. The decoder has access only to the key frames when generating the side information for intermediate frames. This fact introduces a displacement estimation error that depends on several factors: 1) the overall motion complexity; 2) the temporal coherence of the motion field; 3) the temporal distance between successive key frames. Adopting a state-space model and a Kalman filtering framework, we obtain an estimate of the displacement error variance. This is used to determine the rate-distortion function of the overall coding scheme, that takes into account both intra-coded key frames and DVC-coded frames. The proposed model shows that motion-compensated interpolation is unable to achieve the coding efficiency of conventional motion-compensated predictive coding.
Marco Tagliasacchi, Laura Frigerio, Stefano Tubaro
ICIP (3)1
2007 Symmetric Distributed Coding of Stereo Video Sequences
abstract
In this paper we present a novel video coding scheme to compress stereo video sequences. We consider a wireless sensor network scenario, where the sensing nodes cannot communicate with each other and are characterized by limited computational complexity. The joint decoder exploits both the temporal and inter-view correlation to generate the side information. To this end, we propose a fusion algorithm that adaptively selects either the temporal or the inter-view side information on a pixel-by-pixel basis. In addition, the coding algorithm is symmetric with respect to the two cameras. We also propose a practical stopping criterion for turbo decoding that determines when decoding is successful. Experimental results on stereo video sequences show that a coding efficiency gain up to 4dB can be obtained by the proposed scheme at high bit-rates.
Marco Tagliasacchi, Giorgio Prandi, Stefano Tubaro
ICIP (2)1
2007 A genetic algorithm for optical flow estimation
Marco Tagliasacchi
Image Vis. Comput.1
2007 Rate-Distortion Analysis of Motion-Compensated Interpolation at the Decoder in Distributed Video Coding
abstract
This letter analyzes the coding efficiency of distributed video coding (DVC) schemes that perform motion-compensated interpolation at the decoder. The decoder has access only to the key frames, when generating the side information for intermediate frames. Therefore, the true motion field necessary for this operation is not directly available, and the motion vectors must be estimated at the decoder side, thus introducing displacement estimation errors. The accuracy of the motion-compensated interpolation at the decoder depends on several factors: 1, the overall motion complexity; 2, the temporal coherence of the motion field; and 3, the temporal distance between successive key frames. Adopting a state-space model and a Kalman filtering framework, we obtain an estimate of the displacement error variance. This is used to determine the rate-distortion function of the overall coding scheme, that takes into account both intra-coded key frames and DVC-coded frames. The proposed model shows that motion-compensated interpolation is unable to achieve the coding efficiency of conventional motion-compensated predictive coding. In addition, the model provides a good estimate of the group of pictures size that optimizes the coding efficiency. Experimental results on real video sequences validate the results of the proposed model.
Marco Tagliasacchi, Laura Frigerio, Stefano Tubaro
IEEE Signal Process. Lett.1
2006 A Proposal to Suppress the Training Stage in a Coset-Based Distributed Video Codec
abstract
Distributed video coding (DVC) is a coding paradigm that gives the decoder the task to exploit the source statistics to achieve efficient compression. Many approaches to the DVC problem have recently appeared in the literature, including the PRISM codec. Instead of encoding the deterministic quantized prediction error residual, PRISM partitions the quantization lattice into cosets and sends the index of the coset each quantized coefficient belongs to. Estimating the number of cosets is of crucial importance to achieve good coding efficiency. In PRISM, this is determined during an offline training phase. The present work aims at being a starting point for the suppression of the training stage of PRISM at the cost of sending the number of cosets for each DCT coefficient. The statistics of the number of cosets are analyzed to figure out the maximum compression efficiency achievable by entropy coding. Furthermore the paper discusses some techniques that might be used to lower the amount of transmitted bits. Based on these results, directions for future works are proposed
Xavier Artigas, Marco Tagliasacchi, Stefano Tubaro
ICASSP (2)2
2006 Improved Bit Allocation in an Error-Resilient Scheme Based on Distributed Source Coding
abstract
In this work we propose an error-resilient scheme that allows enhancing the robustness of a video stream. Based on distributed source coding (DSC) principles, an auxiliary stream is sent in parallel to the main stream as a redundant representation of the sequence that is used to correct errors at the decoder, thus reducing the impact of drift. In order to perform an optimal bit allocation in the auxiliary stream, the encoder needs to compute a reliable estimate of the expected video distortion observed at the decoder side due to channel loss. This paper proposes an algorithm to calculate the expected distortion of decoded DCT-coefficients (dubbed EDDD) and its application to the bit allocation problem in a DSC based auxiliary stream
Marco Fumagalli, Marco Tagliasacchi, Stefano Tubaro
ICASSP (2)2
2006 Intra Mode Decision Based on Spatio-Temporal Cues in Pixel Domain Wyner-ZIV Video Coding
abstract
Distributed source coding principles have been recently applied to video coding in order to achieve a flexible distribution of the complexity burden between the encoder and the decoder. In this paper we elaborate on a pixel based Wyner-Ziv video codec that shifts all the complexity of the motion estimation phase to the decoder, thus achieving light encoding. We observe that the correlation noise statistics describing the relationship between the frame to be encoded and the side information available at the decoder is not spatially stationary. For this reason we introduce a mode decision scheme either at the encoder or at the decoder in such a way that when the estimated correlation is weak we opt for intra coding on a block-by-block basis. Both spatial and temporal criteria are used to determine whether a block is better intra coded or not
Marco Tagliasacchi, Alan Trapanese, Stefano Tubaro, João Ascenso, Catarina Brites, Fernando Pereira 0001
ICASSP (2)1
2006 Video Coding with Wavelet-Domain Conditional Replenishment and Unequal Error Protection
abstract
A simple and computationally lightweight video coder employing shape-adaptive, embedded intraframe coding and wavelet-domain conditional replenishment is proposed. Robustness to packet losses arises from packetization of the embedded bitstream with unequal error protection which is assigned to the packets with a fast, locally optimal procedure. Experimental results reveal that, when compared to H.264/AVC configured for low-complexity, error-resilient operation, not only does the proposed coder usually produce substantially superior rate-distortion performance as packet losses increase, it also achieves a significantly faster encoding speed.
James E. Fowler, Marco Tagliasacchi, Béatrice Pesquet-Popescu
ICIP2
2006 On the Modeling of Motion in Wyner-Ziv Video Coding
abstract
In the past few years, a number of practical video coding schemes following distributed source coding principles have emerged. One of the main goals of distributed video coding (DVC) is to enable a flexible distribution of the computational complexity between the encoder and the decoder, while approaching the coding efficiency of conventional closed-loop motion-compensated predictive codecs. In this paper we perform a rate-distortion analysis of a well-known Wyner-Ziv architecture, while focusing our attention on the impact of the motion modeling that is used for generating the side information at the decoder. Our analysis is structured according to a Kalman filtering problem and it allows us to compare three different scenarios: motion estimation at the encoder; motion interpolation at the decoder; and motion extrapolation at the decoder.
Marco Tagliasacchi, Stefano Tubaro, Augusto Sarti
ICIP1
2006 Exploiting Spatial Redundancy in Pixel Domain Wyner-Ziv Video Coding
abstract
Distributed video coding is a recent paradigm that enables a flexible distribution of the computational complexity between the encoder and the decoder building on top of distributed source coding principles. In this paper we focus on the scenario where most of the complexity is shifted to the decoder, thus achieving light encoding. We elaborate on a well known pixel based Wyner-Ziv architecture and we improve its coding efficiency by exploiting both spatial and temporal correlation at the decoder side, without the need of performing any transform at the encoder. In order to generate the side information, the decoder adaptively chooses spatial or temporal information, based on the local estimate of the correlation noise. Simulations on test sequences demonstrate that a coding gain of up to +1.8 dB can be obtained with respect to the case that generates the side information by motion interpolation only.
Marco Tagliasacchi, Alan Trapanese, Stefano Tubaro, João Ascenso, Catarina Brites, Fernando Pereira 0001
ICIP1
2006 Motion Estimation by Quadtree Pruning and Merging
abstract
In this paper we propose a rate-distortion optimized motion estimation algorithm that is built upon a quadtree structure. Each node of the quadtree represents a block in the current frame together with its motion vector, and the block size decreases from the root to the leaves. In the first step, the quadtree is pruned according to a rate-distortion criterion in order to obtain blocks of variable sizes. A further rate rebate can be achieved by merging those leaf nodes of the quadtree that can be efficiently represented by the same motion vector. The proposed merging scheme provides a reduction of up to 50% of the rate spent for the motion model with respect to the case that performs pruning only
Marco Tagliasacchi, Mauro Sarchi, Stefano Tubaro
ICME1
2006 Robust wireless video multicast based on a distributed source coding approach
Marco Tagliasacchi, Abhik Majumdar, Kannan Ramchandran, Stefano Tubaro
Signal Process.1
2005 Combining MCTF with distributed source coding
abstract
Motion compensated temporal filtering (MCTF) has proved to be an efficient coding tool in the design of open-loop scalable video codecs. In this paper we propose a MCTF video coding scheme based on lifting where the prediction step is implemented using PRISM (power efficient, robust, high compression syndrome-based multimedia coding), a video coding framework built on distributed source coding principles. We study the effect of integrating the update step at the encoder or at the decoder side. We show that the latter approach allows improving the quality of the side information exploited during decoding. We present the analytical results obtained by modeling the video signal along the motion trajectories as an AR(1) process showing that the update step at the decoder allows to half the contribution of the quantization noise. We also include experimental results with real video data that demonstrate the potential of this approach when the video sequences are coded at low bitrates.
Marco Tagliasacchi, Stefano Tubaro, Augusto Sarti
ICIP (1)1
2004 Scalable coding of variable size blocks motion vectors
abstract
In this paper we discuss an algorithm that is able to provide a scalable (multiresolution) representation of the motion field information. It has been recently demonstrated that, for an open loop wavelet video coder, it is possible to use at the decoder side a scaled version of the original motion information and the residual coefficients computed with the full resolution/quality motion field without incurring into drift. We propose a fully scalable wavelet based video coder that performs motion estimation-compensation in the wavelet domain. In particular the coding scheme is specifically designed for variable size block matching algorithms. In this scenario motion vectors are distributed across an irregular lattice according to a quadtree structure. The developed system allows a scalable representation of the motion field and a flexible allocation of the bit budget between motion and residual data. The simulations that we have carried out show, at low bit-rates, a significant gain of the proposed approach with respect to the case in which the motion information is coded lossless.
Davide Maestroni, Augusto Sarti, Marco Tagliasacchi, Stefano Tubaro
ICIP3
2004 Fast in-band motion estimation with variable size block matching
abstract
In this paper we propose a fast motion estimation technique that works in the wavelet domain. The computational cost of the algorithm turns out to be proportional to the linear size of the search window instead of its area. We complete our proposal with a variable-size block-matching scheme in the wavelet domain. We integrated the motion estimation algorithm in a fully-scalable wavelet in-band prediction coder inspired by the IB-MCTF (in-band motion compensation temporal filtering) proposed in J.C Ye et al., (2003). Comparative tests prove that our coder provides the same quality level than IB-MCTF at a very reduced computational cost. Moreover, although our method turns out to match the performance of MCTF-EZBC (P. Chen, 2003) in terms of PSNR, it clearly outperforms it in terms of perceptual quality, as it is completely free from blocking artifacts.
Davide Maestroni, Augusto Sarti, Marco Tagliasacchi, Stefano Tubaro
ICIP3
2003 Architectural Issues and Solutions in the Development of Data-Intensive Web Applications
Stefano Ceri, Piero Fraternali, Aldo Bongio, Stefano Butti, Roberto Acerbis, Marco Tagliasacchi, Giovanni Toffetti Carughi, Carlo Conserva, Roberto Elli, Fulvio Ciapessoni, Claudio Greppi
CIDR6