Marco Grangetto

dblp:77/2058 · DBLP profile ↗
← Back
129ranked-venue papers
22as first author
32since 2021 · last 2026
0000-0002-2709-7864ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 86 · 19 first-author · 14 since 2021Artificial intelligence and machine learning · 21 · 17 since 2021Systems, architecture and hardware · 11 · 1 since 2021Computer networks · 9 · 3 first-author · 2 since 2021Security and privacy · 3Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Defining an Immersive Environments Production Pipeline for Safety Simulation in Virtual Reality
Alessandro Soliman, Paola Gasbarri, Silvia Meschini, Agata Marta Soccini, Lavinia Chiara Tagliabue, Marco Grangetto
CSEDU (1)6
2026 Towards a validation-less approach for small data: training with neural velocity
abstract
Tuning hyperparameters such as learning rate decay and stop conditions is typically done by assessing loss on a held-out validation set. This work introduces neural velocity (NeVe), the rate of change in neuron transfer functions, as a novel indicator of model convergence. We leverage NeVe to create a dynamic training method where learning rates and stop conditions are adjusted based on neural velocity computed from an auxiliary dataset generated by sampling noise. Our experiments across multiple tasks and architectures show that our approach performs on par with traditional validation-based methods, outperforming them in some tasks. Our method does not require withholding data for validation, making it especially advantageous in data-limited scenarios. This work highlights the potential of neural velocity as a key metric for optimizing neural network training.
Gianluca Dalmasso, Andrea Bragagnolo, Enzo Tartaglione, Attilio Fiandrotti, Marco Grangetto
Neurocomputing5
2026 TEP-ones: A simple yet effective approach for transferability estimation of pruned backbones
abstract
In deep learning, the conventional transfer learning paradigm involves fine-tuning a model pre-trained on a complex source task to adapt it to a simpler target task, capitalizing on abundant training data. Concurrently, the paradigm of neural network pruning has emerged as a powerful strategy for enhancing model efficiency, reducing complexity, and optimizing resource utilization. This paper focuses on pruned model transferability estimation for resource-constraint scenarios, where the goal is to rank the performance of pruned pre-trained models on a downstream task without fine-tuning. To this end, from a formal analysis of the intra-class mutual information between samples belonging to the same target class, we observe that, as pruning increases, a sweet phase naturally rises, where the model benefits from better features at the encoder’s output. From this, we derive a Transferability Estimation for Pruned Backbones (TEP-ones) that eases the choice of which pruned model (without the need to train the classifier) is the best candidate for transfer learning.
Gabriele Spadaro, Andrea Bragagnolo, Riccardo Renzulli, Marco Grangetto, Jhony-Heriberto Giraldo-Zuluaga, Attilio Fiandrotti, Enzo Tartaglione
Neurocomputing4
2026 Capsule networks do not need to model everything
abstract
Capsule networks are biologically inspired neural networks that group neurons into vectors called capsules, each explicitly representing an object or one of its parts. The routing mechanism connects capsules in consecutive layers, forming a hierarchical structure between parts and objects, also known as a parse tree. Capsule networks often attempt to model all elements in an image, requiring large network sizes to handle complexities such as intricate backgrounds or irrelevant objects. However, this comprehensive modeling leads to increased parameter counts and computational inefficiencies. Our goal is to enable capsule networks to focus only on the object of interest, reducing the number of parse trees. We accomplish this with REM (Routing Entropy Minimization), a technique that minimizes the entropy of the parse tree-like structure. REM drives the model parameters distribution towards low entropy configurations through a pruning mechanism, significantly reducing the generation of intra-class parse trees. This empowers capsules to learn more stable and succinct representations with fewer parameters and negligible performance loss.
Riccardo Renzulli, Enzo Tartaglione, Marco Grangetto
Pattern Recognit.3
2026 Anatomical foundation models for brain MRIs
abstract
Deep Learning (DL) in neuroimaging has become increasingly relevant for detecting neurological conditions and neurodegenerative disorders. One of the predominant biomarkers in neuroimaging is represented by brain age, which has been shown to be a good indicator for different conditions, such as Alzheimer’s Disease. Using brain age for weakly supervised pre-training of DL models in transfer learning settings has also recently shown promising results, especially when dealing with data scarcity of different conditions. On the other hand, anatomical information of brain MRIs (e.g. cortical thickness) can provide important information for learning good representations that can be transferred to many downstream tasks. In this work, we propose AnatCL, an anatomical foundation model for structural brain MRIs that (i.) leverages anatomical information in a weakly contrastive learning approach, and (ii.) achieves state-of-the-art performances across many different downstream tasks. To validate our approach we consider 12 different downstream tasks for the diagnosis of different conditions such as Alzheimer’s Disease, autism spectrum disorder, and schizophrenia. Furthermore, we also target the prediction of 10 different clinical assessment scores using structural MRI data. Our findings show that incorporating anatomical information during pre-training leads to more robust and generalizable representations. Pre-trained models can be found at: https://github.com/EIDOSLAB/AnatCL . • We propose AnatCL, a novel brain MRI foundation model; • AnatCL accounts for age and anatomical measures; • We benchmark AnatCL on 22 deep-learning-based phenotyping tasks; • Our method achieves SOTA results on many downstream phenotyping tasks; • We publicly release pre-trained models.
Carlo Alberto Barbano, Matteo Brunello, Benoit Dufumier, Marco Grangetto
Pattern Recognit. Lett.4
2026 Robust brain age estimation from structural MRI with contrastive learning
abstract
Estimating brain age from structural MRI has emerged as a powerful tool for characterizing normative and pathological aging. In this work, we explore contrastive learning as a scalable and robust alternative to L1-supervised approaches for brain age estimation. We introduce a novel contrastive loss function, L e x p , and evaluate it across multiple public neuroimaging datasets comprising over 20,000 scans. Our experiments reveal four key findings. First, scaling pre-training on diverse, multi-site data consistently improves generalization performance, cutting external mean absolute error (MAE) nearly in half. Second, L e x p is robust to site-related confounds, maintaining low scanner-predictability as training size increases. Third, contrastive models reliably capture accelerated aging in patients with cognitive impairment and Alzheimer’s disease, as shown through brain age gap analysis, ROC curves, and longitudinal trends. Lastly, unlike L1-supervised baselines, L e x p maintains a strong correlation between brain age accuracy and downstream diagnostic performance, supporting its potential as a foundation model for neuroimaging. These results position contrastive learning as a promising direction for building generalizable and clinically meaningful brain representations.
Carlo Alberto Barbano, Benoit Dufumier, Edouard Duchesnay, Marco Grangetto, Pietro Gori
Pattern Recognit. Lett.4
2026 Unsupervised contrastive analysis for anomaly detection in brain MRIs via conditional diffusion models
abstract
Contrastive Analysis (CA) detects anomalies by contrasting patterns unique to a target group (e.g., unhealthy subjects) from those in a background group (e.g., healthy subjects). In the context of brain MRIs, existing CA approaches rely on supervised contrastive learning or variational autoencoders (VAEs) using both healthy and unhealthy data, but such reliance on target samples is challenging in clinical settings. Unsupervised Anomaly Detection (UAD) learns a reference representation of healthy anatomy, eliminating the need for target samples. Deviations from this reference distribution can indicate potential anomalies. In this context, diffusion models have been increasingly adopted in UAD due to their superior performance in image generation compared to VAEs. Nonetheless, precisely reconstructing the anatomy of the brain remains a challenge. In this work, we bridge CA and UAD by reformulating contrastive analysis principles for the unsupervised setting. We propose an unsupervised framework to improve the reconstruction quality by training a self-supervised contrastive encoder on healthy images to extract meaningful anatomical features. These features are used to condition a diffusion model to reconstruct the healthy appearance of a given image, enabling interpretable anomaly localization via pixel-wise comparison. We validate our approach through a proof-of-concept on a facial image dataset and further demonstrate its effectiveness on four brain MRI datasets, outperforming baseline methods in anomaly localization on the NOVA benchmark. • Unsupervised framework enhancing reconstruction quality in brain MRIs. • Target-invariant contrastive encoder capturing meaningful anatomical features. • Conditional diffusion model to reconstruct the healthy appearance of a given image. • Outperforms competing methods in anomaly localization on the NOVA benchmark.
Cristiano Patrício, Carlo Alberto Barbano, Attilio Fiandrotti, Riccardo Renzulli, Marco Grangetto, Luís F. Teixeira 0001, João C. Neves 0001
Pattern Recognit. Lett.5
2026 CALICE: Continuous Bitrate Control with Adapted LIC Model
abstract
Learned image compression (LIC) has drawn much attention recently as it outperforms standardized codecs in rate-distortion (RD) efficiency. However, an LIC model is typically trained for a specific RD tradeoff, and achieving a different target rate requires retraining the model and storing the weights as a whole, limiting the practical applicability of LIC. In this article, we introduce CALICE, a framework for achieving continuous bitrate control by plugging into a pre-trained LIC model a set of modular adapters. Unlike similar methods that require a distinct set of adapters for each target rate, our method achieves continuous bitrate control by modulating a single set of adapters via a scalar parameter \(\boldsymbol{\alpha}\) , with a total overhead of less than \(\mathbf{0.35}\boldsymbol{\%}\) of the parameters of the LIC model. This design enables efficient support for multiple distortion objectives by learning lightweight, distortion-aware adapters. We also extend our strategy beyond rate control, demonstrating its ability to provide fine-grained adaptation of perceptual quality along the distortion–perception tradeoff. To our knowledge, this is the first method that jointly addresses rate and perceptual control using a unified, low-cost strategy. We publicly released the code at https://github.com/EIDOSLAB/CALICE .
Gabriele Spadaro, Alberto Presta, Jhony-Heriberto Giraldo-Zuluaga, Attilio Fiandrotti, Marco Grangetto, Enzo Tartaglione
ACM Trans. Multim. Comput. Commun. Appl.5
2025 Towards RISC-V-based HPC: The Italian Pathfinding Activities in the DARE-SGA1 Project
abstract
The European Union’s efforts towards technological sovereignty in High-Performance Computing are driving research and development of RISC-V-based supercomputers. The DARE SGA1 project, in particular, aims to develop chips designed and owned by Europeans. This paper introduces the Italian contribution to DARE SGA1 regarding pathfinding activities toward future RISC-V-based accelerator designs, reliability improvements, system software, and AI and Quantum Chemistry applications.
Giovanni Agosta, Marco Aldinucci, Andrea Bartolini, Laura Bellentani, Andrea Biagioni, Daniele Cesarini, Carlotta Chiarini, Iacopo Colonnelli, Pietro Delugas, Lev Denisov, Ottorino Frezza, Marco Grangetto, Francesca Lo Cicero, Alessandro Lonardo, Michele Martinelli, Andrea Maslov, Mauro Olivieri, Pierpaolo Perticaroli, Luca Pontisso, Cristian Rossi, Davide Rossi 0001, Sergio Saponara, Antonio Sciarappa, Francesco Simula, Matteo Sonza Reorda, Massimo Torquati, Piero Vicini
DSD12
2025 Denoising Diffusion Probabilistic Model for Point Cloud Compression at Low Bit-Rates
abstract
Efficient compression of low-bit-rate point clouds is critical for bandwidth-constrained applications. However, existing techniques mainly focus on high-fidelity reconstruction, requiring many bits for compression. This paper proposes a "Denoising Diffusion Probabilistic Model" (DDPM) architecture for point cloud compression (DDPM-PCC) at low bit-rates. A PointNet encoder produces the condition vector for the generation, which is then quantized via a learnable vector quantizer. This configuration allows to achieve a low bitrates while preserving quality. Experiments on ShapeNet and ModelNet40 show improved rate-distortion at low rates compared to standardized and state-of-the-art approaches. We publicly released the code at https://github.com/EIDOSLAB/DDPM-PCC.
Gabriele Spadaro, Alberto Presta, Jhony-Heriberto Giraldo-Zuluaga, Marco Grangetto, Giuseppe Valenzise, Attilio Fiandrotti, Enzo Tartaglione
ICME4
2025 Neural Velocity for hyperparameter tuning
abstract
Hyperparameter tuning, such as learning rate decay and defining a stopping criterion, often relies on monitoring the validation loss. This paper presents NeVe, a dynamic training approach that adjusts the learning rate and defines the stop criterion based on the novel notion of "neural velocity". The neural velocity measures the rate of change of each neuron’s transfer function and is an indicator of model convergence: sampling neural velocity can be performed even by forwarding noise in the network, reducing the need for a held-out dataset. Our findings show the potential of neural velocity as a key metric for optimizing neural network training efficiently.
Gianluca Dalmasso, Andrea Bragagnolo, Enzo Tartaglione, Attilio Fiandrotti, Marco Grangetto
IJCNN5
2025 Efficient Progressive Image Compression with Variance-Aware Masking
abstract
Learned progressive image compression is gaining momentum as it allows improved image reconstruction as more bits are decoded at the receiver. We propose a progressive image compression method in which an image is first represented as a pair of base-quality and top-quality latent representations. Next, a residual latent representation is encoded as the element-wise difference between the top and base representations. Our scheme enables progressive image compression with element-wise granularity by introducing a masking system that ranks each element of the residual latent representation from most to least important, dividing it into complementary components, which can be transmitted separately to the decoder in order to obtain different reconstruction quality. The masking system does not add further parameters or complexity. At the receiver, any elements of the top latent representation excluded from the transmitted components can be independently replaced with the mean predicted by the hyperprior architecture, ensuring reliable reconstructions at any intermediate quality level. We also in-troduced Rate Enhancement Modules (REMs), which refine the estimation of entropy parameters using already decoded components. We obtain results competitive with state-of-the-art competitors, while significantly reducing computational complexity, decoding time, and number of parameters.
Alberto Presta, Enzo Tartaglione, Attilio Fiandrotti, Marco Grangetto, Pamela C. Cosman
WACV4
2025 WiGNet: Windowed Vision Graph Neural Network
abstract
In recent years, Graph Neural Networks (GNNs) have demonstrated strong adaptability to various real-world challenges, with architectures such as Vision GNN (ViG) achieving state-of-the-art performance in several computer vision tasks. However, their practical applicability is hindered by the computational complexity of constructing the graph, which scales quadratically with the image size. In this paper, we introduce a novel Windowed vision Graph neural Network (WiGNet) model for efficient image processing. WiGNet explores a different strategy from previous works by partitioning the image into windows and constructing a graph within each window. Therefore, our model uses graph convolutions instead of the typical 2D convolution or self-attention mechanism. WiGNet effectively manages computational and memory complexity for large image sizes. We evaluate our method in the ImageNet-1k benchmark dataset and test the adaptability of WiGNet using the CelebA-HQ dataset as a downstream task with higher-resolution images. In both of these scenarios, our method achieves competitive results compared to previous vision GNNs while keeping memory and computational complexity at bay. WiGNet offers a promising solution toward the deployment of vision GNNs in real-world applications. We publicly released the code and pre-trained models at https://github.com/EIDOSLAB/WiGNet.
Gabriele Spadaro, Marco Grangetto, Attilio Fiandrotti, Enzo Tartaglione, Jhony-Heriberto Giraldo-Zuluaga
WACV2
2025 STanH: Parametric Quantization for Variable Rate Learned Image Compression
abstract
In end-to-end learned image compression, encoder and decoder are jointly trained to minimize a R + λD cost function, where λ controls the trade-off between rate of the quantized latent representation and image quality. Unfortunately, a distinct encoder-decoder pair with millions of parameters must be trained for each λ, hence the need to switch encoders and to store multiple encoders and decoders on the user device for every target rate. This paper proposes to exploit a differentiable quantizer designed around a parametric sum of hyperbolic tangents, called STanH, that relaxes the step-wise quantization function. STanH is implemented as a differentiable activation layer with learnable quantization parameters that can be plugged into a pre-trained fixed rate model and refined to achieve different target bitrates. Experimental results show that our method enables variable rate coding with comparable efficiency to the state-of-the-art, yet with significant savings in terms of ease of deployment, training time, and storage costs.
Alberto Presta, Enzo Tartaglione, Attilio Fiandrotti, Marco Grangetto
IEEE Trans. Image Process.4
2024 Domain Adaptation for Learned Image Compression with Supervised Adapters
abstract
In Learned Image Compression (LIC), a model is trained at encoding and decoding images sampled from a source domain, often outperforming traditional codecs on natural images; yet its performance may be far from optimal on images sampled from different domains. In this work, we tackle the problem of adapting a pre-trained model to multiple target domains by plugging into the decoder an adapter module for each of them, including the source one. Each adapter improves the decoder performance on a specific domain, without the model forgetting about the images seen at training time. A gate network computes the weights to optimally blend the contributions from the adapters when the bitstream is decoded. We experimentally validate our method over two state-of-the-art pre-trained models, observing improved rate-distortion efficiency on the target domains without penalties on the source domain. Furthermore, the gate’s ability to find similarities with the learned target domains enables better encoding efficiency also for images outside them.
Alberto Presta, Gabriele Spadaro, Enzo Tartaglione, Attilio Fiandrotti, Marco Grangetto
DCC5
2024 Boost Your NeRF: A Model-Agnostic Mixture of Experts Framework for High Quality and Efficient Rendering
Francesco Di Sario, Riccardo Renzulli, Enzo Tartaglione, Marco Grangetto
ECCV (83)4
2024 Gabic: Graph-Based Attention Block for Image Compression
abstract
While standardized codecs like JPEG and HEVC-intra represent the industry standard in image compression, neural Learned Image Compression (LIC) codecs represent a promising alternative. In detail, integrating attention mechanisms from Vision Transformers into LIC models has shown improved compression efficiency. However, extra efficiency often comes at the cost of aggregating redundant features. This work proposes a Graph-based Attention Block for Image Compression (GABIC), a method to reduce feature redundancy based on a k-Nearest Neighbors enhanced attention mechanism. Our experiments show that GABIC outperforms comparable methods, particularly at high bit rates, enhancing compression performance.
Gabriele Spadaro, Alberto Presta, Enzo Tartaglione, Jhony-Heriberto Giraldo-Zuluaga, Marco Grangetto, Attilio Fiandrotti
ICIP5
2024 ALICE: Adapt your Learnable Image Compression modEl for variable bitrates
abstract
When training a Learned Image Compression model, the loss function is minimized such that the encoder and the decoder attain a target Rate-Distorsion trade-off. Therefore, a distinct model shall be trained and stored at the transmitter and receiver for each target rate, fostering the quest for efficient variable bitrate compression schemes. This paper proposes plugging Low-Rank Adapters into a transformer-based pre-trained LIC model and training them to meet different target rates. With our method, encoding an image at a variable rate is as simple as training the corresponding adapters and plugging them into the frozen pre-trained model. Our experiments show performance comparable with state-of-the-art fixed-rate LIC models at a fraction of the training and deployment cost. We publicly released the code at https://github.com/EIDOSLAB/ALICE.
Gabriele Spadaro, Muhammad Salman Ali, Alberto Presta, Giommaria Pilo, Sung-Ho Bae, Jhony-Heriberto Giraldo-Zuluaga, Attilio Fiandrotti, Marco Grangetto, Enzo Tartaglione
VCIP8
2024 High capacity reversible data hiding in radiographic images with optimal bit allocation
abstract
Abstract This paper extends and improves the performance of a digital reversible watermarking algorithm based on histogram shifting presented in previous works. The considered algorithm exploits the property of image histograms of some kinds of medical images which present many contiguous 0-runs, i.e., a comb structure in the gray level frequencies. In particular, radiographic images exhibit this structure after contrast enhancement during the acquisition process. The previous work suggested performing gray-level histogram shifting according to a local optimization technique. In this paper, we apply combinatorial optimization techniques to entire blocks of contiguous 0-runs using a non-linear objective function transformed to fit a linear optimization algorithm. The obtained results show a meaningful improvement in the payload capacity of the original data-hiding method. A mild Peak Signal-to-Noise Ratio (PSNR) reduction is still acceptable for a qualitative preview of the images, which can be completely restored to their original cover form thanks to the reversibility of the method.
Davide Cavagnino, Alessandro Druetto, Marco Grangetto, Maurizio Lucenteforte
Multim. Tools Appl.3
2024 A lightweight deep learning architecture for malaria parasite-type classification and life cycle stage detection
abstract
Abstract Malaria is an endemic in various tropical countries. The gold standard for disease detection is to examine the blood smears of patients by an expert medical professional to detect malaria parasite called Plasmodium. In the rural areas of underdeveloped countries, with limited infrastructure, a scarcity of healthcare professionals, an absence of sufficient computing devices, and a lack of widespread internet access, this task becomes more challenging. A severe case of malaria can be fatal within one week, so the correct detection of the malaria parasite and its life cycle stage is crucial in treating the disease correctly. Though computer vision-based malaria detection has been adequately explored lately, the malaria life cycle stage classification is still a relatively unexplored field. In this paper, we introduce a fast and robust deep learning methodology to not only classify the malaria parasite-type detection but also the life cycle stage identification of the infected cell. The proposed deep learning architecture is more than twenty times lighter than the widely used DenseNet and has less than 0.4 million parameters, making it a good candidate to be used in the mobile applications of such economically challenged states for malaria detection. We have used four different publicly available malaria datasets to test the proposed architecture and gained significantly better results than the current state of the art on malaria parasite-type and malaria life cycle classification.
Hafiza Ayesha Hoor Chaudhry, Muhammad Shahid Farid, Attilio Fiandrotti, Marco Grangetto
Neural Comput. Appl.4
2023 Unbiased Supervised Contrastive Learning
Carlo Alberto Barbano, Benoit Dufumier, Enzo Tartaglione, Marco Grangetto, Pietro Gori
ICLR4
2023 Disentangling private classes through regularization
abstract
Deep learning models are nowadays broadly deployed to solve an incredibly large variety of tasks. However, little attention has been devoted to connected legal aspects. In 2016, the European Union approved the General Data Protection Regulation which entered into force in 2018. Its main rationale was to protect the privacy and data protection of its citizens by the way of operating the so-called “Data Economy”. As data is the fuel of modern Artificial Intelligence, it is argued that the GDPR can be partly applicable to a series of algorithmic decision-making tasks before a more structured AI Regulation enters into force. In the meantime, AI should not allow undesired information leakage deviating from the purpose for which is created. In this work, we propose DisP, an approach for deep learning models disentangling the information related to some classes we desire to keep private, from the data processed by AI. In particular, DisP is a regularization strategy de-correlating the features belonging to the same private class at training time, hiding the information about private class membership. Our experiments on state-of-the-art deep learning models show the effectiveness of DisP, minimizing the risk of extraction for the classes we desire to keep private.
Enzo Tartaglione, Francesca Gennari, Victor Quétu, Marco Grangetto
Neurocomputing4
2022 Towards Efficient Capsule Networks
abstract
From the moment Neural Networks dominated the scene for image processing, the computational complexity needed to solve the targeted tasks skyrocketed: against such an unsustainable trend, many strategies have been developed, ambitiously targeting performance’s preservation. Promoting sparse topologies, for example, allows the deployment of deep neural networks models on embedded, resource-constrained devices. Recently, Capsule Networks were introduced to enhance explainability of a model, where each capsule is an explicit representation of an object or its parts. These models show promising results on toy datasets, but their low scalability prevents deployment on more complex tasks. In this work, we explore sparsity besides capsule representations to improve their computational efficiency by reducing the number of capsules. We show how pruning with Capsule Network achieves high generalization with less memory requirements, computational effort, and inference and training time.
Riccardo Renzulli, Marco Grangetto
ICIP2
2022 To update or not to update? Neurons at equilibrium in deep models
abstract
Recent advances in deep learning optimization showed that, with some a-posteriori information on fully-trained models, it is possible to match the same performance by simply training a subset of their parameters. Such a discovery has a broad impact from theory to applications, driving the research towards methods to identify the minimum subset of parameters to train without look-ahead information exploitation. However, the methods proposed do not match the state-of-the-art performance, and rely on unstructured sparsely connected models.In this work we shift our focus from the single parameters to the behavior of the whole neuron, exploiting the concept of neuronal equilibrium (NEq). When a neuron is in a configuration at equilibrium (meaning that it has learned a specific input-output relationship), we can halt its update; on the contrary, when a neuron is at non-equilibrium, we let its state evolve towards an equilibrium state, updating its parameters. The proposed approach has been tested on different state-of-the-art learning strategies and tasks, validating NEq and observing that the neuronal equilibrium depends on the specific learning setup.
Andrea Bragagnolo, Enzo Tartaglione, Marco Grangetto
NeurIPS3
2022 LOss-Based SensiTivity rEgulaRization: Towards deep sparse neural networks
Enzo Tartaglione, Andrea Bragagnolo, Attilio Fiandrotti, Marco Grangetto
Neural Networks4
2022 SeReNe: Sensitivity-Based Regularization of Neurons for Structured Sparsity in Neural Networks
abstract
Deep neural networks include millions of learnable parameters, making their deployment over resource-constrained devices problematic. Sensitivity-based regularization of neurons (SeReNe) is a method for learning sparse topologies with a structure, exploiting neural sensitivity as a regularizer. We define the sensitivity of a neuron as the variation of the network output with respect to the variation of the activity of the neuron. The lower the sensitivity of a neuron, the less the network output is perturbed if the neuron output changes. By including the neuron sensitivity in the cost function as a regularization term, we are able to prune neurons with low sensitivity. As entire neurons are pruned rather than single parameters, practical network footprint reduction becomes possible. Our experimental results on multiple network architectures and datasets yield competitive compression ratios with respect to state-of-the-art references.
Enzo Tartaglione, Andrea Bragagnolo, Francesco Odierna, Attilio Fiandrotti, Marco Grangetto
IEEE Trans. Neural Networks Learn. Syst.5
2021 EnD: Entangling and Disentangling Deep Representations for Bias Correction
abstract
Artificial neural networks perform state-of-the-art in an ever-growing number of tasks, and nowadays they are used to solve an incredibly large variety of tasks. There are problems, like the presence of biases in the training data, which question the generalization capability of these models. In this work we propose EnD, a regularization strategy whose aim is to prevent deep models from learning unwanted biases. In particular, we insert an "information bottleneck" at a certain point of the deep neural network, where we disentangle the information about the bias, still letting the useful information for the training task forward-propagating in the rest of the model. One big advantage of EnD is that we do not require additional training complexity (like decoders or extra layers in the model), since it is a regularizer directly applied on the trained model. Our experiments show that EnD effectively improves the generalization on unbiased test sets, and it can be effectively applied on real-case scenarios, like removing hidden biases in the COVID-19 detection from radiographic images.
Enzo Tartaglione, Carlo Alberto Barbano, Marco Grangetto
CVPR3
2021 Capsule Networks with Routing Annealing
Riccardo Renzulli, Enzo Tartaglione, Attilio Fiandrotti, Marco Grangetto
ICANN (1)4
2021 Unitopatho, A Labeled Histopathological Dataset for Colorectal Polyps Classification and Adenoma Dysplasia Grading
abstract
Histopathological characterization of colorectal polyps allows to tailor patients’ management and follow up with the ultimate aim of avoiding or promptly detecting an invasive carcinoma. Colorectal polyps characterization relies on the histological analysis of tissue samples to determine the polyps malignancy and dysplasia grade. Deep neural networks achieve outstanding accuracy in medical patterns recognition, however they require large sets of annotated training images. We introduce UniToPatho, an annotated dataset of 9536 hematoxylin and eosin (H&E) stained patches extracted from 292 whole-slide images, meant for training deep neural networks for colorectal polyps classification and adenomas grading. We present our dataset and provide insights on how to tackle the problem of automatic colorectal polyps characterization by suggesting a multi-resolution deep learning approach.
Carlo Alberto Barbano, Daniele Perlo, Enzo Tartaglione, Attilio Fiandrotti, Luca Bertero, Paola Cassoni, Marco Grangetto
ICIP7
2021 On the Role of Structured Pruning for Neural Network Compression
abstract
This works explores the benefits of structured parameter pruning in the framework of the MPEG standardization efforts for neural network compression. First less relevant parameters are pruned from the network, then remaining parameters are quantized and finally quantized parameters are entropy coded. We consider an unstructured pruning strategy that maximizes the number of pruned parameters at the price of randomly sparse tensors and a structured strategy that prunes fewer parameters yet yields regularly sparse tensors. We show that structured pruning enables better end-to-end compression despite lower pruning ratio because it boosts the efficiency of the arithmetic coder. As a bonus, once decompressed, the network memory footprint is lower as well as its inference time.
Andrea Bragagnolo, Enzo Tartaglione, Attilio Fiandrotti, Marco Grangetto
ICIP4
2021 HEMP: High-order entropy minimization for neural network compression
Enzo Tartaglione, Stéphane Lathuilière, Attilio Fiandrotti, Marco Cagnazzo, Marco Grangetto
Neurocomputing5
2021 On the robustness of three classes of rateless codes against pollution attacks in P2P networks
abstract
Abstract Rateless codes (a.k.a. fountain codes, digital fountain) have found their way in numerous peer-to-peer based applications although their robustness to the so called pollution attack has not been deeply investigated because they have been originally devised as a solution for dealing with block erasures and not for block modification. In this paper we provide an analysis of the intrinsic robustness of three rateless codes algorithms, i.e., random linear network codes (RLNC), Luby transform (LT), and band codes (BC) against intentional data modification. By intrinsic robustness we mean the ability of detecting as soon as possible that modification of at least one equation has occurred as well as the possibility a receiver can decode from the set of equations with and without the modified ones. We focus on bare rateless codes where no additional information is added to equations (e.g., tags) or higher level protocol are used (e.g., verification keys to pre-distribute to receivers) to detect and recover from data modification. We consider several scenarios that combine both random and targeted selection of equations to alter and modification of an equation that can either change the rank of the coding matrix or not. Our analysis reveals that a high percentage of attacks goes undetected unless a minimum code redundancy is achieved, LT codes are the most fragile in virtually all scenarios, RLNC and BC are quite insensitive to the victim selection and type of alteration of chosen equations and exhibit virtually identical robustness although BC offer a low complexity of the decoding algorithm.
Rossano Gaeta, Marco Grangetto
Peer-to-Peer Netw. Appl.2
2020 Pruning Artificial Neural Networks: A Way to Find Well-Generalizing, High-Entropy Sharp Minima
abstract
Recently, a race towards the simplification of deep networks has begun, showing that it is effectively possible to reduce the size of these models with minimal or no performance loss. However, there is a general lack in understanding why these pruning strategies are effective. In this work, we are going to compare and analyze pruned solutions with two different pruning approaches, one-shot and gradual, showing the higher effectiveness of the latter. In particular, we find that gradual pruning allows access to narrow, well-generalizing minima, which are typically ignored when using one-shot approaches. In this work we also propose PSP-entropy, a measure to understand how a given neuron correlates to some specific learned classes. Interestingly, we observe that the features extracted by iteratively-pruned models are less correlated to specific classes, potentially making these models a better fit in transfer learning approaches.
Enzo Tartaglione, Andrea Bragagnolo, Marco Grangetto
ICANN (2)3
2020 Delving in the loss landscape to embed robust watermarks into neural networks
abstract
In the last decade the use of artificial neural networks (ANNs) in many fields like image processing or speech recognition has become a common practice because of their effectiveness to solve complex tasks. However, in such a rush, very little attention has been paid to security aspects. In this work we explore the possibility to embed a watermark into the ANN parameters. We exploit model redundancy and adaptation capacity to lock a subset of its parameters to carry the watermark sequence. The watermark can be extracted in a simple way to claim copyright on models but can be very easily attacked with model fine-tuning. To tackle this culprit we devise a novel watermark aware training strategy. We aim at delving into the loss landscape to find an optimal configuration of the parameters such that we are robust to fine-tuning attacks towards the watermarked parameters. Our experimental results on classical ANN models trained on well-known MNIST and CIFAR-10 datasets show that the proposed approach makes the embedded watermark robust to fine-tuning and compression attacks.
Enzo Tartaglione, Marco Grangetto, Davide Cavagnino, Marco Botta
ICPR2
2020 A non-discriminatory approach to ethical deep learning
abstract
Artificial neural networks perform state-of-the-art in an ever-growing number of tasks, nowadays they are used to solve an incredibly large variety of tasks. However, typical training strategies do not take into account lawful, ethical and discriminatory potential issues the trained ANN models could incur in. In this work we propose NDR, a non-discriminatory regularization strategy to prevent the ANN model to solve the target task using some discriminatory features like, for example, the ethnicity in an image classification task for human faces. In particular, a part of the ANN model is trained to hide the discriminatory information such that the rest of the network focuses in learning the given learning task. Our experiments show that NDR can be exploited to achieve non-discriminatory models with both minimal computational overhead and performance loss.
Enzo Tartaglione, Marco Grangetto
TrustCom2
2020 Graph Laplacian for image anomaly detection
abstract
Abstract Reed–Xiaoli detector (RXD) is recognized as the benchmark algorithm for image anomaly detection; however, it presents known limitations, namely the dependence over the image following a multivariate Gaussian model, the estimation and inversion of a high-dimensional covariance matrix, and the inability to effectively include spatial awareness in its evaluation. In this work, a novel graph-based solution to the image anomaly detection problem is proposed; leveraging the graph Fourier transform, we are able to overcome some of RXD’s limitations while reducing computational cost at the same time. Tests over both hyperspectral and medical images, using both synthetic and real anomalies, prove the proposed technique is able to obtain significant gains over performance by other algorithms in the state of the art.
Francesco Verdoja, Marco Grangetto
Mach. Vis. Appl.2
2019 Post-synaptic Potential Regularization Has Potential
Enzo Tartaglione, Daniele Perlo, Marco Grangetto
ICANN (2)3
2019 Accelerating Spectral Graph Analysis Through Wavefronts of Linear Algebra Operations
abstract
The wavefront pattern captures the unfolding of a parallel computation in which data elements are laid out as a logical multidimensional grid and the dependency graph favours a diagonal sweep across the grid. In the emerging area of spectral graph analysis, the computing often consists in a wavefront running over a tiled matrix, involving expensive linear algebra kernels. While these applications might benefit from parallel heterogeneous platforms (multi-core with GPUs), programming wavefront applications directly with high-performance linear algebra libraries yields code that is complex to write and optimize for the specific application. We advocate a methodology based on two abstractions (linear algebra and parallel pattern-based run-time), that allows to develop portable, self-configuring, and easy-to-profile code on hybrid platforms.
Maurizio Drocco, Paolo Viviani 0001, Iacopo Colonnelli, Marco Aldinucci, Marco Grangetto
PDP5
2019 Virtual Hand Illusion: The Alien Finger Motion Experiment
abstract
In Virtual Reality, the need to understand how subjects perceive their representation is gaining attention and importance. We present a contribution to a better understanding of the sense of embodiment by assessing two of its main components, body ownership and agency, through an experiment involving alien motion. The key aspect of the experimental protocol is to integrate a condition with some personalized alien finger movement while the subject is asked to remain still. Body ownership appears to be significantly reduced, but not agency. We also propose a metric to assess quantitatively that the view of the alien movement induces more finger posture variation compared to the reference context in the still condition.
Agata Marta Soccini, Marco Grangetto, Tetsunari Inamura, Sotaro Shimada
VR2
2019 Robust gait identification using Kinect dynamic skeleton data
Elena Gianaria, Marco Grangetto
Multim. Tools Appl.2
2019 Securing Network Coding Architectures Against Pollution Attacks With Band Codes
abstract
During a pollution attack, malicious nodes purposely transmit bogus data to the honest nodes to cripple the communication. Securing the communication requires identifying and isolating the malicious nodes. However, in network coding (NC) architectures, random recombinations at the nodes increase the probability that honest nodes relay polluted packets. Thus, discriminating between honest and malicious nodes to isolate the latter turns out to be challenging at best. Band codes (BCs) are a family of rateless codes whose coding window size can be adjusted to reduce the probability that honest nodes relay polluted packets. We leverage such a property to design a distributed scheme for identifying the malicious nodes in the network. Each node counts the number of times that each neighbor has been involved in cases of polluted data reception and exchanges such counts with its neighbor nodes. Then, each node computes for each neighbor a discriminative honest score estimating the probability that the neighbor relays clean packets. We model such probability as a function of the BC coding window size, showing its impact on the accuracy and effectiveness of our distributed blacklisting scheme. We experiment distributing a live video feed in a P2P NC system, verifying the accuracy of our model and showing that our scheme allows us to secure the network against pollution attacks recovering near pre-attack video quality.
Attilio Fiandrotti, Rossano Gaeta, Marco Grangetto
IEEE Trans. Inf. Forensics Secur.3
2018 Detection and Tracking of Astral Microtubules in Fluorescence Microscopy Images
abstract
In this paper we explore detection and tracking of astral micro-tubules, a sub-population of microtubules which only exists during and immediately before mitosis and aids in the spindle orientation by connecting it to the cell cortex. Its analysis can be useful to determine the presence of certain diseases, such as brain pathologies and cancer. The proposed algorithm focuses on overcoming the problems regarding fluorescence microscopy images and microtubule behaviour by using various image processing techniques and is then compared with three existing algorithms, tested on consistent sets of images.
Joshua Levine, Marco Grangetto, Marilena Varrecchia, Gabriella Olmo
ICIP2
2018 DOST: a distributed object segmentation tool
Muhammad Shahid Farid, Maurizio Lucenteforte, Marco Grangetto
Multim. Tools Appl.3
2018 Convolutional Neural Network for Intermediate View Enhancement in Multiview Streaming
abstract
Multiview video streaming continues to gain popularity due to the great viewing experience it offers, as well as its availability that has been enabled by increased network throughput and other recent technical developments. User demand for interactive multiview video streaming that provides seamless view switching upon request is also increasing. However, it is a highly challenging task to stream stable and high quality videos that allow real-time scene navigation within the bandwidth constraint. In this paper, a convolutional neural network (ConvNet)-assisted seamless multiview video streaming system is proposed to tackle the challenge. The proposed method solves the problem from two perspectives. First, a ConvNet-assisted multiview representation method is proposed, which provides flexible interactivity without compromising on multiview video compression efficiency. Second, a bit allocation mechanism guided by a navigation model is developed to provide seamless navigation and adapt to network bandwidth fluctuations at the same time. These two blocks work closely to provide an optimized viewing experience to users. They can be integrated into any existing multiview video streaming framework to enhance overall performance. Experimental results demonstrate the effectiveness of the proposed method for seamless multiview streaming.
Li Yu 0004, Tammam Tillo, Jimin Xiao, Marco Grangetto
IEEE Trans. Multim.4
2017 Efficient representation of segmentation contours using chain codes
abstract
Segmentation is one of the most important low-level tasks in image processing as it enables many higher level computer vision tasks like object recognition and tracking. Segmentation can also be exploited for image compression using recent graph-based algorithms, provided that the corresponding contours can be represented efficiently. Transmission of borders is also key to distributed computer vision. In this paper we propose a new chain code tailored to compress segmentation contours. Based on the widely known 3OT, our algorithm is able to encode regions avoiding borders it has already coded once and without the need of any starting point information for each region. We tested our method against three other state of the art chain codes over the BSDS500 dataset, and we demonstrated that the proposed chain code achieves the highest compression ratio, resulting on average in over 27% bit-per-pixel saving.
Francesco Verdoja, Marco Grangetto
ICASSP2
2017 Directional graph weight prediction for image compression
abstract
Graph-based models have recently attracted attention for their potential to enhance transform coding image compression thanks to their capability to efficiently represent discontinuities. Graph transform gets closer to the optimal KLT by using weights that represent inter-pixel correlations but the extra cost to provide such weights can overwhelm the gain, especially in the case of natural images rich of details. In this paper we provide a novel idea to make graph transform adaptive to the actual image content, avoiding the need to encode the graph weights as side information. We show that an approach similar to spatial prediction can be used to effectively predict graph weights in place of pixels; in particular, we propose the design of directional graph weight prediction modes and show the resulting coding gain. The proposed approach can be used jointly with other graph based intra prediction methods to further enhance compression. Our comparative experimental analysis, carried out with a fully fledged still image coding prototype, shows that we are able to achieve significant coding gains.
Francesco Verdoja, Marco Grangetto
ICASSP2
2017 Perceptual quality assessment of 3D synthesized images
abstract
Multiview video plus depth (MVD) is the most popular 3D video format where the texture images contain the color information and the depth maps represent the geometry of the scene. The depth maps are exploited to obtain intermediate views to enable 3D-TV and free-viewpoint applications using the depth image based rendering (DIBR) techniques. DIBR is used to get an estimate of the intermediate views but has to cope with depth errors, occlusions, imprecise camera parameters, re-interpolation, to mention a few issues. Therefore, being able to evaluate the true perceptual quality of synthesized images is of paramount importance for a high quality 3D experience. In this paper, we present a novel algorithm to assess the quality of the synthesized images in the absence of the corresponding references. The algorithm uses the original views from which the virtual image is generated to estimate the distortion induced by the DIBR process. In particular, a block-based perceptual feature matching based on signal phase congruency metric is devised to estimate the synthesis distortion. The experiments worked out on standard DIBR synthesized database show that the proposed algorithm achieves high correlation with the subjective ratings and outperforms the existing 3D quality assessment algorithms.
Muhammad Shahid Farid, Maurizio Lucenteforte, Marco Grangetto
ICME3
2017 Securing Coding-Based Cloud Storage Against Pollution Attacks
abstract
The widespread diffusion of distributed and cloud storage solutions has changed dramatically the way users, system designers, and service providers manage their data. Outsourcing data on remote storage provides indeed many advantages in terms of both capital and operational costs. The security of data outsourced to the cloud, however, still represents one of the major concerns for all stakeholders. Pollution attacks, whereby a set of malicious entities attempt to corrupt stored data, are one of the many risks that affect cloud data security. In this paper we deal with pollution attacks in coding-based block-level cloud storage systems, i.e., systems that use linear codes to fragment, encode, and disperse virtual disk sectors across a set of storage nodes to achieve desired levels of redundancy, and to improve reliability and availability without sacrificing performance. Unfortunately, the effects of a pollution attack on linear coding can be disastrous, since a single polluted fragment can propagate pervasively in the decoding phase, thus hampering the whole sector. In this work we show that, using rateless codes, we can design an early pollution detection algorithm able to spot the presence of an attack while fetching the data from cloud storage during the normal disk reading operations. The alarm triggers a procedure that locates the polluting nodes using the proposed detection mechanism along with statistical inference. The performance of the proposed solution is analyzed under several aspects using both analytical modelling and accurate simulation using real disk traces. Our results show that the proposed approach is very robust and is able to effectively isolate the polluters, even in harsh conditions, provided that enough data redundancy is used.
Cosimo Anglano, Rossano Gaeta, Marco Grangetto
IEEE Trans. Parallel Distributed Syst.3
2016 Kinect-based gait analysis for automatic frailty syndrome assessment
abstract
Smart living and well aging represent key challenges for our society. The precursor state of adverse outcomes that characterize aging has been recognized from scientific community with the frailty syndrome, determined by the loss of physical and psychological capacities. In this paper we define gait and posture indexes that can be effectively and unobtrusively measured using computer vision and RGBD sensors, e.g. the popular MS Kinect. In this study we present preliminary results showing evidence that the proposed approach can pave the way to the design of an automatic and objective tool for detection and early prevention of frailty.
Elena Gianaria, Marco Grangetto, Mattia Roppolo, Anna Mulasso, Emanuela Rabaglietti
ICIP2
2016 Global and local anomaly detectors for tumor segmentation in dynamic pet acquisitions
abstract
In this paper we explore the application of anomaly detection techniques to tumor voxels segmentation. The developed algorithms work on 3-points dynamic FDG-PET acquisitions and leverage on the peculiar anaerobic metabolism that cancer cells experience over time. A few different global or local anomaly detectors are discussed, together with an investigation over two different algorithms aiming to estimate normal tissues' statistical distribution. Finally, all the proposed algorithms are tested on a dataset composed of 9 patients proving that anomaly detectors are able to outperform techniques in the state of the art.
Francesco Verdoja, Barbara Bonafè, Davide Cavagnino, Marco Grangetto, Christian Bracco, Teresio Varetto, Manuela Racca, Michele Stasi
ICIP4
2016 Characterization of Band Codes for Pollution-Resilient Peer-to-Peer Video Streaming
abstract
We provide a comprehensive characterization of band codes (BC) as a resilient-by-design solution to pollution attacks in network coding (NC)-based peer-to-peer live video streaming. Consider one malicious node injecting bogus coded packets into the network: the recombinations at the nodes generate an avalanche of novel coded bogus packets. Therefore, the malicious node can cripple the communication by injecting into the network only a handful of polluted packets. Pollution attacks are typically addressed by identifying and isolating the malicious nodes from the network. Pollution detection is, however, not straightforward in NC as the nodes exchange coded packets. Similarly, malicious nodes identification is complicated by the ambiguity between malicious nodes and nodes that have involuntarily relayed polluted packets. This paper addresses pollution attacks through a radically different approach which relies on BCs. BCs are a family of rateless codes originally designed for controlling the NC decoding complexity in mobile applications. Here, we exploit BCs for the totally different purpose of recombining the packets at the nodes so to avoid that the pollution propagates by adaptively adjusting the coding parameters. Our streaming experiments show that BCs curb the propagation of the pollution and restore the quality of the distributed video stream.
Attilio Fiandrotti, Rossano Gaeta, Marco Grangetto
IEEE Trans. Multim.3
2015 Device-to-Device Content Distribution in Cellular Networks: A User-Centric Collaborative Strategy
abstract
In this paper device-to-device (D2D) communication is proposed as a tool for enhancing the services provided by mobile cellular networks. The technique we discuss relies on the cooperation among the mobile users that participate in the service delivery process, under the control of the base station. In particular, the paper proposes an incentive mechanism encouraging terminals to organize into an optimal number of clusters from the point of view of both bandwidth capacity and power saving. The aspects inducing users to collaborate are analyzed and modelled, e.g., the amount of mobile battery power drained during collaboration, to better design incentives to switch to D2D. The proposed technique allows the base station to estimate the incentives to grant to the mobile terminals to optimize its cost. Through both analysis and simulations, we show that our scheme achieves a significant gain in terms of costs while increasing the bandwidth capacity of the whole cell.
Paolo Castagno, Rossano Gaeta, Marco Grangetto, Matteo Sereno
GLOBECOM3
2015 Objective quality metric for 3D virtual views
abstract
In free-viewpoint television (FTV) framework, due to hardware and bandwidth constraints, only a limited number of viewpoints are generally captured, coded and transmitted; therefore, a large number of views needs to be synthesized at the receiver to grant a really immersive 3D experience. It is thus evident that the estimation of the quality of the synthesized views is of paramount importance. Moreover, quality assessment of the synthesized view is very challenging since the corresponding original views are generally not available either on the encoder (not captured) or the decoder side (not transmitted). To tackle the mentioned issues, this paper presents an algorithm to estimate the quality of the synthesized images in the absence of the corresponding reference images. The algorithm is based upon the cyclopean eye theory. The statistical characteristics of an estimated cyclopean image are compared with the synthesized image to measure its quality. The prediction accuracy and reliability of the proposed technique are tested on standard video dataset compressed with HEVC showing excellent correlation results with respect to state-of-the-art full reference image and video quality metrics.
Muhammad Shahid Farid, Maurizio Lucenteforte, Marco Grangetto
ICIP3
2015 Superpixel-driven graph transform for image compression
abstract
Block-based compression tends to be inefficient when blocks contain arbitrary shaped discontinuities. Recently, graph-based approaches have been proposed to address this issue, but the cost of transmitting graph topology often overcome the gain of such techniques. In this work we propose a new Superpixel-driven Graph Transform (SDGT) that uses clusters of superpixels, which have the ability to adhere nicely to edges in the image, as coding blocks and computes inside these homogeneously colored regions a graph transform which is shape-adaptive. Doing so, only the borders of the regions and the transform coefficients need to be transmitted, in place of all the structure of the graph. The proposed method is finally compared to DCT and the experimental results show how it is able to outperform DCT both visually and in term of PSNR.
Giulia Fracastoro, Francesco Verdoja, Marco Grangetto, Enrico Magli
ICIP3
2015 Pollution-resilient peer-to-peer video streaming with Band Codes
abstract
Band Codes (BC) have been recently proposed as a solution for controlled-complexity random Network Coding (NC) in mobile applications, where energy consumption is a major concern. In this paper, we investigate the potential of BC in a peer-to-peer video streaming scenario where malicious and honest nodes coexists. Malicious nodes launch the so called pollution attack by randomly modifying the content of the coded packets they forward to downstream nodes, preventing honest nodes from correctly recovering the video stream. Whereas in much of the related literature this type of attack is addressed by identifying and isolating the malicious nodes, in this work we propose to address it by adaptively adjusting the coding scheme so to introduce resilience against pollution propagation. We experimentally show the impact of a pollution attack in a defenseless system and in a system where the coding parameters of BC are adaptively modulated following the discovery of polluted packets in the network. We observe that just by tuning the coding parameters, it is possible to reduce the impact of a pollution attack and restore the quality of the video communication.
Attilio Fiandrotti, Rossano Gaeta, Marco Grangetto
ICME3
2015 High capacity reversible data hiding and content protection for radiographic images
Davide Cavagnino, Maurizio Lucenteforte, Marco Grangetto
Signal Process.3
2015 Panorama View With Spatiotemporal Occlusion Compensation for 3D Video Coding
abstract
The future of novel 3D display technologies largely depends on the design of efficient techniques for 3D video representation and coding. Recently, multiple view plus depth video formats have attracted many research efforts since they enable intermediate view estimation and permit to efficiently represent and compress 3D video sequences. In this paper, we present spatiotemporal occlusion compensation with panorama view (STOP), a novel 3D video coding technique based on the creation of a panorama view and occlusion coding in terms of spatiotemporal offsets. The panorama picture represents the most of the visual information acquired from multiple views using a single virtual view, characterized by a larger field of view. Encoding the panorama video with state-of-the-art HECV and representing occlusions with simple spatiotemporal ancillary information STOP achieves high-compression ratio and good visual quality with competitive results with respect to competing techniques. Moreover, STOP enables free viewpoint 3D TV applications whilst allowing legacy display to get a bidimensional service using a standard video codec and simple cropping operations.
Muhammad Shahid Farid, Maurizio Lucenteforte, Marco Grangetto
IEEE Trans. Image Process.3
2015 Simple Countermeasures to Mitigate the Effect of Pollution Attack in Network Coding-Based Peer-to-Peer Live Streaming
abstract
Network coding (NC)-based peer-to-peer (P2P) streaming represents an effective solution to aggregate user capacities and to increase system throughput in live multimedia streaming. Nonetheless, such systems are vulnerable to pollution attacks where a handful of malicious peers can disrupt the communication by transmitting just a few bogus packets which are then recombined and relayed by unaware honest nodes, further spreading the pollution over the network. Whereas previous research focused on malicious nodes identification schemes and pollution-resilient coding, in this paper we show pollution countermeasures which make a standard NC scheme resilient to pollution attacks. Thanks to a simple yet effective analytical model of a reference node collecting packets by malicious and honest neighbors, we demonstrate that: i) packets received earlier are less likely to be polluted, and ii) short generations increase the likelihood to recover a clean generation. Therefore, we propose a recombination scheme where nodes draw packets to be recombined according to their age in the input queue, paired with a decoding scheme able to detect the reception of polluted packets early in the decoding process and short generations. The effectiveness of our approach is experimentally evaluated in a real system we developed and deployed on hundreds to thousands of peers. Experimental evidence shows that, thanks to our simple countermeasures, the effect of a pollution attack is almost canceled and the video quality experienced by the peers is comparable to pre-attack levels.
Attilio Fiandrotti, Rossano Gaeta, Marco Grangetto
IEEE Trans. Multim.3
2015 Exploiting Rateless Codes in Cloud Storage Systems
abstract
Block-level cloud storage (BLCS) offers to users and applications the access to persistent block storage devices (virtual disks) that can be directly accessed and used as if they were raw physical disks. In this paper we devise ENIGMA, an architecture for the back-end of BLCS systems able to provide adequate levels of access and transfer performance, availability, integrity, and confidentiality, for the data it stores. ENIGMA exploits LT rateless codes to store fragments of sectors on storage nodes organized in clusters. We quantitatively evaluate how the various ENIGMA system parameters affect the performance, availability, integrity, and confidentiality of virtual disks. These evaluations are carried out by using both analytical modeling (for availability, integrity, and confidentiality) and discrete event simulation (for performance), and by considering a set of realistic operational scenarios. Our results indicate that it is possible to simultaneously achieve all the objectives set forth for BLCS systems by using ENIGMA, and that a careful choice of the various system parameters is crucial to achieve a good compromise among them. Moreover, they also show that LT coding-based BLCS systems outperform traditional BLCS systems in all the aspects mentioned before.
Cosimo Anglano, Rossano Gaeta, Marco Grangetto
IEEE Trans. Parallel Distributed Syst.3
2014 A panoramic 3D video coding with directional depth aided inpainting
abstract
The success of 3D and free-viewpoint television largely depends on the efficient representation and compression of 3D video in addition to viable rendering methods. This paper presents a novel 3D video coding technique based on the creation of a panorama view to compact the information of a stereoscopic pair. The panorama view represents the information that would be visible to a virtual camera with a larger field of view embracing all the available views. The information in the panorama view is then used to estimate any intermediate view using depth image based rendering. Furthermore, to fill the disocclusions in the reconstructed view a directional depth aided fast marching inpainting technique is presented. The panorama view and corresponding depth map are amenable to standard video compression. In this paper we show that using the novel HEVC standard the proposed 3D video format can be compressed very efficiently.
Muhammad Shahid Farid, Maurizio Lucenteforte, Marco Grangetto
ICIP3
2014 Edge enhancement of depth based rendered images
abstract
Depth image based rendering is a well-known technology for the generation of virtual views in between a limited set of views acquired by a cameras array. Intermediate views are rendered by warping image pixels based on their depth. Nonetheless, depth maps are usually imperfect as they need to be estimated through stereo matching algorithms; moreover, for representation and transmission requirements depth values are obviously quantized. Such depth representation errors translate into a warping error when generating intermediate views thus impacting on the rendered image quality. We observe that depth errors turn to be very critical when they affect the object contours since in such a case they cause significant structural distortion in the warped objects. This paper presents an algorithm to improve the visual quality of the synthesized views by enforcing the shape of the edges in presence of erroneous depth estimates. We show that it is possible to significantly improve the visual quality of the interpolated view by enforcing prior knowledge on the admissible deformations of edges under projective transformation. Both visual and objective results show that the proposed approach is very effective.
Muhammad Shahid Farid, Maurizio Lucenteforte, Marco Grangetto
ICIP3
2014 Automatic method for tumor segmentation from 3-points dynamic PET acquisitions
abstract
In this paper a novel technique to segment tumor voxels in dynamic positron emission tomography (PET) scans is proposed. An innovative anomaly detection tool tailored for 3-points dynamic PET scans is designed. The algorithm allows the identification of tumoral cells in dynamic FDG-PET scans thanks to their peculiar anaerobic metabolism experienced over time. The proposed tool is preliminarily tested on a small dataset showing promising performance as compared to the state of the art in terms of both accuracy and classification errors.
Francesco Verdoja, Marco Grangetto, Christian Bracco, Teresio Varetto, Manuela Racca, Michele Stasi
ICIP2
2014 Exploiting Rateless Codes and Belief Propagation to Infer Identity of Polluters in MANET
abstract
In this paper, we consider a scenario where nodes in a MANET disseminate data chunks using rateless codes. Any node is able to successfully decode any chunk by collecting enough coded blocks from several other nodes without any coordination. We consider the problem of identifying malicious nodes that launch a pollution attack by deliberately modifying the payload of coded blocks before transmitting. It follows that the original chunk can only be obtained if there are no malicious nodes among the chunk providers. In this paper we propose SIEVE, a fully distributed technique to infer the identity of malicious nodes. A node creates what we termed a check whenever a chunk is decoded; a check is a pair composed of the set of other nodes that provided coded blocks used to decode the chunk (the chunk uploaders) and a flag indicating whether the chunk is corrupted or not. SIEVE exploits rateless codes to detect chunk integrity and belief propagation to infer the identity of malicious nodes. In particular, every node autonomously constructs its own bipartite graph (a.k.a. factor graph in the literature) whose vertexes are checks and nodes, respectively. Then, it periodically runs the belief propagation algorithm on its factor graph to infer the probability of other nodes being malicious. We show by running detailed simulations using ns-3 that SIEVE is very accurate and robust under several attack scenarios and deceiving actions. We discuss how the topological properties of the factor graph impacts SIEVE performance and show that nodes speed in the MANET plays a role on the identification accuracy. Furthermore, an interesting trade-off between coding efficiency and SIEVE accuracy, completeness, and reactivity is discovered. We also show that SIEVE is efficient requiring low computational, memory, and communication resources.
Rossano Gaeta, Marco Grangetto, Riccardo Loti
IEEE Trans. Mob. Comput.2
2014 Band Codes for Energy-Efficient Network Coding With Application to P2P Mobile Streaming
abstract
A key problem in network coding (NC) lies in the complexity and energy consumption associated with the packet decoding processes, which hinder its application in mobile environments. Controlling and hence limiting such factors has always been an important but elusive research goal, since the packet degree distribution, which is the main factor driving the complexity, is altered in a non-deterministic way by the random recombinations at the network nodes. In this paper we tackle this problem with a new approach and propose Band Codes (BC), a novel class of network codes specifically designed to preserve the packet degree distribution during packet encoding, recombination and decoding. BC are random codes over GF(2) that exhibit low decoding complexity, feature limited and controlled degree distribution by construction, and hence allow to effectively apply NC even in energy-constrained scenarios. In particular, in this paper we motivate and describe our new design and provide a thorough analysis of its performance. We provide numerical simulations of the BC performance in order to validate the analysis and assess the overhead of BC with respect to a conventional random NC scheme. Moreover, experiment in a real-world application, namely peer-to-peer mobile media streaming using a random-push protocol, show that BC reduce the decoding complexity by a factor of two with negligible increase of the encoding overhead, paving the way for the application of NC to power-constrained devices.
Attilio Fiandrotti, Valerio Bioglio, Marco Grangetto, Rossano Gaeta, Enrico Magli
IEEE Trans. Multim.3
2014 DIP: Distributed Identification of Polluters in P2P Live Streaming
abstract
Peer-to-peer live streaming applications are vulnerable to malicious actions of peers that deliberately modify data to decrease or prevent the fruition of the media (pollution attack). In this article we propose DIP , a fully distributed, accurate, and robust algorithm for the identification of polluters. DIP relies on checks that are computed by peers upon completing reception of all blocks composing a data chunk. A check is a special message that contains the set of peer identifiers that provided blocks of the chunk as well as a bit to signal if the chunk has been corrupted. Checks are periodically transmitted by peers to their neighbors in the overlay network; peers receiving checks use them to maintain a factor graph. This graph is bipartite and an incremental belief propagation algorithm is run on it to compute the probability of a peer being a polluter. Using a prototype deployed over PlanetLab we show by extensive experimentation that DIP allows honest peers to identify polluters with very high accuracy and completeness, even when polluters collude to deceive them. Furthermore, we show that DIP is efficient, requiring low computational, communication, and storage overhead at each peer.
Rossano Gaeta, Marco Grangetto, Lorenzo Bovio
ACM Trans. Multim. Comput. Commun. Appl.2
2014 Rateless Codes and Random Walksfor P2P Resource Discovery in Grids
abstract
Peer-to-peer (P2P) resource location techniques in grid systems have been recently investigated to obtain scalability, reliability, efficiency, fault-tolerance, security, and robustness. Query resolution for locating resources and update information on their own resource status in these systems can be abstracted as the problem of allowing one peer to obtain a local view of global information defined on all peers of a P2P unstructured network. In this paper, the system is represented as a set of nodes connected to form a P2P network where each node holds a piece of information that is required to be communicated to all the participants. Moreover, we assume that the information can dynamically change and that each peer periodically requires to access the values of the data of all other peers. A novel approach based on a continuous flow of control packets exchanged among the nodes using the random walk principle and rateless coding is proposed. An innovative rateless decoding mechanism that is able to cope with asynchronous information updates is also proposed. The performance of the proposed system is evaluated both analytically and experimentally by simulation. The analytical results show that the proposed strategy guarantees quick diffusion of the information and scales well to large networks. Simulations show that the technique is effective also in presence of network and information dynamics.
Valerio Bioglio, Rossano Gaeta, Marco Grangetto, Matteo Sereno
IEEE Trans. Parallel Distributed Syst.3
2013 Depth image based rendering with inverse mapping
abstract
Three-dimensional video has gained much attention during the last decade due its vast applications in cinema, television, animation and virtual reality. The design of intermediate view synthesis algorithms that are efficient both in terms of computational complexity and visual quality is a paramount goal in the fields of 3D free view point television and displays. This papers focuses on the design of a low complexity view synthesis algorithm that produces better quality of the virtual image. A novel view synthesis technique to create a virtual view from two video sequences with corresponding depths is proposed. The technique employs low complexity integer pixel precision warping and a novel approach for hole filling based on inverse mapping. The proposed technique is tested over a number of video sequences and compared with existing state of the art methods, yielding excellent results both in terms of signal to noise ratio and visual quality.
Muhammad Shahid Farid, Maurizio Lucenteforte, Marco Grangetto
MMSP3
2013 Edges shape enforcement for visual enhancement of depth image based rendering
abstract
Depth image based rendering of intermediate views with high visual quality remains a challenging goal in presence of estimated and quantized depth values. Among the other rendering artifacts we observed that edges are usually affected by significant warping errors. In particular, because of depth estimation inaccuracy around object boundaries the edges may completely loose their original shape during the warping process. Nonetheless, edges represent one of the most important cues for the human visual system. In this paper a novel technique aiming at improving the edge rendering is presented. As opposed to previous approaches, the technique exploits only texture information, thus avoiding possible errors in depth estimation. The idea is based on the enforcement of prior knowledge of the edge shape under projective transformation. The proposed algorithm works in two steps: first the damaged edges of the warped image are detected, then these latter are corrected so as to better approximate their shape in the reference view. Finally the corrected edges are rendered within the intermediate image without introducing noticeable texture artifacts. The proposed algorithm has been tested on a variety of standard video sequences exhibiting excellent results in terms of rendered image visual quality.
Muhammad Shahid Farid, Maurizio Lucenteforte, Marco Grangetto
MMSP3
2013 Gait characterization using dynamic skeleton acquisition
abstract
Human gait is an important biometric feature for automatic people recognition. Biometric methodologies are generally intrusive and require the collaboration of the subject in order to perform accurate data acquisition. Gait, instead, can be captured at a distance and without collaboration. This makes it an unobtrusive method for recognizing people in video surveillance systems. In this paper we propose a method to characterize walking gait using three-dimensional skeleton information acquired by the Microsoft Kinect sensor. A set of static and dynamic features correlated to human gait are extracted by the estimated skeleton joint positions. Moreover, we proposed to describe joints positions in a coordinate reference system oriented according to the walking direction to better represents the movement of human body. Using unsupervised clustering over a set of 20 subjects we analyze the effectiveness of the selected features in discriminating people gaits. It turns out that a few dynamic parameters involving the movement of knees, elbows and head are good candidates for robust gait characterization.
Elena Gianaria, Nello Balossino, Marco Grangetto, Maurizio Lucenteforte
MMSP3
2013 A practical Random Network Coding scheme for data distribution on peer-to-peer networks using rateless codes
Valerio Bioglio, Marco Grangetto, Rossano Gaeta, Matteo Sereno
Perform. Evaluation2
2013 Identification of Malicious Nodes in Peer-to-Peer Streaming: A Belief Propagation-Based Technique
abstract
Peer-to-peer streaming has witnessed a great success thanks to the possibility of aggregating resources from all participants. Nevertheless, performance of the entire system may be highly degraded due to the presence of malicious peers that share bogus data on purpose. In this paper, we propose to use a statistical inference technique, namely, belief propagation (BP), to estimate the probability of peers being malicious. The detection algorithm is run by a set of trusted monitor nodes that receives notification messages (checks) from peers whenever they obtain a chunk of data; these checks contain the list of the chunk uploaders and a flag to mark the chunk as polluted or clean. Peers are able to detect if the received chunk is polluted or not but, since multiparty download is employed, they are not capable to identify the source(s) of bogus blocks. This problem definition allows us to define a factor graph of peers and checks on which an incremental version of the belief propagation algorithm is run by the monitor nodes to infer the probability of each peer being a malicious one. We evaluate the accuracy, robustness, and complexity of our technique by running a real peer-to-peer application on PlanetLab. We show that the proposed approach is very accurate and robust against malicious nodes misbehaving (different pollution intensity, presence of fake checks, churning, and total uncooperation from malicious nodes), increasing number and colluding behavior of malicious nodes.
Rossano Gaeta, Marco Grangetto
IEEE Trans. Parallel Distributed Syst.2
2012 Band Codes: Controlled Complexity Network Coding for Peer-to-Peer Video Streaming
abstract
We present Band Codes (BC), a novel class of rate less codes that makes possible to control the computational complexity of Network Coding (NC). NC increases throughput of the networks via packet recombinations at the network nodes. In a NC scenario based on rate less codes, the recombinations at the nodes alter the packet degree distribution selected at the source and increase the computational complexity of the packet decoding process. Unlike other classes of rate less codes, BC preserve the degree distribution of the encoded packets through the recombinations at the nodes. Furthermore, BC enable to control the decoding complexity of each network node independently from the rest of the network. We evaluate BC in a P2P scenario using a purposely designed random-push protocol for live video streaming. The experiments show that BC achieve high encoding efficiency, enable nodes with different computational capabilities to coexist within the same network and reduce the processor load on a real mobile device by nearly 50%.
Attilio Fiandrotti, Valerio Bioglio, Enrico Magli, Marco Grangetto, Rossano Gaeta
ICME4
2012 SIEVE: A Distributed, Accurate, and Robust Technique to Identify Malicious Nodes in Data Dissemination on MANET
abstract
In this paper we consider the following problem: nodes in a MANET must disseminate data chunks using rateless codes but some nodes are assumed to be malicious, i.e., before transmitting a coded packet they may modify its payload. Nodes receiving corrupted coded packets are prevented from correctly decoding the original chunk. We propose SIEVE, a fully distributed technique to identify malicious nodes. SIEVE is based on special messages called checks that nodes periodically transmit. A check contains the list of nodes identifiers that provided coded packets of a chunk as well as a flag to signal if the chunk has been corrupted. SIEVE operates on top of an otherwise reliable architecture and it is based on the construction of a factor graph obtained from the collected checks on which an incremental belief propagation algorithm is run to compute the probability of a node being malicious. Analysis is carried out by detailed simulations using ns-3. We show that SIEVE is very accurate and discuss how nodes speed impacts on its accuracy. We also show SIEVE robustness under several attack scenarios and deceiving actions.
Rossano Gaeta, Marco Grangetto, Riccardo Loti
ICPADS2
2012 An adaptive hybrid CDN/P2P solution for Content Delivery Networks
abstract
Streaming services have grown rapidly in the last few years and providers of video on-demand, such as Netflix or YouTube, are increasing the number of users even more quickly. The majority of these companies implement their services using huge Content Delivery Networks that are as much powerful as expensive, e.g. Amazon and Akamai. In this paper we propose a hybrid CDN/P2P solution that aims at reducing the infrastructural costs exploiting local caching and P2P while guaranteeing an optimal quality of service. The proposed architecture uses a classic CDN complemented by a geographically distributed layer where P2P can be activated exploiting network, content awareness and locality. The performance of the proposed solution is evaluated by means of a prototype implementation that has been deployed using the PlanetLab network and the Amazon AWS cloud services. Our findings show that the proposed approach provides adaptive, flexible, scalable and content centric service to the end users while significantly reducing the infrastructural costs.
Francesco Bronzino, Rossano Gaeta, Marco Grangetto, Giovanni Pau 0001
VCIP3
2012 An HTML5 player for a gstreamer based MPEG DASH client
abstract
The proposed demo shows how a streaming client compliant with MPEG DASH [1] standard and developed using a GStreamer media framework [2], can be integrated into an HTML5 enabled web browser. Through a non-standard interface, the GStreamer adaptive streaming client can be controlled by a JavaScript engine implemented in the web browser.
Emanuele Quacchio, G. Bruno, Marco Grangetto
VCIP3
2012 A study of an hybrid CDN-P2P system over the PlanetLab network
Enrico Baccaglini, Marco Grangetto, Emanuele Quacchio, Simone Zezza
Signal Process. Image Commun.2
2011 Tile format: A novel frame compatible approach for 3D video broadcasting
abstract
The success of the 3D TV broadcasting service largely depends on the technical, economical and user experience aspects. Nevertheless, it is quite clear that one of the keys for the successful deployment of 3D TV is the offer of a service which permits a seamless transition to 3D by exploiting the existing infrastructure and guaranteeing 2D backward compatibility towards legacy receivers, whilst only requiring a moderate increase in the transmission bandwidth. To this end, the DVB has just drafted the first phase of the 3D TV specification and has selected a number of frame compatible format arrangements for the multiplexing of 3D in a standard video sequence. In this paper we propose the novel Tile Format frame compatible arrangement and analyze its advantages over other existing solutions in terms of both image quality and backward compatibility. Furthermore, we detail a real 3D TV broadcasting trial which is currently being conducted in our country and is based upon the proposed video format.
Giovanni Ballocca, Paolo D'Amato, Marco Grangetto, Maurizio Lucenteforte
ICME3
2011 An optimal partial decoding algorithm for rateless codes
abstract
Rateless codes are designed to decode all the input symbols when a certain number of coded symbols have been received. However, it is possible to recover a subset of the input symbols from the actually received coded symbols: this process is called partial decoding and the number of recovered input symbols is termed the intermediate performance of rateless codes. In this paper we study the problem of the optimality of the partial decoding process: we say that a partial decoding algorithm is optimal if, given a rateless code, it is able to maximize the intermediate performance of the code, i.e. it is able to retreive the maximum number of input symbols when a certain number n of coded symbols have been received, for every n. We propose OPD, an optimal partial decoding algorithm for any rateless code, proving its optimality. The proposed algorithm is finally used to analyze the intermediate performance of LT codes.
Valerio Bioglio, Marco Grangetto, Rossano Gaeta, Matteo Sereno
ISIT2
2011 A game theory framework for ISP streaming traffic management
Valerio Bioglio, Rossano Gaeta, Marco Grangetto, Matteo Sereno, Salvatore Spoto
Perform. Evaluation3
2011 Transparent encryption techniques for H.264/AVC and H.264/SVC compressed video
Enrico Magli, Marco Grangetto, Gabriella Olmo
Signal Process.2
2010 Distributed joint source-channel arithmetic coding
abstract
We address distributed source coding with decoder side information, when the decoder observes the source through a noisy channel. Existing approaches employ syndromeor parity-based channel codes. We propose a new approach based on distributed arithmetic coding (DAC).We introduce a DAC with forbidden symbol, which allows to tune the redundancy according to the amount of channel noise. We propose a novel sequential decoder that employs the known side information to decode the corrupted codeword. Experimental results show that the proposed scheme is better than parity-based turbo codes at relatively short block lengths.
Marco Grangetto, Enrico Magli, Gabriella Olmo
ICIP1
2010 Fountains vs Torrents: The P2P ToroVerde Protocol
abstract
In this paper we present ToroVerde, a novel push-based peer-to-peer (P2P) content distribution application exploiting the digital fountain concept through the use of rateless codes. We provide the protocol specification, then describe the simulator and the complete prototype we have developed for Planetlab deployment and testing. To this end, we consider flash crowd and steady arrival patterns as well as highly churning systems to perform a preliminary analysis of the potential advantages of introducing rateless codes. We present results from PlanetLab experiments compared against performance of BitTorrent, that is widely considered as the reference system for content distribution. We also present a few simulation results showing the behavior of ToroVerde as the number of peers in the systems increases. Our results suggest that ToroVerde has the potential of reducing the average download time for small-to-medium sized files in overlays composed of a few hundred peers with a small increase of the communication overhead.
Andrea Magnetto, Salvatore Spoto, Rossano Gaeta, Marco Grangetto, Matteo Sereno
MASCOTS4
2010 Local Access to Sparse and Large Global Information in P2P Networks: A Case for Compressive Sensing
abstract
In this paper we face the following problem: how to provide each peer local access to the full information (not just a summary) that is distributed over all \emph{edges} of an overlay network? How can this be done if local access is performed at a given rate? We focus on \emph{large and sparse} information and we propose to exploit the compressive sensing (CS) theory to efficiently collect and pro-actively disseminate this information across a large overlay network. We devise an approach based on random walks (RW) to spread CS random combinations to participants in a random peer-to-peer (P2P) overlay network. CS allows the peer to compress the RW payload in a distributed fashion: given a constraint on the RW size, e.g., the maximum UDP packet payload size, this amounts to being able to distribute larger information and to guarantee that a large fraction of the global information is obtained by each peer. We analyze the performance of the proposed method by means of a simple (yet accurate) analytical model describing the structure of the so called CS sensing matrix in presence of peer dynamics and communication link failures. We validate our model predictions against a simulator of the system at the peer and network level on different models of random overlay networks. The model we developed can be exploited to select the parameters of the RW and the criteria to build the sensing matrix in order to achieve successful information recovery. Finally, a prototype has been developed and deployed over the PlanetLab network to prove the feasibility of the proposed approach in a realistic environment. Our analysis reveals that the method we propose is feasible, accurate and robust to peer and information dynamics. We also argue that centralized and other distributed approaches, i.e., flooding and gossiping, are unfit in the context we consider.
Rossano Gaeta, Marco Grangetto, Matteo Sereno
Peer-to-Peer Computing2
2010 Sliding-Window Raptor Codes for Efficient Scalable Wireless Video Broadcasting With Unequal Loss Protection
abstract
Digital fountain codes have emerged as a low-complexity alternative to Reed-Solomon codes for erasure correction. The applications of these codes are relevant especially in the field of wireless video, where low encoding and decoding complexity is crucial. In this paper, we introduce a new class of digital fountain codes based on a sliding-window approach applied to Raptor codes. These codes have several properties useful for video applications, and provide better performance than classical digital fountains. Then, we propose an application of sliding-window Raptor codes to wireless video broadcasting using scalable video coding. The rates of the base and enhancement layers, as well as the number of coded packets generated for each layer, are optimized so as to yield the best possible expected quality at the receiver side, and providing unequal loss protection to the different layers according to their importance. The proposed system has been validated in a UMTS broadcast scenario, showing that it improves the end-to-end quality, and is robust towards fluctuations in the packet loss rate.
Pasquale Cataldi, Marco Grangetto, Tammam Tillo, Enrico Magli, Gabriella Olmo
IEEE Trans. Image Process.2
2010 TURINstream: A Totally pUsh, Robust, and effIcieNt P2P Video Streaming Architecture
abstract
This paper presents TURINstream, a novel P2P video streaming architecture designed to jointly achieve low delay, robustness to peer churning, limited protocol overhead, and quality-of-service differentiation based on peers cooperation. Separate control and video overlays are maintained by peers organized in clusters that represent sets of collaborating peers. Clusters are created by means of a distributed algorithm and permit the exploitation of the participant nodes upload capacity. The video is conveyed with a push mechanism by exploiting the advantages of multiple description coding. TURINstream design has been optimized through an event driven overlay simulator able to scale up to tens of thousands of peers. A complete prototype of TURINstream has been developed, deployed, and tested on PlanetLab. We tested our prototype under varying degree of peer churn, flash crowd arrivals, sudden massive departures, and limited upload bandwidth resources. TURINstream fulfills our initial design goals, showing low average connection, startup, and playback delays, high continuity index, low control overhead, and effective quality-of-service differentiation in all tested scenarios.
Andrea Magnetto, Rossano Gaeta, Marco Grangetto, Matteo Sereno
IEEE Trans. Multim.3
2009 Rateless codes network coding for simple and efficient P2P video streaming
abstract
The goal of this paper is the development of network coding solutions able to improve the performance of video streaming applications over peer-to-peer overlays. Recent advances in P2P protocols have shown that rateless codes can be profitably applied to P2P video streaming with several advantages in terms of protocol efficiency and simplification, e.g. push based video delivery, no need of packet reconciliation at the decoder. In this paper existing and novel network coding techniques based on rateless codes are presented and compared, showing that rateless codes, besides simplifying the protocol design, can significantly reduce the startup and playback delays. The proposed protocol is evaluated on real topologies, obtained by crawling the widespread PPLive video streaming application. The reported experimental results show that the proposed protocol significantly reduces the startup and playback delay and allows one to increase the bitrate devoted to the video stream.
Marco Grangetto, Rossano Gaeta, Matteo Sereno
ICME1
2009 Seacast: A protocol for peer-to-peer video streaming supporting multiple description coding
abstract
SEACAST is a peer-to-peer live streaming protocol developed at Politecnico di Torino, which aims at improving current systems in two key areas. The first is the use of fullfledged flow control using RTP/UDP and session signaling. The second is the use of multiple description coding to handle error resilience and user heterogeneity. In this paper we overview SEACAST, highlighting its main innovations, and providing a short summary of performance evaluation over a local testbed at Politecnico di Torino. The results show a definite performance improvement with respect to existing systems, and point out the usefulness of multiple description coding in the peer-to-peer context.
Simone Zezza, Enrico Magli, Gabriella Olmo, Marco Grangetto
ICME4
2009 Analysis of PPLive through active and passive measurements
abstract
The P2P-IPTV is an emerging class of Internet applications that is becoming very popular. The growing popularity of these rather bandwidth demanding multimedia streaming applications has the potential to flood the Internet with a huge amount of traffic. In this paper we present an investigation of the popular P2P-IPTV application PPLive exploiting a measurement strategy that combines both active and passive measures. To this end, we use a crawler that allows the study of the topological characteristics of the overlay of one of the PPLive channels; concurrently, we perform passive measures on a PPlive client we run to join the crawled channel. We successively cross correlate information we obtained from the two measurements to assess the accuracy of the data captured by the crawler. Our results reveal the potentials and the limits of PPLive active measures strategies.
Salvatore Spoto, Rossano Gaeta, Marco Grangetto, Matteo Sereno
IPDPS3
2008 Decoder-driven adaptive distributed arithmetic coding
abstract
We propose a distributed source coding system for data collected by sensor networks. It uses a feedback channel between the sensors and the gateway node (i.e., the joint decoder) but, unlike previous systems, the encoding process is driven by the decoder. Compression is performed using distributed arithmetic coding, which is extended to adaptively estimate the source probabilities. Specifically, the decoder estimates marginal and conditional probabilities, and sends them back to the sensors to drive the distributed arithmetic coding process. This reduces the decoding delay, and potentially eliminates the need of rate-compatible Slepian-Wolf codes.
Marco Grangetto, Enrico Magli, Gabriella Olmo
ICIP1
2008 Redundant Slice Optimal Allocation for H.264 Multiple Description Coding
abstract
In this paper, a novel H.264 multiple description technique is proposed. The coding approach is based on the redundant slice representation option, defined in the H.264 standard. In presence of losses, the redundant representation can be used to replace missing portions of the compressed bit stream, thus yielding a certain degree of error resilience. This paper addresses the creation of two balanced descriptions based on the concept of redundant slices, while keeping full compatibility with the H.264 standard syntax and decoding behavior in case of single description reception. When two descriptions are available still a standard H.264 decoder can be used, given a simple preprocessing of the received compressed bit streams. An analytical setup is employed in order to optimally select the amount of redundancy to be inserted in each frame, taking into account both the transmission condition and the video decoder error propagation. Experimental results demonstrate that the proposed technique favorably compares with other H.264 multiple description approaches.
Tammam Tillo, Marco Grangetto, Gabriella Olmo
IEEE Trans. Circuits Syst. Video Technol.2
2007 Conditional Access to H.264/AVC Video by Means of Redundant Slices
abstract
In this paper a novel conditional access scheme for the distribution of H.264/AVC video is presented. The algorithm permits to cypher the full quality video, while guaranteeing free access to a reduced quality video that can be used as preview. The scheme is based on the novel coding options, introduced in the H.264 standard, and therefore is fully compliant with the standard syntax. The secured video stream exhibits a very small rate overhead and requires limited supplementary computational cost at coding time.
Marco Grangetto, Enrico Magli, Gabriella Olmo
ICIP (6)1
2007 H.264 Multiple Description Coding Based on Redundant Picture Representation
abstract
In this paper a novel H.264 multiple description technique is proposed. The coding approach is based on the redundant slice representation option, defined in the H.264 standard. In presence of losses, the redundant representation can be used to replace missing portions of the compressed bitstream, thus yielding a certain degree of error resilience. This paper addresses the creation of two balanced descriptions based on the concept of redundant slices, while keeping full compatibility with the H.264 standard syntax and decoding behavior. Moreover, a practical algorithm for redundancy tuning as a function of the packet loss rate is provided. Experimental results demonstrate that the proposed technique favorably compares with other H.264 multiple description approaches.
Tammam Tillo, Marco Grangetto, Gabriella Olmo
ICIP (4)2
2007 Sliding-Window Digital Fountain Codes for Streaming of Multimedia Contents
abstract
Digital fountain codes are becoming increasingly important for multimedia communications over networks subject to packet erasures. These codes have significantly lower complexity than Reed-Solomon ones, exhibit high erasure correction performance, and are very well suited to generating multiple equally important descriptions of a source. In this paper we propose an innovative scheme for streaming multimedia contents by using digital fountain codes applied over sliding windows, along with a suitably modified belief-propagation decoder. The use of overlapped windows allows one to have a virtually extended block, which yields superior performance in terms of packet recovery. Simulation results using LT codes show that the proposed algorithm has better performance in terms of efficiency, reliability and memory with respect to fixed-window encoding.
Mattia C. O. Bogino, Pasquale Cataldi, Marco Grangetto, Enrico Magli, Gabriella Olmo
ISCAS3
2007 Symmetric Distributed Arithmetic Coding of Correlated Sources
abstract
We propose a new scheme for symmetric Slepian-Wolf coding of correlated binary sources. Unlike previous designs that employ capacity-achieving channel codes, the proposed scheme is based on arithmetic codes with error correction capability. We define a time-sharing version of a distributed arithmetic coder, and a soft joint decoder. Experimental results on two sources show that, for short block length, the proposed scheme outperforms the symmetric turbo code design in (Stankovic et al., 2006).
Marco Grangetto, Enrico Magli, Gabriella Olmo
MMSP1
2007 Joint Source, Channel Coding, and Secrecy
Enrico Magli, Marco Grangetto, Gabriella Olmo
EURASIP J. Inf. Secur.2
2007 On Modeling Mismatch Errors Induced by Different Quantizers
abstract
In this letter, the mismatch error due to the replacement of a fine with a coarse quantizer is considered, and an analytical model is proposed to describe the related distortion. Simulations show that this model is highly accurate and can be used to estimate the expected distortion of DPCM-based codecs in order to better allocating the rate. For highly correlated sources, this leads to a gain of 1 to 1.5 dB over an exhaustive search method that adopts a uniform redundancy allocation. Moreover, it permits to allocate the redundancy by an easy-to-solve analytical model of the system.
Tammam Tillo, Marco Grangetto, Gabriella Olmo
IEEE Signal Process. Lett.2
2007 Iterative Decoding of Serially Concatenated Arithmetic and Channel Codes With JPEG 2000 Applications
abstract
In this paper, an innovative joint-source channel coding scheme is presented. The proposed approach enables iterative soft decoding of arithmetic codes by means of a soft-in soft- out decoder based on suboptimal search and pruning of a binary tree. An error-resilient arithmetic coder with a forbidden symbol is used in order to improve the performance of the joint source/channel scheme. The performance in the case of transmission across the AWGN channel is evaluated in terms of word error probability and compared to a traditional separated approach. The interleaver gain, the convergence property of the system, and the optimal source/channel rate allocation are investigated. Finally, the practical relevance of the proposed joint decoding approach is demonstrated within the JPEG 2000 coding standard. In particular, an iterative channel and JPEG 2000 decoder is designed and tested in the case of image transmission across the AWGN channel.
Marco Grangetto, Bartolo Scanavino, Gabriella Olmo, Sergio Benedetto
IEEE Trans. Image Process.1
2007 Multiple Description Image Coding Based on Lagrangian Rate Allocation
abstract
In this paper, a novel multiple description coding technique is proposed, based on optimal Lagrangian rate allocation. The method assumes the coded data consists of independently coded blocks. Initially, all the blocks are coded at two different rates. Then blocks are split into two subsets with similar rate distortion characteristics; two balanced descriptions are generated by combining code blocks belonging to the two subsets encoded at opposite rates. A theoretical analysis of the approach is carried out, and the optimal rate distortion conditions are worked out. The method is successfully applied to the JPEG 2000 standard and simulation results show a noticeable performance improvement with respect to state-of-the art algorithms. The proposed technique enables easy tuning of the required coding redundancy. Moreover, the generated streams are fully compatible with Part 1 of the standard.
Tammam Tillo, Marco Grangetto, Gabriella Olmo
IEEE Trans. Image Process.2
2006 Conditional Access to H.264/AVC Video with Drift Control
abstract
In this paper we address the problem of providing conditional access to video sequences, namely, to generate a low-quality video to be used as preview, which can be decoded at full quality if a decryption key is obtained. We propose and investigate the performance of two different techniques, based on smoothing and separate encoding in the compressed domain, and motion vector perturbation. We show that these techniques are able to provide conditional access to different quality levels of H.264/AVC video with very small rate overhead, and that their combination can provide different levels of security towards malicious attacks
Enrico Magli, Marco Grangetto, Gabriella Olmo
ICME2
2006 Scalable Image Retrieval from Distributed Images Database
abstract
In order to store, and retrieve images from large databases, we propose a framework, based on multiple description coding paradigms, that disseminates images over distributed servers. Consequently, decentralized download can be performed, thus reducing links overload and hotspot areas without penalizing downloads speed. Moreover, the tradeoff between system reliability and storage requirement can be achieved by tuning descriptions redundancy, thus providing high flexibility in terms of storage resources, reliability of access, and performance. The scalability of the proposed framework is achieved by the intrinsic progressivity of the multiple description schemes. Moreover, we demonstrate that system can work properly regardless of server crashes
Tammam Tillo, Marco Grangetto, Gabriella Olmo
ICME2
2006 A syntax-preserving error resilience tool for JPEG 2000 based on error correcting arithmetic coding
abstract
JPEG 2000 is the novel ISO standard for image and video coding. Besides its improved coding efficiency, it also provides a few error resilience tools in order to limit the effect of errors in the codestream, which can occur when the compressed image or video data are transmitted over an error-prone channel, as typically occurs in wireless communication scenarios. However, for very harsh channels, these tools often do not provide an adequate degree of error protection. In this paper, we propose a novel error-resilience tool for JPEG 2000, based on the concept of ternary arithmetic coders employing a forbidden symbol. Such coders introduce a controlled degree of redundancy during the encoding process, which can be exploited at the decoder side in order to detect and correct errors. We propose a maximum likelihood and a maximum a posteriori context-based decoder, specifically tailored to the JPEG 2000 arithmetic coder, which are able to carry out both hard and soft decoding of a corrupted code-stream. The proposed decoder extends the JPEG 2000 capabilities in error-prone scenarios, without violating the standard syntax. Extensive simulations on video sequences show that the proposed decoders largely outperform the standard in terms of PSNR and visual quality.
Marco Grangetto, Enrico Magli, Gabriella Olmo
IEEE Trans. Image Process.1
2006 Multimedia Selective Encryption by Means of Randomized Arithmetic Coding
abstract
We propose a novel multimedia security framework based on a modification of the arithmetic coder, which is used by most international image and video coding standards as entropy coding stage. In particular, we introduce a randomized arithmetic coding paradigm, which achieves encryption by inserting some randomization in the arithmetic coding procedure; notably, and unlike previous works on encryption by arithmetic coding, this is done at no expense in terms of coding efficiency. The proposed technique can be applied to any multimedia coder employing arithmetic coding; in this paper we describe an implementation tailored to the JPEG 2000 standard. The proposed approach turns out to be robust towards attempts to estimating the image or discovering the key, and allows very flexible protection procedures at the code-block level, allowing to perform total and selective encryption, as well as conditional access
Marco Grangetto, Enrico Magli, Gabriella Olmo
IEEE Trans. Multim.1
2005 Improved low-complexity intraband lossless compression of hyperspectral images by means of Slepian-Wolf coding
abstract
In remote sensing systems, on-board data compression is a crucial task that has to be carried out with limited computational resources. In this paper we propose a novel lossless compression scheme for multispectral and hyperspectral images, which combines low encoding complexity and high-performance. The encoder is based on distributed source coding concepts, and employs Slepian-Wolf coding of the bitplanes of the CALIC prediction errors to achieve improved performance. Experimental results on AVIRIS data show that the proposed scheme exhibits performance similar to CALIC, and significantly better than JPEG 2000.
Antonello Nonnis, Marco Grangetto, Enrico Magli, Gabriella Olmo, Mauro Barni
ICIP (1)2
2005 Context-Based Distributed Wavelet Video Coding
abstract
In this paper a novel scalable video coder, based on the principle of distributed source coding with side information is proposed. Coding scalability is achieved by means of bitplane coding in the wavelet domain. The distributed coding paradigm is applied to encode the wavelet coefficients of a given frame, by considering those of the previous frame as side information. LDPC syndrome encoding with proper context modeling of the frame correlation allowed us to significantly outperform intra coding obtained with JPEG 2000. Moreover, the proposed approach permits to perform motion compensation at the decoder side, thus opening a new perspective in the field of scalable video coding
Marco Grangetto, Enrico Magli, Gabriella Olmo
MMSP1
2005 Joint source/channel coding and MAP decoding of arithmetic codes
abstract
In this paper, a novel maximum a posteriori (MAP) estimation approach is employed for error correction of arithmetic codes with a forbidden symbol. The system is founded on the principle of joint source channel coding, which allows one to unify the arithmetic decoding and error correction tasks into a single process, with superior performance compared to traditional separated techniques. The proposed system improves the performance in terms of error correction with respect to a separated source and channel coding approach based on convolutional codes, with the additional great advantage of allowing complete flexibility in adjusting the coding rate. The proposed MAP decoder is tested in the case of image transmission across the additive white Gaussian noise channel and compared against standard forward error correction techniques in terms of performance and complexity. Both hard and soft decoding are taken into account, and excellent results in terms of packet error rate and decoded image quality are obtained.
Marco Grangetto, Pamela C. Cosman, Gabriella Olmo
IEEE Trans. Commun.1
2005 Fast code-rate optimization for robust image transmission over lossy packet networks
abstract
In this paper, we propose an efficient method for the allocation of Reed-Solomon codes to source symbols, for unequal loss protection. The proposed formulation recasts the multivariate optimization problem into a univariate one, dramatically reducing the computational complexity. Results are shown for image transmission over lossy packet networks, employing the JPEG2000 and SPIHT encoders. The proposed algorithm exhibits performance equivalent to previous methods, while providing a significant complexity reduction.
Marco Grangetto, Enrico Magli, Gabriella Olmo
IEEE Trans. Commun.1
2005 Concealment of whole-frame losses for wireless low bit-rate video based on multiframe optical flow estimation
abstract
In low bit-rate packet-based video communications, video frames may have very small size, so that each frame fills the payload of a single network packet; thus, packet losses correspond to whole-frame losses, to which the existing error concealment algorithms are badly suited and generally not applicable. In this paper, we deal with the problem of concealment of whole frame-losses, and propose a novel technique which is capable of handling this very critical case. The proposed technique presents other two major innovations with respect to the state-of-the-art: i) it is based on optical flow estimation applied to error concealment and ii) it performs multiframe estimation, thus optimally exploiting the multiple reference frame buffer featured by the most modern video coders such as H.263+ and H.264. If data partitioning is employed, by e.g., sending headers, motion vectors, and coding modes in prioritized packets as can be done in the DiffServ network model, the algorithm is capable of exploiting the motion vectors to improve the error concealment results. The algorithm has been embedded in the H.264 test model software, and tested under both independent and correlated packet loss models with parameters typical of the wireless environment. Results show that the proposed algorithm significantly outperforms other techniques by several dBs in peak signal-to-noise ratio (PSNR), provides good visual quality, and has a rather low complexity, which makes it possible to perform real-time operation with reasonable computational resources.
Stefano Belfiore, Marco Grangetto, Enrico Magli, Gabriella Olmo
IEEE Trans. Multim.2
2004 Joint source-channel iterative decoding of codes
abstract
In this paper an innovative joint source channel coding scheme is presented. The system is based on iterative soft decoding of arithmetic codes, by means of a novel soft-in soft-out decoder based on suboptimal search and pruning of a binary tree. An error resilient arithmetic coder with a forbidden symbol is used in order to improve the performance of the joint source/channel scheme. The performance in the case of transmission across the AWGN channel is evaluated in terms of frame error rate, and compared to a traditional separated approach. Finally the convergence property of the system is analyzed by means of the EXIT chart technique.
Marco Grangetto, Bartolo Scanavino, Gabriella Olmo
ICC1
2004 Error resilient mq coder and map jpeg 2000 decoding
abstract
In this paper a novel error resilient MQ coder for reliable JPEG 2000 image delivery is designed. The proposed coder uses a forbidden symbol in order to force a given amount of redundancy in the codestream. At the decoder side, the presence of the forbidden symbol allows for powerful error correction. Moreover the added redundancy can be easily controlled and the proposed coder is kept backward compatible with MQ. In this work excellent improvements in the case of image transmission across both BSC and AWGN channels are obtained by means of a maximum a posteriori estimation technique.
Marco Grangetto, Enrico Magli, Gabriella Olmo
ICIP1
2004 Multiple description coding with error correction capabilities: an application to motion jpeg 2000
abstract
In the multiple description paradigm, a controllable amount of redundancy is inserted among descriptions, in order to help estimating those ones which are possibly lost due to network congestion. This redundancy can also be exploited in order to correct errors at bit level. In this paper, we propose a novel technique to generate multiple descriptions of video encoded with motion-JPEG 2000, which exploits the inserted extra redundancy also to guarantee error protection in case all descriptions are received, but are possibly affected by bit errors. This method yields excellent performance, since it guarantees not only protection of video information transmitted over non prioritized networks subject to independent packet erasure processes, but also resilience towards the corruption at bit level. Moreover, the generated streams are fully compatible with the part 3 of the JPEG 2000 standard.
Tammam Tillo, Marco Grangetto, Gabriella Olmo
ICIP2
2004 Reliable JPEG 2000 wireless imaging by means of error-correcting MQ coder
abstract
A new error resilience tool is proposed for robust JPEG 2000 imaging over noisy channels. In particular, a modified encoder, based on an MQ arithmetic coder with forbidden symbol, is introduced, along with a maximum likelihood error-correcting MQ decoder. The proposed technique features error detection, error concealment and error correction capability, thus adding new useful functionalities to JPEG 2000. Experimental results show that this technique largely outperforms the standard JPEG 2000 error resilience tools for error concealment and hard/soft channel decoding.
Marco Grangetto, Enrico Magli, Gabriella Olmo
ICME1
2004 Joint despeckling and edge detection of SAR images based on the Mumford-Shah functional
abstract
In this paper, we propose a joint despeckling and edge detection algorithm based on the Mumford-Shah functional, which accomplishes the image filtering and segmentation as a result of an analytical variational problem. This approach turns out to be well suited to jointly despeckle and segment SAR image data; the experimental results demonstrate that the proposed technique yields high quality despeckling without impairing critical image features and with the additional advantage to provide a detailed edge map
Stefano Belfiore, Riccardo Scopigno, Marco Grangetto, Enrico Magli
IGARSS3
2004 Selective encryption of JPEG 2000 images by means of randomized arithmetic coding
abstract
We describe a novel multimedia security framework based on a modification of the arithmetic coder, which is used by most international image and video coding standards as entropy coding stage. In particular, we propose a randomized arithmetic coding paradigm, which achieves encryption by randomly swapping the intervals of the least and most probable symbols in arithmetic coding; moreover, we describe an implementation tailored to the JPEG 2000 standard. The proposed approach turns out to be robust towards attempts to discover the key, and allows very flexible procedures for insertion of redundancy at the codeblock level, allowing to perform total and selective encryption, conditional access, and encryption of regions of interest.
Marco Grangetto, Alberto Grosso, Enrico Magli
MMSP1
2004 A flexible error resilient scheme for JPEG 2000
abstract
Nowadays, wireless multimedia applications are experiencing a rapid growth; in this scenario, challenging obstacles, such as packet losses due to congestion and band limitation along with bit-level error corruption, require the design of novel solutions for robust multimedia delivery. New standards for multimedia applications are incorporating many tools for error resilience; as an example, JPEG 2000 part 11 is explicitly devoted to the wireless applications of the image coding standard. In this paper we address a novel multiple description coding technique, based on post-processing rate allocation of embedded bitstreams. The proposed approach is compliant with the JPEG 2000 standard and has the ability to jointly cope with both packet losses and bit errors. Experimental results show that the designed algorithm significantly outperforms other techniques based on unequal error protection by means of RS codes.
Tammam Tillo, Marco Grangetto, Gabriella Olmo
MMSP2
2004 Ensuring quality of service for image transmission: hybrid loss protection
abstract
We present hybrid loss protection as a new channel coding and packetization scheme for image transmission over nonprioritized lossy packet networks. The scheme employs an interleaver-based structure, and attempts to maximize the expected peak signal-to-noise ratio (PSNR) at the receiver given the constraint that the probability of failure, i.e., the probability that the PSNR of the decoded image is below a given threshold, is upper-bounded by a user-defined value. A new code-allocation algorithm is proposed, which employs Gilbert-Elliot modeling of the network statistics. Experimental results are provided in the case of transmission of images encoded by SPIHT and JPEG 2000 over a wireline, as well as a wireless UMTS-based Internet connection.
Marco Grangetto, Enrico Magli, Gabriella Olmo
IEEE Trans. Image Process.1
2003 Spatio-temporal video error concealment with perceptually optimized mode selection
abstract
We propose a spatio-temporal error concealment algorithm for video transmission in an error-prone environment. The proposed technique employs motion vector estimation, edge-preserving interpolation, and texture analysis/synthesis. It has two main advantages with respect to existing methods, namely: (i) it aims at optimizing the visual quality of the restored video, and not only PSNR; and (ii) it employs an automatic mode selection algorithm in order to decide, on a macroblock basis, whether to use the spatial restoration, the temporal one, or a combination thereof. The algorithm has been applied to H.26L video, providing satisfactory performance over a large set of operating conditions.
Stefano Belfiore, Marco Grangetto, Enrico Magli, Gabriella Olmo
ICASSP (5)2
2003 Error correction by means of arithmetic codes: an application to resilient image transmission
abstract
In this paper, two novel maximum a posteriori (MAP) estimators for the decoding of arithmetic codes in the presence of transmission errors are presented. Trellis search techniques and a forbidden symbol are employed to obtain forward error correction. The proposed system is applied to lossless image compression and transmission across the BSC; the results are compared in terms of both performance and complexity with a traditional separated source and channel coding approach based on convolutional codes.
Marco Grangetto, Gabriella Olmo, Pamela C. Cosman
ICASSP (4)1
2003 Few decoders in the encoder: a low complexity encoding strategy for H.26L
abstract
We propose a reduced complexity technique for the rate-distortion optimization in JVT/H.26L in the presence of packet erasures. It is named "few decoders in the encoder", and is based on the idea of generating a selected number of error patterns in the encoder, so that a limited number of co-decoding processes can be implemented to estimate the transmission distortion term. The correlation amongst packet erasures is taken into account by employing a binary Gilbert model. The proposed algorithm exhibits competitive performance in terms of average PSNR and probability of decoding failure, with very affordable complexity and memory requirements.
Gabriella Olmo, Cristiano Cucco, Marco Grangetto, Enrico Magli
ICASSP (3)3
2003 An error concealment algorithm for streaming video
abstract
A known problem in video streaming is that loss of a packet usually results into loss of a whole video frame. In this paper we propose an error concealment algorithm specifically designed to handle this sort of losses. The technique exploits information in a few past frames (namely the motion vectors) in order to estimate the forward motion vectors of the last received frame. This information is used to project the last frame onto an estimate of the missing frame. The algorithm has been tested on MPEG-2 video, providing very satisfactory results, and outperforming by several dBs in PSNR the concealment technique based on repetition of the last received frame.
Stefano Belfiore, Marco Grangetto, Enrico Magli, Gabriella Olmo
ICIP (3)2
2003 Spatio-temporal video error concealment with perceptually optimized mode selection
abstract
We proposed a spatio-temporal error concealment algorithm for video transmission in an error-prone environment. The proposed technique employs motion vector estimation, edge-preserving interpolation, and texture analysis/synthesis. It has two main advantages with respect to existing methods, namely: i) it aims at optimizing the visual quality of the restored video, and not only PSNR; and ii) it employs an automatic mode selection algorithm in order to decide, on a macroblock basis, whether to use the spatial restoration, the temporal one, or a combination thereof. The algorithm has been applied to H26L video, providing satisfactory performance over a large set of operating conditions.
Stefano Belfiore, Marco Grangetto, Enrico Magli, Gabriella Olmo
ICME2
2003 Comparison of rate allocation strategies for H.264 video transmission over wireless lossy correlated networks
abstract
In this paper we study the problem of transmitting coded video over a UMTS network. We first discuss the statistical characteristics of packet losses in case of RTP/UDP/IP wireless video communication. Then, we propose a new rate allocation algorithm for H.264 video, based on a Gilbert-Elliot model of the packet losses. We compare the proposed algorithm with several other allocation strategies, showing that it achieves satisfactory performance in terms of PSNR. Moreover, we outline the limits of the employed distortion model and outline possible solutions to overcome them.
Stefano Gnavi, Marco Grangetto, Enrico Magli, Gabriella Olmo
ICME2
2003 Spatiotemporal error concealment with optimized mode selection and application to H.264
Stefano Belfiore, Marco Grangetto, Enrico Magli, Gabriella Olmo
Signal Process. Image Commun.2
2002 Robust and edge-preserving video error concealment by coarse-to-fine block replenishment
abstract
In this paper we propose a novel error concealment algorithm for video transmission over wireless networks potentially subject to packet erasures. In particular, we develop a technique for the replenishment of missing macroblocks, which aims at minimizing the impact of the lost data on the resulting video with respect to the human visual system. The proposed algorithm operates three reconstruction stages at different scales, by first recovering smooth large-scale patterns, then large-scale structures, and finally local edges in the lost macroblock. Experimental results show that the proposed algorithm achieves improved visual quality of the reconstructed frames with respect to other state-of-the-art techniques, as well as better PSNR results.
Stefano Belfiore, L. Crisa, Marco Grangetto, Enrico Magli, Gabriella Olmo
ICASSP3
2002 DSP performance comparison between lifting and filter banks for image coding
abstract
The lifting scheme is a very well-known computationally efficient alternative to the filter bank scheme for evaluating the discrete wavelet transform of signals and images. However, the actual computational saving is still a matter of debate. On one hand, theoretical results in the literature report an asymptotic upper-bound of two for very long wavelet filters. On the other hand, it is worth wondering to what extent the architecture of the processor used can actually bias this gain. In this paper we tackle this problem from an implementation perspective, and profile the execution time of the two algorithms on a digital signal processor. Both the real-valued and the integer versions of the wavelet transform are considered. The quantitative results are used to gain some insight on the way the processor architecture affects the algorithms.
Stefano Gnavi, Barbara Penna, Marco Grangetto, Enrico Magli, Gabriella Olmo
ICASSP3
2002 Guaranteeing quality of service for image transmission by means of hybrid loss protection
abstract
In the context of joint source and channel coding, unequal loss protection is often used to make image data more robust to possible packet losses. The allocation of source and code symbols is customarily done so as to maximize the expected PSNR at the receiver. We propose a new objective function, attempting to maximize PSNR given a constraint on the system probability of failure, so that PSNR is constrained to be above a given threshold with a given probability. This leads to the definition of a hybrid loss protection scheme, and the related allocation algorithm, which is able to satisfy this constraint. Experimental results are reported, related to the transmission of JPEG2000-compressed images over the Internet. It is shown that the proposed hybrid approach outperforms existing algorithms in terms of PSNR, while requiring less computational resources.
Marco Grangetto, Enrico Magli, Mauro Marzo, Gabriella Olmo
ICME (2)1
2002 Optimization and implementation of the integer wavelet transform for image coding
abstract
This paper deals with the design and implementation of an image transform coding algorithm based on the integer wavelet transform (IWT). First of all, criteria are proposed for the selection of optimal factorizations of the wavelet filter polyphase matrix to be employed within the lifting scheme. The obtained results lead to the IWT implementations with very satisfactory lossless and lossy compression performance. Then, the effects of finite precision representation of the lifting coefficients on the compression performance are analyzed, showing that, in most cases, a very small number of bits can be employed for the mantissa keeping the performance degradation very limited. Stemming from these results, a VLSI architecture is proposed for the IWT implementation, capable of achieving very high frame rates with moderate gate complexity.
Marco Grangetto, Enrico Magli, Maurizio Martina, Gabriella Olmo
IEEE Trans. Image Process.1
2001 Efficient common-core lossless and lossy image coder based on integer wavelets
Marco Grangetto, Enrico Magli, Gabriella Olmo
Signal Process.1
2000 Minimally non-linear integer wavelets for image coding
abstract
In this paper we deal with the problem of finding high performance factorizations of wavelet filters to be employed within the lifting scheme framework to yield the integer wavelet transform (IWT). A method is proposed, based on the search for the factorization yielding the minimally non-linear iterated graphic function. Results are reported, referring to a set of popular wavelet filters, which show that the obtained implementations lead to IWTs achieving very satisfactory results for both lossy and lossless image compression.
Marco Grangetto, Enrico Magli, Gabriella Olmo
ICASSP1
2000 Finite Precision Wavelets for Image Coding: Lossy and Lossless Compression Performance Evaluation
abstract
This paper investigates the robustness of the wavelet transform, implemented by means of the lifting scheme (LS), with respect to numerical errors in the representation and calculation of transformed coefficients. The study promises to offer important contributions for the understanding of the LS capabilities when specific implementations are considered. This is a topic of growing interest as the new standard JPEG 2000, based on the wavelet transform, is being finalized. Taking into account the effect of finite precision representation can drive both software and hardware implementations with optimized trade off between complexity and performance; moreover the robustness to numerical errors could be an important feature, not usually considered, in order to select the best wavelet filters.
Marco Grangetto, Enrico Magli, Gabriella Olmo
ICIP1