Attilio Fiandrotti

dblp:57/5224 · DBLP profile ↗
← Back
54ranked-venue papers
13as first author
26since 2021 · last 2026
0000-0002-9991-6822ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 10 first-author · 17 since 2021Artificial intelligence and machine learning · 15 · 1 first-author · 11 since 2021Computer networks · 4 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Towards a validation-less approach for small data: training with neural velocity
abstract
Tuning hyperparameters such as learning rate decay and stop conditions is typically done by assessing loss on a held-out validation set. This work introduces neural velocity (NeVe), the rate of change in neuron transfer functions, as a novel indicator of model convergence. We leverage NeVe to create a dynamic training method where learning rates and stop conditions are adjusted based on neural velocity computed from an auxiliary dataset generated by sampling noise. Our experiments across multiple tasks and architectures show that our approach performs on par with traditional validation-based methods, outperforming them in some tasks. Our method does not require withholding data for validation, making it especially advantageous in data-limited scenarios. This work highlights the potential of neural velocity as a key metric for optimizing neural network training.
Gianluca Dalmasso, Andrea Bragagnolo, Enzo Tartaglione, Attilio Fiandrotti, Marco Grangetto
Neurocomputing4
2026 TEP-ones: A simple yet effective approach for transferability estimation of pruned backbones
abstract
In deep learning, the conventional transfer learning paradigm involves fine-tuning a model pre-trained on a complex source task to adapt it to a simpler target task, capitalizing on abundant training data. Concurrently, the paradigm of neural network pruning has emerged as a powerful strategy for enhancing model efficiency, reducing complexity, and optimizing resource utilization. This paper focuses on pruned model transferability estimation for resource-constraint scenarios, where the goal is to rank the performance of pruned pre-trained models on a downstream task without fine-tuning. To this end, from a formal analysis of the intra-class mutual information between samples belonging to the same target class, we observe that, as pruning increases, a sweet phase naturally rises, where the model benefits from better features at the encoder’s output. From this, we derive a Transferability Estimation for Pruned Backbones (TEP-ones) that eases the choice of which pruned model (without the need to train the classifier) is the best candidate for transfer learning.
Gabriele Spadaro, Andrea Bragagnolo, Riccardo Renzulli, Marco Grangetto, Jhony-Heriberto Giraldo-Zuluaga, Attilio Fiandrotti, Enzo Tartaglione
Neurocomputing6
2026 Improving video codec quality with AI-based super-resolution and directional-mode enhancement
Alessandro Artusi, Mattia Angelini, Gabriele Spadaro, Attilio Fiandrotti, Giovanni Ballocca, Alessandra Mosca, Roberto Iacoviello, Leonardo Chiariglione
Multim. Tools Appl.4
2026 Unsupervised contrastive analysis for anomaly detection in brain MRIs via conditional diffusion models
abstract
Contrastive Analysis (CA) detects anomalies by contrasting patterns unique to a target group (e.g., unhealthy subjects) from those in a background group (e.g., healthy subjects). In the context of brain MRIs, existing CA approaches rely on supervised contrastive learning or variational autoencoders (VAEs) using both healthy and unhealthy data, but such reliance on target samples is challenging in clinical settings. Unsupervised Anomaly Detection (UAD) learns a reference representation of healthy anatomy, eliminating the need for target samples. Deviations from this reference distribution can indicate potential anomalies. In this context, diffusion models have been increasingly adopted in UAD due to their superior performance in image generation compared to VAEs. Nonetheless, precisely reconstructing the anatomy of the brain remains a challenge. In this work, we bridge CA and UAD by reformulating contrastive analysis principles for the unsupervised setting. We propose an unsupervised framework to improve the reconstruction quality by training a self-supervised contrastive encoder on healthy images to extract meaningful anatomical features. These features are used to condition a diffusion model to reconstruct the healthy appearance of a given image, enabling interpretable anomaly localization via pixel-wise comparison. We validate our approach through a proof-of-concept on a facial image dataset and further demonstrate its effectiveness on four brain MRI datasets, outperforming baseline methods in anomaly localization on the NOVA benchmark. • Unsupervised framework enhancing reconstruction quality in brain MRIs. • Target-invariant contrastive encoder capturing meaningful anatomical features. • Conditional diffusion model to reconstruct the healthy appearance of a given image. • Outperforms competing methods in anomaly localization on the NOVA benchmark.
Cristiano Patrício, Carlo Alberto Barbano, Attilio Fiandrotti, Riccardo Renzulli, Marco Grangetto, Luís F. Teixeira 0001, João C. Neves 0001
Pattern Recognit. Lett.3
2026 CALICE: Continuous Bitrate Control with Adapted LIC Model
abstract
Learned image compression (LIC) has drawn much attention recently as it outperforms standardized codecs in rate-distortion (RD) efficiency. However, an LIC model is typically trained for a specific RD tradeoff, and achieving a different target rate requires retraining the model and storing the weights as a whole, limiting the practical applicability of LIC. In this article, we introduce CALICE, a framework for achieving continuous bitrate control by plugging into a pre-trained LIC model a set of modular adapters. Unlike similar methods that require a distinct set of adapters for each target rate, our method achieves continuous bitrate control by modulating a single set of adapters via a scalar parameter \(\boldsymbol{\alpha}\) , with a total overhead of less than \(\mathbf{0.35}\boldsymbol{\%}\) of the parameters of the LIC model. This design enables efficient support for multiple distortion objectives by learning lightweight, distortion-aware adapters. We also extend our strategy beyond rate control, demonstrating its ability to provide fine-grained adaptation of perceptual quality along the distortion–perception tradeoff. To our knowledge, this is the first method that jointly addresses rate and perceptual control using a unified, low-cost strategy. We publicly released the code at https://github.com/EIDOSLAB/CALICE .
Gabriele Spadaro, Alberto Presta, Jhony-Heriberto Giraldo-Zuluaga, Attilio Fiandrotti, Marco Grangetto, Enzo Tartaglione
ACM Trans. Multim. Comput. Commun. Appl.4
2025 Denoising Diffusion Probabilistic Model for Point Cloud Compression at Low Bit-Rates
abstract
Efficient compression of low-bit-rate point clouds is critical for bandwidth-constrained applications. However, existing techniques mainly focus on high-fidelity reconstruction, requiring many bits for compression. This paper proposes a "Denoising Diffusion Probabilistic Model" (DDPM) architecture for point cloud compression (DDPM-PCC) at low bit-rates. A PointNet encoder produces the condition vector for the generation, which is then quantized via a learnable vector quantizer. This configuration allows to achieve a low bitrates while preserving quality. Experiments on ShapeNet and ModelNet40 show improved rate-distortion at low rates compared to standardized and state-of-the-art approaches. We publicly released the code at https://github.com/EIDOSLAB/DDPM-PCC.
Gabriele Spadaro, Alberto Presta, Jhony-Heriberto Giraldo-Zuluaga, Marco Grangetto, Giuseppe Valenzise, Attilio Fiandrotti, Enzo Tartaglione
ICME7
2025 Neural Velocity for hyperparameter tuning
abstract
Hyperparameter tuning, such as learning rate decay and defining a stopping criterion, often relies on monitoring the validation loss. This paper presents NeVe, a dynamic training approach that adjusts the learning rate and defines the stop criterion based on the novel notion of "neural velocity". The neural velocity measures the rate of change of each neuron’s transfer function and is an indicator of model convergence: sampling neural velocity can be performed even by forwarding noise in the network, reducing the need for a held-out dataset. Our findings show the potential of neural velocity as a key metric for optimizing neural network training efficiently.
Gianluca Dalmasso, Andrea Bragagnolo, Enzo Tartaglione, Attilio Fiandrotti, Marco Grangetto
IJCNN4
2025 Efficient Progressive Image Compression with Variance-Aware Masking
abstract
Learned progressive image compression is gaining momentum as it allows improved image reconstruction as more bits are decoded at the receiver. We propose a progressive image compression method in which an image is first represented as a pair of base-quality and top-quality latent representations. Next, a residual latent representation is encoded as the element-wise difference between the top and base representations. Our scheme enables progressive image compression with element-wise granularity by introducing a masking system that ranks each element of the residual latent representation from most to least important, dividing it into complementary components, which can be transmitted separately to the decoder in order to obtain different reconstruction quality. The masking system does not add further parameters or complexity. At the receiver, any elements of the top latent representation excluded from the transmitted components can be independently replaced with the mean predicted by the hyperprior architecture, ensuring reliable reconstructions at any intermediate quality level. We also in-troduced Rate Enhancement Modules (REMs), which refine the estimation of entropy parameters using already decoded components. We obtain results competitive with state-of-the-art competitors, while significantly reducing computational complexity, decoding time, and number of parameters.
Alberto Presta, Enzo Tartaglione, Attilio Fiandrotti, Marco Grangetto, Pamela C. Cosman
WACV3
2025 WiGNet: Windowed Vision Graph Neural Network
abstract
In recent years, Graph Neural Networks (GNNs) have demonstrated strong adaptability to various real-world challenges, with architectures such as Vision GNN (ViG) achieving state-of-the-art performance in several computer vision tasks. However, their practical applicability is hindered by the computational complexity of constructing the graph, which scales quadratically with the image size. In this paper, we introduce a novel Windowed vision Graph neural Network (WiGNet) model for efficient image processing. WiGNet explores a different strategy from previous works by partitioning the image into windows and constructing a graph within each window. Therefore, our model uses graph convolutions instead of the typical 2D convolution or self-attention mechanism. WiGNet effectively manages computational and memory complexity for large image sizes. We evaluate our method in the ImageNet-1k benchmark dataset and test the adaptability of WiGNet using the CelebA-HQ dataset as a downstream task with higher-resolution images. In both of these scenarios, our method achieves competitive results compared to previous vision GNNs while keeping memory and computational complexity at bay. WiGNet offers a promising solution toward the deployment of vision GNNs in real-world applications. We publicly released the code and pre-trained models at https://github.com/EIDOSLAB/WiGNet.
Gabriele Spadaro, Marco Grangetto, Attilio Fiandrotti, Enzo Tartaglione, Jhony-Heriberto Giraldo-Zuluaga
WACV3
2025 Robust and efficient airplane cockpit video coding leveraging temporal redundancy
abstract
Abstract Airplane cockpit screens consist of virtual instruments where characters, numbers, and graphics are overlaid on a black or natural background. Recording the cockpit screen allows one to log vital plane data, as aircraft manufacturers do not offer direct access to raw data. However, traditional video codecs struggle at preserving character readability at the required low bit-rates. We showed in a previous work that large rate-distortion gains can be achieved if the characters are encoded as text rather than as pixels. We now leverage temporal redundancy to both achieve robust character recognition and improve encoding efficiency. A convolutional neural network is trained for character classification over synthetic samples augmented with occlusions to gain robustness against overlapping graphics. Further robustness to background occlusions is brought by a probabilistic framework that error-corrects the output of the convolutional neural network. Next, we propose a predictive text coding technique specifically tailored for text in cockpit videos that achieves competitive performance over commodity lossless methods. Experiments with real cockpit video footage show large rate-distortion gains for the proposed method with respect to three different video compression standards. Notably, the H.264/AVC codec retrofitted with our method outperforms H.265/HEVC-SCC and is competitive with the much more complex H.266/VVC while preserving text and graphics. The entire pipeline described in this work has been implemented at Safran Electronics as an embedded avionics system drawing just 2W of power thanks to a combination of software and FPGA implementation.
Iulia Mitrica, Attilio Fiandrotti, Christophe Ruellan, Marco Cagnazzo
Multim. Tools Appl.2
2025 STanH: Parametric Quantization for Variable Rate Learned Image Compression
abstract
In end-to-end learned image compression, encoder and decoder are jointly trained to minimize a R + λD cost function, where λ controls the trade-off between rate of the quantized latent representation and image quality. Unfortunately, a distinct encoder-decoder pair with millions of parameters must be trained for each λ, hence the need to switch encoders and to store multiple encoders and decoders on the user device for every target rate. This paper proposes to exploit a differentiable quantizer designed around a parametric sum of hyperbolic tangents, called STanH, that relaxes the step-wise quantization function. STanH is implemented as a differentiable activation layer with learnable quantization parameters that can be plugged into a pre-trained fixed rate model and refined to achieve different target bitrates. Experimental results show that our method enables variable rate coding with comparable efficiency to the state-of-the-art, yet with significant savings in terms of ease of deployment, training time, and storage costs.
Alberto Presta, Enzo Tartaglione, Attilio Fiandrotti, Marco Grangetto
IEEE Trans. Image Process.3
2024 Find the Lady: Permutation and Re-synchronization of Deep Neural Networks
abstract
Deep neural networks are characterized by multiple symmetrical, equi-loss solutions that are redundant. Thus, the order of neurons in a layer and feature maps can be given arbitrary permutations, without affecting (or minimally affecting) their output. If we shuffle these neurons, or if we apply to them some perturbations (like fine-tuning) can we put them back in the original order i.e. re-synchronize? Is there a possible corruption threat? Answering these questions is important for applications like neural network white-box watermarking for ownership tracking and integrity verification. We advance a method to re-synchronize the order of permuted neurons. Our method is also effective if neurons are further altered by parameter pruning, quantization, and fine-tuning, showing robustness to integrity attacks. Additionally, we provide theoretical and practical evidence for the usual means to corrupt the integrity of the model, resulting in a solution to counter it. We test our approach on popular computer vision datasets and models, and we illustrate the threat and our countermeasure on a popular white-box watermarking method.
Carl De Sousa Trias, Mihai Mitrea, Attilio Fiandrotti, Marco Cagnazzo, Sumanta Chaudhuri, Enzo Tartaglione
AAAI3
2024 Domain Adaptation for Learned Image Compression with Supervised Adapters
abstract
In Learned Image Compression (LIC), a model is trained at encoding and decoding images sampled from a source domain, often outperforming traditional codecs on natural images; yet its performance may be far from optimal on images sampled from different domains. In this work, we tackle the problem of adapting a pre-trained model to multiple target domains by plugging into the decoder an adapter module for each of them, including the source one. Each adapter improves the decoder performance on a specific domain, without the model forgetting about the images seen at training time. A gate network computes the weights to optimally blend the contributions from the adapters when the bitstream is decoded. We experimentally validate our method over two state-of-the-art pre-trained models, observing improved rate-distortion efficiency on the target domains without penalties on the source domain. Furthermore, the gate’s ability to find similarities with the learned target domains enables better encoding efficiency also for images outside them.
Alberto Presta, Gabriele Spadaro, Enzo Tartaglione, Attilio Fiandrotti, Marco Grangetto
DCC4
2024 Gabic: Graph-Based Attention Block for Image Compression
abstract
While standardized codecs like JPEG and HEVC-intra represent the industry standard in image compression, neural Learned Image Compression (LIC) codecs represent a promising alternative. In detail, integrating attention mechanisms from Vision Transformers into LIC models has shown improved compression efficiency. However, extra efficiency often comes at the cost of aggregating redundant features. This work proposes a Graph-based Attention Block for Image Compression (GABIC), a method to reduce feature redundancy based on a k-Nearest Neighbors enhanced attention mechanism. Our experiments show that GABIC outperforms comparable methods, particularly at high bit rates, enhancing compression performance.
Gabriele Spadaro, Alberto Presta, Enzo Tartaglione, Jhony-Heriberto Giraldo-Zuluaga, Marco Grangetto, Attilio Fiandrotti
ICIP6
2024 WaterMAS: Sharpness-Aware Maximization for Neural Network Watermarking
Carl De Sousa Trias, Mihai Mitrea, Attilio Fiandrotti, Marco Cagnazzo, Sumanta Chaudhuri, Enzo Tartaglione
ICPR (5)3
2024 ALICE: Adapt your Learnable Image Compression modEl for variable bitrates
abstract
When training a Learned Image Compression model, the loss function is minimized such that the encoder and the decoder attain a target Rate-Distorsion trade-off. Therefore, a distinct model shall be trained and stored at the transmitter and receiver for each target rate, fostering the quest for efficient variable bitrate compression schemes. This paper proposes plugging Low-Rank Adapters into a transformer-based pre-trained LIC model and training them to meet different target rates. With our method, encoding an image at a variable rate is as simple as training the corresponding adapters and plugging them into the frozen pre-trained model. Our experiments show performance comparable with state-of-the-art fixed-rate LIC models at a fraction of the training and deployment cost. We publicly released the code at https://github.com/EIDOSLAB/ALICE.
Gabriele Spadaro, Muhammad Salman Ali, Alberto Presta, Giommaria Pilo, Sung-Ho Bae, Jhony-Heriberto Giraldo-Zuluaga, Attilio Fiandrotti, Marco Grangetto, Enzo Tartaglione
VCIP7
2024 A lightweight deep learning architecture for malaria parasite-type classification and life cycle stage detection
abstract
Abstract Malaria is an endemic in various tropical countries. The gold standard for disease detection is to examine the blood smears of patients by an expert medical professional to detect malaria parasite called Plasmodium. In the rural areas of underdeveloped countries, with limited infrastructure, a scarcity of healthcare professionals, an absence of sufficient computing devices, and a lack of widespread internet access, this task becomes more challenging. A severe case of malaria can be fatal within one week, so the correct detection of the malaria parasite and its life cycle stage is crucial in treating the disease correctly. Though computer vision-based malaria detection has been adequately explored lately, the malaria life cycle stage classification is still a relatively unexplored field. In this paper, we introduce a fast and robust deep learning methodology to not only classify the malaria parasite-type detection but also the life cycle stage identification of the infected cell. The proposed deep learning architecture is more than twenty times lighter than the widely used DenseNet and has less than 0.4 million parameters, making it a good candidate to be used in the mobile applications of such economically challenged states for malaria detection. We have used four different publicly available malaria datasets to test the proposed architecture and gained significantly better results than the current state of the art on malaria parasite-type and malaria life cycle classification.
Hafiza Ayesha Hoor Chaudhry, Muhammad Shahid Farid, Attilio Fiandrotti, Marco Grangetto
Neural Comput. Appl.3
2022 LOss-Based SensiTivity rEgulaRization: Towards deep sparse neural networks
Enzo Tartaglione, Andrea Bragagnolo, Attilio Fiandrotti, Marco Grangetto
Neural Networks3
2022 SeReNe: Sensitivity-Based Regularization of Neurons for Structured Sparsity in Neural Networks
abstract
Deep neural networks include millions of learnable parameters, making their deployment over resource-constrained devices problematic. Sensitivity-based regularization of neurons (SeReNe) is a method for learning sparse topologies with a structure, exploiting neural sensitivity as a regularizer. We define the sensitivity of a neuron as the variation of the network output with respect to the variation of the activity of the neuron. The lower the sensitivity of a neuron, the less the network output is perturbed if the neuron output changes. By including the neuron sensitivity in the cost function as a regularization term, we are able to prune neurons with low sensitivity. As entire neurons are pruned rather than single parameters, practical network footprint reduction becomes possible. Our experimental results on multiple network architectures and datasets yield competitive compression ratios with respect to state-of-the-art references.
Enzo Tartaglione, Andrea Bragagnolo, Francesco Odierna, Attilio Fiandrotti, Marco Grangetto
IEEE Trans. Neural Networks Learn. Syst.4
2022 Online Learning for Adaptive Video Streaming in Mobile Networks
abstract
In this paper, we propose a novel algorithm for video bitrate adaptation in HTTP Adaptive Streaming (HAS), based on online learning. The proposed algorithm, named Learn2Adapt (L2A) , is shown to provide a robust bitrate adaptation strategy which, unlike most of the state-of-the-art techniques, does not require parameter tuning, channel model assumptions, or application-specific adjustments. These properties make it very suitable for mobile users, who typically experience fast variations in channel characteristics. Experimental results, over real 4G traffic traces, show that L2A improves on the overall Quality of Experience (QoE) and in particular the average streaming bitrate, a result obtained independently of the channel and application scenarios.
Theodoros Karagkioules, Georgios S. Paschos, Nikolaos Liakopoulos, Attilio Fiandrotti, Dimitrios Tsilimantos, Marco Cagnazzo
ACM Trans. Multim. Comput. Commun. Appl.4
2021 Capsule Networks with Routing Annealing
Riccardo Renzulli, Enzo Tartaglione, Attilio Fiandrotti, Marco Grangetto
ICANN (1)3
2021 Unitopatho, A Labeled Histopathological Dataset for Colorectal Polyps Classification and Adenoma Dysplasia Grading
abstract
Histopathological characterization of colorectal polyps allows to tailor patients’ management and follow up with the ultimate aim of avoiding or promptly detecting an invasive carcinoma. Colorectal polyps characterization relies on the histological analysis of tissue samples to determine the polyps malignancy and dysplasia grade. Deep neural networks achieve outstanding accuracy in medical patterns recognition, however they require large sets of annotated training images. We introduce UniToPatho, an annotated dataset of 9536 hematoxylin and eosin (H&E) stained patches extracted from 292 whole-slide images, meant for training deep neural networks for colorectal polyps classification and adenomas grading. We present our dataset and provide insights on how to tackle the problem of automatic colorectal polyps characterization by suggesting a multi-resolution deep learning approach.
Carlo Alberto Barbano, Daniele Perlo, Enzo Tartaglione, Attilio Fiandrotti, Luca Bertero, Paola Cassoni, Marco Grangetto
ICIP4
2021 On the Role of Structured Pruning for Neural Network Compression
abstract
This works explores the benefits of structured parameter pruning in the framework of the MPEG standardization efforts for neural network compression. First less relevant parameters are pruned from the network, then remaining parameters are quantized and finally quantized parameters are entropy coded. We consider an unstructured pruning strategy that maximizes the number of pruned parameters at the price of randomly sparse tensors and a structured strategy that prunes fewer parameters yet yields regularly sparse tensors. We show that structured pruning enables better end-to-end compression despite lower pruning ratio because it boosts the efficiency of the arithmetic coder. As a bonus, once decompressed, the network memory footprint is lower as well as its inference time.
Andrea Bragagnolo, Enzo Tartaglione, Attilio Fiandrotti, Marco Grangetto
ICIP3
2021 HEMP: High-order entropy minimization for neural network compression
Enzo Tartaglione, Stéphane Lathuilière, Attilio Fiandrotti, Marco Cagnazzo, Marco Grangetto
Neurocomputing3
2021 Hybrid dual stream blender for wide baseline view synthesis
Nour Hobloss, Lu Zhang 0037, Stéphane Lathuilière, Marco Cagnazzo, Attilio Fiandrotti
Signal Process. Image Commun.5
2021 Learnable Descriptors for Visual Search
abstract
This work proposes LDVS, a learnable binary local descriptor devised for matching natural images within the MPEG CDVS framework. LDVS descriptors are learned so that they can be sign-quantized and compared using the Hamming distance. The underlying convolutional architecture enjoys a moderate parameters count for operations on mobile devices. Our experiments show that LDVS descriptors perform favorably over comparable learned binary descriptors at patch matching on two different datasets. A complete pair-wise image matching pipeline is then designed around LDVS descriptors, integrating them in the reference CDVS evaluation framework. Experiments show that LDVS descriptors outperform the compressed CDVS SIFT-like descriptors at pair-wise image matching over the challenging CDVS image dataset.
Andrea Migliorati, Attilio Fiandrotti, Gianluca Francini, Riccardo Leonardi
IEEE Trans. Image Process.2
2020 DR2S: Deep Regression with Region Selection for Camera Quality Evaluation
abstract
In this work, we tackle the problem of estimating a camera capability to preserve fine texture details at a given lighting condition. Importantly, our texture preservation measurement should coincide with human perception. Consequently, we formulate our problem as a regression one and we introduce a deep convolutional network to estimate texture quality score. At training time, we use ground-truth quality scores provided by expert human annotators in order to obtain a subjective quality measure. In addition, we propose a region selection method to identify the image regions that are better suited at measuring perceptual quality. Finally, our experimental evaluation shows that our learning-based approach outperforms existing methods and that our region selection algorithm consistently improves the quality estimation.
Marcelin Tworski, Stéphane Lathuilière, Salim Belkarfa, Attilio Fiandrotti, Marco Cagnazzo
ICPR4
2019 Enhancing HEVC Spatial Prediction by Context-based Learning
abstract
Deep generative models have been recently employed to compress images, image residuals or to predict image regions. Based on the observation that state-of-the-art spatial prediction is highly optimized from a rate-distortion point of view, in this work we study how learning-based approaches might be used to further enhance this prediction. To this end, we propose an encoder-decoder convolutional network able to reduce the energy of the residuals of HEVC intra prediction, by leveraging the available context of previously decoded neigh-boring blocks. The proposed context-based prediction enhancement (CBPE) scheme enables to reduce the mean square error of HEVC prediction by 25% on average, without any additional signalling cost in the bitstream.
Attilio Fiandrotti, Andrei I. Purica, Giuseppe Valenzise, Marco Cagnazzo
ICASSP2
2019 Robust license plate recognition using neural networks trained on synthetic images
Tomas Björklund, Attilio Fiandrotti, Mauro Annarumma, Gianluca Francini, Enrico Magli
Pattern Recognit.2
2019 Vehicle joint make and model recognition with multiscale attention windows
Sina Ghassemi, Attilio Fiandrotti, Emanuele Caimotti, Gianluca Francini, Enrico Magli
Signal Process. Image Commun.2
2019 Learning and Adapting Robust Features for Satellite Image Segmentation on Heterogeneous Data Sets
abstract
This paper addresses the problem of training a deep neural network for satellite image segmentation so that it can be deployed over images whose statistics differ from those used for training. For example, in postdisaster damage assessment, the tight time constraints make it impractical to train a network from scratch for each image to be segmented. We propose a convolutional encoder-decoder network able to learn visual representations of increasing semantic level as its depth increases, allowing it to generalize over a wider range of satellite images. Then, we propose two additional methods to improve the network performance over each specific image to be segmented. First, we observe that updating the batch normalization layers' statistics over the target image improves the network performance without human intervention. Second, we show that refining a trained network over a few samples of the image boosts the network performance with minimal human intervention. We evaluate our architecture over three data sets of satellite images, showing the state-of-the-art performance in binary segmentation of previously unseen images and competitive performance with respect to more complex techniques in a multiclass segmentation task.
Sina Ghassemi, Attilio Fiandrotti, Gianluca Francini, Enrico Magli
IEEE Trans. Geosci. Remote. Sens.2
2019 Securing Network Coding Architectures Against Pollution Attacks With Band Codes
abstract
During a pollution attack, malicious nodes purposely transmit bogus data to the honest nodes to cripple the communication. Securing the communication requires identifying and isolating the malicious nodes. However, in network coding (NC) architectures, random recombinations at the nodes increase the probability that honest nodes relay polluted packets. Thus, discriminating between honest and malicious nodes to isolate the latter turns out to be challenging at best. Band codes (BCs) are a family of rateless codes whose coding window size can be adjusted to reduce the probability that honest nodes relay polluted packets. We leverage such a property to design a distributed scheme for identifying the malicious nodes in the network. Each node counts the number of times that each neighbor has been involved in cases of polluted data reception and exchanges such counts with its neighbor nodes. Then, each node computes for each neighbor a discriminative honest score estimating the probability that the neighbor relays clean packets. We model such probability as a function of the BC coding window size, showing its impact on the accuracy and effectiveness of our distributed blacklisting scheme. We experiment distributing a live video feed in a P2P NC system, verifying the accuracy of our model and showing that our scheme allows us to secure the network against pollution attacks recovering near pre-attack video quality.
Attilio Fiandrotti, Rossano Gaeta, Marco Grangetto
IEEE Trans. Inf. Forensics Secur.1
2019 Very Low Bitrate Semantic Compression of Airplane Cockpit Screen Content
abstract
This paper addresses the problem of encoding the video generated by the screen of an airplane cockpit. As other computer screens, cockpit screens consist of computer-generated graphics often atop a natural background. Existing screen content coding schemes fail notably in preserving the readability of textual information at the low bitrates required in avionic applications. We propose a screen coding scheme where textual information is encoded according to the relative semantics rather than in the pixel domain. The encoder localizes textual information, and the semantics of each character are extracted with a convolutional neural network and predictively encoded. Text is then removed via inpainting, and the residual background video is compressed with a standard codec and transmitted to the receiver together with the text semantics. At the decoder side, text is synthesized using the decoded semantics and superimposed over the decoded residual video recovering the original frame. Our proposed scheme offers two key advantages over a semantics-unaware scheme that encodes text in the pixel domain. First, the text readability at the decoder is not compromised by compression artifacts, whereas the relative bitrate is negligible. Second, removal of high-frequency transform coefficients associated with the inpainted text drastically reduces the bitrate of the residual video. Experiments with real cockpit video sequences show BD-rate gains up to 82% and 69% over a reference H.265/HEVC encoder and its screen content coding extension. Moreover, our scheme achieves quasi-errorless character recognition already at very low bitrates, whereas even HEVC-SCC needs at least three or four times more bitrate to achieve a comparable error rate.
Iulia Mitrica, Eric Mercier, Christophe Ruellan, Attilio Fiandrotti, Marco Cagnazzo, Béatrice Pesquet-Popescu
IEEE Trans. Multim.4
2018 Feature Fusion for Robust Patch Matching with Compact Binary Descriptors
abstract
This work addresses the problem of learning compact yet discriminative patch descriptors within a deep learning framework. We observe that features extracted by convolutional layers in the pixel domain are largely complementary to features extracted in a transformed domain. We propose a convolutional network framework for learning binary patch descriptors where pixel domain features are fused with features extracted from the transformed domain. In our framework, while convolutional and transformed features are distinctly extracted, they are fused and provided to a single classifier which thus jointly operates on convolutional and transformed features. We experiment at matching patches from three different dataset, showing that our feature fusion approach outperforms multiple state-of-the-art approaches in terms of accuracy, rate and complexity.
Andrea Migliorati, Attilio Fiandrotti, Gianluca Francini, Skjalg Lepsøy, Riccardo Leonardi
MMSP2
2018 Learning sparse neural networks via sensitivity-driven regularization
abstract
The ever-increasing number of parameters in deep neural networks poses challenges for memory-limited applications. Regularize-and-prune methods aim at meeting these challenges by sparsifying the network weights. In this context we quantify the output sensitivity to the parameters (i.e. their relevance to the network output) and introduce a regularization term that gradually lowers the absolute value of parameters with low sensitivity. Thus, a very large fraction of the parameters approach zero and are eventually set to zero by simple thresholding. Our method surpasses most of the recent techniques both in terms of sparsity and error rates. In some cases, the method reaches twice the sparsity obtained by other techniques at equal error rates.
Enzo Tartaglione, Skjalg Lepsøy, Attilio Fiandrotti, Gianluca Francini
NeurIPS3
2018 CDVSec: Privacy-preserving biometrical user authentication in the cloud with CDVS descriptors
Attilio Fiandrotti, Massimo Mattelliano, Enrico Baccaglini, Paolo Vergori
Pattern Recognit. Lett.1
2017 Automatic license plate recognition with convolutional neural networks trained on synthetic data
abstract
We present an Automatic License Plate Recognition system designed around Convolutional Neural Networks (CNNs) and trained over synthetic plate images. We first design CNNs suitable for plate and character detection, sharing a common architecture and training procedure. Then, we generate synthetic images that account for the varying illumination and pose conditions encountered with real plate images and we use exclusively such synthetic images to train our CNNs. Experiments with real vehicle images captured in natural light with commodity imaging systems show precision and recall in excess of 93% despite our networks are trained exclusively on synthetic images.
Tomas Björklund, Attilio Fiandrotti, Mauro Annarumma, Gianluca Francini, Enrico Magli
MMSP2
2017 Fine-grained vehicle classificationusing deep residual networks with multiscale attention windows
abstract
Fine-grained vehicle classification is a challenging task due to the subtle differences between vehicle classes. Several successful approaches to fine-grained image classification rely on part-based models, where the image is classified according to discriminative object parts. Such approaches require however that parts in the training images be manually annotated, a labor-intensive process. We propose a convolutional architecture realizing a transform network capable of discovering the most discriminative parts of a vehicle at multiple scales. We experimentally show that our architecture outperforms a baseline reference if trained on class labels only, and performs closely to a reference based on a part-model if trained on loose vehicle localization bounding boxes.
Sina Ghassemi, Attilio Fiandrotti, Enrico Magli, Gianluca Francini
MMSP2
2016 Characterization of Band Codes for Pollution-Resilient Peer-to-Peer Video Streaming
abstract
We provide a comprehensive characterization of band codes (BC) as a resilient-by-design solution to pollution attacks in network coding (NC)-based peer-to-peer live video streaming. Consider one malicious node injecting bogus coded packets into the network: the recombinations at the nodes generate an avalanche of novel coded bogus packets. Therefore, the malicious node can cripple the communication by injecting into the network only a handful of polluted packets. Pollution attacks are typically addressed by identifying and isolating the malicious nodes from the network. Pollution detection is, however, not straightforward in NC as the nodes exchange coded packets. Similarly, malicious nodes identification is complicated by the ambiguity between malicious nodes and nodes that have involuntarily relayed polluted packets. This paper addresses pollution attacks through a radically different approach which relies on BCs. BCs are a family of rateless codes originally designed for controlling the NC decoding complexity in mobile applications. Here, we exploit BCs for the totally different purpose of recombining the packets at the nodes so to avoid that the pollution propagates by adaptively adjusting the coding parameters. Our streaming experiments show that BCs curb the propagation of the pollution and restore the quality of the distributed video stream.
Attilio Fiandrotti, Rossano Gaeta, Marco Grangetto
IEEE Trans. Multim.1
2015 Pollution-resilient peer-to-peer video streaming with Band Codes
abstract
Band Codes (BC) have been recently proposed as a solution for controlled-complexity random Network Coding (NC) in mobile applications, where energy consumption is a major concern. In this paper, we investigate the potential of BC in a peer-to-peer video streaming scenario where malicious and honest nodes coexists. Malicious nodes launch the so called pollution attack by randomly modifying the content of the coded packets they forward to downstream nodes, preventing honest nodes from correctly recovering the video stream. Whereas in much of the related literature this type of attack is addressed by identifying and isolating the malicious nodes, in this work we propose to address it by adaptively adjusting the coding scheme so to introduce resilience against pollution propagation. We experimentally show the impact of a pollution attack in a defenseless system and in a system where the coding parameters of BC are adaptively modulated following the discovery of polluted packets in the network. We observe that just by tuning the coding parameters, it is possible to reduce the impact of a pollution attack and restore the quality of the video communication.
Attilio Fiandrotti, Rossano Gaeta, Marco Grangetto
ICME1
2015 Simple Countermeasures to Mitigate the Effect of Pollution Attack in Network Coding-Based Peer-to-Peer Live Streaming
abstract
Network coding (NC)-based peer-to-peer (P2P) streaming represents an effective solution to aggregate user capacities and to increase system throughput in live multimedia streaming. Nonetheless, such systems are vulnerable to pollution attacks where a handful of malicious peers can disrupt the communication by transmitting just a few bogus packets which are then recombined and relayed by unaware honest nodes, further spreading the pollution over the network. Whereas previous research focused on malicious nodes identification schemes and pollution-resilient coding, in this paper we show pollution countermeasures which make a standard NC scheme resilient to pollution attacks. Thanks to a simple yet effective analytical model of a reference node collecting packets by malicious and honest neighbors, we demonstrate that: i) packets received earlier are less likely to be polluted, and ii) short generations increase the likelihood to recover a clean generation. Therefore, we propose a recombination scheme where nodes draw packets to be recombined according to their age in the input queue, paired with a decoding scheme able to detect the reception of polluted packets early in the decoding process and short generations. The effectiveness of our approach is experimentally evaluated in a real system we developed and deployed on hundreds to thousands of peers. Experimental evidence shows that, thanks to our simple countermeasures, the effect of a pollution attack is almost canceled and the video quality experienced by the peers is comparable to pre-attack levels.
Attilio Fiandrotti, Rossano Gaeta, Marco Grangetto
IEEE Trans. Multim.1
2014 Demonstrating the new compact descriptors for visual search (CDVS) standard for image retrieval on mobile devices
abstract
The MPEG CDVS (Compact Descriptors for Visual Search) standard promises to enable effective, bandwidth-efficient, image matching and retrieval. In this paper, we describe the issues related to the implementation of such technology in an app for smartphones and demonstrate its application to the problem of guiding a tourist through a urban photo safari. To the best of our knowledge, this work is also the first to provide preliminary figures of CDVS performance on a mobile device.
Giovanni Ballocca, Attilio Fiandrotti, Marco Gavelli, Massimo Mattelliano, Michele Morello, Alessandra Mosca, Paolo Vergori
ICIP2
2014 Band Codes for Energy-Efficient Network Coding With Application to P2P Mobile Streaming
abstract
A key problem in network coding (NC) lies in the complexity and energy consumption associated with the packet decoding processes, which hinder its application in mobile environments. Controlling and hence limiting such factors has always been an important but elusive research goal, since the packet degree distribution, which is the main factor driving the complexity, is altered in a non-deterministic way by the random recombinations at the network nodes. In this paper we tackle this problem with a new approach and propose Band Codes (BC), a novel class of network codes specifically designed to preserve the packet degree distribution during packet encoding, recombination and decoding. BC are random codes over GF(2) that exhibit low decoding complexity, feature limited and controlled degree distribution by construction, and hence allow to effectively apply NC even in energy-constrained scenarios. In particular, in this paper we motivate and describe our new design and provide a thorough analysis of its performance. We provide numerical simulations of the BC performance in order to validate the analysis and assess the overhead of BC with respect to a conventional random NC scheme. Moreover, experiment in a real-world application, namely peer-to-peer mobile media streaming using a random-push protocol, show that BC reduce the decoding complexity by a factor of two with negligible increase of the encoding overhead, paving the way for the application of NC to power-constrained devices.
Attilio Fiandrotti, Valerio Bioglio, Marco Grangetto, Rossano Gaeta, Enrico Magli
IEEE Trans. Multim.1
2014 Distributed Scheduling for Low-Delay and Loss-Resilient Media Streaming With Network Coding
abstract
Network coding (NC) has been shown to be very effective for collaborative media streaming applications. A pivotal issue in media streaming with NC lies in the packet scheduling policy at the network nodes, which affects the perceived media quality. In this paper, we address the problem of finding the packet scheduling policy that maximizes the number of media segments recovered in the network. We cast this as a distributed minimization problem and propose heuristic solutions that make the proposed framework robust to infrequent or inaccurate feedback information. Moreover, the proposed framework accounts for the properties of layered and multiple description encoded media to provide graceful quality degradation in case of packet losses or lack of upload bandwidth. Experimental results on a local testbed as well as PlanetLab suggest that our scheduling framework achieves better media quality, lower playback delay, and lower bandwidth consumption than a random-push scheme.
Anooq Muzaffar Sheikh, Attilio Fiandrotti, Enrico Magli
IEEE Trans. Multim.2
2013 Feedback-driven network coding for cooperative video streaming
abstract
In this work, we propose a feedback scheme to drive the packet recombination process at the network nodes in a collaborative Network Coding (NC) scenario. Our scheme addresses the issue of determining which symbols are more helpful at the receivers to recover the message and how to accordingly recombine the received packets at the intermediate nodes where the original symbols are not available. We experimentally demonstrate that our scheme increases the coding efficiency and reduces the computational complexity at the decoder in a video communication scenario without using explicit feedback messages.
Attilio Fiandrotti, Valerio Bioglio, Enrico Magli
ICASSP1
2013 Distributed media-aware scheduling for P2P streaming with Network Coding
abstract
We present a distributed packet scheduling scheme for pushbased Peer-to-Peer (P2P) video streaming with Network Coding (NC) over unstructured random overlays. While previous research has shown the potentials of random-push NC for P2P, little attention has been given to the problem of scheduling the packet transmissions at the network nodes. The proposed scheduling scheme exploits the knowledge of the status of the network links and nodes to maximize the number of nodes that are able to recover the media content prior to its playout deadline. Our experiments show a large performance gain with respect to random-push scheduler in terms of better media quality.
Anooq Muzaffar Sheikh, Attilio Fiandrotti, Enrico Magli
ICASSP2
2013 Distributed scheduling for scalable P2P video streaming with network coding
abstract
Previous research has shown the benefits of random-push Network Coding (NC) for P2P video streaming. On the other hand, scalable video coding provides graceful quality adaptation to heterogeneous network conditions. Nevertheless, packet scheduling for scalable media streaming with P2P NC is still a largely unexplored problem. Our ongoing research aims at designing a packet scheduling scheme that maximizes the quality of the video with minimal coordination among peers. In this work, we provide a preliminary description of our scheduling scheme and preliminary performance measurements.
Anooq Muzaffar Sheikh, Attilio Fiandrotti, Enrico Magli
INFOCOM2
2013 PISTA: Parallel Iterative Soft Thresholding algorithm for sparse image recovery
abstract
We present PISTA, a GPU-accelerated Iterative Soft Thresholding (IST) algorithm for sparse image recovery in Compressive Sensing applications. As the time required to recover an image increases with the number of pixels, GPU-acceleration enables to recover even large images in reasonable time. With respect to equivalent methods, IST-like algorithms have lower computational complexity per-iteration and lower memory requirements, plus the operations are inherently suitable for parallelization. Our experiments show that our algorithm enables a significant reduction in the time required to recover an image even over a highly-optimized CPU-only reference.
Attilio Fiandrotti, Sophie M. Fosson, Chiara Ravazzi, Enrico Magli
PCS1
2012 Band Codes: Controlled Complexity Network Coding for Peer-to-Peer Video Streaming
abstract
We present Band Codes (BC), a novel class of rate less codes that makes possible to control the computational complexity of Network Coding (NC). NC increases throughput of the networks via packet recombinations at the network nodes. In a NC scenario based on rate less codes, the recombinations at the nodes alter the packet degree distribution selected at the source and increase the computational complexity of the packet decoding process. Unlike other classes of rate less codes, BC preserve the degree distribution of the encoded packets through the recombinations at the nodes. Furthermore, BC enable to control the decoding complexity of each network node independently from the rest of the network. We evaluate BC in a P2P scenario using a purposely designed random-push protocol for live video streaming. The experiments show that BC achieve high encoding efficiency, enable nodes with different computational capabilities to coexist within the same network and reduce the processor load on a real mobile device by nearly 50%.
Attilio Fiandrotti, Valerio Bioglio, Enrico Magli, Marco Grangetto, Rossano Gaeta
ICME1
2011 Complexity-adaptive Random Network Coding for Peer-to-Peer video streaming
abstract
We present a novel architecture for complexity-adaptive Random Network Coding (RNC) and its application to Peer-to-Peer (P2P) video streaming. Network coding enables the design of simple and effective P2P video distribution systems, however it relies on computationally intensive packet coding operations that may exceed the computational capabilities of power constrained devices. It is hence desirable that the complexity of network coding can be adjusted at every node according to its computational capabilities, so that different classes of nodes can coexist in the network. To this end, we model the computational complexity of network coding as the sum of a packet decoding cost, which is centrally minimized at the encoder, and a packet recoding cost, which is locally controlled by each node. Efficient network coding is achieved exploiting the packet decoding process as a packet pre-recoding stage, hence increasing the chance that transmitted packets are innovative without increasing the recoding cost. Experiments in a P2P video streaming framework show that the proposed design enables the nodes of the network to operate at a wide range of computational complexity levels, while a higher number of low complexity nodes are able to join the network and experience high-quality video.
Attilio Fiandrotti, Simone Zezza, Enrico Magli
MMSP1
2010 Popularity-aware rate allocation in multiview video
abstract
We propose a framework for popularity-driven rate allocation in H.264/MVC-based multi-view video communications when the overall rate and the rate necessary for decoding each view are constrained in the delivery architecture. We formulate a rate allocation optimization problem that takes into account the popularity of each view among the client population and the rate-distortion characteristics of the multi-view sequence so that the performance of the system is maximized in terms of popularity-weighted average quality. We consider the cases where the global bit budget or the decoding rate of each view is constrained. We devise a simple ratevideo- quality model that accounts for the characteristics of interview prediction schemes typical of multi-view video. The video quality model is used for solving the rate allocation problem with the help of an interior point optimization method. We then show through experiments that the proposed rate allocation scheme clearly outperforms baseline solutions in terms of popularity-weighted video quality. In particular, we demonstrate that the joint knowledge of the rate-distortion characteristics of the video content, its coding dependencies, and the popularity factor of each view is key in achieving good coding performance in multi-view video systems.
Attilio Fiandrotti, Jacob Chakareski, Pascal Frossard
VCIP1
2010 Content-adaptive traffic prioritization of spatio-temporal scalable video for robust communications over QoS-provisioned 802.11e networks
Attilio Fiandrotti, Dario Gallucci, Enrico Masala, Juan Carlos De Martin
Signal Process. Image Commun.1
2009 Content-Adaptive Robust H.264/SVC Video Communications over 802.11e Networks
abstract
In this paper we present a low-complexity traffic prioritization strategy for video transmission using the H.264 scalable video coding (SVC) standard over 802.11e wireless networks.The first part of this work focuses on assessing the perceptual impact of data loss in the various enhancement layers using a wide set of H.264/SVC encoded videos.The analysis shows that perceptual impairments are highly correlated with the motion activity in the video sequence.Thus, we propose an adaptive unequal error protection strategy which identifies the most perceptually important parts of the enhancement layers in the video sequence by means of a low complexity macroblock motion analysis process.The algorithm is tested by simulating a realistic 802.11e based home network scenario.Results obtained on a large set of video sequences show that the proposed content-aware traffic prioritization strategy enables PSNR gains up to 2.5 dB as well as noticeable visual quality improvements with respect to a traditional prioritization strategy aiming at minimizing error propagation.
Dario Gallucci, Attilio Fiandrotti, Enrico Masala, Juan Carlos De Martin
AINA2
2008 Traffic Prioritization of H.264/SVC Video over 802.11e Ad Hoc Wireless Networks
abstract
The H.264/SVC video codec extends the H.264/AVC standard with scalability features. In this paper we introduce a traffic prioritization algorithm suitable for the transmission of both H.264/SVC and H.264/AVC video over 802.11e ad hoc wireless networks. The proposed algorithm exploits the traffic prioritization capabilities offered by 802.11e to provide better protection to the most perceptually important parts of a video while achieving efficient network resource usage. We evaluate the algorithm by simulating video transmissions in an ad hoc network scenario. Results show that the H.264/SVC codec particularly benefits from the proposed algorithm, which enables a graceful video quality degradation in congested network conditions, as well as PSNR gains up to 2 dB with respect to the H.264/AVC codec using the same amount of network resources.
Attilio Fiandrotti, Dario Gallucci, Enrico Masala, Enrico Magli
ICCCN1