Gabriele Spadaro

dblp:356/2111 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
9since 2021 · last 2026
0009-0008-4786-1074ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TEP-ones: A simple yet effective approach for transferability estimation of pruned backbones
abstract
In deep learning, the conventional transfer learning paradigm involves fine-tuning a model pre-trained on a complex source task to adapt it to a simpler target task, capitalizing on abundant training data. Concurrently, the paradigm of neural network pruning has emerged as a powerful strategy for enhancing model efficiency, reducing complexity, and optimizing resource utilization. This paper focuses on pruned model transferability estimation for resource-constraint scenarios, where the goal is to rank the performance of pruned pre-trained models on a downstream task without fine-tuning. To this end, from a formal analysis of the intra-class mutual information between samples belonging to the same target class, we observe that, as pruning increases, a sweet phase naturally rises, where the model benefits from better features at the encoder’s output. From this, we derive a Transferability Estimation for Pruned Backbones (TEP-ones) that eases the choice of which pruned model (without the need to train the classifier) is the best candidate for transfer learning.
Gabriele Spadaro, Andrea Bragagnolo, Riccardo Renzulli, Marco Grangetto, Jhony-Heriberto Giraldo-Zuluaga, Attilio Fiandrotti, Enzo Tartaglione
Neurocomputing1
2026 Improving video codec quality with AI-based super-resolution and directional-mode enhancement
Alessandro Artusi, Mattia Angelini, Gabriele Spadaro, Attilio Fiandrotti, Giovanni Ballocca, Alessandra Mosca, Roberto Iacoviello, Leonardo Chiariglione
Multim. Tools Appl.3
2026 CALICE: Continuous Bitrate Control with Adapted LIC Model
abstract
Learned image compression (LIC) has drawn much attention recently as it outperforms standardized codecs in rate-distortion (RD) efficiency. However, an LIC model is typically trained for a specific RD tradeoff, and achieving a different target rate requires retraining the model and storing the weights as a whole, limiting the practical applicability of LIC. In this article, we introduce CALICE, a framework for achieving continuous bitrate control by plugging into a pre-trained LIC model a set of modular adapters. Unlike similar methods that require a distinct set of adapters for each target rate, our method achieves continuous bitrate control by modulating a single set of adapters via a scalar parameter \(\boldsymbol{\alpha}\) , with a total overhead of less than \(\mathbf{0.35}\boldsymbol{\%}\) of the parameters of the LIC model. This design enables efficient support for multiple distortion objectives by learning lightweight, distortion-aware adapters. We also extend our strategy beyond rate control, demonstrating its ability to provide fine-grained adaptation of perceptual quality along the distortion–perception tradeoff. To our knowledge, this is the first method that jointly addresses rate and perceptual control using a unified, low-cost strategy. We publicly released the code at https://github.com/EIDOSLAB/CALICE .
Gabriele Spadaro, Alberto Presta, Jhony-Heriberto Giraldo-Zuluaga, Attilio Fiandrotti, Marco Grangetto, Enzo Tartaglione
ACM Trans. Multim. Comput. Commun. Appl.1
2025 FOLDER: Accelerating Multi-Modal Large Language Models with Enhanced Performance
abstract
Recently, Multi-modal Large Language Models (MLLMs) have shown remarkable effectiveness for multi-modal tasks due to their abilities to generate and understand cross-modal data. However, processing long sequences of visual tokens extracted from visual backbones poses a challenge for deployment in real-time applications. To address this issue, we introduce FOLDER, a simple yet effective plug-and-play module designed to reduce the length of the visual token sequence, mitigating both computational and memory demands during training and inference. Through a comprehensive analysis of the token reduction process, we analyze the information loss introduced by different reduction strategies and develop FOLDER to preserve key information while removing visual redundancy. We showcase the effectiveness of FOLDER by integrating it into the visual backbone of several MLLMs, significantly accelerating the inference phase. Furthermore, we evaluate its utility as a training accelerator or even performance booster for MLLMs. In both contexts, FOLDER achieves comparable or even better performance than the original models, while dramatically reducing complexity by removing up to 70% of visual tokens.
Haicheng Wang, Zhemeng Yu, Gabriele Spadaro, Chen Ju, Victor Quétu, Shuai Xiao 0002, Enzo Tartaglione
ICCV3
2025 Denoising Diffusion Probabilistic Model for Point Cloud Compression at Low Bit-Rates
abstract
Efficient compression of low-bit-rate point clouds is critical for bandwidth-constrained applications. However, existing techniques mainly focus on high-fidelity reconstruction, requiring many bits for compression. This paper proposes a "Denoising Diffusion Probabilistic Model" (DDPM) architecture for point cloud compression (DDPM-PCC) at low bit-rates. A PointNet encoder produces the condition vector for the generation, which is then quantized via a learnable vector quantizer. This configuration allows to achieve a low bitrates while preserving quality. Experiments on ShapeNet and ModelNet40 show improved rate-distortion at low rates compared to standardized and state-of-the-art approaches. We publicly released the code at https://github.com/EIDOSLAB/DDPM-PCC.
Gabriele Spadaro, Alberto Presta, Jhony-Heriberto Giraldo-Zuluaga, Marco Grangetto, Giuseppe Valenzise, Attilio Fiandrotti, Enzo Tartaglione
ICME1
2025 WiGNet: Windowed Vision Graph Neural Network
abstract
In recent years, Graph Neural Networks (GNNs) have demonstrated strong adaptability to various real-world challenges, with architectures such as Vision GNN (ViG) achieving state-of-the-art performance in several computer vision tasks. However, their practical applicability is hindered by the computational complexity of constructing the graph, which scales quadratically with the image size. In this paper, we introduce a novel Windowed vision Graph neural Network (WiGNet) model for efficient image processing. WiGNet explores a different strategy from previous works by partitioning the image into windows and constructing a graph within each window. Therefore, our model uses graph convolutions instead of the typical 2D convolution or self-attention mechanism. WiGNet effectively manages computational and memory complexity for large image sizes. We evaluate our method in the ImageNet-1k benchmark dataset and test the adaptability of WiGNet using the CelebA-HQ dataset as a downstream task with higher-resolution images. In both of these scenarios, our method achieves competitive results compared to previous vision GNNs while keeping memory and computational complexity at bay. WiGNet offers a promising solution toward the deployment of vision GNNs in real-world applications. We publicly released the code and pre-trained models at https://github.com/EIDOSLAB/WiGNet.
Gabriele Spadaro, Marco Grangetto, Attilio Fiandrotti, Enzo Tartaglione, Jhony-Heriberto Giraldo-Zuluaga
WACV1
2024 Domain Adaptation for Learned Image Compression with Supervised Adapters
abstract
In Learned Image Compression (LIC), a model is trained at encoding and decoding images sampled from a source domain, often outperforming traditional codecs on natural images; yet its performance may be far from optimal on images sampled from different domains. In this work, we tackle the problem of adapting a pre-trained model to multiple target domains by plugging into the decoder an adapter module for each of them, including the source one. Each adapter improves the decoder performance on a specific domain, without the model forgetting about the images seen at training time. A gate network computes the weights to optimally blend the contributions from the adapters when the bitstream is decoded. We experimentally validate our method over two state-of-the-art pre-trained models, observing improved rate-distortion efficiency on the target domains without penalties on the source domain. Furthermore, the gate’s ability to find similarities with the learned target domains enables better encoding efficiency also for images outside them.
Alberto Presta, Gabriele Spadaro, Enzo Tartaglione, Attilio Fiandrotti, Marco Grangetto
DCC2
2024 Gabic: Graph-Based Attention Block for Image Compression
abstract
While standardized codecs like JPEG and HEVC-intra represent the industry standard in image compression, neural Learned Image Compression (LIC) codecs represent a promising alternative. In detail, integrating attention mechanisms from Vision Transformers into LIC models has shown improved compression efficiency. However, extra efficiency often comes at the cost of aggregating redundant features. This work proposes a Graph-based Attention Block for Image Compression (GABIC), a method to reduce feature redundancy based on a k-Nearest Neighbors enhanced attention mechanism. Our experiments show that GABIC outperforms comparable methods, particularly at high bit rates, enhancing compression performance.
Gabriele Spadaro, Alberto Presta, Enzo Tartaglione, Jhony-Heriberto Giraldo-Zuluaga, Marco Grangetto, Attilio Fiandrotti
ICIP1
2024 ALICE: Adapt your Learnable Image Compression modEl for variable bitrates
abstract
When training a Learned Image Compression model, the loss function is minimized such that the encoder and the decoder attain a target Rate-Distorsion trade-off. Therefore, a distinct model shall be trained and stored at the transmitter and receiver for each target rate, fostering the quest for efficient variable bitrate compression schemes. This paper proposes plugging Low-Rank Adapters into a transformer-based pre-trained LIC model and training them to meet different target rates. With our method, encoding an image at a variable rate is as simple as training the corresponding adapters and plugging them into the frozen pre-trained model. Our experiments show performance comparable with state-of-the-art fixed-rate LIC models at a fraction of the training and deployment cost. We publicly released the code at https://github.com/EIDOSLAB/ALICE.
Gabriele Spadaro, Muhammad Salman Ali, Alberto Presta, Giommaria Pilo, Sung-Ho Bae, Jhony-Heriberto Giraldo-Zuluaga, Attilio Fiandrotti, Marco Grangetto, Enzo Tartaglione
VCIP1