VLDB 2026 Research / reviewers in the wild / expert
Alessandro Gnutti
dblp:171/7714
· DBLP profile ↗
21ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0002-8308-0776ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 9 first-author · 14 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CSGaussian: Progressive Rate-Distortion Compression and Segmentation for 3D Gaussian SplattingabstractWe present the first unified framework for rate-distortion-optimized compression and segmentation of 3D Gaussian Splatting (3DGS). While 3DGS has proven effective for both real-time rendering and semantic scene understanding, prior works have largely treated these tasks independently, leaving their joint consideration unexplored. Inspired by recent advances in rate-distortion-optimized 3DGS compression, this work integrates semantic learning into the compression pipeline to support decoder-side applications–such as scene editing and manipulation–that extend beyond traditional scene reconstruction and view synthesis. Our scheme features a lightweight implicit neural representation-based hyperprior, enabling efficient entropy coding of both color and semantic attributes while avoiding costly grid-based hyperprior as seen in many prior works. To facilitate compression and segmentation, we further develop compression-guided segmentation learning, consisting of quantization-aware training to enhance feature separability and a quality-aware weighting mechanism to suppress unreliable Gaussian primitives. Extensive experiments on the LERF and 3D-OVS datasets demonstrate that our approach significantly reduces transmission cost while preserving high rendering quality and strong segmentation performance. Yu-Jen Tseng, Chia-Hao Kao, Jing-Zhong Chen, Alessandro Gnutti, Shao-Yuan Lo, Yen-Yu Lin, Wen-Hsiao Peng |
WACV | 4 |
| 2026 | Towards compression-aware iris presentation attack detectionabstractWith the growing integration of biometric recognition systems into high-security and large-scale deployment scenarios, it is becoming increasingly important to ensure their robustness under realistic operational constraints. This implies designing solutions able to withstand potential adversarial threats that could affect their integrity and accountability, and also taking into account requirements of real-world operating systems such as limited availability of bandwidth and memory for data transmission and storage. Hence, the proposed study deals with presentation attack detection (PAD) for iris recognition, evaluating the effectiveness of Transformer-based frameworks at detecting spoofing attacks relying on fabricated or artificial biometric evidences. More specifically, we focus on the effects of image compression on the quality of iris images and on the resulting PAD performance, considering both traditional techniques such as JPEG as well as next-generation learning-based image codecs such as JPEG AI. We then examine the feasibility of mitigating compression-induced performance degradation by fine-tuning the adopted models on compressed images, achieving improvements in terms of half total error rate between 5% and 10% for images compressed at the worst JPEG and JPEG AI qualities. We also evaluate the generalizability of the developed solutions by testing them on learning-based codecs not considered during training, to check whether similar PAD-relevant artifacts are introduced by different compressions. Furthermore, we investigate the redundancy within the embeddings generated by the employed detectors, and demonstrated it is possible to significantly compress them while preserving the achievable PAD performance. Overall, the study provides a systematic analysis of iris PAD under compression constraints, offering insights into model adaptation, cross-codec robustness, and representation efficiency in scenarios where visual data coding plays a central role. Rocco Albano, Filippo Battaglia, Alessandro Gnutti, Emanuele Maiorana, Fabrizio Guerrini, Giuseppe Campobello, Pierangelo Migliorati, Patrizio Campisi |
Signal Process. Image Commun. | 3 |
| 2026 | JPEG AI Compressed Domain Face Detection: A Multi-Scale Bridging PerspectiveabstractLearning-based image coding is showing improved compression efficiency, while also offering a novel advantage in enabling computer vision tasks directly within the compressed domain. The latent representation created by deep learning methods inherently contains all visual features, without a computationally expensive synthesis process at the decoder. This paper is an invited extension of a previous solution for JPEG AI compressed domain face detection that adapts a RetinaFace-based detector to operate directly on the latent tensor. In addition to a former single-scale bridging solution, this work provides a novel multi-scale bridging architecture to enable a more effective multi-scale compressed domain face detection. The results show a significant performance gain, improving accuracy up to 20% for detection of tiny faces on the WIDER FACE dataset compared to single-scale bridging, and further narrowing the gap when compared to detection on uncompressed or JPEG AI decoded images. Furthermore, since the computationally expensive decoding step is bypassed and since the bridges consist of lower-complexity networks, the overall processing cost is significantly reduced. Single and multi-scale bridging, respectively, have about 10% and 32% the complexity of applying pixel domain face detection on decoded images. The proposed architecture is expected to be extended to other multiscale sensitive vision tasks, as JPEG AI is not specifically designed for any single downstream application. Ayman Alkhateeb, Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi, João Ascenso, Fernando Pereira 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | Learning Optimal Linear Block Transform by Rate Distortion MinimizationabstractThe rise of deep learning has spurred advancements in image compression, with end-to-end learned systems gaining traction. However, their adoption in standard frameworks is limited, as they require a major overhaul of existing hardware designed for traditional methods. Moreover, their computational complexity, especially on the decoder side, remains significantly higher than conventional codecs. Consequently, optimizing traditional codecs remains a key research focus. Alessandro Gnutti, Chia-Hao Kao, Wen-Hsiao Peng, Riccardo Leonardi |
DCC | 1 |
| 2025 | MH-LVC: Multi-Hypothesis Temporal Prediction for Learned Conditional Residual Video Coding
Huu-Tai Phung, Zong-Lin Gao, Yi-Chen Yao, Kuan-Wei Ho, Yi-Hsin Chen, Yu-Hsiang Lin, Alessandro Gnutti, Wen-Hsiao Peng |
ICCV | 7 |
| 2025 | Bridging Compressed Image Latents and Multimodal Large Language ModelsabstractThis paper presents the first-ever study of adapting compressed image latents to suit the needs of downstream vision tasks that adopt Multimodal Large Language Models (MLLMs). MLLMs have extended the success of large language models to modalities (e.g. images) beyond text, but their billion scale hinders deployment on resource-constrained end devices. While cloud-hosted MLLMs could be available, transmitting raw, uncompressed images captured by end devices to the cloud requires an efficient image compression system. To address this, we focus on emerging neural image compression and propose a novel framework with a lightweight transform-neck and a surrogate loss to adapt compressed image latents for MLLM-based vision tasks.
Given the huge scale of MLLMs, our framework excludes the entire downstream MLLM except part of its visual encoder from training our system. This stands out from most existing coding for machine approaches that involve downstream networks in training and thus could be impractical when the networks are MLLMs. The proposed framework is general in that it is applicable to various MLLMs, neural image codecs, and multiple application scenarios, where the neural image codec can be (1) pre-trained for human perception without updating, (2) fully updated for joint human and machine perception, or (3) fully updated for only machine perception.
Extensive experiments on different neural image codecs and various MLLMs show that our method achieves great rate-accuracy performance with much less complexity. Chia-Hao Kao, Cheng Chien, Yu-Jen Tseng, Yi-Hsin Chen, Alessandro Gnutti, Shao-Yuan Lo, Wen-Hsiao Peng, Riccardo Leonardi |
ICLR | 5 |
| 2025 | Rate Distortion Learned Transform For Image Compression
Alessandro Gnutti, Chia-Hao Kao, Wen-Hsiao Peng, Riccardo Leonardi |
PCS | 1 |
| 2025 | Learned Transcoding for Neural Image Compression
Chia-Hao Kao, Andrea Migliorati, Alessandro Gnutti, Enrico Magli, Riccardo Leonardi |
PCS | 3 |
| 2025 | Variable-Size Symmetry-Based Graph Fourier Transforms for Image Compression
Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi, Antonio Ortega |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Lidar Depth Map Guided Image Compression ModelabstractThe incorporation of LiDAR technology into some high-end smartphones has unlocked numerous possibilities across various applications, including photography, image restoration, augmented reality, and more. In this paper, we introduce a novel direction that harnesses LiDAR depth maps to enhance the compression of the corresponding RGB camera images. To the best of our knowledge, this represents the initial exploration in this particular research direction. Specifically, we propose a Transformer-based learned image compression system capable of achieving variable-rate compression using a single model while utilizing the LiDAR depth map as supplementary information for both the encoding and decoding processes. Experimental results demonstrate that integrating LiDAR yields an average PSNR gain of 0.83 dB and an average bitrate reduction of 16% as compared to its absence. Alessandro Gnutti, Stefano Della Fiore, Mattia Savardi, Yi-Hsin Chen, Riccardo Leonardi, Wen-Hsiao Peng |
ICIP | 1 |
| 2024 | JPEG AI Compressed Domain Face DetectionabstractLearning-based image coding has achieved competitive performance in terms of compression efficiency, while also gaining a key advantage in the ability to carry out computer vision tasks directly in the compressed domain. In fact, the latent representation which is generated using deep learning techniques may natively encapsulate all visual features needed for processing tasks, thereby eliminating the need to perform the expensive synthesis transform process at the decoder side. In this paper, it is proposed to perform face detection using the latent code present in the JPEG AI architecture. First, some experiments show how decoded images can be efficiently processed for face detection without retraining, albeit with some performance degradation. Then, for the first time a compressed domain RetinaFace-based detector applied to JPEG AI latent representations is competitively proposed. The performance achieved is comparable to the performance of the original RetinaFace applied to the reconstructed JPEG AI images, while reducing computational complexity since it bypasses the image decoding process. It is expected that this approach might be extended to other vision tasks since the JPEG AI representation format is not tailored specifically for any computer vision task. Ayman Alkhateeb, Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi, João Ascenso, Fernando Pereira 0001 |
MMSP | 2 |
| 2024 | Transformer-Based Learned Image Compression for Joint Decoding and DenoisingabstractThis work introduces a Transformer-based image compression system. It has the flexibility to switch between the standard image reconstruction and the denoising reconstruction from a single compressed bitstream. Instead of training separate decoders for these tasks, we incorporate two add-on modules to adapt a pre-trained image decoder from performing the standard image reconstruction to joint decoding and denoising. Our scheme adopts a two-pronged approach. It features a latent refinement module to refine the latent representation of a noisy input image for reconstructing a noise-free image. Additionally, it incorporates an instance-specific prompt generator that adapts the decoding process to improve on the latent refinement. Experimental results show that our method achieves a similar level of denoising quality to training a separate decoder for joint decoding and denoising at the expense of only a modest increase in the decoder's model size and computational complexity. Yi-Hsin Chen, Kuan-Wei Ho, Shiau-Rung Tsai, Guan-Hsun Lin, Alessandro Gnutti, Wen-Hsiao Peng, Riccardo Leonardi |
PCS | 5 |
| 2022 | CANF-VC: Conditional Augmented Normalizing Flows for Video Compression
Yung-Han Ho, Chih-Peng Chang, Alessandro Gnutti, Wen-Hsiao Peng |
ECCV (16) | 4 |
| 2022 | Machine learning techniques for MRI feature-based detection of frontotemporal lobar degenerationabstractMaking a diagnosis of neurodegenerative diseases at an early stage is one of the most significant challenges of modern neuroscience. Although this family of diseases remains without a cure, the effectiveness of their medical treatment largely relies on the timing of their detection. For certain groups of diseases, such as Fronto-Temporal Dementia (FTD), trained professionals can effectively reach a correct diagnosis through the visual analysis of Magnetic Resonance Imaging, in its functional (fMRI) or raw (MRI) version. However, this operation is time-consuming and may be subject to personal interpretation. In this paper, we explore the performance of a group of machine learning algorithms to formulate a correct FTD diagnosis, in order to provide medical professionals with a supporting tool. The dataset consists of MRI data acquired on 30 subjects, and the experiments are carried out by investigating different fMRI techniques based on a Multi-Voxel Pattern Analysis (MVPA) approach. The results obtained show high accuracy in identifying FTD in elderly patients when Support Vector Machine and Random Forest techniques are used, with outcomes varying based on the fMRI methods. Tatiana Pilipenko, Alessandro Gnutti, Andrea Silvestri, Ivan Serina, Riccardo Leonardi |
KES | 2 |
| 2021 | Symmetry-Based Graph Fourier Transforms: Are They Optimal For Image Compression?abstractTraditional block-based transforms are based on applying a single transform to all blocks. As an alternative, better performance in image and video processing and representation can be achieved by choosing one among a discrete set of transforms for each block. As an example, our recently proposed set of multiple transforms called Symmetry-Based Graph Fourier Transforms (SBGFTs) have shown good performance in terms of energy compaction, improving HEVC intra coding performance when used to replace the Discrete Cosine Transform (DCT). This paper further explores the performance of the SBGFTs in a multiple transforms, non-linear approximation perspective, by comparing them with two alternative sets of orthogonal transforms, namely, the Karhunen-Loève Transform (KLT) and the Sparse Orthonormal Transform (SOT). Experimental results confirm that SBGFTs achieve superior representation ability in this context as well, suggesting that they could assume a central role in image compression. Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi, Antonio Ortega |
ICIP | 1 |
| 2021 | Combining Appearance and Gradient Information for Image Symmetry DetectionabstractThis work addresses the challenging problem of reflection symmetry detection in unconstrained environments. Starting from the understanding on how the visual cortex manages planar symmetry detection, it is proposed to treat the problem in two stages: i) the design of a stable metric that extracts subsets of consistently oriented candidate segments, whenever the underlying 2D signal appearance exhibits definite near symmetric correspondences; ii) the ranking of such segments on the basis of the surrounding gradient orientation specularity, in order to reflect real symmetric object boundaries. Since these operations are related to the way the human brain performs planar symmetry detection, a better correspondence can be established between the outcomes of the proposed algorithm and a human-constructed ground truth. When compared to the testing sets used in recent symmetry detection competitions, a remarkable performance gain can be observed. In additional, further validation has been achieved by conducting perceptual validation experiments with users on a newly built dataset. Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi |
IEEE Trans. Image Process. | 1 |
| 2020 | 2D Discrete Mirror Transform for Image Non-Linear ApproximationabstractIn this paper, a new 2D transform named Discrete Mirror Transform (DMT) is presented. The DMT is computed by decomposing a signal into its even and odd parts around an optimal location in a given direction so that the signal energy is maximally split between the two components. After minimizing the information required to regenerate the original signal by removing redundant structures, the process is iterated leading the signal energy to distribute into a continuously smaller set of coefficients. The DMT can be displayed as a binary tree, where each node represents the single (even or odd) signal derived from the decomposition in the previous level. An optimized version of the DMT (ODMT) is also introduced, by exploiting the possibility to choose different directions at which performing the decomposition. Experimental simulations have been carried out in order to test the sparsity properties of the DMT and ODMT when applied on images: referring to both transforms, the results show a superior performance with respect to the popular Discrete Cosine Transform (DCT) and Discrete Wavelet Transform (DWT) in terms of non-linear approximation. Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi |
ICPR | 1 |
| 2019 | Iterative Mirror Decomposition for Signal RepresentationabstractIn this paper it is shown how to describe any finite-energy continuous or discrete signal through an ordered set of positions to uniquely represent it. This is obtained by designing an iterative decomposition through a series of mirror operations around those positions. The purpose is to find at any step of the decomposition the location that provides for the maximum decoupling between the even and odd components of the signal with respect to it. The algorithm can then be iterated at infinity determining a sequence of positions. The per location information determines the optimal energy decoupling strategy at each stage providing remarkable sparsity in the representation. Thanks to the sparsity of the resulting representation, experimental simulations demonstrate superior approximation capabilities of this proposed non-linear mirror transform. Fabrizio Guerrini, Alessandro Gnutti, Riccardo Leonardi |
ICASSP | 2 |
| 2019 | Coding of Image Intra Prediction Residuals Using Symmetric GraphsabstractThe Discrete Cosine Transform (DCT) is widely deployed by modern image and video coding standards such as JPEG and H.26x. In most cases, the DCT is applied in a separable manner to rows and columns, which limits its ability to represent signals with diagonal orientation. As an alternative, non-separable transforms can represent signals with different orientations, but are significantly more computationally complex. To address this problem, in this paper we propose a set of non-separable Symmetry-Based Graph Fourier Transforms (SBGFTs), whose symmetric structures lead to a faster implementation. We study a practical image coding scenario that exploits the proposed SBGFTs, where for each intra predicted image residual block the optimal graph is chosen by solving a graph-based Rate-Distortion (R-D) problem. Experimental results indicate a coding efficiency higher than JPEG and JPEG2000. Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi, Antonio Ortega |
ICIP | 1 |
| 2018 | Symmetry-Based Graph Fourier Transforms for Image RepresentationabstractIt is well-known that the application of the Discrete Cosine Transform (DCT) in transform coding schemes is justified by the fact that it belongs to a family of transforms asymptotically equivalent to the Karhunen-Loeve Transform (KLT) of a first order Markov process. However, when the pixel-to-pixel correlation is low the DCT does not provide a compression performance comparable with the KLT. In this paper, we propose a set of symmetry-based Graph Fourier Transforms (GFT) whose associated graphs present a totally or partially symmetric grid. We show that this family of transforms well represents both natural images and residual signals outperforming the DCT in terms of energy compaction. We also investigate how to reduce the cardinality of the set of transforms through an analysis that studies the relation between efficient symmetry-based GFTs and the directional modes used in H.265 standard. Experimental results indicate that coding efficiency is high. Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi, Antonio Ortega |
ICIP | 1 |
| 2017 | Even/odd decomposition made sparse: A fingerprint to hidden patterns
Fabrizio Guerrini, Alessandro Gnutti, Riccardo Leonardi |
Signal Process. | 2 |