Diego Valsesia

dblp:136/4988 · DBLP profile ↗
← Back
52ranked-venue papers
27as first author
26since 2021 · last 2026
0000-0003-1997-2910ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 13 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 8 first-author · 12 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 8 since 2021Computer networks · 1 · 1 first-authorSecurity and privacy · 1 · 1 first-authorTheory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 A multiplication-free neural architecture for image restoration
abstract
Deep neural networks have established themselves as the dominant approach for image restoration problems such as deblurring and denoising. However, the high computational complexity of state-of-the-art designs prevents their effective use in resource-constrained scenarios such as edge devices, where power-efficient inference is key. In this paper, we present a state-of-the-art backbone neural network design for image restoration, called MuFIR (Multiplication-Free Image Restoration), that is entirely devoid of multiplication operations. When coupled with suitable hardware implementations, the proposed concept enables fast and low-complexity inference by requiring only integer additions and bitshifts. This is made possible by several ingredients proposed in this work, namely ternary weight quantization to eliminate multiplications in the main network layers, careful use of novel normalizations to ensure stability of the ternarized architecture, and quantization of specific parameters and activations to combinations of powers of two, to remove the remaining multiplications. This is coupled with an annealed training procedure which progressively transforms a conventional network into our multiplication-free design. We experimentally show that, despite the all-integer operations and the lack of multiplications, MuFIR achieves performance close to that of full-precision models in terms of deblurring and denoising quality.
Luca Dordoni, Diego Valsesia, Enrico Magli
Neurocomputing2
2025 DreamCache: Finetuning-Free Lightweight Personalized Image Generation via Feature Caching
abstract
Personalized image generation requires text-to-image generative models that capture the core features of a reference subject to allow for controlled generation across different contexts. Existing methods face challenges due to complex training requirements, high inference costs, limited flexibility, or a combination of these issues. In this paper, we introduce DreamCache, a scalable approach for efficient and high-quality personalized image generation. By caching a small number of reference image features from a subset of layers and a single timestep of the pretrained diffusion denoiser, DreamCache enables dynamic modulation of the generated image features through lightweight, trained conditioning adapters. DreamCache achieves state-of-the-art image and text alignment, utilizing an order of magnitude fewer extra parameters, and is both more computationally effective and versatile than existing models.1
Emanuele Aiello, Umberto Michieli, Diego Valsesia, Mete Ozay, Enrico Magli
CVPR3
2025 A Novel Method and Dataset for Depth-Guided Image Deblurring From Smartphone Lidar
abstract
Modern smartphones are equipped with Lidar sensors providing depth-sensing capabilities. Recent works have shown that this complementary sensor allows to improve various tasks in image processing, including deblurring. However, there is a current lack of datasets with realistic blurred images and paired mobile Lidar depth maps to further study the topic. At the same time, there is also a lack of blind zero-shot methods that can deblur a real image using the depth guidance without requiring extensive training sets of paired data. In this paper, we propose an image deblurring method based on denoising diffusion models that can leverage the Lidar depth guidance and does not require training data with paired Lidar depth maps. We also present the first dataset with real blurred images with corresponding Lidar depth maps and sharp ground truth images, acquired with an Apple iPhone 15 Pro, for the purpose of studying Lidar-guided deblurring. Experimental results on this novel dataset show that Lidar guidance is effective and the proposed method outperforms state-of-the-art deblurring methods in terms of perceptual quality.
Antonio Montanaro, Diego Valsesia
ICIP2
2025 Modeling Uncertainty for Gaussian Splatting
abstract
We present stochastic Gaussian splatting (SGS): the first framework for uncertainty estimation using Gaussian splatting (GS). GS recently advanced the novel-view synthesis field by achieving impressive reconstruction quality at a fraction of the computational cost of neural radiance fields (NeRFs). However, contrary to the latter, it still lacks the ability to provide information about the confidence associated with their outputs. To address this limitation, in this brief, we introduce a variational inference (VI)-based approach that seamlessly integrates uncertainty prediction into the common rendering pipeline of GS. In addition, we introduce the area under sparsification error (AUSE) as a new term in the loss function, enabling optimization of uncertainty estimation alongside image reconstruction. Experimental results on the three different datasets demonstrate that our method outperforms existing approaches in terms of both image rendering quality and uncertainty estimation accuracy. Overall, our framework equips practitioners with valuable insights into the reliability of synthesized views, facilitating safer decision-making in real-world applications.
Luca Savant Aira, Diego Valsesia, Enrico Magli
IEEE Trans. Neural Networks Learn. Syst.2
2024 Degradation-Aware Self-Supervised Multi-Temporal Super-Resolution
abstract
Learning deep super-resolution models without the need for ground truth data at a higher resolution is critical for satellite imaging applications. This is either due to the lack of existing images at better resolution for certain target wavelengths or the existence of significant domain gaps between the images of different satellites. In this paper, we propose a method and neural network architecture for a multi-image super-resolution problem, where each image in the input stack might be affected by a different degradation process. A test-time finetuning procedure allows to dynamically account for the degradations observed for a specific set of LR inputs, improving over baseline results.
Matteo Impieri, Diego Valsesia, Tiziano Bianchi, Enrico Magli
IGARSS2
2024 Hybrid Recurrent-Attentive Neural Network for Onboard Predictive Hyperspectral Image Compression
abstract
AI-based compression is gaining popularity for traditional photos and videos. However, such techniques do not typically scale well to the task of compressing hyperspectral images, and may have computational requirements in terms of memory usage and total floating point operations that are prohibitive for usage onboard of satellites. In this paper, we explore the design of a predictive compression method based on a novel neural network design, called LineRWKV. Our neural network predictor works in a line-by-line fashion limiting memory and computational requirements thanks to a recurrent inference mechanism. However, in contrast to classic recurrent networks, it relies on an attention operation that can be parallelized for training, akin to Transformers, unlocking efficient training on large datasets, which is critical to learn complex predictors. In our preliminary results, we show that LineRWKV significantly outperforms the state-of-the-art CCSDS-123 standard and has competitive throughput.
Diego Valsesia, Tiziano Bianchi, Enrico Magli
IGARSS1
2024 MotionCraft: Physics-Based Zero-Shot Video Generation
abstract
Generating videos with realistic and physically plausible motion is one of the main recent challenges in computer vision. While diffusion models are achieving compelling results in image generation, video diffusion models are limited by heavy training and huge models, resulting in videos that are still biased to the training dataset. In this work we propose MotionCraft, a new zero-shot video generator to craft physics-based and realistic videos. MotionCraft is able to warp the noise latent space of an image diffusion model, such as Stable Diffusion, by applying an optical flow derived from a physics simulation. We show that warping the noise latent space results in coherent application of the desired motion while allowing the model to generate missing elements consistent with the scene evolution, which would otherwise result in artefacts or missing content if the flow was applied in the pixel space. We compare our method with the state-of-the-art Text2Video-Zero reporting qualitative and quantitative improvements, demonstrating the effectiveness of our approach to generate videos with finely-prescribed complex motion dynamics.
Antonio Montanaro, Luca Savant Aira, Emanuele Aiello, Diego Valsesia, Enrico Magli
NeurIPS4
2024 Onboard Deep Lossless and Near-Lossless Predictive Coding of Hyperspectral Images With Line-Based Attention
abstract
Deep learning methods have traditionally been difficult to apply to compression of hyperspectral images onboard spacecrafts due to the large computational complexity needed to achieve adequate representational power, as well as the lack of suitable datasets for training and testing. In this article, we depart from the traditional autoencoder approach, and we design a predictive neural network, called line receptance weighted key value (LineRWKV), which works recursively line by line to limit memory consumption. In order to achieve that, we adopt a novel hybrid attentive-recursive operation that combines the representational advantages of Transformers with the linear complexity and recursive implementation of recurrent neural networks (RNNs). The compression algorithm performs the prediction of each pixel using LineRWKV, followed by entropy coding of the residual. Experiments on multiple datasets show that LineRWKV is highly memory-efficient, significantly outperforms state-of-the-art deep learning methods, and is the first deep learning approach to outperform CCSDS-123.0-B-2 at lossless and near-lossless compression. Promising throughput results are also evaluated on a 7-W embedded system.
Diego Valsesia, Tiziano Bianchi, Enrico Magli
IEEE Trans. Geosci. Remote. Sens.1
2023 Multi-Level Fusion for Burst Super-Resolution with Deep Permutation-Invariant Conditioning
abstract
Developing deep learning techniques for super-resolving bursts of images acquired by mobile cameras is a topic that has recently gained significant interest. This topic fits the general problem of learning-based multi-image super-resolution (SR), which, contrary to its sibling single-image SR, has so far received little attention despite its potential. In this work, we introduce a neural network architecture for burst SR, called MLB-FuseNet (Multi-Level Burst Fusion Network), that is capable of extracting features in a manner that is invariant to permutations in the burst and to progressively condition features extracted from a reference image. Permutation invariance is desirable as it is known that the order of images in a burst does not matter in this problem, but its study has so far been neglected. Moreover, we also introduce a module exploiting a polyphase decomposition to improve feature extraction from mosaiced raw images. Results show an improvement over the state of the art on the BurstSR dataset – a recent and popular benchmark for this problem.
Martina Cilia, Diego Valsesia, Giulia Fracastoro, Enrico Magli
ICASSP2
2023 Towards Hyperbolic Regularizers For Point Cloud Part Segmentation
abstract
Hyperbolic neural networks are emerging as an effective technique to better capture hierarchical representations of many data types, from text to images and, recently, point clouds. In this paper, we extend our earlier work, that showed how to use regularizers in the hyperbolic space to improve performance of point cloud classification models, to the problem of part segmentation. This requires careful modeling of the hierarchical relationships between parts and whole point cloud to properly control the hyperbolic geometry of the feature space produced by the neural network. We show how the proposed method improves the performance of commonly used neural network architectures, reaching state-of-the-art performance on the part segmentation task.
Antonio Montanaro, Diego Valsesia, Enrico Magli
ICASSP2
2023 Onboard Processing Capabilities of an Earth Observation Compressive Sensing Payload
abstract
In this paper, we explore the onboard processing capabilities of an optical Earth observation instrument operating under the principles of compressed sensing, currently under preliminary study. In particular, we focus on two main aspects for onboard operations: i) how to process measurements in a computationally-efficient way to obtain previews of the reconstructed image that can be easily used by downstream inference algorithms; ii) the possibility of having simultaneous compression and encryption by proper management of the pseudorandom patterns used for the sensing matrix and measurements
Tiziano Bianchi, Martina Cilia, Enrico Magli, Andrea Migliorati, Nicola Prette, Diego Valsesia
IGARSS6
2023 Diffraction Efficiency-Aware Reconstruction for Compressive Sensing in the Mid-Infrared
abstract
Compressive sensing has established itself as a novel imaging paradigm. In this paper, we analyze the behavior of a a compressive instrument based on spatial light modulators (SLM), operating in the mid-infrared. We show that, contrary to the well-studied visible and near-infrared wavelengths, mid-infrared poses modeling challenges due to non-negligible SLM diffraction effects. We show a way to model such effect analytically and to account for them in the reconstruction process, leading to improved reconstruction quality.
Tiziano Bianchi, Donatella Guzzi, Cinzia Lastri, Enrico Magli, Vanni Nardino, Lorenzo Palombi, Nicola Prette, Valentina Raimondi, Diego Valsesia
IGARSS9
2023 Proba-V Multi-Temporal Super-Resolution Guided by Sentinel-2
abstract
Multi-image super-resolution (MISR) is a technique used to increase the spatial resolution of images acquired by remote sensing platforms by combining the images acquired through multiple revisits. Supervised training of MISR models requires collecting high-resolution images to be used as ground truth. Except for a few special cases, this involves acquiring images from a different satellite, resulting in a shift in the optical and radiometric characteristics with respect to the sensor to be super-resolved. In this paper, we explore the use of Sentinel-2 images to train a MISR model for Proba-V images and highlight the challenges of this pursuit.
Gabriele Inzerillo, Diego Valsesia, Enrico Magli, Fabrizio Niro, Erminia De Grandis
IGARSS2
2023 Towards Unsupervised Multi-Temporal Satellite Image Super-Resolution
abstract
Multi-temporal super-resolution (SR) whereby a number of images of the same scene acquired at different times are fused to enhance its spatial resolution has recently enjoyed great success thanks to advances in deep learning methods. However, the literature has so far focused on supervised training approaches that require the availability of high-resolution (HR) images at the target resolution. This is a significant limitation because such imagery may not exist, might be difficult to source or exhibit domain gaps such as different spectral bands or radiometric characteristics. Unsupervised training approaches that do not require imagery beyond the input low resolution are needed to overcome this limitation. This paper presents a first analysis of the problem, taking inspiration from the literature on blind single-image SR, but also focusing on the uniqueness of multi-temporal satellite images. Our preliminary results show that it is indeed possible to develop accurate deep learning models for multi-temporal SR without HR images.
Nicola Prette, Diego Valsesia, Tiziano Bianchi, Enrico Magli
IGARSS2
2023 RAN-GNNs: Breaking the Capacity Limits of Graph Neural Networks
Diego Valsesia, Giulia Fracastoro, Enrico Magli
IEEE Trans. Neural Networks Learn. Syst.1
2022 Signal Compression via Neural Implicit Representations
abstract
Existing end-to-end signal compression schemes using neural networks are largely based on an autoencoder-like structure, where a universal encoding function creates a compact latent space and the signal representation in this space is quantized and stored. Recently, advances from the field of 3D graphics have shown the possibility of building implicit representation networks, i.e., neural networks returning the value of a signal at a given query coordinate. In this paper, we propose using neural implicit representations as a novel paradigm for signal compression with neural networks, where the compact representation of the signal is defined by the very weights of the network. We discuss how this compression framework works, how to include priors in the design, and highlight interesting connections with transform coding. While the framework is general, and still lacks maturity, we already show very competitive performance on the task of compressing point cloud attributes, which is notoriously challenging due to the irregularity of the domain, but becomes trivial in the proposed framework.
Francesca Pistilli, Diego Valsesia, Giulia Fracastoro, Enrico Magli
ICASSP2
2022 Exploring the Solution Space of Linear Inverse Problems with GAN Latent Geometry
abstract
Inverse problems consist in reconstructing signals from incomplete sets of measurements and their performance is highly dependent on the quality of the prior knowledge encoded via regularization. While traditional approaches focus on obtaining a unique solution, an emerging trend considers exploring multiple feasibile solutions. In this paper, we propose a method to generate multiple reconstructions that fit both the measurements and a data-driven prior learned by a generative adversarial network. In particular, we show that, starting from an initial solution, it is possible to find directions in the latent space of the generative model that are null to the forward operator, and thus keep consistency with the measurements, while inducing significant perceptual change. Our exploration approach allows to generate multiple solutions to the inverse problem an order of magnitude faster than existing approaches; we show results on image super-resolution and inpainting problems.
Antonio Montanaro, Diego Valsesia, Enrico Magli
ICIP2
2022 Super-Resolved Multi-Temporal Segmentation with Deep Permutation-Invariant Networks
abstract
Multi-image super-resolution from multi-temporal satellite acquisitions of a scene has recently enjoyed great success thanks to new deep learning models. In this paper, we go beyond classic image reconstruction at a higher resolution by studying a super-resolved inference problem, namely semantic segmentation at a spatial resolution higher than the one of sensing platform. We expand upon recently proposed models exploiting temporal permutation invariance with a multi-resolution fusion module able to infer the rich semantic information needed by the segmentation task. The model presented in this paper has recently won the AI4EO challenge on Enhanced Sentinel 2 Agriculture.
Diego Valsesia, Enrico Magli
IGARSS1
2022 Cross-modal Learning for Image-Guided Point Cloud Shape Completion
abstract
In this paper we explore the recent topic of point cloud completion, guided by an auxiliary image. We show how it is possible to effectively combine the information from the two modalities in a localized latent space, thus avoiding the need for complex point cloud reconstruction methods from single views used by the state-of-the-art. We also investigate a novel self-supervised setting where the auxiliary image provides a supervisory signal to the training process by using a differentiable renderer on the completed point cloud to measure fidelity in the image space. Experiments show significant improvements over state-of-the-art supervised methods for both unimodal and multimodal completion. We also show the effectiveness of the self-supervised approach which outperforms a number of supervised methods and is competitive with the latest supervised models only exploiting point cloud information.
Emanuele Aiello, Diego Valsesia, Enrico Magli
NeurIPS2
2022 Rethinking the compositionality of point clouds through regularization in the hyperbolic space
abstract
Point clouds of 3D objects exhibit an inherent compositional nature where simple parts can be assembled into progressively more complex shapes to form whole objects. Explicitly capturing such part-whole hierarchy is a long-sought objective in order to build effective models, but its tree-like nature has made the task elusive. In this paper, we propose to embed the features of a point cloud classifier into the hyperbolic space and explicitly regularize the space to account for the part-whole hierarchy. The hyperbolic space is the only space that can successfully embed the tree-like nature of the hierarchy. This leads to substantial improvements in the performance of state-of-art supervised models for point cloud classification.
Antonio Montanaro, Diego Valsesia, Enrico Magli
NeurIPS2
2022 Semi-Supervised Learning for Joint SAR and Multispectral Land Cover Classification
abstract
Semi-supervised learning techniques are gaining popularity due to their capability of building models that are effective, even when scarce amounts of labeled data are available. In this paper, we present a framework and specific tasks for self-supervised pretraining ofmultichannelmodels, such as the fusion of multispectral and synthetic aperture radar images. We show that the proposed self-supervised approach is highly effective at learning features that correlate with the labels for land cover classification. This is enabled by an explicit design of pretraining tasks which promotes bridging the gaps between sensing modalities and exploiting the spectral characteristics of the input. In a semi-supervised setting, when limited labels are available, using the proposed self-supervised pretraining, followed by supervised finetuning for land cover classification with SAR and multispectral data, outperforms conventional approaches such as purely supervised learning, initialization from training on ImageNet and other recent self-supervised approaches.
Antonio Montanaro, Diego Valsesia, Giulia Fracastoro, Enrico Magli
IEEE Geosci. Remote. Sens. Lett.2
2022 Speckle2Void: Deep Self-Supervised SAR Despeckling With Blind-Spot Convolutional Neural Networks
abstract
Information extraction from synthetic aperture radar (SAR) images is heavily impaired by speckle noise, and hence, despeckling is a crucial preliminary step in scene analysis algorithms. The recent success of deep learning envisions a new generation of despeckling techniques that could outperform classical model-based methods. However, current deep learning approaches to despeckling require supervision for training, whereas clean SAR images are impossible to obtain. In the literature, this issue is tackled by resorting to either synthetically speckled optical images, which exhibit different properties with respect to true SAR images, or multitemporal SAR images, which are difficult to acquire or fuse accurately. In this article, inspired by recent works on blind-spot denoising networks, we propose a self-supervised Bayesian despeckling method. The proposed method is trained by employing only noisy SAR images and can, therefore, learn features of real SAR images rather than synthetic data. Experiments show that the performance of the proposed approach is very close to the supervised training approach on synthetic data and superior on real data in both quantitative and visual assessments.
Andrea Bordone Molini, Diego Valsesia, Giulia Fracastoro, Enrico Magli
IEEE Trans. Geosci. Remote. Sens.2
2022 Permutation Invariance and Uncertainty in Multitemporal Image Super-Resolution
abstract
Recent advances have shown how deep neural networks can be extremely effective at super-resolving remote sensing imagery, starting from a multitemporal collection of low-resolution images. However, existing models have neglected the issue of temporal permutation, whereby the temporal ordering of the input images does not carry any relevant information for the super-resolution task and causes such models to be inefficient with the, often scarce, ground truth data that available for training. Thus, models ought not to learn feature extractors that rely on temporal ordering. In this paper, we show how building a model that is fully invariant to temporal permutation significantly improves performance and data efficiency. Moreover, we study how to quantify the uncertainty of the super-resolved image so that the final user is informed on the local quality of the product. We show how uncertainty correlates with temporal variation in the series, and how quantifying it further improves model performance. Experiments on the Proba-V challenge dataset show significant improvements over the state of the art without the need for self-ensembling, as well as improved data efficiency, reaching the performance of the challenge winner with just 25% of the training data.
Diego Valsesia, Enrico Magli
IEEE Trans. Geosci. Remote. Sens.1
2021 Denoise and Contrast for Category Agnostic Shape Completion
abstract
In this paper, we present a deep learning model that exploits the power of self-supervision to perform 3D point cloud completion, estimating the missing part and a context region around it. Local and global information are encoded in a combined embedding. A denoising pretext task provides the network with the needed local cues, decoupled from the high-level semantics and naturally shared over multiple classes. On the other hand, contrastive learning maximizes the agreement between variants of the same shape with different missing portions, thus producing a representation which captures the global appearance of the shape. The combined embedding inherits category-agnostic properties from the chosen pretext tasks. Differently from existing approaches, this allows to better generalize the completion properties to new categories unseen at training time. Moreover, while decoding the obtained joint representation, we better blend the reconstructed missing part with the partial shape by paying attention to its known surrounding region and reconstructing this frame as auxiliary objective. Our extensive experiments and detailed ablation on the ShapeNet dataset show the effectiveness of each part of the method with new state of the art results. Our quantitative and qualitative analysis confirms how our approach is able to work on novel categories without relying neither on classification and shape symmetry priors, nor on adversarial training procedures.
Antonio Alliegro, Diego Valsesia, Giulia Fracastoro, Enrico Magli, Tatiana Tommasi
CVPR2
2021 Spatial Light Modulator-Based Architecture to Implement a Super-Resolved Compressive Instrument for Earth Observation
abstract
Due to a growing interest for imagery with high spatial and spectral resolution, Earth Observation sensors are producing increasing amounts of data. This poses a severe challenge in terms of computational, memory and transmission requirements. In order to overcome these limitations, a fascinating approach is the implementation of a compressive sensing architecture. In this paper, we present an instrumental concept based on the use of a spatial light modulator to implement a super-resolved, compressive demonstrator of an instrument aimed at Earth Observation in the visible and medium infrared spectral regions from geostationary platform.
Valentina Raimondi, Luigi Acampora, Gabriele Amato, Massimo Baldi, Dirk Berndt, Alberto Bianchi, Tiziano Bianchi, Donato Borrelli, Valentina Colcelli, Chiara Corti, Francesco Corti, Marco Corti, Nick Cox, Ulrike A. Dauderstädt, Peter Dürr, Sara Francés González, Paolo Frosini, Donatella Guzzi, Jessica Huntingford, Detlef Kunze, Demetrio Labate, Nicolas Lamquin, Cinzia Lastri, Enrico Magli, Vanni Nardino, Christophe Pache, Lorenzo Palombi, Irene Pettinelli, Giuseppe Pilato, Alexandre Pollini, Leopoldo Rossini, Enrico Suetta, Davide Taricco, Diego Valsesia, Michael Wagner 0028
IGARSS34
2021 Learning Localized Representations of Point Clouds With Graph-Convolutional Generative Adversarial Networks
abstract
Point clouds are an important type of geometric data generated by 3D acquisition devices, and have widespread use in computer graphics and vision. However, learning representations for point clouds is particularly challenging due to their nature as being an unordered collection of points irregularly distributed in 3D space. Recently, supervised and semisupervised problems for point clouds leveraged graph convolution, a generalization of the convolution operation for data defined over graphs. This operation has been shown to be very successful at extracting localized features from point clouds. In this paper, we study the unsupervised problem of a generative model exploiting graph convolution. Employing graph convolution operations in generative models is not straightforward and it poses some unique challenges. In particular, we focus on the generator of a GAN, where the graph is not known in advance as it is the very output of the generator. We show that the proposed architecture can learn to generate the graph and the features simultaneously. We also study the problem of defining an upsampling layer in the graph-convolutional generator, proposing two methods that respectively learn to exploit a multi-resolution or self-similarity prior to sample the data distribution.
Diego Valsesia, Giulia Fracastoro, Enrico Magli
IEEE Trans. Multim.1
2020 Learning Graph-Convolutional Representations for Point Cloud Denoising
Francesca Pistilli, Giulia Fracastoro, Diego Valsesia, Enrico Magli
ECCV (20)3
2020 Deepsum++: Non-Local Deep Neural Network for Super-Resolution of Unregistered Multitemporal Images
abstract
Deep learning methods for super-resolution of a remote sensing scene from multiple unregistered low-resolution images have recently gained attention thanks to a challenge proposed by the European Space Agency. This paper presents an evolution of the winner of the challenge, showing how incorporating non-local information in a convolutional neural network allows to exploit self-similar patterns that provide enhanced regularization of the super-resolution problem. Experiments on the dataset of the challenge show improved performance over the state-of-the-art, which does not exploit non-local information.
Andrea Bordone Molini, Diego Valsesia, Giulia Fracastoro, Enrico Magli
IGARSS2
2020 Towards Deep Unsupervised Sar Despeckling with Blind-Spot Convolutional Neural Networks
abstract
SAR despeckling is a problem of paramount importance in remote sensing, since it represents the first step of many scene analysis algorithms. Recently, deep learning techniques have outperformed classical model-based despeckling algorithms. However, such methods require clean ground truth images for training, thus resorting to synthetically speckled optical images since clean SAR images cannot be acquired. In this paper, inspired by recent works on blind-spot denoising networks, we propose a self-supervised Bayesian despeckling method. The proposed method is trained employing only noisy images and can therefore learn features of real SAR images rather than synthetic data. We show that the performance of the proposed network is very close to the supervised training approach on synthetic data and competitive on real data.
Andrea Bordone Molini, Diego Valsesia, Giulia Fracastoro, Enrico Magli
IGARSS2
2020 Detection of Solar Coronal Mass Ejections from Raw Images with Deep Convolutional Neural Networks
abstract
Coronal Mass Ejections (CMEs) are massive releases of plasma from the solar corona. When the charged material is ejected towards the Earth, it can cause geomagnetic storms and severely damage electronic equipment and power grids. Early detection of CMEs is therefore crucial for damage containment. In this paper, we study detection of CMEs from sequential images of the solar corona acquired by a satellite. A low-complexity deep neural network is trained to process the raw images, ideally directly on the satellite, in order to provide early alerts.
Diego Valsesia, Andrea Grippi, Enrico Magli, Roberto Susino, Daniele Telloni, Gianalfredo Nicolini, Marta Casti, Angelo Fabio Mulone, Rosario Messineo
IGARSS1
2020 NIR image colorization with graph-convolutional neural networks
abstract
Colorization of near-infrared (NIR) images is a challenging problem due to the different material properties at the infared wavelenghts, thus reducing the correlation with visible images. In this paper, we study how graph-convolutional neural networks allow exploiting a more powerful inductive bias than standard CNNs, in the form of non-local self-similiarity. Its impact is evaluated by showing how training with mean squared error only as loss leads to poor results with a standard CNN, while the graph-convolutional network produces significantly sharper and more realistic colorizations.
Diego Valsesia, Giulia Fracastoro, Enrico Magli
VCIP1
2020 DeepSUM: Deep Neural Network for Super-Resolution of Unregistered Multitemporal Images
abstract
Recently, convolutional neural networks (CNNs) have been successfully applied to many remote sensing problems. However, deep learning techniques for multi-image super-resolution (SR) from multitemporal unregistered imagery have received little attention so far. This article proposes a novel CNN-based technique that exploits both spatial and temporal correlations to combine multiple images. This novel framework integrates the spatial registration task directly inside the CNN, and allows one to exploit the representation learning capabilities of the network to enhance registration accuracy. The entire SR process relies on a single CNN with three main stages: shared 2-D convolutions to extract high-dimensional features from the input images; a subnetwork proposing registration filters derived from the high-dimensional feature representations; 3-D convolutions for slow fusion of the features from multiple images. The whole network can be trained end-to-end to recover a single high-resolution image from multiple unregistered low-resolution images. The method presented in this article is the winner of the PROBA-V SR challenge issued by the European Space Agency (ESA).
Andrea Bordone Molini, Diego Valsesia, Giulia Fracastoro, Enrico Magli
IEEE Trans. Geosci. Remote. Sens.2
2020 Deep Graph-Convolutional Image Denoising
abstract
Non-local self-similarity is well-known to be an effective prior for the image denoising problem. However, little work has been done to incorporate it in convolutional neural networks, which surpass non-local model-based methods despite only exploiting local information. In this paper, we propose a novel end-to-end trainable neural network architecture employing layers based on graph convolution operations, thereby creating neurons with non-local receptive fields. The graph convolution operation generalizes the classic convolution to arbitrary graphs. In this work, the graph is dynamically computed from similarities among the hidden features of the network, so that the powerful representation learning capabilities of the network are exploited to uncover self-similar patterns. We introduce a lightweight Edge-Conditioned Convolution which addresses vanishing gradient and over-parameterization issues of this particular graph convolution. Extensive experiments show state-of-the-art performance with improved qualitative and quantitative results on both synthetic Gaussian noise and real noise.
Diego Valsesia, Giulia Fracastoro, Enrico Magli
IEEE Trans. Image Process.1
2019 Image Denoising with Graph-Convolutional Neural Networks
abstract
Recovering an image from a noisy observation is a key problem in signal processing. Recently, it has been shown that data-driven approaches employing convolutional neural networks can outperform classical model-based techniques, because they can capture more powerful and discriminative features. However, since these methods are based on convolutional operations, they are only capable of exploiting local similarities without taking into account non-local self-similarities. In this paper we propose a convolutional neural network that employs graph-convolutional layers in order to exploit both local and non-local similarities. The graph-convolutional layers dynamically construct neighborhoods in the feature space to detect latent correlations in the feature maps produced by the hidden layers. The experimental results show that the proposed architecture outperforms classical convolutional neural networks for the denoising task.
Diego Valsesia, Giulia Fracastoro, Enrico Magli
ICIP1
2019 Learning Localized Generative Models for 3D Point Clouds via Graph Convolution
Diego Valsesia, Giulia Fracastoro, Enrico Magli
ICLR (Poster)1
2019 Analysis of SparseHash: An efficient embedding of set-similarity via sparse projections
Diego Valsesia, Sophie M. Fosson, Chiara Ravazzi, Tiziano Bianchi, Enrico Magli
Pattern Recognit. Lett.1
2019 High-Throughput Onboard Hyperspectral Image Compression With Ground-Based CNN Reconstruction
abstract
Compression of hyperspectral images onboard of spacecrafts is a tradeoff between the limited computational resources and the ever-growing spatial and spectral resolution of the optical instruments. As such, it requires low-complexity algorithms with good rate-distortion performance and high throughput. In recent years, the Consultative Committee for Space Data Systems (CCSDS) has focused on lossless and near-lossless compression approaches based on predictive coding, resulting in the recently published CCSDS 123.0-B-2 recommended standard. While the in-loop reconstruction of quantized prediction residuals provides excellent rate-distortion performance for the near-lossless operating mode, it significantly constrains the achievable throughput due to data dependencies. In this paper, we study the performance of a faster method based on the prequantization of the image followed by a lossless predictive compressor. While this is well known to be suboptimal, one can exploit powerful signal models to reconstruct the image at the ground segment, recovering part of the suboptimality. In particular, we show that convolutional neural networks can be used for this task and that they can recover the whole SNR drop incurred at a bit rate of 2 bits per pixel.
Diego Valsesia, Enrico Magli
IEEE Trans. Geosci. Remote. Sens.1
2017 Fast and Lightweight Rate Control for Onboard Predictive Coding of Hyperspectral Images
abstract
Predictive coding is attractive for compression of hyperspectral images onboard of spacecrafts in light of the excellent rate-distortion performance and low complexity of recent schemes. In this letter, we propose a rate control algorithm and integrate it in a lossy extension to the CCSDS-123 lossless compression recommendation. The proposed rate algorithm overhauls our previous scheme by being orders of magnitude faster and simpler to implement, while still providing the same accuracy in terms of output rate and comparable or better image quality.
Diego Valsesia, Enrico Magli
IEEE Geosci. Remote. Sens. Lett.1
2017 Binary Adaptive Embeddings From Order Statistics of Random Projections
abstract
We use some of the largest order statistics of the random projections of a reference signal to construct a binary embedding that is adapted to signals correlated with such signal. The embedding is characterized from the analytical standpoint and shown to provide improved performance on tasks such as classification in a reduced-dimensionality space.
Diego Valsesia, Enrico Magli
IEEE Signal Process. Lett.1
2017 User Authentication via PRNU-Based Physical Unclonable Functions
abstract
Multifactor user authentication systems enhance security by augmenting passwords with the verification of additional pieces of information such as the possession of a particular device. This paper presents an innovative user authentication scheme that verifies the possession of one’s smartphone by uniquely identifying its camera. High-frequency components of the photo-response nonuniformity of the optical sensor are extracted from raw images and used as a weak physical unclonable function. A novel scheme for efficient transmission and server-side verification is also designed based on adaptive random projections and on an innovative fuzzy extractor using polar codes. The security of the system is thoroughly analyzed under different attack scenarios both theoretically and experimentally.
Diego Valsesia, Giulio Coluccia, Tiziano Bianchi, Enrico Magli
IEEE Trans. Inf. Forensics Secur.1
2016 Universal encoding of multispectral images
abstract
We propose a new method for low-complexity compression of multispectral images. We develop on a novel approach to coding signals with side information based on recent advances in compressed sensing and universal scalar quantization. Our approach can be interpreted as a variation of quantized compressed sensing, where the most significant bits are discarded at the encoder and recovered at the decoder from the side information. The image is reconstructed using weighted total variation minimization, incorporating side information in the weights while enforcing consistency with the recovered quantized coefficient values. Our experiments validate our approach and confirm the improvements in rate-distortion performance.
Diego Valsesia, Petros Boufounos
ICASSP1
2016 Multispectral image compression using universal vector quantization
abstract
We propose a new method for low-complexity compression of multispectral images based on universal vector quantization. Our approach generalizes the recently developed theory of universal scalar quantization to vector quantization, and uses it in the context of distributed coding. We exploit the availability of side information on the decoder to reduce the encoding rate of a vector quantizer, applied to compressed measurements of the image. The encoding reuses quantization labels to label multiple quantization cells and leverages the side information to select the correct cell at the decoder. The image is reconstructed using weighted total variation minimization, incorporating side information in the weights while enforcing consistency with the recovered quantization cell.
Diego Valsesia, Petros Boufounos
ITW1
2015 Scale-robust compressive camera fingerprint matching with random projections
abstract
Recently, we demonstrated that random projections can provide an extremely compact representation of a camera fingerprint without significantly affecting the matching performance. In this paper, we propose a new construction that makes random projections of camera fingerprints scale-robust. The proposed method maps the compressed fingerprint of a rescaled image to the compressed fingerprint of the original image, rescaled by the same factor. In this way, fingerprints obtained from rescaled images can be directly matched in the compressed domain, which is much more efficient than existing scale-robust approaches. Experimental results on the publicly available Dresden database show that the proposed technique is robust to a wide range of scale transformations. Moreover, robustness can be further improved by providing reference scales in the database, with a small additional storage cost.
Diego Valsesia, Giulio Coluccia, Tiziano Bianchi, Enrico Magli
ICASSP1
2015 Image retrieval based on compressed camera sensor fingerprints
abstract
Image retrieval is the process of finding images from a large collection, satisfying a user-specified criterion. Content-based retrieval has been the traditional paradigm, in which one wishes to find images whose content is similar to a query. In this paper we explore a novel criterion for image search, based on forensic principles. We address the problem of retrieving all the photos in a collection that have been acquired by a specific device which is presented to the system as a query. This is an important forensic problem, whose solution could be very useful for detecting improper usage of pictures. We do not rely on metadata such as Exif headers because they can be unavailable, or easily manipulated, and in most cases cannot identify the specific device. We rely instead on a forensic tool called Photo Response Non-Uniformity (PRNU), which constitutes a reliable fingerprint of a camera sensor. We examine recent advances in compression of such fingerprints, which allow to address the previously unexplored image retrieval problem on large scales.
Diego Valsesia, Giulio Coluccia, Tiziano Bianchi, Enrico Magli
ICME1
2015 Graded Quantization for Multiple Description Coding of Compressive Measurements
abstract
Compressed sensing (CS) is an emerging paradigm for acquisition of compressed representations of a sparse signal. Its low complexity is appealing for resource-constrained scenarios like sensor networks. However, such scenarios are often coupled with unreliable communication channels and providing robust transmission of the acquired data to a receiver is an issue. Multiple description coding (MDC) effectively combats channel losses for systems without feedback, thus raising the interest in developing MDC methods explicitly designed for the CS framework, and exploiting its properties. We propose a method called Graded Quantization (CS-GQ) that leverages the democratic property of compressive measurements to effectively implement MDC, and we provide methods to optimize its performance. A novel decoding algorithm based on the alternating directions method of multipliers is derived to reconstruct signals from a limited number of received descriptions. Simulations are performed to assess the performance of CS-GQ against other methods in presence of packet losses. The proposed method is successful at providing robust coding of CS measurements and outperforms other schemes for the considered test metrics.
Diego Valsesia, Giulio Coluccia, Enrico Magli
IEEE Trans. Commun.1
2015 Large-Scale Image Retrieval Based on Compressed Camera Identification
abstract
Retrieving pictures from large collections according to a specific criterion is an increasingly relevant task. An important , but so far overlooked, such criterion is the retrieval of pictures acquired by a specific camera. Instead of relying on metadata , which can be absent or easily manipulated, a forensic tool is exploited, namely the photo response non-uniformity (PRNU) of the camera sensor. Recent works showed that random projections can be used to significantly compress the PRNU, enabling operation on very large scales, previously impossible due to the size of the PRNU and to the complexity of the matching operations. In this paper, we propose efficient techniques for management and retrieval of images employing the PRNU, and test them on a database of 1174 cameras and half a million pictures downloaded from the Internet.
Diego Valsesia, Giulio Coluccia, Tiziano Bianchi, Enrico Magli
IEEE Trans. Multim.1
2014 Compressive signal processing with circulant sensing matrices
abstract
Compressive sensing achieves effective dimensionality reduction of signals, under a sparsity constraint, by means of a small number of random measurements acquired through a sensing matrix. In a signal processing system, the problem arises of processing the random projections directly, without first reconstructing the signal. In this paper, we show that circulant sensing matrices allow to perform a variety of classical signal processing tasks such as filtering, interpolation, registration, transforms, and so forth, directly in the compressed domain and in an exact fashion, i.e., without relying on estimators as proposed in the existing literature. The advantage of the techniques presented in this paper is to enable direct measurement-to-measurement transformations, without the need of costly recovery procedures.
Diego Valsesia, Enrico Magli
ICASSP1
2014 A hardware-friendly architecture for onboard rate-controlled predictive coding of hyperspectral and multispectral images
abstract
In this paper we propose an efficient architecture for onboard implementation of rate-controlled predictive lossy compression of hyperspectral and multispectral images. In particular, we consider the recent state-of-the-art rate control algorithm for onboard predictive compression [1], and propose an architecture addressing two fundamental aspects of its hardware implementation. Specifically, this architecture overcomes the serial nature of the algorithm, as well as the large memory requirements of the entropy coding stage, achieving a pipelined implementation suitable for high-throughput onboard implementation, at a negligible cost in terms of coding efficiency.
Diego Valsesia, Enrico Magli
ICIP1
2014 A Novel Rate Control Algorithm for Onboard Predictive Coding of Multispectral and Hyperspectral Images
abstract
Predictive coding is attractive for compression on board of spacecraft due to its low computational complexity, modest memory requirements, and the ability to accurately control quality on a pixel-by-pixel basis. Traditionally, predictive compression focused on the lossless and near-lossless modes of operation, where the maximum error can be bounded but the rate of the compressed image is variable. Rate control is considered a challenging problem for predictive encoders due to the dependencies between quantization and prediction in the feedback loop and the lack of a signal representation that packs the signal's energy into few coefficients. In this paper, we show that it is possible to design a rate control scheme intended for onboard implementation. In particular, we propose a general framework to select quantizers in each spatial and spectral region of an image to achieve the desired target rate while minimizing distortion. The rate control algorithm allows achieving lossy near-lossless compression and any in-between type of compression, e.g., lossy compression with a near-lossless constraint. While this framework is independent of the specific predictor used, in order to show its performance, in this paper, we tailor it to the predictor adopted by the CCSDS-123 lossless compression standard, obtaining an extension that allows performing lossless, near-lossless, and lossy compression in a single package. We show that the rate controller has excellent performance in terms of accuracy in the output rate, rate-distortion characteristics, and is extremely competitive with respect to state-of-the-art transform coding.
Diego Valsesia, Enrico Magli
IEEE Trans. Geosci. Remote. Sens.1
2013 Graded quantization: Democracy for multiple descriptions in compressed sensing
abstract
The compressed sensing paradigm allows to efficiently represent sparse signals by means of their linear measurements. However, the problem of transmitting these measurements to a receiver over a channel potentially prone to packet losses has received little attention so far. In this paper, we propose novel methods to generate multiple descriptions from compressed sensing measurements to increase the robustness over unreliable channels. In particular, we exploit the democracy property of compressive measurements to generate descriptions in a simple manner by partitioning the measurement vector and properly allocating bit-rate, outperforming classical methods like the multiple description scalar quantizer. In addition, we propose a modified version of the Basis Pursuit Denoising recovery procedure that is specifically tailored to the proposed methods. Experimental results show significant performance gains with respect to existing methods.
Diego Valsesia, Giulio Coluccia, Enrico Magli
ICASSP1
2013 Smoothness-constrained image recovery from block-based random projections
abstract
In this paper we address the problem of visual quality of images reconstructed from block-wise random projections. Independent reconstruction of the blocks can severely affect visual quality, by displaying artifacts along block borders. We propose a method to enforce smoothness across block borders by modifying the sensing and reconstruction process so as to employ partially overlapping blocks. The proposed algorithm accomplishes this by computing a fast preview from the blocks, whose purpose is twofold. On one hand, it allows to enforce a set of constraints to drive the reconstruction algorithm towards a smooth solution, imposing the similarity of block borders. On the other hand, the preview is used as a predictor of the entire block, allowing to recover the prediction error, only. The quality improvement over the result of independent reconstruction can be easily assessed both visually and in terms of PSNR and SSIM index.
Giulio Coluccia, Diego Valsesia, Enrico Magli
MMSP2
2013 Spatially scalable compressed image sensing with hybrid transform and inter-layer prediction model
abstract
Compressive imaging is an emerging application of compressed sensing, devoted to acquisition, encoding and reconstruction of images using random projections as measurements. In this paper we propose a novel method to provide a scalable encoding of an image acquired by means of compressed sensing techniques. Two bit-streams are generated to provide two distinct quality levels: a low-resolution base layer and full-resolution enhancement layer. In the proposed method we exploit a fast preview of the image at the encoder in order to perform inter-layer prediction and encode the prediction residuals only. The proposed method successfully provides resolution and quality scalability with modest complexity and it provides gains in the quality of the reconstructed images with respect to separate encoding of the quality layers. Remarkably, we also show that the scheme can also provide significant gains with respect to a direct, non-scalable system, thus accomplishing two features at once: scalability and improved reconstruction performance.
Diego Valsesia, Enrico Magli
MMSP1