Pietro Zanuttigh

dblp:18/797 · DBLP profile ↗
← Back
63ranked-venue papers
5as first author
31since 2021 · last 2026
0000-0002-9502-2389ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 44 · 4 first-author · 18 since 2021Artificial intelligence and machine learning · 31 · 20 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 K-Merge: Online Continual Merging of Adapters for On-device Large Language Models
abstract
Donald Shenaj, Ondrej Bohdal, Taha Ceritli, Mete Ozay, Pietro Zanuttigh, Umberto Michieli. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Donald Shenaj, Ondrej Bohdal, Taha Ceritli, Mete Ozay, Pietro Zanuttigh, Umberto Michieli
ACL (1)5
2026 Federated Medical Image Classification Under Class and Domain Imbalance Exploiting Synthetic Sample Generation
Martina Pavan, Matteo Caligiuri, Francesco Barbato, Pietro Zanuttigh
ICPR (8)4
2026 FedPromo: Federated Lightweight Proxy Models at the Edge for Fine-Grained Image Classification With Foundation Models
abstract
Federated learning (FL) is an established paradigm for training deep learning models on decentralized data, particularly relevant in Internet of Things (IoT) scenarios. However, as model sizes grow, conventional FL approaches require significant computational resources, which may not be feasible for resource-constrained IoT devices. We introduce FedPromo, a novel framework that enables efficient adaptation of classification heads trained in a federated way to large-scale foundation models (FMs) stored on a central server without requiring explicit data sharing. Instead of directly training the large model on client devices, FedPromo optimizes lightweight proxy models via FL, reducing computational overhead, energy consumption, and bandwidth usage while maintaining privacy. We evaluate our method in cross-domain fine-grained image classification via a two-stage process: server-side knowledge distillation (KD) aligns representations of a large-scale FM with those of a compact counterpart, then the compact model encoder is deployed to devices for local classifier learning. These classifiers are aggregated and transferred back to the FM, enabling learning of fine-grained representations without direct access to user data. Through novel regularization strategies, our framework enables decentralized multidomain learning, balancing performance, privacy, and resource efficiency for wide-scale IoT deployment. Extensive experiments on five image classification benchmarks and feature representation analysis demonstrate that FedPromo outperforms existing methods when employing mobile-targeted efficient architectures, making it suitable for IoT devices such as smart home devices, industrial sensors, and phones. Across domains, FedPromo enjoys average gains of 4.8% (9.1%) top-1 (top-5) accuracy on the clients and 7.5% (9.7%) on the server with respect to the best competitors.
Matteo Caligiuri, Francesco Barbato, Donald Shenaj, Umberto Michieli, Pietro Zanuttigh
IEEE Internet Things J.5
2026 RECALL+: Adversarial web-based replay for continual learning in semantic segmentation
Chang Liu 0047, Giulia Rizzoli, Francesco Barbato, Andrea Maracani, Marco Toldo, Umberto Michieli, Pietro Zanuttigh
Image Vis. Comput.8
2026 FlyAwareV2: A multimodal cross-domain UAV dataset for urban scene understanding
abstract
The development of computer vision algorithms for Unmanned Aerial Vehicle (UAV) applications in urban environments heavily relies on the availability of large-scale datasets with accurate annotations. However, collecting and annotating real-world UAV data is extremely challenging and costly. To address this limitation, we present FlyAwareV2, a novel multimodal dataset encompassing both real and synthetic UAV imagery tailored for urban scene understanding tasks. Building upon the recently introduced SynDrone and FlyAware datasets, FlyAwareV2 introduces several new key contributions: (1) Multimodal data (RGB, depth, semantic labels) across diverse environmental conditions including varying weather and daytime; (2) Depth maps for real samples computed via state-of-the-art monocular depth estimation; (3) Benchmarks for RGB and multimodal semantic segmentation on standard architectures; (4) Studies on synthetic-to-real domain adaptation to assess the generalization capabilities of models trained on the synthetic data. With its rich set of annotations and environmental diversity, FlyAwareV2 provides a valuable resource for research on UAV-based 3D urban scene understanding. Dataset link: https://medialab.dei.unipd.it/paper_data/FlyAwareV2 • We introduce a multimodal dataset for aerial imaging across diverse environmental conditions including varying weather and daytime. • The dataset encompasses both real and synthetic data including depth maps for real samples computed via state-of-the-art monocular depth estimation. • Benchmarks for RGB and multimodal semantic segmentation on standard architectures are provided. • We also evaluate performances of synthetic-to-real domain adaptation to assess the generalization capabilities of models trained on the synthetic data.
Francesco Barbato, Matteo Caligiuri, Pietro Zanuttigh
Signal Process. Image Commun.3
2025 MultimodalStudio: A Heterogeneous Sensor Dataset and Framework for Neural Rendering across Multiple Imaging Modalities
abstract
Neural Radiance Fields (NeRF) have shown impressive performances in the rendering of 3D scenes from arbitrary viewpoints. While RGB images are widely preferred for training volume rendering models, the interest in other radiance modalities is also growing. However, the capability of the underlying implicit neural models to learn and transfer information across heterogeneous imaging modalities has seldom been explored, mostly due to the limited training data availability. For this purpose, we present MultimodalStudio (MMS): it encompasses MMS-DATA and MMS-FW. MMS-DATA is a multimodal multi-view dataset containing 32 scenes acquired with 5 different imaging modalities: RGB, monochrome, near-infrared, polarization and multi-spectral. MMS-FW is a novel modular multimodal NeRF framework designed to handle multimodal raw data and able to support an arbitrary number of multi-channel devices. Through extensive experiments, we demonstrate that MMS-FW trained on MMS-DATA can transfer information between different imaging modalities and produce higher quality renderings than using single modalities alone. We publicly release the dataset and the framework, to promote the research on multimodal volume rendering and beyond.
Federico Lincetto, Gianluca Agresti, Mattia Rossi, Pietro Zanuttigh
CVPR4
2025 LoRA.rar: Learning to Merge LoRAs via Hypernetworks for Subject-Style Conditioned Image Generation
abstract
Recent advancements in image generation models have enabled personalized image creation with both user-defined subjects (content) and styles. Prior works achieved personalization by merging corresponding low-rank adapters (LoRAs) through optimization-based methods, which are computationally demanding and unsuitable for real-time use on resource-constrained devices like smartphones. To address this, we introduce LoRA$.$rar, a method that not only improves image quality but also achieves a remarkable speedup of over $4000\times$ in the merging process. We collect a dataset of style and subject LoRAs and pre-train a hypernetwork on a diverse set of content-style LoRA pairs, learning an efficient merging strategy that generalizes to new, unseen content-style pairs, enabling fast, high-quality personalization. Moreover, we identify limitations in existing evaluation metrics for content-style quality and propose a new protocol using multimodal large language models (MLLMs) for more accurate assessment. Our method significantly outperforms the current state of the art in both content and style fidelity, as validated by MLLM assessments and human evaluations.
Donald Shenaj, Ondrej Bohdal, Mete Ozay, Pietro Zanuttigh, Umberto Michieli
ICCV4
2025 When Cars Meet Drones: Hyperbolic Federated Learning for Source-Free Domain Adaptation in Adverse Weather
abstract
In Federated Learning (FL), multiple clients collaboratively train a global model without sharing private data. In semantic segmentation, the Federated source Free Domain Adaptation (FFREEDA) setting is of particular interest, where clients undergo unsupervised training after supervised pretraining at the server side. While few recent works address FL for autonomous vehicles, intrinsic real-world challenges such as the presence of adverse weather conditions and the existence of different autonomous agents are still unexplored. To bridge this gap, we address both problems and introduce a new federated semantic segmentation setting where both car and drone clients co-exist and collaborate. Specifically, we propose a novel approach for this setting which exploits a batch-norm weather-aware strategy to dynamically adapt the model to the different weather conditions, while hyperbolic space prototypes are used to align the heterogeneous client representations. Finally, we introduce FLYAWARE, the first semantic segmentation dataset with adverse weather data for aerial vehicles.
Giulia Rizzoli, Matteo Caligiuri, Donald Shenaj, Francesco Barbato, Pietro Zanuttigh
WACV5
2025 From open-vocabulary to vocabulary-free semantic segmentation
abstract
Open-vocabulary semantic segmentation enables models to identify novel object categories beyond their training data. While this flexibility represents a significant advancement, current approaches still rely on manually specified class names as input, creating an inherent bottleneck in real-world applications. This work proposes a Vocabulary-Free Semantic Segmentation pipeline, eliminating the need for predefined class vocabularies. Specifically, we address the chicken-and-egg problem where users need knowledge of all potential objects within a scene to identify them, yet the purpose of segmentation is often to discover these objects. The proposed approach leverages Vision–Language Models to automatically recognize objects and generate appropriate class names, aiming to solve the challenge of class specification and naming quality. Through extensive experiments on several public datasets, we highlight the crucial role of the text encoder in model performance, particularly when the image text classes are paired with generated descriptions. Despite the challenges introduced by the sensitivity of the segmentation text encoder to false negatives within the class tagging process, which adds complexity to the task, we demonstrate that our fully automated pipeline significantly enhances vocabulary-free segmentation accuracy across diverse real-world scenarios. Code is available at https://github.com/klarareichard/open-vocab2free-seg . • Propose a novel two-stage pipeline using an image tagger and a class-specific decoder. • Setting a new benchmark for Vocabulary-Free Semantic Segmentation. • Show the impact of enriched text inputs on the encoder assuming a perfect tagger. • Analyze the influence of undetected objects and false detections on the segmentation.
Klara Reichard, Giulia Rizzoli, Stefano Gasperini, Lukas Hoyer, Pietro Zanuttigh, Nassir Navab, Federico Tombari
Pattern Recognit. Lett.5
2024 Learning from the Web: Language Drives Weakly-Supervised Incremental Learning for Semantic Segmentation
Chang Liu 0047, Giulia Rizzoli, Pietro Zanuttigh, Fu Li 0002
ECCV (17)3
2024 HoloADMM: High-Quality Holographic Complex Field Recovery
Mazen Mel, Paul Springer, Pietro Zanuttigh, Haitao Zhou, Alexander Gatto
ECCV (71)3
2024 Continual Road-Scene Semantic Segmentation Via Feature-Aligned Symmetric Multi-Modal Network
abstract
State-of-the-art multimodal semantic segmentation strategies combining LiDAR and color data are usually designed on top of asymmetric information-sharing schemes and assume that both modalities are always available. This strong assumption may not hold in real-world scenarios, where sensors are prone to failure or can face adverse conditions that make the acquired information unreliable. This problem is exacerbated when continual learning scenarios are considered since they have stringent data reliability constraints. In this work, we re-frame the task of multimodal semantic segmentation by enforcing a tightly coupled feature representation and a symmetric information-sharing scheme, which allows our approach to work even when one of the input modalities is missing. We also introduce an ad-hoc class-incremental continual learning scheme, proving our approach’s effectiveness and reliability even in safety-critical settings, such as autonomous driving. We evaluate our approach on the SemanticKITTI dataset, achieving impressive performances.
Francesco Barbato, Elena Camuffo, Simone Milani, Pietro Zanuttigh
ICIP4
2024 ALERT-Transformer: Bridging Asynchronous and Synchronous Machine Learning for Real-Time Event-based Spatio-Temporal Data
abstract
We seek to enable classic processing of continuous ultra-sparse spatiotemporal data generated by event-based sensors with dense machine learning models. We propose a novel hybrid pipeline composed of asynchronous sensing and synchronous processing that combines several ideas: (1) an embedding based on PointNet models -- the ALERT module -- that can continuously integrate new and dismiss old events thanks to a leakage mechanism, (2) a flexible readout of the embedded data that allows to feed any downstream model with always up-to-date features at any sampling rate, (3) exploiting the input sparsity in a patch-based approach inspired by Vision Transformer to optimize the efficiency of the method. These embeddings are then processed by a transformer model trained for object and gesture recognition. Using this approach, we achieve performances at the state-of-the-art with a lower latency than competitors. We also demonstrate that our asynchronous model can operate at any desired sampling rate.
Carmen Martin-Turrero, Maxence Bouvier, Manuel Breitenstein, Pietro Zanuttigh, Vincent Parret
ICML4
2024 Cross-Architecture Auxiliary Feature Space Translation for Efficient Few-Shot Personalized Object Detection
abstract
Recent years have seen object detection robotic systems deployed in several personal devices (e.g., home robots and appliances). This has highlighted a challenge in their design, i.e., they cannot efficiently update their knowledge to distinguish between general classes and user-specific instances (e.g., a dog vs. user’s dog). We refer to this challenging task as Instance-level Personalized Object Detection (IPOD). The personalization task requires many samples for model tuning and optimization in a centralized server, raising privacy concerns. An alternative is provided by approaches based on recent large-scale Foundation Models, but their compute costs preclude on-device applications. In our work we tackle both problems at the same time, designing a Few-Shot IPOD strategy called AuXFT. We introduce a conditional coarse-to-fine few-shot learner to refine the coarse predictions made by an efficient object detector, showing that using an off-the-shelf model leads to poor personalization due to neural collapse. Therefore, we introduce a Translator block that generates an auxiliary feature space where features generated by a self-supervised model (e.g., DINOv2) are distilled without impacting the performance of the detector. We validate AuXFT on three publicly available datasets and one in-house benchmark designed for the IPOD task, achieving remarkable gains in all considered scenarios with excellent time-complexity trade-off: AuXFT reaches a performance of 80% its upper bound at just 32% of the inference time, 13% of VRAM and 19% of the model size.
Francesco Barbato, Umberto Michieli, Ji Joong Moon, Pietro Zanuttigh, Mete Ozay
IROS4
2024 A Modular System for Enhanced Robustness of Multimedia Understanding Networks via Deep Parametric Estimation
abstract
Performance degradation caused by corrupted multimedia samples is a critical challenge for machine learning models. Previously, three groups of approaches have been proposed to tackle this issue: i) enhancer and denoiser modules to improve the quality of the noisy data, ii) data augmentation approaches, and iii) domain adaptation strategies. All have drawbacks limiting applicability; the first requires paired clean-corrupted data for training and has an high computational cost, while the others can only be used on the same task they were trained on. In this paper, we propose SyMPIE to solve these shortcomings, designing a small, modular, and efficient system to enhance input data for robust downstream multimedia understanding with minimal computational cost. Our SyMPIE is pre-trained on an upstream task/network that should not match the downstream ones and does not need paired clean-corrupted samples. Our key insight is that most input corruptions found in real-world tasks can be modeled through global operations on color channels of images or spatial filters with small kernels. We validate our approach on multiple datasets and tasks, such as image classification (on ImageNetC, ImageNetC-Bar, VizWiz, and a newly proposed mixed corruption benchmark named ImageNetC-mixed) and semantic segmentation (on Cityscapes, ACDC, and DarkZurich) with consistent improvements of about 5% relative accuracy gain across the board1.
Francesco Barbato, Umberto Michieli, Mehmet Kerim Yucel, Pietro Zanuttigh, Mete Ozay
MMSys4
2024 Learning With Style: Continual Semantic Segmentation Across Tasks and Domains
abstract
Deep learning models dealing with image understanding in real-world settings must be able to adapt to a wide variety of tasks across different domains. Domain adaptation and class incremental learning deal with domain and task variability separately, whereas their unified solution is still an open problem. We tackle both facets of the problem together, taking into account the semantic shift within both input and label spaces. We start by formally introducing continual learning under task and domain shift. Then, we address the proposed setup by using style transfer techniques to extend knowledge across domains when learning incremental tasks and a robust distillation framework to effectively recollect task knowledge under incremental domain shift. The devised framework (LwS, Learning with Style) is able to generalize incrementally acquired task knowledge across all the domains encountered, proving to be robust against catastrophic forgetting. Extensive experimental evaluation on multiple autonomous driving datasets shows how the proposed method outperforms existing approaches, which prove to be ill-equipped to deal with continual semantic segmentation under both task and domain shift.
Marco Toldo, Umberto Michieli, Pietro Zanuttigh
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Road scenes segmentation across different domains by disentangling latent representations
abstract
Abstract Deep learning models obtain impressive accuracy in road scene understanding; however, they need a large number of labeled samples for their training. Additionally, such models do not generalize well to environments where the statistical properties of data do not perfectly match those of training scenes, and this can be a significant problem for intelligent vehicles. Hence, domain adaptation approaches have been introduced to transfer knowledge acquired on a label-abundant source domain to a related label-scarce target domain. In this work, we design and carefully analyze multiple latent space-shaping regularization strategies that work together to reduce the domain shift. More in detail, we devise a feature clustering strategy to increase domain alignment, a feature perpendicularity constraint to space apart features belonging to different semantic classes, including those not present in the current batch, and a feature norm alignment strategy to separate active and inactive channels. In addition, we propose a novel evaluation metric to capture the relative performance of an adapted model with respect to supervised training. We validate our framework in driving scenarios, considering both synthetic-to-real and real-to-real adaptation, outperforming previous feature-level state-of-the-art methods on multiple road scenes benchmarks.
Francesco Barbato, Umberto Michieli, Marco Toldo, Pietro Zanuttigh
Vis. Comput.4
2024 End-to-end learning for joint depth and image reconstruction from diffracted rotation
abstract
Abstract Monocular depth estimation is an open challenge due to the ill-posed nature of the problem at hand. Deep learning techniques proved capable of producing acceptable depth estimation accuracy but the lack of robust depth cues within RGB images severally limits their performance. Coded aperture-based methods using phase and amplitude masks encode strong depth cues within 2D images by means of depth-dependent Point Spread Functions (PSFs) at the price of a reduced image quality. In this paper, we propose a novel end-to-end learning approach for depth from diffracted rotation. A phase mask that produces a Rotating Point Spread Function (RPSF) as a function of defocus is jointly optimized with the weights of a depth estimation neural network. To this aim, we introduce a differentiable physical model of the aperture mask and exploit an accurate simulation of the camera imaging pipeline. Our approach requires a significantly less complex model and less training data, yet it outperforms existing methods for monocular depth estimation on indoor benchmarks. In addition, we address the image degradation problem by incorporating a non-blind and nonuniform image deblurring module to recover the sharp all-in-focus image from its blurred counterpart.
Mazen Mel, Muhammad Siddiqui, Pietro Zanuttigh
Vis. Comput.3
2023 Exploiting Multiple Priors for Neural 3D Indoor Reconstruction
Federico Lincetto, Gianluca Agresti, Mattia Rossi, Pietro Zanuttigh
BMVC4
2023 DepthFormer: Multimodal Positional Encodings and Cross-Input Attention for Transformer-based Segmentation Networks
abstract
Most approaches for semantic segmentation use only information from color cameras to parse the scenes, yet recent advancements show that using depth data allows to further improve performances. In this work, we focus on transformer-based deep learning architectures, that have achieved state-of-the-art performances on the segmentation task, and we propose to employ depth information by embedding it in the positional encoding. Effectively, we extend the network to multimodal data without adding any parameters and in a natural way that exploits the strength of transformers’ self-attention modules. We also investigate the idea of performing cross-modality operations inside the attention module, swapping the key inputs between the depth and color branches. Our approach consistently improves performances on the Cityscapes benchmark.
Francesco Barbato, Giulia Rizzoli, Pietro Zanuttigh
ICASSP3
2023 Learning Across Domains and Devices: Style-Driven Source-Free Domain Adaptation in Clustered Federated Learning
abstract
Federated Learning (FL) has recently emerged as a possible way to tackle the domain shift in real-world Semantic Segmentation (SS) without compromising the private nature of the collected data. However, most of the existing works on FL unrealistically assume labeled data in the re-mote clients. Here we propose a novel task (FFreeDA) in which the clients’ data is unlabeled and the server accesses a source labeled dataset for pre-training only. To solve FFreeDA, we propose LADD, which leverages the knowledge of the pre-trained model by employing self-supervision with ad-hoc regularization techniques for local training and introducing a novel federated clustered aggregation scheme based on the clients’ style. Our experiments show that our algorithm is able to efficiently tackle the new task out-performing existing approaches. The code is available at https://github.com/Erosinho13/LADD.
Donald Shenaj, Eros Fanì, Marco Toldo, Debora Caldarola, Antonio Tavera, Umberto Michieli, Marco Ciccone, Pietro Zanuttigh, Barbara Caputo
WACV8
2023 SELMA: SEmantic Large-Scale Multimodal Acquisitions in Variable Weather, Daytime and Viewpoints
abstract
Accurate scene understanding from multiple sensors mounted on cars is a key requirement for autonomous driving systems. Nowadays, this task is mainly performed through data-hungry deep learning techniques that need very large amounts of data to be trained. Due to the high cost of performing segmentation labeling, many synthetic datasets have been proposed. However, most of them miss the multi-sensor nature of the data, and do not capture the significant changes introduced by the variation of daytime and weather conditions. To fill these gaps, we introduce SELMA, a novel synthetic dataset for semantic segmentation that contains more than 30K unique waypoints acquired from 24 different sensors including RGB, depth, semantic cameras and LiDARs, in 27 different weather and daytime conditions, for a total of more than 20M samples. SELMA is based on CARLA, an open-source simulator for generating synthetic data in autonomous driving scenarios, that we modified to increase the variability and the diversity in the scenes and class sets, and to align it with other benchmark datasets. As shown by the experimental evaluation, SELMA allows the efficient training of standard and multi-modal deep learning architectures, and achieves remarkable results on real-world data. SELMA is free and publicly available, thus supporting open science and research.
Paolo Testolina, Francesco Barbato, Umberto Michieli, Marco Giordani, Pietro Zanuttigh, Michele Zorzi
IEEE Trans. Intell. Transp. Syst.5
2022 Joint Reconstruction and Super Resolution of Hyper-Spectral CTIS Images
Mazen Mel, Alexander Gatto, Pietro Zanuttigh
BMVC3
2022 Edge-Aware Graph Matching Network for Part-Based Semantic Segmentation
abstract
Abstract Semantic segmentation of parts of objects is a marginally explored and challenging task in which multiple instances of objects and multiple parts within those objects must be recognized in an image. We introduce a novel approach (GMENet) for this task combining object-level context conditioning, part-level spatial relationships, and shape contour information. The first target is achieved by introducing a class-conditioning module that enforces class-level semantics when learning the part-level ones. Thus, intermediate-level features carry object-level prior to the decoding stage. To tackle part-level ambiguity and spatial relationships among parts we exploit an adjacency graph-based module that aims at matching the spatial relationships between parts in the ground truth and predicted maps. Last, we introduce an additional module to further leverage edges localization. Besides testing our framework on the already used Pascal-Part-58 and Pascal-Person-Part benchmarks, we further introduce two novel benchmarks for large-scale part parsing, i.e., a more challenging version of Pascal-Part with 108 classes and the ADE20K-Part benchmark with 544 parts. GMENet achieves state-of-the-art results in all the considered tasks and furthermore allows to improve object-level segmentation accuracy.
Umberto Michieli, Pietro Zanuttigh
Int. J. Comput. Vis.2
2022 Continual coarse-to-fine domain adaptation in semantic segmentation
Donald Shenaj, Francesco Barbato, Umberto Michieli, Pietro Zanuttigh
Image Vis. Comput.4
2022 Reframing control methods for parameters optimization in adversarial image generation
Qamar Alfalouji, Piergiorgio Sartor, Pietro Zanuttigh
Neural Networks3
2022 Unsupervised Domain Adaptation of Deep Networks for ToF Depth Refinement
abstract
Depth maps acquired with ToF cameras have a limited accuracy due to the high noise level and to the multi-path interference. Deep networks can be used for refining ToF depth, but their training requires real world acquisitions with ground truth, which is complex and expensive to collect. A possible workaround is to train networks on synthetic data, but the domain shift between the real and synthetic data reduces the performances. In this paper, we propose three approaches to perform unsupervised domain adaptation of a depth denoising network from synthetic to real data. These approaches are respectively acting at the input, at the feature and at the output level of the network. The first approach uses domain translation networks to transform labeled synthetic ToF data into a representation closer to real data, that is then used to train the denoiser. The second approach tries to align the network internal features related to synthetic and real data. The third approach uses an adversarial loss, implemented with a discriminator trained to recognize the ground truth statistic, to train the denoiser on unlabeled real data. Experimental results show that the considered approaches are able to outperform other state-of-the-art techniques and achieve superior denoising performances.
Gianluca Agresti, Henrik Schäfer, Piergiorgio Sartor, Yalcin Incesu, Pietro Zanuttigh
IEEE Trans. Pattern Anal. Mach. Intell.5
2021 Continual Semantic Segmentation via Repulsion-Attraction of Sparse and Disentangled Latent Representations
abstract
Deep neural networks suffer from the major limitation of catastrophic forgetting old tasks when learning new ones. In this paper we focus on class incremental continual learning in semantic segmentation, where new categories are made available over time while previous training data is not retained. The proposed continual learning scheme shapes the latent space to reduce forgetting whilst improving the recognition of novel classes. Our framework is driven by three novel components which we also combine on top of existing techniques effortlessly. First, prototypes matching enforces latent space consistency on old classes, constraining the encoder to produce similar latent representation for previously seen classes in the subsequent steps. Second, features sparsification allows to make room in the latent space to accommodate novel classes. Finally, contrastive learning is employed to cluster features according to their semantics while tearing apart those of different classes. Extensive evaluation on the Pascal VOC2012 and ADE20K datasets demonstrates the effectiveness of our approach, significantly outperforming state-of-the-art methods.
Umberto Michieli, Pietro Zanuttigh
CVPR2
2021 RECALL: Replay-based Continual Learning in Semantic Segmentation
abstract
Deep networks allow to obtain outstanding results in semantic segmentation, however they need to be trained in a single shot with a large amount of data. Continual learning settings where new classes are learned in incremental steps and previous training data is no longer available are challenging due to the catastrophic forgetting phenomenon. Existing approaches typically fail when several incremental steps are performed or in presence of a distribution shift of the background class. We tackle these issues by recreating no longer available data for the old classes and outlining a content inpainting scheme on the background class. We propose two sources for replay data. The first resorts to a generative adversarial network to sample from the class space of past learning steps. The second relies on web-crawled data to retrieve images containing examples of old classes from online databases. In both scenarios no samples of past steps are stored, thus avoiding privacy concerns. Replay data are then blended with new samples during the incremental steps. Our approach, RECALL, outperforms state-of-the-art methods.
Andrea Maracani, Umberto Michieli, Marco Toldo, Pietro Zanuttigh
ICCV4
2021 Unsupervised Domain Adaptation in Semantic Segmentation via Orthogonal and Clustered Embeddings
abstract
Deep learning frameworks allowed for a remarkable advancement in semantic segmentation, but the data hungry nature of convolutional networks has rapidly raised the demand for adaptation techniques able to transfer learned knowledge from label-abundant domains to unlabeled ones. In this paper we propose an effective Unsupervised Domain Adaptation (UDA) strategy, based on a feature clustering method that captures the different semantic modes of the feature distribution and groups features of the same class into tight and well-separated clusters. Furthermore, we introduce two novel learning objectives to enhance the discriminative clustering performance: an orthogonality loss forces spaced out individual representations to be orthogonal, while a sparsity loss reduces class-wise the number of active feature channels. The joint effect of these modules is to regularize the structure of the feature space. Extensive evaluations in the synthetic-to-real scenario show that we achieve state-of-the-art performance.
Marco Toldo, Umberto Michieli, Pietro Zanuttigh
WACV3
2021 Knowledge distillation for incremental learning in semantic segmentation
Umberto Michieli, Pietro Zanuttigh
Comput. Vis. Image Underst.2
2020 GMNet: Graph Matching Network for Large Scale Part Semantic Segmentation in the Wild
Umberto Michieli, Edoardo Borsato, Luca Rossi 0008, Pietro Zanuttigh
ECCV (8)4
2020 Semi-supervised Deep Learning Techniques for Spectrum Reconstruction
abstract
State-of-the-art approaches for the estimation of hyperspectral images (HSI) from RGB data are mostly based on deep learning techniques but due to the lack of training data their performances are limited to uncommon scenarios where a large hyperspectral database is available. In this work we present a family of novel deep learning schemes for hyperspectral data estimation able to work when the hyperspectral information at our disposal is limited. Firstly, we introduce a learning scheme exploiting a physical model based on the backward mapping to the RGB space and total variation regularization that can be trained with a limited amount of HSI images. Then, we propose a novel semi-supervised learning scheme able to work even with just a few pixels labeled with hyperspectral information. Finally, we show that the approach can be extended to a transfer learning scenario. The proposed techniques allow to reach impressive performances while requiring only some HSI images or just a few pixels for the training.
Adriano Simonetto, Pietro Zanuttigh, Vincent Parret, Piergiorgio Sartor, Alexander Gatto
ICPR2
2020 Unsupervised Domain Adaptation with Multiple Domain Discriminators and Adaptive Self-Training
abstract
Unsupervised Domain Adaptation (UDA) aims at improving the generalization capability of a model trained on a source domain to perform well on a target domain for which no labeled data is available. In this paper, we consider the semantic segmentation of urban scenes and we propose an approach to adapt a deep neural network trained on synthetic data to real scenes addressing the domain shift between the two different data distributions. We introduce a novel UDA framework where a standard supervised loss on labeled synthetic data is supported by an adversarial module and a self-training strategy aiming at aligning the two domain distributions. The adversarial module is driven by a couple of fully convolutional discriminators dealing with different domains: the first discriminates between ground truth and generated maps, while the second between segmentation maps coming from synthetic or real world data. The self-training module exploits the confidence estimated by the discriminators on unlabeled data to select the regions used to reinforce the learning process. Furthermore, the confidence is thresholded with an adaptive mechanism based on the per-class overall confidence. Experimental results prove the effectiveness of the proposed strategy in adapting a segmentation network trained on synthetic datasets like GTA5 and SYNTHIA, to real world datasets like Cityscapes and Mapillary.
Teo Spadotto, Marco Toldo, Umberto Michieli, Pietro Zanuttigh
ICPR4
2020 Unsupervised domain adaptation for mobile semantic segmentation based on cycle consistency and feature alignment
Marco Toldo, Umberto Michieli, Gianluca Agresti, Pietro Zanuttigh
Image Vis. Comput.4
2019 Unsupervised Domain Adaptation for ToF Data Denoising With Adversarial Learning
abstract
Time-of-Flight data is typically affected by a high level of noise and by artifacts due to Multi-Path Interference (MPI). While various traditional approaches for ToF data improvement have been proposed, machine learning techniques have seldom been applied to this task, mostly due to the limited availability of real world training data with depth ground truth. In this paper, we avoid to rely on labeled real data in the learning framework. A Coarse-Fine CNN, able to exploit multi-frequency ToF data for MPI correction, is trained on synthetic data with ground truth in a supervised way. In parallel, an adversarial learning strategy, based on the Generative Adversarial Networks (GAN) framework, is used to perform an unsupervised pixel-level domain adaptation from synthetic to real world data, exploiting unlabeled real world acquisitions. Experimental results demonstrate that the proposed approach is able to effectively denoise real world data and to outperform state-of-the-art techniques.
Gianluca Agresti, Henrik Schäfer, Piergiorgio Sartor, Pietro Zanuttigh
CVPR4
2018 Joint segmentation of color and depth data based on splitting and merging driven by surface fitting
Giampaolo Pagnutti, Pietro Zanuttigh
Image Vis. Comput.2
2018 Head-mounted gesture controlled interface for human-computer interaction
Alvise Memo, Pietro Zanuttigh
Multim. Tools Appl.2
2017 Deep learning for 3D shape classification from multiple depth maps
abstract
This paper proposes a novel approach for the classification of 3D shapes exploiting deep learning techniques. The proposed algorithm starts by constructing a set of depth maps by rendering the input 3D shape from different viewpoints. Then the depth maps are fed to a multi-branch Convolutional Neural Network. Each branch of the network takes in input one of the depth maps and produces a classification vector by using 5 convolutional layers of progressively reduced resolution. The various classification vectors are finally fed to a linear classifier that combines the outputs of the various branches and produces the final classification. Experimental results on the Princeton ModelNet database show how the proposed approach allows to obtain a high classification accuracy and outperforms several state-of-the-art approaches.
Pietro Zanuttigh, Ludovico Minto
ICIP1
2017 Segmentation and semantic labelling of RGBD data with convolutional neural networks and surface fitting
abstract
We present an approach for segmentation and semantic labelling of RGBD data exploiting together geometrical cues and deep learning techniques. An initial over‐segmentation is performed using spectral clustering and a set of non‐uniform rational B‐spline surfaces is fitted on the extracted segments. Then a convolutional neural network (CNN) receives in input colour and geometry data together with surface fitting parameters. The network is made of nine convolutional stages followed by a softmax classifier and produces a vector of descriptors for each sample. In the next step, an iterative merging algorithm recombines the output of the over‐segmentation into larger regions matching the various elements of the scene. The couples of adjacent segments with higher similarity according to the CNN features are candidate to be merged and the surface fitting accuracy is used to detect which couples of segments belong to the same surface. Finally, a set of labelled segments is obtained by combining the segmentation output with the descriptors from the CNN. Experimental results show how the proposed approach outperforms state‐of‐the‐art methods and provides an accurate segmentation and labelling.
Giampaolo Pagnutti, Ludovico Minto, Pietro Zanuttigh
IET Comput. Vis.3
2016 Reliable Fusion of ToF and Stereo Depth Driven by Confidence Measures
Giulio Marin, Pietro Zanuttigh, Stefano Mattoccia
ECCV (7)2
2016 Depth map coding with elastic contours and 3D surface prediction
abstract
Depth maps are typically made of smooth regions separated by sharp edges. Following this rationale, this paper presents a novel coding scheme where depth data is represented by a set of contours defining the various regions together with a compact representation of the values inside each region. The proposed coding scheme is based on elastic curves, which make possible to compactly represent the contours exploiting also the temporal consistency in different frames. A 3D surface prediction algorithm is then used to obtain an accurate estimation of the depth field from the coded contours and a subsampled version of the data. Finally, an ad-hoc coding strategy for the low resolution data and the prediction residuals is presented. Experimental results prove how the proposed approach is able to obtain a very high coding efficiency outperforming the HEVC coder at medium-low bitrates.
Marco Calemme, Pietro Zanuttigh, Simone Milani, Marco Cagnazzo, Béatrice Pesquet-Popescu
ICIP2
2016 3D scanning of cultural heritage with consumer depth cameras
Enrico Cappelletto, Pietro Zanuttigh, Guido M. Cortelazzo
Multim. Tools Appl.2
2016 Hand gesture recognition with jointly calibrated Leap Motion and depth sensor
Giulio Marin, Fabio Dominio, Pietro Zanuttigh
Multim. Tools Appl.3
2015 Compression of photo collections using geometrical information
abstract
This paper proposes a novel scheme for the joint compression of photo collections framing the same object or scene. The proposed approach starts by locating corresponding features in the various images and then exploits a Structure from Motion algorithm to estimate the geometric relationships between the various images and their viewpoints. Then it uses 3D information and warping to predict images one from the other. Furthermore, graph algorithms are used to compute minimum weight topologies and identify the ordering of the input images that maximizes the efficiency of prediction. The obtained data is fed to a modified HEVC coder to perform the compression. Experimental results show that the proposed scheme outperforms competing solutions and can be efficiently employed for the storage of large image collections in the virtual exploration of architectural landmarks or in photo sharing websites.
Simone Milani, Pietro Zanuttigh
ICME2
2015 Performance evaluation of the 1st and 2nd generation Kinect for multimedia applications
abstract
Microsoft Kinect had a key role in the development of consumer depth sensors being the device that brought depth acquisition to the mass market. Despite the success of this sensor, with the introduction of the second generation, Microsoft has completely changed the technology behind the sensor from structured light to Time-Of-Flight. This paper presents a comparison of the data provided by the first and second generation Kinect in order to explain the achievements that have been obtained with the switch of technology. After an accurate analysis of the accuracy of the two sensors under different conditions, two sample applications, i.e., 3D reconstruction and people tracking, are presented and used to compare the performance of the two sensors.
Simone Zennaro, Matteo Munaro, Simone Milani, Pietro Zanuttigh, Andrea Bernardi, Stefano Ghidoni, Emanuele Menegatti
ICME4
2015 Probabilistic ToF and Stereo Data Fusion Based on Mixed Pixels Measurement Models
abstract
This paper proposes a method for fusing data acquired by a ToF camera and a stereo pair based on a model for depth measurement by ToF cameras which accounts also for depth discontinuity artifacts due to the mixed pixel effect. Such model is exploited within both a ML and a MAP-MRF frameworks for ToF and stereo data fusion. The proposed MAP-MRF framework is characterized by site-dependent range values, a rather important feature since it can be used both to improve the accuracy and to decrease the computational complexity of standard MAP-MRF approaches. This paper, in order to optimize the site dependent global cost function characteristic of the proposed MAP-MRF approach, also introduces an extension to Loopy Belief Propagation which can be used in other contexts. Experimental data validate the proposed ToF measurements model and the effectiveness of the proposed fusion techniques.
Carlo Dal Mutto, Pietro Zanuttigh, Guido M. Cortelazzo
IEEE Trans. Pattern Anal. Mach. Intell.2
2014 Hand gesture recognition with leap motion and kinect devices
abstract
The recent introduction of novel acquisition devices like the Leap Motion and the Kinect allows to obtain a very informative description of the hand pose that can be exploited for accurate gesture recognition. This paper proposes a novel hand gesture recognition scheme explicitly targeted to Leap Motion data. An ad-hoc feature set based on the positions and orientation of the fingertips is computed and fed into a multi-class SVM classifier in order to recognize the performed gestures. A set of features is also extracted from the depth computed from the Kinect and combined with the Leap Motion ones in order to improve the recognition performance. Experimental results present a comparison between the accuracy that can be obtained from the two devices on a subset of the American Manual Alphabet and show how, by combining the two features sets, it is possible to achieve a very high accuracy in real-time.
Giulio Marin, Fabio Dominio, Pietro Zanuttigh
ICIP3
2014 Scene segmentation from depth and color data driven by surface fitting
abstract
Scene segmentation is a very challenging problem for which color information alone is often not sufficient. Recently the introduction of consumer depth cameras has opened the way to novel approaches exploiting depth data. This paper proposes a novel segmentation scheme that exploits the joint usage of color and depth data together with a 3D surface estimation scheme. Firstly a set of multi-dimensional vectors is built from color and geometry information and normalized cuts spectral clustering is applied to them in order to coarsely segment the scene. Then a NURBS model is fitted on each of the computed segments. The accuracy of the fitting is used as a measure of the plausibility that the segment represents a single surface or object. Segments that do not represent a single surface are split again into smaller regions and the process is iterated until the optimal segmentation is obtained. Experimental results show how the proposed method allows to obtain an accurate and reliable scene segmentation.
Giampaolo Pagnutti, Pietro Zanuttigh
ICIP2
2014 Combining multiple depth-based descriptors for hand gesture recognition
Fabio Dominio, Mauro Donadeo, Pietro Zanuttigh
Pattern Recognit. Lett.3
2013 Handheld scanning with 3D cameras
abstract
Novel real-time depth acquisition devices, like Microsoft Kinect, allow a very fast and simple acquisition of 3D views. These sensors have been used in many applications but their employment for 3D scanning purposes is a very challenging task due to the limited accuracy and reliability of their data. In this paper we present a 3D reconstruction pipeline explicitly targeted to the Kinect. The proposed scheme aims at obtaining a reliable reconstruction that is not affected by the limiting issues of these cameras and is at the same time simple and fast in order to allow to use the Kinect sensor as an handheld scanner. In order to achieve these targets a novel algorithm for the extraction of salient points exploiting both depth and color data is firstly used. The extracted points are then used within a modified version of the ICP algorithm that exploits both geometry and color distances to precisely align the views produced by the Kinect even when the geometry information is not sufficient to constrain the registration. Experimental results show how the proposed approach is able to produce reliable 3D reconstructions from the Kinect data.
Enrico Cappelletto, Pietro Zanuttigh, Guido M. Cortelazzo
MMSP2
2013 Combining color and shape descriptors for 3D model retrieval
Giuliano Pasqualotto, Pietro Zanuttigh, Guido M. Cortelazzo
Signal Process. Image Commun.2
2013 Scalable Coding of Depth Maps With R-D Optimized Embedding
abstract
Recent work on depth map compression has revealed the importance of incorporating a description of discontinuity boundary geometry into the compression scheme. We propose a novel compression strategy for depth maps that incorporates geometry information while achieving the goals of scalability and embedded representation. Our scheme involves two separate image pyramid structures, one for breakpoints and the other for sub-band samples produced by a breakpoint-adaptive transform. Breakpoints capture geometric attributes, and are amenable to scalable coding. We develop a rate-distortion optimization framework for determining the presence and precision of breakpoints in the pyramid representation. We employ a variation of the EBCOT scheme to produce embedded bit-streams for both the breakpoint and sub-band data. Compared to JPEG 2000, our proposed scheme enables the same the scalability features while achieving substantially improved rate-distortion performance at the higher bit-rate range and comparable performance at the lower rates.
Reji Mathew, David S. Taubman, Pietro Zanuttigh
IEEE Trans. Image Process.3
2012 Highly Scalable Coding of Depth Maps with Arc Breakpoints
abstract
Recent work highlights the importance of incorporating geometry information into the compression of depth maps. For many applications, features such as resolution scalability and embedded coding are also highly desirable. JPEG 2000 offers these scalability features but suffers from poor compression performance in the vicinity of strong discontinuities. We propose a novel compression strategy for depth maps that incorporates geometry information while retaining the highly scalable coding properties of JPEG 2000. Our scheme involves two separate image pyramid structures, one for arc breakpoints and other for sub-band samples produced by a breakpoint-adaptive transform. Breakpoints capture geometric attributes and are also amenable to scalable coding. We develop an R-D optimization framework for the breakpoint data. We also use a variation of the EBCOT scheme to produce embedded bit-streams for both the breakpoint and sub-band data, allowing them to be independently and incrementally sequenced based on R-D considerations.
Reji Mathew, Pietro Zanuttigh, David S. Taubman
DCC2
2012 Pairwise similarities for scene segmentation combining color and depth data
Filippo Bergamasco, Andrea Albarelli, Andrea Torsello, M. Favaro, Pietro Zanuttigh
ICPR5
2012 Scalable depth maps with R-D optimized embedding
abstract
Recent work has highlighted the importance of incorporating geometry information into the compression of depth maps. In prior approaches however the geometry information is not resolution scalable nor amenable to embedded coding. In this paper we propose a novel compression strategy for depth maps that incorporates geometry information while achieving the goals of scalability and embedded representation. Our scheme involves two separate image pyramid structures, one for breakpoints and other for sub-band samples produced by a breakpoint-adaptive transform. Breakpoints capture geometric attributes and are amenable to scalable coding. We develop an R-D optimization framework for the breakpoint data. We also use a variation of the EBCOT scheme to produce embedded bit-streams for both the breakpoint and sub-band data, allowing them to be independently and incrementally sequenced based on R-D considerations.
Reji Mathew, David S. Taubman, Pietro Zanuttigh
MMSP3
2011 Efficient depth map compression exploiting segmented color data
abstract
3D video representations usually associate to each view a depth map with the corresponding geometric information. Many compression schemes have been proposed for multi-view video and for depth data, but the exploitation of the correlation between the two representations to enhance compression performances is still an open research issue. This paper presents a novel compression scheme that exploits a segmentation of the color data to predict the shape of the different surfaces in the depth map. Then each segment is approximated with a parameterized plane. In case the approximation is sufficiently accurate for the target bit rate, the surface coefficients are compressed and transmitted. Otherwise, the region is coded using a standard H.264/AVC Intra coder. Experimental results show that the proposed scheme permits to outperformthe standardH.264/AVC Intra codec on depth data and can be effectively included into multi-view plus depth compression schemes.
Simone Milani, Pietro Zanuttigh, Marco Zamarin, Søren Forchhammer
ICME2
2010 A novel multi-view image coding scheme based on view-warping and 3D-DCT
Marco Zamarin, Simone Milani, Pietro Zanuttigh, Guido M. Cortelazzo
J. Vis. Commun. Image Represent.3
2008 Analysis of compressed depth and image streaming on unreliable networks
abstract
This paper explores the issues connected to the transmission of three dimensional scenes over unreliable networks such as the wireless ones. It analyzes the effect of the loss of compressed data packets in a typical image-based rendering scenario, where a set of compressed images together with the corresponding depth maps are transmitted and used to generate the views required from the user at client side. The different impact on the rendered views of the geometry and texture packets is analyzed in detail, taking into account also the position of the lost packets and the warping operation. Finally we will discuss how to exploit this results in the design of an efficient network protocol for the transmission of 3D models.
Pietro Zanuttigh, Andrea Zanella, Guido M. Cortelazzo
ISCC1
2006 Server Policies For Interactive Transmission Of 3D Scenes
abstract
We consider an interactive client-server application for remote browsing of 3D scenes. The information about texture and geometry is available at server side in the form of scalably compressed images and depthmaps, corresponding to a multitude of original image views. Image and depth components are both open to augmentation as more content becomes available. During the interactive browsing experience, the server allocates the available bandwidth between the delivery of new elements from the various original view bit-streams and new elements from the original geometry bit-streams. We propose a rate-distortion criterion to decide the best transmission policy for the server, since the best solution is not always to send the nearest original view image to the one which the client is rendering. We also outline how the JPIP standard for interactive transmission of JPEG2000 images can be exploited for remote exploration of 3D scenes
Pietro Zanuttigh, Nicola Brusco, David S. Taubman, Guido M. Cortelazzo
MMSP1
2006 A novel framework for the interactive transmission of 3D scenes
Pietro Zanuttigh, Nicola Brusco, David S. Taubman, Guido M. Cortelazzo
Signal Process. Image Commun.1
2005 Greedy non-linear approximation of the plenoptic function for interactive transmission of 3D scenes
abstract
We consider an interactive browsing environment, with greedy optimization of a current view, conditioned on the availability of previously transmitted information for other (possibly nearby) views, and subject to a transmission budget constraint. Texture information is available at a server in the form of scalably compressed images, corresponding to a multitude of original image views. Surface geometry is also represented at the server in a scalable fashion. At any point in the interactive browsing experience, the server must decide how to allocate transmission resources between the delivery of new elements from the various original view bit-streams and new elements from the geometry bit-stream. The proposed framework may be interpreted as a greedy strategy for non-linear approximation of the plenoptic function, since it considers both view sampling and rate-distortion criteria. We particularly elaborate upon a novel geometry- and distortion-sensitive strategy for blending the information available from different views at the client.
Pietro Zanuttigh, Nicola Brusco, David S. Taubman, Guido M. Cortelazzo
ICIP (1)1
2004 Content-based retrieval of 3D models based on multiple aspects
abstract
This paper presents an approach for aspect-based retrieval of 3D models or circular views of real objects. The objects in our database are described by a set of images, which are taken varying the pose of a constant angle along a given plane. A set of parameters is computed for each image, taking into account general features like bounding boxes, contours, and mass distribution. The number of parameters of each image is then reduced using principal component analysis, obtaining a low dimensional feature space. We assume that the query is in the same form of objects in our database, and thus it can undergo the same feature extraction process. The system then retrieves the most similar objects computing the L/sub 1/ distance in the features space. Preliminary results are presented with a database of 100 different objects, consisting of 2000 different views.
Nicola Orio, Pietro Zanuttigh, Guido M. Cortelazzo
MMSP2