Francesco Barbato

dblp:289/7301 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Federated Medical Image Classification Under Class and Domain Imbalance Exploiting Synthetic Sample Generation
Martina Pavan, Matteo Caligiuri, Francesco Barbato, Pietro Zanuttigh
ICPR (8)3
2026 FedPromo: Federated Lightweight Proxy Models at the Edge for Fine-Grained Image Classification With Foundation Models
abstract
Federated learning (FL) is an established paradigm for training deep learning models on decentralized data, particularly relevant in Internet of Things (IoT) scenarios. However, as model sizes grow, conventional FL approaches require significant computational resources, which may not be feasible for resource-constrained IoT devices. We introduce FedPromo, a novel framework that enables efficient adaptation of classification heads trained in a federated way to large-scale foundation models (FMs) stored on a central server without requiring explicit data sharing. Instead of directly training the large model on client devices, FedPromo optimizes lightweight proxy models via FL, reducing computational overhead, energy consumption, and bandwidth usage while maintaining privacy. We evaluate our method in cross-domain fine-grained image classification via a two-stage process: server-side knowledge distillation (KD) aligns representations of a large-scale FM with those of a compact counterpart, then the compact model encoder is deployed to devices for local classifier learning. These classifiers are aggregated and transferred back to the FM, enabling learning of fine-grained representations without direct access to user data. Through novel regularization strategies, our framework enables decentralized multidomain learning, balancing performance, privacy, and resource efficiency for wide-scale IoT deployment. Extensive experiments on five image classification benchmarks and feature representation analysis demonstrate that FedPromo outperforms existing methods when employing mobile-targeted efficient architectures, making it suitable for IoT devices such as smart home devices, industrial sensors, and phones. Across domains, FedPromo enjoys average gains of 4.8% (9.1%) top-1 (top-5) accuracy on the clients and 7.5% (9.7%) on the server with respect to the best competitors.
Matteo Caligiuri, Francesco Barbato, Donald Shenaj, Umberto Michieli, Pietro Zanuttigh
IEEE Internet Things J.2
2026 RECALL+: Adversarial web-based replay for continual learning in semantic segmentation
Chang Liu 0047, Giulia Rizzoli, Francesco Barbato, Andrea Maracani, Marco Toldo, Umberto Michieli, Pietro Zanuttigh
Image Vis. Comput.3
2026 FlyAwareV2: A multimodal cross-domain UAV dataset for urban scene understanding
abstract
The development of computer vision algorithms for Unmanned Aerial Vehicle (UAV) applications in urban environments heavily relies on the availability of large-scale datasets with accurate annotations. However, collecting and annotating real-world UAV data is extremely challenging and costly. To address this limitation, we present FlyAwareV2, a novel multimodal dataset encompassing both real and synthetic UAV imagery tailored for urban scene understanding tasks. Building upon the recently introduced SynDrone and FlyAware datasets, FlyAwareV2 introduces several new key contributions: (1) Multimodal data (RGB, depth, semantic labels) across diverse environmental conditions including varying weather and daytime; (2) Depth maps for real samples computed via state-of-the-art monocular depth estimation; (3) Benchmarks for RGB and multimodal semantic segmentation on standard architectures; (4) Studies on synthetic-to-real domain adaptation to assess the generalization capabilities of models trained on the synthetic data. With its rich set of annotations and environmental diversity, FlyAwareV2 provides a valuable resource for research on UAV-based 3D urban scene understanding. Dataset link: https://medialab.dei.unipd.it/paper_data/FlyAwareV2 • We introduce a multimodal dataset for aerial imaging across diverse environmental conditions including varying weather and daytime. • The dataset encompasses both real and synthetic data including depth maps for real samples computed via state-of-the-art monocular depth estimation. • Benchmarks for RGB and multimodal semantic segmentation on standard architectures are provided. • We also evaluate performances of synthetic-to-real domain adaptation to assess the generalization capabilities of models trained on the synthetic data.
Francesco Barbato, Matteo Caligiuri, Pietro Zanuttigh
Signal Process. Image Commun.1
2025 When Cars Meet Drones: Hyperbolic Federated Learning for Source-Free Domain Adaptation in Adverse Weather
abstract
In Federated Learning (FL), multiple clients collaboratively train a global model without sharing private data. In semantic segmentation, the Federated source Free Domain Adaptation (FFREEDA) setting is of particular interest, where clients undergo unsupervised training after supervised pretraining at the server side. While few recent works address FL for autonomous vehicles, intrinsic real-world challenges such as the presence of adverse weather conditions and the existence of different autonomous agents are still unexplored. To bridge this gap, we address both problems and introduce a new federated semantic segmentation setting where both car and drone clients co-exist and collaborate. Specifically, we propose a novel approach for this setting which exploits a batch-norm weather-aware strategy to dynamically adapt the model to the different weather conditions, while hyperbolic space prototypes are used to align the heterogeneous client representations. Finally, we introduce FLYAWARE, the first semantic segmentation dataset with adverse weather data for aerial vehicles.
Giulia Rizzoli, Matteo Caligiuri, Donald Shenaj, Francesco Barbato, Pietro Zanuttigh
WACV4
2024 Continual Road-Scene Semantic Segmentation Via Feature-Aligned Symmetric Multi-Modal Network
abstract
State-of-the-art multimodal semantic segmentation strategies combining LiDAR and color data are usually designed on top of asymmetric information-sharing schemes and assume that both modalities are always available. This strong assumption may not hold in real-world scenarios, where sensors are prone to failure or can face adverse conditions that make the acquired information unreliable. This problem is exacerbated when continual learning scenarios are considered since they have stringent data reliability constraints. In this work, we re-frame the task of multimodal semantic segmentation by enforcing a tightly coupled feature representation and a symmetric information-sharing scheme, which allows our approach to work even when one of the input modalities is missing. We also introduce an ad-hoc class-incremental continual learning scheme, proving our approach’s effectiveness and reliability even in safety-critical settings, such as autonomous driving. We evaluate our approach on the SemanticKITTI dataset, achieving impressive performances.
Francesco Barbato, Elena Camuffo, Simone Milani, Pietro Zanuttigh
ICIP1
2024 Cross-Architecture Auxiliary Feature Space Translation for Efficient Few-Shot Personalized Object Detection
abstract
Recent years have seen object detection robotic systems deployed in several personal devices (e.g., home robots and appliances). This has highlighted a challenge in their design, i.e., they cannot efficiently update their knowledge to distinguish between general classes and user-specific instances (e.g., a dog vs. user’s dog). We refer to this challenging task as Instance-level Personalized Object Detection (IPOD). The personalization task requires many samples for model tuning and optimization in a centralized server, raising privacy concerns. An alternative is provided by approaches based on recent large-scale Foundation Models, but their compute costs preclude on-device applications. In our work we tackle both problems at the same time, designing a Few-Shot IPOD strategy called AuXFT. We introduce a conditional coarse-to-fine few-shot learner to refine the coarse predictions made by an efficient object detector, showing that using an off-the-shelf model leads to poor personalization due to neural collapse. Therefore, we introduce a Translator block that generates an auxiliary feature space where features generated by a self-supervised model (e.g., DINOv2) are distilled without impacting the performance of the detector. We validate AuXFT on three publicly available datasets and one in-house benchmark designed for the IPOD task, achieving remarkable gains in all considered scenarios with excellent time-complexity trade-off: AuXFT reaches a performance of 80% its upper bound at just 32% of the inference time, 13% of VRAM and 19% of the model size.
Francesco Barbato, Umberto Michieli, Ji Joong Moon, Pietro Zanuttigh, Mete Ozay
IROS1
2024 A Modular System for Enhanced Robustness of Multimedia Understanding Networks via Deep Parametric Estimation
abstract
Performance degradation caused by corrupted multimedia samples is a critical challenge for machine learning models. Previously, three groups of approaches have been proposed to tackle this issue: i) enhancer and denoiser modules to improve the quality of the noisy data, ii) data augmentation approaches, and iii) domain adaptation strategies. All have drawbacks limiting applicability; the first requires paired clean-corrupted data for training and has an high computational cost, while the others can only be used on the same task they were trained on. In this paper, we propose SyMPIE to solve these shortcomings, designing a small, modular, and efficient system to enhance input data for robust downstream multimedia understanding with minimal computational cost. Our SyMPIE is pre-trained on an upstream task/network that should not match the downstream ones and does not need paired clean-corrupted samples. Our key insight is that most input corruptions found in real-world tasks can be modeled through global operations on color channels of images or spatial filters with small kernels. We validate our approach on multiple datasets and tasks, such as image classification (on ImageNetC, ImageNetC-Bar, VizWiz, and a newly proposed mixed corruption benchmark named ImageNetC-mixed) and semantic segmentation (on Cityscapes, ACDC, and DarkZurich) with consistent improvements of about 5% relative accuracy gain across the board1.
Francesco Barbato, Umberto Michieli, Mehmet Kerim Yucel, Pietro Zanuttigh, Mete Ozay
MMSys1
2024 Road scenes segmentation across different domains by disentangling latent representations
abstract
Abstract Deep learning models obtain impressive accuracy in road scene understanding; however, they need a large number of labeled samples for their training. Additionally, such models do not generalize well to environments where the statistical properties of data do not perfectly match those of training scenes, and this can be a significant problem for intelligent vehicles. Hence, domain adaptation approaches have been introduced to transfer knowledge acquired on a label-abundant source domain to a related label-scarce target domain. In this work, we design and carefully analyze multiple latent space-shaping regularization strategies that work together to reduce the domain shift. More in detail, we devise a feature clustering strategy to increase domain alignment, a feature perpendicularity constraint to space apart features belonging to different semantic classes, including those not present in the current batch, and a feature norm alignment strategy to separate active and inactive channels. In addition, we propose a novel evaluation metric to capture the relative performance of an adapted model with respect to supervised training. We validate our framework in driving scenarios, considering both synthetic-to-real and real-to-real adaptation, outperforming previous feature-level state-of-the-art methods on multiple road scenes benchmarks.
Francesco Barbato, Umberto Michieli, Marco Toldo, Pietro Zanuttigh
Vis. Comput.1
2023 DepthFormer: Multimodal Positional Encodings and Cross-Input Attention for Transformer-based Segmentation Networks
abstract
Most approaches for semantic segmentation use only information from color cameras to parse the scenes, yet recent advancements show that using depth data allows to further improve performances. In this work, we focus on transformer-based deep learning architectures, that have achieved state-of-the-art performances on the segmentation task, and we propose to employ depth information by embedding it in the positional encoding. Effectively, we extend the network to multimodal data without adding any parameters and in a natural way that exploits the strength of transformers’ self-attention modules. We also investigate the idea of performing cross-modality operations inside the attention module, swapping the key inputs between the depth and color branches. Our approach consistently improves performances on the Cityscapes benchmark.
Francesco Barbato, Giulia Rizzoli, Pietro Zanuttigh
ICASSP1
2023 SELMA: SEmantic Large-Scale Multimodal Acquisitions in Variable Weather, Daytime and Viewpoints
abstract
Accurate scene understanding from multiple sensors mounted on cars is a key requirement for autonomous driving systems. Nowadays, this task is mainly performed through data-hungry deep learning techniques that need very large amounts of data to be trained. Due to the high cost of performing segmentation labeling, many synthetic datasets have been proposed. However, most of them miss the multi-sensor nature of the data, and do not capture the significant changes introduced by the variation of daytime and weather conditions. To fill these gaps, we introduce SELMA, a novel synthetic dataset for semantic segmentation that contains more than 30K unique waypoints acquired from 24 different sensors including RGB, depth, semantic cameras and LiDARs, in 27 different weather and daytime conditions, for a total of more than 20M samples. SELMA is based on CARLA, an open-source simulator for generating synthetic data in autonomous driving scenarios, that we modified to increase the variability and the diversity in the scenes and class sets, and to align it with other benchmark datasets. As shown by the experimental evaluation, SELMA allows the efficient training of standard and multi-modal deep learning architectures, and achieves remarkable results on real-world data. SELMA is free and publicly available, thus supporting open science and research.
Paolo Testolina, Francesco Barbato, Umberto Michieli, Marco Giordani, Pietro Zanuttigh, Michele Zorzi
IEEE Trans. Intell. Transp. Syst.2
2022 Continual coarse-to-fine domain adaptation in semantic segmentation
Donald Shenaj, Francesco Barbato, Umberto Michieli, Pietro Zanuttigh
Image Vis. Comput.2