EDBT 2026 Demo / reviewers in the wild / expert
Umberto Michieli
dblp:217/1611
· DBLP profile ↗
39ranked-venue papers
10as first author
35since 2021 · last 2026
0000-0003-2666-4342ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 7 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 5 first-author · 18 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | K-Merge: Online Continual Merging of Adapters for On-device Large Language ModelsabstractDonald Shenaj, Ondrej Bohdal, Taha Ceritli, Mete Ozay, Pietro Zanuttigh, Umberto Michieli. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Donald Shenaj, Ondrej Bohdal, Taha Ceritli, Mete Ozay, Pietro Zanuttigh, Umberto Michieli |
ACL (1) | 6 |
| 2026 | FedPromo: Federated Lightweight Proxy Models at the Edge for Fine-Grained Image Classification With Foundation ModelsabstractFederated learning (FL) is an established paradigm for training deep learning models on decentralized data, particularly relevant in Internet of Things (IoT) scenarios. However, as model sizes grow, conventional FL approaches require significant computational resources, which may not be feasible for resource-constrained IoT devices. We introduce FedPromo, a novel framework that enables efficient adaptation of classification heads trained in a federated way to large-scale foundation models (FMs) stored on a central server without requiring explicit data sharing. Instead of directly training the large model on client devices, FedPromo optimizes lightweight proxy models via FL, reducing computational overhead, energy consumption, and bandwidth usage while maintaining privacy. We evaluate our method in cross-domain fine-grained image classification via a two-stage process: server-side knowledge distillation (KD) aligns representations of a large-scale FM with those of a compact counterpart, then the compact model encoder is deployed to devices for local classifier learning. These classifiers are aggregated and transferred back to the FM, enabling learning of fine-grained representations without direct access to user data. Through novel regularization strategies, our framework enables decentralized multidomain learning, balancing performance, privacy, and resource efficiency for wide-scale IoT deployment. Extensive experiments on five image classification benchmarks and feature representation analysis demonstrate that FedPromo outperforms existing methods when employing mobile-targeted efficient architectures, making it suitable for IoT devices such as smart home devices, industrial sensors, and phones. Across domains, FedPromo enjoys average gains of 4.8% (9.1%) top-1 (top-5) accuracy on the clients and 7.5% (9.7%) on the server with respect to the best competitors. Matteo Caligiuri, Francesco Barbato, Donald Shenaj, Umberto Michieli, Pietro Zanuttigh |
IEEE Internet Things J. | 4 |
| 2026 | RECALL+: Adversarial web-based replay for continual learning in semantic segmentation
Chang Liu 0047, Giulia Rizzoli, Francesco Barbato, Andrea Maracani, Marco Toldo, Umberto Michieli, Pietro Zanuttigh |
Image Vis. Comput. | 6 |
| 2025 | DreamCache: Finetuning-Free Lightweight Personalized Image Generation via Feature CachingabstractPersonalized image generation requires text-to-image generative models that capture the core features of a reference subject to allow for controlled generation across different contexts. Existing methods face challenges due to complex training requirements, high inference costs, limited flexibility, or a combination of these issues. In this paper, we introduce DreamCache, a scalable approach for efficient and high-quality personalized image generation. By caching a small number of reference image features from a subset of layers and a single timestep of the pretrained diffusion denoiser, DreamCache enables dynamic modulation of the generated image features through lightweight, trained conditioning adapters. DreamCache achieves state-of-the-art image and text alignment, utilizing an order of magnitude fewer extra parameters, and is both more computationally effective and versatile than existing models.1 Emanuele Aiello, Umberto Michieli, Diego Valsesia, Mete Ozay, Enrico Magli |
CVPR | 2 |
| 2025 | Efficient Compositional Multi-tasking for On-device Large Language ModelsabstractAdapter parameters provide a mechanism to modify the behavior of machine learning models and have gained significant popularity in the context of large language models (LLMs) and generative AI.These parameters can be merged to support multiple tasks via a process known as task merging.However, prior work on merging in LLMs, particularly in natural language processing, has been limited to scenarios where each test example addresses only a single task.In this paper, we focus on on-device settings and study the problem of text-based compositional multi-tasking, where each test example involves the simultaneous execution of multiple tasks.For instance, generating a translated summary of a long text requires solving both translation and summarization tasks concurrently.To facilitate research in this setting, we propose a benchmark comprising four practically relevant compositional tasks.We also present an efficient method (Learnable Calibration) tailored for on-device applications, where computational resources are limited, emphasizing the need for solutions that are both resourceefficient and high-performing.Our contributions lay the groundwork for advancing the capabilities of LLMs in real-world multi-tasking scenarios, expanding their applicability to complex, resource-constrained use cases.Project page: Ondrej Bohdal, Mete Ozay, Ji Joong Moon, Kyeng-Hun Lee, Hyeonmok Ko, Umberto Michieli |
EMNLP | 6 |
| 2025 | HydraOpt: Navigating the Efficiency-Performance Trade-off of Adapter MergingabstractLarge language models (LLMs) often leverage adapters, such as low-rank-based adapters, to achieve strong performance on downstream tasks.However, storing a separate adapter for each task significantly increases memory requirements, posing a challenge for resourceconstrained environments such as mobile devices.Although model merging techniques can reduce storage costs, they typically result in substantial performance degradation.In this work, we introduce HydraOpt, a new model merging technique that capitalizes on the inherent similarities between the matrices of low-rank adapters.Unlike existing methods that produce a fixed trade-off between storage size and performance, HydraOpt allows us to navigate this spectrum of efficiency and performance.Our experiments show that Hy-draOpt significantly reduces storage size (48% reduction) compared to storing all adapters, while achieving competitive performance (0.2-1.8% drop).Furthermore, it outperforms existing merging techniques in terms of performance at the same or slightly worse storage efficiency. Taha Ceritli, Ondrej Bohdal, Mete Ozay, Ji Joong Moon, Kyeng-Hun Lee, Hyeonmok Ko, Umberto Michieli |
EMNLP | 7 |
| 2025 | Controllable Forgetting Mechanism for Few-Shot Class-Incremental LearningabstractClass-incremental learning in the context of limited personal labeled samples (few-shot) is critical for numerous real-world applications, such as smart home devices. A key challenge in these scenarios is balancing the trade-off between adapting to new, personalized classes and maintaining the performance of the model on the original, base classes. Fine-tuning the model on novel classes often leads to the phenomenon of catastrophic forgetting, where the accuracy of base classes declines unpredictably and significantly. In this paper, we propose a simple yet effective mechanism to address this challenge by controlling the trade-off between novel and base class accuracy. We specifically target the ultra-low-shot scenario, where only a single example is available per novel class. Our approach introduces a Novel Class Detection (NCD) rule, which adjusts the degree of forgetting a priori while simultaneously enhancing performance on novel classes. We demonstrate the versatility of our solution by applying it to state-of-the-art Few-Shot Class-Incremental Learning (FSCIL) methods, showing consistent improvements across different settings. To better quantify the trade-off between novel and base class performance, we introduce new metrics: NCR@2FOR and NCR@5FOR. Our approach achieves up to a 30% improvement in novel class accuracy on the CIFAR100 dataset (1-shot, 1 novel class) while maintaining a controlled base class forgetting rate of 2%. Kirill Paramonov, Mete Ozay, Eunju Yang, Ji Joong Moon, Umberto Michieli |
ICASSP | 5 |
| 2025 | LoRA.rar: Learning to Merge LoRAs via Hypernetworks for Subject-Style Conditioned Image GenerationabstractRecent advancements in image generation models have enabled personalized image creation with both user-defined subjects (content) and styles. Prior works achieved personalization by merging corresponding low-rank adapters (LoRAs) through optimization-based methods, which are computationally demanding and unsuitable for real-time use on resource-constrained devices like smartphones. To address this, we introduce LoRA$.$rar, a method that not only improves image quality but also achieves a remarkable speedup of over $4000\times$ in the merging process. We collect a dataset of style and subject LoRAs and pre-train a hypernetwork on a diverse set of content-style LoRA pairs, learning an efficient merging strategy that generalizes to new, unseen content-style pairs, enabling fast, high-quality personalization. Moreover, we identify limitations in existing evaluation metrics for content-style quality and propose a new protocol using multimodal large language models (MLLMs) for more accurate assessment. Our method significantly outperforms the current state of the art in both content and style fidelity, as validated by MLLM assessments and human evaluations. Donald Shenaj, Ondrej Bohdal, Mete Ozay, Pietro Zanuttigh, Umberto Michieli |
ICCV | 5 |
| 2025 | Continual Error Correction on Low-Resource DevicesabstractThe proliferation of AI models in everyday devices has highlighted a critical challenge: prediction errors that degrade user experience. While existing solutions focus on error detection, they rarely provide efficient correction mechanisms, especially for resource-constrained devices. We present a novel system enabling users to correct AI misclassifications through few-shot learning, requiring minimal computational resources and storage. Our approach combines server-side foundation model training with on-device prototype-based classification, enabling efficient error correction through prototype updates rather than model retraining. The system consists of two key components: (1) a server-side pipeline that leverages knowledge distillation to transfer robust feature representations from foundation models to device-compatible architectures, and (2) a device-side mechanism that enables ultra-efficient error correction through prototype adaptation. We demonstrate our system's effectiveness on both image classification and object detection tasks, achieving over 50% error correction in one-shot scenarios on Food-101 and Flowers-102 datasets while maintaining minimal forgetting (less than 0.02%) and negligible computational overhead. Our implementation, validated through an Android demonstration app, proves the system's practicality in real-world scenarios. Kirill Paramonov, Mete Ozay, Aristeidis Mystakidis, Nikolaos Tsalikidis, Dimitrios Sotos, Anastasios Drosou, Dimitrios Tzovaras, Kiseok Chang, Sangdok Mo, Namwoong Kim, Woojong Yoo, Ji Joong Moon, Umberto Michieli |
MMSys | 14 |
| 2025 | Brain network science modelling of sparse neural networks enables Transformers and LLMs to perform as fully connectedabstractThis study aims to enlarge our current knowledge on the application of brain-inspired network science principles for training artificial neural networks (ANNs) with sparse connectivity. Dynamic sparse training (DST) emulates the synaptic turnover of real brain networks, reducing the computational demands of training and inference in ANNs. However, existing DST methods face difficulties in maintaining peak performance at high connectivity sparsity levels. The Cannistraci-Hebb training (CHT) is a brain-inspired method that is used in DST for growing synaptic connectivity in sparse neural networks. CHT leverages a gradient-free, topology-driven link regrowth mechanism, which has been shown to achieve ultra-sparse (1\% connectivity or lower) advantage across various tasks compared to fully connected networks. Yet, CHT suffers two main drawbacks: (i) its time complexity is $\mathcal{O}(N\cdot d^3)$- N node network size, d node degree - hence it can be efficiently applied only to ultra-sparse networks. (ii) it rigidly selects top link prediction scores, which is inappropriate for the early training epochs, when the network topology presents many unreliable connections. Here, we design the first brain-inspired network model - termed bipartite receptive field (BRF) - to initialize the connectivity of sparse artificial neural networks. Then, we propose a matrix multiplication GPU-friendly approximation of the CH link predictor, which reduces the computational complexity to $\mathcal{O}(N^3)$, enabling a fast implementation of link prediction in large-scale models. Moreover, we introduce the Cannistraci-Hebb training soft rule (CHTs), which adopts a flexible strategy for sampling connections in both link removal and regrowth, balancing the exploration and exploitation of network topology. Additionally, we propose a sigmoid-based gradual density decay strategy, leading to an advanced framework referred to as CHTss. Empirical results show that BRF offers performance advantages over previous network science models. Using 1\% of connections, CHTs outperforms fully connected networks in MLP architectures on visual classification tasks, compressing some networks to less than 30\% of the nodes. Using 5\% of the connections, CHTss outperforms fully connected networks in two Transformer-based machine translation tasks. Finally, with only 30\% of the connections, both CHTs and CHTss achieve superior performance over other dynamic sparse training methods, and perform on par with—or even surpass—their fully connected counterparts in language modeling across various sparsity levels within the LLaMA model family. The code is available at: https://github.com/biomedical-cybernetics/Cannistraci-Hebb-Training-Soft-Rule-. Yingtao Zhang, Diego Cerretti, Jialin Zhao 0004, Ziheng Liao, Umberto Michieli, Carlo V. Cannistraci |
NeurIPS | 6 |
| 2025 | Adaptive Cannistraci-Hebb Network Automata Modelling of Complex Networks for Path-based Link PredictionabstractMany complex networks have partially observed or evolving connectivity, making link prediction a fundamental task.
Topological link prediction infers missing links using only network topology, with applications in social, biological, and technological systems.
The Cannistraci-Hebb (CH) theory provides a topological formulation of Hebbian learning, grounded on two pillars:
(1) the **minimization of external links** within local communities, and
(2) the **path-based definition of local communities** that capture homophilic (similarity-driven) interactions via paths of length 2 and synergetic (diversity-driven) interactions via paths of length 3.
Building on this, we introduce the Cannistraci-Hebb Adaptive (CHA) network automata, an adaptive learning machine that automatically selects the optimal CH rule and path length to model each network.
CHA unifies theoretical interpretability and data-driven adaptivity, bridging physics-inspired network science and machine intelligence.
Across 1,269 networks from 14 domains, CHA consistently surpasses state-of-the-art methods—including SPM, SBM, graph embedding methods, and message-passing graph neural networks—while revealing the mechanistic principles governing link formation. Our code is available at https://github.com/biomedical-cybernetics/Cannistraci_Hebb_network_automata. Jialin Zhao 0004, Alessandro Muscoloni, Umberto Michieli, Yingtao Zhang, Carlo V. Cannistraci |
NeurIPS | 3 |
| 2025 | Learning From Mistakes: Self-Regularizing Hierarchical Representations in Point Cloud Semantic SegmentationabstractRecent advances in autonomous robotic technologies have highlighted the growing need for precise environmental analysis. Point cloud semantic segmentation has gained attention to accomplish fine-grained scene understanding by acting directly on raw content provided by sensors. Recent solutions showed how different learning techniques can be used to improve the performance of the model, without any architectural or dataset change. Following this trend, we present a coarse-to-fine setup that LEArns from classification mistaKes (LEAK) derived from a standard model. First, classes are clustered into macro groups according to mutual prediction errors; then, the learning process is regularized by: (1) aligning class-conditional prototypical feature representation for both fine and coarse classes, (2) weighting instances with a per-class fairness index. Our LEAK approach is very general and can be seamlessly applied on top of any segmentation architecture; indeed, experimental results showed that it enables state-of-the-art performances on different architectures, datasets and tasks, while ensuring more balanced class-wise results and faster convergence. Elena Camuffo, Umberto Michieli, Simone Milani |
IEEE Trans. Multim. | 2 |
| 2024 | HOP to the Next Tasks and Domains for Continual Learning in NLPabstractContinual Learning (CL) aims to learn a sequence of problems (i.e., tasks and domains) by transferring knowledge acquired on previous problems, whilst avoiding forgetting of past ones. Different from previous approaches which focused on CL for one NLP task or domain in a specific use-case, in this paper, we address a more general CL setting to learn from a sequence of problems in a unique framework. Our method, HOP, permits to hop across tasks and domains by addressing the CL problem along three directions: (i) we employ a set of adapters to generalize a large pre-trained model to unseen problems, (ii) we compute high-order moments over the distribution of embedded representations to distinguish independent and correlated statistics across different tasks and domains, (iii) we process this enriched information with auxiliary heads specialized for each end problem. Extensive experimental campaign on 4 NLP applications, 5 benchmarks and 2 CL setups demonstrates the effectiveness of our HOP. Umberto Michieli, Mete Ozay |
AAAI | 1 |
| 2024 | Deep Neural Network Models Trained with a Fixed Random Classifier Transfer Better Across DomainsabstractThe recently discovered Neural collapse (NC) phenomenon states that the last-layer weights of Deep Neural Networks (DNN), converge to the so-called Equiangular Tight Frame (ETF) simplex, at the terminal phase of their training. This ETF geometry is equivalent to vanishing within-class variability of the last layer activations. Inspired by NC properties, we explore in this paper the transferability of DNN models trained with their last layer weight fixed according to ETF. This enforces class separation by eliminating class covariance information, effectively providing implicit regularization. We show that DNN models trained with such a fixed classifier significantly improve transfer performance, particularly on out-of-domain datasets. On a broad range of fine-grained image classification datasets, our approach outperforms i) baseline methods that do not perform any covariance regularization (up to 22%), as well as ii) methods that explicitly whiten covariance of activations throughout training (up to 19%). Our findings suggest that DNNs trained with fixed ETF classifiers offer a powerful mechanism for improving transfer learning across domains. Hafiz Tiomoko Ali, Umberto Michieli, Ji Joong Moon, Mete Ozay |
ICASSP | 2 |
| 2024 | FFT-Based Selection and Optimization of Statistics for Robust Recognition of Severely Corrupted ImagesabstractImproving model robustness in case of corrupted images is among the key challenges to enable robust vision systems on smart devices, such as robotic agents. Particularly, robust test-time performance is imperative for most of the applications. This paper presents a novel approach to improve robustness of any classification model, especially on severely corrupted images. Our method (FROST) employs high-frequency features to detect input image corruption type, and select layer-wise feature normalization statistics. FROST provides the state-of-the-art results for different models and datasets, outperforming competitors on ImageNet-C by up to 37.1% relative gain, improving baseline of 40.9% mCE on severe corruptions. Elena Camuffo, Umberto Michieli, Ji Joong Moon, Mete Ozay |
ICASSP | 2 |
| 2024 | Object-Conditioned Bag of Instances for Few-Shot Personalized Instance RecognitionabstractNowadays, users demand for increased personalization of vision systems to localize and identify personal instances of objects (e.g., my dog rather than dog) from a few-shot dataset only. Despite outstanding results of deep networks on classical label-abundant benchmarks (e.g., those of the latest YOLOv8 model for standard object detection), they struggle to maintain within-class variability to represent different instances rather than object categories only. We construct an Object-conditioned Bag of Instances (OBoI) based on multiorder statistics of extracted features, where generic object detection models are extended to search and identify personal instances from the OBoI’s metric space, without need for backpropagation. By relying on multi-order statistics, OBoI achieves consistent superior accuracy in distinguishing different instances. In the results, we achieve 77.1% personal object recognition accuracy in case of 18 personal instances, showing about 12% relative gain over the state of the art. Umberto Michieli, Ji Joong Moon, Mete Ozay |
ICASSP | 1 |
| 2024 | Cross-Architecture Auxiliary Feature Space Translation for Efficient Few-Shot Personalized Object DetectionabstractRecent years have seen object detection robotic systems deployed in several personal devices (e.g., home robots and appliances). This has highlighted a challenge in their design, i.e., they cannot efficiently update their knowledge to distinguish between general classes and user-specific instances (e.g., a dog vs. user’s dog). We refer to this challenging task as Instance-level Personalized Object Detection (IPOD). The personalization task requires many samples for model tuning and optimization in a centralized server, raising privacy concerns. An alternative is provided by approaches based on recent large-scale Foundation Models, but their compute costs preclude on-device applications. In our work we tackle both problems at the same time, designing a Few-Shot IPOD strategy called AuXFT. We introduce a conditional coarse-to-fine few-shot learner to refine the coarse predictions made by an efficient object detector, showing that using an off-the-shelf model leads to poor personalization due to neural collapse. Therefore, we introduce a Translator block that generates an auxiliary feature space where features generated by a self-supervised model (e.g., DINOv2) are distilled without impacting the performance of the detector. We validate AuXFT on three publicly available datasets and one in-house benchmark designed for the IPOD task, achieving remarkable gains in all considered scenarios with excellent time-complexity trade-off: AuXFT reaches a performance of 80% its upper bound at just 32% of the inference time, 13% of VRAM and 19% of the model size. Francesco Barbato, Umberto Michieli, Ji Joong Moon, Pietro Zanuttigh, Mete Ozay |
IROS | 2 |
| 2024 | Enhanced Model Robustness to Input Corruptions by Per-corruption Adaptation of Normalization StatisticsabstractDeveloping a reliable vision system is a fundamental challenge for robotic technologies (e.g., indoor service robots and outdoor autonomous robots) which can ensure reliable navigation even in challenging environments such as adverse weather conditions (e.g., fog, rain), poor lighting conditions (e.g., over/under exposure), or sensor degradation (e.g., blurring, noise), and can guarantee high performance in safety-critical functions. Current solutions proposed to improve model robustness usually rely on generic data augmentation techniques or employ costly test-time adaptation methods. In addition, most approaches focus on addressing a single vision task (typically, image recognition) utilising synthetic data. In this paper, we introduce Per-corruption Adaptation of Normalization statistics (PAN) to enhance the model robustness of vision systems. Our approach entails three key components: (i) a corruption type identification module, (ii) dynamic adjustment of normalization layer statistics based on identified corruption type, and (iii) real-time update of these statistics according to input data. PAN can integrate seamlessly with any convolutional model for enhanced accuracy in several robot vision tasks. In our experiments, PAN obtains robust performance improvement on challenging real-world corrupted image datasets (e.g., OpenLoris, ExDark, ACDC), where most of the current solutions tend to fail. Moreover, PAN outperforms the baseline models by 20-30% on synthetic benchmarks in object recognition tasks. Elena Camuffo, Umberto Michieli, Simone Milani, Ji Joong Moon, Mete Ozay |
IROS | 2 |
| 2024 | Swiss DINO: Efficient and Versatile Vision Framework for On-device Personal Object SearchabstractIn this paper, we address a recent trend in robotic home appliances to include vision systems on personal devices, capable of personalizing the appliances on the fly. In particular, we formulate and address an important technical task of personal object search, which involves localization and identification of personal items of interest on images captured by robotic appliances, with each item referenced only by a few annotated images. The task is crucial for robotic home appliances and mobile systems, which need to process personal visual scenes or to operate with particular personal objects (e.g., for grasping or navigation). In practice, personal object search presents two main technical challenges. First, a robot vision system needs to be able to distinguish between many fine-grained classes, in the presence of occlusions and clutter. Second, the strict resource requirements for the on-device system restrict the usage of most state-of-the-art methods for few-shot learning and often prevent on-device adaptation. In this work, we propose Swiss DINO: a simple yet effective framework for one-shot personal object search based on the recent DINOv2 transformer model, which was shown to have strong zero-shot generalization properties. Swiss DINO handles challenging on-device personalized scene understanding requirements and does not require any adaptation training. We show significant improvement (up to 55%) in segmentation and recognition accuracy compared to the common lightweight solutions, and significant footprint reduction of backbone inference time (up to 100×) and GPU consumption (up to 10×) compared to the heavy transformer-based solutions1. Kirill Paramonov, Jia-Xing Zhong, Umberto Michieli, Ji Joong Moon, Mete Ozay |
IROS | 3 |
| 2024 | A Modular System for Enhanced Robustness of Multimedia Understanding Networks via Deep Parametric EstimationabstractPerformance degradation caused by corrupted multimedia samples is a critical challenge for machine learning models. Previously, three groups of approaches have been proposed to tackle this issue: i) enhancer and denoiser modules to improve the quality of the noisy data, ii) data augmentation approaches, and iii) domain adaptation strategies. All have drawbacks limiting applicability; the first requires paired clean-corrupted data for training and has an high computational cost, while the others can only be used on the same task they were trained on. In this paper, we propose SyMPIE to solve these shortcomings, designing a small, modular, and efficient system to enhance input data for robust downstream multimedia understanding with minimal computational cost. Our SyMPIE is pre-trained on an upstream task/network that should not match the downstream ones and does not need paired clean-corrupted samples. Our key insight is that most input corruptions found in real-world tasks can be modeled through global operations on color channels of images or spatial filters with small kernels. We validate our approach on multiple datasets and tasks, such as image classification (on ImageNetC, ImageNetC-Bar, VizWiz, and a newly proposed mixed corruption benchmark named ImageNetC-mixed) and semantic segmentation (on Cityscapes, ACDC, and DarkZurich) with consistent improvements of about 5% relative accuracy gain across the board1. Francesco Barbato, Umberto Michieli, Mehmet Kerim Yucel, Pietro Zanuttigh, Mete Ozay |
MMSys | 2 |
| 2024 | Learning With Style: Continual Semantic Segmentation Across Tasks and DomainsabstractDeep learning models dealing with image understanding in real-world settings must be able to adapt to a wide variety of tasks across different domains. Domain adaptation and class incremental learning deal with domain and task variability separately, whereas their unified solution is still an open problem. We tackle both facets of the problem together, taking into account the semantic shift within both input and label spaces. We start by formally introducing continual learning under task and domain shift. Then, we address the proposed setup by using style transfer techniques to extend knowledge across domains when learning incremental tasks and a robust distillation framework to effectively recollect task knowledge under incremental domain shift. The devised framework (LwS, Learning with Style) is able to generalize incrementally acquired task knowledge across all the domains encountered, proving to be robust against catastrophic forgetting. Extensive experimental evaluation on multiple autonomous driving datasets shows how the proposed method outperforms existing approaches, which prove to be ill-equipped to deal with continual semantic segmentation under both task and domain shift. Marco Toldo, Umberto Michieli, Pietro Zanuttigh |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Road scenes segmentation across different domains by disentangling latent representationsabstractAbstract Deep learning models obtain impressive accuracy in road scene understanding; however, they need a large number of labeled samples for their training. Additionally, such models do not generalize well to environments where the statistical properties of data do not perfectly match those of training scenes, and this can be a significant problem for intelligent vehicles. Hence, domain adaptation approaches have been introduced to transfer knowledge acquired on a label-abundant source domain to a related label-scarce target domain. In this work, we design and carefully analyze multiple latent space-shaping regularization strategies that work together to reduce the domain shift. More in detail, we devise a feature clustering strategy to increase domain alignment, a feature perpendicularity constraint to space apart features belonging to different semantic classes, including those not present in the current batch, and a feature norm alignment strategy to separate active and inactive channels. In addition, we propose a novel evaluation metric to capture the relative performance of an adapted model with respect to supervised training. We validate our framework in driving scenarios, considering both synthetic-to-real and real-to-real adaptation, outperforming previous feature-level state-of-the-art methods on multiple road scenes benchmarks. Francesco Barbato, Umberto Michieli, Marco Toldo, Pietro Zanuttigh |
Vis. Comput. | 2 |
| 2023 | Fully Automated Scan-to-BIM Via Point Cloud Instance SegmentationabstractDigital reconstruction through Building Information Models (BIM) is a valuable methodology for documenting and analyzing existing buildings. Its pipeline starts with geometric acquisition. (e.g., via photogrammetry or laser scanning) for accurate point cloud collection. However, the acquired data are noisy and unstructured, and the creation of a semantically-meaningful BIM representation requires a huge computational effort, as well as expensive and time-consuming human annotations. In this paper, we propose a fully automated scan-to-BIM pipeline. The approach relies on: (i) our dataset (HePIC), acquired from two large buildings and annotated at a point-wise semantic level based on existent BIM models; (ii) a novel ad hoc deep network (BIM-Net++) for semantic segmentation, whose output is then processed to extract instance information necessary to recreate BIM objects; (iii) novel model pre-training and class re-weighting to eliminate the need for a large amount of labeled data and human intervention. D. Campagnolo, Elena Camuffo, Umberto Michieli, Paolo Borin, Simone Milani |
ICIP | 3 |
| 2023 | A Model for Every User and Budget: Label-Free and Personalized Mixed-Precision Quantization
Edward Fish, Umberto Michieli, Mete Ozay |
INTERSPEECH | 2 |
| 2023 | Online Continual Learning in Keyword Spotting for Low-Resource Devices via Pooling High-Order Temporal Statistics
Umberto Michieli, Pablo Peso Parada, Mete Ozay |
INTERSPEECH | 1 |
| 2023 | Online Continual Learning for Robust Indoor Object RecognitionabstractVision systems mounted on home robots need to interact with unseen classes in changing environments. Robots have limited computational resources, labelled data and storage capability. These requirements pose some unique challenges: models should adapt without forgetting past knowledge in a data- and parameter-efficient way. We characterize the problem as few-shot (FS) online continual learning (OCL), where robotic agents learn from a non-repeated stream of few-shot data updating only a few model parameters. Additionally, such models experience variable conditions at test time, where objects may appear in different poses (e.g., horizontal or vertical) and environments (e.g., day or night). To improve robustness of CL agents, we propose RobOCLe, which; 1) constructs an enriched feature space computing high order statistical moments from the embedded features of samples; and 2) computes similarity between high order statistics of the samples on the enriched feature space, and predicts their class labels. We evaluate robustness of CL models to train/test augmentations in various cases. We show that different moments allow RobOCLe to capture different properties of deformations, providing higher robustness with no decrease of inference speed. Umberto Michieli, Mete Ozay |
IROS | 1 |
| 2023 | Learning Across Domains and Devices: Style-Driven Source-Free Domain Adaptation in Clustered Federated LearningabstractFederated Learning (FL) has recently emerged as a possible way to tackle the domain shift in real-world Semantic Segmentation (SS) without compromising the private nature of the collected data. However, most of the existing works on FL unrealistically assume labeled data in the re-mote clients. Here we propose a novel task (FFreeDA) in which the clients’ data is unlabeled and the server accesses a source labeled dataset for pre-training only. To solve FFreeDA, we propose LADD, which leverages the knowledge of the pre-trained model by employing self-supervision with ad-hoc regularization techniques for local training and introducing a novel federated clustered aggregation scheme based on the clients’ style. Our experiments show that our algorithm is able to efficiently tackle the new task out-performing existing approaches. The code is available at https://github.com/Erosinho13/LADD. Donald Shenaj, Eros Fanì, Marco Toldo, Debora Caldarola, Antonio Tavera, Umberto Michieli, Marco Ciccone, Pietro Zanuttigh, Barbara Caputo |
WACV | 6 |
| 2023 | Federated Learning via Attentive Margin of Semantic Feature RepresentationsabstractFederated learning (FL) in Internet of Things (IoT) systems enables distributed model training using a large corpus of decentralized training data dispersed among multiple IoT clients. In this distributed setting, system and statistical heterogeneity, in the form of highly imbalanced, and nonindependent and identically distributed (non-i.i.d.) data stored on multiple devices, are likely to hinder model training. Existing methods aggregate models disregarding the internal representations being learned, which yet play an essential role to solve the pursued task, especially in the case of deep learning modules. To leverage feature representations in an FL framework, we introduce a method, called FedMargin, which computes client deviations using margins over feature representations learned on distributed data, and applies them to drive federated optimization via an attention mechanism. Local and aggregated margins are jointly exploited, taking into account local representation shift and representation discrepancy with the global model. In addition, we propose three methods to analyse statistical properties of feature representations learned in FL, in order to elucidate the relationship between accuracy, margins, and feature discrepancy of FL models. In experimental analyses, FedMargin demonstrates state-of-the-art accuracy and convergence rate across image classification and semantic segmentation benchmarks by enabling maximum margin training of FL models. Moreover, FedMargin reduces the uncertainty of predictions of FL models compared to the baseline. In this work, we also evaluate FL models on dense prediction tasks, such as semantic segmentation, proving the versatility of the proposed approach. Umberto Michieli, Marco Toldo, Mete Ozay |
IEEE Internet Things J. | 1 |
| 2023 | SELMA: SEmantic Large-Scale Multimodal Acquisitions in Variable Weather, Daytime and ViewpointsabstractAccurate scene understanding from multiple sensors mounted on cars is a key requirement for autonomous driving systems. Nowadays, this task is mainly performed through data-hungry deep learning techniques that need very large amounts of data to be trained. Due to the high cost of performing segmentation labeling, many synthetic datasets have been proposed. However, most of them miss the multi-sensor nature of the data, and do not capture the significant changes introduced by the variation of daytime and weather conditions. To fill these gaps, we introduce SELMA, a novel synthetic dataset for semantic segmentation that contains more than 30K unique waypoints acquired from 24 different sensors including RGB, depth, semantic cameras and LiDARs, in 27 different weather and daytime conditions, for a total of more than 20M samples. SELMA is based on CARLA, an open-source simulator for generating synthetic data in autonomous driving scenarios, that we modified to increase the variability and the diversity in the scenes and class sets, and to align it with other benchmark datasets. As shown by the experimental evaluation, SELMA allows the efficient training of standard and multi-modal deep learning architectures, and achieves remarkable results on real-world data. SELMA is free and publicly available, thus supporting open science and research. Paolo Testolina, Francesco Barbato, Umberto Michieli, Marco Giordani, Pietro Zanuttigh, Michele Zorzi |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Edge-Aware Graph Matching Network for Part-Based Semantic SegmentationabstractAbstract Semantic segmentation of parts of objects is a marginally explored and challenging task in which multiple instances of objects and multiple parts within those objects must be recognized in an image. We introduce a novel approach (GMENet) for this task combining object-level context conditioning, part-level spatial relationships, and shape contour information. The first target is achieved by introducing a class-conditioning module that enforces class-level semantics when learning the part-level ones. Thus, intermediate-level features carry object-level prior to the decoding stage. To tackle part-level ambiguity and spatial relationships among parts we exploit an adjacency graph-based module that aims at matching the spatial relationships between parts in the ground truth and predicted maps. Last, we introduce an additional module to further leverage edges localization. Besides testing our framework on the already used Pascal-Part-58 and Pascal-Person-Part benchmarks, we further introduce two novel benchmarks for large-scale part parsing, i.e., a more challenging version of Pascal-Part with 108 classes and the ADE20K-Part benchmark with 544 parts. GMENet achieves state-of-the-art results in all the considered tasks and furthermore allows to improve object-level segmentation accuracy. Umberto Michieli, Pietro Zanuttigh |
Int. J. Comput. Vis. | 1 |
| 2022 | Continual coarse-to-fine domain adaptation in semantic segmentation
Donald Shenaj, Francesco Barbato, Umberto Michieli, Pietro Zanuttigh |
Image Vis. Comput. | 3 |
| 2021 | Continual Semantic Segmentation via Repulsion-Attraction of Sparse and Disentangled Latent RepresentationsabstractDeep neural networks suffer from the major limitation of catastrophic forgetting old tasks when learning new ones. In this paper we focus on class incremental continual learning in semantic segmentation, where new categories are made available over time while previous training data is not retained. The proposed continual learning scheme shapes the latent space to reduce forgetting whilst improving the recognition of novel classes. Our framework is driven by three novel components which we also combine on top of existing techniques effortlessly. First, prototypes matching enforces latent space consistency on old classes, constraining the encoder to produce similar latent representation for previously seen classes in the subsequent steps. Second, features sparsification allows to make room in the latent space to accommodate novel classes. Finally, contrastive learning is employed to cluster features according to their semantics while tearing apart those of different classes. Extensive evaluation on the Pascal VOC2012 and ADE20K datasets demonstrates the effectiveness of our approach, significantly outperforming state-of-the-art methods. Umberto Michieli, Pietro Zanuttigh |
CVPR | 1 |
| 2021 | RECALL: Replay-based Continual Learning in Semantic SegmentationabstractDeep networks allow to obtain outstanding results in semantic segmentation, however they need to be trained in a single shot with a large amount of data. Continual learning settings where new classes are learned in incremental steps and previous training data is no longer available are challenging due to the catastrophic forgetting phenomenon. Existing approaches typically fail when several incremental steps are performed or in presence of a distribution shift of the background class. We tackle these issues by recreating no longer available data for the old classes and outlining a content inpainting scheme on the background class. We propose two sources for replay data. The first resorts to a generative adversarial network to sample from the class space of past learning steps. The second relies on web-crawled data to retrieve images containing examples of old classes from online databases. In both scenarios no samples of past steps are stored, thus avoiding privacy concerns. Replay data are then blended with new samples during the incremental steps. Our approach, RECALL, outperforms state-of-the-art methods. Andrea Maracani, Umberto Michieli, Marco Toldo, Pietro Zanuttigh |
ICCV | 2 |
| 2021 | Unsupervised Domain Adaptation in Semantic Segmentation via Orthogonal and Clustered EmbeddingsabstractDeep learning frameworks allowed for a remarkable advancement in semantic segmentation, but the data hungry nature of convolutional networks has rapidly raised the demand for adaptation techniques able to transfer learned knowledge from label-abundant domains to unlabeled ones. In this paper we propose an effective Unsupervised Domain Adaptation (UDA) strategy, based on a feature clustering method that captures the different semantic modes of the feature distribution and groups features of the same class into tight and well-separated clusters. Furthermore, we introduce two novel learning objectives to enhance the discriminative clustering performance: an orthogonality loss forces spaced out individual representations to be orthogonal, while a sparsity loss reduces class-wise the number of active feature channels. The joint effect of these modules is to regularize the structure of the feature space. Extensive evaluations in the synthetic-to-real scenario show that we achieve state-of-the-art performance. Marco Toldo, Umberto Michieli, Pietro Zanuttigh |
WACV | 2 |
| 2021 | Knowledge distillation for incremental learning in semantic segmentation
Umberto Michieli, Pietro Zanuttigh |
Comput. Vis. Image Underst. | 1 |
| 2020 | GMNet: Graph Matching Network for Large Scale Part Semantic Segmentation in the Wild
Umberto Michieli, Edoardo Borsato, Luca Rossi 0008, Pietro Zanuttigh |
ECCV (8) | 1 |
| 2020 | Unsupervised Domain Adaptation with Multiple Domain Discriminators and Adaptive Self-TrainingabstractUnsupervised Domain Adaptation (UDA) aims at improving the generalization capability of a model trained on a source domain to perform well on a target domain for which no labeled data is available. In this paper, we consider the semantic segmentation of urban scenes and we propose an approach to adapt a deep neural network trained on synthetic data to real scenes addressing the domain shift between the two different data distributions. We introduce a novel UDA framework where a standard supervised loss on labeled synthetic data is supported by an adversarial module and a self-training strategy aiming at aligning the two domain distributions. The adversarial module is driven by a couple of fully convolutional discriminators dealing with different domains: the first discriminates between ground truth and generated maps, while the second between segmentation maps coming from synthetic or real world data. The self-training module exploits the confidence estimated by the discriminators on unlabeled data to select the regions used to reinforce the learning process. Furthermore, the confidence is thresholded with an adaptive mechanism based on the per-class overall confidence. Experimental results prove the effectiveness of the proposed strategy in adapting a segmentation network trained on synthetic datasets like GTA5 and SYNTHIA, to real world datasets like Cityscapes and Mapillary. Teo Spadotto, Marco Toldo, Umberto Michieli, Pietro Zanuttigh |
ICPR | 3 |
| 2020 | Unsupervised domain adaptation for mobile semantic segmentation based on cycle consistency and feature alignment
Marco Toldo, Umberto Michieli, Gianluca Agresti, Pietro Zanuttigh |
Image Vis. Comput. | 2 |
| 2018 | Game Theoretic Analysis of Road User Safety Scenarios Involving Autonomous VehiclesabstractInteractions between pedestrians, cyclists, and human-driven vehicles have become a major concern for traffic safety over the years. The upcoming age of autonomous vehicles will further raise major problems on whether self-driving cars can accurately avoid accidents; on the other hand, usability issues arise on whether human-driven cars and pedestrians can dominate the road at the expense of the autonomous vehicles that will be programmed to avoid accidents. This paper proposes some game theoretical models applied to traffic scenarios, where the strategic interaction between a pedestrian and an autonomous vehicle is analyzed. The games have been simulated to demonstrate the theoretical analysis and the predicted behaviors. These investigations can shed new lights on how urban traffic regulations and inter-vehicle communications could be required to allow for a general improved management of traffic in the presence of autonomous vehicles. Umberto Michieli, Leonardo Badia |
PIMRC | 1 |