Marco Toldo

dblp:24/11263 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
8since 2021 · last 2026
0000-0003-0954-1201ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 4 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RECALL+: Adversarial web-based replay for continual learning in semantic segmentation
Chang Liu 0047, Giulia Rizzoli, Francesco Barbato, Andrea Maracani, Marco Toldo, Umberto Michieli, Pietro Zanuttigh
Image Vis. Comput.5
2024 Learning With Style: Continual Semantic Segmentation Across Tasks and Domains
abstract
Deep learning models dealing with image understanding in real-world settings must be able to adapt to a wide variety of tasks across different domains. Domain adaptation and class incremental learning deal with domain and task variability separately, whereas their unified solution is still an open problem. We tackle both facets of the problem together, taking into account the semantic shift within both input and label spaces. We start by formally introducing continual learning under task and domain shift. Then, we address the proposed setup by using style transfer techniques to extend knowledge across domains when learning incremental tasks and a robust distillation framework to effectively recollect task knowledge under incremental domain shift. The devised framework (LwS, Learning with Style) is able to generalize incrementally acquired task knowledge across all the domains encountered, proving to be robust against catastrophic forgetting. Extensive experimental evaluation on multiple autonomous driving datasets shows how the proposed method outperforms existing approaches, which prove to be ill-equipped to deal with continual semantic segmentation under both task and domain shift.
Marco Toldo, Umberto Michieli, Pietro Zanuttigh
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Road scenes segmentation across different domains by disentangling latent representations
abstract
Abstract Deep learning models obtain impressive accuracy in road scene understanding; however, they need a large number of labeled samples for their training. Additionally, such models do not generalize well to environments where the statistical properties of data do not perfectly match those of training scenes, and this can be a significant problem for intelligent vehicles. Hence, domain adaptation approaches have been introduced to transfer knowledge acquired on a label-abundant source domain to a related label-scarce target domain. In this work, we design and carefully analyze multiple latent space-shaping regularization strategies that work together to reduce the domain shift. More in detail, we devise a feature clustering strategy to increase domain alignment, a feature perpendicularity constraint to space apart features belonging to different semantic classes, including those not present in the current batch, and a feature norm alignment strategy to separate active and inactive channels. In addition, we propose a novel evaluation metric to capture the relative performance of an adapted model with respect to supervised training. We validate our framework in driving scenarios, considering both synthetic-to-real and real-to-real adaptation, outperforming previous feature-level state-of-the-art methods on multiple road scenes benchmarks.
Francesco Barbato, Umberto Michieli, Marco Toldo, Pietro Zanuttigh
Vis. Comput.3
2023 Learning Across Domains and Devices: Style-Driven Source-Free Domain Adaptation in Clustered Federated Learning
abstract
Federated Learning (FL) has recently emerged as a possible way to tackle the domain shift in real-world Semantic Segmentation (SS) without compromising the private nature of the collected data. However, most of the existing works on FL unrealistically assume labeled data in the re-mote clients. Here we propose a novel task (FFreeDA) in which the clients’ data is unlabeled and the server accesses a source labeled dataset for pre-training only. To solve FFreeDA, we propose LADD, which leverages the knowledge of the pre-trained model by employing self-supervision with ad-hoc regularization techniques for local training and introducing a novel federated clustered aggregation scheme based on the clients’ style. Our experiments show that our algorithm is able to efficiently tackle the new task out-performing existing approaches. The code is available at https://github.com/Erosinho13/LADD.
Donald Shenaj, Eros Fanì, Marco Toldo, Debora Caldarola, Antonio Tavera, Umberto Michieli, Marco Ciccone, Pietro Zanuttigh, Barbara Caputo
WACV3
2023 Federated Learning via Attentive Margin of Semantic Feature Representations
abstract
Federated learning (FL) in Internet of Things (IoT) systems enables distributed model training using a large corpus of decentralized training data dispersed among multiple IoT clients. In this distributed setting, system and statistical heterogeneity, in the form of highly imbalanced, and nonindependent and identically distributed (non-i.i.d.) data stored on multiple devices, are likely to hinder model training. Existing methods aggregate models disregarding the internal representations being learned, which yet play an essential role to solve the pursued task, especially in the case of deep learning modules. To leverage feature representations in an FL framework, we introduce a method, called FedMargin, which computes client deviations using margins over feature representations learned on distributed data, and applies them to drive federated optimization via an attention mechanism. Local and aggregated margins are jointly exploited, taking into account local representation shift and representation discrepancy with the global model. In addition, we propose three methods to analyse statistical properties of feature representations learned in FL, in order to elucidate the relationship between accuracy, margins, and feature discrepancy of FL models. In experimental analyses, FedMargin demonstrates state-of-the-art accuracy and convergence rate across image classification and semantic segmentation benchmarks by enabling maximum margin training of FL models. Moreover, FedMargin reduces the uncertainty of predictions of FL models compared to the baseline. In this work, we also evaluate FL models on dense prediction tasks, such as semantic segmentation, proving the versatility of the proposed approach.
Umberto Michieli, Marco Toldo, Mete Ozay
IEEE Internet Things J.2
2022 Bring Evanescent Representations to Life in Lifelong Class Incremental Learning
abstract
In Class Incremental Learning (CIL), a classification model is progressively trained at each incremental step on an evolving dataset of new classes, while at the same time, it is required to preserve knowledge of all the classes ob-served so far. Prototypical representations can be lever-aged to model feature distribution for the past data and in-ject information of former classes in later incremental steps without resorting to stored exemplars. However, if not up-dated, those representations become increasingly outdated as the incremental learning progresses with new classes. To address the aforementioned problems, we propose a frame-work which aims to (i) model the semantic drift by learning the relationship between representations of past and novel classes among incremental steps, and (ii) estimate the feature drift, defined as the evolution of the represen-tations learned by models at each incremental step. Se-mantic and feature drifts are then jointly exploited to infer up-to-date representations of past classes (evanescent rep-resentations), and thereby infuse past knowledge into incre-mental training. We experimentally evaluate our framework achieving exemplar-free SotA results on multiple bench-marks. In the ablation study, we investigate nontrivial relationships between evanescent representations and models.
Marco Toldo, Mete Ozay
CVPR1
2021 RECALL: Replay-based Continual Learning in Semantic Segmentation
abstract
Deep networks allow to obtain outstanding results in semantic segmentation, however they need to be trained in a single shot with a large amount of data. Continual learning settings where new classes are learned in incremental steps and previous training data is no longer available are challenging due to the catastrophic forgetting phenomenon. Existing approaches typically fail when several incremental steps are performed or in presence of a distribution shift of the background class. We tackle these issues by recreating no longer available data for the old classes and outlining a content inpainting scheme on the background class. We propose two sources for replay data. The first resorts to a generative adversarial network to sample from the class space of past learning steps. The second relies on web-crawled data to retrieve images containing examples of old classes from online databases. In both scenarios no samples of past steps are stored, thus avoiding privacy concerns. Replay data are then blended with new samples during the incremental steps. Our approach, RECALL, outperforms state-of-the-art methods.
Andrea Maracani, Umberto Michieli, Marco Toldo, Pietro Zanuttigh
ICCV3
2021 Unsupervised Domain Adaptation in Semantic Segmentation via Orthogonal and Clustered Embeddings
abstract
Deep learning frameworks allowed for a remarkable advancement in semantic segmentation, but the data hungry nature of convolutional networks has rapidly raised the demand for adaptation techniques able to transfer learned knowledge from label-abundant domains to unlabeled ones. In this paper we propose an effective Unsupervised Domain Adaptation (UDA) strategy, based on a feature clustering method that captures the different semantic modes of the feature distribution and groups features of the same class into tight and well-separated clusters. Furthermore, we introduce two novel learning objectives to enhance the discriminative clustering performance: an orthogonality loss forces spaced out individual representations to be orthogonal, while a sparsity loss reduces class-wise the number of active feature channels. The joint effect of these modules is to regularize the structure of the feature space. Extensive evaluations in the synthetic-to-real scenario show that we achieve state-of-the-art performance.
Marco Toldo, Umberto Michieli, Pietro Zanuttigh
WACV1
2020 Unsupervised Domain Adaptation with Multiple Domain Discriminators and Adaptive Self-Training
abstract
Unsupervised Domain Adaptation (UDA) aims at improving the generalization capability of a model trained on a source domain to perform well on a target domain for which no labeled data is available. In this paper, we consider the semantic segmentation of urban scenes and we propose an approach to adapt a deep neural network trained on synthetic data to real scenes addressing the domain shift between the two different data distributions. We introduce a novel UDA framework where a standard supervised loss on labeled synthetic data is supported by an adversarial module and a self-training strategy aiming at aligning the two domain distributions. The adversarial module is driven by a couple of fully convolutional discriminators dealing with different domains: the first discriminates between ground truth and generated maps, while the second between segmentation maps coming from synthetic or real world data. The self-training module exploits the confidence estimated by the discriminators on unlabeled data to select the regions used to reinforce the learning process. Furthermore, the confidence is thresholded with an adaptive mechanism based on the per-class overall confidence. Experimental results prove the effectiveness of the proposed strategy in adapting a segmentation network trained on synthetic datasets like GTA5 and SYNTHIA, to real world datasets like Cityscapes and Mapillary.
Teo Spadotto, Marco Toldo, Umberto Michieli, Pietro Zanuttigh
ICPR2
2020 Unsupervised domain adaptation for mobile semantic segmentation based on cycle consistency and feature alignment
Marco Toldo, Umberto Michieli, Gianluca Agresti, Pietro Zanuttigh
Image Vis. Comput.1
2012 Low-delay peer-to-peer media streaming based on network coding over randomized multicast trees
abstract
We introduce randomized multicast trees (RMT), an overlay topology designed for low-delay media streaming using network coding. RMTs improve on tree-based overlays in terms of start-up delay. We develop a push-based streaming system that leverages network coding to efficiently distribute the information in the overlay without using buffer maps, followed by a short pull stage to recover from packet losses, and appropriate management procedures to handle ungraceful peers departures.We report performance results of the proposed system, and compare it with an optimized pull system, and with an existing peer-to-peer system employing network coding, showing a significant performance improvement in terms of delay and resiliency to peers’ dynamics and packet losses.
Marco Toldo, Enrico Magli
IEEE Trans. Multim.1
2010 A resilient and low-delay P2P streaming system based on network coding with random multicast trees
abstract
Network coding is known to provide increased throughput and reduced delay for communications over networks. In this paper we propose a peer-to-peer video streaming system that exploits network coding in order to achieve low start-up delay, high streaming rate, and high resiliency to peers' dynamics. In particular, we introduce the concept of random multicast trees as overlay topology. This topology offers all benefits of tree-based overlays, notably a short start-up delay, but is much more efficient at distributing data and recovering from ungraceful peers departures. We develop a push-based streaming system that leverages network coding to efficiently distribute the information in the overlay without using buffer maps. We show performance results of the proposed system and compare it with an optimized pull systems based on Coolstreaming, showing significant improvement.
Marco Toldo, Enrico Magli
MMSP1