VLDB 2026 Research / reviewers in the wild / expert
Antonio Carta
dblp:178/6658
· DBLP profile ↗
15ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0002-0003-2323ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 4 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Replay-free Online Continual Learning with Self-Supervised MultiPatchesabstractOnline Continual Learning (OCL) methods train a model on a non-stationary data stream where only a few examples are available at a time, often leveraging replay strategies.However, usage of replay is sometimes forbidden, especially in applications with strict privacy regulations.Therefore, we propose Continual MultiPatches (CMP), an effective plugin for existing OCL self-supervised learning strategies that avoids the use of replay samples.CMP generates multiple patches from a single example and projects them into a shared feature space, where patches coming from the same example are pushed together without collapsing into a single point.CMP surpasses replay and other SSL-based strategies on OCL streams, challenging the role of replay as a go-to solution for self-supervised OCL.Code available at https://github.com/giacomo-cgn/cmp . Giovanni A. Cignoni, Andrea Cossu, Alexandra Gomez-Villa, Joost van de Weijer 0001, Antonio Carta |
ESANN | 5 |
| 2025 | Online Curvature-Aware Replay: Leveraging 2nd Order Information for Online Continual Learning
Edoardo Urettini, Antonio Carta |
ICML | 2 |
| 2024 | GAS-Norm: Score-Driven Adaptive Normalization for Non-Stationary Time Series Forecasting in Deep LearningabstractDespite their popularity, deep neural networks (DNNs) applied to time series forecasting often fail to beat simpler statistical models. One of the main causes of this suboptimal performance is the data non-stationarity present in many processes. In particular, changes in the mean and variance of the input data can disrupt the predictive capability of a DNN. In this paper, we first show how DNN forecasting models fail in simple non-stationary settings. We then introduce GAS-Norm, a novel methodology for adaptive time series normalization and forecasting based on the combination of a Generalized Autoregressive Score (GAS) model and a Deep Neural Network. The GAS approach encompasses a score-driven family of models that estimate the mean and variance at each new observation, providing updated statistics to normalize the input data of the deep model. The output of the DNN is eventually denormalized using the statistics forecasted by the GAS model, resulting in a hybrid approach that leverages the strengths of both statistical modeling and deep learning. The adaptive normalization improves the performance of the model in non-stationary settings. The proposed approach is model-agnostic and can be applied to any DNN forecasting model. To empirically validate our proposal, we first compare GAS-Norm with other state-of-the-art normalization methods. We then combine it with state-of-the-art DNN forecasting models and test them on real-world datasets from the Monash open-access forecasting repository. Results show that deep forecasting models improve their performance in 21 out of 25 settings when combined with GAS-Norm compared to other normalization methods. Edoardo Urettini, Daniele Atzeni, Reshawn Ramjattan, Antonio Carta |
CIKM | 4 |
| 2024 | Projected Latent Distillation for Data-Agnostic Consolidation in distributed continual learningabstractIn continual learning applications on-the-edge multiple self-centered devices (SCD) learn different local tasks independently, with each SCD only optimizing its own task. Can we achieve (almost) zero-cost collaboration between different devices? We formalize this problem as a Distributed Continual Learning (DCL) scenario, where SCDs greedily adapt to their own local tasks and a separate continual learning (CL) model perform a sparse and asynchronous consolidation step that combines the SCD models sequentially into a single multi-task model without using the original data. Unfortunately, current CL methods are not directly applicable to this scenario. We propose Data-Agnostic Consolidation (DAC), a novel double knowledge distillation method which performs distillation in the latent space via a novel Projected Latent Distillation loss. Experimental results show that DAC enables forward transfer between SCDs and reaches state-of-the-art accuracy on Split CIFAR100, CORe50 and Split TinyImageNet, both in single device and distributed CL scenarios. Somewhat surprisingly, a single out-of-distribution image is sufficient as the only source of data for DAC. Antonio Carta, Andrea Cossu, Vincenzo Lomonaco, Davide Bacciu, Joost van de Weijer 0001 |
Neurocomputing | 1 |
| 2024 | Continual pre-training mitigates forgetting in language and visionabstractPre-trained models are commonly used in Continual Learning to initialize the model before training on the stream of non-stationary data. However, pre-training is rarely applied during Continual Learning. We investigate the characteristics of the Continual Pre-Training scenario, where a model is continually pre-trained on a stream of incoming data and only later fine-tuned to different downstream tasks. We introduce an evaluation protocol for Continual Pre-Training which monitors forgetting against a Forgetting Control dataset not present in the continual stream. We disentangle the impact on forgetting of 3 main factors: the input modality (NLP, Vision), the architecture type (Transformer, ResNet) and the pre-training protocol (supervised, self-supervised). Moreover, we propose a Sample-Efficient Pre-training method (SEP) that speeds up the pre-training phase. We show that the pre-training protocol is the most important factor accounting for forgetting. Surprisingly, we discovered that self-supervised continual pre-training in both NLP and Vision is sufficient to mitigate forgetting without the use of any Continual Learning strategy. Other factors, like model depth, input modality and architecture type are not as crucial. • Continual Pre-Training incrementally acquires knowledge from unstructured data streams. • Self-Supervised Continual Pre-Training effectively mitigates forgetting. • The representation drift is reduced by Self-Supervised Continual Pre-Training. • Performance on domain-specific tasks can be improved with a limited amount of data. Andrea Cossu, Antonio Carta, Lucia C. Passaro, Vincenzo Lomonaco, Tinne Tuytelaars, Davide Bacciu |
Neural Networks | 2 |
| 2023 | A Simple Recipe to Meta-Learn Forward and Backward TransferabstractMeta-learning holds the potential to provide a general and explicit solution to tackle interference and forgetting in continual learning. However, many popular algorithms introduce expensive and unstable optimization processes with new key hyper-parameters and requirements, hindering their applicability. We propose a new, general, and simple meta-learning algorithm for continual learning (SiM4C) that explicitly optimizes to minimize forgetting and facilitate forward transfer. We show our method is stable, introduces only minimal computational overhead, and can be integrated with any memory-based continual learning algorithm in only a few lines of code. SiM4C meta-learns how to effectively continually learn even on very long task sequences, largely outperforming prior meta-approaches. Naively integrating with existing memory-based algorithms, we also record universal performance benefits and state-of-the-art results across different visual classification benchmarks without introducing new hyper-parameters. Edoardo Cetin, Antonio Carta, Oya Çeliktutan |
ICCV | 2 |
| 2023 | Avalanche: A PyTorch Library for Deep Continual LearningabstractContinual learning is the problem of learning from a nonstationary stream of data, a fundamental issue for sustainable and efficient training of deep neural networks over time. Unfortunately, deep learning libraries only provide primitives for offline training, assuming that model's architecture and data are fixed. Avalanche is an open source library maintained by the ContinualAI non-profit organization that extends PyTorch by providing first-class support for dynamic architectures, streams of datasets, and incremental training and evaluation methods. Avalanche provides a large set of predefined benchmarks and training algorithms and it is easy to extend and modular while supporting a wide range of continual learning scenarios. Documentation is available at https://avalanche.continualai.org. Antonio Carta, Lorenzo Pellegrini, Andrea Cossu, Hamed Hemati, Vincenzo Lomonaco |
J. Mach. Learn. Res. | 1 |
| 2022 | Sample Condensation in Online Continual LearningabstractOnline Continual learning is a challenging learning scenario where the model must learn from a non-stationary stream of data where each sample is seen only once. The main challenge is to incrementally learn while avoiding catastrophic forgetting, namely the problem of forgetting previously acquired knowledge while learning from new data. A popular solution in these scenario is to use a small memory to retain old data and rehearse them over time. Unfortunately, due to the limited memory size, the quality of the memory will deteriorate over time. In this paper we propose OLCGM, a novel replay-based continual learning strategy that uses knowledge condensation techniques to continuously compress the memory and achieve a better use of its limited size. The sample condensation step compresses old samples, instead of removing them like other replay strategies. As a result, the experiments show that, whenever the memory budget is limited compared to the complexity of the data, OLCGM improves the final accuracy compared to state-of-the-art replay strategies. Mattia Sangermano, Antonio Carta, Andrea Cossu, Davide Bacciu |
IJCNN | 2 |
| 2021 | Continual Learning with Echo State NetworksabstractContinual Learning (CL) refers to a learning setup where data is non stationary and the model has to learn without forgetting existing knowledge.The study of CL for sequential patterns revolves around trained recurrent networks.In this work, instead, we introduce CL in the context of Echo State Networks (ESNs), where the recurrent component is kept fixed.We provide the first evaluation of catastrophic forgetting in ESNs and we highlight the benefits in using CL strategies which are not applicable to trained recurrent models.Our results confirm the ESN as a promising model for CL and open to its use in streaming scenarios.* This work has been partially supported by the H2020 TEACHING Andrea Cossu, Davide Bacciu, Antonio Carta, Claudio Gallicchio, Vincenzo Lomonaco |
ESANN | 3 |
| 2021 | Encoding-based memory for recurrent neural networks
Antonio Carta, Alessandro Sperduti, Davide Bacciu |
Neurocomputing | 1 |
| 2021 | Continual learning for recurrent neural networks: An empirical evaluation
Andrea Cossu, Antonio Carta, Vincenzo Lomonaco, Davide Bacciu |
Neural Networks | 2 |
| 2020 | Learning Style-Aware Symbolic Music Representations by Adversarial AutoencodersabstractWe address the challenging open problem of learning an effective latent space for symbolic music data in generative music modeling. We focus on leveraging adversarial regularization as a flexible and natural mean to imbue variational autoencoders with context information concerning music genre and style. Through the paper, we show how Gaussian mixtures taking into account music metadata information can be used as an effective prior for the autoencoder latent space, introducing the first Music Adversarial Autoencoder (MusAE). The empirical analysis on a large scale benchmark shows that our model has a higher reconstruction accuracy than state-of-the-art models based on standard variational autoencoders. It is also able to create realistic interpolations between two musical sequences, smoothly changing the dynamics of the different tracks. Experiments show that the model can organise its latent space accordingly to low-level properties of the musical pieces, as well as to embed into the latent variables the high-level genre information injected from the prior distribution to increase its overall performance. This allows us to perform changes to the generated pieces in a principled way. Andrea Valenti, Antonio Carta, Davide Bacciu |
ECAI | 2 |
| 2020 | Continual Learning with Gated Incremental Memories for sequential data processingabstractThe ability to learn in dynamic, nonstationary environments without forgetting previous knowledge, also known as Continual Learning (CL), is a key enabler for scalable and trustworthy deployments of adaptive solutions. While the importance of continual learning is largely acknowledged in machine vision and reinforcement learning problems, this is mostly under-documented for sequence processing tasks. This work proposes a Recurrent Neural Network (RNN) model for CL that is able to deal with concept drift in input distribution without forgetting previously acquired knowledge. We also implement and test a popular CL approach, Elastic Weight Consolidation (EWC), on top of two different types of RNNs. Finally, we compare the performances of our enhanced architecture against EWC and RNNs on a set of standard CL benchmarks, adapted to the sequential data processing scenario. Results show the superior performance of our architecture and highlight the need for special solutions designed to address CL in RNNs. Andrea Cossu, Antonio Carta, Davide Bacciu |
IJCNN | 2 |
| 2020 | Incremental Training of a Recurrent Neural Network Exploiting a Multi-scale Dynamic Memory
Antonio Carta, Alessandro Sperduti, Davide Bacciu |
ECML/PKDD (1) | 1 |
| 2019 | Linear Memory Networks
Davide Bacciu, Antonio Carta, Alessandro Sperduti |
ICANN (1) | 2 |