Andrea Cossu

dblp:262/6262 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
17since 2021 · last 2026
0000-0002-4874-8830ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 8 first-author · 17 since 2021
YearPublicationVenuePosition
2026 Random Unicycle Network (RUN!): supercharging harmonic oscillator networks via non-holonomic constraints
abstract
Motivated by advances in physical reservoir computing, we seek models that retain the modularity of echo state networks while enriching their internal dynamics.Recent studies have demonstrated that oscillator networks can achieve this balance, although their simple harmonic nature may limit their expressiveness.Here, we investigate the idea of augmenting harmonic oscillators with non-holonomic (velocitylevel) constraints, known to induce rich, nonlocal behaviors.We implement these constraints intrinsically within each dynamical unit, yielding a model equivalent to the unicycle -the canonical representation of the simplest vehicle.We test the model on three time-series classification benchmarks, achieving competitive or superior accuracy compared to the state of the art, with reservoirs as small as 20 unicycles.
Mariano Ramírez Montero, Andrea Ceni, Andrea Cossu, Davide Bacciu, Claudio Gallicchio, Cosimo Della Santina
ESANN3
2026 A practical guide to streaming continual learning
Andrea Cossu, Federico Giannini, Giacomo Ziffer, Alessio Bernardo, Alexander Gepperth, Emanuele Della Valle, Barbara Hammer, Davide Bacciu
Neurocomputing1
2025 Replay-free Online Continual Learning with Self-Supervised MultiPatches
abstract
Online Continual Learning (OCL) methods train a model on a non-stationary data stream where only a few examples are available at a time, often leveraging replay strategies.However, usage of replay is sometimes forbidden, especially in applications with strict privacy regulations.Therefore, we propose Continual MultiPatches (CMP), an effective plugin for existing OCL self-supervised learning strategies that avoids the use of replay samples.CMP generates multiple patches from a single example and projects them into a shared feature space, where patches coming from the same example are pushed together without collapsing into a single point.CMP surpasses replay and other SSL-based strategies on OCL streams, challenging the role of replay as a go-to solution for self-supervised OCL.Code available at https://github.com/giacomo-cgn/cmp .
Giovanni A. Cignoni, Andrea Cossu, Alexandra Gomez-Villa, Joost van de Weijer 0001, Antonio Carta
ESANN2
2025 Don't drift away: Advances and Applications of Streaming and Continual Learning
abstract
Non-stationary environments subject to concept drift require the design of adaptive models that can continuously learn and update.Two primary research communities have emerged to address this challenge: Continual Learning (CL) and Streaming Machine Learning (SML).CL manages virtual drifts by learning new concepts without forgetting past knowledge, while SML focuses on real drifts, rapidly adapting to evolving data distributions.However, a unified approach is needed to balance adaptation and knowledge retention.Streaming Continual Learning (SCL) bridges the gap between CL and SML, ensuring models retain useful past information while efficiently adapting to new data.We explore key challenges in SCL, including handling temporal dependencies in data streams and adapting latent representations for personalization and knowledge editing.Additionally, we identify promising SCL benchmarks which can foster and promote a unified research effort between CL and SML. 35
Andrea Cossu, Davide Bacciu, Alessio Bernardo, Emanuele Della Valle, Alexander Gepperth, Federico Giannini, Barbara Hammer, Giacomo Ziffer
ESANN1
2025 Lifelong Evolution of Swarms
abstract
Adapting to task changes without forgetting previous knowledge is a key skill for intelligent systems, and a crucial aspect of lifelong learning. Swarm controllers, however, are typically designed for specific tasks, lacking the ability to retain knowledge across changing tasks. Lifelong learning, on the other hand, focuses on individual agents with limited insights into the emergent abilities of a collective like a swarm. To address this gap, we introduce a lifelong evolutionary framework for swarms, where a population of swarm controllers is evolved in a dynamic environment that incrementally presents novel tasks. This requires evolution to find controllers that quickly adapt to new tasks while retaining knowledge of previous ones, as they may reappear in the future. We discover that the population inherently preserves information about previous tasks, and it can reuse it to foster adaptation and mitigate forgetting. In contrast, the top-performing individual for a given task catastrophically forgets previous tasks. To mitigate this phenomenon, we design a regularization process for the evolutionary algorithm, reducing forgetting in top-performing individuals. Evolving swarms in a lifelong fashion raises fundamental questions on the current state of deep lifelong learning and on the robustness of swarm controllers in dynamic environments.
Lorenzo Leuzzi, Davide Bacciu, Sabine Hauert, Andrea Cossu
GECCO5
2024 Random Oscillators Network for Time Series Processing
abstract
We introduce the Random Oscillators Network (RON), a physically-inspired recurrent model derived from a network of heterogeneous oscillators. Unlike traditional recurrent neural networks, RON keeps the connections between oscillators untrained by leveraging on smart random initialisations, leading to exceptional computational efficiency. A rigorous theoretical analysis finds the necessary and sufficient conditions for the stability of RON, highlighting the natural tendency of RON to lie at the edge of stability, a regime of configurations offering particularly powerful and expressive models. Through an extensive empirical evaluation on several benchmarks, we show four main advantages of RON. 1) RON shows excellent long-term memory and sequence classification ability, outperforming other randomised approaches. 2) RON outperforms fully-trained recurrent models and state-of-the-art randomised models in chaotic time series forecasting. 3) RON provides expressive internal representations even in a small parametrisation regime making it amenable to be deployed on low-powered devices and at the edge. 4) RON is up to two orders of magnitude faster than fully-trained models.
Andrea Ceni, Andrea Cossu, Maximilian Stölzle, Jingyue Liu 0001, Cosimo Della Santina, Davide Bacciu, Claudio Gallicchio
AISTATS2
2024 Towards Deep Continual Workspace Monitoring: Performance Evaluation of CL Strategies for Object Detection in Working Sites
abstract
Object detection plays a crucial role in computer-based monitoring tasks, where the adaptability of object detection algorithms to complex and dynamic backgrounds is essential for achieving accurate and stable detection performance.Despite the effectiveness of state-of-the-art object detectors, continual object detection remains a significant challenge in real-world applications.In this study, we utilized a dataset tailored for continual object detection in diverse working environments.Using this dataset, a task-incremental and task-agnostic continual learning scenario was established in which each experience, corresponding to object detection sub-datasets collected from different work sites.Common baseline continual learning (CL) strategies were employed throughout the continual training process to evaluate their efficacy.Our findings, consistent with the CL literature, underscore replay-based strategies as the top performers, assessed across both task-aware and task-agnostic settings.Additionally, zero-shot object detection demonstrates notably lower performance compared to the best-performing CL strategies, emphasizing the critical importance of CL strategies in maintaining consistent detection performance and adapting to new environments and work sites.
Asli Çelik, Oguzhan Urhan, Andrea Cossu, Vincenzo Lomonaco
ESANN3
2024 Enhancing Echo State Networks with Gradient-based Explainability Methods
abstract
Recurrent Neural Networks are effective for analyzing temporal data, such as time series, but they often require costly and time-intensive training.Echo State Networks simplify the training process by using a fixed recurrent layer, the reservoir, and a trainable output layer, the readout.In sequence classification problems, the readout typically receives only the final state of the reservoir.However, averaging all states can sometimes be beneficial.In this work, we assess whether a weighted average of hidden states can enhance the Echo State Network performance.To this end, we propose a gradient-based, explainable technique to guide the contribution of each hidden state towards the final prediction.We show that our approach outperforms the naive average, as well as other baselines, in time series classification, particularly on noisy data.
Francesco Spinnato, Andrea Cossu, Riccardo Guidotti, Andrea Ceni, Claudio Gallicchio, Davide Bacciu
ESANN2
2024 Projected Latent Distillation for Data-Agnostic Consolidation in distributed continual learning
abstract
In continual learning applications on-the-edge multiple self-centered devices (SCD) learn different local tasks independently, with each SCD only optimizing its own task. Can we achieve (almost) zero-cost collaboration between different devices? We formalize this problem as a Distributed Continual Learning (DCL) scenario, where SCDs greedily adapt to their own local tasks and a separate continual learning (CL) model perform a sparse and asynchronous consolidation step that combines the SCD models sequentially into a single multi-task model without using the original data. Unfortunately, current CL methods are not directly applicable to this scenario. We propose Data-Agnostic Consolidation (DAC), a novel double knowledge distillation method which performs distillation in the latent space via a novel Projected Latent Distillation loss. Experimental results show that DAC enables forward transfer between SCDs and reaches state-of-the-art accuracy on Split CIFAR100, CORe50 and Split TinyImageNet, both in single device and distributed CL scenarios. Somewhat surprisingly, a single out-of-distribution image is sufficient as the only source of data for DAC.
Antonio Carta, Andrea Cossu, Vincenzo Lomonaco, Davide Bacciu, Joost van de Weijer 0001
Neurocomputing2
2024 Drifting explanations in continual learning
abstract
Continual Learning (CL) trains models on streams of data, with the aim of learning new information without forgetting previous knowledge. However, many of these models lack interpretability, making it difficult to understand or explain how they make decisions. This lack of interpretability becomes even more challenging given the non-stationary nature of the data streams in CL. Furthermore, CL strategies aimed at mitigating forgetting directly impact the learned representations. We study the behavior of different explanation methods in CL and propose CLEX (ContinuaL EXplanations), an evaluation protocol to robustly assess the change of explanations in Class-Incremental scenarios, where forgetting is pronounced. We observed that models with similar predictive accuracy do not generate similar explanations. Replay-based strategies, well-known to be some of the most effective ones in class-incremental scenarios, are able to generate explanations that are aligned to the ones of a model trained offline. On the contrary, naive fine-tuning often results in degenerate explanations that drift from the ones of an offline model. Finally, we discovered that even replay strategies do not always operate at best when applied to fully-trained recurrent models. Instead, randomized recurrent models (leveraging on an untrained recurrent component) clearly reduce the drift of the explanations. This discrepancy between fully-trained and randomized recurrent models, previously known only in the context of their predictive continual performance, is more general, including also continual explanations.
Andrea Cossu, Francesco Spinnato, Riccardo Guidotti, Davide Bacciu
Neurocomputing1
2024 Continual pre-training mitigates forgetting in language and vision
abstract
Pre-trained models are commonly used in Continual Learning to initialize the model before training on the stream of non-stationary data. However, pre-training is rarely applied during Continual Learning. We investigate the characteristics of the Continual Pre-Training scenario, where a model is continually pre-trained on a stream of incoming data and only later fine-tuned to different downstream tasks. We introduce an evaluation protocol for Continual Pre-Training which monitors forgetting against a Forgetting Control dataset not present in the continual stream. We disentangle the impact on forgetting of 3 main factors: the input modality (NLP, Vision), the architecture type (Transformer, ResNet) and the pre-training protocol (supervised, self-supervised). Moreover, we propose a Sample-Efficient Pre-training method (SEP) that speeds up the pre-training phase. We show that the pre-training protocol is the most important factor accounting for forgetting. Surprisingly, we discovered that self-supervised continual pre-training in both NLP and Vision is sufficient to mitigate forgetting without the use of any Continual Learning strategy. Other factors, like model depth, input modality and architecture type are not as crucial. • Continual Pre-Training incrementally acquires knowledge from unstructured data streams. • Self-Supervised Continual Pre-Training effectively mitigates forgetting. • The representation drift is reduced by Self-Supervised Continual Pre-Training. • Performance on domain-specific tasks can be improved with a limited amount of data.
Andrea Cossu, Antonio Carta, Lucia C. Passaro, Vincenzo Lomonaco, Tinne Tuytelaars, Davide Bacciu
Neural Networks1
2023 A Protocol for Continual Explanation of SHAP
abstract
Continual Learning trains models on a stream of data, with the aim of learning new information without forgetting previous knowledge.Given the dynamic nature of such environments, explaining the predictions of these models can be challenging.We study the behavior of SHAP values explanations in Continual Learning and propose an evaluation protocol to robustly assess the change of explanations in Class-Incremental scenarios.We observed that, while Replay strategies enforce the stability of SHAP values in feedforward/convolutional models, they are not able to do the same with fully-trained recurrent models.We show that alternative recurrent approaches, like randomized recurrent models, are more effective in keeping the explanations stable over time.
Andrea Cossu, Francesco Spinnato, Riccardo Guidotti, Davide Bacciu
ESANN1
2023 Avalanche: A PyTorch Library for Deep Continual Learning
abstract
Continual learning is the problem of learning from a nonstationary stream of data, a fundamental issue for sustainable and efficient training of deep neural networks over time. Unfortunately, deep learning libraries only provide primitives for offline training, assuming that model's architecture and data are fixed. Avalanche is an open source library maintained by the ContinualAI non-profit organization that extends PyTorch by providing first-class support for dynamic architectures, streams of datasets, and incremental training and evaluation methods. Avalanche provides a large set of predefined benchmarks and training algorithms and it is easy to extend and modular while supporting a wide range of continual learning scenarios. Documentation is available at https://avalanche.continualai.org.
Antonio Carta, Lorenzo Pellegrini, Andrea Cossu, Hamed Hemati, Vincenzo Lomonaco
J. Mach. Learn. Res.3
2022 Continual Learning for Human State Monitoring
abstract
Continual Learning (CL) on time series data represents a promising but under-studied avenue for real-world applications.We propose two new CL benchmarks for Human State Monitoring.We carefully designed the benchmarks to mirror real-world environments in which new subjects are continuously added.We conducted an empirical evaluation to assess the ability of popular CL strategies to mitigate forgetting in our benchmarks.Our results show that, possibly due to the domainincremental properties of our benchmarks, forgetting can be easily tackled even with a simple finetuning and that existing strategies struggle in accumulating knowledge over a fixed, held-out, test subject.* This work has been partially
Federico Matteoni, Andrea Cossu, Claudio Gallicchio, Vincenzo Lomonaco, Davide Bacciu
ESANN2
2022 Sample Condensation in Online Continual Learning
abstract
Online Continual learning is a challenging learning scenario where the model must learn from a non-stationary stream of data where each sample is seen only once. The main challenge is to incrementally learn while avoiding catastrophic forgetting, namely the problem of forgetting previously acquired knowledge while learning from new data. A popular solution in these scenario is to use a small memory to retain old data and rehearse them over time. Unfortunately, due to the limited memory size, the quality of the memory will deteriorate over time. In this paper we propose OLCGM, a novel replay-based continual learning strategy that uses knowledge condensation techniques to continuously compress the memory and achieve a better use of its limited size. The sample condensation step compresses old samples, instead of removing them like other replay strategies. As a result, the experiments show that, whenever the memory budget is limited compared to the complexity of the data, OLCGM improves the final accuracy compared to state-of-the-art replay strategies.
Mattia Sangermano, Antonio Carta, Andrea Cossu, Davide Bacciu
IJCNN3
2021 Continual Learning with Echo State Networks
abstract
Continual Learning (CL) refers to a learning setup where data is non stationary and the model has to learn without forgetting existing knowledge.The study of CL for sequential patterns revolves around trained recurrent networks.In this work, instead, we introduce CL in the context of Echo State Networks (ESNs), where the recurrent component is kept fixed.We provide the first evaluation of catastrophic forgetting in ESNs and we highlight the benefits in using CL strategies which are not applicable to trained recurrent models.Our results confirm the ESN as a promising model for CL and open to its use in streaming scenarios.* This work has been partially supported by the H2020 TEACHING
Andrea Cossu, Davide Bacciu, Antonio Carta, Claudio Gallicchio, Vincenzo Lomonaco
ESANN1
2021 Continual learning for recurrent neural networks: An empirical evaluation
Andrea Cossu, Antonio Carta, Vincenzo Lomonaco, Davide Bacciu
Neural Networks1
2020 Continual Learning with Gated Incremental Memories for sequential data processing
abstract
The ability to learn in dynamic, nonstationary environments without forgetting previous knowledge, also known as Continual Learning (CL), is a key enabler for scalable and trustworthy deployments of adaptive solutions. While the importance of continual learning is largely acknowledged in machine vision and reinforcement learning problems, this is mostly under-documented for sequence processing tasks. This work proposes a Recurrent Neural Network (RNN) model for CL that is able to deal with concept drift in input distribution without forgetting previously acquired knowledge. We also implement and test a popular CL approach, Elastic Weight Consolidation (EWC), on top of two different types of RNNs. Finally, we compare the performances of our enhanced architecture against EWC and RNNs on a set of standard CL benchmarks, adapted to the sequential data processing scenario. Results show the superior performance of our architecture and highlight the need for special solutions designed to address CL in RNNs.
Andrea Cossu, Antonio Carta, Davide Bacciu
IJCNN1