Alessandro Betti

dblp:180/7658 · DBLP profile ↗
← Back
22ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0002-9052-8743ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 State-space modeling in long sequence processing: a survey on recurrence in the transformer era
abstract
Effectively learning from sequential data is a longstanding goal of Artificial Intelligence, especially in the case of long sequences. From the dawn of Machine Learning, several researchers have pursued algorithms and architectures capable of processing sequences of patterns, retaining information about past inputs while still leveraging future data, without losing precious long-term dependencies and correlations. While such an ultimate goal is inspired by the human hallmark of continuous real-time processing of sensory information, several solutions have simplified the learning paradigm by artificially limiting the processed context or dealing with sequences of limited length, given in advance. These solutions were further emphasized by the ubiquity of Transformers, which initially overshadowed the role of Recurrent Neural Nets. However, recurrent networks are currently experiencing a strong recent revival due to the growing popularity of (deep) State-Space models and novel instances of large-context Transformers, which are both based on recurrent computations that aim to go beyond several limits of currently ubiquitous technologies. The fast development of Large Language Models has renewed the interest in efficient solutions to process data over time. This survey provides an in-depth summary of the latest approaches that are based on recurrent models for sequential data processing. A complete taxonomy of recent trends in architectural and algorithmic solutions is reported and discussed, guiding researchers in this appealing research field. The emerging picture suggests that there is room for exploring novel routes, constituted by learning algorithms that depart from the standard Backpropagation Through Time, towards a more realistic scenario where patterns are effectively processed online, leveraging local-forward computations, and opening new directions for research on this topic.
Matteo Tiezzi, Michele Casoni, Alessandro Betti, Marco Gori, Stefano Melacci
Neural Networks3
2025 A Conditional Generative Diffusion Model for Spatio-Temporal Data
abstract
Diffusion models have become widely used for generating text, image, video, and audio. In recent years, these models have also been introduced into the time series domain for tasks such as forecasting, imputation, and generation. An interesting application is conditional generation with respect to metadata, enabling the synthesis of data sequences that match specified conditions. In the present work, we focus on spatio-temporal data and more precisely on time series data that also exhibit a spatial nature—for instance, measurements from sensor networks, traffic flows, mobile network usage across different areas. However, current approaches for time series conditional generation often neglect spatial autocorrelation. In this work, we extend the well-known DIFFWAVE model to address this challenge by directly taking into account the spatial nature of the data in the denoising process of the diffusion model. We evaluate our approach on a large and complex real-world dataset from the NET-MOB 2023 data challenge, which collects mobile network usage of different mobile applications across urban areas. Our results demonstrate that, in addition to achieving competitive performance across all evaluated metrics, our approach is also able to correctly capture the spatial autocorrelation present in the real data.
Giulio Loddi, Alessandro Betti, Fabio Pinelli
ECAI2
2025 Stability of State and Costate Dynamics in Continuous Time Recurrent Neural Networks
abstract
The notion of stability plays a crucial role in ensuring the safe development of a model in a lifelong learning context.This paper investigates the fundamental aspects of stability in a class of continuous-time recurrent neural networks which include both state and costate variables.The latter are directly inherited from optimal control theory, and they act as adjoint variables closely related to gradient terms.Stability is investigated both in terms of state and of costate dynamics, showing the key conditions that must be satisfied to produce bounded dynamics in the forward and learning stages.* This work was
Alessandro Betti, Marco Gori, Stefano Melacci
ESANN1
2025 Perpetual Generation: Online Learning of Linear State-Space Models from a Single Stream
Michele Casoni, Tommaso Guidi, Stefano Melacci, Alessandro Betti, Marco Gori
ICANN (1)4
2025 Generative System Dynamics in Recurrent Neural Networks
abstract
In this study, we investigate the continuous time dynamics of Recurrent Neural Networks (RNNs), focusing on systems with nonlinear activation functions. The objective of this work is to identify conditions under which RNNs exhibit perpetual oscillatory behavior, without converging to static fixed points. We establish that skew-symmetric weight matrices are fundamental to enable stable limit cycles in both linear and nonlinear configurations. We further demonstrate that hyperbolic tangent-like activation functions (odd, bounded, and continuous) preserve these oscillatory dynamics by ensuring motion invariants in state space. Numerical simulations showcase how nonlinear activation functions not only maintain limit cycles, but also enhance the numerical stability of the system integration process, mitigating those instabilities that are commonly associated with the forward Euler method. The experimental results of this analysis highlight practical considerations for designing neural architectures capable of capturing complex temporal dependencies, i.e., strategies for enhancing memorization skills in recurrent models.
Michele Casoni, Tommaso Guidi, Alessandro Betti, Stefano Melacci, Marco Gori
IJCNN3
2025 Position Paper: A new Perspective on Online Continual Learning
abstract
We consider the setting of online/streaming continual learning. In particular, we analyze the implications of the consolidated continual learning evaluation setting in an online scenario. As the data stream potentially extends infinitely, the evaluation set also grows boundlessly, necessitating a redefinition of model evaluation. In the streaming learning literature, the model is usually evaluated with the test-then-train method that does not require a separate evaluation set at all, or with continuous reevaluation that considers partially delayed labels [1]. In this work, we propose to model the data-generating process as random walks on concept graphs, i.e. graphs that define the possible transitions between the concepts to learn. This offers a precise formalization of the problem of learning on (possibly) infinitely long data streams with concept drifts and recurring concepts. In such a setting, we also discuss alternative possibilities that blend the continual learning setting with the traditional online learning one. We propose a framework that generalizes the commonly considered online continual learning scenario (i.e. corresponds to a specific topology of the concept graph). Finally, we discuss how different topologies of data streams, derived from our framework, could inspire new research directions in continual learning.
Nicolò Navarin, Alessandro Betti, Marco Gori
IJCNN2
2024 Neural Time-Reversed Generalized Riccati Equation
abstract
Optimal control deals with optimization problems in which variables steer a dynamical system, and its outcome contributes to the objective function. Two classical approaches to solving these problems are Dynamic Programming and the Pontryagin Maximum Principle. In both approaches, Hamiltonian equations offer an interpretation of optimality through auxiliary variables known as costates. However, Hamiltonian equations are rarely used due to their reliance on forward-backward algorithms across the entire temporal domain. This paper introduces a novel neural-based approach to optimal control. Neural networks are employed not only for implementing state dynamics but also for estimating costate variables. The parameters of the latter network are determined at each time step using a newly introduced local policy referred to as the time-reversed generalized Riccati equation. This policy is inspired by a result discussed in the Linear Quadratic (LQ) problem, which we conjecture stabilizes state dynamics. We support this conjecture by discussing experimental results from a range of optimal control case studies.
Alessandro Betti, Michele Casoni, Marco Gori, Simone Marullo, Stefano Melacci, Matteo Tiezzi
AAAI1
2024 Bridging Continual Learning of Motion and Self-Supervised Representations
abstract
Efficiently learning unsupervised pixel-wise visual representations is crucial for training agents that can perceive their environment without relying on heavy human supervision or abundant annotated data. Motivated by recent work that promotes motion as a key source of information in representation learning, we propose a novel instance of contrastive criterions over time and space. In our architecture, pixel-wise motion field and representations are extracted by neural models, trained from scratch in an integrated fashion. Learning proceeds online over time, exploiting also a momentum-based moving average to update the feature extractor, without replaying any large buffers of past data. Experiments on real-world videos and on a recently introduced benchmark, with photorealistic streams generated from a 3D environment, confirm that the proposed model can learn to estimate motion and jointly develop representations. Our model nicely encodes the variable appearance of the visual information in space and time, significantly overcoming a recent approach and it also compares favourably with convolutional and Transformer-based networks, offline-pre-trained on large collections of supervised and unsupervised images.
Matteo Tiezzi, Simone Marullo, Alessandro Betti, Michele Casoni, Stefano Melacci
ECAI3
2024 Nature-Inspired Local Propagation
abstract
The spectacular results achieved in machine learning, including the recent advances in generative AI, rely on large data collections. On the opposite, intelligent processes in nature arises without the need for such collections, but simply by on-line processing of the environmental information. In particular, natural learning processes rely on mechanisms where data representation and learning are intertwined in such a way to respect spatiotemporal locality. This paper shows that such a feature arises from a pre-algorithmic view of learning that is inspired by related studies in Theoretical Physics. We show that the algorithmic interpretation of the derived “laws of learning”, which takes the structure of Hamiltonian equations, reduces to Backpropagation when the speed of propagation goes to infinity. This opens the doors to machine learning studies based on full on-line information processing that are based on the replacement of Backpropagation with the proposed spatiotemporal local algorithm.
Alessandro Betti, Marco Gori
NeurIPS1
2023 Knowledge-Driven Active Learning
abstract
In the last few years, Deep Learning models have become increasingly popular. However, their deployment is still precluded in those contexts where the amount of supervised data is limited and manual labelling expensive. Active learning strategies aim at solving this problem by requiring supervision only on few unlabelled samples, which improve the most model performances after adding them to the training set. Most strategies are based on uncertain sample selection, and even often restricted to samples lying close to the decision boundary. Here we propose a very different approach, taking into consideration domain knowledge. Indeed, in the case of multi-label classification, the relationships among classes offer a way to spot incoherent predictions, i.e., predictions where the model may most likely need supervision. We have developed a framework where first-order-logic knowledge is converted into constraints and their violation is checked as a natural guide for sample selection. We empirically demonstrate that knowledge-driven strategy outperforms standard strategies, particularly on those datasets where domain knowledge is complete. Furthermore, we show how the proposed approach enables discovering data distributions lying far from training data. Finally, the proposed knowledge-driven strategy can be also easily used in object-detection problems where standard uncertainty-based techniques are difficult to apply.
Gabriele Ciravegna, Frédéric Precioso, Alessandro Betti, Kevin Mottin, Marco Gori
ECML/PKDD (1)3
2023 Local propagation of visual stimuli in focus of attention
abstract
Fast reactions to changes in the surrounding visual environment require efficient attention mechanisms to reallocate computational resources to most relevant locations in the visual field. While current computational models keep improving their predictive ability thanks to the increasing availability of data, they still struggle approximating the effectiveness and efficiency exhibited by foveated animals. In this paper, we present a biologically-plausible computational model of focus of attention that exhibits spatiotemporal locality and that is very well-suited for parallel and distributed implementations. Attention emerges as a wave propagation process originated by visual stimuli corresponding to details and motion information. The resulting field obeys the principle of "inhibition of return" so as not to get stuck in potential holes. An accurate experimentation of the model shows that it achieves top level performance in scanpath prediction tasks. This can easily be understood at the light of a theoretical result that we establish in the paper, where we prove that as the velocity of wave propagation goes to infinity, the proposed model reduces to recently proposed state of the art gravitational models of focus of attention.
Lapo Faggi, Alessandro Betti, Dario Zanca, Stefano Melacci, Marco Gori
Neurocomputing2
2022 PARTIME: Scalable and Parallel Processing Over Time with Deep Neural Networks
abstract
In this paper, we present PARTIME, a software library written in Python and based on PyTorch, designed specifically to speed up neural networks whenever data is continuously streamed over time, for both learning and inference. Existing libraries are designed to exploit data-level parallelism, assuming that samples are batched, a condition that is not naturally met in applications that are based on streamed data. Differently, PARTIME starts processing each data sample at the time in which it becomes available from the stream. PARTIME wraps the code that implements a feed-forward multi-layer network and it distributes the layer-wise processing among multiple devices, such as Graphics Processing Units (GPUs). Thanks to its pipeline-based computational scheme, PARTIME allows the devices to perform computations in parallel. At inference time this results in scaling capabilities that are theoretically linear with respect to the number of devices. During the learning stage, PARTIME can leverage the non-i.i.d. nature of the streamed data with samples that are smoothly evolving over time for efficient gradient computations. Experiments are performed in order to empirically compare PARTIME with classic non-parallel neural computations in online learning, distributing operations on up to 8 NVIDIA GPUs, showing significant speedups that are almost linear in the number of devices, mitigating the impact of the data transfer overhead.
Enrico Meloni, Lapo Faggi, Simone Marullo, Alessandro Betti, Matteo Tiezzi, Marco Gori, Stefano Melacci
ICMLA4
2022 Stochastic Coherence Over Attention Trajectory For Continuous Learning In Video Streams
abstract
Devising intelligent agents able to live in an environment and learn by observing the surroundings is a longstanding goal of Artificial Intelligence. From a bare Machine Learning perspective, challenges arise when the agent is prevented from leveraging large fully-annotated dataset, but rather the interactions with supervisory signals are sparsely distributed over space and time. This paper proposes a novel neural-network-based approach to progressively and autonomously develop pixel-wise representations in a video stream. The proposed method is based on a human-like attention mechanism that allows the agent to learn by observing what is moving in the attended locations. Spatio-temporal stochastic coherence along the attention trajectory, paired with a contrastive term, leads to an unsupervised learning criterion that naturally copes with the considered setting. Differently from most existing works, the learned representations are used in open-set class-incremental classification of each frame pixel, relying on few supervisions. Our experiments leverage 3D virtual environments and they show that the proposed agents can learn to distinguish objects just by observing the video stream. Inheriting features from state-of-the art models is not as powerful as one might expect.
Matteo Tiezzi, Simone Marullo, Lapo Faggi, Enrico Meloni, Alessandro Betti, Stefano Melacci
IJCAI5
2022 Foveated Neural Computation
Matteo Tiezzi, Simone Marullo, Alessandro Betti, Enrico Meloni, Lapo Faggi, Marco Gori, Stefano Melacci
ECML/PKDD (3)3
2020 Developing Constrained Neural Units Over Time
abstract
In this paper we present a foundational study on a constrained method that defines learning problems with Neural Networks in the context of the principle of least cognitive action, which very much resembles the principle of least action in mechanics. Starting from a general approach to enforce constraints into the dynamical laws of learning, this work focuses on an alternative way of defining Neural Networks, that is different from the majority of existing approaches. In particular, the structure of the neural architecture is defined by means of a special class of constraints that are extended also to the interaction with data, leading to "architectural" and "input-related" constraints, respectively. The proposed theory is cast into the time domain, in which data are presented to the network in an ordered manner, that makes this study an important step toward alternative ways of processing continuous streams of data with Neural Networks. The connection with the classic Backpropagation-based update rule of the weights of networks is discussed, showing that there are conditions under which our approach degenerates to Backpropagation. Moreover, the theory is experimentally evaluated on a simple problem that allows us to deeply study several aspects of the theory itself and to show the soundness of the model.
Alessandro Betti, Marco Gori, Simone Marullo, Stefano Melacci
IJCNN1
2020 Local Propagation in Constraint-based Neural Networks
abstract
In this paper we study a constraint-based representation of neural network architectures. We cast the learning problem in the Lagrangian framework and we investigate a simple optimization procedure that is well suited to fulfil the so-called architectural constraints, learning from the available supervisions. The computational structure of the proposed Local Propagation (LP) algorithm is based on the search for saddle points in the adjoint space composed of weights, neural outputs, and Lagrange multipliers. All the updates of the model variables are locally performed, so that LP is fully parallelizable over the neural units, circumventing the classic problem of gradient vanishing in deep networks. The implementation of popular neural models is described in the context of LP, together with those conditions that trace a natural connection with Backpropagation. We also investigate the setting in which we tolerate bounded violations of the architectural constraints, and we provide experimental evidence that LP is a feasible approach to train shallow and deep networks, opening the road to further investigations on more complex architectures, easily describable by constraints.
Giuseppe Marra, Matteo Tiezzi, Stefano Melacci, Alessandro Betti, Marco Maggini, Marco Gori
IJCNN4
2020 Focus of Attention Improves Information Transfer in Visual Features
abstract
Unsupervised learning from continuous visual streams is a challenging problem that cannot be naturally and efficiently managed in the classic batch-mode setting of computation. The information stream must be carefully processed accordingly to an appropriate spatio-temporal distribution of the visual data, while most approaches of learning commonly assume uniform probability density. In this paper we focus on unsupervised learning for transferring visual information in a truly online setting by using a computational model that is inspired to the principle of least action in physics. The maximization of the mutual information is carried out by a temporal process which yields online estimation of the entropy terms. The model, which is based on second-order differential equations, maximizes the information transfer from the input to a discrete space of symbols related to the visual features of the input, whose computation is supported by hidden neurons. In order to better structure the input probability distribution, we use a human-like focus of attention model that, coherently with the information maximization model, is also based on second-order differential equations. We provide experimental results to support the theory by showing that the spatio-temporal filtering induced by the focus of attention allows the system to globally transfer more information from the input stream over the focused areas and, in some contexts, over the whole frames with respect to the unfiltered case that yields uniform probability distributions.
Matteo Tiezzi, Stefano Melacci, Alessandro Betti, Marco Maggini, Marco Gori
NeurIPS3
2020 Learning visual features under motion invariance
abstract
Humans are continuously exposed to a stream of visual data with a natural temporal structure. However, most successful computer vision algorithms work at image level, completely discarding the precious information carried by motion. In this paper, we claim that processing visual streams naturally leads to formulate the motion invariance principle, which enables the construction of a new theory of learning that originates from variational principles, just like in physics. Such principled approach is well suited for a discussion on a number of interesting questions that arise in vision, and it offers a well-posed computational scheme for the discovery of convolutional filters over the retina. Differently from traditional convolutional networks, which need massive supervision, the proposed theory offers a truly new scenario for the unsupervised processing of video signals, where features are extracted in a multi-layer architecture with motion invariance. While the theory enables the implementation of novel computer vision systems, it also sheds light on the role of information-based principles to drive possible biological solutions.
Alessandro Betti, Marco Gori, Stefano Melacci
Neural Networks1
2020 Cognitive Action Laws: The Case of Visual Features
Alessandro Betti, Marco Gori, Stefano Melacci
IEEE Trans. Neural Networks Learn. Syst.1
2019 Motion Invariance in Visual Environments
abstract
The puzzle of computer vision might find new challenging solutions when we realize that most successful methods are working at image level, which is remarkably more difficult than processing directly visual streams, just as it happens in nature. In this paper, we claim that the processing of a stream of frames naturally leads to formulate the motion invariance principle, which enables the construction of a new theory of visual learning based on convolutional features. The theory addresses a number of intriguing questions that arise in natural vision, and offers a well-posed computational scheme for the discovery of convolutional filters over the retina. They are driven by the Euler- Lagrange differential equations derived from the principle of least cognitive action, that parallels the laws of mechanics. Unlike traditional convolutional networks, which need massive supervision, the proposed theory offers a truly new scenario in which feature learning takes place by unsupervised processing of video signals. An experimental report of the theory is presented where we show that features extracted under motion invariance yield an improvement that can be assessed by measuring information-based indexes.
Alessandro Betti, Marco Gori, Stefano Melacci
IJCAI1
2018 Classification of Heterogenous M2M/IoT Traffic Based on C-plane and U-plane Data
abstract
This paper is motivated by the observation that M2M/IoT traffic is rather heterogeneous. By using traffic data collected in a real scenario, we prove that some M2M devices might generate a signaling traffic more similar to the traffic of a smartphone than to the traffic of a traditional M2M device used for metering applications. This makes difficult the task of a mobile operator to understand the impact of the introduction of M2M/IoT devices into the network. The paper presents a classification of the M2M/IoT heterogeneous world in three classes. The classification has been performed using the C-plane attributes and its effectiveness has been assessed by performing a clustering analysis over the U-plane data, which represent the ground truth. We found a good matching between the results of the classification task carried out by using the C-plane data and the results of the clustering task carried out by exploiting the U-plane data. This means that, by using the C-plane data (simpler with respect to using U-plane data) the network operator can understand with a good accuracy which type of M2M devices are active on the network, and what are the applications that they run and the data traffic that they generate. Therefore, the results may be useful for a proper dimensioning and management of the evolution of an EPC network.
Simone Di Domenico, Mauro De Sanctis, Ernestina Cianca, Lorenzo Silvestri, Vito Curcuru, Alessandro Betti
PIMRC6
2016 The principle of least cognitive action
Alessandro Betti, Marco Gori
Theor. Comput. Sci.1