Davide Bacciu

dblp:07/6626 · DBLP profile ↗
← Back
149ranked-venue papers
56as first author
89since 2021 · last 2026
0000-0001-5213-2468ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 141 · 55 first-author · 82 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Random Unicycle Network (RUN!): supercharging harmonic oscillator networks via non-holonomic constraints
abstract
Motivated by advances in physical reservoir computing, we seek models that retain the modularity of echo state networks while enriching their internal dynamics.Recent studies have demonstrated that oscillator networks can achieve this balance, although their simple harmonic nature may limit their expressiveness.Here, we investigate the idea of augmenting harmonic oscillators with non-holonomic (velocitylevel) constraints, known to induce rich, nonlocal behaviors.We implement these constraints intrinsically within each dynamical unit, yielding a model equivalent to the unicycle -the canonical representation of the simplest vehicle.We test the model on three time-series classification benchmarks, achieving competitive or superior accuracy compared to the state of the art, with reservoirs as small as 20 unicycles.
Mariano Ramírez Montero, Andrea Ceni, Andrea Cossu, Davide Bacciu, Claudio Gallicchio, Cosimo Della Santina
ESANN4
2026 Emotion Recognition in Multimodal Social Data
abstract
Emotion recognition on social media is often approached in unimodal or single-label settings, despite the multimodal nature of online communication.This paper presents a study of multilabel emotion recognition from paired text-image data.We evaluate vision-language encoders and compare them with strong unimodal baselines and a zero-shot multimodal LLM.A simple multimodal classifier built on CLIP achieves the most reliable performance.Data-centric additions such as emoji transcription, caption augmentation, and pseudo-labelling offer limited gains, whereas calibrated decision thresholds have a consistent effect.The results highlight the value of visual cues and show limitations of recent VLMs.
Lucia C. Passaro, Davide Amadei, Davide Bacciu
ESANN3
2026 Boundary-Constrained Diffusion Models for Floorplan Generation: Balancing Realism and Diversity
abstract
Diffusion models have become widely popular for automated floorplan generation, producing highly realistic layouts conditioned on user-defined constraints.However, optimizing for perceptual metrics such as the Frećhet Inception Distance (FID) causes limited design diversity.To address this, we propose the Diversity Score (DS), a metric that quantifies layout diversity under fixed constraints.Moreover, to improve geometric consistency, we introduce a Boundary Cross-Attention (BCA) module that enables conditioning on building boundaries.Our experiments show that BCA significantly improves boundary adherence, while prolonged training drives diversity collapse undiagnosed by FID, revealing a critical trade-off between realism and diversity.Out-Of-Distribution evaluations further demonstrate the models' reliance on dataset priors, emphasizing the need for generative systems that explicitly balance fidelity, diversity, and generalization in architectural design tasks.
Leonardo Stoppani, Davide Bacciu, Shahab Mokarizadeh
ESANN2
2026 Sparse assemblies of recurrent neural networks with stability guarantees
Andrea Ceni, Valerio De Caro, Davide Bacciu, Claudio Gallicchio
Neurocomputing3
2026 A practical guide to streaming continual learning
Andrea Cossu, Federico Giannini, Giacomo Ziffer, Alessio Bernardo, Alexander Gepperth, Emanuele Della Valle, Barbara Hammer, Davide Bacciu
Neurocomputing8
2026 Informed machine learning for complex data
abstract
Machine Learning (ML) has become a central force in Artificial Intelligence, driving major breakthroughs in applications that handle increasingly complex data, from images and text sequences to graph structures. While new architectures such as Transformers and Graph Neural Networks continue to redefine performance benchmarks in various domains, these predominantly data-driven methods often neglect critical domain knowledge, practical constraints, and broader contextual factors. This oversight diminishes their trustworthiness and restricts their impact in real-world settings. In this paper, we discuss the need for a more informed approach to ML for complex data. Specifically, we advocate for solutions that explicitly integrate structural awareness to capture underlying relationships in the data, incorporate key technical requirements to ensure safety and compliance with industry standards, embed environmental considerations to promote sustainability and resource efficiency, adhere to established physical principles, and uphold ethical and societal values. By weaving these dimensions together, informed ML can bridge the gap between purely data-centric methods and the nuanced demands of practical applications. We show how this integrated framework not only strengthens model performance but also ensures that ML solutions remain trustworthy, efficient, and sensitive to human ecological, ethical, and regulatory imperatives. Our discussion underscores the transformative potential of Informed ML to drive innovation across diverse domains, setting a new benchmark for responsible and high-impact ML system design.
Luca Oneto, Nicolò Navarin, Alessio Micheli, Luca Pasa, Claudio Gallicchio, Davide Bacciu, Davide Anguita
Neurocomputing6
2025 On Oversquashing in Graph Neural Networks Through the Lens of Dynamical Systems
abstract
A common problem in Message-Passing Neural Networks is oversquashing -- the limited ability to facilitate effective information flow between distant nodes. Oversquashing is attributed to the exponential decay in information transmission as node distances increase. This paper introduces a novel perspective to address oversquashing, leveraging dynamical systems properties of global and local non-dissipativity, that enable the maintenance of a constant information flow rate. We present SWAN, a uniquely parameterized GNN model with antisymmetry both in space and weight domains, as a means to obtain non-dissipativity. Our theoretical analysis asserts that by implementing these properties, SWAN offers an enhanced ability to transmit information over extended distances. Empirical evaluations on synthetic and real-world benchmarks that emphasize long-range interactions validate the theoretical understanding of SWAN, and its ability to mitigate oversquashing.
Alessio Gravina, Moshe Eliasof, Claudio Gallicchio, Davide Bacciu, Carola-Bibiane Schönlieb
AAAI4
2025 Fast: Similarity-Based Knowledge Transfer for Efficient Policy Learning
abstract
Transfer Learning (TL) offers the potential to accelerate learning by transferring knowledge across tasks. However, it faces critical challenges such as negative transfer, domain adaptation and inefficiency in selecting solid source policies. These issues often represent critical problems in evolving domains, i.e. game development, where scenarios transform and agents must adapt. The continuous release of new agents is costly and inefficient. In this work we challenge the key issues in TL to improve knowledge transfer, agents performance across tasks and reduce computational costs. The proposed methodology, called FAST - Framework for Adaptive Similarity-based Transfer, leverages visual frames and textual descriptions to create a latent representation of tasks dynamics, that is exploited to estimate similarity between environments. The similarity scores guides our method in choosing candidate policies from which transfer abilities to simplify learning of novel tasks. Experimental results, over multiple racing tracks, demonstrate that FAST achieves competitive final performance compared to learning-from-scratch methods while requiring significantly less training steps. These findings highlight the potential of embedding-driven task similarity estimations.
Alessandro Capurso, Elia Piccoli, Davide Bacciu
CoG3
2025 Foundation and Generative Models for Graphs
abstract
The rapidly evolving field of machine learning for graphstructured data gathered significant attention due to its ability to preserve critical information inherent in complex data structures.As a result, significant efforts have been dedicated to designing advanced architectures and foundational models optimized for graph-based operations.Research in this area explores methodologies for graph representation learning and graph generation, incorporating probabilistic models such as variational autoencoders and normalizing flows.Despite increasing interest from researchers as well as their efforts in solving graph-related problems, several issues and areas remain to be addressed to improve model generalization and reliability.This tutorial reviews foundational concepts and challenges in graph representation, structure learning, and graph generation, while also summarizing the contributions accepted for publication in the special session on this topic at the 33th European
Davide Bacciu, Federico Errica, Stefano Moro, Luca Pasa, Davide Rigoni 0001, Daniele Zambon
ESANN1
2025 Towards Adaptive and Stable Compositional Assemblies of Recurrent Neural Network Modules
abstract
Recurrent neural networks (RNNs) are computational models regarded as dynamical systems.Modularity is a key ingredient of complex systems.Thus, the composition of RNN modules provides a simple paradigm for building complex computational models, with the potential to approach the human brain capability.We devise strategies for training RNNs assembled into a larger RNN of RNNs, provided with theoretical guarantees of stability that hold during training for the composed global network.Experiments on pixel-by-pixel image classification benchmarks prove the effectiveness of this approach.* This work has been
Valerio De Caro, Andrea Ceni, Davide Bacciu, Claudio Gallicchio
ESANN3
2025 Don't drift away: Advances and Applications of Streaming and Continual Learning
abstract
Non-stationary environments subject to concept drift require the design of adaptive models that can continuously learn and update.Two primary research communities have emerged to address this challenge: Continual Learning (CL) and Streaming Machine Learning (SML).CL manages virtual drifts by learning new concepts without forgetting past knowledge, while SML focuses on real drifts, rapidly adapting to evolving data distributions.However, a unified approach is needed to balance adaptation and knowledge retention.Streaming Continual Learning (SCL) bridges the gap between CL and SML, ensuring models retain useful past information while efficiently adapting to new data.We explore key challenges in SCL, including handling temporal dependencies in data streams and adapting latent representations for personalization and knowledge editing.Additionally, we identify promising SCL benchmarks which can foster and promote a unified research effort between CL and SML. 35
Andrea Cossu, Davide Bacciu, Alessio Bernardo, Emanuele Della Valle, Alexander Gepperth, Federico Giannini, Barbara Hammer, Giacomo Ziffer
ESANN2
2025 Generalized Stochastic Pooling
abstract
Pooling layers play a critical role in Convolutional Neural Networks by reducing spatial dimensions and enhancing translation invariance.While conventional methods like max pooling and average pooling are effective, they can respectively amplify noise or dilute important features.Stochastic pooling introduces probabilistic sampling to improve generalization but is susceptible to biases from outliers, often mimicking max pooling in such cases.To address these limitations, we propose a generalization of stochastic pooling that introduces a tunable parameter to control the balance between uniform sampling, stochastic pooling, and max pooling.Experiments on multiple datasets demonstrate that uniform sampling outperforms the biased one, achieving a favorable trade-off between regularization and performance.
Francesco Landolfi, Davide Bacciu
ESANN2
2025 Towards Efficient Molecular Property Optimization with Graph Energy Based Models
abstract
Optimizing chemical properties is a challenging task due to the vastness and complexity of chemical space.Here, we present a generative energy-based architecture for implicit chemical property optimization, designed to efficiently generate molecules that satisfy target properties without explicit conditional generation.We use Graph Energy Based Models and a training approach that does not require property labels.We validated our approach on well-established chemical benchmarks, showing superior results to state-of-the-art methods and demonstrating robustness and efficiency towards de novo drug design.
Luca Miglior, Lorenzo Simone, Marco Podda, Davide Bacciu
ESANN4
2025 Lifelong Evolution of Swarms
abstract
Adapting to task changes without forgetting previous knowledge is a key skill for intelligent systems, and a crucial aspect of lifelong learning. Swarm controllers, however, are typically designed for specific tasks, lacking the ability to retain knowledge across changing tasks. Lifelong learning, on the other hand, focuses on individual agents with limited insights into the emergent abilities of a collective like a swarm. To address this gap, we introduce a lifelong evolutionary framework for swarms, where a population of swarm controllers is evolved in a dynamic environment that incrementally presents novel tasks. This requires evolution to find controllers that quickly adapt to new tasks while retaining knowledge of previous ones, as they may reappear in the future. We discover that the population inherently preserves information about previous tasks, and it can reuse it to foster adaptation and mitigate forgetting. In contrast, the top-performing individual for a given task catastrophically forgets previous tasks. To mitigate this phenomenon, we design a regularization process for the evolutionary algorithm, reducing forgetting in top-performing individuals. Evolving swarms in a lifelong fashion raises fundamental questions on the current state of deep lifelong learning and on the robustness of swarm controllers in dynamic environments.
Lorenzo Leuzzi, Davide Bacciu, Sabine Hauert, Andrea Cossu
GECCO2
2025 Real-Time and Personalized Product Recommendations for Large E-Commerce Platforms
Matteo Tolloso, Davide Bacciu, Shahab Mokarizadeh, Marco Varesi
ICANN (3)2
2025 Port-Hamiltonian Architectural Bias for Long-Range Propagation in Deep Graph Networks
abstract
The dynamics of information diffusion within graphs is a critical open issue that heavily influences graph representation learning, especially when considering long-range propagation. This calls for principled approaches that control and regulate the degree of propagation and dissipation of information throughout the neural flow. Motivated by this, we introduce port-Hamiltonian Deep Graph Networks, a novel framework that models neural information flow in graphs by building on the laws of conservation of Hamiltonian dynamical systems. We reconcile under a single theoretical and practical framework both non-dissipative long-range propagation and non-conservative behaviors, introducing tools from mechanical systems to gauge the equilibrium between the two components. Our approach can be applied to general message-passing architectures, and it provides theoretical guarantees on information conservation in time. Empirical results prove the effectiveness of our port-Hamiltonian scheme in pushing simple graph convolutional architectures to state-of-the-art performance in long-range benchmarks.
Simon Heilig, Alessio Gravina, Alessandro Trenta, Claudio Gallicchio, Davide Bacciu
ICLR5
2025 Graph Adaptive Autoregressive Moving Average Models
abstract
Graph State Space Models (SSMs) have recently been introduced to enhance Graph Neural Networks (GNNs) in modeling long-range interactions. Despite their success, existing methods either compromise on permutation equivariance or limit their focus to pairwise interactions rather than sequences. Building on the connection between Autoregressive Moving Average (ARMA) and SSM, in this paper, we introduce GRAMA, a Graph Adaptive method based on a learnable ARMA framework that addresses these limitations. By transforming from static to sequential graph data, GRAMA leverages the strengths of the ARMA framework, while preserving permutation equivariance. Moreover, GRAMA incorporates a selective attention mechanism for dynamic learning of ARMA coefficients, enabling efficient and flexible long-range information propagation. We also establish theoretical connections between GRAMA and Selective SSMs, providing insights into its ability to capture long-range dependencies. Experiments on 26 synthetic and real-world datasets demonstrate that GRAMA consistently outperforms backbone models and performs competitively with state-of-the-art methods.
Moshe Eliasof, Alessio Gravina, Andrea Ceni, Claudio Gallicchio, Davide Bacciu, Carola-Bibiane Schönlieb
ICML5
2025 Non-Dissipative Graph Propagation for Non-Local Community Detection
abstract
Community detection in graphs aims to cluster nodes into meaningful groups, a task particularly challenging in heterophilic graphs, where nodes sharing similarities and membership to the same community are typically distantly connected. This is particularly evident when this task is tackled by graph neural networks, since they rely on an inherently local message passing scheme to learn the node representations that serve to cluster nodes into communities. In this work, we argue that the ability to propagate long-range information during message passing is key to effectively perform community detection in heterophilic graphs. To this end, we introduce the Unsupervised Antisymmetric Graph Neural Network (uAGNN), a novel unsupervised community detection approach leveraging non-dissipative dynamical systems to ensure stability and to propagate long-range information effectively. By employing antisymmetric weight matrices, uAGNN captures both local and global graph structures, overcoming the limitations posed by heterophilic scenarios. Extensive experiments across ten datasets demonstrate uAGNN’s superior performance in high and medium heterophilic settings, where traditional methods fail to exploit long-range dependencies. These results highlight uAGNN’s potential as a powerful tool for unsupervised community detection in diverse graph environments.
Will Leeney, Alessio Gravina, Davide Bacciu
IJCNN3
2025 Return of ChebNet: Understanding and Improving an Overlooked GNN on Long Range Tasks
abstract
ChebNet, one of the earliest spectral GNNs, has largely been overshadowed by Message Passing Neural Networks (MPNNs), which gained popularity for their simplicity and effectiveness in capturing local graph structure. Despite their success, MPNNs are limited in their ability to capture long-range dependencies between nodes. This has led researchers to adapt MPNNs through *rewiring* or make use of *Graph Transformers*, which compromise the computational efficiency that characterized early spatial message passing architectures, and typically disregard the graph structure. Almost a decade after its original introduction, we revisit ChebNet to shed light on its ability to model distant node interactions. We find that out-of-box, ChebNet already shows competitive advantages relative to classical MPNNs and GTs on long-range benchmarks, while maintaining good scalability properties for high-order polynomials. However, we uncover that this polynomial expansion leads ChebNet to an unstable regime during training. To address this limitation, we cast ChebNet as a stable and non-dissipative dynamical system, which we coin Stable-ChebNet. Our Stable-ChebNet model allows for stable information propagation, and has controllable dynamics which do not require the use of eigendecompositions, positional encodings, or graph rewiring. Across several benchmarks, Stable-ChebNet achieves near state-of-the-art performance.
Ali Hariri, Alvaro Arroyo, Alessio Gravina, Moshe Eliasof, Carola-Bibiane Schönlieb, Davide Bacciu, Xiaowen Dong 0001, Kamyar Azizzadenesheli, Pierre Vandergheynst
NeurIPS6
2025 Graph Diffusion that can Insert and Delete
abstract
Generative models of graphs based on discrete Denoising Diffusion Probabilistic Models (DDPMs) offer a principled approach to molecular generation by systematically removing structural noise through iterative atom and bond adjustments. However, existing formulations are fundamentally limited by their inability to adapt the graph size (that is, the number of atoms) during the diffusion process, severely restricting their effectiveness in conditional generation scenarios such as property-driven molecular design, where the targeted property often correlates with the molecular size. In this paper, we reformulate the noising and denoising processes to support monotonic insertion and deletion of nodes. The resulting model, which we call GrIDDD, dynamically grows or shrinks the chemical graph during generation. GrIDDD matches or exceeds the performance of existing graph diffusion models on molecular property targeting despite being trained on a more difficult problem. Furthermore, when applied to molecular optimization, GrIDDD exhibits competitive performance compared to specialized optimization models. This work paves the way for size-adaptive molecular generation with graph diffusion.
Matteo Ninniri, Marco Podda, Davide Bacciu
NeurIPS3
2025 Deferring Concept Bottleneck Models: Learning to Defer Interventions to Inaccurate Experts
abstract
Concept Bottleneck Models (CBMs) are interpretable machine learning models that ground their predictions on human-understandable concepts, allowing for targeted interventions in their decision-making process. However, when intervened on, CBMs assume the availability of humans that can identify the need to intervene and always provide correct interventions. Both assumptions are unrealistic and impractical, considering labor costs and human error-proneness. In contrast, Learning to Defer (L2D) extends supervised learning by allowing machine learning models to identify cases where a human is more likely to be correct than the model, thus leading to deferring systems with improved performance. In this work, we gain inspiration from L2D and propose Deferring CBMs (DCBMs), a novel framework that allows CBMs to learn when an intervention is needed. To this end, we model DCBMs as a composition of deferring systems and derive a consistent L2D loss to train them. Moreover, by relying on a CBM architecture, DCBMs can explain the reasons for deferring on the final task. Our results show that DCBMs can achieve high predictive performance and interpretability by deferring only when needed.
Andrea Pugnana, Riccardo Massidda, Francesco Giannini, Pietro Barbiero, Mateo Espinosa Zarlenga, Roberto Pellungrini, Gabriele Dominici, Fosca Giannotti, Davide Bacciu
NeurIPS9
2025 SONAR: Long-Range Graph Propagation Through Information Waves
abstract
Capturing effective long-range information propagation remains a fundamental yet challenging problem in graph representation learning. Motivated by this, we introduce SONAR, a novel GNN architecture inspired by the dynamics of wave propagation in continuous media. SONAR models information flow on graphs as oscillations governed by the wave equation, allowing it to maintain effective propagation dynamics over long distances. By integrating adaptive edge resistances and state-dependent external forces, our method balances conservative and non-conservative behaviors, improving the ability to learn more complex dynamics. We provide a rigorous theoretical analysis of SONAR's energy conservation and information propagation properties, demonstrating its capacity to address the long-range propagation problem. Extensive experiments on synthetic and real-world benchmarks confirm that SONAR achieves state-of-the-art performance, particularly on tasks requiring long-range information exchange.
Alessandro Trenta, Alessio Gravina, Davide Bacciu
NeurIPS3
2025 ECG synthesis for cardiac arrhythmias: Integrating self-supervised learning and generative adversarial networks
Lorenzo Simone, Davide Bacciu, Vincenzo Gervasi
Artif. Intell. Medicine2
2024 Quasi-Orthogonal ECG-Frank XYZ Transformation with Energy-Based Models and Clinical Text
Lorenzo Simone, Davide Bacciu, Vincenzo Gervasi
AIME (2)2
2024 Random Oscillators Network for Time Series Processing
abstract
We introduce the Random Oscillators Network (RON), a physically-inspired recurrent model derived from a network of heterogeneous oscillators. Unlike traditional recurrent neural networks, RON keeps the connections between oscillators untrained by leveraging on smart random initialisations, leading to exceptional computational efficiency. A rigorous theoretical analysis finds the necessary and sufficient conditions for the stability of RON, highlighting the natural tendency of RON to lie at the edge of stability, a regime of configurations offering particularly powerful and expressive models. Through an extensive empirical evaluation on several benchmarks, we show four main advantages of RON. 1) RON shows excellent long-term memory and sequence classification ability, outperforming other randomised approaches. 2) RON outperforms fully-trained recurrent models and state-of-the-art randomised models in chaotic time series forecasting. 3) RON provides expressive internal representations even in a small parametrisation regime making it amenable to be deployed on low-powered devices and at the edge. 4) RON is up to two orders of magnitude faster than fully-trained models.
Andrea Ceni, Andrea Cossu, Maximilian Stölzle, Jingyue Liu 0001, Cosimo Della Santina, Davide Bacciu, Claudio Gallicchio
AISTATS6
2024 I Know How: Combining Prior Policies to Solve New Tasks
abstract
Multi-Task Reinforcement Learning aims at developing agents that are able to continually evolve and adapt to new scenarios. However, this goal is challenging to achieve due to the phenomenon of catastrophic forgetting and the high demand of computational resources. Learning from scratch for each new task is not a viable or sustainable option, and thus agents should be able to collect and exploit prior knowledge while facing new problems. While several methodologies have attempted to address the problem from different perspectives, they lack a common structure. In this work, we propose a new framework, I Know How (IKH), which provides a common formalization. Our methodology focuses on modularity and compositionality of knowledge in order to achieve and enhance agent’s ability to learn and adapt efficiently to dynamic environments. To support our framework definition, we present a simple application of it in a simulated driving environment and compare its performance with that of state-of-the-art approaches.
Malio Li, Elia Piccoli, Vincenzo Lomonaco, Davide Bacciu
CoG4
2024 Generalizing Convolution to Point Clouds
abstract
Convolution, a fundamental operation in deep learning for structured grid data like images, cannot be directly applied to point clouds due to their irregular and unordered nature.Many approaches in literature that perform convolution on point clouds achieve this by designing a convolutional operator from scratch, often with little resemblance to the one used on images.We present two point cloud convolutions that naturally follow from the convolution in its standard definition popular with images.We do so by relaxing the indexing of the kernel weights with a "soft" dictionary that resembles the attention mechanism of the transformers.Finally, experimental results demonstrate the effectiveness of the proposed relaxations on two benchmark point cloud classification tasks. Generalizing Convolution to Point CloudsGiven the tensors
Davide Bacciu, Francesco Landolfi
ESANN1
2024 ADLER - An efficient Hessian-based strategy for adaptive learning rate
abstract
We derive a sound positive semi-definite approximation of the Hessian of deep models for which Hessian-vector products are easily computable.This enables us to provide an adaptive SGD learning rate strategy based on the minimization of the local quadratic approximation, which requires just twice the computation of a single SGD run, but performs comparably with grid search on SGD learning rates on different model architectures (CNN with and without residual connections) on classification tasks, which makes the algorithm a promising first step toward obtaining hyperparameter-free optimization of deep learning models, and also reduces the energy impact of training.We also compare the novel approximation with the Gauss-Newton approximation.
Dario Balboni, Davide Bacciu
ESANN2
2024 Large-Scale Continuous Structure Learning from Time-Series Data
abstract
Structure learning is the problem of recovering from data a Directed Acyclic Graph (DAG) of the interactions among variables.By enforcing a differentiable acyclicity constraint on the adjacency matrix of the graph, existing methods solve this problem as an optimization problem and have been recently extended to time-series data.Due to the cubic computational complexity of existing acyclicity constraints, their application is limited to a few variables.In this paper, we introduce svarcosmo, an optimization-based structure learning method for time-series data that builds upon recent developments on unconstrained but provably acyclic models.We empirically show on both simulated and real data that svarcosmo correctly recovers the underlying DAG in significantly less time, enabling optimization-based structure learning on high-dimensional data.
Filippo Michelis, Riccardo Massidda, Davide Bacciu
ESANN3
2024 Sequential Continual Pre-Training for Neural Machine Translation
abstract
We explore continual pre-training for Neural Machine Translation within a continual learning framework.We introduce a setting where new languages are gradually added to pre-trained models across multiple training experiences.These pre-trained models are subsequently fine-tuned on downstream translation tasks.We compare mBART and mT5 pre-training objectives using four European Languages.Our findings demonstrate that sequentially adding languages during pre-training effectively mitigates catastrophic forgetting and minimally impacts downstream task performance. * Work supported by PNRR-M4C2
Niko Dalla Noce, Michele Resta, Davide Bacciu
ESANN3
2024 Informed Machine Learning for Complex Data
abstract
In the contemporary era of data-driven decision-making, the application of Machine Learning (ML) on complex data (e.g., images, text, sequences, trees, and graphs) has become increasingly pivotal (e.g., Large Language Models and Graph Neural Networks).In this context, there is a gap between purely data-driven models and domain-specific knowledge, requirements, and expertise.In particular, this domain specificity needs to be integrated into the ML models to improve learning generalization, sustainability, trustworthiness, reliability, security, and safety.This additional knowledge can assume different forms, e.g.: software developers require ML to comply with many technical requirements, companies require ML to comply with economic and environmental sustainability, domain experts require ML to be aligned with physical and logical laws, and society requires ML to be aligned with ethical principles.This special session gathers valuable contributions and early findings in the field of Informed ML for Complex Data.Our main objective is to showcase the potential and limitations of new ideas, improvements, or the blending of ML and other research areas in solving real-world problems.
Luca Oneto, Nicolò Navarin, Alessio Micheli, Luca Pasa, Claudio Gallicchio, Davide Bacciu, Davide Anguita
ESANN6
2024 Enhancing Echo State Networks with Gradient-based Explainability Methods
abstract
Recurrent Neural Networks are effective for analyzing temporal data, such as time series, but they often require costly and time-intensive training.Echo State Networks simplify the training process by using a fixed recurrent layer, the reservoir, and a trainable output layer, the readout.In sequence classification problems, the readout typically receives only the final state of the reservoir.However, averaging all states can sometimes be beneficial.In this work, we assess whether a weighted average of hidden states can enhance the Echo State Network performance.To this end, we propose a gradient-based, explainable technique to guide the contribution of each hidden state towards the final prediction.We show that our approach outperforms the naive average, as well as other baselines, in time series classification, particularly on noisy data.
Francesco Spinnato, Andrea Cossu, Riccardo Guidotti, Andrea Ceni, Claudio Gallicchio, Davide Bacciu
ESANN6
2024 TEACHING Platform for Human-Centric Autonomous Applications: Design and Overview
abstract
The TEACHING project enhances AI applications in pervasive environments via Humanistic Intelligence, fostering synergy between humans and Cyber-Physical Systems of Systems (CPSoS). Here, we present the TEACHING Platform, a microservice-based framework providing the technological advancements to represent humans and CPSoS as containerized software models that interact to mutually empower each other.
Valerio De Caro, Christos Chronis, Massimo Coppola, Vincenzo Lomonaco, Claudio Gallicchio, Konstantinos Tserpes, Davide Bacciu
HPDC7
2024 Constraint-Free Structure Learning with Smooth Acyclic Orientations
abstract
The structure learning problem consists of fitting data generated by a Directed Acyclic Graph (DAG) to correctly reconstruct its arcs. In this context, differentiable approaches constrain or regularize an optimization problem with a continuous relaxation of the acyclicity property. The computational cost of evaluating graph acyclicity is cubic on the number of nodes and significantly affects scalability. In this paper, we introduce COSMO, a constraint-free continuous optimization scheme for acyclic structure learning. At the core of our method lies a novel differentiable approximation of an orientation matrix parameterized by a single priority vector. Differently from previous works, our parameterization fits a smooth orientation matrix and the resulting acyclic adjacency matrix without evaluating acyclicity at any step. Despite this absence, we prove that COSMO always converges to an acyclic solution. In addition to being asymptotically faster, our empirical analysis highlights how COSMO performance on graph reconstruction compares favorably with competing structure learning methods.
Riccardo Massidda, Francesco Landolfi, Martina Cinquini, Davide Bacciu
ICLR4
2024 Long Range Propagation on Continuous-Time Dynamic Graphs
abstract
Learning Continuous-Time Dynamic Graphs (C-TDGs) requires accurately modeling spatio-temporal information on streams of irregularly sampled events. While many methods have been proposed recently, we find that most message passing-, recurrent- or self-attention-based methods perform poorly on *long-range* tasks. These tasks require correlating information that occurred "far" away from the current event, either spatially (higher-order node information) or along the time dimension (events occurred in the past). To address long-range dependencies, we introduce Continuous-Time Graph Anti-Symmetric Network (CTAN). Grounded within the ordinary differential equations framework, our method is designed for efficient propagation of information. In this paper, we show how CTAN's (i) long-range modeling capabilities are substantiated by theoretical findings and how (ii) its empirical performance on synthetic long-range benchmarks and real-world benchmarks is superior to other methods. Our results motivate CTAN's ability to propagate long-range information in C-TDGs as well as the inclusion of long-range tasks as part of temporal graph models evaluation.
Alessio Gravina, Giulio Lovisotto, Claudio Gallicchio, Davide Bacciu, Claas Grohnfeldt
ICML4
2024 Temporal Graph ODEs for Irregularly-Sampled Time Series
Alessio Gravina, Daniele Zambon, Davide Bacciu, Cesare Alippi
IJCAI3
2024 Multi-Relational Graph Neural Network for Out-of-Domain Link Prediction
abstract
Dynamic multi-relational graphs are an expressive relational representation for data enclosing entities and relations of different types, and where relationships are allowed to vary in time. Addressing predictive tasks over such data requires the ability to find structure embeddings that capture the diversity of the relationships involved, as well as their dynamic evolution. In this work, we establish a novel class of challenging tasks for dynamic multi-relational graphs involving out-of-domain link prediction, where the relationship being predicted is not available in the input graph. We then introduce a novel Graph Neural Network model, named GOOD, designed specifically to tackle the out-of-domain generalization problem. GOOD introduces a novel design concept for multi-relation embedding aggregation, based on the idea that good representations are such when it is possible to disentangle the mixing proportions of the different relational embeddings that have produced it. We also propose five benchmarks based on two retail domains, where we show that GOOD can effectively generalize predictions out of known relationship types and achieve state-of-the-art results. Most importantly, we provide insights into problems where out-of-domain prediction might be preferred to an in-domain formulation, that is, where the relationship to be predicted has very few positive examples.
Asma Sattar, Georgios Deligiorgis, Marco Trincavelli, Davide Bacciu
IJCNN4
2024 ChemAlgebra: Algebraic Reasoning on Chemical Reactions
abstract
While showing impressive performance on various kinds of learning tasks, it is yet unclear whether deep learning models have the ability to robustly tackle reasoning tasks. Measuring the robustness of reasoning in modern machine learning models such as Transformers is challenging as one needs to provide a task that cannot be easily shortcut by exploiting spurious statistical correlations in the data, while operating on complex objects and constraints. To address this issue, we propose ChemAlgebra, a benchmark for measuring the reasoning capabilities of Transformer models through the prediction of stoichiometrically-balanced chemical reactions. ChemAlgebra requires manipulating sets of complex discrete objects – molecules represented as formulas or graphs – under algebraic constraints such as the mass preservation principle. We believe that ChemAlgebra can serve as a useful test bed for the next generation of machine reasoning models and as a promoter of their development.
Andrea Valenti, Davide Bacciu, Antonio Vergari
IJCNN2
2024 Self-generated Replay Memories for Continual Neural Machine Translation
abstract
Michele Resta, Davide Bacciu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Michele Resta, Davide Bacciu
NAACL-HLT2
2024 Classifier-Free Graph Diffusion for Molecular Property Targeting
Matteo Ninniri, Marco Podda, Davide Bacciu
ECML/PKDD (4)3
2024 Learning Causal Abstractions of Linear Structural Causal Models
abstract
The need for modelling causal knowledge at different levels of granularity arises in several settings. Causal Abstraction provides a framework for formalizing this problem by relating two Structural Causal Models at different levels of detail. Despite increasing interest in applying causal abstraction, e.g. in the interpretability of large machine learning models, the graphical and parametrical conditions under which a causal model can abstract another are not known. Furthermore, learning causal abstractions from data is still an open problem. In this work, we tackle both issues for linear causal models with linear abstraction functions. First, we characterize how the low-level coefficients and the abstraction function determine the high-level coefficients and how the high-level model constrains the causal ordering of low-level variables. Then, we apply our theoretical results to learn high-level and low-level causal models and their abstraction function from observational data. In particular, we introduce Abs-LiNGAM, a method that leverages the constraints induced by the learned high-level model and the abstraction function to speedup the recovery of the larger low-level model, under the assumption of non-Gaussian noise terms. In simulated settings, we show the effectiveness of learning causal abstractions from data and the potential of our method in improving scalability of causal discovery.
Riccardo Massidda, Sara Magliacane, Davide Bacciu
UAI3
2024 Projected Latent Distillation for Data-Agnostic Consolidation in distributed continual learning
abstract
In continual learning applications on-the-edge multiple self-centered devices (SCD) learn different local tasks independently, with each SCD only optimizing its own task. Can we achieve (almost) zero-cost collaboration between different devices? We formalize this problem as a Distributed Continual Learning (DCL) scenario, where SCDs greedily adapt to their own local tasks and a separate continual learning (CL) model perform a sparse and asynchronous consolidation step that combines the SCD models sequentially into a single multi-task model without using the original data. Unfortunately, current CL methods are not directly applicable to this scenario. We propose Data-Agnostic Consolidation (DAC), a novel double knowledge distillation method which performs distillation in the latent space via a novel Projected Latent Distillation loss. Experimental results show that DAC enables forward transfer between SCDs and reaches state-of-the-art accuracy on Split CIFAR100, CORe50 and Split TinyImageNet, both in single device and distributed CL scenarios. Somewhat surprisingly, a single out-of-distribution image is sufficient as the only source of data for DAC.
Antonio Carta, Andrea Cossu, Vincenzo Lomonaco, Davide Bacciu, Joost van de Weijer 0001
Neurocomputing4
2024 Drifting explanations in continual learning
abstract
Continual Learning (CL) trains models on streams of data, with the aim of learning new information without forgetting previous knowledge. However, many of these models lack interpretability, making it difficult to understand or explain how they make decisions. This lack of interpretability becomes even more challenging given the non-stationary nature of the data streams in CL. Furthermore, CL strategies aimed at mitigating forgetting directly impact the learned representations. We study the behavior of different explanation methods in CL and propose CLEX (ContinuaL EXplanations), an evaluation protocol to robustly assess the change of explanations in Class-Incremental scenarios, where forgetting is pronounced. We observed that models with similar predictive accuracy do not generate similar explanations. Replay-based strategies, well-known to be some of the most effective ones in class-incremental scenarios, are able to generate explanations that are aligned to the ones of a model trained offline. On the contrary, naive fine-tuning often results in degenerate explanations that drift from the ones of an offline model. Finally, we discovered that even replay strategies do not always operate at best when applied to fully-trained recurrent models. Instead, randomized recurrent models (leveraging on an untrained recurrent component) clearly reduce the drift of the explanations. This discrepancy between fully-trained and randomized recurrent models, previously known only in the context of their predictive continual performance, is more general, including also continual explanations.
Andrea Cossu, Francesco Spinnato, Riccardo Guidotti, Davide Bacciu
Neurocomputing4
2024 Continual pre-training mitigates forgetting in language and vision
abstract
Pre-trained models are commonly used in Continual Learning to initialize the model before training on the stream of non-stationary data. However, pre-training is rarely applied during Continual Learning. We investigate the characteristics of the Continual Pre-Training scenario, where a model is continually pre-trained on a stream of incoming data and only later fine-tuned to different downstream tasks. We introduce an evaluation protocol for Continual Pre-Training which monitors forgetting against a Forgetting Control dataset not present in the continual stream. We disentangle the impact on forgetting of 3 main factors: the input modality (NLP, Vision), the architecture type (Transformer, ResNet) and the pre-training protocol (supervised, self-supervised). Moreover, we propose a Sample-Efficient Pre-training method (SEP) that speeds up the pre-training phase. We show that the pre-training protocol is the most important factor accounting for forgetting. Surprisingly, we discovered that self-supervised continual pre-training in both NLP and Vision is sufficient to mitigate forgetting without the use of any Continual Learning strategy. Other factors, like model depth, input modality and architecture type are not as crucial. • Continual Pre-Training incrementally acquires knowledge from unstructured data streams. • Self-Supervised Continual Pre-Training effectively mitigates forgetting. • The representation drift is reduced by Self-Supervised Continual Pre-Training. • Performance on domain-specific tasks can be improved with a limited amount of data.
Andrea Cossu, Antonio Carta, Lucia C. Passaro, Vincenzo Lomonaco, Tinne Tuytelaars, Davide Bacciu
Neural Networks6
2024 Deep Learning for Dynamic Graphs: Models and Benchmarks
abstract
Recent progress in research on deep graph networks (DGNs) has led to a maturation of the domain of learning on graphs. Despite the growth of this research field, there are still important challenges that are yet unsolved. Specifically, there is an urge of making DGNs suitable for predictive tasks on real-world systems of interconnected entities, which evolve over time. With the aim of fostering research in the domain of dynamic graphs, first, we survey recent advantages in learning both temporal and spatial information, providing a comprehensive overview of the current state-of-the-art in the domain of representation learning for dynamic graphs. Second, we conduct a fair performance comparison among the most popular proposed approaches on node- and edge-level tasks, leveraging rigorous model selection and assessment for all the methods, thus establishing a sound baseline for evaluating new architectures and approaches.
Alessio Gravina, Davide Bacciu
IEEE Trans. Neural Networks Learn. Syst.2
2024 IEEE Transactions on Neural Networks and Learning Systems Special Issue on Causal Discovery and Causality-Inspired Machine Learning
abstract
Causality is a fundamental notion in science and engineering. It has attracted much interest across research communities in statistics, machine learning (ML), healthcare, and artificial intelligence (AI), and is becoming increasingly recognized as a vital research area. One of the fundamental problems in causality is how to find the causal structure or the underlying causal model. Accordingly, one focus of this Special Issue is oncausal discovery, i.e., how can we discover causal structure over a set of variables from observational data with automated procedures? Besides learning causality, another focus is on using causality to help understand and advance ML, that is, causality-inspired ML.
Kun Zhang 0001, Ilya Shpitser, Sara Magliacane, Davide Bacciu, Fei Wu 0001, Changshui Zhang, Peter Spirtes
IEEE Trans. Neural Networks Learn. Syst.4
2023 Generalizing Downsampling from Regular Data to Graphs
abstract
Downsampling produces coarsened, multi-resolution representations of data and it is used, for example, to produce lossy compression and visualization of large images, reduce computational costs, and boost deep neural representation learning. Unfortunately, due to their lack of a regular structure, there is still no consensus on how downsampling should apply to graphs and linked data. Indeed reductions in graph data are still needed for the goals described above, but reduction mechanisms do not have the same focus on preserving topological structures and properties, while allowing for resolution-tuning, as is the case in regular data downsampling. In this paper, we take a step in this direction, introducing a unifying interpretation of downsampling in regular and graph data. In particular, we define a graph coarsening mechanism which is a graph-structured counterpart of controllable equispaced coarsening mechanisms in regular data. We prove theoretical guarantees for distortion bounds on path lengths, as well as the ability to preserve key topological properties in the coarsened graphs. We leverage these concepts to define a graph pooling mechanism that we empirically assess in graph classification tasks, providing a greedy algorithm that allows efficient parallel implementation on GPUs, and showing that it compares favorably against pooling methods in literature.
Davide Bacciu, Alessio Conte, Francesco Landolfi
AAAI1
2023 ECGAN: Self-supervised Generative Adversarial Network for Electrocardiography
Lorenzo Simone, Davide Bacciu
AIME2
2023 Graph Representation Learning
abstract
In a broad range of real-world machine learning applications, representing examples as graphs is crucial to avoid a loss of information.For this reason, in the last few years, the definition of machine learning methods, particularly neural networks, for graph-structured inputs has been gaining increasing attention.In particular, Deep Graph Networks (DGNs) are nowadays the most commonly adopted models to learn a representation that can be used to address different tasks related to nodes, edges, or even entire graphs.This tutorial paper reviews fundamental concepts and open challenges of graph representation learning and summarizes the contributions that have been accepted for publication to the ESANN 2023 special session on the topic.
Davide Bacciu, Federico Errica, Alessio Micheli, Nicolò Navarin, Luca Pasa, Marco Podda, Daniele Zambon
ESANN1
2023 Communication-Efficient Ridge Regression in Federated Echo State Networks
abstract
Federated Echo State Networks represent an efficient methodology for learning in pervasive environments with private temporal data due to the low computational cost required by the learning phase.In this paper, we propose Partial Federated Ridge Regression (pFedRR), an approximate, communication-efficient version of the exact method for learning the readout in a federated setting.Each client compresses the local statistics to be exchanged with the server via an importance-based method, which selects the most relevant neurons with respect to the local distribution.We evaluate the methodology on two Human State Monitoring benchmarks, and results show that the importance-based selection of the information significantly reduces the communication cost, while acting as a regularization method to improve the generalization capabilities.
Valerio De Caro, Antonio Di Mauro, Davide Bacciu, Claudio Gallicchio
ESANN3
2023 Improving Fairness via Intrinsic Plasticity in Echo State Networks
abstract
Artificial Intelligence, and in particular Machine Learning, has become ubiquitous in today's society, both revolutionizing and impacting society as a whole.However, it can also lead to algorithmic bias and unfair results, especially when sensitive information is involved.This paper addresses the problem of algorithmic fairness in Machine Learning for temporal data, focusing on ensuring that sensitive time-dependent information does not unfairly influence the outcome of a classifier.In particular, we focus on a class of training-efficient recurrent neural models called Echo State Networks, and show, for the first time, how to leverage local unsupervised adaptation of the internal dynamics in order to build fairer classifiers.Experimental results on real-world problems from physiological sensor data demonstrate the potential of the proposal.
Andrea Ceni, Davide Bacciu, Valerio De Caro, Claudio Gallicchio, Luca Oneto
ESANN2
2023 A Protocol for Continual Explanation of SHAP
abstract
Continual Learning trains models on a stream of data, with the aim of learning new information without forgetting previous knowledge.Given the dynamic nature of such environments, explaining the predictions of these models can be challenging.We study the behavior of SHAP values explanations in Continual Learning and propose an evaluation protocol to robustly assess the change of explanations in Class-Incremental scenarios.We observed that, while Replay strategies enforce the stability of SHAP values in feedforward/convolutional models, they are not able to do the same with fully-trained recurrent models.We show that alternative recurrent approaches, like randomized recurrent models, are more effective in keeping the explanations stable over time.
Andrea Cossu, Francesco Spinnato, Riccardo Guidotti, Davide Bacciu
ESANN4
2023 Hidden Markov Models for Temporal Graph Representation Learning
abstract
We propose the Hidden Markov Model for temporal Graphs, a deep and fully probabilistic model for learning in the domain of dynamic time-varying graphs.We extend hidden Markov models for sequences to the graph domain by stacking probabilistic layers that perform efficient message passing and learn representations for the individual nodes.We evaluate the goodness of the learned representations on temporal node prediction tasks, and we observe promising results compared to neural approaches.
Federico Errica, Alessio Gravina, Davide Bacciu, Alessio Micheli
ESANN3
2023 A Tropical View of Graph Neural Networks
abstract
Learning dynamic programming algorithms with Graph Neural Networks (GNNs) is a research direction which is increasingly gaining popularity.Prior work has demonstrated that in order to learn such algorithms, it is necessary to have an "alignment" between the neural architecture and the dynamics of the target algorithms, and that GNNs align, in fact, with dynamic programming.Here, we provide a different view of this alignment, studying it through the lens of tropical algebra.We show that GNNs can approximate dynamic programming algorithms up to arbitrary precision, provided that their input and output are appropriately pre-and post-processed.
Francesco Landolfi, Davide Bacciu, Danilo Numeroso
ESANN2
2023 Anti-Symmetric DGN: a stable architecture for Deep Graph Networks
Alessio Gravina, Davide Bacciu, Claudio Gallicchio
ICLR2
2023 Dual Algorithmic Reasoning
Danilo Numeroso, Davide Bacciu, Petar Velickovic
ICLR2
2023 Graph-based Polyphonic Multitrack Music Generation
abstract
Graphs can be leveraged to model polyphonic multitrack symbolic music, where notes, chords and entire sections may be linked at different levels of the musical hierarchy by tonal and rhythmic relationships. Nonetheless, there is a lack of works that consider graph representations in the context of deep learning systems for music generation. This paper bridges this gap by introducing a novel graph representation for music and a deep Variational Autoencoder that generates the structure and the content of musical graphs separately, one after the other, with a hierarchical architecture that matches the structural priors of music. By separating the structure and content of musical graphs, it is possible to condition generation by specifying which instruments are played at certain times. This opens the door to a new form of human-computer interaction in the context of music co-creation. After training the model on existing MIDI datasets, the experiments show that the model is able to generate appealing short and long musical sequences and to realistically interpolate between them, producing music that is tonally and rhythmically consistent. Finally, the visualization of the embeddings shows that the model is able to organize its latent space in accordance with known musical concepts.
Emanuele Cosenza, Andrea Valenti, Davide Bacciu
IJCAI3
2023 Continual adaptation of federated reservoirs in pervasive environments
Valerio De Caro, Claudio Gallicchio, Davide Bacciu
Neurocomputing3
2023 Graph Neural Network for Context-Aware Recommendation
abstract
Abstract Recommendation problems are naturally tackled as a link prediction task in a bipartite graph between user and item nodes, labelled with rating information on edges. To provide personal recommendations and improve the performance of the recommender system, it is necessary to integrate side information along with user-item interactions. The integration of context is a key success factor in recommendation systems because it allows catering for user preferences and opinions, especially when this pertains to the circumstances surrounding the interaction between users and items. In this paper, we propose a context-aware Graph Convolutional Matrix Completion which captures structural information and integrates the user’s opinion on items along with the surrounding context on edges and static features of user and item nodes. Our graph encoder produces user and item representations with respect to context, features and opinion. The decoder takes the aggregated embeddings to predict the user-item score considering the surrounding context. We have evaluated the performance of our model on 14 five publicly available datasets and compared it with state-of-the-art algorithms. Throughout this we show how it can effectively integrate user opinion along with surrounding context to produce a final node representation which is aware of the favourite circumstances of the particular node.
Asma Sattar, Davide Bacciu
Neural Process. Lett.2
2023 Modeling Mood Polarity and Declaration Occurrence by Neural Temporal Point Processes
abstract
Neural point processes provide the flexibility needed to deal with time series of heterogeneous nature within the robust framework of point processes. This aspect is of particular relevance when dealing with real-world data, mixing generative processes characterized by radically different distributions and sampling. This brief discusses a neural point process approach for health and behavioral data, comprising both sparse events coming from user subjective declarations as well as fast-flowing time series from wearable sensors. We propose and empirically validate different neural architectures and we assess the effect of including input sources of different nature. The empirical analysis is built on the top of a challenging original dataset, never published before, and collected as part of a real-world experiment in an uncontrolled setting. Results show the potential of neural point processes both in terms of predicting the next event type as well as in predicting the time to next user interaction.
Davide Bacciu, Davide Morelli, Vlad Pandelea
IEEE Trans. Neural Networks Learn. Syst.1
2023 Explaining Deep Graph Networks via Input Perturbation
abstract
Deep graph networks (DGNs) are a family of machine learning models for structured data which are finding heavy application in life sciences (drug repurposing, molecular property predictions) and on social network data (recommendation systems). The privacy and safety-critical nature of such domains motivates the need for developing effective explainability methods for this family of models. So far, progress in this field has been challenged by the combinatorial nature and complexity of graph structures. In this respect, we present a novel local explanation framework specifically tailored to graph data and DGNs. Our approach leverages reinforcement learning to generate meaningful local perturbations of the input graph, whose prediction we seek an interpretation for. These perturbed data points are obtained by optimizing a multiobjective score taking into account similarities both at a structural level as well as at the level of the deep model outputs. By this means, we are able to populate a set of informative neighboring samples for the query graph, which is then used to fit an interpretable model for the predictive behavior of the deep network locally to the query graph prediction. We show the effectiveness of the proposed explainer by a qualitative analysis on two chemistry datasets, TOX21 and Estimated SOLubility (ESOL) and by quantitative results on a benchmark dataset for explanations, CYCLIQ.
Davide Bacciu, Danilo Numeroso
IEEE Trans. Neural Networks Learn. Syst.1
2022 An Empirical Verification of Wide Networks Theory
Dario Balboni, Davide Bacciu
BMVC2
2022 Deep Features for CBIR with Scarce Data using Hebbian Learning
abstract
Features extracted from Deep Neural Networks (DNNs) have proven to be very effective in the context of Content Based Image Retrieval (CBIR). Recently, biologically inspired Hebbian learning algorithms have shown promises for DNN training. In this contribution, we study the performance of such algorithms in the development of feature extractors for CBIR tasks. Specifically, we consider a semi-supervised learning strategy in two steps: first, an unsupervised pre-training stage is performed using Hebbian learning on the image dataset; second, the network is fine-tuned using supervised Stochastic Gradient Descent (SGD) training. For the unsupervised pre-training stage, we explore the nonlinear Hebbian Principal Component Analysis (HPCA) learning rule. For the supervised fine-tuning stage, we assume sample efficiency scenarios, in which the amount of labeled samples is just a small fraction of the whole dataset. Our experimental analysis, conducted on the CIFAR10 and CIFAR100 datasets, shows that, when few labeled samples are available, our Hebbian approach provides relevant improvements compared to various alternative methods.
Gabriele Lagani, Davide Bacciu, Claudio Gallicchio, Fabrizio Falchi, Claudio Gennaro, Giuseppe Amato 0001
CBMI2
2022 Learning Image Captioning as a Structured Transduction Task
Davide Bacciu, Davide Serramazza
EANN1
2022 Deep Learning for Graphs
abstract
The flourishing field of deep learning for graphs relies on the layered computation of representations from graph-structured input data.Message passing is the most common strategy for such processing of graphs, based on an efficient information exchange among the connected nodes via a local and iterative procedure.Representations learned in this way can be used to address different tasks related to nodes, edges, or even entire graphs.This tutorial paper reviews fundamental concepts and open challenges of deep learning for graphs and summarizes the contributions that have been accepted for publication to the ESANN 2022 special session on the topic.
Davide Bacciu, Federico Errica, Nicolò Navarin, Luca Pasa, Daniele Zambon
ESANN1
2022 Federated Adaptation of Reservoirs via Intrinsic Plasticity
abstract
We propose a novel algorithm for performing federated learning with Echo State Networks (ESNs) in a client-server scenario.In particular, our proposal focuses on the adaptation of reservoirs by combining Intrinsic Plasticity with Federated Averaging.The former is a gradientbased method for adapting the reservoir's non-linearity in a local and unsupervised manner, while the latter provides the framework for learning in the federated scenario.We evaluate our approach on real-world datasets from human monitoring, in comparison with the previous approach for federated ESNs existing in literature.Results show that adapting the reservoir with our algorithm provides a significant improvement on the performance of the global model.
Valerio De Caro, Claudio Gallicchio, Davide Bacciu
ESANN3
2022 Continual Learning for Human State Monitoring
abstract
Continual Learning (CL) on time series data represents a promising but under-studied avenue for real-world applications.We propose two new CL benchmarks for Human State Monitoring.We carefully designed the benchmarks to mirror real-world environments in which new subjects are continuously added.We conducted an empirical evaluation to assess the ability of popular CL strategies to mitigate forgetting in our benchmarks.Our results show that, possibly due to the domainincremental properties of our benchmarks, forgetting can be easily tackled even with a simple finetuning and that existing strategies struggle in accumulating knowledge over a fixed, held-out, test subject.* This work has been partially
Federico Matteoni, Andrea Cossu, Claudio Gallicchio, Vincenzo Lomonaco, Davide Bacciu
ESANN5
2022 Continual Incremental Language Learning for Neural Machine Translation
abstract
The paper provides an experimental investigation of the phenomena of catastrophic forgetting for Neural Machine Translation systems.We introduce and describe the continual incremental language learning setting and its analogy with the classical continual learning scenario.The experiments measure the performance loss of a naive incremental training strategy against a jointly trained baseline, and we show the mitigating effect of the replay strategy.To this end, we also introduce a prioritized replay buffer strategy informed by the specific application domain. 85
Michele Resta, Davide Bacciu
ESANN2
2022 Modular Representations for Weak Disentanglement
abstract
The recently introduced weakly disentangled representations proposed to relax some constraints of the previous definitions of disentanglement, in exchange for more flexibility.However, at the moment, weak disentanglement can only be achieved by increasing the amount of supervision as the number of factors of variations of the data increase.In this paper, we introduce modular representations for weak disentanglement, a novel method that allows to keep the amount of supervised information constant with respect the number of generative factors.The experiments shows that models using modular representations can increase their performance with respect to previous work without the need of additional supervision.* The work has been partially supported by the EU H2020 TAILOR project (n.952215).
Andrea Valenti, Davide Bacciu
ESANN2
2022 The Infinite Contextual Graph Markov Model
abstract
The Contextual Graph Markov Model (CGMM) is a deep, unsupervised, and probabilistic model for graphs that is trained incrementally on a layer-by-layer basis. As with most Deep Graph Networks, an inherent limitation is the need to perform an extensive model selection to choose the proper size of each layer’s latent representation. In this paper, we address this problem by introducing the Infinite Contextual Graph Markov Model (iCGMM), the first deep Bayesian nonparametric model for graph learning. During training, iCGMM can adapt the complexity of each layer to better fit the underlying data distribution. On 8 graph classification tasks, we show that iCGMM: i) successfully recovers or improves CGMM’s performances while reducing the hyper-parameters’ search space; ii) performs comparably to most end-to-end supervised methods. The results include studies on the importance of depth, hyper-parameters, and compression of the graph embeddings. We also introduce a novel approximated inference procedure that better deals with larger graph topologies.
Daniele Castellana, Federico Errica, Davide Bacciu, Alessio Micheli
ICML3
2022 Sample Condensation in Online Continual Learning
abstract
Online Continual learning is a challenging learning scenario where the model must learn from a non-stationary stream of data where each sample is seen only once. The main challenge is to incrementally learn while avoiding catastrophic forgetting, namely the problem of forgetting previously acquired knowledge while learning from new data. A popular solution in these scenario is to use a small memory to retain old data and rehearse them over time. Unfortunately, due to the limited memory size, the quality of the memory will deteriorate over time. In this paper we propose OLCGM, a novel replay-based continual learning strategy that uses knowledge condensation techniques to continuously compress the memory and achieve a better use of its limited size. The sample condensation step compresses old samples, instead of removing them like other replay strategies. As a result, the experiments show that, whenever the memory budget is limited compared to the complexity of the data, OLCGM improves the final accuracy compared to state-of-the-art replay strategies.
Mattia Sangermano, Antonio Carta, Andrea Cossu, Davide Bacciu
IJCNN4
2022 Leveraging Relational Information for Learning Weakly Disentangled Representations
abstract
Disentanglement is a difficult property to enforce in neural representations. This might be due, in part, to a formalization of the disentanglement problem that focuses too heavily on separating relevant factors of variation of the data in single isolated dimensions of the neural representation. We argue that such a definition might be too restrictive and not necessarily beneficial in terms of downstream tasks. In this work, we present an alternative view over learning (weakly) disentangled representations, which leverages concepts from relational learning. We identify the regions of the latent space that correspond to specific instances of generative factors, and we learn the relationships among these regions in order to perform controlled changes to the latent codes. We also introduce a compound generative model that implements such a weak disentanglement approach. Our experiments shows that the learned representations can separate the relevant factors of variation in the data, while preserving the information needed for effectively generating high quality data samples.
Andrea Valenti, Davide Bacciu
IJCNN2
2022 Knowledge-Driven Interpretation of Convolutional Neural Networks
Riccardo Massidda, Davide Bacciu
ECML/PKDD (1)2
2022 A tensor framework for learning in structured domains
Daniele Castellana, Davide Bacciu
Neurocomputing2
2022 FADER: Fast adversarial example rejection
Francesco Crecchi, Marco Melis, Angelo Sotgiu, Davide Bacciu, Battista Biggio
Neurocomputing4
2022 Inductive-transductive learning for very sparse fashion graphs
Haris Dukic, Shahab Mokarizadeh, Georgios Deligiorgis, Pierpaolo Sepe, Davide Bacciu, Marco Trincavelli
Neurocomputing5
2022 Topographic mapping for quality inspection and intelligent filtering of smart-bracelet data
Davide Bacciu, Gioele Bertoncini, Davide Morelli
Neural Comput. Appl.1
2022 Controlling astrocyte-mediated synaptic pruning signals for schizophrenia drug repurposing with deep graph networks
abstract
Schizophrenia is a debilitating psychiatric disorder, leading to both physical and social morbidity. Worldwide 1% of the population is struggling with the disease, with 100,000 new cases annually only in the United States. Despite its importance, the goal of finding effective treatments for schizophrenia remains a challenging task, and previous work conducted expensive large-scale phenotypic screens. This work investigates the benefits of Machine Learning for graphs to optimize drug phenotypic screens and predict compounds that mitigate abnormal brain reduction induced by excessive glial phagocytic activity in schizophrenia subjects. Given a compound and its concentration as input, we propose a method that predicts a score associated with three possible compound effects, i.e., reduce, increase, or not influence phagocytosis. We leverage a high-throughput screening to prove experimentally that our method achieves good generalization capabilities. The screening involves 2218 compounds at five different concentrations. Then, we analyze the usability of our approach in a practical setting, i.e., prioritizing the selection of compounds in the SWEETLEAD library. We provide a list of 64 compounds from the library that have the most potential clinical utility for glial phagocytosis mitigation. Lastly, we propose a novel approach to computationally validate their utility as possible therapies for schizophrenia.
Alessio Gravina, Jennifer L. Wilson, Davide Bacciu, Kevin Grimes, Corrado Priami
PLoS Comput. Biol.3
2021 Deep learning for graphs
abstract
Deep learning for graphs encompasses all those neural models endowed with multiple layers of computation operating on data represented as graphs.The most common building blocks of these models are graph encoding layers, which compute a vector embedding for each node in a graph using message-passing operators.In this paper, we provide an overview of the key concepts in the field, point towards open questions, and frame the contributions of the ESANN 2021 special session into the broader context of deep learning for graphs.
Davide Bacciu, Filippo Maria Bianchi, Benjamin Paaßen, Cesare Alippi
ESANN1
2021 Continual Learning with Echo State Networks
abstract
Continual Learning (CL) refers to a learning setup where data is non stationary and the model has to learn without forgetting existing knowledge.The study of CL for sequential patterns revolves around trained recurrent networks.In this work, instead, we introduce CL in the context of Echo State Networks (ESNs), where the recurrent component is kept fixed.We provide the first evaluation of catastrophic forgetting in ESNs and we highlight the benefits in using CL strategies which are not applicable to trained recurrent models.Our results confirm the ESN as a promising model for CL and open to its use in streaming scenarios.* This work has been partially supported by the H2020 TEACHING
Andrea Cossu, Davide Bacciu, Antonio Carta, Claudio Gallicchio, Vincenzo Lomonaco
ESANN2
2021 Inductive learning for product assortment graph completion
abstract
Global retailers have assortments that contain hundreds of thousands of products that can be linked by several types of relationships like style compatibility, "bought together", "watched together", etc. Graphs are a natural representation for assortments, where products are nodes and relations are edges.Relations like style compatibility are often produced by a manual process and therefore do not cover uniformly the whole graph.We propose to use inductive learning to enhance a graph encoding style compatibility of a fashion assortment, leveraging rich node information comprising textual descriptions and visual data.Then, we show how the proposed graph enhancement improves substantially the performance on transductive tasks with a minor impact on graph sparsity.
Marco Trincavelli, Haris Dukic, Georgios Deligiorgis, Pierpaolo Sepe, Davide Bacciu
ESANN5
2021 Calliope - A Polyphonic Music Transformer
abstract
The polyphonic nature of music makes the application of deep learning to music modelling a challenging task.On the other hand, the Transformer architecture seems to be a good fit for this kind of data.In this work, we present Calliope, a novel autoencoder model based on Transformers for the efficient modelling of multi-track sequences of polyphonic music.The experiments show that our model is able to improve the state of the art on musical sequence reconstruction and generation, with remarkably good results especially on long sequences.* This research was partially supported by H2020 TAILOR (GA 952215)
Andrea Valenti, Stefano Berti, Davide Bacciu
ESANN3
2021 Graph Mixture Density Networks
abstract
We introduce the Graph Mixture Density Networks, a new family of machine learning models that can fit multimodal output distributions conditioned on graphs of arbitrary topology. By combining ideas from mixture models and graph representation learning, we address a broader class of challenging conditional density estimation problems that rely on structured data. In this respect, we evaluate our method on a new benchmark application that leverages random graphs for stochastic epidemic simulations. We show a significant improvement in the likelihood of epidemic outcomes when taking into account both multimodality and structure. The empirical analysis is complemented by two real-world regression tasks showing the effectiveness of our approach in modeling the output prediction uncertainty. Graph Mixture Density Networks open appealing research opportunities in the study of structure-dependent phenomena that exhibit non-trivial conditional output distributions.
Federico Errica, Davide Bacciu, Alessio Micheli
ICML2
2021 Modeling Edge Features with Deep Bayesian Graph Networks
abstract
We propose an extension of the Contextual Graph Markov Model, a deep and probabilistic machine learning model for graphs, to model the distribution of edge features. Our approach is architectural, as we introduce an additional Bayesian network mapping edge features into discrete states to be used by the original model. In doing so, we are also able to build richer graph representations even in the absence of edge features, which is confirmed by the performance improvements on standard graph classification benchmarks. Moreover, we successfully test our proposal in a graph regression scenario where edge features are of fundamental importance, and we show that the learned edge representation provides substantial performance improvements against the original model on three link prediction tasks. By keeping the computational complexity linear in the number of edges, the proposed model is amenable to large-scale graph processing.
Daniele Atzeni, Davide Bacciu, Federico Errica, Alessio Micheli
IJCNN2
2021 Graphgen-redux: a Fast and Lightweight Recurrent Model for labeled Graph Generation
abstract
The problem of labeled graph generation is gaining attention in the Deep Learning community. The task is challenging due to the sparse and discrete nature of graph spaces. Several approaches have been proposed in the literature, most of which require to transform the graphs into sequences that encode their structure and labels and to learn the distribution of such sequences through an auto-regressive generative model. Among this family of approaches, we focus on the Graphgen model. The preprocessing phase of Graphgen transforms graphs into unique edge sequences called Depth-First Search (DFS) codes, such that two isomorphic graphs are assigned the same DFS code. Each element of a DFS code is associated with a graph edge: specifically, it is a quintuple comprising one node identifier for each of the two endpoints, their node labels, and the edge label. Graphgen learns to generate such sequences auto-regressively and models the probability of each component of the quintuple independently. While effective, the independence assumption made by the model is too loose to capture the complex label dependencies of real-world graphs precisely. By introducing a novel graph preprocessing approach, we are able to process the labeling information of both nodes and edges jointly. The corresponding model, which we term Graphgen-redux, improves upon the generative performances of Graphgen in a wide range of datasets of chemical and social graphs. In addition, it uses approximately 78% fewer parameters than the vanilla variant and requires 50% fewer epochs of training on average.
Davide Bacciu, Marco Podda
IJCNN1
2021 Federated Reservoir Computing Neural Networks
abstract
A critical aspect in Federated Learning is the aggregation strategy for the combination of multiple models, trained on the edge, into a single model that incorporates all the knowledge in the federation. Common Federated Learning approaches for Recurrent Neural Networks (RNNs) do not provide guarantees on the predictive performance of the aggregated model. In this paper we show how the use of Echo State Networks (ESNs), which are efficient state-of-the-art RNN models for time-series processing, enables a form of federation that is optimal in the sense that it produces models mathematically equivalent to the corresponding centralized model. Furthermore, the proposed method is compliant with privacy constraints. The proposed method, which we denote as Incremental Federated Learning, is experimentally evaluated against an averaging strategy on two datasets for human state and activity recognition.
Davide Bacciu, Daniele Di Sarli, Pouria Faraji, Claudio Gallicchio, Alessio Micheli
IJCNN1
2021 K-plex cover pooling for graph neural networks
abstract
Abstract Graph pooling methods provide mechanisms for structure reduction that are intended to ease the diffusion of context between nodes further in the graph, and that typically leverage community discovery mechanisms or node and edge pruning heuristics. In this paper, we introduce a novel pooling technique which borrows from classical results in graph theory that is non-parametric and generalizes well to graphs of different nature and connectivity patterns. Our pooling method, namedKPlexPool, builds on the concepts of graph covers andk-plexes, i.e. pseudo-cliques where each node can miss up toklinks. The experimental evaluation on benchmarks on molecular and social graph classification shows thatKPlexPoolachieves state of the art performances against both parametric and non-parametric pooling methods in the literature, despite generating pooled graphs based solely on topological information.
Davide Bacciu, Alessio Conte, Roberto Grossi, Francesco Landolfi, Andrea Marino 0001
Data Min. Knowl. Discov.1
2021 Encoding-based memory for recurrent neural networks
Antonio Carta, Alessandro Sperduti, Davide Bacciu
Neurocomputing3
2021 Continual learning for recurrent neural networks: An empirical evaluation
Andrea Cossu, Antonio Carta, Vincenzo Lomonaco, Davide Bacciu
Neural Networks4
2020 A Deep Generative Model for Fragment-Based Molecule Generation
abstract
Molecule generation is a challenging open problem in cheminformatics. Currently, deep generative approaches addressing the challenge belong to two broad categories, differing in how molecules are represented. One approach encodes molecular graphs as strings of text, and learns their corresponding character-based language model. Another, more expressive, approach operates directly on the molecular graph. In this work, we address two limitations of the former: generation of invalid and duplicate molecules. To improve validity rates, we develop a language model for small molecular substructures called fragments, loosely inspired by the well-known paradigm of Fragment-Based Drug Design. In other words, we generate molecules fragment by fragment, instead of atom by atom. To improve uniqueness rates, we present a frequency-based masking strategy that helps generate molecules with infrequent fragments. We show experimentally that our model largely outperforms other language model-based competitors, reaching state-of-the-art performances typical of graph-based approaches. Moreover, generated molecules display molecular properties similar to those in the training sample, even in absence of explicit task-specific supervision.
Marco Podda, Davide Bacciu, Alessio Micheli
AISTATS2
2020 Learning from Non-Binary Constituency Trees via Tensor Decomposition
abstract
Processing sentence constituency trees in binarised form is a common and popular approach in literature.However, constituency trees are non-binary by nature.The binarisation procedure changes deeply the structure, furthering constituents that instead are close.In this work, we introduce a new approach to deal with non-binary constituency trees which leverages tensor-based models.In particular, we show how a powerful composition function based on the canonical tensor decomposition can exploit such a rich structure.A key point of our approach is the weight sharing constraint imposed on the factor matrices, which allows limiting the number of model parameters.Finally, we introduce a Tree-LSTM model which takes advantage of this composition function and we experimentally assess its performance on different NLP tasks.
Daniele Castellana, Davide Bacciu
COLING2
2020 Learning Style-Aware Symbolic Music Representations by Adversarial Autoencoders
abstract
We address the challenging open problem of learning an effective latent space for symbolic music data in generative music modeling. We focus on leveraging adversarial regularization as a flexible and natural mean to imbue variational autoencoders with context information concerning music genre and style. Through the paper, we show how Gaussian mixtures taking into account music metadata information can be used as an effective prior for the autoencoder latent space, introducing the first Music Adversarial Autoencoder (MusAE). The empirical analysis on a large scale benchmark shows that our model has a higher reconstruction accuracy than state-of-the-art models based on standard variational autoencoders. It is also able to create realistic interpolations between two musical sequences, smoothly changing the dynamics of the different tracks. Experiments show that the model can organise its latent space accordingly to low-level properties of the musical pieces, as well as to embed into the latent variables the high-level genre information injected from the prior distribution to increase its overall performance. This allows us to perform changes to the generated pieces in a principled way.
Andrea Valenti, Antonio Carta, Davide Bacciu
ECAI3
2020 Tensor Decompositions in Deep Learning
Davide Bacciu, Danilo P. Mandic
ESANN1
2020 Tensor Decompositions in Recursive Neural Networks for Tree-Structured Data
Daniele Castellana, Davide Bacciu
ESANN2
2020 Perplexity-free Parametric t-SNE
Francesco Crecchi, Cyril de Bodt, Michel Verleysen, John A. Lee 0001, Davide Bacciu
ESANN5
2020 Theoretically Expressive and Edge-aware Graph Learning
Federico Errica, Davide Bacciu, Alessio Micheli
ESANN2
2020 Biochemical Pathway Robustness Prediction with Graph Neural Networks
Marco Podda, Alessio Micheli, Davide Bacciu, Paolo Milazzo
ESANN3
2020 A Fair Comparison of Graph Neural Networks for Graph Classification
Federico Errica, Marco Podda, Davide Bacciu, Alessio Micheli
ICLR3
2020 Generalising Recursive Neural Models by Tensor Decomposition
abstract
Most machine learning models for structured data encode the structural knowledge of a node by leveraging simple aggregation functions (in neural models, typically a weighted sum) of the information in the node's neighbourhood. Nevertheless, the choice of simple context aggregation functions, such as the sum, can be widely sub-optimal. In this work we introduce a general approach to model aggregation of structural context leveraging a tensor-based formulation. We show how the exponential growth in the size of the parameter space can be controlled through an approximation based on the Tucker tensor decomposition. This approximation allows limiting the parameters space size, decoupling it from its strict relation with the size of the hidden encoding space. By this means, we can effectively regulate the trade-off between expressivity of the encoding, controlled by the hidden size, computational complexity and model generalisation, influenced by parameterisation. Finally, we introduce a new Tensorial Tree-LSTM derived as an instance of our framework and we use it to experimentally assess our working hypotheses on tree classification scenarios.
Daniele Castellana, Davide Bacciu
IJCNN2
2020 Continual Learning with Gated Incremental Memories for sequential data processing
abstract
The ability to learn in dynamic, nonstationary environments without forgetting previous knowledge, also known as Continual Learning (CL), is a key enabler for scalable and trustworthy deployments of adaptive solutions. While the importance of continual learning is largely acknowledged in machine vision and reinforcement learning problems, this is mostly under-documented for sequence processing tasks. This work proposes a Recurrent Neural Network (RNN) model for CL that is able to deal with concept drift in input distribution without forgetting previously acquired knowledge. We also implement and test a popular CL approach, Elastic Weight Consolidation (EWC), on top of two different types of RNNs. Finally, we compare the performances of our enhanced architecture against EWC and RNNs on a set of standard CL benchmarks, adapted to the sequential data processing scenario. Results show the superior performance of our architecture and highlight the need for special solutions designed to address CL in RNNs.
Andrea Cossu, Antonio Carta, Davide Bacciu
IJCNN3
2020 Incremental Training of a Recurrent Neural Network Exploiting a Multi-scale Dynamic Memory
Antonio Carta, Alessandro Sperduti, Davide Bacciu
ECML/PKDD (1)3
2020 ROS-Neuro Integration of Deep Convolutional Autoencoders for EEG Signal Compression in Real-time BCIs
abstract
Typical EEG-based BCI applications require the computation of complex functions over the noisy EEG channels to be carried out in an efficient way. Deep learning algorithms are capable of learning flexible nonlinear functions directly from data, and their constant processing latency is perfect for their deployment into online BCI systems. However, it is crucial for the jitter of the processing system to be as low as possible, in order to avoid unpredictable behaviour that can ruin the system's overall usability. In this paper, we present a novel encoding method, based on on deep convolutional autoencoders, that is able to perform efficient compression of the raw EEG inputs. We deploy our model in a ROS-Neuro node, thus making it suitable for the integration in ROS-based BCI and robotic systems in real world scenarios. The experimental results show that our system is capable to generate meaningful compressed encoding preserving to original information contained in the raw input. They also show that the ROS-Neuro node is able to produce such encodings at a steady rate, with minimal jitter. We believe that our system can represent an important step towards the development of an effective BCI processing pipeline fully standardized in ROS-Neuro framework.
Andrea Valenti, Michele Barsotti, Raffaello Brondi, Davide Bacciu, Luca Ascari
SMC4
2020 Measuring the effects of confounders in medical supervised classification problems: the Confounding Index (CI)
Elisa Ferrari, Alessandra Retico, Davide Bacciu
Artif. Intell. Medicine3
2020 Edge-based sequential graph generation with recurrent neural networks
Davide Bacciu, Alessio Micheli, Marco Podda
Neurocomputing1
2020 Probabilistic Learning on Graphs via Contextual Architectures
abstract
We propose a novel methodology for representation learning on graph-structured data, in which a stack of Bayesian Networks learns different distributions of a vertex's neighbourhood. Through an incremental construction policy and layer-wise training, we can build deeper architectures with respect to typical graph convolutional neural networks, with benefits in terms of context spreading between vertices. First, the model learns from graphs via maximum likelihood estimation without using target labels. Then, a supervised readout is applied to the learned graph embeddings to deal with graph classification and vertex classification tasks, showing competitive results against neural models for graphs. The computational complexity is linear in the number of edges, facilitating learning on large scale data sets. By studying how depth affects the performances of our model, we discover that a broader context generally improves performances. In turn, this leads to a critical analysis of some benchmarks used in literature.
Davide Bacciu, Federico Errica, Alessio Micheli
J. Mach. Learn. Res.1
2020 A gentle introduction to deep learning for graphs
Davide Bacciu, Federico Errica, Alessio Micheli, Marco Podda
Neural Networks1
2020 Augmenting Recurrent Neural Networks Resilience by Dropout
abstract
This brief discusses the simple idea that dropout regularization can be used to efficiently induce resiliency to missing inputs at prediction time in a generic neural network. We show how the approach can be effective on tasks where imputation strategies often fail, namely, involving recurrent neural networks and scenarios where whole sequences of input observations are missing. The experimental analysis provides an assessment of the accuracy-resiliency tradeoff in multiple recurrent models, including reservoir computing methods, and comprising real-world ambient intelligence and biomedical time series.
Davide Bacciu, Francesco Crecchi
IEEE Trans. Neural Networks Learn. Syst.1
2019 Societal Issues in Machine Learning: When Learning from Data is Not Enough
Davide Bacciu, Battista Biggio, Paulo J. G. Lisboa, José D. Martín, Luca Oneto, Alfredo Vellido
ESANN1
2019 Graph generation by sequential edge prediction
Davide Bacciu, Alessio Micheli, Marco Podda
ESANN1
2019 Detecting Black-box Adversarial Examples through Nonlinear Dimensionality Reduction
Francesco Crecchi, Davide Bacciu, Battista Biggio
ESANN2
2019 Linear Memory Networks
Davide Bacciu, Antonio Carta, Alessandro Sperduti
ICANN (1)1
2019 Bayesian Tensor Factorisation for Bottom-up Hidden Tree Markov Models
abstract
Bottom-Up Hidden Tree Markov Model is a highly expressive model for tree-structured data. Unfortunately, it cannot be used in practice due to the intractable size of its state-transition matrix. We propose a new approximation which lies on the Tucker factorisation of tensors. The probabilistic interpretation of such approximation allows us to define a new probabilistic model for tree-structured data. Hence, we define the new approximated model and we derive its learning algorithm. Then, we empirically assess the effective power of the new model evaluating it on two different tasks. In both cases, our model outperforms the other approximated model known in the literature.
Daniele Castellana, Davide Bacciu
IJCNN2
2019 An ambient intelligence approach for learning in smart robotic environments
abstract
Abstract Smart robotic environments combine traditional (ambient) sensing devices and mobile robots. This combination extends the type of applications that can be considered, reduces their complexity, and enhances the individual values of the devices involved by enabling new services that cannot be performed by a single device. To reduce the amount of preparation and preprogramming required for their deployment in real‐world applications, it is important to make these systems self‐adapting. The solution presented in this paper is based upon a type of compositional adaptation where (possibly multiple) plans of actions are created through planning and involve the activation of pre‐existing capabilities. All the devices in the smart environment participate in a pervasive learning infrastructure, which is exploited to recognize which plans of actions are most suited to the current situation. The system is evaluated in experiments run in a real domestic environment, showing its ability to proactively and smoothly adapt to subtle changes in the environment and in the habits and preferences of their user(s), in presence of appropriately defined performance measuring functions.
Davide Bacciu, Maurizio Di Rocco, Mauro Dragone, Claudio Gallicchio, Alessio Micheli, Alessandro Saffiotti
Comput. Intell.1
2019 Bayesian mixtures of Hidden Tree Markov Models for structured data clustering
Davide Bacciu, Daniele Castellana
Neurocomputing1
2018 Mixture of Hidden Markov Model as Tree Encoder
Davide Bacciu, Daniele Castellana
ESANN1
2018 Bioinformatics and medicine in the era of deep learning
Davide Bacciu, Paulo J. G. Lisboa, José D. Martín, Ruxandra Stoean, Alfredo Vellido
ESANN1
2018 Contextual Graph Markov Model: A Deep and Generative Approach to Graph Processing
abstract
We introduce the Contextual Graph Markov Model, an approach combining ideas from generative models and neural networks for the processing of graph data. It founds on a constructive methodology to build a deep architecture comprising layers of probabilistic models that learn to encode the structured information in an incremental fashion. Context is diffused in an efficient and scalable way across the graph vertexes and edges. The resulting graph encoding is used in combination with discriminative models to address structure classification benchmarks.
Davide Bacciu, Federico Errica, Alessio Micheli
ICML1
2018 Concentric ESN: Assessing the Effect of Modularity in Cycle Reservoirs
abstract
The paper introduces concentric Echo State Network, an approach to design reservoir topologies that tries to bridge the gap between deterministically constructed simple cycle models and deep reservoir computing approaches. We show how to modularize the reservoir into simple unidirectional and concentric cycles with pairwise bidirectional jump connections between adjacent loops. We provide a preliminary experimental assessment showing how concentric reservoirs yield to superior predictive accuracy and memory capacity with respect to single cycle reservoirs and deep reservoir models.
Davide Bacciu, Andrea Bongiorno
IJCNN1
2018 Randomized neural networks for preference learning with physiological data
Davide Bacciu, Michele Colombo, Davide Morelli, David Plans
Neurocomputing1
2018 Generative Kernels for Tree-Structured Data
abstract
This paper presents a family of methods for the design of adaptive kernels for tree-structured data that exploits the summarization properties of hidden states of hidden Markov models for trees. We introduce a compact and discriminative feature space based on the concept of hidden states multisets and we discuss different approaches to estimate such hidden state encoding. We show how it can be used to build an efficient and general tree kernel based on Jaccard similarity. Furthermore, we derive an unsupervised convolutional generative kernel using a topology induced on the Markov states by a tree topographic mapping. This paper provides an extensive empirical assessment on a variety of structured data learning tasks, comparing the predictive accuracy and computational efficiency of state-of-the-art generative, adaptive, and syntactical tree kernels. The results show that the proposed generative approach has a good tradeoff between computational complexity and predictive performance, in particular when considering the soft matching introduced by the topographic mapping.
Davide Bacciu, Alessio Micheli, Alessandro Sperduti
IEEE Trans. Neural Networks Learn. Syst.1
2017 ELM Preference Learning for Physiological Data
Davide Bacciu, Michele Colombo, Davide Morelli, David Plans
ESANN1
2017 DropIn: Making reservoir computing neural networks robust to missing inputs by dropout
abstract
The paper presents a novel, principled approach to train recurrent neural networks from the Reservoir Computing family that are robust to missing part of the input features at prediction time. By building on the ensembling properties of Dropout regularization, we propose a methodology, named DropIn, which efficiently trains a neural model as a committee machine of subnetworks, each capable of predicting with a subset of the original input features. We discuss the application of the DropIn methodology in the context of Reservoir Computing models and targeting applications characterized by input sources that are unreliable or prone to be disconnected, such as in pervasive wireless sensor networks and ambient intelligence. We provide an experimental assessment using real-world data from such application domains, showing how the Dropin methodology allows to maintain predictive performances comparable to those of a model without missing features, even when 20%–50% of the inputs are not available.
Davide Bacciu, Francesco Crecchi, Davide Morelli
IJCNN1
2017 A learning system for automatic Berg Balance Scale score estimation
Davide Bacciu, Stefano Chessa, Claudio Gallicchio, Alessio Micheli, Luca Pedrelli, Erina Ferro, Luigi Fortunati, Davide La Rosa, Filippo Palumbo, Federico Vozzi, Oberdan Parodi
Eng. Appl. Artif. Intell.1
2016 A reservoir activation kernel for trees
Davide Bacciu, Claudio Gallicchio, Alessio Micheli
ESANN1
2016 Detecting Socialization Events in Ageing People: The Experience of the DOREMI Project
abstract
The detection of socialization events is useful to build indicators about social isolation of people, which is an important indicator in e-health applications. On the other hand, it is rather difficult to achieve with non-invasive solutions. This paper reports about the currently work-in-progress on the technological solution for the detection of socialization events adopted in the DOREMI project.
Davide Bacciu, Stefano Chessa, Erina Ferro, Luigi Fortunati, Claudio Gallicchio, Davide La Rosa, Miguel Llorente, Alessio Micheli, Filippo Palumbo, Oberdan Parodi, Andrea Valenti, Federico Vozzi
Intelligent Environments1
2016 Unsupervised feature selection for sensor time-series in pervasive computing applications
Davide Bacciu
Neural Comput. Appl.1
2015 ESNigma: efficient feature selection for echo state networks
Davide Bacciu, Filippo Benedetti, Alessio Micheli
ESANN1
2015 A cognitive robotic ecology approach to self-configuring and evolving AAL systems
Mauro Dragone, Giuseppe Amato 0001, Davide Bacciu, Stefano Chessa, Sonya A. Coleman, Maurizio Di Rocco, Claudio Gallicchio, Claudio Gennaro, Héctor Lozano Peiteado, Liam P. Maguire, T. Martin McGinnity, Alessio Micheli, Gregory M. P. O'Hare, Arantxa Rentería, Alessandro Saffiotti, Claudio Vairo, Philip J. Vance
Eng. Appl. Artif. Intell.3
2014 An Iterative Feature Filter for Sensor Timeseries in Pervasive Computing Applications
Davide Bacciu
EANN1
2014 Modeling Bi-directional Tree Contexts by Generative Transductions
Davide Bacciu, Alessio Micheli, Alessandro Sperduti
ICONIP (1)1
2014 Integrating bi-directional contexts in a generative kernel for trees
abstract
Context is essential to evaluate an atomic piece of information composing an articulated structured sample. A particular context captures different structural information with respect to an alternative context. The paper introduces a generative kernel that easily and effectively combines the structural information captured by generative tree models characterized by different contextual capabilities. The proposed approach exploits the idea of hidden states multisets to realize a tree encoding that takes into account both the summarized information on the path leading to a node (i.e. a top-down context) as well as the information on how substructures are composed to create a subtree rooted on a node (bottom-up context). An thorough experimental analysis is provided, showing that the bi-directional approach incorporating top-down and bottom-up contexts yields to superior performances with respect to the unidirectional contexts alone, achieving state of the art results on challenging tree classification benchmarks.
Davide Bacciu, Alessio Micheli, Alessandro Sperduti
IJCNN1
2014 An experimental characterization of reservoir computing in ambient assisted living applications
Davide Bacciu, Paolo Barsocchi, Stefano Chessa, Claudio Gallicchio, Alessio Micheli
Neural Comput. Appl.1
2013 An input-output hidden Markov model for tree transductions
Davide Bacciu, Alessio Micheli, Alessandro Sperduti
Neurocomputing1
2013 Compositional Generative Mapping for Tree-Structured Data - Part II: Topographic Projection Model
abstract
We introduce GTM-SD (Generative Topographic Mapping for Structured Data), which is the first compositional generative model for topographic mapping of tree-structured data. GTM-SD exploits a scalable bottom-up hidden-tree Markov model that was introduced in Part I of this paper to achieve a recursive topographic mapping of hierarchical information. The proposed model allows efficient exploitation of contextual information from shared substructures by a recursive upward propagation on the tree structure which distributes substructure information across the topographic map. Compared to its noncompositional generative counterpart, GTM-SD is shown to allow the topographic mapping of the full sample tree, which includes a projection onto the lattice of all the distinct subtrees rooted in each of its nodes. Experimental results show that the continuous projection space generated by the smooth topographic mapping of GTM-SD yields a finer grained discrimination of the sample structures with respect to the state-of-the-art recursive neural network approach.
Davide Bacciu, Alessio Micheli, Alessandro Sperduti
IEEE Trans. Neural Networks Learn. Syst.1
2012 Input-Output Hidden Markov Models for trees
Davide Bacciu, Alessio Micheli, Alessandro Sperduti
ESANN1
2012 A Generative Multiset Kernel for Structured Data
Davide Bacciu, Alessio Micheli, Alessandro Sperduti
ICANN (1)1
2012 Compositional Generative Mapping for Tree-Structured Data - Part I: Bottom-Up Probabilistic Modeling of Trees
abstract
We introduce a novel compositional (recursive) probabilistic model for trees that defines an approximated bottom-up generative process from the leaves to the root of a tree. The proposed model defines contextual state transitions from the joint configuration of the children to the parent nodes. We argue that the bottom-up context postulates different probabilistic assumptions with respect to a top-down approach, leading to different representational capabilities. We discuss classes of applications that are best suited to a bottom-up approach. In particular, the bottom-up context is shown to better correlate and model the co-occurrence of substructures among the child subtrees of internal nodes. A mixed memory approximation is introduced to factorize the joint children-to-parent state transition matrix as a mixture of pairwise transitions. The proposed approach is the first practical bottom-up generative model for tree-structured data that maintains the same computational class of its top-down counterpart. Comparative experimental analyses exploiting synthetic and real-world datasets show that the proposed model can deal with deep structures better than a top-down generative model. The model is also shown to better capture structural information from real-world data comprising trees with a large out-degree. The proposed bottom-up model can be used as a fundamental building block for the development of other new powerful models.
Davide Bacciu, Alessio Micheli, Alessandro Sperduti
IEEE Trans. Neural Networks Learn. Syst.1
2011 Adaptive tree kernel by multinomial generative topographic mapping
abstract
Learning the kernel function from data is a challenging open issue in structured data processing. In the paper, we propose a novel adaptive kernel, defined over a generative learning model, that exploits a novel multinomial extension of the Generative Topographic Mapping for Structured Data (GTM-SD). We show how the proposed kernel effectively exploits the GTM-SD continuity and smoothness properties to provide dense kernels characterized by an high discriminative power even with small topographic maps. Experimental evaluations on challenging structured XML document repositories show the effectiveness of the proposed approach against state-of-the-art syntactic and adaptive convolutional kernels.
Davide Bacciu, Alessio Micheli, Alessandro Sperduti
IJCNN1
2011 Clustering of protein expression data: a benchmark of statistical and neural approaches
Ian H. Jarman, Terence A. Etchells, Davide Bacciu, Jonathan M. Garibaldi, Ian O. Ellis, Paulo J. G. Lisboa
Soft Comput.3
2010 Bottom-Up Generative Modeling of Tree-Structured Data
Davide Bacciu, Alessio Micheli, Alessandro Sperduti
ICONIP (1)1
2010 Compositional generative mapping of structured data
abstract
We introduce a compositional generative model for topographic mapping of tree-structured data. It exploits a scalable bottom-up hidden tree Markov model to achieve a recursive topographic mapping of hierarchical information. The model allows for an efficient exploitation of contextual information from shared substructures by recursive upward propagation on the tree structure and by allowing it to distribute across the map. Experimental results show that the model yields to a topographically ordered mapping of the substructures in the input data.
Davide Bacciu, Alessio Micheli, Alessandro Sperduti
IJCNN1
2009 Patient stratification with competing risks by multivariate Fisher distance
abstract
Early characterization of patients with respect to their predicted response to treatment is a fundamental step towards the delivery of effective, personalized care. Starting from the results of a time-to-event model with competing risks using the framework of partial logistic artificial neural networks with automatic relevance determination (PLANNCR-ARD), we discuss an effective semi-supervised approach to patient stratification with application to Acute Myeloid Leukaemia (AML) data (n = 509) acquired prospectively by the GIMEMA consortium. Multiple prognostic indices provided by the survival model are exploited to build a metric based on the Fisher information matrix. Cluster number estimation is then performed in the Fisher-induced affine space, yielding to the discovery of a stratification of the patients into groups characterized by significantly different mortality risks following induction therapy in AML. The proposed model is shown to be able to cluster the input data, while promoting specificity of both target outcomes, namely Complete Remission (CR) and Induction Death (ID). This generic clustering methodology generates an affine transformation of the data space that is coherent with the prognostic information predicted by the PLANNCR-ARD model.
Davide Bacciu, Ian H. Jarman, Terence A. Etchells, Paulo J. G. Lisboa
IJCNN1
2009 Expansive competitive learning for kernel vector quantization
Davide Bacciu, Antonina Starita
Pattern Recognit. Lett.1
2008 Are Model-Based Clustering and Neural Clustering Consistent? A Case Study from Bioinformatics
Davide Bacciu, Elia Biganzoli, Paulo J. G. Lisboa, Antonina Starita
KES (2)1
2008 Competitive Repetition Suppression (CoRe) Clustering: A Biologically Inspired Learning Model With Application to Robust Clustering
abstract
Determining a compact neural coding for a set of input stimuli is an issue that encompasses several biological memory mechanisms as well as various artificial neural network models. In particular, establishing the optimal network structure is still an open problem when dealing with unsupervised learning models. In this paper, we introduce a novel learning algorithm, named competitive repetition-suppression (CoRe) learning, inspired by a cortical memory mechanism called repetition suppression (RS). We show how such a mechanism is used, at various levels of the cerebral cortex, to generate compact neural representations of the visual stimuli. From the general CoRe learning model, we derive a clustering algorithm, named CoRe clustering, that can automatically estimate the unknown cluster number from the data without using a priori information concerning the input distribution. We illustrate how CoRe clustering, besides its biological plausibility, posses strong theoretical properties in terms of robustness to noise and outliers, and we provide an error function describing CoRe learning dynamics. Such a description is used to analyze CoRe relationships with the state-of-the art clustering models and to highlight CoRe similitude with rival penalized competitive learning (RPCL), showing how CoRe extends such a model by strengthening the rival penalization estimation by means of loss functions from robust statistics.
Davide Bacciu, Antonina Starita
IEEE Trans. Neural Networks1
2007 Convergence Behavior of Competitive Repetition-Suppression Clustering
Davide Bacciu, Antonina Starita
ICONIP (1)1
2007 A Robust Bio-Inspired Clustering Algorithm for the Automatic Determination of Unknown Cluster Number
abstract
The paper introduces a robust clustering algorithm that can automatically determine the unknown cluster number from noisy data without any a-priori information. We show how our clustering algorithm can be derived from a general learning theory, named CoRe learning, that models a cortical memory mechanism called repetition suppression. Moreover, we describe CoRe clustering relationships with Rival Penalized Competitive Learning (RPCL), showing how CoRe extends this model by strengthening the rival penalization estimation by means of robust loss functions. Finally, we present the results of simulations concerning the unsupervised segmentation of noisy images.
Davide Bacciu, Antonina Starita
IJCNN1
2006 Competitive Repetition-suppression (CoRe) Learning
Davide Bacciu, Antonina Starita
ICANN (1)1
2004 A RLWPR network for learning the internal model of an anthropomorphic robot arm
abstract
Studies of human motor control suggest that humans develop internal models of the arm during the execution of voluntary movements. In particular, the internal model consists of the inverse dynamic model of the musculoskeletal system and intervenes in the feedforward loop of the motor control system to improve reactivity and stability in rapid movements. In this paper, an interaction control scheme inspired by biological motor control is resumed, i.e. the coactivation-based compliance control in the joint space (Zollo, L, et al., 2003), and a feedforward module capable of online learning the manipulator inverse dynamics is presented. A novel recurrent learning paradigm is proposed which derives from an interesting functional equivalence between locally weighted regression networks and Takagi-Sugeno-Kang fuzzy systems. The proposed learning paradigm has been named recurrent locally weighted regression networks and strengthens the computational power of feedforward locally weighted regression networks. Simulation results are reported to validate the control scheme.
Davide Bacciu, Loredana Zollo, Eugenio Guglielmelli, Fabio Leoni, Antonina Starita
IROS1