Francesco Alesiani

dblp:122/8256 · DBLP profile ↗
← Back
19ranked-venue papers
8as first author
12since 2021 · last 2025
0000-0003-4413-7247ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 3 since 2021Computer networks · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Adaptive Message Passing: A General Framework to Mitigate Oversmoothing, Oversquashing, and Underreaching
abstract
Long-range interactions are essential for the correct description of complex systems in many scientific fields. The price to pay for including them in the calculations, however, is a dramatic increase in the overall computational costs. Recently, deep graph networks have been employed as efficient, data-driven models for predicting properties of complex systems represented as graphs. These models rely on a message passing strategy that should, in principle, capture long-range information without explicitly modeling the corresponding interactions. In practice, most deep graph networks cannot really model long-range dependencies due to the intrinsic limitations of (synchronous) message passing, namely oversmoothing, oversquashing, and underreaching. This work proposes a general framework that learns to mitigate these limitations: within a variational inference framework, we endow message passing architectures with the ability to adapt their depth and filter messages along the way. With theoretical and empirical arguments, we show that this strategy better captures long-range interactions, by competing with the state of the art on five node and graph prediction datasets.
Federico Errica, Henrik Christiansen, Viktor Zaverkin, Takashi Maruyama, Mathias Niepert, Francesco Alesiani
ICML6
2024 Higher-Rank Irreducible Cartesian Tensors for Equivariant Message Passing
abstract
The ability to perform fast and accurate atomistic simulations is crucial for advancing the chemical sciences. By learning from high-quality data, machine-learned interatomic potentials achieve accuracy on par with ab initio and first-principles methods at a fraction of their computational cost. The success of machine-learned interatomic potentials arises from integrating inductive biases such as equivariance to group actions on an atomic system, e.g., equivariance to rotations and reflections. In particular, the field has notably advanced with the emergence of equivariant message passing. Most of these models represent an atomic system using spherical tensors, tensor products of which require complicated numerical coefficients and can be computationally demanding. Cartesian tensors offer a promising alternative, though state-of-the-art methods lack flexibility in message-passing mechanisms, restricting their architectures and expressive power. This work explores higher-rank irreducible Cartesian tensors to address these limitations. We integrate irreducible Cartesian tensor products into message-passing neural networks and prove the equivariance and traceless property of the resulting layers. Through empirical evaluations on various benchmark data sets, we consistently observe on-par or better performance than that of state-of-the-art spherical and Cartesian models.
Viktor Zaverkin, Francesco Alesiani, Takashi Maruyama, Federico Errica, Henrik Christiansen, Makoto Takamoto, Nicolas Weber, Mathias Niepert
NeurIPS2
2023 Implicit Bilevel Optimization: Differentiating through Bilevel Optimization Programming
abstract
Bilevel Optimization Programming is used to model complex and conflicting interactions between agents, for example in Robust AI or Privacy preserving AI. Integrating bilevel mathematical programming within deep learning is thus an essential objective for the Machine Learning community. Previously proposed approaches only consider single-level programming. In this paper, we extend existing single-level optimization programming approaches and thus propose Differentiating through Bilevel Optimization Programming (BiGrad) for end-to-end learning of models that use Bilevel Programming as a layer. BiGrad has wide applicability and can be used in modern machine learning frameworks. BiGrad is applicable to both continuous and combinatorial Bilevel optimization problems. We describe a class of gradient estimators for the combinatorial case which reduces the requirements in terms of computation complexity; for the case of the continuous variable, the gradient computation takes advantage of the push-back approach (i.e. vector-jacobian product) for an efficient implementation. Experiments show that the BiGrad successfully extends existing single-level approaches to Bilevel Programming.
Francesco Alesiani
AAAI1
2023 Learning Neural PDE Solvers with Parameter-Guided Channel Attention
abstract
Scientific Machine Learning (SciML) is concerned with the development of learned emulators of physical systems governed by partial differential equations (PDE). In application domains such as weather forecasting, molecular dynamics, and inverse design, ML-based surrogate models are increasingly used to augment or replace inefficient and often non-differentiable numerical simulation algorithms. While a number of ML-based methods for approximating the solutions of PDEs have been proposed in recent years, they typically do not adapt to the parameters of the PDEs, making it difficult to generalize to PDE parameters not seen during training. We propose a Channel Attention guided by PDE Parameter Embeddings (CAPE) component for neural surrogate models and a simple yet effective curriculum learning strategy. The CAPE module can be combined with any neural PDE solvers allowing them to adapt to unseen PDE parameters. The curriculum learning strategy provides a seamless transition between teacher-forcing and fully auto-regressive training. We compare CAPE in conjunction with the curriculum learning strategy using a PDE benchmark and obtain consistent and significant improvements over the baseline models. The experiments also show several advantages of CAPE, such as its increased ability to generalize to unseen PDE parameters without large increases inference time and parameter count. An implementation of the method and experiments are available at https://anonymous.4open.science/r/CAPE-ML4Sci-145B.
Makoto Takamoto, Francesco Alesiani, Mathias Niepert
ICML2
2023 Gated information bottleneck for generalization in sequential environments
Francesco Alesiani, Shujian Yu
Knowl. Inf. Syst.1
2022 Modular-Relatedness for Continual Learning
Ammar Shaker, Francesco Alesiani, Shujian Yu
IDA2
2022 PDEBench: An Extensive Benchmark for Scientific Machine Learning
abstract
Machine learning-based modeling of physical systems has experienced increased interest in recent years. Despite some impressive progress, there is still a lack of benchmarks for Scientific ML that are easy to use but still challenging and repre- sentative of a wide range of problems. We introduce PDEBENCH, a benchmark suite of time-dependent simulation tasks based on Partial Differential Equations (PDEs). PDEBENCH comprises both code and data to benchmark the performance of novel machine learning models against both classical numerical simulations and machine learning baselines. Our proposed set of benchmark problems con- tribute the following unique features: (1) A much wider range of PDEs compared to existing benchmarks, ranging from relatively common examples to more real- istic and difficult problems; (2) much larger ready-to-use datasets compared to prior work, comprising multiple simulation runs across a larger number of ini- tial and boundary conditions and PDE parameters; (3) more extensible source codes with user-friendly APIs for data generation and baseline results with popular machine learning models (FNO, U-Net, PINN, Gradient-Based Inverse Method). PDEBENCH allows researchers to extend the benchmark freely for their own pur- poses using a standardized API and to compare the performance of new models to existing baseline methods. We also propose new evaluation metrics with the aim to provide a more holistic understanding of learning methods in the context of Scientific ML. With those metrics we identify tasks which are challenging for recent ML methods and propose these tasks as future challenges for the community. The code is available at https://github.com/pdebench/PDEBench.
Makoto Takamoto, Timothy Praditia, Raphael Leiteritz, Daniel MacKinlay, Francesco Alesiani, Dirk Pflüger, Mathias Niepert
NeurIPS5
2022 Principle of relevant information for graph sparsification
abstract
Graph sparsification aims to reduce the number of edges of a graph while maintaining its structural properties. In this paper, we propose the first general and effective information-theoretic formulation of graph sparsification, by taking inspiration from the Principle of Relevant Information (PRI). To this end, we extend the PRI from a standard scalar random variable setting to structured data (i.e., graphs). Our Graph-PRI objective is achieved by operating on the graph Laplacian, made possible by expressing the graph Laplacian of a subgraph in terms of a sparse edge selection vector w. We provide both theoretical and empirical justifications on the validity of our Graph-PRI approach. We also analyze its analytical solutions in a few special cases. We finally present three representative real-world applications, namely graph sparsification, graph regularized multi-task learning, and medical imaging-derived brain network classification, to demonstrate the effectiveness, the versatility and the enhanced interpretability of our approach over prevalent sparsification techniques. Code of Graph-PRI is available at https://github.com/SJYuCNEL/PRI-Graphs.
Shujian Yu, Francesco Alesiani, Wenzhe Yin, Robert Jenssen, José C. Príncipe
UAI2
2021 Measuring Dependence with Matrix-based Entropy Functional
abstract
Measuring the dependence of data plays a central role in statistics and machine learning. In this work, we summarize and generalize the main idea of existing information-theoretic dependence measures into a higher-level perspective by the Shearer's inequality. Based on our generalization, we then propose two measures, namely the matrix-based normalized total correlation and the matrix-based normalized dual total correlation, to quantify the dependence of multiple variables in arbitrary dimensional space, without explicit estimation of the underlying data distributions. We show that our measures are differentiable and statistically more powerful than prevalent ones. We also show the impact of our measures in four different machine learning problems, namely the gene regulatory network inference, the robust machine learning under covariate shift and non-Gaussian noises, the subspace outlier detection, and the understanding of the learning dynamics of convolutional neural networks, to demonstrate their utilities, advantages, as well as implications to those problems.
Shujian Yu, Francesco Alesiani, Robert Jenssen, José C. Príncipe
AAAI2
2021 Gated Information Bottleneck for Generalization in Sequential Environments
abstract
Deep neural networks suffer from poor generalization to unseen environments when the underlying data distribution is different from that in the training set. By learning minimum sufficient representations from training data, the information bottleneck (IB) approach has demonstrated its effectiveness to improve generalization in different AI applications. In this work, we propose a new neural network-based IB approach, termed gated information bottleneck (GIB), that dynamically drops spurious correlations and progressively selects the most task-relevant features across different environments by a trainable soft mask (on raw features). GIB enjoys a simple and tractable objective, without any variational approximation or distributional assumption. We empirically demonstrate the superiority of GIB over other popular neural network-based IB approaches in adversarial robustness and out-of-distribution (OOD) detection. Meanwhile, we also establish the connection between IB theory and invariant causal representation learning, and observed that GIB demonstrates appealing performance when different environments arrive sequentially, a more practical scenario where invariant risk minimization (IRM) fails.
Francesco Alesiani, Shujian Yu
ICDM1
2021 Reinforcement Learning for Route Optimization with Robustness Guarantees
abstract
Application of deep learning to NP-hard combinatorial optimization problems is an emerging research trend, and a number of interesting approaches have been published over the last few years. In this work we address robust optimization, which is a more complex variant where a max-min problem is to be solved. We obtain robust solutions by solving the inner minimization problem exactly and apply Reinforcement Learning to learn a heuristic for the outer problem. The minimization term in the inner objective represents an obstacle to existing RL-based approaches, as its value depends on the full solution in a non-linear manner and cannot be evaluated for partial solutions constructed by the agent over the course of each episode. We overcome this obstacle by defining the reward in terms of the one-step advantage over a baseline policy whose role can be played by any fast heuristic for the given problem. The agent is trained to maximize the total advantage, which, as we show, is equivalent to the original objective. We validate our approach by solving min-max versions of standard benchmarks for the Capacitated Vehicle Routing and the Traveling Salesperson Problem, where our agents obtain near-optimal solutions and improve upon the baselines.
Tobias Jacobs, Francesco Alesiani, Gülcin Ermis
IJCAI2
2021 Bilevel Continual Learning
abstract
Continual Learning (CL) studies the problem of learning a sequence of tasks, one at a time, such that the learning of each new task does not lead to the deterioration in performance on the previously seen ones while exploiting previously learned features. This paper presents Bilevel Continual Learning (BiCL), a general framework for continual learning that fuses bilevel optimization and recent advances in meta-learning for deep neural networks. BiCL is able to train both deep discriminative and generative models under the conservative setting of the online continual learning. Experimental results show that BiCL provides competitive performance in terms of accuracy for the current task while reducing the effect of catastrophic forgetting.
Ammar Shaker, Francesco Alesiani, Shujian Yu, Wenzhe Yin
IJCNN2
2020 Measuring the Discrepancy between Conditional Distributions: Methods, Properties and Applications
abstract
We propose a simple yet powerful test statistic to quantify the discrepancy between two conditional distributions. The new statistic avoids the explicit estimation of the underlying distributions in high-dimensional space and it operates on the cone of symmetric positive semidefinite (SPS) matrix using the Bregman matrix divergence. Moreover, it inherits the merits of the correntropy function to explicitly incorporate high-order statistics in the data. We present the properties of our new statistic and illustrate its connections to prior art. We finally show the applications of our new statistic on three different machine learning problems, namely the multi-task learning over graphs, the concept drift detection, and the information-theoretic feature selection, to demonstrate its utility and advantage. Code of our statistic is available at https://bit.ly/BregmanCorrentropy.
Shujian Yu, Ammar Shaker, Francesco Alesiani, José C. Príncipe
IJCAI3
2020 Towards Interpretable Multi-task Learning Using Bilevel Programming
Francesco Alesiani, Shujian Yu, Ammar Shaker, Wenzhe Yin
ECML/PKDD (2)1
2019 Efficient and Scalable Multi-Task Regression on Massive Number of Tasks
abstract
Many real-world large-scale regression problems can be formulated as Multi-task Learning (MTL) problems with a massive number of tasks, as in retail and transportation domains. However, existing MTL methods still fail to offer both the generalization performance and the scalability for such problems. Scaling up MTL methods to problems with a tremendous number of tasks is a big challenge. Here, we propose a novel algorithm, named Convex Clustering Multi-Task regression Learning (CCMTL), which integrates with convex clustering on the k-nearest neighbor graph of the prediction models. Further, CCMTL efficiently solves the underlying convex problem with a newly proposed optimization method. CCMTL is accurate, efficient to train, and empirically scales linearly in the number of tasks. On both synthetic and real-world datasets, the proposed CCMTL outperforms seven state-of-the-art (SoA) multi-task learning methods in terms of prediction accuracy as well as computational efficiency. On a real-world retail dataset with 23,812 tasks, CCMTL requires only around 30 seconds to train on a single thread, while the SoA methods need up to hours or even days.
Francesco Alesiani, Ammar Shaker
AAAI2
2018 On Learning From Inaccurate and Incomplete Traffic Flow Data
abstract
Today, we live in an era where pervasive sensor networks both collect and broadcast rich digital footprints about the human mobility. However, most of this data often comes in an incomplete and/or inaccurate fashion. In this paper, we propose a knowledge discovery framework to handle such issues in the context of automatic incident detection systems fed with traffic flow data. This framework operates in three steps: 1) it clusters sensors with a novel multi-criteria distance metric tailored for this purpose, followed by a heuristic rule that labels the abnormal groups; 2) then, a spatial cross-correlation framework identifies seasonal and individual abnormal readings to perform a more fine-grained filtering; and 3) finally, we propose a novel fundamental diagram that discovers the critical density of a given road section/spot on a data-driven fashion that is resistant to both outliers and noise within the input data. Large-scale experiments were conducted over traffic flow data provided by a major Asian highway operator. The obtained results illustrate well the contributions of this framework: it drastically reduces the noise within the raw data, and it also allows determining reliable definitions of traffic states (congestion/no congestion) on a completely automated way.
Francesco Alesiani, Luís Moreira-Matias, Mahsa Faizrahnemoon
IEEE Trans. Intell. Transp. Syst.1
2016 Remote Testimony: How to Trust an Autonomous Vehicle
abstract
With the advance of Car2X communication models, vehicles become a virtual part of intelligent traffic infrastructures, requiring remote interaction of safety-critical components. The availability of such an integrated system could provide high advantage for the mobility of persons and goods. In this scenario, one of the main concerns with intelligent cars and their network interconnectivity is the susceptibility to cyber-physical attacks. We present a novel concept to establish trust, reliability and safety in intelligent vehicles. Our idea allows a peer in the network to receive a testimony that a vehicle's safety-critical components have not been compromised and as a mater of fact communicate reliable and trustful data. We describe a remote testimony architecture and demonstrate its immunity in the crucial case of an adversary having full control over the vehicle's central units. Our architecture refrains from the heavy and expensive machinery of trusted computing solutions. It makes minimal trust assumptions on the underlying hardware and requires from sensors tamper resistance and provision of a digital signature.
Francesco Alesiani, Sebastian Gajek
VTC Spring1
2003 Differential space-time CDMA with turbo decoding
abstract
Sstarting from recent results on MIMO channels, we design a new transceiver for a code division multiaccess scenario with multiantenna transmitter and receiver. The proposed structure is based on a combination of trellis modulation, turbo codes and differential unitary space-time modulation.
Francesco Alesiani, Alberto Tarable
GLOBECOM1
2001 Performance of adaptive modulation techniques in the UMTS system
abstract
We study the performance of a UNITS downlink, as achieved by a single-user (RAKE) and a multiuser (linear MMSE) receiver with adaptive modulation and multicode in the presence of multipath fading. We advocate an adaptive transmission scheme that varies the number of virtual users and the modulation spectral efficiency in order to optimize the system throughput. Two multipath channel models are considered, exhibiting 3 and 9 paths, respectively. Over the more favorable multipath channel, the multiuser MMSE receiver yields a significant performance enhancement with respect to the RAKE receiver.
Francesco Alesiani, Ezio Biglieri, Giorgio Taricco, Emanuele Viterbo
GLOBECOM1