VLDB 2026 Research / reviewers in the wild / expert
Andrea Borghesi
dblp:135/5392
· DBLP profile ↗
29ranked-venue papers
9as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 12 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Federated transfer learning for anomaly detection in HPC systems: First real-world validation on a tier-0 supercomputerabstract• First real-world FTL for anomaly detection on Tier-0 supercomputer. • Validated on 100 Marconi100 nodes under three learning paradigms. • FTL boosts F1 by up to 0.50 on nodes not in federated training. • Top-N vs Random-N shows diversity can outperform performance-based selection. • Enables privacy-preserving, scalable HPC fault detection without raw data sharing. High-Performance Computing (HPC) systems increasingly require intelligent, scalable anomaly detection to ensure operational reliability. However, conventional centralized approaches often struggle with data privacy constraints, poor generalization across heterogeneous nodes, and limited scalability. This study presents the first real-world application of federated transfer learning (FTL) for anomaly detection in a production-grade Tier-0 supercomputer. By combining federated learning with transfer learning, the proposed framework enables decentralized model training and personalized adaptation to unseen nodes, without accessing raw data. We validate the approach using two large-scale telemetry datasets collected from 100 nodes of the Marconi100 supercomputer, evaluating its effectiveness across supervised, semi-supervised, and unsupervised learning paradigms. Results show that FTL consistently improves anomaly detection performance on nodes that did not participate in federated training, with F1-score gains reaching up to 0.50. These improvements demonstrate the framework’s ability to generalize across non-identically distributed data and maintain detection accuracy under real-world conditions. This work establishes FTL as a scalable, privacy-preserving solution for fault detection in HPC environments. Its practical deployment on production hardware confirms its readiness for real-time monitoring applications in large-scale, heterogeneous computing systems. Emmen Farooq, Michela Milano, Andrea Borghesi |
Expert Syst. Appl. | 3 |
| 2026 | An online algorithm for power consumption prediction of HPC workloadabstractAs modern High-Performance Computing (HPC) systems push the boundaries of computational capabilities, their power consumption becomes a serious threat to environmental and energy sustainability. In such a context, accurate prediction of the jobs’ power consumption is instrumental to develop efficient power management strategies acting at the system level. To this end, in this paper, we present an online prediction algorithm to predict job power consumption in a production HPC system, prior to job execution. Our solution employs machine learning tools, and it is able to predict the minimum, average and maximum power consumption of a job, aggregated per node throughout its execution. Our approach leverages only information which is available at the time of job submission, and it is validated on two datasets extracted from production supercomputers, namely F-DATA from Supercomputer Fugaku and PM100 from Marconi100. Our experimental results show that our prediction algorithm outperforms state-of-the-art techniques, and it can accurately predict job power consumption, by obtaining an error of less than 12% on F-DATA and less than 22% on PM100. Francesco Antici, Andrea Borghesi, Zeynep Kiziltan, Jens Domke, Andrea Bartolini |
Future Gener. Comput. Syst. | 2 |
| 2026 | Federated LSTM autoencoders for time series anomaly detection in production-scale HPC systemsabstractHigh-Performance Computing (HPC) systems are becoming increasingly vulnerable to anomalies as their scale and complexity grow. In this work, we propose a federated learning (FL) framework that integrates Long Short-Term Memory (LSTM) autoencoders for time series anomaly detection, allowing decentralized model training without sharing raw data. Using real telemetry from the Marconi100 Tier-0 supercomputer, our approach improves the average F1-score from 0.388 to 0.867 (+123 %) and the AUC from 0.334 to 0.808 (+142 %). It also cuts the training data requirement by a factor of 15, reducing the collection period from 4.5 months to just 1.25 weeks. These improvements are consistent across unsupervised, semi-supervised, and supervised settings, and significance testing with the Wilcoxon signed-rank test confirms they are statistically robust ( p < 0.01). To our knowledge, this is the first comprehensive evaluation of FL-based LSTM autoencoders for anomaly detection in real HPC environments. Emmen Farooq, Michela Milano, Andrea Borghesi |
Knowl. Based Syst. | 3 |
| 2025 | Machine learning approaches to predict the execution time of the meteorological simulation software COSMOabstractAbstract Predicting the execution time of weather forecast models is a complex task, since these models are usually performed on High Performance Computing systems that require large computing capabilities. Indeed, a reliable prediction can imply several benefits, by allowing for an improved planning of the model execution, a better allocation of available resources, and the identification of possible anomalies. However, to make such predictions is usually hard, since there is a scarcity of datasets that benchmark the existing meteorological simulation models. In this work, we focus on the runtime predictions of the execution of the COSMO (COnsortium for SMall-scale MOdeling) weather forecasting model used at the Hydro-Meteo-Climate Structure of the Regional Agency for the Environment and Energy Prevention Emilia-Romagna. We show how a plethora of Machine Learning approaches can obtain accurate runtime predictions of this complex model, by designing a new well-defined benchmark for this application task. Indeed, our contribution is twofold: 1) the creation of a large public dataset reporting the runtime of COSMO run under a variety of different configurations; 2) a comparative study of ML models, which greatly outperform the current state-of-practice used by the domain experts. This data collection represents an essential initial benchmark for this application field, and a useful resource for analyzing the model performance: better accuracy in runtime predictions could help facility owners to improve job scheduling and resource allocation of the entire system; while for a final user, a posteriori analysis could help to identify anomalous runs. Allegra De Filippo, Emanuele Di Giacomo, Andrea Borghesi |
J. Intell. Inf. Syst. | 3 |
| 2024 | LSTM-Based Unsupervised Anomaly Detection in High-Performance Computing: A Federated Learning ApproachabstractHigh-Performance Computing (HPC) systems are intricate machines that must be run at maximum efficiency to justify their high cost and to minimize environmental impact. Any anomalies that hinder the smooth operation of supercomputing nodes are a significant issue in modern HPC systems. Therefore, the development of automated anomaly detection methods is a crucial area of research within the HPC domain. Machine Learning (ML) models have shown great success in identifying anomalies on individual nodes, especially as contemporary super-computers are outfitted with advanced monitoring systems that provide large datasets for training. However, the potential to combine data from various nodes and to utilize collective ML models remains largely unexplored. Federated Learning (FL) presents a promising approach by enabling individual models to share and learn from one another. Although FL has been employed in areas like healthcare and IoT, its application in HPC is still novel. This study explores how FL can be leveraged to enhance anomaly detection in HPC systems. Using data from a real-world supercomputer, the approach has shown significant promise, boosting the average F1-score from 0.307 to 0.815, and the average AUC from 0.368 to 0.77. Moreover, FL drastically reduces the time required to gather sufficient data for training, allowing faster deployment of detection models. Traditional ML models typically need about 4.5 months of data to perform effectively, but FL can achieve the same with only 1.2 weeks of data, resulting in a 15-fold reduction in data requirements. Emmen Farooq, Andrea Borghesi |
IEEE Big Data | 2 |
| 2024 | Unveiling Computer Chess Evolution: Can Machine Learning Detect Historical Trends?
Andrea Borghesi, Paolo Ciancarini, Angelo Di Iorio, Gianluca Moro |
ICEC | 1 |
| 2024 | Harnessing federated learning for anomaly detection in supercomputer nodes
Emmen Farooq, Michela Milano, Andrea Borghesi |
Future Gener. Comput. Syst. | 3 |
| 2024 | GRAAFE: GRaph Anomaly Anticipation Framework for Exascale HPC systems
Martin Molan, Mohsen Seyedkazemi Ardebili, Junaid Ahmed Khan, Francesco Beneventi, Daniele Cesarini, Andrea Borghesi, Andrea Bartolini |
Future Gener. Comput. Syst. | 6 |
| 2023 | A Federated Learning Approach for Anomaly Detection in High Performance ComputingabstractHigh Performance Computing (HPC) systems are complex machines that need to be operated at their maximum potential to recoup their investment cost and to mitigate their environmental impact. Anomalous conditions hindering the correct usage of the supercomputing nodes are a significant problem. Hence, the development of automated anomaly detection techniques remains a vital area of research. Machine Learning (ML) models demonstrated to be good at detecting anomalies on individual nodes. However, the potential of combining data from multiple computing nodes and associated ML models has not been explored yet. Federated Learning (FL) can address this shortcoming, by allowing individual models to learn from each other. This paper applies FL to improve the performance of anomaly detection models for HPC systems. The approach has been validated on data from an actual supercomputer, obtaining an improvement in the average f-score from 0.31 to 0.84. We also show how FL can significantly shorten the data collection period needed to create a training set. While ML models need, on average, 4.5 months of training data, FL reduces the training set size to 1.2 weeks - a 15x reduction. Emmen Farooq, Andrea Borghesi |
ICTAI | 2 |
| 2023 | MusiComb: a Sample-based Approach to Music Generation Through ConstraintsabstractRecent developments in the field of deep learning have steered research on music generation systems towards a massive use of large end-to-end neural architectures. The capability of these systems to produce convincing outputs has been extensively proven. Nonetheless, they usually come with several drawbacks, such as a low degree of user control, a lack of global structure, and the inherent impossibility of online generation due to high computational costs. Our contribution is two-fold: first, we identify these limitations and show how they have been discussed and partially addressed in the existing literature; then, we propose a novel music generation approach aimed at overcoming such limitations, by properly combining a set of samples under user-defined constraints. We model our task as a job-shop problem, and we show that interesting results can be obtained at very low computational costs. Our framework is genre-independent as it deals with samples metadata rather then individual notes, even though additional genre-specific constraint could be introduced by users to meet their stylistic requirements. Luca Giuliani, Francesco Ballerini, Allegra De Filippo, Andrea Borghesi |
ICTAI | 4 |
| 2023 | Towards Symbiotic Creativity: A Methodological Approach to Compare Human and AI Robotic Dance CreationsabstractArtificial Intelligence (AI) has gradually attracted attention in the field of artistic creation, resulting in a debate on the evaluation of AI artistic outputs. However, there is a lack of common criteria for objective artistic evaluation both of human and AI creations. This is a frequent issue in the field of dance, where different performance metrics focus either on evaluating human or computational skills separately. This work proposes a methodological approach for the artistic evaluation of both AI and human artistic creations in the field of robotic dance. First, we define a series of common initial constraints to create robotic dance choreographies in a balanced initial setting, in collaboration with a group of human dancers and choreographer. Then, we compare both creation processes through a human audience evaluation. Finally, we investigate which choreography aspects (e.g., the music genre) have the largest impact on the evaluation, and we provide useful guidelines and future research directions for the analysis of interconnections between AI and human dance creation. Allegra De Filippo, Luca Giuliani, Eleonora Mancini, Andrea Borghesi, Paola Mello, Michela Milano |
IJCAI | 4 |
| 2023 | RUAD: Unsupervised anomaly detection in HPC systems
Martin Molan, Andrea Borghesi, Daniele Cesarini, Luca Benini, Andrea Bartolini |
Future Gener. Comput. Syst. | 2 |
| 2023 | ExaMon-X: A Predictive Maintenance Framework for Automatic Monitoring in Industrial IoT SystemsabstractIn recent years, the Industrial Internet of Things (IIoT) has led to significant steps forward in many industries, thanks to the exploitation of several technologies, ranging from Big Data processing to artificial intelligence (AI). Among the various IIoT scenarios, large-scale data centers can reap significant benefits from adopting Big Data analytics and AI-boosted approaches since these technologies can allow effective predictive maintenance. However, most of the off-the-shelf currently available solutions are not ideally suited to the high-performance computing (HPC) context, e.g., they do not sufficiently take into account the very heterogeneous data sources and the privacy issues that hinder the adoption of the cloud solution, or they do not fully exploit the computing capabilities available in loco in a supercomputing facility. In this article, we tackle this issue, and we propose an IIoT holistic and vertical framework for predictive maintenance in supercomputers. The framework is based on a big lightweight data monitoring infrastructure, specialized databases suited for heterogeneous data, and a set of high-level AI-based functionalities tailored to HPC actors’ specific needs. We present the deployment and assess the usage of this framework in several in-production HPC systems. Andrea Borghesi, Alessio Burrello, Andrea Bartolini |
IEEE Internet Things J. | 1 |
| 2022 | Semi-supervised anomaly detection on a Tier-0 HPC systemabstractAutomated and data-driven methodologies are being introduced to assist system administrators in managing increasingly complex modern HPC systems. Anomaly detection (AD) is an integral part of improving the overall availability as it eases the system administrators' burden and reduces the time between an anomaly and its resolution. This work improves upon the current state-of-the-art (SoA) AD model by considering temporal dependencies in the data and including long-short term memory cells in the architecture of the AD model. The proposed model is evaluated on a complete ten-month history of a Tier-0 system (Marconi100 from CINECA consisting of 985 nodes). The proposed model achieves an area under the curve (AUC) of 0.758, improving upon the state-of-the-art approach that achieves an AUC of 0.747. Martin Molan, Andrea Borghesi, Luca Benini, Andrea Bartolini |
CF | 2 |
| 2022 | Analysing Supercomputer Nodes Behaviour with the Latent Representation of Deep Learning Models
Martin Molan, Andrea Borghesi, Luca Benini, Andrea Bartolini |
Euro-Par | 2 |
| 2022 | HADA: An automated tool for hardware dimensioning of AI applications
Allegra De Filippo, Andrea Borghesi, Andrea Boscarino, Michela Milano |
Knowl. Based Syst. | 2 |
| 2022 | Anomaly Detection and Anticipation in High Performance Computing SystemsabstractIn their quest toward Exascale, High Performance Computing (HPC) systems are rapidly becoming larger and more complex, together with the issues concerning their maintenance. Luckily, many current HPC systems are endowed with data monitoring infrastructures that characterize the system state, and whose data can be used to train Deep Learning (DL) anomaly detection models, a very popular research area. However, the lack of labels describing the state of the system is a wide-spread issue, as annotating data is a costly task, generally falling on human system administrators and thus does not scale toward exascale. In this article we investigate the possibility to extract labels from a service monitoring tool (Nagios) currently used by HPC system administrators to flag the nodes which undergo maintenance operations. This allows to automatically annotate data collected by a fine-grained monitoring infrastructure; this labelled data is then used to train and validate a DL model for anomaly detection. We conduct the experimental evaluation on a tier-0 production supercomputer hosted at CINECA, Bologna, Italy. The results reveal that the DL model can accurately detect the real failures, and, moreover, it canpredictthe insurgency of anomalies, by systematically anticipating the actual labels (i.e., the moment when system administrators realize when an anomalous event happened); the average advance time computed on historical traces is around 45 minutes. The proposed technology can be easily scaled toward exascale systems to easy their maintenance. Andrea Borghesi, Martin Molan, Michela Milano, Andrea Bartolini |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | IoTwins: Design and Implementation of a Platform for the Management of Digital Twins in Industrial ScenariosabstractWith the increase of the volume of data produced by IoT devices, there is a growing demand of applications capable of elaborating data anywhere along the IoT-to-Cloud path (Edge/Fog). In industrial environments, strict real-time constraints require computation to run as close to the data origin as possible (e.g., IoT Gateway or Edge nodes), whilst batch-wise tasks such as Big Data analytics and Machine Learning model training are advised to run on the Cloud, where computing resources are abundant. The H2020 IoTwins project leverages the digital twin concept to implement virtual representation of physical assets (e.g., machine parts, machines, production/control processes) and deliver a software platform that will help enterprises, and in particular SMEs, to build highly innovative, AI-based services that exploit the potential of IoT/Edge/Cloud computing paradigms. In this paper, we discuss the design principles of the IoTwins reference architecture, delving into technical details of its components and offered functionalities, and propose an exemplary software implementation. Andrea Borghesi, Giuseppe Di Modica, Paolo Bellavista, Varun Gowtham, Alexander Willner, Daniel Nehls, Florian Kintzler, Stephan Cejka, Simone Rossi Tisbeni, Alessandro Costantini, Matteo Galletti, Marica Antonacci, Jean Christian Ahouangonou |
CCGRID | 1 |
| 2021 | BS-Net: Learning COVID-19 pneumonia severity on a large chest X-ray dataset
Alberto Signoroni, Mattia Savardi, Sergio Benini, Nicola Adami, Riccardo Leonardi, Paolo Gibellini, Filippo Vaccher, Marco Ravanelli, Andrea Borghesi, Roberto Maroldi, Davide Farina |
Medical Image Anal. | 9 |
| 2020 | Combining learning and optimization for transprecision computingabstractThe growing demands of the worldwide IT infrastructure stress the need for reduced power consumption, which is addressed in so-called transprecision computing by improving energy efficiency at the expense of precision. For example, reducing the number of bits for some floating-point operations leads to higher efficiency, but also to a non-linear decrease of the computation accuracy. Depending on the application, small errors can be tolerated, thus allowing to fine-tune the precision of the computation. Finding the optimal precision for all variables in respect of an error bound is a complex task, which is tackled in the literature via heuristics. In this paper, we report on a first attempt to address the problem by combining a Mathematical Programming (MP) model and a Machine Learning (ML) model, following the Empirical Model Learning methodology. The ML model learns the relation between variables precision and the output error; this information is then embedded in the MP focused on minimizing the number of bits. An additional refinement phase is then added to improve the quality of the solution. The experimental results demonstrate an average speedup of 6.5% and a 3% increase in solution quality compared to the state-of-the-art. In addition, experiments on a hardware platform capable of mixed-precision arithmetic (PULPissimo) show the benefits of the proposed approach, with energy savings of around 40% compared to fixed-precision. Andrea Borghesi, Giuseppe Tagliavini, Michele Lombardi 0001, Luca Benini, Michela Milano |
CF | 1 |
| 2020 | A machine learning approach to online fault classification in HPC systems
Alessio Netti, Zeynep Kiziltan, Özalp Babaoglu, Alina Sîrbu, Andrea Bartolini, Andrea Borghesi |
Future Gener. Comput. Syst. | 6 |
| 2020 | Countdown Slack: A Run-Time Library to Reduce Energy Footprint in Large-Scale MPI ApplicationsabstractThe power consumption of supercomputers is a major challenge for system owners, users, and society. It limits the capacity of system installations, it requires large cooling infrastructures, and it is the cause of a large carbon footprint. Reducing power during application execution without changing the application source code or increasing time-to-completion is highly desirable in real-life high-performance computing scenarios. The power management run-time frameworks proposed in the last decade are based on the assumption that the duration of communication and application phases in an MPI application can be predicted and used at run-time to trade-off communication slack with power consumption. In this article, we first show that this assumption is too general and leads to mispredictions, slowing down applications, thereby jeopardizing the claimed benefits. We then propose a new approach based on (i) the separation of communication phases and slack during MPI calls and (ii) a timeout algorithm to cope with the hardware power management latency, which jointly makes it possible to achieve performance-neutral power saving in MPI applications without requiring labor-intensive and risky application source code modifications. We validate our approach in a tier-1 production environment with widely adopted scientific applications. Our approach has a time-to-completion overhead lower than 1 percent, while it successfully exploits slack in communication phases to achieve an average energy saving of 10 percent. If we focus on a large-scale application runs, the proposed approach achieves 22 percent energy saving with an overhead of only 0.4 percent. With respect to state-of-the-art approaches, COUNTDOWN Slack is the only that always leads to an energy saving with negligible overhead (<; 3 percent). Daniele Cesarini, Andrea Bartolini, Andrea Borghesi, Carlo Cavazzoni, Mathieu Luisier, Luca Benini |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2019 | Anomaly Detection Using Autoencoders in High Performance Computing SystemsabstractAnomaly detection in supercomputers is a very difficult problem due to the big scale of the systems and the high number of components. The current state of the art for automated anomaly detection employs Machine Learning methods or statistical regression models in a supervised fashion, meaning that the detection tool is trained to distinguish among a fixed set of behaviour classes (healthy and unhealthy states).We propose a novel approach for anomaly detection in HighPerformance Computing systems based on a Machine (Deep) Learning technique, namely a type of neural network called autoencoder. The key idea is to train a set of autoencoders to learn the normal (healthy) behaviour of the supercomputer nodes and, after training, use them to identify abnormal conditions. This is different from previous approaches which where based on learning the abnormal condition, for which there are much smaller datasets (since it is very hard to identify them to begin with).We test our approach on a real supercomputer equipped with a fine-grained, scalable monitoring infrastructure that can provide large amount of data to characterize the system behaviour. The results are extremely promising: after the training phase to learn the normal system behaviour, our method is capable of detecting anomalies that have never been seen before with a very good accuracy (values ranging between 88% and 96%). Andrea Borghesi, Andrea Bartolini, Michele Lombardi 0001, Michela Milano, Luca Benini |
AAAI | 1 |
| 2019 | Online Fault Classification in HPC Systems Through Machine Learning
Alessio Netti, Zeynep Kiziltan, Özalp Babaoglu, Alina Sîrbu, Andrea Bartolini, Andrea Borghesi |
Euro-Par | 6 |
| 2019 | A semisupervised autoencoder-based approach for anomaly detection in high performance computing systems
Andrea Borghesi, Andrea Bartolini, Michele Lombardi 0001, Michela Milano, Luca Benini |
Eng. Appl. Artif. Intell. | 1 |
| 2018 | The D.A.V.I.D.E. big-data-powered fine-grain power and performance monitoring supportabstractOn the race toward exascale supercomputing systems are facing important challenges which limit the efficiency of the system. Among all, power and energy consumption fueled by the end of Dennard's scaling start to show their impact on limiting supercomputers peak performance and cost effectiveness. Andrea Bartolini, Andrea Borghesi, Antonio Libri, Francesco Beneventi, Daniele Gregori, Simone Tinti, Cosimo Gianfreda, Piero Altoe |
CF | 2 |
| 2015 | Power Capping in High Performance Computing Systems
Andrea Borghesi, Francesca Collina, Michele Lombardi 0001, Michela Milano, Luca Benini |
CP | 1 |
| 2014 | Proactive Workload Dispatching on the EURORA Supercomputer
Andrea Bartolini, Andrea Borghesi, Thomas Bridi, Michele Lombardi 0001, Michela Milano |
CP | 2 |
| 2013 | Simulation Of Incentive Mechanisms For Renewable Energy PoliciesabstractDesigning sustainable energy policies has a strong impact on economy, society and environment. Beside a planning activity, policy makers are called to design a number of implementation instruments to enforce their plans. They encompass subsidies, fiscal incentives, feed in tariffs to name a few. Understanding the impact of these instruments on the energy market is essential to select the most efficient one. We propose in this paper a multi-agent simulator that mimics the adoption of photovoltaic as a consequence of a number of implementation instruments. The simulator mainly considers economic evaluations in the agent decision-making procedure, but we are aware also social aspects play an important role and they are subject of current research. Andrea Borghesi, Michela Milano, Marco Gavanelli, Tony Woods |
ECMS | 1 |