EDBT 2026 Demo / reviewers in the wild / expert
Alessio Netti
dblp:211/5921
· DBLP profile ↗
9ranked-venue papers
8as first author
5since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 8 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Mixed precision support in HPC applications: What about reliability?
Alessio Netti, Patrik Omland, Michael Paulitsch, Jorge Parra, Gustavo Espinosa, Udit Kumar Agarwal, Abraham Chan, Karthik Pattabiraman |
J. Parallel Distributed Comput. | 1 |
| 2023 | HPC Hardware Design Reliability Benchmarking With HDFITabstractChips pack ever more, ever smaller transistors. Fault rates increase in turn and become more concerning, particularly at the scale ofHigh-Performance Computing(HPC) systems: on one hand, hardware fault protection is costly - more than 10% silicon area for floating-point units; on the other, HPC users expect correct application output after the anticipated time of computation, but workloads are seldom bit-reproducible and tolerances in output are allowed for. Benign hardware faults causing errors within these tolerances are therefore acceptable: however, with abstract reliability targets such as ’undetected failures per time,’ current HPC system design does not allow for pursuing trade-offs between reliability and performance with respect to faults. To address the above, we propose a user-centric reliability benchmark to specify HPC system reliability targets, allowing for better performance optimizations in hardware design, while meeting HPC user expectations. Our open-sourceHardware Design Fault Injection Toolkit(HDFIT) enables - for the first time - end-to-end hardware design reliability experiments: from netlist-level fault injection to application output error. In a proof of concept we present an HPCgeneral matrix multiply(GEMM) reliability study, targeting a series of popular applications, and using HDFIT to benchmark an open-source GEMM accelerator. Patrik Omland, Alessio Netti, Andrea Baldovin, Michael Paulitsch, Gustavo Espinosa, Jorge Parra, Gereon Hinz, Alois C. Knoll |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | Operational Data Analytics in practice: Experiences from design to deployment in production HPC environments
Alessio Netti, Michael Ott 0001, Carla Guillén, Daniele Tafani, Martin Schulz 0001 |
Parallel Comput. | 1 |
| 2021 | A Conceptual Framework for HPC Operational Data AnalyticsabstractThis paper provides a broad framework for understanding trends in Operational Data Analytics (ODA) for High-Performance Computing (HPC) facilities. The goal of ODA is to allow for the continuous monitoring, archiving, and analysis of near real-time performance data, providing immediately actionable information for multiple operational uses. In this work, we combine two models to provide a comprehensive HPC ODA framework: one is an evolutionary model of analytics capabilities that consists of four types, which are descriptive, diagnostic, predictive and prescriptive, while the other is a four-pillar model for energy-efficient HPC operations that covers facility, system hardware, system software, and applications. This new framework is then overlaid with a description of current development and production deployments of ODA within leading-edge HPC facilities. Finally, we perform a comprehensive survey of ODA works and classify them according to our framework, in order to demonstrate its effectiveness. Alessio Netti, Woong Shin, Michael Ott 0001, Torsten Wilde, Natalie J. Bates |
CLUSTER | 1 |
| 2021 | Correlation-wise Smoothing: Lightweight Knowledge Extraction for HPC Monitoring DataabstractModern High-Performance Computing (HPC) and data center operators rely more and more on data analytics techniques to improve the efficiency and reliability of their operations. They employ models that ingest time-series monitoring sensor data and transform it into actionable knowledge for system tuning: a process known as Operational Data Analytics (ODA). However, monitoring data has a high dimensionality, is hardware-dependent and difficult to interpret. This, coupled with the strict requirements of ODA, makes most traditional data mining methods impractical and in turn renders this type of data cumbersome to process. Most current ODA solutions use ad-hoc processing methods that are not generic, are sensible to the sensors' features and are not fit for visualization. In this paper we propose a novel method, called Correlation-wise Smoothing (CS), to extract descriptive signatures from time-series monitoring data in a generic and lightweight way. Our CS method exploits correlations between data dimensions to form groups and produces image-like signatures that can be easily manipulated, visualized and compared. We evaluate the CS method on HPC-ODA, a collection of datasets that we release with this work, and show that it leads to the same performance as most state-of-the-art methods while producing signatures that are up to ten times smaller and up to ten times faster, while gaining visualizability, portability across systems and clear scaling properties. Alessio Netti, Daniele Tafani, Michael Ott 0001, Martin Schulz 0001 |
IPDPS | 1 |
| 2020 | DCDB Wintermute: Enabling Online and Holistic Operational Data Analytics on HPC SystemsabstractAs we approach the exascale era, the size and complexity of HPC systems continues to increase, raising concerns about their manageability and sustainability. For this reason, more and more HPC centers are experimenting with fine-grained monitoring coupled with Operational Data Analytics (ODA) to optimize efficiency and effectiveness of system operations. However, while monitoring is a common reality in HPC, there is no well-stated and comprehensive list of requirements, nor matching frameworks, to support holistic and online ODA. This leads to insular ad-hoc solutions, each addressing only specific aspects of the problem. Alessio Netti, Micha Müller, Carla Guillén, Michael Ott 0001, Daniele Tafani, Gence Ozer, Martin Schulz 0001 |
HPDC | 1 |
| 2020 | A machine learning approach to online fault classification in HPC systems
Alessio Netti, Zeynep Kiziltan, Özalp Babaoglu, Alina Sîrbu, Andrea Bartolini, Andrea Borghesi |
Future Gener. Comput. Syst. | 1 |
| 2019 | Online Fault Classification in HPC Systems Through Machine Learning
Alessio Netti, Zeynep Kiziltan, Özalp Babaoglu, Alina Sîrbu, Andrea Bartolini, Andrea Borghesi |
Euro-Par | 1 |
| 2019 | From facility to application sensor data: modular, continuous and holistic monitoring with DCDBabstractToday's HPC installations are highly-complex systems, and their complexity will only increase as we move to exascale and beyond. At each layer, from facilities to systems, from runtimes to applications, a wide range of tuning decisions must be made in order to achieve efficient operation. This, however, requires systematic and continuous monitoring of system and user data. While many insular solutions exist, a system for holistic and facility-wide monitoring is still lacking in the current HPC ecosystem. Alessio Netti, Micha Müller, Axel Auweter, Carla Guillén, Michael Ott 0001, Daniele Tafani, Martin Schulz 0001 |
SC | 1 |