Gustavo Espinosa

dblp:340/0130 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0003-0401-7355ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware reliability and fault tolerance · 35% High-performance computing · 35% Distributed systems · 20%

Topics — the 5 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › supercomputing
exascale computing
0.912025
Fine-grained Automated Failure Management for Extreme-Scale GPU Accelerated Systems · SC 2025
Distributed systems
fault tolerance
0.912025
Fine-grained Automated Failure Management for Extreme-Scale GPU Accelerated Systems · SC 2025
Hardware reliability and fault tolerance
fault injection
0.712023
HPC Hardware Design Reliability Benchmarking With HDFIT · IEEE Trans. Parallel Distributed Syst. 2023
GPUs and heterogeneous computing › GPU computing
GPU-accelerated systems
0.312025
Fine-grained Automated Failure Management for Extreme-Scale GPU Accelerated Systems · SC 2025
Hardware accelerators and domain-specific architectures › sparse matrix multiplication accelerator
matrix multiplication accelerator
0.212023
HPC Hardware Design Reliability Benchmarking With HDFIT · IEEE Trans. Parallel Distributed Syst. 2023

Methods — techniques the papers use, named apart from their topics

failure analysis · 0.9reliability benchmarking · 0.7fault injection · 0.7
YearPublicationVenuePosition
2025 Fine-grained Automated Failure Management for Extreme-Scale GPU Accelerated Systems
abstract
As high-performance computing (HPC) systems scale in size, system wide hardware failure rates increase. Historical data from previous large-scale HPC installations illustrate this trend, with the mean time between failures (MTBF) decreasing steadily over the past decade. Recent studies from artificial intelligence and machine-learning (AI/ML) training extrapolate MTBF declining even further for future GPU accelerated systems. As MTBF decreases, mean time to repair (MTTR) becomes more pronounced, highlighting the need for efficient recovery strategies.
Yonatan Levitt, Richard Barella, Sam Zeltner, Thomas Musta, Lance C. Cheney, Gustavo Espinosa, Olivier Franza, Balazs Gerofi
SC6
2023 Mixed precision support in HPC applications: What about reliability?
Alessio Netti, Patrik Omland, Michael Paulitsch, Jorge Parra, Gustavo Espinosa, Udit Kumar Agarwal, Abraham Chan, Karthik Pattabiraman
J. Parallel Distributed Comput.6
2023 HPC Hardware Design Reliability Benchmarking With HDFIT
abstract
Chips pack ever more, ever smaller transistors. Fault rates increase in turn and become more concerning, particularly at the scale ofHigh-Performance Computing(HPC) systems: on one hand, hardware fault protection is costly - more than 10% silicon area for floating-point units; on the other, HPC users expect correct application output after the anticipated time of computation, but workloads are seldom bit-reproducible and tolerances in output are allowed for. Benign hardware faults causing errors within these tolerances are therefore acceptable: however, with abstract reliability targets such as ’undetected failures per time,’ current HPC system design does not allow for pursuing trade-offs between reliability and performance with respect to faults. To address the above, we propose a user-centric reliability benchmark to specify HPC system reliability targets, allowing for better performance optimizations in hardware design, while meeting HPC user expectations. Our open-sourceHardware Design Fault Injection Toolkit(HDFIT) enables - for the first time - end-to-end hardware design reliability experiments: from netlist-level fault injection to application output error. In a proof of concept we present an HPCgeneral matrix multiply(GEMM) reliability study, targeting a series of popular applications, and using HDFIT to benchmark an open-source GEMM accelerator.
Patrik Omland, Alessio Netti, Andrea Baldovin, Michael Paulitsch, Gustavo Espinosa, Jorge Parra, Gereon Hinz, Alois C. Knoll
IEEE Trans. Parallel Distributed Syst.6