Don E. Maxwell

dblp:63/5068 · also Don Maxwell · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
1since 2021 · last 2021
0000-0002-3794-5687ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 1 since 2021Security and privacy · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
GPUs and heterogeneous computing · 31% High-performance computing · 29% Hardware reliability and fault tolerance · 29%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 15 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU reliability
0.932020
GPU lifetimes on titan supercomputer: survival analysis and reliability · SC 2020
Reliability lessons learned from GPU experience with the Titan supercomputer at Oak Ridge leadership computing facility · SC 2015
Understanding GPU errors on large-scale HPC systems and the implications for system design and operation · HPCA 2015
Hardware reliability and fault tolerance
failure analysis
0.412020
GPU lifetimes on titan supercomputer: survival analysis and reliability · SC 2020
Hardware reliability and fault tolerance › reliability analysis
lifetime analysis
0.412020
GPU lifetimes on titan supercomputer: survival analysis and reliability · SC 2020
High-performance computing
supercomputing
0.322020
Reliability lessons learned from GPU experience with the Titan supercomputer at Oak Ridge leadership computing facility · SC 2015
GPU lifetimes on titan supercomputer: survival analysis and reliability · SC 2020
GPUs and heterogeneous computing
GPU scheduling
0.312018
GPU age-aware scheduling to improve the reliability of leadership jobs on Titan · SC 2018
Embedded and real-time systems › real-time scheduling
reliability-aware scheduling
0.312018
GPU age-aware scheduling to improve the reliability of leadership jobs on Titan · SC 2018
High-performance computing › supercomputing
supercomputer deployment
0.312018
The design, deployment, and evaluation of the CORAL pre-exascale systems · SC 2018
Hardware reliability and fault tolerance
soft errors
0.322015
Reliability lessons learned from GPU experience with the Titan supercomputer at Oak Ridge leadership computing facility · SC 2015
Understanding GPU errors on large-scale HPC systems and the implications for system design and operation · HPCA 2015
GPUs and heterogeneous computing › multi-GPU computing
GPU cluster
0.212015
Reliability lessons learned from GPU experience with the Titan supercomputer at Oak Ridge leadership computing facility · SC 2015
Hardware reliability and fault tolerance
system reliability
0.112020
GPU lifetimes on titan supercomputer: survival analysis and reliability · SC 2020
Performance modeling and evaluation
benchmarking
0.112018
The design, deployment, and evaluation of the CORAL pre-exascale systems · SC 2018
High-performance computing
performance optimization at scale
0.112008
New algorithm to enable 400+ TFlop/s sustained performance in simulations of disorder effects in high-Tc superconductors · SC 2008
Emerging computing paradigms › quantum computing › quantum simulation
quantum monte carlo simulation
0.112008
New algorithm to enable 400+ TFlop/s sustained performance in simulations of disorder effects in high-Tc superconductors · SC 2008
High-performance computing
scientific computing systems
0.112008
New algorithm to enable 400+ TFlop/s sustained performance in simulations of disorder effects in high-Tc superconductors · SC 2008
Computational science and engineering
computational physics
0.012008
New algorithm to enable 400+ TFlop/s sustained performance in simulations of disorder effects in high-Tc superconductors · SC 2008

Methods — techniques the papers use, named apart from their topics

time between failures analysis · 0.4statistical survival analysis · 0.4reliability analysis · 0.2neutron beam experiments · 0.2field data analysis · 0.2error characterization · 0.2mixed single-/double precision · 0.2delayed monte carlo updates · 0.2
YearPublicationVenuePosition
2021 Understanding failures through the lifetime of a top-level supercomputer
Elvis Rojas, Esteban Meneses, Terry R. Jones, Don E. Maxwell
J. Parallel Distributed Comput.4
2020 Towards a Model to Estimate the Reliability of Large-Scale Hybrid Supercomputers
Elvis Rojas, Esteban Meneses, Terry R. Jones, Don E. Maxwell
Euro-Par4
2020 GPU lifetimes on titan supercomputer: survival analysis and reliability
abstract
The Cray XK7 Titan was the top supercomputer system in the world for a long time and remained critically important throughout its nearly seven year life. It was an interesting machine from a reliability viewpoint as most of its power came from 18,688 GPUs whose operation was forced to execute three rework cycles, two on the GPU mechanical assembly and one on the GPU circuitboards. We write about the last rework cycle and a reliability analysis of over 100,000 years of GPU lifetimes during Titan’s 6-year-long productive period. Using time between failures analysis and statistical survival analysis techniques, we find that GPU reliability is dependent on heat dissipation to an extent that strongly correlates with detailed nuances of the cooling architecture and job scheduling. We describe the history, data collection, cleaning, and analysis and give recommendations for future supercomputing systems. We make the data and our analysis codes publicly available.
George Ostrouchov, Don E. Maxwell, Rizwan A. Ashraf, Christian Engelmann, Mallikarjun Shankar, James H. Rogers
SC2
2019 Analyzing a Five-Year Failure Record of a Leadership-Class Supercomputer
abstract
Extreme-scale computing systems are required to solve some of the grand challenges in science and technology. From astrophysics to molecular biology, supercomputers are an essential tool to accelerate scientific discovery. However, large computing systems are prone to failures due to their complexity. It is crucial to develop an understanding of how these systems fail to design reliable supercomputing platforms for the future. This paper examines a five-year failure and workload record of a leadership-class supercomputer. To the best of our knowledge, five years represents the vast majority of the lifespan of a supercomputer. This is the first time such analysis is performed on a top 10 modern supercomputer. We performed a failure categorization and found out that: i) most errors are GPUrelated, with roughly 37% of them being double-bit errors on the cards; ii) failures are not evenly spread across the physical machine, with room temperature presumably playing a major role; and iii) software errors of the system bring down several nodes concurrently. Our failure rate analysis unveils that: i) the system consistently degrades, being at least twice as reliable at the beginning, compared to the end of the period; ii) Weibull distribution closely fits the mean-time-between-failure data; and iii) hardware and software errors show a markedly different pattern. Finally, we correlated failure and workload records to reveal that: i) failure and workload records are weakly correlated, except for certain types of failures when segmented by the hours of the day; ii) several categories of failures make jobs crash within the first minutes of execution; and iii) a significant fraction of failed jobs exhaust the requested time with a disregard of when the failure occurred during execution. Index Terms-Fault tolerance, resilience, failure analysis, high performance computing.
Elvis Rojas, Esteban Meneses, Terry R. Jones, Don E. Maxwell
SBAC-PAD4
2019 Are we witnessing the spectre of an HPC meltdown?
abstract
Summary We measure and analyze the performance observed when running applications and benchmarks before and after the Meltdown and Spectre fixes have been applied to the Cray supercomputers and supporting systems at the Oak Ridge Leadership Computing Facility (OLCF). Of particular interest is the effect of these fixes on applications selected from the OLCF portfolio when running at scale. This comprehensive study presents results from experiments run on Titan, Eos, Cumulus, and Percival supercomputers at the OLCF. The results from this study are useful for HPC users running on Cray supercomputers and serve to better understand the impact that these two vulnerabilities have on diverse HPC workloads at scale.
Verónica G. Vergara Larrea, Michael J. Brim, Wayne Joubert, Swen Böhm, Matthew B. Baker, Oscar R. Hernandez, Sarp Oral, James Simmons, Don E. Maxwell
Concurr. Comput. Pract. Exp.9
2018 The design, deployment, and evaluation of the CORAL pre-exascale systems
Sudharshan S. Vazhkudai, Bronis R. de Supinski, Arthur S. Bland, Al Geist, James C. Sexton, James A. Kahle, Christopher Zimmer 0001, Scott Atchley, Sarp Oral, Don E. Maxwell, Verónica G. Vergara Larrea, Adam Bertsch, Robin Goldstone, Wayne Joubert, Christopher M. Chambreau, David Appelhans, Robert Blackmore, Ben Casses, George Chochia, Gene Davison, Matthew Ezell, Thomas Gooding, Elsa Gonsiorowski, Leopold Grinberg, Bill Hanson, Bill Hartner, Ian Karlin, Matthew L. Leininger, Dustin Leverman, Chris Marroquin, Adam Moody, Martin Ohmacht, Ramesh Pankajakshan, Fernando Pizzano, James H. Rogers, Bryan S. Rosenburg, Drew Schmidt, Mallikarjun Shankar, Feiyi Wang, Py Watson, Bob Walkup, Lance D. Weems, Junqi Yin
SC10
2018 GPU age-aware scheduling to improve the reliability of leadership jobs on Titan
Christopher Zimmer 0001, Don E. Maxwell, Stephen Taylor McNally, Scott Atchley, Sudharshan S. Vazhkudai
SC2
2015 Understanding and Exploiting Spatial Properties of System Failures on Extreme-Scale HPC Systems
abstract
As we approach exascale, the scientific simulations are expected to experience more interruptions due to increased system failures. Designing better HPC resilience techniques requires understanding the key characteristics of system failures on these systems. While temporal properties of system failures on HPC systems have been well-investigated, there is limited understanding about the spatial characteristics of system failures and its impact on the resilience mechanisms. Therefore, we examine the spatial characteristics and behavior of system failures. We investigate the interaction between spatial and temporal characteristics of failures and its implications for system operations and resilience mechanisms on large-scale HPC systems. We show that system failures have "spatial locality" at different granularity in the system, study impact of different failure-types, and investigate the correlation among different failure-types. Finally, we propose a novel scheme that exploits the spatial locality in failures to improve application and system performance. Our evaluation shows that the proposed scheme significantly improves the system performance in a dynamic and production-level HPC system.
Saurabh Gupta 0002, Devesh Tiwari, Christopher Jantzi, James H. Rogers, Don E. Maxwell
DSN5
2015 Understanding GPU errors on large-scale HPC systems and the implications for system design and operation
abstract
Increase in graphics hardware performance and improvements in programmability has enabled GPUs to evolve from a graphics-specific accelerator to a general-purpose computing device. Titan, the world's second fastest supercomputer for open science in 2014, consists of more dum 18,000 GPUs that scientists from various domains such as astrophysics, fusion, climate, and combustion use routinely to run large-scale simulations. Unfortunately, while the performance efficiency of GPUs is well understood, their resilience characteristics in a large-scale computing system have not been fully evaluated. We present a detailed study to provide a thorough understanding of GPU errors on a large-scale GPU-enabled system. Our data was collected from the Titan supercomputer at the Oak Ridge Leadership Computing Facility and a GPU cluster at the Los Alamos National Laboratory. We also present results from our extensive neutron-beam tests, conducted at Los Alamos Neutron Science Center (LANSCE) and at ISIS (Rutherford Appleron Laboratories, UK), to measure the resilience of different generations of GPUs. We present several findings from our field data and neutron-beam experiments, and discuss the implications of our results for future GPU architects, current and future HPC computing facilities, and researchers focusing on GPU resilience.
Devesh Tiwari, Saurabh Gupta 0002, James H. Rogers, Don E. Maxwell, Paolo Rech, Sudharshan S. Vazhkudai, Daniel Oliveira 0002, Dave Londo, Nathan DeBardeleben, Philippe Olivier Alexandre Navaux, Luigi Carro, Arthur S. Bland
HPCA4
2015 Reliability lessons learned from GPU experience with the Titan supercomputer at Oak Ridge leadership computing facility
abstract
The high computational capability of graphics processing units (GPUs) is enabling and driving the scientific discovery process at large-scale. The world's second fastest supercomputer for open science, Titan, has more than 18,000 GPUs that computational scientists use to perform scientific simulations and data analysis. Understanding of GPU reliability characteristics, however, is still in its nascent stage since GPUs have only recently been deployed at large-scale. This paper presents a detailed study of GPU errors and their impact on system operations and applications, describing experiences with the 18,688 GPUs on the Titan supercomputer as well as lessons learned in the process of efficient operation of GPUs at scale. These experiences are helpful to HPC sites which already have large-scale GPU clusters or plan to deploy GPUs in the future.
Devesh Tiwari, Saurabh Gupta 0002, George Gallarno, Jim Rogers, Don E. Maxwell
SC5
2008 New algorithm to enable 400+ TFlop/s sustained performance in simulations of disorder effects in high-Tc superconductors
abstract
Staggering computational and algorithmic advances in recent years now make possible systematic Quantum Monte Carlo (QMC) simulations of high temperature (high-Tc) superconductivity in a microscopic model, the two dimensional (2D) Hubbard model, with parameters relevant to the cuprate materials. Here we report the algorithmic and computational advances that enable us to study the effect of disorder and nano-scale inhomogeneities on the pair-formation and the superconducting transition temperature necessary to understand real materials. The simulation code is written with a generic and extensible approach and is tuned to perform well at scale. Significant algorithmic improvements have been made to make effective use of current supercomputing architectures. By implementing delayed Monte Carlo updates and a mixed single-/double precision mode, we are able to dramatically increase the efficiency of the code. On the Cray XT4 systems of the Oak Ridge National Laboratory (ORNL), for example, we currently run production jobs on 31 thousand processors and thereby routinely achieve a sustained performance that exceeds 200 TFlop/s. On a system with 49 thousand processors we achieved a sustained performance of 409 TFlop/s. We present a study of how random disorder in the effective Coulomb interaction strength affects the superconducting transition temperature in the Hubbard model.
Gonzalo Alvarez 0001, Michael S. Summers, Don E. Maxwell, Markus Eisenbach 0002, Jeremy S. Meredith, Jeffrey M. Larkin, John M. Levesque, Thomas A. Maier, Paul R. C. Kent, Eduardo F. D'Azevedo, Thomas C. Schulthess
SC3