Arthur S. Bland

dblp:42/122 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
High-performance computing · 52% Storage systems · 17% GPUs and heterogeneous computing · 15%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › supercomputing
supercomputer deployment
0.312018
The design, deployment, and evaluation of the CORAL pre-exascale systems · SC 2018
GPUs and heterogeneous computing
GPU reliability
0.212015
Understanding GPU errors on large-scale HPC systems and the implications for system design and operation · HPCA 2015
Storage systems › file systems › distributed file system
parallel file system
0.212014
Best Practices and Lessons Learned from Deploying and Operating Large-Scale Data-Centric Parallel File Systems · SC 2014
High-performance computing
supercomputing
0.212014
Best Practices and Lessons Learned from Deploying and Operating Large-Scale Data-Centric Parallel File Systems · SC 2014
Performance modeling and evaluation
benchmarking
0.122018
The design, deployment, and evaluation of the CORAL pre-exascale systems · SC 2018
Early evaluation of the IBM p690 · SC 2002
Hardware reliability and fault tolerance
soft errors
0.112015
Understanding GPU errors on large-scale HPC systems and the implications for system design and operation · HPCA 2015
Storage systems
storage reliability
0.112014
Best Practices and Lessons Learned from Deploying and Operating Large-Scale Data-Centric Parallel File Systems · SC 2014
Performance modeling and evaluation › system-level analysis
architecture evaluation
0.012002
Early evaluation of the IBM p690 · SC 2002
High-performance computing › supercomputing
supercomputing systems
0.012002
Early evaluation of the IBM p690 · SC 2002

Methods — techniques the papers use, named apart from their topics

neutron beam experiments · 0.2field data analysis · 0.2technology evaluation · 0.2benchmarking · 0.2
YearPublicationVenuePosition
2018 The design, deployment, and evaluation of the CORAL pre-exascale systems
Sudharshan S. Vazhkudai, Bronis R. de Supinski, Arthur S. Bland, Al Geist, James C. Sexton, James A. Kahle, Christopher Zimmer 0001, Scott Atchley, Sarp Oral, Don E. Maxwell, Verónica G. Vergara Larrea, Adam Bertsch, Robin Goldstone, Wayne Joubert, Christopher M. Chambreau, David Appelhans, Robert Blackmore, Ben Casses, George Chochia, Gene Davison, Matthew Ezell, Thomas Gooding, Elsa Gonsiorowski, Leopold Grinberg, Bill Hanson, Bill Hartner, Ian Karlin, Matthew L. Leininger, Dustin Leverman, Chris Marroquin, Adam Moody, Martin Ohmacht, Ramesh Pankajakshan, Fernando Pizzano, James H. Rogers, Bryan S. Rosenburg, Drew Schmidt, Mallikarjun Shankar, Feiyi Wang, Py Watson, Bob Walkup, Lance D. Weems, Junqi Yin
SC3
2015 Understanding GPU errors on large-scale HPC systems and the implications for system design and operation
abstract
Increase in graphics hardware performance and improvements in programmability has enabled GPUs to evolve from a graphics-specific accelerator to a general-purpose computing device. Titan, the world's second fastest supercomputer for open science in 2014, consists of more dum 18,000 GPUs that scientists from various domains such as astrophysics, fusion, climate, and combustion use routinely to run large-scale simulations. Unfortunately, while the performance efficiency of GPUs is well understood, their resilience characteristics in a large-scale computing system have not been fully evaluated. We present a detailed study to provide a thorough understanding of GPU errors on a large-scale GPU-enabled system. Our data was collected from the Titan supercomputer at the Oak Ridge Leadership Computing Facility and a GPU cluster at the Los Alamos National Laboratory. We also present results from our extensive neutron-beam tests, conducted at Los Alamos Neutron Science Center (LANSCE) and at ISIS (Rutherford Appleron Laboratories, UK), to measure the resilience of different generations of GPUs. We present several findings from our field data and neutron-beam experiments, and discuss the implications of our results for future GPU architects, current and future HPC computing facilities, and researchers focusing on GPU resilience.
Devesh Tiwari, Saurabh Gupta 0002, James H. Rogers, Don E. Maxwell, Paolo Rech, Sudharshan S. Vazhkudai, Daniel Oliveira 0002, Dave Londo, Nathan DeBardeleben, Philippe Olivier Alexandre Navaux, Luigi Carro, Arthur S. Bland
HPCA12
2014 Best Practices and Lessons Learned from Deploying and Operating Large-Scale Data-Centric Parallel File Systems
abstract
The Oak Ridge Leadership Computing Facility (OLCF) has deployed multiple large-scale parallel file systems (PFS) to support its operations. During this process, OLCF acquired significant expertise in large-scale storage system design, file system software development, technology evaluation, benchmarking, procurement, deployment, and operational practices. Based on the lessons learned from each new PFS deployment, OLCF improved its operating procedures, and strategies. This paper provides an account of our experience and lessons learned in acquiring, deploying, and operating large-scale parallel file systems. We believe that these lessons will be useful to the wider HPC community.
Sarp Oral, James Simmons, Jason Hill, Dustin Leverman, Feiyi Wang, Matthew Ezell, Ross G. Miller, Douglas Fuller, Raghul Gunasekaran, Youngjae Kim 0001, Saurabh Gupta 0002, Devesh Tiwari, Sudharshan S. Vazhkudai, James H. Rogers, David Dillow, Galen M. Shipman, Arthur S. Bland
SC17
2002 Early evaluation of the IBM p690
abstract
Oak Ridge National Laboratory recently received 27 32-way IBM pSeries 690 SMP nodes. In this paper, we describe our initial evaluation of the p690 architecture, focusing on the performance of benchmarks and applications that are representative of the expected production workload.
Patrick H. Worley, Thomas H. Dunigan, Mark R. Fahey, James B. White III, Arthur S. Bland
SC5