Richard F. Barrett

dblp:33/4794 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
High-performance computing · 50% Performance modeling and evaluation · 39% Energy-efficient computing · 6%

Topics — the 2 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation
benchmarking
0.232008
Early evaluation of IBM BlueGene/P · SC 2008
Cray XT4: an early evaluation for petascale scientific simulation · SC 2007
Performance evaluation of the cray XT3 configured with dual core opteron processors · PPoPP 2007
High-performance computing › supercomputing
supercomputing systems
0.222008
Early evaluation of IBM BlueGene/P · SC 2008
Performance evaluation of the cray XT3 configured with dual core opteron processors · PPoPP 2007

Methods — techniques the papers use, named apart from their topics

microbenchmarks · 0.2application benchmarks · 0.2performance evaluation · 0.1
YearPublicationVenuePosition
2017 PeaPaw: Performance and Energy-Aware Partitioning of Workload on Heterogeneous Platforms
abstract
Performance and energy are two major concerns for application development on heterogeneous platforms. It is challenging for application developers to fully exploit the performance/energy potential of heterogeneous platforms. One reason is the lack of reliable prediction of the system’s performance/energy before application implementation. Another reason is that a heterogeneous platform presents a large design space for workload partitioning between different processors. To reduce such development cost, this article proposes a framework, PeaPaw, to assist application developers to identify a workload partition (WP) that has high potential leading to high performance or energy efficiency before actual implementation. The PeaPaw framework includes both analytical performance/energy models and two sets of workload partitioning guidelines. Based on the design goal, application developers can obtain a workload partitioning guideline from PeaPaw for a given platform and follow it to design one or multiple WPs for a given workload. Then PeaPaw can be used to estimate the performance/energy of the designed WPs, and the WP with the best estimated performance/energy can be selected for actual implementation. To demonstrate the effectiveness of PeaPaw, we have conducted three case studies. Results from these case studies show that PeaPaw can faithfully estimate the performance/energy relationships of WPs and provide effective workload partitioning guidelines.
Li Tang 0007, Richard F. Barrett, Jeanine E. Cook, Xiaobo Sharon Hu
ACM Trans. Design Autom. Electr. Syst.2
2015 Enabling Tractable Exploration of the Performance of Adaptive Mesh Refinement
abstract
A broad range of physical phenomena in science and engineering can be explored using finite difference and volume based application codes. Incorporating Adaptive Mesh Refinement (AMR) into these codes focuses attention on the most critical parts of a simulation, enabling increased numerical accuracy of the solution while limiting memory consumption. However, adaptivity comes at the cost of increased runtime complexity, which is particularly challenging on emerging and expected future architectures. In order to explore the design space offered by new computing environments, we have developed a proxy application called miniAMR. MiniAMR exposes a range of the important issues that will significantly impact the performance potential of full application codes. In this paper, we describe miniAMR, demonstrate what is designed to represent in a full application code, and illustrate how it can be used to exploit future high performance computing architectures. To ensure an accurate understanding of what miniAMR is intended to represent, we compare it with CTH, a shock hydrodynamics code in heavy use throughout several computational science and engineering communities.
Courtenay T. Vaughan, Richard F. Barrett
CLUSTER2
2015 Assessing a mini-application as a performance proxy for a finite element method engineering application
abstract
Summary The performance of a large‐scale, production‐quality science and engineering application (‘app’) is often dominated by a small subset of the code. Even within that subset, computational and data access patterns are often repeated, so that an even smaller portion can represent the performance‐impacting features. If application developers, parallel computing experts, and computer architects can together identify this representative subset and then develop a small mini‐application (‘miniapp’) that can capture these primary performance characteristics, then this miniapp can be used to both improve the performance of the app as well as provide a tool for co‐design for the high‐performance computing community. However, a critical question is whether a miniapp can effectively capture key performance behavior of an app. This study provides a comparison of an implicit finite element semiconductor device modeling app on unstructured meshes with an implicit finite element miniapp on unstructured meshes. The goal is to assess whether the miniapp is predictive of the performance of the app. Single compute node performance will be compared, as well as scaling up to 16,000 cores. Results indicate that the miniapp can be reasonably predictive of the performance characteristics of the app for a single iteration of the solver on a single compute node. Published 2015. This article is a U.S. Government work and is in the public domain in the USA.
Paul T. Lin, Michael A. Heroux, Richard F. Barrett, Alan B. Williams
Concurr. Comput. Pract. Exp.3
2015 Assessing the role of mini-applications in predicting key performance characteristics of scientific and engineering applications
Richard F. Barrett, Paul S. Crozier, Douglas Doerfler, Michael A. Heroux, Paul T. Lin, Heidi Thornquist, Timothy G. Trucano, Courtenay T. Vaughan
J. Parallel Distributed Comput.1
2014 Exascale design space exploration and co-design
Sudip S. Dosanjh, Richard F. Barrett, Douglas Doerfler, Simon D. Hammond, Karl S. Hemmert, Michael A. Heroux, Paul T. Lin, Kevin T. Pedretti, Arun Rodrigues, Timothy G. Trucano, Justin Luitjens
Future Gener. Comput. Syst.2
2013 GPU acceleration of Data Assembly in Finite Element Methods and its energy implications
abstract
The Finite Element Method (FEM) is a numerical technique widely used in finding approximate solutions for many scientific and engineering problems. The Data Assembly (DA) stage in FEM can take up to 50% of the total FEM execution time. Accelerating DA with Graphics Processing Units (GPUs) presents challenges due to DA's mixed compute-intensive and memory-intensive workloads. This paper uses a representative finite element mini-application to explore DA acceleration on CPU+GPU platforms. Implementations based on different thread, kernel and task design approaches are developed and compared. Their performance and energy consumption are measured on four CPU+GPU and two CPU only platforms. The results show that (i) the performance and energy for different implementations on the same platform can vary significantly but the performance and energy trends are the same, and (ii) there exist performance and energy tradeoffs across some platforms if the best implementation is chosen for each of the platforms.
Li Tang 0007, Xiaobo Sharon Hu, Danny Ziyi Chen, Michael T. Niemier, Richard F. Barrett, Simon D. Hammond, Genie Hsieh
ASAP5
2012 Toward codesign in high performance computing systems
abstract
Preparations for exascale computing have led to the realization that computing environments will be significantly different from those that provide petascale capabilities. This change is driven by energy constraints, which has compelled hardware architects to design systems that will require a significant re-thinking of how application algorithms are selected and implemented. The "codesign" principle may offer a common basis for application and system developers as well as architects to work synergistically towards achieving exascale computing. This paper aims to introduce to the embedded system design community the unique challenges and opportunities as well as exciting developments in exascale HPC system codesign. Given the success of adopting codesign practices in the embedded system design area, this effort should be mutually beneficial to both communities.
Richard F. Barrett, Xiaobo Sharon Hu, Sudip S. Dosanjh, Steven G. Parker, Michael A. Heroux, John Shalf
ICCAD1
2012 Application-driven analysis of two generations of capability computing: the transition to multicore processors
abstract
SUMMARY Multicore processors form the basis of most traditional high performance parallel processing architectures. Early experiences with these computers showed significant performance problems, both with regard to computation and inter‐process communication. The transition from Purple, an IBM POWER5‐based machine, to Cielo, a Cray XE6, as the main capability computing platform for the United States Department of Energy's Advanced Simulation and Computing campaign provides an opportunity to reexamine these issues after experiences with a few generations of multicore‐based machines. Experiences with Purple identified some important characteristics that led to strong performance of complex scientific application programs at very large scales. Herein, we compare the performance of some Advanced Simulation and Computing mission critical applications at capability scale across this transition to multicore processors. Copyright © 2012 John Wiley & Sons, Ltd.
Mahesh Rajan, Courtenay T. Vaughan, Douglas Doerfler, Richard F. Barrett, Paul T. Lin, Kevin T. Pedretti, Karl S. Hemmert
Concurr. Comput. Pract. Exp.4
2011 Achieving Exascale Computing through Hardware/Software Co-design
Sudip S. Dosanjh, Richard F. Barrett, Michael A. Heroux, Arun Rodrigues
EuroMPI2
2009 Impact of Quad-Core Cray XT4 System and Software Stack on Scientific Computation
Sadaf R. Alam, Richard F. Barrett, Heike Jagode, Jeffery A. Kuehn, Stephen W. Poole, Ramanan Sankaran
Euro-Par2
2009 Performance Characterization of a Hierarchical MPI Implementation on Large-scale Distributed-memory Platforms
abstract
The building blocks of emerging Petascale massively parallel processing (MPP) systems are multi-core processors with four or more cores as a single processing element and a customized network interface. The resulting memory and communication hierarchy of these platforms are now exposed to application developers and end users by creating a hierarchical or multi-core aware message-passing (MPI) programming interface and by providing a handful of runtime, tunable parameters that allows mapping and control of MPI tasks and message handling. We characterize performance of MPI communication patterns and present strategies for optimizing applications performance on Cray XT series systems that are composed of contemporary AMD processors and a proprietary network infrastructure. We highlight dependencies in its memory and network subsystems, which could influence production-level applications performance. We demonstrate that MPI micro-benchmarks could mislead an application developer or end user since these benchmarks often do not expose the interplay between memory allocation and usage in the user space, which depends on the number of tasks or cores and workload characteristics. Our studies show performance improvements compared to the default options for our target scientific benchmarks and production-level applications.
Sadaf R. Alam, Richard F. Barrett, Jeffery A. Kuehn, Stephen W. Poole
ICPP2
2009 Performance analysis and projections for Petascale applications on Cray XT series systems
abstract
The Petascale Cray XT5 system at the Oak Ridge National Laboratory (ORNL) Leadership Computing Facility (LCF) shares a number of system and software features with its predecessor, the Cray XT4 system including the quad-core AMD processor and a multi-core aware MPI library. We analyze performance of scalable scientific applications on the quad-core Cray XT4 system as part of the early system access using a combination of micro-benchmarks and Petascale ready applications. Particularly, we evaluate impact of key changes that occurred during the dual-core to quad-core processor upgrade on applications behavior and provide projections for the next-generation massively-parallel platforms with multi-core processors, specifically for proposed Petascale Cray XT5 system. We compare and contrast the quad-core XT4 system features with the upcoming XT5 system and discuss strategies for improving scaling and performance for our target applications.
Sadaf R. Alam, Richard F. Barrett, Jeffery A. Kuehn, Stephen W. Poole
IPDPS2
2008 Early evaluation of IBM BlueGene/P
abstract
BlueGene/P (BG/P) is the second generation BlueGene architecture from IBM, succeeding BlueGene/L (BG/L). BG/P is a system-on-a-chip (SoC) design that uses four PowerPC 450 cores operating at 850 MHz with a double precision, dual pipe floating point unit per core. These chips are connected with multiple interconnection networks including a 3-D torus, a global collective network, and a global barrier network. The design is intended to provide a highly scalable, physically dense system with relatively low power requirements per flop. In this paper, we report on our examination of BG/P, presented in the context of a set of important scientific applications, and as compared to other major large scale supercomputers in use today. Our investigation confirms that BG/P has good scalability with an expected lower performance per processor when compared to the Cray XT4's Opteron. We also find that BG/P uses very low power per floating point operation for certain kernels, yet it has less of a power advantage when considering science driven metrics for mission applications.
Sadaf R. Alam, Richard F. Barrett, M. Bast, Mark R. Fahey, Jeffery A. Kuehn, Collin McCurdy, James H. Rogers, Philip C. Roth, Ramanan Sankaran, Jeffrey S. Vetter, Patrick H. Worley, Weikuan Yu
SC2
2007 Performance evaluation of the cray XT3 configured with dual core opteron processors
abstract
No abstract available.
Richard F. Barrett, Sadaf R. Alam, Jeffrey S. Vetter
PPoPP1
2007 Cray XT4: an early evaluation for petascale scientific simulation
abstract
The scientific simulation capabilities of next generation high-end computing technology will depend on striking a balance among memory, processor, I/O, and local and global network performance across the breadth of the scientific simulation space. The Cray XT4 combines commodity AMD dual core Opteron processor technology with the second generation of Cray's custom communication accelerator in a system design whose balance is claimed to be driven by the demands of scientific simulation. This paper presents an evaluation of the Cray XT4 using micro-benchmarks to develop a controlled understanding of individual system components, providing the context for analyzing and comprehending the performance of several petascale-ready applications. Results gathered from several strategic application domains are compared with observations on the previous generation Cray XT3 and other high-end computing systems, demonstrating performance improvements across a wide variety of application benchmark problems.
Sadaf R. Alam, Jeffery A. Kuehn, Richard F. Barrett, Jeffrey M. Larkin, Mark R. Fahey, Ramanan Sankaran, Patrick H. Worley
SC3