Michael A. Clark

dblp:90/5374 · also Mike A. Clark · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
0since 2021 · last 2019
0000-0001-5211-2002ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
High-performance computing · 54% Interconnection networks and networks-on-chip · 16% Parallel and multicore computing · 10%

Topics — the 20 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
scientific computing systems
0.732018
Simulating the weak death of the Neutron in a femtoscale universe with near-exascale computing · SC 2018
Accelerating lattice QCD multigrid on GPUs using fine-grained parallelization · SC 2016
Scaling lattice QCD beyond 100 GPUs · SC 2011
High-performance computing › scientific computing systems
lattice quantum chromodynamics
0.632018
Simulating the weak death of the Neutron in a femtoscale universe with near-exascale computing · SC 2018
Scaling lattice QCD beyond 100 GPUs · SC 2011
Parallelizing the QUDA Library for Multi-GPU Calculations in Lattice Quantum Chromodynamics · SC 2010
Interconnection networks and networks-on-chip › cluster interconnect
infiniband
0.412019
An evaluation of the CORAL interconnects · SC 2019
Performance modeling and evaluation › benchmarking
interconnect benchmarking
0.412019
An evaluation of the CORAL interconnects · SC 2019
Interconnection networks and networks-on-chip › high-speed networks
supercomputer interconnect
0.412019
An evaluation of the CORAL interconnects · SC 2019
High-performance computing
performance optimization at scale
0.422018
Simulating the weak death of the Neutron in a femtoscale universe with near-exascale computing · SC 2018
Scaling lattice QCD beyond 100 GPUs · SC 2011
High-performance computing › supercomputing
exascale computing
0.312018
Simulating the weak death of the Neutron in a femtoscale universe with near-exascale computing · SC 2018
Parallel and multicore computing › parallelization strategies
fine-grained parallelization
0.212016
Accelerating lattice QCD multigrid on GPUs using fine-grained parallelization · SC 2016
GPUs and heterogeneous computing
GPU computing
0.212016
Accelerating lattice QCD multigrid on GPUs using fine-grained parallelization · SC 2016
High-performance computing › sparse linear solver
multigrid solvers
0.212016
Accelerating lattice QCD multigrid on GPUs using fine-grained parallelization · SC 2016
Memory systems › memory access optimization
cache blocking
0.112011
High-performance lattice QCD for multi-core based parallel systems using a cache-friendly hybrid threaded-MPI approach · SC 2011
GPUs and heterogeneous computing
GPU-accelerated scientific computing
0.112011
Scaling lattice QCD beyond 100 GPUs · SC 2011
Parallel and multicore computing › parallelization strategies
multi-dimensional parallelism
0.112011
Scaling lattice QCD beyond 100 GPUs · SC 2011
High-performance computing › supercomputing
supercomputing systems
0.112019
An evaluation of the CORAL interconnects · SC 2019
Parallel and multicore computing
parallel programming models
0.112016
Accelerating lattice QCD multigrid on GPUs using fine-grained parallelization · SC 2016
High-performance computing
stencil computation
0.112016
Accelerating lattice QCD multigrid on GPUs using fine-grained parallelization · SC 2016
Processor architecture and microarchitecture
instruction set architecture
0.012004
QCDOC: A 10 Teraflops Computer for Tightly-Coupled Calculations · SC 2004
High-performance computing
supercomputing
0.012004
QCDOC: A 10 Teraflops Computer for Tightly-Coupled Calculations · SC 2004
High-performance computing › numerical linear algebra › preconditioner
domain decomposition preconditioner
0.012011
Scaling lattice QCD beyond 100 GPUs · SC 2011
GPUs and heterogeneous computing
multi-GPU computing
0.012010
Parallelizing the QUDA Library for Multi-GPU Calculations in Lattice Quantum Chromodynamics · SC 2010

Methods — techniques the papers use, named apart from their topics

communication benchmarking · 0.4monte carlo simulation · 0.3lattice QCD · 0.3hierarchical multigrid · 0.2fine-grained parallelism mapping · 0.2MPI · 0.2threading · 0.1multi-dimensional parallelization · 0.1domain-decomposed preconditioner · 0.1cache blocking · 0.1
YearPublicationVenuePosition
2019 An evaluation of the CORAL interconnects
abstract
The US Department of Energy deployed the Summit and Sierra supercomputers with the latest state-of-the-art network interconnect technology in 2018 and both systems entered production in 2019. In this paper, we provide an in-depth assessment of the systems' network interconnects that are based on Enhanced Data Rate (EDR) 100 Gb/s Mellanox InfiniBand. Both systems use second-generation EDR Host Channel Adapters (HCAs) and switches with several new features such as Adaptive Routing (AR), switch-based collectives, and HCA-based tag matching. Although based on the same components, Summit's network is "non-blocking" (i.e., a fully provisioned Clos network) and Sierra's network has a 2:1 taper between the racks and aggregation switches. We evaluate the two systems' interconnects using traditional communication benchmarks as well as production applications. We find that the new Adaptive Routing dramatically improves performance but the other new features still need improvement.
Christopher Zimmer 0001, Scott Atchley, Ramesh Pankajakshan, Brian E. Smith, Ian Karlin, Matthew L. Leininger, Adam Bertsch, Brian S. Ryujin, Jason Burmark, André Walker-Loud, Michael A. Clark, Olga Pearce
SC11
2018 Simulating the weak death of the Neutron in a femtoscale universe with near-exascale computing
Evan Berkowitz, Michael A. Clark, Arjun Singh Gambhir, Kenneth S. McElvain, Amy N. Nicholson, Enrico Rinaldi, Pavlos Vranas, André Walker-Loud, Chia-Cheng Chang, Bálint Joó, Thorsten Kurth, Konstantinos Orginos
SC2
2016 Accelerating lattice QCD multigrid on GPUs using fine-grained parallelization
abstract
The past decade has witnessed a dramatic acceleration of lattice quantum chromodynamics calculations in nuclear and particle physics. This has been due to both significant progress in accelerating the iterative linear solvers using multigrid algorithms, and due to the throughput improvements brought by GPUs. Deploying hierarchical algorithms optimally on GPUs is non-trivial owing to the lack of parallelism on the coarse grids, and as such, these advances have not proved multiplicative. Using the QUDA library, we demonstrate that by exposing all sources of parallelism that the underlying stencil problem possesses, and through appropriate mapping of this parallelism to the GPU architecture, we can achieve high efficiency even for the coarsest of grids. Results are presented for the Wilson-Clover discretization, where we demonstrate up to 10x speedup over present state-of-the-art GPU-accelerated methods on Titan. Finally, we look to the future, and consider the software implications of our findings.
Michael A. Clark, Bálint Joó, Alexei Strelchenko, Michael Cheng, Arjun Singh Gambhir, Richard C. Brower
SC1
2014 A Framework for Lattice QCD Calculations on GPUs
abstract
Computing platforms equipped with accelerators like GPUs have proven to provide great computational power. However, exploiting such platforms for existing scientific applications is not a trivial task. Current GPU programming frameworks such as CUDA C/C++ require low-level programming from the developer in order to achieve high performance code. As a result porting of applications to GPUs is typically limited to time-dominant algorithms and routines, leaving the remainder not accelerated which can open a serious Amdahl's law issue. The Lattice QCD application Chroma allows us to explore a different porting strategy. The layered structure of the software architecture logically separates the data-parallel from the application layer. The QCD Data-Parallel software layer provides data types and expressions with stencil-like operations suitable for lattice field theory. Chroma implements algorithms in terms of this high-level interface. Thus by porting the low-level layer one effectively ports the whole application layer in one swing. The QDP-JIT/PTX library, our reimplementation of the low-level layer, provides a framework for Lattice QCD calculations for the CUDA architecture. The complete software interface is supported and thus applications can be run unaltered on GPU-based parallel computers. This reimplementation was possible due to the availability of a JIT compiler which translates an assembly language (PTX) to GPU code. The existing expression templates enabled us to employ compile-time computations in order to build code generators and to automate the memory management for CUDA. Our implementation has allowed us to deploy the full Chroma gauge-generation program on large scale GPU-based machines such as Titan and Blue Waters and accelerate the calculation by more than an order of magnitude.
Frank Tobias Winter, Michael A. Clark, Robert G. Edwards, Bálint Joó
IPDPS2
2011 Scaling lattice QCD beyond 100 GPUs
abstract
Over the past five years, graphics processing units (GPUs) have had a transformational effect on numerical lattice quantum chromodynamics (LQCD) calculations in nuclear and particle physics. While GPUs have been applied with great success to the post-Monte Carlo "analysis" phase which accounts for a substantial fraction of the workload in a typical LQCD calculation, the initial Monte Carlo "gauge field generation" phase requires capability-level supercomputing, corresponding to O(100) GPUs or more. Such strong scaling has not been previously achieved. In this contribution, we demonstrate that using a multi-dimensional parallelization strategy and a domain-decomposed preconditioner allows us to scale into this regime. We present results for two popular discretizations of the Dirac operator, Wilson-clover and improved staggered, employing up to 256 GPUs on the Edge cluster at Lawrence Livermore National Laboratory.
Ronald Babich, Michael A. Clark, Bálint Joó, Guochun Shi, Richard C. Brower, Steven A. Gottlieb
SC2
2011 High-performance lattice QCD for multi-core based parallel systems using a cache-friendly hybrid threaded-MPI approach
abstract
Lattice Quantum Chromo-dynamics (LQCD) is a computationally challenging problem that solves the discretized Dirac equation in the presence of an SU(3) gauge field. Its key operation is a matrix-vector product, known as the Dslash operator. We have developed a novel multicore architecture-friendly implementation of the Wilson-Dslash operator which delivers 75 Gflops (single-precision) on an Intel® Xeon® Processor X5680 achieving 60% computational efficiency for datasets that fit in the last-level cache. For datasets larger than the last-level cache, this performance drops to 50 Gflops. Our performance is 2-3X higher than a well-known implementation from the Chroma software suite when running on the same hardware platform. The novel implementation of LQCD reported in this paper is based on recently published the 3.5D spatial and 4.5D temporal tiling schemes. Both blocking schemes significantly reduce LQCD external memory bandwidth requirements, delivering a more compute-bound implementation. The performance advantage of our schemes will become more significant as the gap between compute flops and external memory bandwidth continues to grow. We demonstrate very good cluster-level scalability of our implementation: for a lattice of 323 x 256 sites, we achieve over 4 Tflops when strong-scaled to a 128 node system (1536 cores total). For the same lattice size, a full Conjugate Gradients Wilson-Dslash operator, achieves 2.95 Tflops.
Mikhail Smelyanskiy, Karthikeyan Vaidyanathan, Jee W. Choi, Bálint Joó, Jatin Chhugani, Michael A. Clark, Pradeep Dubey
SC6
2010 Parallelizing the QUDA Library for Multi-GPU Calculations in Lattice Quantum Chromodynamics
abstract
Graphics Processing Units (GPUs) are having a transformational effect on numerical lattice quantum chromo- dynamics (LQCD) calculations of importance in nuclear and particle physics. The QUDA library provides a package of mixed precision sparse matrix linear solvers for LQCD applications, supporting single GPUs based on NVIDIA's Compute Unified Device Architecture (CUDA). This library, interfaced to the QDP++/Chroma framework for LQCD calculations, is currently in production use on the "9g" cluster at the Jefferson Laboratory, enabling unprecedented price/performance for a range of problems in LQCD. Nevertheless, memory constraints on current GPU devices limit the problem sizes that can be tackled. In this contribution we describe the parallelization of the QUDA library onto multiple GPUs using MPI, including strategies for the overlapping of communication and computation. We report on both weak and strong scaling for up to 32 GPUs interconnected by InfiniBand, on which we sustain in excess of 4 Tflops.
Ronald Babich, Michael A. Clark, Bálint Joó
SC2
2004 QCDOC: A 10 Teraflops Computer for Tightly-Coupled Calculations
abstract
Numerical simulations of the strong nuclear force, known as quantum chromodynamics or QCD, have proven to be a demanding, forefront problem in high-performance computing. In this report, we describe a new computer, QCDOC (QCD On a Chip), designed for optimal price/performance in the study of QCD. QCDOC uses a six-dimensional, low-latency mesh network to connect processing nodes, each of which includes a single custom ASIC, designed by our collaboration and built by IBM, plus DDR SDRAM. Each node has a peak speed of 1Gigaflops and two 12,288node, 10+ Teraflops machines are to be completed in the fall of 2004. Currently, a 512 node machine is running, delivering efficiencies as high as 45% of peak on the conjugate gradient solvers that dominate our calculations and a 4096-node machine with a cost of $1.6M is under construction. This should give us a price/performance less than $1per sustained Megaflops.
Peter A. Boyle, Dong Chen 0005, Norman H. Christ, Michael A. Clark, Saul D. Cohen, Zhihua Dong, Alan Gara, Bálint Joó, Chulwoo Jung, Ludmila A. Levkova, Xiaodong Liao, Guofeng Liu, Robert D. Mawhinney, Shigemi Ohta, Konstantin Petrov, Tilo Wettig, Azusa Yamaguchi, Calin Cristian
SC4