EDBT 2026 Demo / reviewers in the wild / expert
Michael A. Clark
dblp:90/5374 · also Mike A. Clark
· DBLP profile ↗
8ranked-venue papers
1as first author
0since 2021 · last 2019
0000-0001-5211-2002ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
High-performance computing · 54% Interconnection networks and networks-on-chip · 16% Parallel and multicore computing · 10% |
Topics — the 20 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing
scientific computing systems |
0.7 | 3 | 2018 | Simulating the weak death of the Neutron in a femtoscale universe with near-exascale computing · SC 2018 Accelerating lattice QCD multigrid on GPUs using fine-grained parallelization · SC 2016 Scaling lattice QCD beyond 100 GPUs · SC 2011 |
High-performance computing › scientific computing systems
lattice quantum chromodynamics |
0.6 | 3 | 2018 | Simulating the weak death of the Neutron in a femtoscale universe with near-exascale computing · SC 2018 Scaling lattice QCD beyond 100 GPUs · SC 2011 Parallelizing the QUDA Library for Multi-GPU Calculations in Lattice Quantum Chromodynamics · SC 2010 |
Interconnection networks and networks-on-chip › cluster interconnect
infiniband |
0.4 | 1 | 2019 | An evaluation of the CORAL interconnects · SC 2019 |
Performance modeling and evaluation › benchmarking
interconnect benchmarking |
0.4 | 1 | 2019 | An evaluation of the CORAL interconnects · SC 2019 |
Interconnection networks and networks-on-chip › high-speed networks
supercomputer interconnect |
0.4 | 1 | 2019 | An evaluation of the CORAL interconnects · SC 2019 |
High-performance computing
performance optimization at scale |
0.4 | 2 | 2018 | Simulating the weak death of the Neutron in a femtoscale universe with near-exascale computing · SC 2018 Scaling lattice QCD beyond 100 GPUs · SC 2011 |
High-performance computing › supercomputing
exascale computing |
0.3 | 1 | 2018 | Simulating the weak death of the Neutron in a femtoscale universe with near-exascale computing · SC 2018 |
Parallel and multicore computing › parallelization strategies
fine-grained parallelization |
0.2 | 1 | 2016 | Accelerating lattice QCD multigrid on GPUs using fine-grained parallelization · SC 2016 |
GPUs and heterogeneous computing
GPU computing |
0.2 | 1 | 2016 | Accelerating lattice QCD multigrid on GPUs using fine-grained parallelization · SC 2016 |
High-performance computing › sparse linear solver
multigrid solvers |
0.2 | 1 | 2016 | Accelerating lattice QCD multigrid on GPUs using fine-grained parallelization · SC 2016 |
Memory systems › memory access optimization
cache blocking |
0.1 | 1 | 2011 | High-performance lattice QCD for multi-core based parallel systems using a cache-friendly hybrid threaded-MPI approach · SC 2011 |
GPUs and heterogeneous computing
GPU-accelerated scientific computing |
0.1 | 1 | 2011 | Scaling lattice QCD beyond 100 GPUs · SC 2011 |
Parallel and multicore computing › parallelization strategies
multi-dimensional parallelism |
0.1 | 1 | 2011 | Scaling lattice QCD beyond 100 GPUs · SC 2011 |
High-performance computing › supercomputing
supercomputing systems |
0.1 | 1 | 2019 | An evaluation of the CORAL interconnects · SC 2019 |
Parallel and multicore computing
parallel programming models |
0.1 | 1 | 2016 | Accelerating lattice QCD multigrid on GPUs using fine-grained parallelization · SC 2016 |
High-performance computing
stencil computation |
0.1 | 1 | 2016 | Accelerating lattice QCD multigrid on GPUs using fine-grained parallelization · SC 2016 |
Processor architecture and microarchitecture
instruction set architecture |
0.0 | 1 | 2004 | QCDOC: A 10 Teraflops Computer for Tightly-Coupled Calculations · SC 2004 |
High-performance computing
supercomputing |
0.0 | 1 | 2004 | QCDOC: A 10 Teraflops Computer for Tightly-Coupled Calculations · SC 2004 |
High-performance computing › numerical linear algebra › preconditioner
domain decomposition preconditioner |
0.0 | 1 | 2011 | Scaling lattice QCD beyond 100 GPUs · SC 2011 |
GPUs and heterogeneous computing
multi-GPU computing |
0.0 | 1 | 2010 | Parallelizing the QUDA Library for Multi-GPU Calculations in Lattice Quantum Chromodynamics · SC 2010 |
Methods — techniques the papers use, named apart from their topics
communication benchmarking · 0.4monte carlo simulation · 0.3lattice QCD · 0.3hierarchical multigrid · 0.2fine-grained parallelism mapping · 0.2MPI · 0.2threading · 0.1multi-dimensional parallelization · 0.1domain-decomposed preconditioner · 0.1cache blocking · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | An evaluation of the CORAL interconnectsabstractThe US Department of Energy deployed the Summit and Sierra supercomputers with the latest state-of-the-art network interconnect technology in 2018 and both systems entered production in 2019. In this paper, we provide an in-depth assessment of the systems' network interconnects that are based on Enhanced Data Rate (EDR) 100 Gb/s Mellanox InfiniBand. Both systems use second-generation EDR Host Channel Adapters (HCAs) and switches with several new features such as Adaptive Routing (AR), switch-based collectives, and HCA-based tag matching. Although based on the same components, Summit's network is "non-blocking" (i.e., a fully provisioned Clos network) and Sierra's network has a 2:1 taper between the racks and aggregation switches. We evaluate the two systems' interconnects using traditional communication benchmarks as well as production applications. We find that the new Adaptive Routing dramatically improves performance but the other new features still need improvement. Christopher Zimmer 0001, Scott Atchley, Ramesh Pankajakshan, Brian E. Smith, Ian Karlin, Matthew L. Leininger, Adam Bertsch, Brian S. Ryujin, Jason Burmark, André Walker-Loud, Michael A. Clark, Olga Pearce |
SC | 11 |
| 2018 | Simulating the weak death of the Neutron in a femtoscale universe with near-exascale computing
Evan Berkowitz, Michael A. Clark, Arjun Singh Gambhir, Kenneth S. McElvain, Amy N. Nicholson, Enrico Rinaldi, Pavlos Vranas, André Walker-Loud, Chia-Cheng Chang, Bálint Joó, Thorsten Kurth, Konstantinos Orginos |
SC | 2 |
| 2016 | Accelerating lattice QCD multigrid on GPUs using fine-grained parallelizationabstractThe past decade has witnessed a dramatic acceleration of lattice quantum chromodynamics calculations in nuclear and particle physics. This has been due to both significant progress in accelerating the iterative linear solvers using multigrid algorithms, and due to the throughput improvements brought by GPUs. Deploying hierarchical algorithms optimally on GPUs is non-trivial owing to the lack of parallelism on the coarse grids, and as such, these advances have not proved multiplicative. Using the QUDA library, we demonstrate that by exposing all sources of parallelism that the underlying stencil problem possesses, and through appropriate mapping of this parallelism to the GPU architecture, we can achieve high efficiency even for the coarsest of grids. Results are presented for the Wilson-Clover discretization, where we demonstrate up to 10x speedup over present state-of-the-art GPU-accelerated methods on Titan. Finally, we look to the future, and consider the software implications of our findings. Michael A. Clark, Bálint Joó, Alexei Strelchenko, Michael Cheng, Arjun Singh Gambhir, Richard C. Brower |
SC | 1 |
| 2014 | A Framework for Lattice QCD Calculations on GPUsabstractComputing platforms equipped with accelerators like GPUs have proven to provide great computational power. However, exploiting such platforms for existing scientific applications is not a trivial task. Current GPU programming frameworks such as CUDA C/C++ require low-level programming from the developer in order to achieve high performance code. As a result porting of applications to GPUs is typically limited to time-dominant algorithms and routines, leaving the remainder not accelerated which can open a serious Amdahl's law issue. The Lattice QCD application Chroma allows us to explore a different porting strategy. The layered structure of the software architecture logically separates the data-parallel from the application layer. The QCD Data-Parallel software layer provides data types and expressions with stencil-like operations suitable for lattice field theory. Chroma implements algorithms in terms of this high-level interface. Thus by porting the low-level layer one effectively ports the whole application layer in one swing. The QDP-JIT/PTX library, our reimplementation of the low-level layer, provides a framework for Lattice QCD calculations for the CUDA architecture. The complete software interface is supported and thus applications can be run unaltered on GPU-based parallel computers. This reimplementation was possible due to the availability of a JIT compiler which translates an assembly language (PTX) to GPU code. The existing expression templates enabled us to employ compile-time computations in order to build code generators and to automate the memory management for CUDA. Our implementation has allowed us to deploy the full Chroma gauge-generation program on large scale GPU-based machines such as Titan and Blue Waters and accelerate the calculation by more than an order of magnitude. Frank Tobias Winter, Michael A. Clark, Robert G. Edwards, Bálint Joó |
IPDPS | 2 |
| 2011 | Scaling lattice QCD beyond 100 GPUsabstractOver the past five years, graphics processing units (GPUs) have had a transformational effect on numerical lattice quantum chromodynamics (LQCD) calculations in nuclear and particle physics. While GPUs have been applied with great success to the post-Monte Carlo "analysis" phase which accounts for a substantial fraction of the workload in a typical LQCD calculation, the initial Monte Carlo "gauge field generation" phase requires capability-level supercomputing, corresponding to O(100) GPUs or more. Such strong scaling has not been previously achieved. In this contribution, we demonstrate that using a multi-dimensional parallelization strategy and a domain-decomposed preconditioner allows us to scale into this regime. We present results for two popular discretizations of the Dirac operator, Wilson-clover and improved staggered, employing up to 256 GPUs on the Edge cluster at Lawrence Livermore National Laboratory. Ronald Babich, Michael A. Clark, Bálint Joó, Guochun Shi, Richard C. Brower, Steven A. Gottlieb |
SC | 2 |
| 2011 | High-performance lattice QCD for multi-core based parallel systems using a cache-friendly hybrid threaded-MPI approachabstractLattice Quantum Chromo-dynamics (LQCD) is a computationally challenging problem that solves the discretized Dirac equation in the presence of an SU(3) gauge field. Its key operation is a matrix-vector product, known as the Dslash operator. We have developed a novel multicore architecture-friendly implementation of the Wilson-Dslash operator which delivers 75 Gflops (single-precision) on an Intel® Xeon® Processor X5680 achieving 60% computational efficiency for datasets that fit in the last-level cache. For datasets larger than the last-level cache, this performance drops to 50 Gflops. Our performance is 2-3X higher than a well-known implementation from the Chroma software suite when running on the same hardware platform. The novel implementation of LQCD reported in this paper is based on recently published the 3.5D spatial and 4.5D temporal tiling schemes. Both blocking schemes significantly reduce LQCD external memory bandwidth requirements, delivering a more compute-bound implementation. The performance advantage of our schemes will become more significant as the gap between compute flops and external memory bandwidth continues to grow. We demonstrate very good cluster-level scalability of our implementation: for a lattice of 323 x 256 sites, we achieve over 4 Tflops when strong-scaled to a 128 node system (1536 cores total). For the same lattice size, a full Conjugate Gradients Wilson-Dslash operator, achieves 2.95 Tflops. Mikhail Smelyanskiy, Karthikeyan Vaidyanathan, Jee W. Choi, Bálint Joó, Jatin Chhugani, Michael A. Clark, Pradeep Dubey |
SC | 6 |
| 2010 | Parallelizing the QUDA Library for Multi-GPU Calculations in Lattice Quantum ChromodynamicsabstractGraphics Processing Units (GPUs) are having a transformational effect on numerical lattice quantum chromo- dynamics (LQCD) calculations of importance in nuclear and particle physics. The QUDA library provides a package of mixed precision sparse matrix linear solvers for LQCD applications, supporting single GPUs based on NVIDIA's Compute Unified Device Architecture (CUDA). This library, interfaced to the QDP++/Chroma framework for LQCD calculations, is currently in production use on the "9g" cluster at the Jefferson Laboratory, enabling unprecedented price/performance for a range of problems in LQCD. Nevertheless, memory constraints on current GPU devices limit the problem sizes that can be tackled. In this contribution we describe the parallelization of the QUDA library onto multiple GPUs using MPI, including strategies for the overlapping of communication and computation. We report on both weak and strong scaling for up to 32 GPUs interconnected by InfiniBand, on which we sustain in excess of 4 Tflops. Ronald Babich, Michael A. Clark, Bálint Joó |
SC | 2 |
| 2004 | QCDOC: A 10 Teraflops Computer for Tightly-Coupled CalculationsabstractNumerical simulations of the strong nuclear force, known as quantum chromodynamics or QCD, have proven to be a demanding, forefront problem in high-performance computing. In this report, we describe a new computer, QCDOC (QCD On a Chip), designed for optimal price/performance in the study of QCD. QCDOC uses a six-dimensional, low-latency mesh network to connect processing nodes, each of which includes a single custom ASIC, designed by our collaboration and built by IBM, plus DDR SDRAM. Each node has a peak speed of 1Gigaflops and two 12,288node, 10+ Teraflops machines are to be completed in the fall of 2004. Currently, a 512 node machine is running, delivering efficiencies as high as 45% of peak on the conjugate gradient solvers that dominate our calculations and a 4096-node machine with a cost of $1.6M is under construction. This should give us a price/performance less than $1per sustained Megaflops. Peter A. Boyle, Dong Chen 0005, Norman H. Christ, Michael A. Clark, Saul D. Cohen, Zhihua Dong, Alan Gara, Bálint Joó, Chulwoo Jung, Ludmila A. Levkova, Xiaodong Liao, Guofeng Liu, Robert D. Mawhinney, Shigemi Ohta, Konstantin Petrov, Tilo Wettig, Azusa Yamaguchi, Calin Cristian |
SC | 4 |