Ronald Babich

dblp:49/8432 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
0since 2021 · last 2011
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 60% GPUs and heterogeneous computing · 22% Parallel and multicore computing · 17%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › scientific computing systems
lattice quantum chromodynamics
0.222011
Scaling lattice QCD beyond 100 GPUs · SC 2011
Parallelizing the QUDA Library for Multi-GPU Calculations in Lattice Quantum Chromodynamics · SC 2010
GPUs and heterogeneous computing
GPU-accelerated scientific computing
0.112011
Scaling lattice QCD beyond 100 GPUs · SC 2011
Parallel and multicore computing › parallelization strategies
multi-dimensional parallelism
0.112011
Scaling lattice QCD beyond 100 GPUs · SC 2011
High-performance computing
scientific computing systems
0.112011
Scaling lattice QCD beyond 100 GPUs · SC 2011
High-performance computing › numerical linear algebra › preconditioner
domain decomposition preconditioner
0.012011
Scaling lattice QCD beyond 100 GPUs · SC 2011
High-performance computing
performance optimization at scale
0.012011
Scaling lattice QCD beyond 100 GPUs · SC 2011
GPUs and heterogeneous computing
multi-GPU computing
0.012010
Parallelizing the QUDA Library for Multi-GPU Calculations in Lattice Quantum Chromodynamics · SC 2010

Methods — techniques the papers use, named apart from their topics

multi-dimensional parallelization · 0.1domain-decomposed preconditioner · 0.1mixed precision · 0.1MPI · 0.1
YearPublicationVenuePosition
2011 Scaling lattice QCD beyond 100 GPUs
abstract
Over the past five years, graphics processing units (GPUs) have had a transformational effect on numerical lattice quantum chromodynamics (LQCD) calculations in nuclear and particle physics. While GPUs have been applied with great success to the post-Monte Carlo "analysis" phase which accounts for a substantial fraction of the workload in a typical LQCD calculation, the initial Monte Carlo "gauge field generation" phase requires capability-level supercomputing, corresponding to O(100) GPUs or more. Such strong scaling has not been previously achieved. In this contribution, we demonstrate that using a multi-dimensional parallelization strategy and a domain-decomposed preconditioner allows us to scale into this regime. We present results for two popular discretizations of the Dirac operator, Wilson-clover and improved staggered, employing up to 256 GPUs on the Edge cluster at Lawrence Livermore National Laboratory.
Ronald Babich, Michael A. Clark, Bálint Joó, Guochun Shi, Richard C. Brower, Steven A. Gottlieb
SC1
2010 Parallelizing the QUDA Library for Multi-GPU Calculations in Lattice Quantum Chromodynamics
abstract
Graphics Processing Units (GPUs) are having a transformational effect on numerical lattice quantum chromo- dynamics (LQCD) calculations of importance in nuclear and particle physics. The QUDA library provides a package of mixed precision sparse matrix linear solvers for LQCD applications, supporting single GPUs based on NVIDIA's Compute Unified Device Architecture (CUDA). This library, interfaced to the QDP++/Chroma framework for LQCD calculations, is currently in production use on the "9g" cluster at the Jefferson Laboratory, enabling unprecedented price/performance for a range of problems in LQCD. Nevertheless, memory constraints on current GPU devices limit the problem sizes that can be tackled. In this contribution we describe the parallelization of the QUDA library onto multiple GPUs using MPI, including strategies for the overlapping of communication and computation. We report on both weak and strong scaling for up to 32 GPUs interconnected by InfiniBand, on which we sustain in excess of 4 Tflops.
Ronald Babich, Michael A. Clark, Bálint Joó
SC1