Steven A. Gottlieb

dblp:46/7222 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
0since 2021 · last 2011
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 60% GPUs and heterogeneous computing · 20% Parallel and multicore computing · 20%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › scientific computing systems
lattice quantum chromodynamics
0.122011
Scaling lattice QCD beyond 100 GPUs · SC 2011
Lattice QCD on the IBM Scalable POWERParallel Systems SP2 · SC 1995
High-performance computing
scientific computing systems
0.122011
Scaling lattice QCD beyond 100 GPUs · SC 2011
Lattice QCD on the IBM Scalable POWERParallel Systems SP2 · SC 1995
GPUs and heterogeneous computing
GPU-accelerated scientific computing
0.112011
Scaling lattice QCD beyond 100 GPUs · SC 2011
Parallel and multicore computing › parallelization strategies
multi-dimensional parallelism
0.112011
Scaling lattice QCD beyond 100 GPUs · SC 2011
High-performance computing
performance optimization at scale
0.122011
Scaling lattice QCD beyond 100 GPUs · SC 2011
Lattice QCD on the IBM Scalable POWERParallel Systems SP2 · SC 1995
High-performance computing › numerical linear algebra › preconditioner
domain decomposition preconditioner
0.012011
Scaling lattice QCD beyond 100 GPUs · SC 2011
High-performance computing
code optimization
0.011995
Lattice QCD on the IBM Scalable POWERParallel Systems SP2 · SC 1995

Methods — techniques the papers use, named apart from their topics

multi-dimensional parallelization · 0.1domain-decomposed preconditioner · 0.1sparse system inversion · 0.0
YearPublicationVenuePosition
2011 Design of MILC Lattice QCD Application for GPU Clusters
abstract
We present an implementation of the improved staggered quark action lattice QCD computation designed for execution on a GPU cluster. The parallelization strategy is based on dividing the space-time lattice along the time dimension and distributing the sub-lattices among the GPU cluster nodes. We provide a mixed-precision floating-point GPU implementation of the multi-mass conjugate gradient solver. Our single GPU implementation of the conjugate gradient solver achieves a 9x performance improvement over the highly optimized code executed on a state-of-the-art eight-core CPU node. The overall application executes almost six times faster on a GPU-enabled cluster vs. a conventional multi-core cluster. The developed code is currently used for running production QCD calculations with electromagnetic corrections.
Guochun Shi, Steven A. Gottlieb, Aaron Torok, Volodymyr V. Kindratenko
IPDPS2
2011 Scaling lattice QCD beyond 100 GPUs
abstract
Over the past five years, graphics processing units (GPUs) have had a transformational effect on numerical lattice quantum chromodynamics (LQCD) calculations in nuclear and particle physics. While GPUs have been applied with great success to the post-Monte Carlo "analysis" phase which accounts for a substantial fraction of the workload in a typical LQCD calculation, the initial Monte Carlo "gauge field generation" phase requires capability-level supercomputing, corresponding to O(100) GPUs or more. Such strong scaling has not been previously achieved. In this contribution, we demonstrate that using a multi-dimensional parallelization strategy and a domain-decomposed preconditioner allows us to scale into this regime. We present results for two popular discretizations of the Dirac operator, Wilson-clover and improved staggered, employing up to 256 GPUs on the Edge cluster at Lawrence Livermore National Laboratory.
Ronald Babich, Michael A. Clark, Bálint Joó, Guochun Shi, Richard C. Brower, Steven A. Gottlieb
SC6
1995 Lattice QCD on the IBM Scalable POWERParallel Systems SP2
abstract
A 512 node IBM Scalable POWERParallel Systems SP2 was installed at the Cornell Theory Center in October 1994. During the past couple of months we have been porting and optimizing code for carrying out lattice QCD calculations. Present performance is far from ideal, however, and optimization efforts are still under way. The rate limiting step in our code involves a rather generic inversion of a large, sparse system, based on a partial differential equation in a multidimensional space. The insights we have gained so far may be useful in diagnosing performance in a wide class of applications.
Claude Bernard, Carleton DeTar, Steven A. Gottlieb, Urs M. Heller, James E. Hetrick, Naruhito Ishizuka, Leo Kärkkäinen, Steven R. Lantz, Kari Rummukainen, Robert L. Sugar, Doug Toussaint, Matthew Wingate
SC3