EDBT 2026 Demo / reviewers in the wild / expert
Steven A. Gottlieb
dblp:46/7222
· DBLP profile ↗
3ranked-venue papers
0as first author
0since 2021 · last 2011
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
High-performance computing · 60% GPUs and heterogeneous computing · 20% Parallel and multicore computing · 20% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing › scientific computing systems
lattice quantum chromodynamics |
0.1 | 2 | 2011 | Scaling lattice QCD beyond 100 GPUs · SC 2011 Lattice QCD on the IBM Scalable POWERParallel Systems SP2 · SC 1995 |
High-performance computing
scientific computing systems |
0.1 | 2 | 2011 | Scaling lattice QCD beyond 100 GPUs · SC 2011 Lattice QCD on the IBM Scalable POWERParallel Systems SP2 · SC 1995 |
GPUs and heterogeneous computing
GPU-accelerated scientific computing |
0.1 | 1 | 2011 | Scaling lattice QCD beyond 100 GPUs · SC 2011 |
Parallel and multicore computing › parallelization strategies
multi-dimensional parallelism |
0.1 | 1 | 2011 | Scaling lattice QCD beyond 100 GPUs · SC 2011 |
High-performance computing
performance optimization at scale |
0.1 | 2 | 2011 | Scaling lattice QCD beyond 100 GPUs · SC 2011 Lattice QCD on the IBM Scalable POWERParallel Systems SP2 · SC 1995 |
High-performance computing › numerical linear algebra › preconditioner
domain decomposition preconditioner |
0.0 | 1 | 2011 | Scaling lattice QCD beyond 100 GPUs · SC 2011 |
High-performance computing
code optimization |
0.0 | 1 | 1995 | Lattice QCD on the IBM Scalable POWERParallel Systems SP2 · SC 1995 |
Methods — techniques the papers use, named apart from their topics
multi-dimensional parallelization · 0.1domain-decomposed preconditioner · 0.1sparse system inversion · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2011 | Design of MILC Lattice QCD Application for GPU ClustersabstractWe present an implementation of the improved staggered quark action lattice QCD computation designed for execution on a GPU cluster. The parallelization strategy is based on dividing the space-time lattice along the time dimension and distributing the sub-lattices among the GPU cluster nodes. We provide a mixed-precision floating-point GPU implementation of the multi-mass conjugate gradient solver. Our single GPU implementation of the conjugate gradient solver achieves a 9x performance improvement over the highly optimized code executed on a state-of-the-art eight-core CPU node. The overall application executes almost six times faster on a GPU-enabled cluster vs. a conventional multi-core cluster. The developed code is currently used for running production QCD calculations with electromagnetic corrections. Guochun Shi, Steven A. Gottlieb, Aaron Torok, Volodymyr V. Kindratenko |
IPDPS | 2 |
| 2011 | Scaling lattice QCD beyond 100 GPUsabstractOver the past five years, graphics processing units (GPUs) have had a transformational effect on numerical lattice quantum chromodynamics (LQCD) calculations in nuclear and particle physics. While GPUs have been applied with great success to the post-Monte Carlo "analysis" phase which accounts for a substantial fraction of the workload in a typical LQCD calculation, the initial Monte Carlo "gauge field generation" phase requires capability-level supercomputing, corresponding to O(100) GPUs or more. Such strong scaling has not been previously achieved. In this contribution, we demonstrate that using a multi-dimensional parallelization strategy and a domain-decomposed preconditioner allows us to scale into this regime. We present results for two popular discretizations of the Dirac operator, Wilson-clover and improved staggered, employing up to 256 GPUs on the Edge cluster at Lawrence Livermore National Laboratory. Ronald Babich, Michael A. Clark, Bálint Joó, Guochun Shi, Richard C. Brower, Steven A. Gottlieb |
SC | 6 |
| 1995 | Lattice QCD on the IBM Scalable POWERParallel Systems SP2abstractA 512 node IBM Scalable POWERParallel Systems SP2 was installed at the Cornell Theory Center in October 1994. During the past couple of months we have been porting and optimizing code for carrying out lattice QCD calculations. Present performance is far from ideal, however, and optimization efforts are still under way. The rate limiting step in our code involves a rather generic inversion of a large, sparse system, based on a partial differential equation in a multidimensional space. The insights we have gained so far may be useful in diagnosing performance in a wide class of applications. Claude Bernard, Carleton DeTar, Steven A. Gottlieb, Urs M. Heller, James E. Hetrick, Naruhito Ishizuka, Leo Kärkkäinen, Steven R. Lantz, Kari Rummukainen, Robert L. Sugar, Doug Toussaint, Matthew Wingate |
SC | 3 |