Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Guochun Shi

dblp:64/4279 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
0since 2021 · last 2012
0000-0002-2160-0449ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 56% GPUs and heterogeneous computing · 22% Parallel and multicore computing · 22%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU-accelerated scientific computing
0.112011
Scaling lattice QCD beyond 100 GPUs · SC 2011
High-performance computing › scientific computing systems
lattice quantum chromodynamics
0.112011
Scaling lattice QCD beyond 100 GPUs · SC 2011
Parallel and multicore computing › parallelization strategies
multi-dimensional parallelism
0.112011
Scaling lattice QCD beyond 100 GPUs · SC 2011
High-performance computing
scientific computing systems
0.112011
Scaling lattice QCD beyond 100 GPUs · SC 2011
High-performance computing › numerical linear algebra › preconditioner
domain decomposition preconditioner
0.012011
Scaling lattice QCD beyond 100 GPUs · SC 2011
High-performance computing
performance optimization at scale
0.012011
Scaling lattice QCD beyond 100 GPUs · SC 2011

Methods — techniques the papers use, named apart from their topics

multi-dimensional parallelization · 0.1domain-decomposed preconditioner · 0.1
YearPublicationVenuePosition
2012 Fine-grain parallelism using multi-core, Cell/BE, and GPU Systems
Frederico Pratas, Pedro Trancoso, Leonel Sousa, Alexandros Stamatakis, Guochun Shi, Volodymyr V. Kindratenko
Parallel Comput.5
2011 GPU acceleration of an image characterization algorithm for document similarity analysis
abstract
This paper aims to provide GPU acceleration of an decision support for selecting software and hardware architecture for content-based document comparison. We evaluate Java, C, CUDA C and OpenCL implementations of an image characterization algorithm used for content-based document comparison on a CPU and NVIDIA and AMD graphics processing units (GPUs). Based on our experimental results, we conclude that the original Java implementation of the image characterization algorithm running on a CPU-based architecture can be accelerated by a factor of 6 if the Java code is re-implemented in C, or by a factor of almost 16 if the Java code is re-implemented in CUDA C and run on NVIDIA GTX 480 GPU hardware. We also provide a power efficiency analysis.
Guochun Shi, Volodymyr V. Kindratenko, Rob Kooper, Peter Bajcsy
AICCSA1
2011 Design of MILC Lattice QCD Application for GPU Clusters
abstract
We present an implementation of the improved staggered quark action lattice QCD computation designed for execution on a GPU cluster. The parallelization strategy is based on dividing the space-time lattice along the time dimension and distributing the sub-lattices among the GPU cluster nodes. We provide a mixed-precision floating-point GPU implementation of the multi-mass conjugate gradient solver. Our single GPU implementation of the conjugate gradient solver achieves a 9x performance improvement over the highly optimized code executed on a state-of-the-art eight-core CPU node. The overall application executes almost six times faster on a GPU-enabled cluster vs. a conventional multi-core cluster. The developed code is currently used for running production QCD calculations with electromagnetic corrections.
Guochun Shi, Steven A. Gottlieb, Aaron Torok, Volodymyr V. Kindratenko
IPDPS1
2011 Scaling lattice QCD beyond 100 GPUs
abstract
Over the past five years, graphics processing units (GPUs) have had a transformational effect on numerical lattice quantum chromodynamics (LQCD) calculations in nuclear and particle physics. While GPUs have been applied with great success to the post-Monte Carlo "analysis" phase which accounts for a substantial fraction of the workload in a typical LQCD calculation, the initial Monte Carlo "gauge field generation" phase requires capability-level supercomputing, corresponding to O(100) GPUs or more. Such strong scaling has not been previously achieved. In this contribution, we demonstrate that using a multi-dimensional parallelization strategy and a domain-decomposed preconditioner allows us to scale into this regime. We present results for two popular discretizations of the Dirac operator, Wilson-clover and improved staggered, employing up to 256 GPUs on the Edge cluster at Lawrence Livermore National Laboratory.
Ronald Babich, Michael A. Clark, Bálint Joó, Guochun Shi, Richard C. Brower, Steven A. Gottlieb
SC4
2010 Direct self-consistent field computations on GPU clusters
abstract
We present an implementation of one of the direct self-consistent-field (DSCF) calculation techniques, the restricted Hartree-Fock method, on a high-performance computing cluster outfitted with graphics processing units (GPUs) and demonstrate its effectiveness and scalability up to 128 cluster nodes on molecules of as many as 1,732 atoms. We discuss the overall parallel application architecture that relies on message passing interface for distributing workload among GPU cluster nodes and POSIX threads to manage the use of GPUs internal to each node. This approach of combining coarse and fine-grain parallelism on a distributed memory system allows to perform DSCF calculations on molecules that up until now have been unattainable due to the excessive computational requirements.
Guochun Shi, Volodymyr V. Kindratenko, Ivan S. Ufimtsev, Todd J. Martínez
IPDPS1
2009 GPU clusters for high-performance computing
abstract
Large-scale GPU clusters are gaining popularity in the scientific computing community. However, their deployment and production use are associated with a number of new challenges. In this paper, we present our efforts to address some of the challenges with building and running GPU clusters in HPC environments. We touch upon such issues as balanced cluster architecture, resource sharing in a cluster environment, programming models, and applications for GPU clusters.
Volodymyr V. Kindratenko, Jeremy Enos, Guochun Shi, Michael T. Showerman, Galen Wesley Arnold, John E. Stone, James C. Phillips, Wen-Mei W. Hwu
CLUSTER3
2008 Implementation of NAMD molecular dynamics non-bonded force-field on the cell broadband engine processor
abstract
We present results of porting an important kernel of a production molecular dynamics simulation program, NAMD, to the Cell/B.E. processor. The non-bonded force-field kernel, as implemented in the NAMD SPEC 2006 CPU benchmark, has been implemented. Both single-precision and double-precision floating-point kernel variations are considered, and performance results obtained on the Cell/B.E., as well as several other platforms, are reported. Our results obtained on a 3.2 GHz Cell/B.E. blade show linear speedups when using multiple synergistic processing elements.
Guochun Shi, Volodymyr V. Kindratenko
IPDPS1