Everett H. Phillips

dblp:45/9082 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 74% GPUs and heterogeneous computing · 20% Performance modeling and evaluation · 6%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Environmental and earth informatics · 100%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
large-scale training
0.312018
Exascale deep learning for climate analytics · SC 2018
High-performance computing › numerical linear algebra
dense matrix multiplication
0.112011
Fast implementation of DGEMM on Fermi GPU · SC 2011
GPUs and heterogeneous computing
GPU computing
0.112011
Fast implementation of DGEMM on Fermi GPU · SC 2011
Environmental and earth informatics › climate science
climate data analysis
0.112018
Exascale deep learning for climate analytics · SC 2018

Methods — techniques the papers use, named apart from their topics

distributed training · 0.7deep learning · 0.7vector memory operations · 0.1software pipelining · 0.1instruction scheduling · 0.1
YearPublicationVenuePosition
2018 Exascale deep learning for climate analytics
Thorsten Kurth, Sean Treichler, Joshua Romero, Mayur Mudigonda, Nathan Luehr, Everett H. Phillips, Ankur Mahesh, Michael A. Matheson, Jack Deslippe, Massimiliano Fatica, Prabhat, Michael Houston
SC6
2011 Fast implementation of DGEMM on Fermi GPU
abstract
In this paper we present a thorough experience on tuning double-precision matrix-matrix multiplication (DGEM-M) on the Fermi GPU architecture. We choose an optimal algorithm with blocking in both shared memory and registers to satisfy the constraints of the Fermi memory hierarchy. Our optimization strategy is further guided by a performance modeling based on micro-architecture benchmarks. Our optimizations include software pipelining, use of vector memory operations, and instruction scheduling. Our best CUDA algorithm achieves comparable performance with the latest CUBLAS library. We further improve upon this with an implementation in the native machine language, leading to 20% increase in performance. That is, the achieved peak performance (efficiency) is improved from 302Gflop/s (58%) to 362Gflop/s (70%).
Guangming Tan, Linchuan Li, Sean Triechle, Everett H. Phillips, Yungang Bao, Ninghui Sun
SC4
2010 Implementing the Himeno benchmark with CUDA on GPU clusters
abstract
This paper describes the use of CUDA to accelerate the Himeno benchmark on clusters with GPUs. The implementation is designed to optimize memory bandwidth utilization. Our approach achieves over 83% of the theoretical peak bandwidth on a NVIDIA Tesla C1060 GPU and performs at over 50 GFlops. A multi-GPU implementation that utilizes MPI alongside CUDA streams to overlap GPU execution with data transfers allows linear scaling and performs at over 800 GFlops on a cluster with 16 GPUs. The paper presents the optimizations required to achieve this level of performance.
Everett H. Phillips, Massimiliano Fatica
IPDPS1