Benjamin Welton

dblp:64/10367 · DBLP profile ↗
← Back
5ranked-venue papers
5as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
GPUs and heterogeneous computing · 36% High-performance computing · 26% Performance modeling and evaluation · 25%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU computing
0.522019
Diogenes: looking for an honest CPU/GPU performance measurement tool · SC 2019
Mr. Scan: extreme scale density-based clustering using a tree-based network of GPGPU nodes · SC 2013
Performance modeling and evaluation
performance analysis tools
0.412019
Diogenes: looking for an honest CPU/GPU performance measurement tool · SC 2019
Data mining
clustering
0.212013
Mr. Scan: extreme scale density-based clustering using a tree-based network of GPGPU nodes · SC 2013
Data mining › clustering › density-based clustering
DBSCAN
0.212013
Mr. Scan: extreme scale density-based clustering using a tree-based network of GPGPU nodes · SC 2013
Data mining › clustering
density-based clustering
0.212013
Mr. Scan: extreme scale density-based clustering using a tree-based network of GPGPU nodes · SC 2013
Distributed systems › distributed algorithms
distributed clustering
0.212013
Mr. Scan: extreme scale density-based clustering using a tree-based network of GPGPU nodes · SC 2013
High-performance computing › supercomputing
extreme scale computing
0.212013
Mr. Scan: extreme scale density-based clustering using a tree-based network of GPGPU nodes · SC 2013
High-performance computing
performance optimization
0.112019
Diogenes: looking for an honest CPU/GPU performance measurement tool · SC 2019
High-performance computing
scientific computing systems
0.112019
Diogenes: looking for an honest CPU/GPU performance measurement tool · SC 2019

Methods — techniques the papers use, named apart from their topics

multi-stage/multi-run instrumentation · 0.4feed-forward measurement · 0.4hybrid parallel implementation · 0.3MRNet · 0.3GPGPU · 0.3
YearPublicationVenuePosition
2020 Identifying and (automatically) remedying performance problems in CPU/GPU applications
abstract
GPU accelerators have become common on today's leadership-class computing platforms. Effective exploitation of the additional parallelism offered by GPUs is fraught with challenges. A key performance challenge faced by developers is how to limit the time consumed by synchronizations between the CPU and GPU. We introduce the extended feed-forward measurement (FFM) performance tool that provides an automated detection of synchronization problems, identifies if the synchronization problem is a component of a larger construct that exhibits a problem beyond an individual synchronization operation, identifies remedies that can correct the issue, and in some cases automatically applies remedies to problems exhibited by larger constructs. The extended FFM performance tool identifies three causes of unnecessary synchronizations: a problem caused by a single operation, a problem caused by memory management issues, and a problem caused by a memory transfer. The extended FFM model prescribes remedies for each construct and can automatically apply remedies for memory management and memory transfer cause problems. We created an implementation of the extended FFM performance tool and employed it to identify and automatically correct problems in three real-world scientific applications, resulting in an automatically obtained reduction in execution time between 9% and 43%.
Benjamin Welton, Barton P. Miller
ICS1
2019 Diogenes: looking for an honest CPU/GPU performance measurement tool
abstract
GPU accelerators have become common on today's leadership-class computing platforms. Exploiting the additional parallelism offered by GPUs is fraught with challenges. A key performance challenge faced by developers is how to limit the time consumed by synchronization and memory transfers between the CPU and GPU. We introduce the feed-forward measurement (FFM) performance tool model that automates the identification of unnecessary or inefficient synchronization and memory transfer, providing an estimate of potential benefit if the problem were fixed. FFM uses a new multi-stage/multi-run instrumentation model that adjusts instrumentation based application behavior on prior runs, guiding FFM to problematic GPU operations that were previously unknown. The collected data feeds a new analysis model that gives an accurate estimate of potential benefit of fixing the problem. We created an implementation of FFM called Diogenes that we have used to identify problems in four real-world scientific applications.
Benjamin Welton, Barton P. Miller
SC1
2018 Exposing Hidden Performance Opportunities in High Performance GPU Applications
abstract
Leadership class systems with nodes containing many-core accelerators, such as GPUs, have the potential to increase the performance of applications. Effectively exploiting the parallelism provided by many-core accelerators requires developers to identify where accelerator parallelization would provide benefit and ensuring efficient interaction between the CPU and accelerator. In the abstract, these issues appear straightforward and well understood. However, we have found that significant untapped performance opportunities exist in these areas even in well-known, heavily optimized, real world applications created by experienced GPU developers. These untapped performance opportunities exist because accelerated libraries can create unexpected synchronization delay and memory transfer requests, interaction between accelerated libraries can cause unexpected inefficiencies when combined, and vectorization opportunities can be hidden by the structure of the program. In applications we have studied (Qball, QBox, Hoomd-blue, LAMMPs, and cuIBM), exploiting these opportunities resulted in reduction of their execution time by 18%-87%. In this work, we provide concrete evidence of the existence and impact that these performance issues have on real world applications today. We characterize the missed performance opportunities we have identified by their underlying cause and describe a preliminary design of detection methods that can be used by performance tools to identify these missed opportunities.
Benjamin Welton, Barton P. Miller
CCGrid1
2013 Mr. Scan: extreme scale density-based clustering using a tree-based network of GPGPU nodes
abstract
Density-based clustering algorithms are a widely-used class of data mining techniques that can find irregularly shaped clusters and cluster data without prior knowledge of the number of clusters it contains. DBSCAN is the most well-known density-based clustering algorithm. We introduce our version of DBSCAN, called Mr. Scan, which uses a hybrid parallel implementation that combines the MRNet tree-based distribution network with GPGPU-equipped nodes. Mr. Scan avoids the problems of existing implementations by effectively partitioning the point space and by optimizing DBSCAN's computation over dense data regions. We tested Mr. Scan on both a geolocated Twitter dataset and image data obtained from the Sloan Digital Sky Survey. At its largest scale, Mr. Scan clustered 6.5 billion points from the Twitter dataset on 8,192 GPU nodes on Cray Titan in 17.3 minutes. All other parallel DBSCAN implementations have only demonstrated the ability to cluster up to 100 million points.
Benjamin Welton, Evan Samanas, Barton P. Miller
SC1
2011 Improving I/O Forwarding Throughput with Data Compression
abstract
While network bandwidth is steadily increasing, it is doing so at a much slower rate than the corresponding increase in CPU performance. This trend has widened the gap between CPU and network speed. In this paper, we investigate improvements to I/O performance by exploiting this gap. We harness idle CPU resources to compress network traffic, reducing the amount of data transferred over the network and increasing effective network bandwidth. We created a set of compression services within the I/O Forwarding Scalability Layer. These services transparently compress and decompress data as it is transferred over the network. We studied the effect of the compression services on a variety of data sets and conducted experiments on a high-performance computing cluster.
Benjamin Welton, Dries Kimpe, Jason Cope, Christina M. Patrick, Kamil Iskra, Robert B. Ross
CLUSTER1