Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Putt Sakdhnagool

dblp:49/11416 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
0since 2021 · last 2020
0000-0002-7925-0525ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
GPUs and heterogeneous computing · 38% Parallel and multicore computing · 24% High-performance computing · 23%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization › accelerator compilation
GPU compiler optimization
0.412019
Optimizing GPU programs by register demotion: poster · PPoPP 2019
GPUs and heterogeneous computing
GPU resource management
0.412019
Optimizing GPU programs by register demotion: poster · PPoPP 2019
GPUs and heterogeneous computing › GPU resource management
GPU virtualization
0.312017
Pagoda: Fine-Grained GPU Resource Virtualization for Narrow Tasks · PPoPP 2017
Parallel and multicore computing
parallel computing
0.312017
Massively parallel 3D image reconstruction · SC 2017
Cloud and datacenter computing
resource management
0.312017
Pagoda: Fine-Grained GPU Resource Virtualization for Narrow Tasks · PPoPP 2017
High-performance computing
scientific computing systems
0.312017
Massively parallel 3D image reconstruction · SC 2017
High-performance computing › scientific computing systems
X-ray CT reconstruction
0.312017
Massively parallel 3D image reconstruction · SC 2017
GPUs and heterogeneous computing
heterogeneous cluster computing
0.212015
HeteroDoop: A MapReduce Programming System for Accelerator Clusters · HPDC 2015
Parallel and multicore computing › data-parallel programming
mapreduce
0.212015
HeteroDoop: A MapReduce Programming System for Accelerator Clusters · HPDC 2015
Memory systems
shared memory
0.112019
Optimizing GPU programs by register demotion: poster · PPoPP 2019
Image and video processing
super-resolution
0.112017
Massively parallel 3D image reconstruction · SC 2017
Parallel and multicore computing
task scheduling
0.112017
Pagoda: Fine-Grained GPU Resource Virtualization for Narrow Tasks · PPoPP 2017
GPUs and heterogeneous computing
GPU computing
0.112015
HeteroDoop: A MapReduce Programming System for Accelerator Clusters · HPDC 2015

Methods — techniques the papers use, named apart from their topics

compiler transformation · 0.8assembly code transformation · 0.8model-based iterative reconstruction · 0.6GPU virtualization · 0.3tail scheduling · 0.2runtime system · 0.2optimizing compiler · 0.2
YearPublicationVenuePosition
2020 Quantum Dynamics at Scale: Ultrafast Control of Emergent Functional Materials
abstract
Confluence of extreme-scale quantum dynamics simulations (i.e. [email protected]) and cutting-edge x-ray free-electron laser experiments are revolutionizing materials science. An archetypal example is the exciting concept of using picosecond light pulses to control emergent material properties on demand in atomically-thin layered materials. This paper describes efforts to scale our quantum molecular dynamics engine toward the United States' first exaflop/s computer, under an Aurora Early Science Program project named "Metascalable layered material genome". Key algorithmic and computing techniques incorporated are: (1) globally-scalable and locally-fast solvers within a linear-scaling divide-conquer-recombine algorithmic framework; (2) algebraic 'BLASification' of computational kernels; and (3) data alignment and loop restructuring, along with register and cache blocking, for enhanced vectorization and efficient memory access. The resulting weak-scaling parallel efficiency was 0.93 on 131,072 Intel Xeon Phi cores for a 56.6 million atom (or 169 million valence-electron) system, whereas the various code transformations achieved 5-fold speedup. The optimized simulation engine allowed us for the first time to establish a significant effect of substrate on the dynamics of layered material upon electronic excitation.
Subodh Tiwari, Aravind Krishnamoorthy, Pankaj Rajak, Putt Sakdhnagool, Manaschai Kunaseth, Fuyuki Shimojo, Shogo Fukushima, Aiichiro Nakano, Ye Luo 0001, Rajiv K. Kalia, Ken-ichi Nomura, Priya Vashishta
HPC Asia4
2019 Optimizing GPU programs by register demotion: poster
abstract
GPU utilization, measured as occupancy, is limited by the parallel threads' combined usage of on-chip resources. If the resource demand cannot be met, GPUs will reduce the number of concurrent threads, impacting the program performance. We have observed that registers are the occupancy limiters while shared metmory tends to be underused. The de facto approach spills excessive registers to the out-of-chip memory, ignoring the shared memory and leaving the on-chip resources underutilized. To mitigate the register demand, our work presents a novel compiler technique, called register demotion, that allows data in the register to be placed into the underutilized shared memory by transforming the GPU assembly code (SASS). Register demotion achieves up to 18% speedup over the nvcc compiler, with a geometric mean of 7%.
Putt Sakdhnagool, Amit Sabne, Rudolf Eigenmann
PPoPP1
2019 Comparative analysis of coprocessors
Putt Sakdhnagool, Amit Sabne, Rudolf Eigenmann
Concurr. Comput. Pract. Exp.1
2017 Pagoda: Fine-Grained GPU Resource Virtualization for Narrow Tasks
abstract
Massively multithreaded GPUs achieve high throughput by running thousands of threads in parallel. To fully utilize the hardware, workloads spawn work to the GPU in bulk by launching large tasks, where each task is a kernel that contains thousands of threads that occupy the entire GPU.
Tsung Tai Yeh, Amit Sabne, Putt Sakdhnagool, Rudolf Eigenmann, Timothy G. Rogers
PPoPP3
2017 Massively parallel 3D image reconstruction
abstract
Computed Tomographic (CT) image reconstruction is an important technique used in a wide range of applications. Among reconstruction methods, Model-Based Iterative Reconstruction (MBIR) is known to produce much higher quality CT images; however, the high computational requirements of MBIR greatly restrict their application. Currently, MBIR speed is primarily limited by irregular data access patterns, the difficulty of effective parallelization, and slow algorithmic convergence.
Xiao Wang 0004, Amit Sabne, Putt Sakdhnagool, Sherman J. Kisner, Charles A. Bouman, Samuel P. Midkiff
SC3
2016 POSTER: Pagoda: A Runtime System to Maximize GPU Utilization in Data Parallel Tasks with Limited Parallelism
abstract
Massively multithreaded GPUs achieve high throughput by running thousands of threads in parallel. To fully utilize the hardware, contemporary workloads spawn work to the GPU in bulk by launching large tasks, where each task is a kernel that contains thousands of threads that occupy the entire GPU.
Tsung Tai Yeh, Amit Sabne, Putt Sakdhnagool, Rudolf Eigenmann, Timothy G. Rogers
PACT3
2015 HeteroDoop: A MapReduce Programming System for Accelerator Clusters
abstract
The deluge of data has inspired big-data processing frameworks that span across large clusters. Frameworks for MapReduce, a state-of-the-art programming model, have primarily made use of the CPUs in distributed systems, leaving out computationally powerful accelerators such as GPUs. This paper presents HeteroDoop, a MapReduce framework that employs both CPUs and GPUs in a cluster. HeteroDoop offers the following novel features: (i) a small set of directives can be placed on an existing sequential, CPU-only program, expressing MapReduce semantics; (ii) an optimizing compiler translates the directive-augmented program into a GPU code; (iii) a runtime system assists the compiler in handling MapReduce semantics on the GPU; and (iv) a tail scheduling scheme minimizes job execution time in light of disparate processing capabilities of CPUs and GPUs. This paper addresses several challenges that need to be overcome in order to support these features. HeteroDoop is built on top of the state-of-the-art, CPU-only Hadoop MapReduce framework, inheriting its functionality. Evaluation results of HeteroDoop on recent hardware indicate that usage of even a single GPU per node can improve performance by up to 2.78x, with a geometric mean of 1.6x across our benchmarks, compared to a CPU-only Hadoop, running on a cluster with 20-core CPUs.
Amit Sabne, Putt Sakdhnagool, Rudolf Eigenmann
HPDC2
2013 Scaling large-data computations on multi-GPU accelerators
abstract
Modern supercomputers rely on accelerators to speed up highly parallel workloads. Intricate programming models, limited device memory sizes and overheads of data transfers between CPU and accelerator memories are among the open challenges that restrict the widespread use of accelerators. First, this paper proposes a mechanism and an implementation to automatically pipeline the CPU-GPU memory channel so as to overlap the GPU computation with the memory copies, alleviating the data transfer overhead. Second, in doing so, the paper presents a technique called Computation Splitting, COSP, that caters to arbitrary device memory sizes and automatically manages to run out-of-card OpenMP-like applications on GPUs. Third, a novel adaptive runtime tuning mechanism is proposed to automatically select the pipeline stage size so as to gain the best possible performance. The mechanism adapts to the underlying hardware in the starting phase of a program and chooses the pipeline stage size. The techniques are implemented in a system that is able to translate an input OpenMP program to multiple GPUs attached to the same host CPU. Experimentation on a set of nine benchmarks shows that, on average, the pipelining scheme improves the performance by 1.49x, while limiting the runtime tuning overhead to 3% of the execution time.
Amit Sabne, Putt Sakdhnagool, Rudolf Eigenmann
ICS2