EDBT 2026 Demo / reviewers in the wild / expert
Putt Sakdhnagool
dblp:49/11416
· DBLP profile ↗
8ranked-venue papers
2as first author
0since 2021 · last 2020
0000-0002-7925-0525ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
GPUs and heterogeneous computing · 38% Parallel and multicore computing · 24% High-performance computing · 23% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization › accelerator compilation
GPU compiler optimization |
0.4 | 1 | 2019 | Optimizing GPU programs by register demotion: poster · PPoPP 2019 |
GPUs and heterogeneous computing
GPU resource management |
0.4 | 1 | 2019 | Optimizing GPU programs by register demotion: poster · PPoPP 2019 |
GPUs and heterogeneous computing › GPU resource management
GPU virtualization |
0.3 | 1 | 2017 | Pagoda: Fine-Grained GPU Resource Virtualization for Narrow Tasks · PPoPP 2017 |
Parallel and multicore computing
parallel computing |
0.3 | 1 | 2017 | Massively parallel 3D image reconstruction · SC 2017 |
Cloud and datacenter computing
resource management |
0.3 | 1 | 2017 | Pagoda: Fine-Grained GPU Resource Virtualization for Narrow Tasks · PPoPP 2017 |
High-performance computing
scientific computing systems |
0.3 | 1 | 2017 | Massively parallel 3D image reconstruction · SC 2017 |
High-performance computing › scientific computing systems
X-ray CT reconstruction |
0.3 | 1 | 2017 | Massively parallel 3D image reconstruction · SC 2017 |
GPUs and heterogeneous computing
heterogeneous cluster computing |
0.2 | 1 | 2015 | HeteroDoop: A MapReduce Programming System for Accelerator Clusters · HPDC 2015 |
Parallel and multicore computing › data-parallel programming
mapreduce |
0.2 | 1 | 2015 | HeteroDoop: A MapReduce Programming System for Accelerator Clusters · HPDC 2015 |
Memory systems
shared memory |
0.1 | 1 | 2019 | Optimizing GPU programs by register demotion: poster · PPoPP 2019 |
Image and video processing
super-resolution |
0.1 | 1 | 2017 | Massively parallel 3D image reconstruction · SC 2017 |
Parallel and multicore computing
task scheduling |
0.1 | 1 | 2017 | Pagoda: Fine-Grained GPU Resource Virtualization for Narrow Tasks · PPoPP 2017 |
GPUs and heterogeneous computing
GPU computing |
0.1 | 1 | 2015 | HeteroDoop: A MapReduce Programming System for Accelerator Clusters · HPDC 2015 |
Methods — techniques the papers use, named apart from their topics
compiler transformation · 0.8assembly code transformation · 0.8model-based iterative reconstruction · 0.6GPU virtualization · 0.3tail scheduling · 0.2runtime system · 0.2optimizing compiler · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Quantum Dynamics at Scale: Ultrafast Control of Emergent Functional MaterialsabstractConfluence of extreme-scale quantum dynamics simulations (i.e. [email protected]) and cutting-edge x-ray free-electron laser experiments are revolutionizing materials science. An archetypal example is the exciting concept of using picosecond light pulses to control emergent material properties on demand in atomically-thin layered materials. This paper describes efforts to scale our quantum molecular dynamics engine toward the United States' first exaflop/s computer, under an Aurora Early Science Program project named "Metascalable layered material genome". Key algorithmic and computing techniques incorporated are: (1) globally-scalable and locally-fast solvers within a linear-scaling divide-conquer-recombine algorithmic framework; (2) algebraic 'BLASification' of computational kernels; and (3) data alignment and loop restructuring, along with register and cache blocking, for enhanced vectorization and efficient memory access. The resulting weak-scaling parallel efficiency was 0.93 on 131,072 Intel Xeon Phi cores for a 56.6 million atom (or 169 million valence-electron) system, whereas the various code transformations achieved 5-fold speedup. The optimized simulation engine allowed us for the first time to establish a significant effect of substrate on the dynamics of layered material upon electronic excitation. Subodh Tiwari, Aravind Krishnamoorthy, Pankaj Rajak, Putt Sakdhnagool, Manaschai Kunaseth, Fuyuki Shimojo, Shogo Fukushima, Aiichiro Nakano, Ye Luo 0001, Rajiv K. Kalia, Ken-ichi Nomura, Priya Vashishta |
HPC Asia | 4 |
| 2019 | Optimizing GPU programs by register demotion: posterabstractGPU utilization, measured as occupancy, is limited by the parallel threads' combined usage of on-chip resources. If the resource demand cannot be met, GPUs will reduce the number of concurrent threads, impacting the program performance. We have observed that registers are the occupancy limiters while shared metmory tends to be underused. The de facto approach spills excessive registers to the out-of-chip memory, ignoring the shared memory and leaving the on-chip resources underutilized. To mitigate the register demand, our work presents a novel compiler technique, called register demotion, that allows data in the register to be placed into the underutilized shared memory by transforming the GPU assembly code (SASS). Register demotion achieves up to 18% speedup over the nvcc compiler, with a geometric mean of 7%. Putt Sakdhnagool, Amit Sabne, Rudolf Eigenmann |
PPoPP | 1 |
| 2019 | Comparative analysis of coprocessors
Putt Sakdhnagool, Amit Sabne, Rudolf Eigenmann |
Concurr. Comput. Pract. Exp. | 1 |
| 2017 | Pagoda: Fine-Grained GPU Resource Virtualization for Narrow TasksabstractMassively multithreaded GPUs achieve high throughput by running thousands of threads in parallel. To fully utilize the hardware, workloads spawn work to the GPU in bulk by launching large tasks, where each task is a kernel that contains thousands of threads that occupy the entire GPU. Tsung Tai Yeh, Amit Sabne, Putt Sakdhnagool, Rudolf Eigenmann, Timothy G. Rogers |
PPoPP | 3 |
| 2017 | Massively parallel 3D image reconstructionabstractComputed Tomographic (CT) image reconstruction is an important technique used in a wide range of applications. Among reconstruction methods, Model-Based Iterative Reconstruction (MBIR) is known to produce much higher quality CT images; however, the high computational requirements of MBIR greatly restrict their application. Currently, MBIR speed is primarily limited by irregular data access patterns, the difficulty of effective parallelization, and slow algorithmic convergence. Xiao Wang 0004, Amit Sabne, Putt Sakdhnagool, Sherman J. Kisner, Charles A. Bouman, Samuel P. Midkiff |
SC | 3 |
| 2016 | POSTER: Pagoda: A Runtime System to Maximize GPU Utilization in Data Parallel Tasks with Limited ParallelismabstractMassively multithreaded GPUs achieve high throughput by running thousands of threads in parallel. To fully utilize the hardware, contemporary workloads spawn work to the GPU in bulk by launching large tasks, where each task is a kernel that contains thousands of threads that occupy the entire GPU. Tsung Tai Yeh, Amit Sabne, Putt Sakdhnagool, Rudolf Eigenmann, Timothy G. Rogers |
PACT | 3 |
| 2015 | HeteroDoop: A MapReduce Programming System for Accelerator ClustersabstractThe deluge of data has inspired big-data processing frameworks that span across large clusters. Frameworks for MapReduce, a state-of-the-art programming model, have primarily made use of the CPUs in distributed systems, leaving out computationally powerful accelerators such as GPUs. This paper presents HeteroDoop, a MapReduce framework that employs both CPUs and GPUs in a cluster. HeteroDoop offers the following novel features: (i) a small set of directives can be placed on an existing sequential, CPU-only program, expressing MapReduce semantics; (ii) an optimizing compiler translates the directive-augmented program into a GPU code; (iii) a runtime system assists the compiler in handling MapReduce semantics on the GPU; and (iv) a tail scheduling scheme minimizes job execution time in light of disparate processing capabilities of CPUs and GPUs. This paper addresses several challenges that need to be overcome in order to support these features. HeteroDoop is built on top of the state-of-the-art, CPU-only Hadoop MapReduce framework, inheriting its functionality. Evaluation results of HeteroDoop on recent hardware indicate that usage of even a single GPU per node can improve performance by up to 2.78x, with a geometric mean of 1.6x across our benchmarks, compared to a CPU-only Hadoop, running on a cluster with 20-core CPUs. Amit Sabne, Putt Sakdhnagool, Rudolf Eigenmann |
HPDC | 2 |
| 2013 | Scaling large-data computations on multi-GPU acceleratorsabstractModern supercomputers rely on accelerators to speed up highly parallel workloads. Intricate programming models, limited device memory sizes and overheads of data transfers between CPU and accelerator memories are among the open challenges that restrict the widespread use of accelerators. First, this paper proposes a mechanism and an implementation to automatically pipeline the CPU-GPU memory channel so as to overlap the GPU computation with the memory copies, alleviating the data transfer overhead. Second, in doing so, the paper presents a technique called Computation Splitting, COSP, that caters to arbitrary device memory sizes and automatically manages to run out-of-card OpenMP-like applications on GPUs. Third, a novel adaptive runtime tuning mechanism is proposed to automatically select the pipeline stage size so as to gain the best possible performance. The mechanism adapts to the underlying hardware in the starting phase of a program and chooses the pipeline stage size. The techniques are implemented in a system that is able to translate an input OpenMP program to multiple GPUs attached to the same host CPU. Experimentation on a set of nine benchmarks shows that, on average, the pipelining scheme improves the performance by 1.49x, while limiting the runtime tuning overhead to 3% of the execution time. Amit Sabne, Putt Sakdhnagool, Rudolf Eigenmann |
ICS | 2 |