Jeongho Nah

dblp:91/10957 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
GPUs and heterogeneous computing · 44% High-performance computing · 36% Parallel and multicore computing · 21%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
heterogeneous cluster computing
0.212015
Accelerating LINPACK with MPI-OpenCL on Clusters of Multi-GPU Nodes · IEEE Trans. Parallel Distributed Syst. 2015
High-performance computing › numerical linear algebra
linpack
0.212015
Accelerating LINPACK with MPI-OpenCL on Clusters of Multi-GPU Nodes · IEEE Trans. Parallel Distributed Syst. 2015
GPUs and heterogeneous computing › multi-GPU computing
multi-GPU server
0.212015
Accelerating LINPACK with MPI-OpenCL on Clusters of Multi-GPU Nodes · IEEE Trans. Parallel Distributed Syst. 2015
High-performance computing
supercomputing
0.212015
Accelerating LINPACK with MPI-OpenCL on Clusters of Multi-GPU Nodes · IEEE Trans. Parallel Distributed Syst. 2015
GPUs and heterogeneous computing › heterogeneous cluster computing
heterogeneous CPU-GPU cluster
0.112012
OpenCL as a unified programming model for heterogeneous CPU/GPU clusters · PPoPP 2012
Parallel and multicore computing › parallel programming models
unified programming model
0.112012
OpenCL as a unified programming model for heterogeneous CPU/GPU clusters · PPoPP 2012
Parallel and multicore computing
MPI
0.112015
Accelerating LINPACK with MPI-OpenCL on Clusters of Multi-GPU Nodes · IEEE Trans. Parallel Distributed Syst. 2015
Parallel and multicore computing
parallel programming models
0.112015
Accelerating LINPACK with MPI-OpenCL on Clusters of Multi-GPU Nodes · IEEE Trans. Parallel Distributed Syst. 2015
High-performance computing
cluster computing
0.012012
OpenCL as a unified programming model for heterogeneous CPU/GPU clusters · PPoPP 2012

Methods — techniques the papers use, named apart from their topics

blocked LU decomposition · 0.2OpenCL · 0.2source-to-source translation · 0.1OpenCL runtime · 0.1
YearPublicationVenuePosition
2015 Accelerating LINPACK with MPI-OpenCL on Clusters of Multi-GPU Nodes
abstract
OpenCL is an open standard to write parallel applications for heterogeneous computing systems. Since its usage is restricted to a single operating system instance, programmers need to use a mix of OpenCL and MPI to program a heterogeneous cluster. In this paper, we introduce an MPI-OpenCL implementation of the LINPACK benchmark for a cluster with multi-GPU nodes. The LINPACK benchmark is one of the most widely used benchmark applications for evaluating high performance computing systems. Our implementation is based on High Performance LINPACK (HPL) and uses the blocked LU decomposition algorithm. We address that optimizations aimed at reducing the overhead of CPUs are necessary to overcome the performance gap between the CPUs and the multiple GPUs. Our LINPACK implementation achieves 93.69 Tflops (46 percent of the theoretical peak) on the target cluster with 49 nodes, each node containing two eight-core CPUs and four GPUs.
Gangwon Jo, Jeongho Nah, Jaejin Lee
IEEE Trans. Parallel Distributed Syst.2
2013 An OpenCL optimizing compiler for reconfigurable processors
abstract
This paper presents simple and efficient optimization techniques for an OpenCL compiler that targets reconfigurable processors. The target architecture consists of a generalpurpose processor core and an embedded reconfigurable accelerator with vector units. The accelerator is able to switch its architecture between the VLIW mode and the Coarse Grained Reconfigurable Array (CGRA) mode to achieve high performance. One big problem of this architecture is programming difficulty and OpenCL can be a good solution. However, since OpenCL does not guarantee performance portability, hardware dependent optimization is still necessary. Hence, we develop an OpenCL compiler framework that exploits the mode switching capability and vector units. To measure the effectiveness of the techniques, we have implemented the OpenCL framework and evaluate their performance with fourteen OpenCL benchmark applications.
Jeongho Nah, Hongjune Kim, Seok Joong Hwang, Donghoon Yoo, Jaejin Lee
FPT1
2012 SnuCL: an OpenCL framework for heterogeneous CPU/GPU clusters
abstract
In this paper, we propose SnuCL, an OpenCL framework for heterogeneous CPU/GPU clusters. We show that the original OpenCL semantics naturally fits to the heterogeneous cluster programming environment, and the framework achieves high performance and ease of programming. The target cluster architecture consists of a designated, single host node and many compute nodes. They are connected by an interconnection network, such as Gigabit Ethernet and InfiniBand switches. Each compute node is equipped with multicore CPUs and multiple GPUs. A set of CPU cores or each GPU becomes an OpenCL compute device. The host node executes the host program in an OpenCL application. SnuCL provides a system image running a single operating system instance for heterogeneous CPU/GPU clusters to the user. It allows the application to utilize compute devices in a compute node as if they were in the host node. No communication API, such as the MPI library, is required in the application source. SnuCL also provides collective communication extensions to OpenCL to facilitate manipulating memory objects. With SnuCL, an OpenCL application becomes portable not only between heterogeneous devices in a single node, but also between compute devices in the cluster environment. We implement SnuCL and evaluate its performance using eleven OpenCL benchmark applications.
Jeongho Nah, Gangwon Jo, Jaejin Lee
ICS4
2012 OpenCL as a unified programming model for heterogeneous CPU/GPU clusters
abstract
In this paper, we propose an OpenCL framework for heterogeneous CPU/GPU clusters, and show that the framework achieves both high performance and ease of programming. The framework provides an illusion of a single system for the user. It allows the application to utilize multiple heterogeneous compute devices, such as multicore CPUs and GPUs, in a remote node as if they were in a local node. No communication API, such as the MPI library, is required in the application source. We implement the OpenCL framework and evaluate its performance on a heterogeneous CPU/GPU cluster that consists of one host node and nine compute nodes using eleven OpenCL benchmark applications.
Jeongho Nah, Gangwon Jo, Jaejin Lee
PPoPP4