EDBT 2026 Demo / reviewers in the wild / expert
Jeongho Nah
dblp:91/10957
· DBLP profile ↗
4ranked-venue papers
1as first author
0since 2021 · last 2015
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
GPUs and heterogeneous computing · 44% High-performance computing · 36% Parallel and multicore computing · 21% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing
heterogeneous cluster computing |
0.2 | 1 | 2015 | Accelerating LINPACK with MPI-OpenCL on Clusters of Multi-GPU Nodes · IEEE Trans. Parallel Distributed Syst. 2015 |
High-performance computing › numerical linear algebra
linpack |
0.2 | 1 | 2015 | Accelerating LINPACK with MPI-OpenCL on Clusters of Multi-GPU Nodes · IEEE Trans. Parallel Distributed Syst. 2015 |
GPUs and heterogeneous computing › multi-GPU computing
multi-GPU server |
0.2 | 1 | 2015 | Accelerating LINPACK with MPI-OpenCL on Clusters of Multi-GPU Nodes · IEEE Trans. Parallel Distributed Syst. 2015 |
High-performance computing
supercomputing |
0.2 | 1 | 2015 | Accelerating LINPACK with MPI-OpenCL on Clusters of Multi-GPU Nodes · IEEE Trans. Parallel Distributed Syst. 2015 |
GPUs and heterogeneous computing › heterogeneous cluster computing
heterogeneous CPU-GPU cluster |
0.1 | 1 | 2012 | OpenCL as a unified programming model for heterogeneous CPU/GPU clusters · PPoPP 2012 |
Parallel and multicore computing › parallel programming models
unified programming model |
0.1 | 1 | 2012 | OpenCL as a unified programming model for heterogeneous CPU/GPU clusters · PPoPP 2012 |
Parallel and multicore computing
MPI |
0.1 | 1 | 2015 | Accelerating LINPACK with MPI-OpenCL on Clusters of Multi-GPU Nodes · IEEE Trans. Parallel Distributed Syst. 2015 |
Parallel and multicore computing
parallel programming models |
0.1 | 1 | 2015 | Accelerating LINPACK with MPI-OpenCL on Clusters of Multi-GPU Nodes · IEEE Trans. Parallel Distributed Syst. 2015 |
High-performance computing
cluster computing |
0.0 | 1 | 2012 | OpenCL as a unified programming model for heterogeneous CPU/GPU clusters · PPoPP 2012 |
Methods — techniques the papers use, named apart from their topics
blocked LU decomposition · 0.2OpenCL · 0.2source-to-source translation · 0.1OpenCL runtime · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | Accelerating LINPACK with MPI-OpenCL on Clusters of Multi-GPU NodesabstractOpenCL is an open standard to write parallel applications for heterogeneous computing systems. Since its usage is restricted to a single operating system instance, programmers need to use a mix of OpenCL and MPI to program a heterogeneous cluster. In this paper, we introduce an MPI-OpenCL implementation of the LINPACK benchmark for a cluster with multi-GPU nodes. The LINPACK benchmark is one of the most widely used benchmark applications for evaluating high performance computing systems. Our implementation is based on High Performance LINPACK (HPL) and uses the blocked LU decomposition algorithm. We address that optimizations aimed at reducing the overhead of CPUs are necessary to overcome the performance gap between the CPUs and the multiple GPUs. Our LINPACK implementation achieves 93.69 Tflops (46 percent of the theoretical peak) on the target cluster with 49 nodes, each node containing two eight-core CPUs and four GPUs. Gangwon Jo, Jeongho Nah, Jaejin Lee |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2013 | An OpenCL optimizing compiler for reconfigurable processorsabstractThis paper presents simple and efficient optimization techniques for an OpenCL compiler that targets reconfigurable processors. The target architecture consists of a generalpurpose processor core and an embedded reconfigurable accelerator with vector units. The accelerator is able to switch its architecture between the VLIW mode and the Coarse Grained Reconfigurable Array (CGRA) mode to achieve high performance. One big problem of this architecture is programming difficulty and OpenCL can be a good solution. However, since OpenCL does not guarantee performance portability, hardware dependent optimization is still necessary. Hence, we develop an OpenCL compiler framework that exploits the mode switching capability and vector units. To measure the effectiveness of the techniques, we have implemented the OpenCL framework and evaluate their performance with fourteen OpenCL benchmark applications. Jeongho Nah, Hongjune Kim, Seok Joong Hwang, Donghoon Yoo, Jaejin Lee |
FPT | 1 |
| 2012 | SnuCL: an OpenCL framework for heterogeneous CPU/GPU clustersabstractIn this paper, we propose SnuCL, an OpenCL framework for heterogeneous CPU/GPU clusters. We show that the original OpenCL semantics naturally fits to the heterogeneous cluster programming environment, and the framework achieves high performance and ease of programming. The target cluster architecture consists of a designated, single host node and many compute nodes. They are connected by an interconnection network, such as Gigabit Ethernet and InfiniBand switches. Each compute node is equipped with multicore CPUs and multiple GPUs. A set of CPU cores or each GPU becomes an OpenCL compute device. The host node executes the host program in an OpenCL application. SnuCL provides a system image running a single operating system instance for heterogeneous CPU/GPU clusters to the user. It allows the application to utilize compute devices in a compute node as if they were in the host node. No communication API, such as the MPI library, is required in the application source. SnuCL also provides collective communication extensions to OpenCL to facilitate manipulating memory objects. With SnuCL, an OpenCL application becomes portable not only between heterogeneous devices in a single node, but also between compute devices in the cluster environment. We implement SnuCL and evaluate its performance using eleven OpenCL benchmark applications. Jeongho Nah, Gangwon Jo, Jaejin Lee |
ICS | 4 |
| 2012 | OpenCL as a unified programming model for heterogeneous CPU/GPU clustersabstractIn this paper, we propose an OpenCL framework for heterogeneous CPU/GPU clusters, and show that the framework achieves both high performance and ease of programming. The framework provides an illusion of a single system for the user. It allows the application to utilize multiple heterogeneous compute devices, such as multicore CPUs and GPUs, in a remote node as if they were in a local node. No communication API, such as the MPI library, is required in the application source. We implement the OpenCL framework and evaluate its performance on a heterogeneous CPU/GPU cluster that consists of one host node and nine compute nodes using eleven OpenCL benchmark applications. Jeongho Nah, Gangwon Jo, Jaejin Lee |
PPoPP | 4 |