EDBT 2026 Demo / reviewers in the wild / expert
Gangwon Jo
dblp:11/10954
· DBLP profile ↗
10ranked-venue papers
3as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorSecurity and privacy · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
GPUs and heterogeneous computing · 42% High-performance computing · 20% Parallel and multicore computing · 18% | |
| Network and information security
1 paper |
Network security · 100% |
Topics — the 16 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing › multi-GPU computing
GPU cluster |
0.5 | 1 | 2021 | SnuRHAC: A Runtime for Heterogeneous Accelerator Clusters with CUDA Unified Memory · HPDC 2021 |
Parallel and multicore computing
parallel programming runtimes |
0.5 | 1 | 2021 | SnuRHAC: A Runtime for Heterogeneous Accelerator Clusters with CUDA Unified Memory · HPDC 2021 |
High-performance computing
cluster computing |
0.4 | 3 | 2021 | A distributed OpenCL framework using redundant computation and data replication · PLDI 2016 SnuRHAC: A Runtime for Heterogeneous Accelerator Clusters with CUDA Unified Memory · HPDC 2021 OpenCL as a unified programming model for heterogeneous CPU/GPU clusters · PPoPP 2012 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.4 | 1 | 2020 | SOFF: An OpenCL High-Level Synthesis Framework for FPGAs · ISCA 2020 |
Electronic design automation
high-level synthesis |
0.4 | 1 | 2020 | SOFF: An OpenCL High-Level Synthesis Framework for FPGAs · ISCA 2020 |
GPUs and heterogeneous computing › CPU-GPU heterogeneous computing
CPU-GPU heterogeneous systems |
0.2 | 1 | 2016 | PIPSEA: A Practical IPsec Gateway on Embedded APUs · CCS 2016 |
GPUs and heterogeneous computing
heterogeneous programming models |
0.2 | 1 | 2016 | A distributed OpenCL framework using redundant computation and data replication · PLDI 2016 |
GPUs and heterogeneous computing
heterogeneous cluster computing |
0.2 | 1 | 2015 | Accelerating LINPACK with MPI-OpenCL on Clusters of Multi-GPU Nodes · IEEE Trans. Parallel Distributed Syst. 2015 |
High-performance computing › numerical linear algebra
linpack |
0.2 | 1 | 2015 | Accelerating LINPACK with MPI-OpenCL on Clusters of Multi-GPU Nodes · IEEE Trans. Parallel Distributed Syst. 2015 |
GPUs and heterogeneous computing › multi-GPU computing
multi-GPU server |
0.2 | 1 | 2015 | Accelerating LINPACK with MPI-OpenCL on Clusters of Multi-GPU Nodes · IEEE Trans. Parallel Distributed Syst. 2015 |
High-performance computing
supercomputing |
0.2 | 1 | 2015 | Accelerating LINPACK with MPI-OpenCL on Clusters of Multi-GPU Nodes · IEEE Trans. Parallel Distributed Syst. 2015 |
GPUs and heterogeneous computing › heterogeneous cluster computing
heterogeneous CPU-GPU cluster |
0.1 | 1 | 2012 | OpenCL as a unified programming model for heterogeneous CPU/GPU clusters · PPoPP 2012 |
Parallel and multicore computing › parallel programming models
unified programming model |
0.1 | 1 | 2012 | OpenCL as a unified programming model for heterogeneous CPU/GPU clusters · PPoPP 2012 |
Program analysis
static analysis |
0.1 | 1 | 2017 | POSTER: MAPA: An Automatic Memory Access Pattern Analyzer for GPU Applications · PPoPP 2017 |
Parallel and multicore computing
MPI |
0.1 | 1 | 2015 | Accelerating LINPACK with MPI-OpenCL on Clusters of Multi-GPU Nodes · IEEE Trans. Parallel Distributed Syst. 2015 |
Parallel and multicore computing
parallel programming models |
0.1 | 1 | 2015 | Accelerating LINPACK with MPI-OpenCL on Clusters of Multi-GPU Nodes · IEEE Trans. Parallel Distributed Syst. 2015 |
Methods — techniques the papers use, named apart from their topics
OpenCL · 0.7symbolic analysis · 0.6runtime pattern selection · 0.6static prefetching · 0.5page fault handling · 0.5dynamic prefetching · 0.5pipelining · 0.4memory subsystem synthesis · 0.4high-level synthesis · 0.4zero-copy packet i/o · 0.2data replication · 0.2DPDK · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | SnuRHAC: A Runtime for Heterogeneous Accelerator Clusters with CUDA Unified MemoryabstractThis paper proposes a framework called SnuRHAC, which provides an illusion of a single GPU for the multiple GPUs in a cluster. Under SnuRHAC, a CUDA program designed to use a single GPU can utilize multiple GPUs in a cluster without any source code modification. SnuRHAC automatically distributes workload to multiple GPUs in a cluster and manages data across the nodes. To manage data efficiently, SnuRHAC extends CUDA Unified Memory and exploits its page fault mechanism. We also propose two prefetching techniques to fully exploit UM and to maximize performance. Static prefetching allows SnuRHAC to prefetch data by statically analyzing CUDA kernels. Dynamic prefetching complements static prefetching. SnuRHAC enforces an application to run on a single GPU if it is not suitable for multiple GPUs. We evaluate the performance of SnuRHAC using 18 benchmark applications from various sources. The evaluation result shows that while SnuRHAC significantly improves ease-of-programming, it shows scalable performance for the cluster environment depending on the application characteristics. Daeyoung Park, Gangwon Jo, Jungho Park, Jaejin Lee |
HPDC | 3 |
| 2020 | SOFF: An OpenCL High-Level Synthesis Framework for FPGAsabstractRecently, OpenCL has been emerging as a programming model for energy-efficient FPGA accelerators. However, the state-of-the-art OpenCL frameworks for FPGAs suffer from poor performance and usability. This paper proposes a high-level synthesis framework of OpenCL for FPGAs, called SOFF. It automatically synthesizes a datapath to execute many OpenCL kernel threads in a pipelined manner. It also synthesizes an efficient memory subsystem for the datapath based on the characteristics of OpenCL kernels. Unlike previous high-level synthesis techniques, we propose a formal way to handle variable latency instructions, complex control flows, OpenCL barriers, and atomic operations that appear in real-world OpenCL kernels. SOFF is the first OpenCL framework that correctly compiles and executes all applications in the SPEC ACCEL benchmark suite except three applications that require more FPGA resources than are available. In addition, SOFF achieves the speedup of 1.33 over Intel FPGA SDK for OpenCL without any explicit user annotation or source code modification. Gangwon Jo, Heehoon Kim, Jeesoo Lee, Jaejin Lee |
ISCA | 1 |
| 2018 | Accelerated Code Generator for Processing Ocean Color Remote Sensing Data on GpuabstractAs the satellite data gradually increases, the processing time of the algorithm for remote sensing also increases. It is now essential to make efforts to improve the processing performance in various accelerators such as Graphics Processing Unit (GPU). However, it is very difficult to use programming models that make programs to run on accelerators. Especially for scientists, easy programming and performance are often more important than providing many optimization functions and directives. To meet these requirements, we developed the accelerated code generator based on Open Computing Language (OpenCL) for processing ocean color remote sensing data. We easily applied our generator to atmospheric correction program for GOCI data as a case study and the program's performance achieved about 29.2x that is as good as the hand-written OpenCL program. We look forward to many scientists that want to find a tool mentioned above taking advantage of our generator. Jae-Moo Heo, Gangwon Jo, Hee-Jeong Han, Hyun Yang |
IGARSS | 2 |
| 2017 | POSTER: MAPA: An Automatic Memory Access Pattern Analyzer for GPU ApplicationsabstractVarious existing optimization and memory consistency management techniques for GPU applications rely on memory access patterns of kernels. However, they suffer from poor practicality because they require explicit user interventions to extract kernel memory access patterns. This paper proposes an automatic memory-access-pattern analysis framework called MAPA. MAPA is based on a source-level analysis technique derived from traditional symbolic analyses and a run-time pattern selection technique. The experimental results show that MAPA properly analyzes 116 real-world OpenCL kernels from Rodinia and Parboil. Gangwon Jo, Jaejin Lee |
PPoPP | 1 |
| 2016 | PIPSEA: A Practical IPsec Gateway on Embedded APUsabstractAccelerated Processing Unit (APU) is a heterogeneous multicore processor that contains general-purpose CPU cores and a GPU in a single chip. It also supports Heterogeneous System Architecture (HSA) that provides coherent physically-shared memory between the CPU and the GPU. In this paper, we present the design and implementation of a high-performance IPsec gateway using a low-cost commodity embedded APU. The HSA supported by the APUs eliminates the data copy overhead between the CPU and the GPU, which is unavoidable in the previous discrete GPU approaches. The gateway is implemented in OpenCL to exploit the GPU and uses zero-copy packet I/O APIs in DPDK. The IPsec gateway handles the real-world network traffic where each packet has a different workload. The proposed packet scheduling algorithm significantly improves GPU utilization for such traffic. It works not only for APUs but also for discrete GPUs. With three CPU cores and one GPU in the APU, the IPsec gateway achieves a throughput of 10.36 Gbps with an average latency of 2.79 ms to perform AES-CBC+HMAC-SHA1 for incoming packets of 1024 bytes. Jung-Ho Park, Wookeun Jung, Gangwon Jo, Ilkoo Lee, Jaejin Lee |
CCS | 3 |
| 2016 | A distributed OpenCL framework using redundant computation and data replicationabstractApplications written solely in OpenCL or CUDA cannot execute on a cluster as a whole. Most previous approaches that extend these programming models to clusters are based on a common idea: designating a centralized host node and coordinating the other nodes with the host for computation. However, the centralized host node is a serious performance bottleneck when the number of nodes is large. In this paper, we propose a scalable and distributed OpenCL framework called SnuCL-D for large-scale clusters. SnuCL-D's remote device virtualization provides an OpenCL application with an illusion that all compute devices in a cluster are confined in a single node. To reduce the amount of control-message and data communication between nodes, SnuCL-D replicates the OpenCL host program execution and data in each node. We also propose a new OpenCL host API function and a queueing optimization technique that significantly reduce the overhead incurred by the previous centralized approaches. To show the effectiveness of SnuCL-D, we evaluate SnuCL-D with a microbenchmark and eleven benchmark applications on a large-scale CPU cluster and a medium-scale GPU cluster. Gangwon Jo, Jaejin Lee |
PLDI | 2 |
| 2015 | Accelerating LINPACK with MPI-OpenCL on Clusters of Multi-GPU NodesabstractOpenCL is an open standard to write parallel applications for heterogeneous computing systems. Since its usage is restricted to a single operating system instance, programmers need to use a mix of OpenCL and MPI to program a heterogeneous cluster. In this paper, we introduce an MPI-OpenCL implementation of the LINPACK benchmark for a cluster with multi-GPU nodes. The LINPACK benchmark is one of the most widely used benchmark applications for evaluating high performance computing systems. Our implementation is based on High Performance LINPACK (HPL) and uses the blocked LU decomposition algorithm. We address that optimizations aimed at reducing the overhead of CPUs are necessary to overcome the performance gap between the CPUs and the multiple GPUs. Our LINPACK implementation achieves 93.69 Tflops (46 percent of the theoretical peak) on the target cluster with 49 nodes, each node containing two eight-core CPUs and four GPUs. Gangwon Jo, Jeongho Nah, Jaejin Lee |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2013 | Automatic OpenCL work-group size selection for multicore CPUsabstractIn this paper, we address the effect of the work-group size on the performance of OpenCL kernels. We propose a profiling-based algorithm that finds a good work-group size, in terms of performance, for the target multicore CPU architecture. Our algorithm reduces misses in the private L1 data cache and achieves load balancing between cores. It exploits the polyhedral model to estimate the working-set size and the number of cache misses for a parameterized work-group size of the OpenCL kernel. Based on the profiling information, it heuristically searches the space of parameterized work-group sizes. Our virtually-extended index space helps to increase the probability to find a better work-group size. We implement our work-group size selection algorithm as a development tool that consists of a code generator and a search library. The code generator extracts the polytope of each memory reference from the kernel code and generates a function that simplifies polytopes using the run-time information and invokes search library routines. The search library calculates the working-set size using the polytopes and finds a proper work-group size. We evaluate our approach using 31 OpenCL kernels on four different multicore CPUs. We compare its accuracy and search time to those of an exhaustive search method. Experimental results show that our tool is, on average, 1566 times faster than the exhaustive search and selects a work-group size whose performance is the same as or comparable to that of the exhaustive search. Gangwon Jo, Jaejin Lee |
PACT | 3 |
| 2012 | SnuCL: an OpenCL framework for heterogeneous CPU/GPU clustersabstractIn this paper, we propose SnuCL, an OpenCL framework for heterogeneous CPU/GPU clusters. We show that the original OpenCL semantics naturally fits to the heterogeneous cluster programming environment, and the framework achieves high performance and ease of programming. The target cluster architecture consists of a designated, single host node and many compute nodes. They are connected by an interconnection network, such as Gigabit Ethernet and InfiniBand switches. Each compute node is equipped with multicore CPUs and multiple GPUs. A set of CPU cores or each GPU becomes an OpenCL compute device. The host node executes the host program in an OpenCL application. SnuCL provides a system image running a single operating system instance for heterogeneous CPU/GPU clusters to the user. It allows the application to utilize compute devices in a compute node as if they were in the host node. No communication API, such as the MPI library, is required in the application source. SnuCL also provides collective communication extensions to OpenCL to facilitate manipulating memory objects. With SnuCL, an OpenCL application becomes portable not only between heterogeneous devices in a single node, but also between compute devices in the cluster environment. We implement SnuCL and evaluate its performance using eleven OpenCL benchmark applications. Jeongho Nah, Gangwon Jo, Jaejin Lee |
ICS | 5 |
| 2012 | OpenCL as a unified programming model for heterogeneous CPU/GPU clustersabstractIn this paper, we propose an OpenCL framework for heterogeneous CPU/GPU clusters, and show that the framework achieves both high performance and ease of programming. The framework provides an illusion of a single system for the user. It allows the application to utilize multiple heterogeneous compute devices, such as multicore CPUs and GPUs, in a remote node as if they were in a local node. No communication API, such as the MPI library, is required in the application source. We implement the OpenCL framework and evaluate its performance on a heterogeneous CPU/GPU cluster that consists of one host node and nine compute nodes using eleven OpenCL benchmark applications. Jeongho Nah, Gangwon Jo, Jaejin Lee |
PPoPP | 5 |