Joo Hwan Lee

dblp:40/2645 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
1since 2021 · last 2021
0000-0003-3989-2109ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Performance modeling and evaluation · 23% GPUs and heterogeneous computing · 21% Storage systems · 20%
Databases, data mining, and information retrieval
1 paper
Query processing and optimization · 100%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › parallel algorithms › sorting
bitonic sort
0.512021
NASCENT: Near-Storage Acceleration of Database Sort on SmartSSD · FPGA 2021
Storage systems › computational storage
near-storage computing
0.512021
NASCENT: Near-Storage Acceleration of Database Sort on SmartSSD · FPGA 2021
GPUs and heterogeneous computing › GPU memory access
memory divergence
0.312018
StaleLearn: Learning Acceleration with Asynchronous Synchronization Between Model Replicas on PIM · IEEE Trans. Computers 2018
Memory systems
processing-in-memory
0.312018
StaleLearn: Learning Acceleration with Asynchronous Synchronization Between Model Replicas on PIM · IEEE Trans. Computers 2018
Performance modeling and evaluation › processor performance modeling › accelerator performance modeling
GPU performance modeling
0.212014
GPUMech: GPU Performance Modeling Technique Based on Interval Analysis · MICRO 2014
Performance modeling and evaluation › numerical algorithms
interval analysis
0.212014
GPUMech: GPU Performance Modeling Technique Based on Interval Analysis · MICRO 2014
Performance modeling and evaluation
simulation
0.212014
GPUMech: GPU Performance Modeling Technique Based on Interval Analysis · MICRO 2014
GPUs and heterogeneous computing
multi-GPU computing
0.112011
Achieving a single compute device image in OpenCL for multiple GPUs · PPoPP 2011
GPUs and heterogeneous computing
GPU architecture
0.112014
GPUMech: GPU Performance Modeling Technique Based on Interval Analysis · MICRO 2014
Cloud and datacenter computing › resource allocation
workload allocation
0.012011
Achieving a single compute device image in OpenCL for multiple GPUs · PPoPP 2011

Methods — techniques the papers use, named apart from their topics

bitonic sort · 1.0FPGA-based in-situ processing · 1.0stale value tolerance · 0.7asynchronous synchronization · 0.7interval analysis · 0.2functional simulation · 0.2CPI stack generation · 0.2source-to-source translation · 0.1run-time memory access range analysis · 0.1
YearPublicationVenuePosition
2021 NASCENT: Near-Storage Acceleration of Database Sort on SmartSSD
abstract
As the size of data generated every day grows dramatically, the computational bottleneck of computer systems has been shifted toward the storage devices. Thanks to recent developments in storage devices, the interface between the storage and the computational platforms has become the main limitation as it provides limited bandwidth which does not scale when the number of storage devices increases. Interconnect networks limit the performance of the system when independent operations are executing on different storage devices since they do not provide simultaneous accesses to all the storage devices. Offloading the computations to the storage devices eliminates the burden of data transfer from the interconnects. Emerging as a nascent computing trend, near storage computing offloads a portion of computation to the storage devices to accelerate the big data applications. In this paper, we propose a near storage accelerator for database sort, NASCENT, which utilizes Samsung SmartSSD, an NVMe flash drive with an on-board FPGA chip that processes data in-situ. We propose, to the best of our knowledge, the first near storage database sort based on bitonic sort which considers the specifications of the storage devices to increase the scalability of computer systems as the number of storage devices increases. NASCENT improves both performance and energy efficiency as the number of storage devices increases. With 12 SmartSSDs, NASCENT is 7.6x (147.2x) faster and 5.6x (131.4x) more energy efficient than the FPGA (CPU) baseline.
Sahand Salamat, Armin Haj Aboutalebi, Behnam Khaleghi, Joo Hwan Lee, Yang-Seok Ki, Tajana Rosing
FPGA4
2020 A Generic FPGA Accelerator for Minimum Storage Regenerating Codes
abstract
Erasure coding is widely used in storage systems to achieve fault tolerance while minimizing the storage overhead. Recently, Minimum Storage Regenerating (MSR) codes are emerging to minimize repair bandwidth while maintaining the storage efficiency. Traditionally, erasure coding is implemented in the storage software stacks, which hinders normal operations and blocks resources that could be serving other user needs due to poor cache performance and costs high CPU and memory utilizations. In this paper, we propose a generic FPGA accelerator for MSR codes encoding/decoding which maximizes the computation parallelism and minimizes the data movement between off-chip DRAM and the on-chip SRAM buffers. To demonstrate the efficiency of our proposed accelerator, we implemented the encoding/decoding algorithms for a specific MSR code called Zigzag code on Xilinx VCU1525 acceleration card. Our evaluation shows our proposed accelerator can achieve ~2.4-3.1x better throughput and ~4.2-5.7x better power efficiency compared to the state-of-art multi-core CPU implementation and ~2.8-3.3x better throughput and ~4.2-5.3x better power efficiency compared to a modern GPU accelerator.
Joo Hwan Lee, Rekha Pitchumani, Yang-Seok Ki, A. L. Narasimha Reddy, Paul Gratz
ASP-DAC2
2020 On the Limits of Parallelizing Convolutional Neural Networks on GPUs
abstract
GPUs are currently the platform of choice for training neural networks. However, training a deep neural network (DNN) is a time-consuming process even on GPUs because of the massive number of parameters that have to be learned. As a result, accelerating DNN training has been an area of significant research in the last couple of years.
Behnam Pourghassemi, Joo Hwan Lee, Aparna Chandramowlishwaran
SPAA3
2019 Empirical Investigation of Stale Value Tolerance on Parallel RNN Training
abstract
The objective of this paper is to provide a detailed understanding of stale value tolerance of parallel training. During parallel training, multiple workers read-and-modify shared model parameters multiple times, incurring multiple data transactions between workers, most of which are redundant due to the stale value tolerant characteristic of training. While considerable effort has tried to reduce the excessive data communication by utilizing stale value tolerance, there is a lack of detailed understanding of stale value tolerance and its dependence on multiple design choices in training of neural networks. This ambiguity has prevented domain experts from designing systems that take full advantage of the performance potential by leveraging stale value tolerance. This paper investigates how communication reduction affects the progress of parallel training for recurrent neural networks (RNN). We investigate stale value tolerance of RNN training by varying the update density, activation functions, and learning rate.
Joo Hwan Lee, Hyesoon Kim
ISPASS1
2018 StaleLearn: Learning Acceleration with Asynchronous Synchronization Between Model Replicas on PIM
abstract
GPU has become popular with a large amount of parallelism found in learning. While the GPU has been effective for many learning tasks, still many GPU learning applications have low execution efficiency due to sparse data. Sparse data induces divergent memory accesses with low locality, thereby consuming a large fraction of execution time transferring data across the memory hierarchy. Although a considerable effort has been devoted to reducing the memory divergence, iterative-convergent learning provides a unique opportunity to achieve full potential in modern GPUs that it allows different threads to continue computation using stale values. In this paper, we propose StaleLearn, a learning acceleration mechanism to reduce the memory divergence overhead of GPU learning by utilizing the stale value tolerance of the iterative-convergent learning. Based on the stale value tolerance, StaleLearn transforms the problem of divergent memory accesses into the synchronization problem by replicating the model and reduces the synchronization overhead by asynchronous synchronization on Processor-in-Memory (PIM). The stale value tolerance enables a clear task decomposition between the GPU and PIM, which can effectively exploit parallelism between PIM and GPU. On average, our approach accelerates representative GPU learning applications by 3.17 times with existing PIM proposals.
Joo Hwan Lee, Hyesoon Kim
IEEE Trans. Computers1
2015 BSSync: Processing Near Memory for Machine Learning Workloads with Bounded Staleness Consistency Models
abstract
Parallel machine learning workloads have become prevalent in numerous application domains. Many of these workloads are iterative convergent, allowing different threads to compute in an asynchronous manner, relaxing certain read-after-write data dependencies to use stale values. While considerable effort has been devoted to reducing the communication latency between nodes by utilizing asynchronous parallelism, inefficient utilization of relaxed consistency models within a single node have caused parallel implementations to have low execution efficiency. The long latency and serialization caused by atomic operations have a significant impact on performance. The data communication is not overlapped with the main computation, which reduces execution efficiency. The inefficiency comes from the data movement between where they are stored and where they are processed. In this work, we propose Bounded Staled Sync (BSSync), a hardware support for the bounded staleness consistency model, which accompanies simple logic layers in the memory hierarchy. BSSync overlaps the long latency atomic operation with the main computation, targeting iterative convergent machine learning workloads. Compared to previous work that allows staleness for read operations, BSSync utilizes staleness for write operations, allowing stale-writes. We demonstrate the benefit of the proposed scheme for representative machine learning workloads. On average, our approach outperforms the baseline asynchronous parallel implementation by 1.33x times.
Joo Hwan Lee, Jaewoong Sim, Hyesoon Kim
PACT1
2014 GPUMech: GPU Performance Modeling Technique Based on Interval Analysis
abstract
GPU has become a first-order computing plat-form. Nonetheless, not many performance modeling techniques have been developed for architecture studies. Several GPU analytical performance models have been proposed, but they mostly target application optimizations rather than the study of different architecture design options. Interval analysis is a relatively accurate performance modeling technique, which traverses the instruction trace and uses functional simulators, e.g., Cache simulator, to track the stall events that cause performance loss. It shows hundred times of speedup compared to detailed timing simulations and better accuracy compared to pure analytical models. However, previous techniques are limited to CPUs and not applicable to multithreaded architectures. In this work, we propose GPU Mech, an interval analysis-based performance modeling technique for GPU architectures. GPU Mech models multithreading and resource contentions caused by memory divergence. We compare GPU Mech with a detailed timing simulator and show that on average, GPU Mechhas 13.2% error for modeling the round-robin scheduling policy and 14.0% error for modeling the greedy-then-oldest policy while achieving a 97x faster simulation speed. In addition, GPU Mech generates CPI stacks, which help hardware/software developers to visualize performance bottlenecks of a kernel.
Jen-Cheng Huang, Joo Hwan Lee, Hyesoon Kim, Hsien-Hsin S. Lee
MICRO2
2011 Achieving a single compute device image in OpenCL for multiple GPUs
abstract
In this paper, we propose an OpenCL framework that combines multiple GPUs and treats them as a single compute device. Providing a single virtual compute device image to the user makes an OpenCL application written for a single GPU portable to the platform that has multiple GPU devices. It also makes the application exploit full computing power of the multiple GPU devices and the total amount of GPU memories available in the platform. Our OpenCL framework automatically distributes at run-time the OpenCL kernel written for a single GPU into multiple CUDA kernels that execute on the multiple GPU devices. It applies a run-time memory access range analysis to the kernel by performing a sampling run and identifies an optimal workload distribution for the kernel. To achieve a single compute device image, the runtime maintains virtual device memory that is allocated in the main memory. The OpenCL runtime treats the memory as if it were the memory of a single GPU device and keeps it consistent to the memories of the multiple GPU devices. Our OpenCL-C-to-C translator generates the sampling code from the OpenCL kernel code and OpenCL-C-to-CUDA-C translator generates the CUDA kernel code for the distributed OpenCL kernel. We show the effectiveness of our OpenCL framework by implementing the OpenCL runtime and two source-to-source translators. We evaluate its performance with a system that contains 8 GPUs using 11 OpenCL benchmark applications.
Honggyu Kim, Joo Hwan Lee, Jaejin Lee
PPoPP3