Keunsoo Kim

dblp:120/9078 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 3 first-authorSoftware engineering, systems software and programming languages · 4Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
11 papers
Processor architecture and microarchitecture · 32% GPUs and heterogeneous computing · 24% Storage systems · 10%
Computer networks
1 paper
Edge and fog computing · 87% Cellular and mobile networks · 13%

Topics — the 30 heaviest of 39, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture › out-of-order execution
register renaming
0.722019
OverCome: Coarse-Grained Instruction Commit with Handover Register Renaming · IEEE Trans. Computers 2019
WIR: Warp Instruction Reuse to Minimize Repeated Computations in GPUs · HPCA 2018
GPUs and heterogeneous computing › GPU scheduling
warp scheduling
0.622019
Adaptive Cooperation of Prefetching and Warp Scheduling on GPUs · IEEE Trans. Computers 2019
APRES: Improving Cache Efficiency by Exploiting Load Characteristics on GPUs · ISCA 2016
Memory systems › cache
prefetching
0.532019
Adaptive Cooperation of Prefetching and Warp Scheduling on GPUs · IEEE Trans. Computers 2019
APRES: Improving Cache Efficiency by Exploiting Load Characteristics on GPUs · ISCA 2016
Warped-preexecution: A GPU pre-execution approach for improving latency hiding · HPCA 2016
GPUs and heterogeneous computing
GPU power management
0.522017
Improving Energy Efficiency of GPUs through Data Compression and Compressed Execution · IEEE Trans. Computers 2017
Warped-compression: enabling power efficient GPUs through register compression · ISCA 2015
Storage systems › computational storage
in-storage computing
0.412020
REACT: Scalable and High-Performance Regular Expression Pattern Matching Accelerator for In-Storage Processing · IEEE Trans. Parallel Distributed Syst. 2020
Hardware accelerators and domain-specific architectures › pattern matching accelerator
regular expression matching accelerator
0.412020
REACT: Scalable and High-Performance Regular Expression Pattern Matching Accelerator for In-Storage Processing · IEEE Trans. Parallel Distributed Syst. 2020
Storage systems
adaptive prefetching
0.412019
Adaptive Cooperation of Prefetching and Warp Scheduling on GPUs · IEEE Trans. Computers 2019
Processor architecture and microarchitecture › out-of-order execution
instruction window
0.412019
OverCome: Coarse-Grained Instruction Commit with Handover Register Renaming · IEEE Trans. Computers 2019
Processor architecture and microarchitecture › register file
physical register file
0.412019
OverCome: Coarse-Grained Instruction Commit with Handover Register Renaming · IEEE Trans. Computers 2019
Processor architecture and microarchitecture
reorder buffer
0.412019
OverCome: Coarse-Grained Instruction Commit with Handover Register Renaming · IEEE Trans. Computers 2019
GPUs and heterogeneous computing
GPU microarchitecture
0.312018
WIR: Warp Instruction Reuse to Minimize Repeated Computations in GPUs · HPCA 2018
Processor architecture and microarchitecture › dynamic optimization
instruction reuse
0.312018
WIR: Warp Instruction Reuse to Minimize Repeated Computations in GPUs · HPCA 2018
Processor architecture and microarchitecture
latency hiding
0.322016
Warped-preexecution: A GPU pre-execution approach for improving latency hiding · HPCA 2016
APRES: Improving Cache Efficiency by Exploiting Load Characteristics on GPUs · ISCA 2016
Parallel and multicore computing
thread-level parallelism
0.322016
Virtual Thread: Maximizing Thread-Level Parallelism beyond GPU Scheduling Limit · ISCA 2016
APRES: Improving Cache Efficiency by Exploiting Load Characteristics on GPUs · ISCA 2016
Storage systems
data compression
0.312017
Improving Energy Efficiency of GPUs through Data Compression and Compressed Execution · IEEE Trans. Computers 2017
GPUs and heterogeneous computing › GPU microarchitecture
GPU register file
0.312017
Improving Energy Efficiency of GPUs through Data Compression and Compressed Execution · IEEE Trans. Computers 2017
Energy-efficient computing › memory energy efficiency
register file energy reduction
0.312017
Improving Energy Efficiency of GPUs through Data Compression and Compressed Execution · IEEE Trans. Computers 2017
Memory systems › memory interference
cache contention
0.212016
APRES: Improving Cache Efficiency by Exploiting Load Characteristics on GPUs · ISCA 2016
Cloud and datacenter computing › resource allocation
dynamic resource allocation
0.212016
Warped-Slicer: Efficient Intra-SM Slicing through Dynamic Resource Partitioning for GPU Multiprogramming · ISCA 2016
GPUs and heterogeneous computing
GPU architecture
0.212016
Virtual Thread: Maximizing Thread-Level Parallelism beyond GPU Scheduling Limit · ISCA 2016
GPUs and heterogeneous computing
GPU computing
0.212016
Warped-preexecution: A GPU pre-execution approach for improving latency hiding · HPCA 2016
GPUs and heterogeneous computing
GPU sharing
0.212016
Warped-Slicer: Efficient Intra-SM Slicing through Dynamic Resource Partitioning for GPU Multiprogramming · ISCA 2016
Processor architecture and microarchitecture
instruction-level parallelism
0.212016
Warped-preexecution: A GPU pre-execution approach for improving latency hiding · HPCA 2016
Processor architecture and microarchitecture › speculative execution
pre-execution
0.212016
Warped-preexecution: A GPU pre-execution approach for improving latency hiding · HPCA 2016
Processor architecture and microarchitecture › many-core architecture
streaming multiprocessor
0.212016
Warped-Slicer: Efficient Intra-SM Slicing through Dynamic Resource Partitioning for GPU Multiprogramming · ISCA 2016
Parallel and multicore computing › parallel scheduling
thread scheduling
0.212016
Virtual Thread: Maximizing Thread-Level Parallelism beyond GPU Scheduling Limit · ISCA 2016
Edge and fog computing › mobile edge computing
computation offloading
0.212015
Network Variation and Fault Tolerant Performance Acceleration in Mobile Devices with Simultaneous Remote Execution · IEEE Trans. Computers 2015
Edge and fog computing
remote execution
0.212015
Network Variation and Fault Tolerant Performance Acceleration in Mobile Devices with Simultaneous Remote Execution · IEEE Trans. Computers 2015
Distributed systems
fault tolerance
0.212015
Network Variation and Fault Tolerant Performance Acceleration in Mobile Devices with Simultaneous Remote Execution · IEEE Trans. Computers 2015
Hardware reliability and fault tolerance
network fault tolerance
0.212015
Network Variation and Fault Tolerant Performance Acceleration in Mobile Devices with Simultaneous Remote Execution · IEEE Trans. Computers 2015

Methods — techniques the papers use, named apart from their topics

parallel processing architecture · 0.4data access scheduling · 0.4warp grouping · 0.4history-based approach · 0.4dynamic prefetching · 0.4cycle-accurate simulation · 0.3compressed execution · 0.3warp pre-execution mode · 0.2register renaming · 0.2pre-load into l1 cache · 0.2simultaneous remote execution · 0.2
YearPublicationVenuePosition
2020 REACT: Scalable and High-Performance Regular Expression Pattern Matching Accelerator for In-Storage Processing
abstract
This article proposes REACT, a regular expression matching accelerator, which can be embedded in a modern Solid-State Drive (SSD) and a novel data access scheduling algorithm for high matching throughput. Specifically, REACT, including our data access scheduling algorithm, increases the utilization of SSD and the degree of internal memory parallelism for pattern matching processes. While the low-level flash exhibits long latency, modern SSDs in practice achieve high I/O performance by utilizing the massive internal parallelism at the system-level. However, exploiting the parallelism is limited for pattern matching since the subblocks, which constitute an input data and can be placed in multiple flash pages, should be tested in a sequence to process the input correctly. This limitation can induce low utilization of the accelerator. To address this challenge, the proposed REACT simultaneously processes multiple input streams with a parallel processing architecture to maximize matching throughput by hiding the long and irregular latency. The scheduling algorithm finds a data stream which requires a sub-block in closest time and prioritizes the access request to reduce the data stall of REACT. REACT achieves maximum 22.6 percent of matching throughput improvement on a 16channel high-performance SSD compared to the accelerator without the proposed scheduling algorithm.
Won Seob Jeong, Changmin Lee 0002, Keunsoo Kim, Myung Kuk Yoon, Won Jeon, Myoungsoo Jung, Won Woo Ro
IEEE Trans. Parallel Distributed Syst.3
2019 OverCome: Coarse-Grained Instruction Commit with Handover Register Renaming
abstract
Coarse-grained instruction commit mechanisms enabled the effective size of the instruction window to be as large as possible by committing a group of instructions atomically. Within a group, the reorder buffer (ROB) and physical registerfile (PRF) entries are conservatively managed, and thus the instruction window can handle more in-flight instructions beyond the hardware limit. However, previous approaches have suffered from high storage requirements for managing group information and unbalanced lifetime of instruction window resources, i.e., the ROB and PRF. In this paper, we propose an OverCome microarchitecture based on a history-based approach to address these problems. First, OverCome retains the conservative allocation of the ROB regardless of the group size limit, thereby providing high scalability. Second, it handles the information of numerous groups with a low storage cost. These two techniques achieve a significant reduction in the pressure on the ROB; thus, a new bottleneck arises: the pressure on the PRF. To address this issue, we propose a novel register renaming technique to reduce the lifetime of physical registers to a large extent, by tightly coupling the early release and lazy allocation schemes. Thus, the proposed design strikes a balance between the ROB and PRF requirements. Detailed evaluation of the proposed techniques on a state-of-the-art superscalar processor shows that our proposals augment the effective size of the instruction window by more than 4×, with a net overhead of less than 3 percent of the core area.
Ipoom Jeong, Changmin Lee 0002, Keunsoo Kim, Won Woo Ro
IEEE Trans. Computers3
2019 Adaptive Cooperation of Prefetching and Warp Scheduling on GPUs
abstract
This paper proposes a new architecture, called Adaptive PREfetching and Scheduling (APRES), which improves cache efficiency of GPUs. APRES relies on the observation that GPU loads tend to have either high locality or strided access patterns across warps. APRES schedules warps so that as many cache hits are generated as possible before the generation of any cache miss. Without directly predicting future cache hits/misses for each warp, APRES creates a warp group that will execute the same static load shortly and prioritizes the grouped warps. If the first executed warp in the group hits the cache, grouped warps are likely to access the same cache lines. Unless, APRES considers the load as a strided type and generates prefetch requests for the grouped warps. In addition, APRES includes a new dynamic L1 prefetch and data cache partitioning to reduce contentions between demand-fetched and prefetched lines. In our evaluation, APRES achieves 27.8 percent performance improvement.
Yunho Oh, Keunsoo Kim, Myung Kuk Yoon, Jong Hyun Park, Yongjun Park 0001, Murali Annavaram, Won Woo Ro
IEEE Trans. Computers2
2018 WIR: Warp Instruction Reuse to Minimize Repeated Computations in GPUs
abstract
Warp instructions with an identical arithmetic operation on same input values produce the identical computation results. This paper proposes warp instruction reuse to allow such repeated warp instructions to reuse previous computation results instead of actually executing the instructions. Bypassing register reading, functional unit, and register writing operations improves energy efficiency. This reuse technique is especially beneficial for GPUs since a GPU warp register is usually as wide as thousands of bits. In addition, we propose warp register reuse which allows identical warp register values to share a single physical register through register renaming. The register reuse technique enables to see if different logical warp registers have an identical value by only looking at their physical warp register IDs. Based on this observation, warp register reuse helps to perform all necessary operations for warp instruction reuse with register IDs, which is substantially more efficient than directly manipulating register values. Performance evaluation shows that 20.5% SM energy and 10.7% GPU energy can be saved by allowing 18.7% of warp instructions to reuse prior results.
Keunsoo Kim, Won Woo Ro
HPCA1
2017 Improving Energy Efficiency of GPUs through Data Compression and Compressed Execution
abstract
GPU design trends show that the register file size will continue to increase to enable even more thread level parallelism. As a result register file consumes a large fraction of the total GPU chip power. This paper explores register file data compression for GPUs to improve power efficiency. Compression reduces the width of the register file read and write operations, which in turn reduces dynamic power. This work is motivated by the observation that the register values of threads within the same warp are similar, namely the arithmetic differences between two successive thread registers is small. Compression exploits the value similarity by removing data redundancy of register values. Without decompressing operand values some instructions can be processed inside register file, which enables to further save energy by minimizing data movement and processing in power hungry main execution unit. Evaluation results show that the proposed techniques save 25 percent of the total register file energy consumption and 21 percent of the total execution unit energy consumption with negligible performance impact.
Sangpil Lee, Keunsoo Kim, Gunjae Koo, Hyeran Jeon, Murali Annavaram, Won Woo Ro
IEEE Trans. Computers2
2016 Warped-preexecution: A GPU pre-execution approach for improving latency hiding
abstract
This paper presents a pre-execution approach for improving GPU performance, called P-mode (pre-execution mode). GPUs utilize a number of concurrent threads for hiding processing delay of operations. However, certain long-latency operations such as off-chip memory accesses often take hundreds of cycles and hence leads to stalls even in the presence of thread concurrency and fast thread switching capability. It is unclear if adding more threads can improve latency tolerance due to increased memory contention. Further, adding more threads increases on-chip storage demands. Instead we propose that when a warp is stalled on a long-latency operation it enters P-mode. In P-mode, a warp continues to fetch and decode successive instructions to identify any independent instruction that is not on the long latency dependence chain. These independent instructions are then pre-executed. To tackle write-after-write and write-after-read hazards, during P-mode output values are written to renamed physical registers. We exploit the register file underutilization to re-purpose a few unused registers to store the P-mode results. When a warp is switched from P-mode to normal execution mode it reuses pre-executed results by reading the renamed registers. Any global load operation in P-mode is transformed into a pre-load which fetches data into the L1 cache to reduce future memory access penalties. Our evaluation results show 23% performance improvement for memory intensive applications, without negatively impacting other application categories.
Keunsoo Kim, Sangpil Lee, Myung Kuk Yoon, Gunjae Koo, Won Woo Ro, Murali Annavaram
HPCA1
2016 APRES: Improving Cache Efficiency by Exploiting Load Characteristics on GPUs
abstract
Long memory latency and limited throughput become performance bottlenecks of GPGPU applications. The latency takes hundreds of cycles which is difficult to be hidden by simply interleaving tens of warp execution. While cache hierarchy helps to reduce memory system pressure, massive Thread-Level Parallelism (TLP) often causes excessive cache contention. This paper proposes Adaptive PREfetching and Scheduling (APRES) to improve GPU cache efficiency. APRES relies on the following observations. First, certain static load instructions tend to generate memory addresses having very high locality. Second, although loads have no locality, the access addresses still can show highly strided access pattern. Third, the locality behavior tends to be consistent regardless of warp ID. APRES schedules warps so that as many cache hits generated as possible before any cache misses generated. This is to minimize cache thrashing when many warps are contending for a cache line. However, to realize this operation, it is required to predict which warp will hit the cache in the near future. Without directly predicting future cache hit/miss for each warp, APRES creates a group of warps that will execute the same load instruction in the near future. Based on the third observation, we expect the locality behavior is consistent over all warps in the group. If the first executed warp in the group hits the cache, then the load is considered as a high locality type, and APRES prioritizes all warps in the group. Group prioritization leads to consecutive cache hits, because the grouped warps are likely to access the same cache line. If the first warp missed the cache, then the load is considered as a strided type, and APRES generates prefetch requests for the other warps in the group. After that, APRES prioritizes prefetch targeted warps so that the demand requests are merged to Miss Status Holding Register (MSHR) or prefetched lines can be accessed. On memory-intensive applications, APRES achieves 31.7% performance improvement compared to the baseline GPU and 7.2% additional speedup compared to the best combination of existing warp scheduling and prefetching methods.
Yunho Oh, Keunsoo Kim, Myung Kuk Yoon, Jong Hyun Park, Yongjun Park 0001, Won Woo Ro, Murali Annavaram
ISCA2
2016 Warped-Slicer: Efficient Intra-SM Slicing through Dynamic Resource Partitioning for GPU Multiprogramming
abstract
As technology scales, GPUs are forecasted to incorporate an ever-increasing amount of computing resources to support thread-level parallelism. But even with the best effort, exposing massive thread-level parallelism from a single GPU kernel, particularly from general purpose applications, is going to be a difficult challenge. In some cases, even if there is sufficient thread-level parallelism in a kernel, there may not be enough available memory bandwidth to support such massive concurrent thread execution. Hence, GPU resources may be underutilized as more general purpose applications are ported to execute on GPUs. In this paper, we explore multiprogramming GPUs as a way to resolve the resource underutilization issue. There is a growing hardware support for multiprogramming on GPUs. Hyper-Q has been introduced in the Kepler architecture which enables multiple kernels to be invoked via tens of hardware queue streams. Spatial multitasking has been proposed to partition GPU resources across multiple kernels. But the partitioning is done at the coarse granularity of streaming multiprocessors (SMs) where each kernel is assigned to a subset of SMs. In this paper, we advocate for partitioning a single SM across multiple kernels, which we term as intra-SM slicing. We explore various intra-SM slicing strategies that slice resources within each SM to concurrently run multiple kernels on the SM. Our results show that there is not one intra-SM slicing strategy that derives the best performance for all application pairs. We propose Warped-Slicer, a dynamic intra-SM slicing strategy that uses an analytical method for calculating the SM resource partitioning across different kernels that maximizes performance. The model relies on a set of short online profile runs to determine how each kernel's performance varies as more thread blocks from each kernel are assigned to an SM. The model takes into account the interference effect of shared resource usage across multiple kernels. The model is also computationally efficient and can determine the resource partitioning quickly to enable dynamic decision making as new kernels enter the system. We demonstrate that the proposed Warped-Slicer approach improves performance by 23% over the baseline multiprogramming approach with minimal hardware overhead.
Qiumin Xu, Hyeran Jeon, Keunsoo Kim, Won Woo Ro, Murali Annavaram
ISCA3
2016 Virtual Thread: Maximizing Thread-Level Parallelism beyond GPU Scheduling Limit
abstract
Modern GPUs require tens of thousands of concurrent threads to fully utilize the massive amount of processing resources. However, thread concurrency in GPUs can be diminished either due to shortage of thread scheduling structures (scheduling limit), such as available program counters and single instruction multiple thread stacks, or due to shortage of on-chip memory (capacity limit), such as register file and shared memory. Our evaluations show that in practice concurrency in many general purpose applications running on GPUs is curtailed by the scheduling limit rather than the capacity limit. Maximizing the utilization of on-chip memory resources without unduly increasing the scheduling complexity is a key goal of this paper. This paper proposes a Virtual Thread (VT) architecture which assigns Cooperative Thread Arrays (CTAs) up to the capacity limit, while ignoring the scheduling limit. However, to reduce the logic complexity of managing more threads concurrently, we propose to place CTAs into active and inactive states, such that the number of active CTAs still respects the scheduling limit. When all the warps in an active CTA hit a long latency stall, the active CTA is context switched out and the next ready CTA takes its place. We exploit the fact that both active and inactive CTAs still fit within the capacity limit which obviates the need to save and restore large amounts of CTA state. Thus VT significantly reduces performance penalties of CTA swapping. By swapping between active and inactive states, VT can exploit higher degree of thread level parallelism without increasing logic complexity. Our simulation results show that VT improves performance by 23.9% on average.
Myung Kuk Yoon, Keunsoo Kim, Sangpil Lee, Won Woo Ro, Murali Annavaram
ISCA2
2016 Server side, play buffer based quality control for adaptive media streaming
Keunsoo Kim, Benjamin Y. Cho, Won Woo Ro
Multim. Tools Appl.1
2015 Warped-compression: enabling power efficient GPUs through register compression
abstract
This paper presents Warped-Compression, a warp-level register compression scheme for reducing GPU power consumption. This work is motivated by the observation that the register values of threads within the same warp are similar, namely the arithmetic differences between two successive thread registers is small. Removing data redundancy of register values through register compression reduces the effective register width, thereby enabling power reduction opportunities. GPU register files are huge as they are necessary to keep concurrent execution contexts and to enable fast context switching. As a result register file consumes a large fraction of the total GPU chip power. GPU design trends show that the register file size will continue to increase to enable even more thread level parallelism. To reduce register file data redundancy warped-compression uses low-cost and implementation-efficient base-delta-immediate (BDI) compression scheme, that takes advantage of banked register file organization used in GPUs. Since threads within a warp write values with strong similarity, BDI can quickly compress and decompress by selecting either a single register, or one of the register banks, as the primary base and then computing delta values of all the other registers, or banks. Warped-compression can be used to reduce both dynamic and leakage power. By compressing register values, each warp-level register access activates fewer register banks, which leads to reduction in dynamic power. When fewer banks are used to store the register content, leakage power can be reduced by power gating the unused banks. Evaluation results show that register compression saves 25% of the total register file power consumption.
Sangpil Lee, Keunsoo Kim, Gunjae Koo, Hyeran Jeon, Won Woo Ro, Murali Annavaram
ISCA2
2015 Network Variation and Fault Tolerant Performance Acceleration in Mobile Devices with Simultaneous Remote Execution
abstract
As mobile applications provide increasingly richer features to end users, it has become imperative to overcome the constraints of a resource-limited mobile hardware. Remote execution is one promising technique to resolve this important problem. Using this technique, the computation intensive part of the workload is migrated to resource-rich servers, and then once the computation is completed, the results can be returned to the client devices. To enable this operation, strong wireless connectivity is required. However, unstable wireless connections are the staple of real-life. This makes performance unpredictable, sometimes offsetting the benefits brought by this technique and leading to performance degradation. To address this problem, in this paper, we present a Simultaneous Remote Execution (SRE) model for mobile devices. Our SRE model performs concurrent executions both locally and remotely. Therefore, the worst-case execution time on fluctuating network condition is significantly reduced. In addition, SRE provides inherent tolerance for abrupt network failure. We designed and implemented an SRE-based offloading system consisting of a real smartphone and a remote server connected via 3G and Wifi networks. The experimental results under various real-life network variation scenarios show that SRE outperforms the alternative schemes in highly fluctuating network environments.
Keunsoo Kim, Benjamin Y. Cho, Won Woo Ro, Jean-Luc Gaudiot
IEEE Trans. Computers1