EDBT 2026 Demo / reviewers in the wild / expert
Keunsoo Kim
dblp:120/9078
· DBLP profile ↗
12ranked-venue papers
4as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 3 first-authorSoftware engineering, systems software and programming languages · 4Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
11 papers |
Processor architecture and microarchitecture · 32% GPUs and heterogeneous computing · 24% Storage systems · 10% | |
| Computer networks
1 paper |
Edge and fog computing · 87% Cellular and mobile networks · 13% |
Topics — the 30 heaviest of 39, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture › out-of-order execution
register renaming |
0.7 | 2 | 2019 | OverCome: Coarse-Grained Instruction Commit with Handover Register Renaming · IEEE Trans. Computers 2019 WIR: Warp Instruction Reuse to Minimize Repeated Computations in GPUs · HPCA 2018 |
GPUs and heterogeneous computing › GPU scheduling
warp scheduling |
0.6 | 2 | 2019 | Adaptive Cooperation of Prefetching and Warp Scheduling on GPUs · IEEE Trans. Computers 2019 APRES: Improving Cache Efficiency by Exploiting Load Characteristics on GPUs · ISCA 2016 |
Memory systems › cache
prefetching |
0.5 | 3 | 2019 | Adaptive Cooperation of Prefetching and Warp Scheduling on GPUs · IEEE Trans. Computers 2019 APRES: Improving Cache Efficiency by Exploiting Load Characteristics on GPUs · ISCA 2016 Warped-preexecution: A GPU pre-execution approach for improving latency hiding · HPCA 2016 |
GPUs and heterogeneous computing
GPU power management |
0.5 | 2 | 2017 | Improving Energy Efficiency of GPUs through Data Compression and Compressed Execution · IEEE Trans. Computers 2017 Warped-compression: enabling power efficient GPUs through register compression · ISCA 2015 |
Storage systems › computational storage
in-storage computing |
0.4 | 1 | 2020 | REACT: Scalable and High-Performance Regular Expression Pattern Matching Accelerator for In-Storage Processing · IEEE Trans. Parallel Distributed Syst. 2020 |
Hardware accelerators and domain-specific architectures › pattern matching accelerator
regular expression matching accelerator |
0.4 | 1 | 2020 | REACT: Scalable and High-Performance Regular Expression Pattern Matching Accelerator for In-Storage Processing · IEEE Trans. Parallel Distributed Syst. 2020 |
Storage systems
adaptive prefetching |
0.4 | 1 | 2019 | Adaptive Cooperation of Prefetching and Warp Scheduling on GPUs · IEEE Trans. Computers 2019 |
Processor architecture and microarchitecture › out-of-order execution
instruction window |
0.4 | 1 | 2019 | OverCome: Coarse-Grained Instruction Commit with Handover Register Renaming · IEEE Trans. Computers 2019 |
Processor architecture and microarchitecture › register file
physical register file |
0.4 | 1 | 2019 | OverCome: Coarse-Grained Instruction Commit with Handover Register Renaming · IEEE Trans. Computers 2019 |
Processor architecture and microarchitecture
reorder buffer |
0.4 | 1 | 2019 | OverCome: Coarse-Grained Instruction Commit with Handover Register Renaming · IEEE Trans. Computers 2019 |
GPUs and heterogeneous computing
GPU microarchitecture |
0.3 | 1 | 2018 | WIR: Warp Instruction Reuse to Minimize Repeated Computations in GPUs · HPCA 2018 |
Processor architecture and microarchitecture › dynamic optimization
instruction reuse |
0.3 | 1 | 2018 | WIR: Warp Instruction Reuse to Minimize Repeated Computations in GPUs · HPCA 2018 |
Processor architecture and microarchitecture
latency hiding |
0.3 | 2 | 2016 | Warped-preexecution: A GPU pre-execution approach for improving latency hiding · HPCA 2016 APRES: Improving Cache Efficiency by Exploiting Load Characteristics on GPUs · ISCA 2016 |
Parallel and multicore computing
thread-level parallelism |
0.3 | 2 | 2016 | Virtual Thread: Maximizing Thread-Level Parallelism beyond GPU Scheduling Limit · ISCA 2016 APRES: Improving Cache Efficiency by Exploiting Load Characteristics on GPUs · ISCA 2016 |
Storage systems
data compression |
0.3 | 1 | 2017 | Improving Energy Efficiency of GPUs through Data Compression and Compressed Execution · IEEE Trans. Computers 2017 |
GPUs and heterogeneous computing › GPU microarchitecture
GPU register file |
0.3 | 1 | 2017 | Improving Energy Efficiency of GPUs through Data Compression and Compressed Execution · IEEE Trans. Computers 2017 |
Energy-efficient computing › memory energy efficiency
register file energy reduction |
0.3 | 1 | 2017 | Improving Energy Efficiency of GPUs through Data Compression and Compressed Execution · IEEE Trans. Computers 2017 |
Memory systems › memory interference
cache contention |
0.2 | 1 | 2016 | APRES: Improving Cache Efficiency by Exploiting Load Characteristics on GPUs · ISCA 2016 |
Cloud and datacenter computing › resource allocation
dynamic resource allocation |
0.2 | 1 | 2016 | Warped-Slicer: Efficient Intra-SM Slicing through Dynamic Resource Partitioning for GPU Multiprogramming · ISCA 2016 |
GPUs and heterogeneous computing
GPU architecture |
0.2 | 1 | 2016 | Virtual Thread: Maximizing Thread-Level Parallelism beyond GPU Scheduling Limit · ISCA 2016 |
GPUs and heterogeneous computing
GPU computing |
0.2 | 1 | 2016 | Warped-preexecution: A GPU pre-execution approach for improving latency hiding · HPCA 2016 |
GPUs and heterogeneous computing
GPU sharing |
0.2 | 1 | 2016 | Warped-Slicer: Efficient Intra-SM Slicing through Dynamic Resource Partitioning for GPU Multiprogramming · ISCA 2016 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.2 | 1 | 2016 | Warped-preexecution: A GPU pre-execution approach for improving latency hiding · HPCA 2016 |
Processor architecture and microarchitecture › speculative execution
pre-execution |
0.2 | 1 | 2016 | Warped-preexecution: A GPU pre-execution approach for improving latency hiding · HPCA 2016 |
Processor architecture and microarchitecture › many-core architecture
streaming multiprocessor |
0.2 | 1 | 2016 | Warped-Slicer: Efficient Intra-SM Slicing through Dynamic Resource Partitioning for GPU Multiprogramming · ISCA 2016 |
Parallel and multicore computing › parallel scheduling
thread scheduling |
0.2 | 1 | 2016 | Virtual Thread: Maximizing Thread-Level Parallelism beyond GPU Scheduling Limit · ISCA 2016 |
Edge and fog computing › mobile edge computing
computation offloading |
0.2 | 1 | 2015 | Network Variation and Fault Tolerant Performance Acceleration in Mobile Devices with Simultaneous Remote Execution · IEEE Trans. Computers 2015 |
Edge and fog computing
remote execution |
0.2 | 1 | 2015 | Network Variation and Fault Tolerant Performance Acceleration in Mobile Devices with Simultaneous Remote Execution · IEEE Trans. Computers 2015 |
Distributed systems
fault tolerance |
0.2 | 1 | 2015 | Network Variation and Fault Tolerant Performance Acceleration in Mobile Devices with Simultaneous Remote Execution · IEEE Trans. Computers 2015 |
Hardware reliability and fault tolerance
network fault tolerance |
0.2 | 1 | 2015 | Network Variation and Fault Tolerant Performance Acceleration in Mobile Devices with Simultaneous Remote Execution · IEEE Trans. Computers 2015 |
Methods — techniques the papers use, named apart from their topics
parallel processing architecture · 0.4data access scheduling · 0.4warp grouping · 0.4history-based approach · 0.4dynamic prefetching · 0.4cycle-accurate simulation · 0.3compressed execution · 0.3warp pre-execution mode · 0.2register renaming · 0.2pre-load into l1 cache · 0.2simultaneous remote execution · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | REACT: Scalable and High-Performance Regular Expression Pattern Matching Accelerator for In-Storage ProcessingabstractThis article proposes REACT, a regular expression matching accelerator, which can be embedded in a modern Solid-State Drive (SSD) and a novel data access scheduling algorithm for high matching throughput. Specifically, REACT, including our data access scheduling algorithm, increases the utilization of SSD and the degree of internal memory parallelism for pattern matching processes. While the low-level flash exhibits long latency, modern SSDs in practice achieve high I/O performance by utilizing the massive internal parallelism at the system-level. However, exploiting the parallelism is limited for pattern matching since the subblocks, which constitute an input data and can be placed in multiple flash pages, should be tested in a sequence to process the input correctly. This limitation can induce low utilization of the accelerator. To address this challenge, the proposed REACT simultaneously processes multiple input streams with a parallel processing architecture to maximize matching throughput by hiding the long and irregular latency. The scheduling algorithm finds a data stream which requires a sub-block in closest time and prioritizes the access request to reduce the data stall of REACT. REACT achieves maximum 22.6 percent of matching throughput improvement on a 16channel high-performance SSD compared to the accelerator without the proposed scheduling algorithm. Won Seob Jeong, Changmin Lee 0002, Keunsoo Kim, Myung Kuk Yoon, Won Jeon, Myoungsoo Jung, Won Woo Ro |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2019 | OverCome: Coarse-Grained Instruction Commit with Handover Register RenamingabstractCoarse-grained instruction commit mechanisms enabled the effective size of the instruction window to be as large as possible by committing a group of instructions atomically. Within a group, the reorder buffer (ROB) and physical registerfile (PRF) entries are conservatively managed, and thus the instruction window can handle more in-flight instructions beyond the hardware limit. However, previous approaches have suffered from high storage requirements for managing group information and unbalanced lifetime of instruction window resources, i.e., the ROB and PRF. In this paper, we propose an OverCome microarchitecture based on a history-based approach to address these problems. First, OverCome retains the conservative allocation of the ROB regardless of the group size limit, thereby providing high scalability. Second, it handles the information of numerous groups with a low storage cost. These two techniques achieve a significant reduction in the pressure on the ROB; thus, a new bottleneck arises: the pressure on the PRF. To address this issue, we propose a novel register renaming technique to reduce the lifetime of physical registers to a large extent, by tightly coupling the early release and lazy allocation schemes. Thus, the proposed design strikes a balance between the ROB and PRF requirements. Detailed evaluation of the proposed techniques on a state-of-the-art superscalar processor shows that our proposals augment the effective size of the instruction window by more than 4×, with a net overhead of less than 3 percent of the core area. Ipoom Jeong, Changmin Lee 0002, Keunsoo Kim, Won Woo Ro |
IEEE Trans. Computers | 3 |
| 2019 | Adaptive Cooperation of Prefetching and Warp Scheduling on GPUsabstractThis paper proposes a new architecture, called Adaptive PREfetching and Scheduling (APRES), which improves cache efficiency of GPUs. APRES relies on the observation that GPU loads tend to have either high locality or strided access patterns across warps. APRES schedules warps so that as many cache hits are generated as possible before the generation of any cache miss. Without directly predicting future cache hits/misses for each warp, APRES creates a warp group that will execute the same static load shortly and prioritizes the grouped warps. If the first executed warp in the group hits the cache, grouped warps are likely to access the same cache lines. Unless, APRES considers the load as a strided type and generates prefetch requests for the grouped warps. In addition, APRES includes a new dynamic L1 prefetch and data cache partitioning to reduce contentions between demand-fetched and prefetched lines. In our evaluation, APRES achieves 27.8 percent performance improvement. Yunho Oh, Keunsoo Kim, Myung Kuk Yoon, Jong Hyun Park, Yongjun Park 0001, Murali Annavaram, Won Woo Ro |
IEEE Trans. Computers | 2 |
| 2018 | WIR: Warp Instruction Reuse to Minimize Repeated Computations in GPUsabstractWarp instructions with an identical arithmetic operation on same input values produce the identical computation results. This paper proposes warp instruction reuse to allow such repeated warp instructions to reuse previous computation results instead of actually executing the instructions. Bypassing register reading, functional unit, and register writing operations improves energy efficiency. This reuse technique is especially beneficial for GPUs since a GPU warp register is usually as wide as thousands of bits. In addition, we propose warp register reuse which allows identical warp register values to share a single physical register through register renaming. The register reuse technique enables to see if different logical warp registers have an identical value by only looking at their physical warp register IDs. Based on this observation, warp register reuse helps to perform all necessary operations for warp instruction reuse with register IDs, which is substantially more efficient than directly manipulating register values. Performance evaluation shows that 20.5% SM energy and 10.7% GPU energy can be saved by allowing 18.7% of warp instructions to reuse prior results. Keunsoo Kim, Won Woo Ro |
HPCA | 1 |
| 2017 | Improving Energy Efficiency of GPUs through Data Compression and Compressed ExecutionabstractGPU design trends show that the register file size will continue to increase to enable even more thread level parallelism. As a result register file consumes a large fraction of the total GPU chip power. This paper explores register file data compression for GPUs to improve power efficiency. Compression reduces the width of the register file read and write operations, which in turn reduces dynamic power. This work is motivated by the observation that the register values of threads within the same warp are similar, namely the arithmetic differences between two successive thread registers is small. Compression exploits the value similarity by removing data redundancy of register values. Without decompressing operand values some instructions can be processed inside register file, which enables to further save energy by minimizing data movement and processing in power hungry main execution unit. Evaluation results show that the proposed techniques save 25 percent of the total register file energy consumption and 21 percent of the total execution unit energy consumption with negligible performance impact. Sangpil Lee, Keunsoo Kim, Gunjae Koo, Hyeran Jeon, Murali Annavaram, Won Woo Ro |
IEEE Trans. Computers | 2 |
| 2016 | Warped-preexecution: A GPU pre-execution approach for improving latency hidingabstractThis paper presents a pre-execution approach for improving GPU performance, called P-mode (pre-execution mode). GPUs utilize a number of concurrent threads for hiding processing delay of operations. However, certain long-latency operations such as off-chip memory accesses often take hundreds of cycles and hence leads to stalls even in the presence of thread concurrency and fast thread switching capability. It is unclear if adding more threads can improve latency tolerance due to increased memory contention. Further, adding more threads increases on-chip storage demands. Instead we propose that when a warp is stalled on a long-latency operation it enters P-mode. In P-mode, a warp continues to fetch and decode successive instructions to identify any independent instruction that is not on the long latency dependence chain. These independent instructions are then pre-executed. To tackle write-after-write and write-after-read hazards, during P-mode output values are written to renamed physical registers. We exploit the register file underutilization to re-purpose a few unused registers to store the P-mode results. When a warp is switched from P-mode to normal execution mode it reuses pre-executed results by reading the renamed registers. Any global load operation in P-mode is transformed into a pre-load which fetches data into the L1 cache to reduce future memory access penalties. Our evaluation results show 23% performance improvement for memory intensive applications, without negatively impacting other application categories. Keunsoo Kim, Sangpil Lee, Myung Kuk Yoon, Gunjae Koo, Won Woo Ro, Murali Annavaram |
HPCA | 1 |
| 2016 | APRES: Improving Cache Efficiency by Exploiting Load Characteristics on GPUsabstractLong memory latency and limited throughput become performance bottlenecks of GPGPU applications. The latency takes hundreds of cycles which is difficult to be hidden by simply interleaving tens of warp execution. While cache hierarchy helps to reduce memory system pressure, massive Thread-Level Parallelism (TLP) often causes excessive cache contention. This paper proposes Adaptive PREfetching and Scheduling (APRES) to improve GPU cache efficiency. APRES relies on the following observations. First, certain static load instructions tend to generate memory addresses having very high locality. Second, although loads have no locality, the access addresses still can show highly strided access pattern. Third, the locality behavior tends to be consistent regardless of warp ID. APRES schedules warps so that as many cache hits generated as possible before any cache misses generated. This is to minimize cache thrashing when many warps are contending for a cache line. However, to realize this operation, it is required to predict which warp will hit the cache in the near future. Without directly predicting future cache hit/miss for each warp, APRES creates a group of warps that will execute the same load instruction in the near future. Based on the third observation, we expect the locality behavior is consistent over all warps in the group. If the first executed warp in the group hits the cache, then the load is considered as a high locality type, and APRES prioritizes all warps in the group. Group prioritization leads to consecutive cache hits, because the grouped warps are likely to access the same cache line. If the first warp missed the cache, then the load is considered as a strided type, and APRES generates prefetch requests for the other warps in the group. After that, APRES prioritizes prefetch targeted warps so that the demand requests are merged to Miss Status Holding Register (MSHR) or prefetched lines can be accessed. On memory-intensive applications, APRES achieves 31.7% performance improvement compared to the baseline GPU and 7.2% additional speedup compared to the best combination of existing warp scheduling and prefetching methods. Yunho Oh, Keunsoo Kim, Myung Kuk Yoon, Jong Hyun Park, Yongjun Park 0001, Won Woo Ro, Murali Annavaram |
ISCA | 2 |
| 2016 | Warped-Slicer: Efficient Intra-SM Slicing through Dynamic Resource Partitioning for GPU MultiprogrammingabstractAs technology scales, GPUs are forecasted to incorporate an ever-increasing amount of computing resources to support thread-level parallelism. But even with the best effort, exposing massive thread-level parallelism from a single GPU kernel, particularly from general purpose applications, is going to be a difficult challenge. In some cases, even if there is sufficient thread-level parallelism in a kernel, there may not be enough available memory bandwidth to support such massive concurrent thread execution. Hence, GPU resources may be underutilized as more general purpose applications are ported to execute on GPUs. In this paper, we explore multiprogramming GPUs as a way to resolve the resource underutilization issue. There is a growing hardware support for multiprogramming on GPUs. Hyper-Q has been introduced in the Kepler architecture which enables multiple kernels to be invoked via tens of hardware queue streams. Spatial multitasking has been proposed to partition GPU resources across multiple kernels. But the partitioning is done at the coarse granularity of streaming multiprocessors (SMs) where each kernel is assigned to a subset of SMs. In this paper, we advocate for partitioning a single SM across multiple kernels, which we term as intra-SM slicing. We explore various intra-SM slicing strategies that slice resources within each SM to concurrently run multiple kernels on the SM. Our results show that there is not one intra-SM slicing strategy that derives the best performance for all application pairs. We propose Warped-Slicer, a dynamic intra-SM slicing strategy that uses an analytical method for calculating the SM resource partitioning across different kernels that maximizes performance. The model relies on a set of short online profile runs to determine how each kernel's performance varies as more thread blocks from each kernel are assigned to an SM. The model takes into account the interference effect of shared resource usage across multiple kernels. The model is also computationally efficient and can determine the resource partitioning quickly to enable dynamic decision making as new kernels enter the system. We demonstrate that the proposed Warped-Slicer approach improves performance by 23% over the baseline multiprogramming approach with minimal hardware overhead. Qiumin Xu, Hyeran Jeon, Keunsoo Kim, Won Woo Ro, Murali Annavaram |
ISCA | 3 |
| 2016 | Virtual Thread: Maximizing Thread-Level Parallelism beyond GPU Scheduling LimitabstractModern GPUs require tens of thousands of concurrent threads to fully utilize the massive amount of processing resources. However, thread concurrency in GPUs can be diminished either due to shortage of thread scheduling structures (scheduling limit), such as available program counters and single instruction multiple thread stacks, or due to shortage of on-chip memory (capacity limit), such as register file and shared memory. Our evaluations show that in practice concurrency in many general purpose applications running on GPUs is curtailed by the scheduling limit rather than the capacity limit. Maximizing the utilization of on-chip memory resources without unduly increasing the scheduling complexity is a key goal of this paper. This paper proposes a Virtual Thread (VT) architecture which assigns Cooperative Thread Arrays (CTAs) up to the capacity limit, while ignoring the scheduling limit. However, to reduce the logic complexity of managing more threads concurrently, we propose to place CTAs into active and inactive states, such that the number of active CTAs still respects the scheduling limit. When all the warps in an active CTA hit a long latency stall, the active CTA is context switched out and the next ready CTA takes its place. We exploit the fact that both active and inactive CTAs still fit within the capacity limit which obviates the need to save and restore large amounts of CTA state. Thus VT significantly reduces performance penalties of CTA swapping. By swapping between active and inactive states, VT can exploit higher degree of thread level parallelism without increasing logic complexity. Our simulation results show that VT improves performance by 23.9% on average. Myung Kuk Yoon, Keunsoo Kim, Sangpil Lee, Won Woo Ro, Murali Annavaram |
ISCA | 2 |
| 2016 | Server side, play buffer based quality control for adaptive media streaming
Keunsoo Kim, Benjamin Y. Cho, Won Woo Ro |
Multim. Tools Appl. | 1 |
| 2015 | Warped-compression: enabling power efficient GPUs through register compressionabstractThis paper presents Warped-Compression, a warp-level register compression scheme for reducing GPU power consumption. This work is motivated by the observation that the register values of threads within the same warp are similar, namely the arithmetic differences between two successive thread registers is small. Removing data redundancy of register values through register compression reduces the effective register width, thereby enabling power reduction opportunities. GPU register files are huge as they are necessary to keep concurrent execution contexts and to enable fast context switching. As a result register file consumes a large fraction of the total GPU chip power. GPU design trends show that the register file size will continue to increase to enable even more thread level parallelism. To reduce register file data redundancy warped-compression uses low-cost and implementation-efficient base-delta-immediate (BDI) compression scheme, that takes advantage of banked register file organization used in GPUs. Since threads within a warp write values with strong similarity, BDI can quickly compress and decompress by selecting either a single register, or one of the register banks, as the primary base and then computing delta values of all the other registers, or banks. Warped-compression can be used to reduce both dynamic and leakage power. By compressing register values, each warp-level register access activates fewer register banks, which leads to reduction in dynamic power. When fewer banks are used to store the register content, leakage power can be reduced by power gating the unused banks. Evaluation results show that register compression saves 25% of the total register file power consumption. Sangpil Lee, Keunsoo Kim, Gunjae Koo, Hyeran Jeon, Won Woo Ro, Murali Annavaram |
ISCA | 2 |
| 2015 | Network Variation and Fault Tolerant Performance Acceleration in Mobile Devices with Simultaneous Remote ExecutionabstractAs mobile applications provide increasingly richer features to end users, it has become imperative to overcome the constraints of a resource-limited mobile hardware. Remote execution is one promising technique to resolve this important problem. Using this technique, the computation intensive part of the workload is migrated to resource-rich servers, and then once the computation is completed, the results can be returned to the client devices. To enable this operation, strong wireless connectivity is required. However, unstable wireless connections are the staple of real-life. This makes performance unpredictable, sometimes offsetting the benefits brought by this technique and leading to performance degradation. To address this problem, in this paper, we present a Simultaneous Remote Execution (SRE) model for mobile devices. Our SRE model performs concurrent executions both locally and remotely. Therefore, the worst-case execution time on fluctuating network condition is significantly reduced. In addition, SRE provides inherent tolerance for abrupt network failure. We designed and implemented an SRE-based offloading system consisting of a real smartphone and a remote server connected via 3G and Wifi networks. The experimental results under various real-life network variation scenarios show that SRE outperforms the alternative schemes in highly fluctuating network environments. Keunsoo Kim, Benjamin Y. Cho, Won Woo Ro, Jean-Luc Gaudiot |
IEEE Trans. Computers | 1 |