VLDB 2026 Research / reviewers in the wild / expert
Jaekyu Lee
dblp:02/7555
· DBLP profile ↗
17ranked-venue papers
7as first author
6since 2021 · last 2025
0000-0002-0574-5381ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 7 first-author · 3 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Let-Me-In: (Still) Employing In-pointer Bounds Metadata for Fine-grained GPU Memory SafetyabstractThe importance of ensuring the robustness of GPU systems has grown significantly, especially as GPUs have become vital in critical decision-making systems such as autonomous driving and medical diagnostics. However, GPU programming languages, primarily based on $\mathrm{C} / \mathrm{C}++$, inherit memory vulnerabilities that threaten the robustness of GPU applications. The heterogeneous GPU memory hierarchy makes it more difficult to find effective universal solutions. While several studies have proposed advanced GPU memory safety mechanisms, they still grapple with significant challenges, including substantial metadata storage and access overhead, elevated hardware implementation costs, and limited security coverage, particularly regarding fine-grained memory safety. We address this issue with Let-Me-In (LMI), a fine-grained memory safety mechanism specifically designed for GPUs. LMI features an efficient hardware bounds-checking mechanism that ensures negligible impact on performance and hardware costs, even in scenarios where thousands of concurrent threads perform memory operations across buffers in heap and local memory. This is achieved by aligning memory allocation to powers of two and performing static analysis to identify and mark pointer arithmetic instructions. This approach also enables storing metadata inside the unused upper bits of pointers, which are shrinking due to the expansion of the virtual memory address space. The unique characteristics of GPU programs make this approach feasible, unlike in CPU programs, where the inherent complexity of programs poses challenges. Our evaluation shows that LMI incurs only negligible hardware and performance overhead, making it a practical and efficient solution for enhancing GPU memory safety. Euijun Chung, Seonjin Na, Yonghae Kim, Jaekyu Lee, Hyesoon Kim |
HPCA | 6 |
| 2025 | RV-CURE: A RISC-V Capability Architecture for Full Memory SafetyabstractMemory-safety violations remain persistent in the real world. Although a tagged-pointer concept has demonstrated significant practical potential, prior work has shown scalability limitations in both performance and security.In this paper, we revisit the tagged-pointer design based on our observation that a pointer tag, stored in a pointer address, can be associated with security metadata and used as a hash to look up a hash table that stores associated metadata. To realize our idea as a new tagging-based memory-capability model, we investigate a hardware-software co-design approach. First, we develop a generalized tagging method, data-pointer tagging (DPT), to ensure full memory safety. DPT assigns a 16-bit tag to each memory object and associates that tag with the object’s capability metadata. On a memory access, DPT then performs a capability check using its associated metadata and validates the access. Furthermore, we design a RISC-V capability architecture, RV-CURE, that implements hardware extensions for DPT and thus enables robust, efficient capability enforcement. Altogether, we prototype a RISC-V evaluation framework, in which we launch FPGA instances running the Linux OS and conduct a full-system simulation. Our evaluation shows that RV-CURE imposes 9.5–19.6% runtime overhead for the SPEC 2017 C/C++ workloads while ensuring strong memory safety. Yonghae Kim, Anurag Kar, Jaekyu Lee, Hyesoon Kim |
IEEE Trans. Computers | 4 |
| 2022 | Securing GPU via region-based bounds checkingabstractGraphics processing units (GPUs) have become essential general-purpose computing platforms to accelerate a wide range of workloads, such as deep learning, scientific, and high-performance computing (HPC) applications. However, recent memory corruption attacks, such as buffer overflow, exposed security vulnerabilities in GPUs. We demonstrate that out-of-bounds writes are reproducible on an Nvidia GPU, which can enable other security attacks. Yonghae Kim, Jiashen Cao, Euna Kim, Jaekyu Lee, Hyesoon Kim |
ISCA | 5 |
| 2022 | Microarchitectural Performance Evaluation of AV1 Video Encoding WorkloadsabstractVideo encoding/decoding is an extremely relevant workload in our society today. Videos account for a significant percentage of the world’s network traffic, which is expected only to be growing. Thus, it is important to understand these workloads to optimize the hardware to handle them better.This paper explores the reasons for the large runtimes taken by AV1 encoding workloads. We discover that the runtime of the SVT-AV1 encoder is significantly higher than other encoders because it requires a larger number of instructions to encode the same video, rather than any significant microarchitectural inefficiencies. We also compare the thread scaling of SVT-AV1 against other codecs and observe that SVT-AV1 contains the highest degree of parallelism of the tested encoders. Steffen Jensen, Jaekyu Lee, Dam Sunwoo, Matthew Horsnell, Lizy Kurian John |
ISPASS | 2 |
| 2022 | Channel Sampler in Hyperspectral Images for Vehicle DetectionabstractSince hyperspectral images (HSIs) contain visual information of multiple wavelengths, invisible signals to human eyes can also be detected. Therefore, it can be widely used for target object detection in bad weather and disaster environments. However, the channel dimension of the HSI is very large, and thus it is very inefficient to apply the existing object detector naively. In this letter, we present a lightweight convolutional neural network (CNN)-based channel sampler to estimate the importance score of each channel in the HSI. Based on the importance score of each channel, we can generate single-channel images that achieve the best object detection performance, as well as analyze the impact of the wavelength in the HSI on object detection performance. The proposed sampler is trained by a self-supervised adversarial learning method that recovers the original input HSI from the generated single-channel image. Therefore, our channel sampler can be seamlessly combined with any existing detectors. For experiments, we build a hyperspectral dataset for vehicle detection and then show the effectiveness of our method through various ablation studies. Geonsoo Lee, Jaekyu Lee, Jeonghyun Baek, Hoseong Kim, Donghyeon Cho |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2021 | Re-establishing Fetch-Directed Instruction Prefetching: An Industry PerspectiveabstractInstruction prefetching can play a pivotal role in improving the performance of workloads with large instruction footprints and frequent, costly frontend stalls. In particular, Fetch Directed Prefetching (FDP) is an effective technique to mitigate frontend stalls since it leverages existing branch prediction resources in a processor and incurs very little hardware overhead. Modern processors have been trending towards provisioning more frontend resources, which bodes well for FDP as it requires these resources to be effective. However, recent academic research has been using outdated and less than optimal frontend baselines that employ smaller structures, resulting in equivocal outcomes. This paper presents a detailed FDP microarchitecture and evaluates two improvements, better branch history management and post-fetch correction. Our mechanism provides a 41.0% speedup over the baseline (no prefetching, no FDP) with only 195 bytes of hardware overhead and outperforms the 1st Instruction Prefetching Championship (IPC-1) winners that had a 128KB storage budget. We believe that our FDP-based frontend design can serve as a new reference baseline for instruction prefetching research to bridge the gap between academia and industry. Yasuo Ishii, Jaekyu Lee, Krishnendra Nathella, Dam Sunwoo |
ISPASS | 2 |
| 2020 | Hardware-based Always-On Heap Memory SafetyabstractMemory safety violations, caused by illegal use of pointers in unsafe programming languages such as C and C++, have been a major threat to modern computer systems. However, implementing a low-overhead yet robust runtime memory safety solution is still challenging. Various hardware-based mechanisms have been proposed, but their significant hardware requirements have limited their feasibility, and their performance overhead is too high to be an always-on solution. In this paper, we propose AOS, a low-overhead always-on heap memory safety solution that implements a novel bounds-checking mechanism. We identify that the major challenges of existing bounds-checking approaches are 1) the extra instruction overhead for memory checking and metadata propagation and 2) the complex metadata addressing. To address these challenges, using Arm PA primitives, we leverage unused upper bits of a pointer to store a key and have it propagated along with the pointer address, eliminating propagation overhead. Then, we use the embedded key to index a hashed bounds table to achieve efficient metadata management. We also introduce a micro-architectural unit to remove the need for memory checking instructions. We show that AOS overcomes all the aforementioned challenges and demonstrate its feasibility as an efficient runtime memory safety solution. Our evaluation for SPEC 2006 workloads shows an 8.4% performance overhead on average. Yonghae Kim, Jaekyu Lee, Hyesoon Kim |
MICRO | 2 |
| 2020 | Securing Branch Predictors with Two-Level EncryptionabstractModern processors rely on various speculative mechanisms to meet performance demand. Branch predictors are one of the most important micro-architecture components to deliver performance. However, they have been under heavy scrutiny because of recent side-channel attacks. Branch predictors are indexed using the PC and recent branch histories. An adversary can manipulate these parameters to access and control the same branch predictor entry that a victim uses. Recent Spectre attacks exploit this to set up speculative-execution-based security attacks. In this article, we aim to mitigate branch predictor side-channels using two-level encryption. At the first level, we randomize the set-index by encrypting the PC using a per-context secret key. At the second level, we encrypt the data in each branch predictor entry. While periodic key changes make the branch predictor more secure, performance degradation can be significant. To alleviate performance degradation, we propose a practical set update mechanism that also considers parallelism in multi-banked branch predictors. We show that our mechanism exhibits only 1.0% and 0.2% performance degradation while changing keys every 10K and 50K cycles, respectively, which is much lower than other state-of-the-art approaches. Jaekyu Lee, Yasuo Ishii, Dam Sunwoo |
ACM Trans. Archit. Code Optim. | 1 |
| 2019 | TwohandsMusic: Multitask Learning-Based Egocentric Piano-Playing Gesture Recognition System for Two HandsabstractWe present TwohandsMusic, a new real-time system for recognizing egocentric piano-playing gestures on planar objects by using a depth camera. Existing methods have usually recognized single tap gestures of one hand using a sensor installed in front of or under the user's hand. In contrast, we consider recognizing multi-tap gestures of both hands using a depth camera installed near the user's head. Our approach consists of two steps: hand detection and gesture recognition. At the hand detection step, we detect both hands using a 2DCNN (Convolutional Neural Network), called SegNet, and generate cropped hand images, which is to be used in the next step. In the gesture recognition step, we estimate 3D hand poses and classify multi-tap gestures simultaneously using a 3DCNN with multitask learning, called MusicNet. For training and validating of our system, we collect 85K dataset including tapping chords and show improved results over state-of-the-art methods. Kyeongeun Seo, Hyeonjoong Cho, Daewoong Choi, Sangyub Lee 0001, Jaekyu Lee, Jae-Jin Ko |
ICIP | 5 |
| 2015 | GREEN Cache: Exploiting the Disciplined Memory Model of OpenCL on GPUsabstractAs various graphics processing unit architectures are deployed across broad computing spectrum from a hand-held or embedded device to a high-performance computing server, OpenCL becomes the de facto standard programming environment for general-purpose computing on graphics processing units. Unlike its CPU counterpart, OpenCL has several distinct features such as its disciplined memory model, which is partially inherited from conventional 3D graphics programming models. On the other hand, due to ever increasing memory bandwidth pressure and low power requirement, the capacity of on-chip caches in GPUs keeps increasing overtime. Given such trends, we believe that we have interesting programming model/architecture co-optimization opportunities, in particular, how to energy-efficiently utilize large on-chip caches for GPUs. In this paper, as a showcase, we study the characteristics of the OpenCL memory model and propose a technique called GPU Region-aware energy-efficient non-inclusive cache hierarchy, or GREEN cache hierarchy. With the GREEN cache, our simulation results show that we can save 56 percent of dynamic energy in the L1 cache, 39 percent of dynamic energy in the L2 cache, and 50 percent of leakage energy in the L2 cache with practically no performance degradation and off-chip access increases. Jaekyu Lee, Dong Hyuk Woo, Hyesoon Kim, Mani Azimi |
IEEE Trans. Computers | 1 |
| 2013 | Design space exploration of on-chip ring interconnection for a CPU-GPU heterogeneous architecture
Jaekyu Lee, Hyesoon Kim, Sudhakar Yalamanchili |
J. Parallel Distributed Comput. | 1 |
| 2013 | Adaptive virtual channel partitioning for network-on-chip in heterogeneous architecturesabstractCurrent heterogeneous chip-multiprocessors (CMPs) integrate a GPU architecture on a die. However, the heterogeneity of this architecture inevitably exerts different pressures on shared resource management due to differing characteristics of CPU and GPU cores. We consider how to efficiently share on-chip resources between cores within the heterogeneous system, in particular the on-chip network. Heterogeneous architectures use an on-chip interconnection network to access shared resources such as last-level cache tiles and memory controllers, and this type of on-chip network will have a significant impact on performance. In this article, we propose a feedback-directed virtual channel partitioning (VCP) mechanism for on-chip routers to effectively share network bandwidth between CPU and GPU cores in a heterogeneous architecture. VCP dedicates a few virtual channels to CPU and GPU applications with separate injection queues. The proposed mechanism balances on-chip network bandwidth for applications running on CPU and GPU cores by adaptively choosing the best partitioning configuration. As a result, our mechanism improves system throughput by 15% over the baseline across 39 heterogeneous workloads. Jaekyu Lee, Hyesoon Kim, Sudhakar Yalamanchili |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2012 | TAP: A TLP-aware cache management policy for a CPU-GPU heterogeneous architectureabstractCombining CPUs and GPUs on the same chip has become a popular architectural trend. However, these heterogeneous architectures put more pressure on shared resource management. In particular, managing the last-level cache (LLC) is very critical to performance. Lately, many researchers have proposed several shared cache management mechanisms, including dynamic cache partitioning and promotion-based cache management, but no cache management work has been done on CPU-GPU heterogeneous architectures. Sharing the LLC between CPUs and GPUs brings new challenges due to the different characteristics of CPU and GPGPU applications. Unlike most memory-intensive CPU benchmarks that hide memory latency with caching, many GPGPU applications hide memory latency by combining thread-level parallelism (TLP) and caching. In this paper, we propose a TLP-aware cache management policy for CPU-GPU heterogeneous architectures. We introduce a core-sampling mechanism to detect how caching affects the performance of a GPGPU application. Inspired by previous cache management schemes, Utility-based Cache Partitioning (UCP) and Re-Reference Interval Prediction (RRIP), we propose two new mechanisms: TAP-UCP and TAP-RRIP. TAP-UCP improves performance by 5% over UCP and 11% over LRU on 152 heterogeneous workloads, and TAP-RRIP improves performance by 9% over RRIP and 12% over LRU. Jaekyu Lee, Hyesoon Kim |
HPCA | 1 |
| 2012 | FLEXclusion: Balancing cache capacity and on-chip bandwidth via Flexible ExclusionabstractExclusive last-level caches (LLCs) reduce memory accesses by effectively utilizing cache capacity. However, they require excessive on-chip bandwidth to support frequent insertions of cache lines on eviction from upper-level caches. Non-inclusive caches, on the other hand, have the advantage of using the on-chip bandwidth more effectively but suffer from a higher miss rate. Traditionally, the decision to use the cache as exclusive or non-inclusive is made at design time. However, the best option for a cache organization depends on application characteristics, such as working set size and the amount of traffic consumed by LLC insertions. This paper proposes FLEXclusion, a design that dynamically selects between exclusion and non-inclusion depending on workload behavior. With FLEXclusion, the cache behaves like an exclusive cache when the application benefits from extra cache capacity, and it acts as a non-inclusive cache when additional cache capacity is not useful, so that it can reduce on-chip bandwidth. FLEXclusion leverages the observation that both non-inclusion and exclusion rely on similar hardware support, so our proposal can be implemented with negligible hardware changes. Our evaluations show that a FLEXclusive cache reduces the on-chip LLC insertion traffic by 72.6% compared to an exclusive design and improves performance by 5.9% compared to a non-inclusive design. Jaewoong Sim, Jaekyu Lee, Moinuddin K. Qureshi, Hyesoon Kim |
ISCA | 2 |
| 2012 | When Prefetching Works, When It Doesn't, and WhyabstractIn emerging and future high-end processor systems, tolerating increasing cache miss latency and properly managing memory bandwidth will be critical to achieving high performance. Prefetching, in both hardware and software, is among our most important available techniques for doing so; yet, we claim that prefetching is perhaps also the least well-understood. Thus, the goal of this study is to develop a novel, foundational understanding of both the benefits and limitations of hardware and software prefetching. Our study includes: source code-level analysis, to help in understanding the practical strengths and weaknesses of compiler- and software-based prefetching; a study of the synergistic and antagonistic effects between software and hardware prefetching; and an evaluation of hardware prefetching training policies in the presence of software prefetching requests. We use both simulation and measurement on real systems. We find, for instance, that although there are many opportunities for compilers to prefetch much more aggressively than they currently do, there is also a tangible risk of interference with training existing hardware prefetching mechanisms. Taken together, our observations suggest new research directions for cooperative hardware/software prefetching. Jaekyu Lee, Hyesoon Kim, Richard W. Vuduc |
ACM Trans. Archit. Code Optim. | 1 |
| 2010 | Many-Thread Aware Prefetching Mechanisms for GPGPU ApplicationsabstractWe consider the problem of how to improve memory latency tolerance in massively multithreaded GPGPUs when the thread-level parallelism of an application is not sufficient to hide memory latency. One solution used in conventional CPU systems is prefetching, both in hardware and software. However, we show that straightforwardly applying such mechanisms to GPGPU systems does not deliver the expected performance benefits and can in fact hurt performance when not used judiciously. This paper proposes new hardware and software prefetching mechanisms tailored to GPGPU systems, which we refer to as many-thread aware prefetching (MT-prefetching) mechanisms. Our software MT-prefetching mechanism, called inter-thread prefetching, exploits the existence of common memory access behavior among fine-grained threads. For hardware MT-prefetching, we describe a scalable prefetcher training algorithm along with a hardware-based inter-thread prefetching mechanism. In some cases, blindly applying prefetching degrades performance. To reduce such negative effects, we propose an adaptive prefetch throttling scheme, which permits automatic GPGPU application- and hardware-specific adjustment. We show that adaptation reduces the negative effects of prefetching and can even improve performance. Overall, compared to the state-of-the-art software and hardware prefetching, our MT-prefetching improves performance on average by 16%(software pref.)/15% (hardware pref.) on our benchmarks. Jaekyu Lee, Nagesh B. Lakshminarayana, Hyesoon Kim, Richard W. Vuduc |
MICRO | 1 |
| 2009 | Age based scheduling for asymmetric multiprocessorsabstractAsymmetric (or Heterogeneous) Multiprocessors are becoming popular in the current era of multi-cores due to their power efficiency and potential performance and energy efficiency. However, scheduling of multithreaded applications in Asymmetric Multiprocessors is still a challenging problem. Scheduling algorithms for Asymmetric Multiprocessors must not only be aware of asymmetry in processor performance, but have to consider the characteristics of application threads also. Nagesh B. Lakshminarayana, Jaekyu Lee, Hyesoon Kim |
SC | 2 |