VLDB 2026 Research / reviewers in the wild / expert
Yoongu Kim
dblp:70/4287
· DBLP profile ↗
14ranked-venue papers
4as first author
1since 2021 · last 2022
0000-0002-7062-3356ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 4 first-authorSoftware engineering, systems software and programming languages · 4 · 2 first-authorSecurity and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
13 papers |
Memory systems · 82% Performance modeling and evaluation · 8% Hardware reliability and fault tolerance · 6% | |
| Network and information security
1 paper |
Hardware security and side channels · 100% |
Topics — the 26 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
DRAM |
2.0 | 9 | 2022 | Half-Double: Hammering From the Next Row Over · USENIX Security Symposium 2022 Adaptive-latency DRAM: Optimizing DRAM timing for the common-case · HPCA 2015 The efficacy of error mitigation techniques for DRAM retention failures: a comparative experimental study · SIGMETRICS 2014 |
Memory systems › DRAM
rowhammer |
0.8 | 2 | 2022 | Half-Double: Hammering From the Next Row Over · USENIX Security Symposium 2022 Flipping bits in memory without accessing them: An experimental study of DRAM disturbance errors · ISCA 2014 |
Hardware security and side channels › fault attacks › fault injection attack
rowhammer attack |
0.6 | 1 | 2022 | Half-Double: Hammering From the Next Row Over · USENIX Security Symposium 2022 |
Memory systems › memory controller
memory scheduling |
0.4 | 4 | 2013 | MISE: Providing performance predictability and improving fairness in shared main memory systems · HPCA 2013 Thread Cluster Memory Scheduling: Exploiting Differences in Memory Access Behavior · MICRO 2010 ATLAS: A scalable and high-performance scheduling algorithm for multiple memory controllers · HPCA 2010 |
Memory systems
memory controller |
0.3 | 2 | 2013 | Linearly compressed pages: a low-complexity, low-latency main memory compression framework · MICRO 2013 ATLAS: A scalable and high-performance scheduling algorithm for multiple memory controllers · HPCA 2010 |
Performance modeling and evaluation › workload characterization
memory characterization |
0.2 | 1 | 2015 | Adaptive-latency DRAM: Optimizing DRAM timing for the common-case · HPCA 2015 |
Hardware reliability and fault tolerance
error mitigation |
0.2 | 1 | 2014 | The efficacy of error mitigation techniques for DRAM retention failures: a comparative experimental study · SIGMETRICS 2014 |
Hardware reliability and fault tolerance
memory reliability |
0.2 | 1 | 2014 | Flipping bits in memory without accessing them: An experimental study of DRAM disturbance errors · ISCA 2014 |
Memory systems › DRAM › refresh management
refresh scheduling |
0.2 | 1 | 2014 | Improving DRAM performance by parallelizing refreshes with accesses · HPCA 2014 |
Memory systems › memory compression
cache compression |
0.2 | 1 | 2013 | Linearly compressed pages: a low-complexity, low-latency main memory compression framework · MICRO 2013 |
Memory systems › DRAM
DRAM architecture |
0.2 | 1 | 2013 | Tiered-latency DRAM: A low latency and low cost DRAM architecture · HPCA 2013 |
Memory systems › DRAM
DRAM refresh |
0.2 | 1 | 2013 | An experimental study of data retention behavior in modern DRAM devices: implications for retention time profiling mechanisms · ISCA 2013 |
Memory systems › DRAM › DRAM latency reduction
low-latency DRAM |
0.2 | 1 | 2013 | Tiered-latency DRAM: A low latency and low cost DRAM architecture · HPCA 2013 |
Memory systems › memory compression
main memory compression |
0.2 | 1 | 2013 | Linearly compressed pages: a low-complexity, low-latency main memory compression framework · MICRO 2013 |
Memory systems
processing-in-memory |
0.2 | 1 | 2013 | RowClone: fast and energy-efficient in-DRAM bulk data copy and initialization · MICRO 2013 |
Performance modeling and evaluation › performance prediction
slowdown estimation |
0.2 | 1 | 2013 | MISE: Providing performance predictability and improving fairness in shared main memory systems · HPCA 2013 |
Memory systems › DRAM › DRAM microarchitecture
subarray-level parallelism |
0.1 | 1 | 2012 | A case for exploiting subarray-level parallelism (SALP) in DRAM · ISCA 2012 |
Performance modeling and evaluation
workload characterization |
0.1 | 2 | 2013 | An experimental study of data retention behavior in modern DRAM devices: implications for retention time profiling mechanisms · ISCA 2013 ATLAS: A scalable and high-performance scheduling algorithm for multiple memory controllers · HPCA 2010 |
Energy-efficient computing › memory energy efficiency
DRAM energy efficiency |
0.1 | 1 | 2014 | Improving DRAM performance by parallelizing refreshes with accesses · HPCA 2014 |
Electronic design automation › hardware verification and test
memory testing |
0.1 | 1 | 2014 | The efficacy of error mitigation techniques for DRAM retention failures: a comparative experimental study · SIGMETRICS 2014 |
Memory systems
cache management |
0.0 | 1 | 2013 | Tiered-latency DRAM: A low latency and low cost DRAM architecture · HPCA 2013 |
Memory systems › cache
DRAM cache |
0.0 | 1 | 2013 | Tiered-latency DRAM: A low latency and low cost DRAM architecture · HPCA 2013 |
Energy-efficient computing
memory energy efficiency |
0.0 | 1 | 2013 | RowClone: fast and energy-efficient in-DRAM bulk data copy and initialization · MICRO 2013 |
Processor architecture and microarchitecture
chip multiprocessor |
0.0 | 1 | 2010 | Thread Cluster Memory Scheduling: Exploiting Differences in Memory Access Behavior · MICRO 2010 |
Memory systems › memory interference
memory contention |
0.0 | 1 | 2010 | Thread Cluster Memory Scheduling: Exploiting Differences in Memory Access Behavior · MICRO 2010 |
Performance modeling and evaluation › workload characterization
multiprogrammed workloads |
0.0 | 1 | 2010 | ATLAS: A scalable and high-performance scheduling algorithm for multiple memory controllers · HPCA 2010 |
Methods — techniques the papers use, named apart from their topics
FPGA-based testing platform · 0.4runtime timing reconfiguration · 0.2workload evaluation · 0.2simulation · 0.2experimental study · 0.2experimental characterization · 0.2software-managed cache · 0.2isolation transistor · 0.2hardware-managed cache · 0.2bitline segmentation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Half-Double: Hammering From the Next Row Over
Andreas Kogler, Jonas Juffinger, Salman Qazi, Yoongu Kim, Moritz Lipp, Nicolas Boichat, Eric Shiu, Mattias Nissler, Daniel Gruss |
USENIX Security Symposium | 4 |
| 2015 | Adaptive-latency DRAM: Optimizing DRAM timing for the common-caseabstractIn current systems, memory accesses to a DRAM chip must obey a set of minimum latency restrictions specified in the DRAM standard. Such timing parameters exist to guarantee reliable operation. When deciding the timing parameters, DRAM manufacturers incorporate a very large margin as a provision against two worst-case scenarios. First, due to process variation, some outlier chips are much slower than others and cannot be operated as fast. Second, chips become slower at higher temperatures, and all chips need to operate reliably at the highest supported (i.e., worst-case) DRAM temperature (85° C). In this paper, we show that typical DRAM chips operating at typical temperatures (e.g., 55° C) are capable of providing a much smaller access latency, but are nevertheless forced to operate at the largest latency of the worst-case. Our goal in this paper is to exploit the extra margin that is built into the DRAM timing parameters to improve performance. Using an FPGA-based testing platform, we first characterize the extra margin for 115 DRAM modules from three major manufacturers. Our results demonstrate that it is possible to reduce four of the most critical timing parameters by a minimum/maximum of 17.3%/54.8% at 55°C without sacrificing correctness. Based on this characterization, we propose Adaptive-Latency DRAM (AL-DRAM), a mechanism that adoptively reduces the timing parameters for DRAM modules based on the current operating condition. AL-DRAM does not require any changes to the DRAM chip or its interface. We evaluate AL-DRAM on a real system that allows us to reconfigure the timing parameters at runtime. We show that AL-DRAM improves the performance of memory-intensive workloads by an average of 14% without introducing any errors. We discuss and show why AL-DRAM does not compromise reliability. We conclude that dynamically optimizing the DRAM timing parameters can reliably improve system performance. Donghyuk Lee, Yoongu Kim, Gennady Pekhimenko, Samira Manabi Khan, Vivek Seshadri, Kevin Kai-Wei Chang, Onur Mutlu |
HPCA | 2 |
| 2014 | Improving DRAM performance by parallelizing refreshes with accessesabstractModern DRAM cells are periodically refreshed to prevent data loss due to leakage. Commodity DDR (double data rate) DRAM refreshes cells at the rank level. This degrades performance significantly because it prevents an entire DRAM rank from serving memory requests while being refreshed. DRAM designed for mobile platforms, LPDDR (low power DDR) DRAM, supports an enhanced mode, called per-bank refresh, that refreshes cells at the bank level. This enables a bank to be accessed while another in the same rank is being refreshed, alleviating part of the negative performance impact of refreshes. Unfortunately, there are two shortcomings of per-bank refresh employed in today's systems. First, we observe that the perbank refresh scheduling scheme does not exploit the full potential of overlapping refreshes with accesses across banks because it restricts the banks to be refreshed in a sequential round-robin order. Second, accesses to a bank that is being refreshed have to wait. To mitigate the negative performance impact of DRAM refresh, we propose two complementary mechanisms, DARP (Dynamic Access Refresh Parallelization) and SARP (Subarray Access Refresh Parallelization). The goal is to address the drawbacks of per-bank refresh by building more efficient techniques to parallelize refreshes and accesses within DRAM. First, instead of issuing per-bank refreshes in a round-robin order, as it is done today, DARP issues per-bank refreshes to idle banks in an out-of-order manner. Furthermore, DARP proactively schedules refreshes during intervals when a batch of writes are draining to DRAM. Second, SARP exploits the existence of mostly-independent subarrays within a bank. With minor modifications to DRAM organization, it allows a bank to serve memory accesses to an idle subarray while another subarray is being refreshed. Extensive evaluations on a wide variety of workloads and systems show that our mechanisms improve system performance (and energy efficiency) compared to three state-of-the-art refresh policies and the performance benefit increases as DRAM density increases. Kevin Kai-Wei Chang, Donghyuk Lee, Zeshan Chishti, Alaa R. Alameldeen, Chris Wilkerson, Yoongu Kim, Onur Mutlu |
HPCA | 6 |
| 2014 | Flipping bits in memory without accessing them: An experimental study of DRAM disturbance errorsabstractMemory isolation is a key property of a reliable and secure computing system-an access to one memory address should not have unintended side effects on data stored in other addresses. However, as DRAM process technology scales down to smaller dimensions, it becomes more difficult to prevent DRAM cells from electrically interacting with each other. In this paper, we expose the vulnerability of commodity DRAM chips to disturbance errors. By reading from the same address in DRAM, we show that it is possible to corrupt data in nearby addresses. More specifically, activating the same row in DRAM corrupts data in nearby rows. We demonstrate this phenomenon on Intel and AMD systems using a malicious program that generates many DRAM accesses. We induce errors in most DRAM modules (110 out of 129) from three major DRAM manufacturers. From this we conclude that many deployed systems are likely to be at risk. We identify the root cause of disturbance errors as the repeated toggling of a DRAM row's wordline, which stresses inter-cell coupling effects that accelerate charge leakage from nearby rows. We provide an extensive characterization study of disturbance errors and their behavior using an FPGA-based testing platform. Among our key findings, we show that (i) it takes as few as 139K accesses to induce an error and (ii) up to one in every 1.7K cells is susceptible to errors. After examining various potential ways of addressing the problem, we propose a low-overhead solution to prevent the errors. Yoongu Kim, Ross Daly, Jeremie S. Kim, Chris Fallin, Ji-Hye Lee, Donghyuk Lee, Chris Wilkerson, Konrad Lai, Onur Mutlu |
ISCA | 1 |
| 2014 | The efficacy of error mitigation techniques for DRAM retention failures: a comparative experimental studyabstractAs DRAM cells continue to shrink, they become more susceptible to retention failures. DRAM cells that permanently exhibit short retention times are fairly easy to identify and repair through the use of memory tests and row and column redundancy. However, the retention time of many cells may vary over time due to a property called Variable Retention Time (VRT). Since these cells intermittently transition between failing and non-failing states, they are particularly difficult to identify through memory tests alone. In addition, the high temperature packaging process may aggravate this problem as the susceptibility of cells to VRT increases after the assembly of DRAM chips. A promising alternative to manufacture-time testing is to detect and mitigate retention failures after the system has become operational. Such a system would require mechanisms to detect and mitigate retention failures in the field, but would be responsive to retention failures introduced after system assembly and could dramatically reduce the cost of testing, enabling much longer tests than are practical with manufacturer testing equipment. Samira Manabi Khan, Donghyuk Lee, Yoongu Kim, Alaa R. Alameldeen, Chris Wilkerson, Onur Mutlu |
SIGMETRICS | 3 |
| 2013 | Tiered-latency DRAM: A low latency and low cost DRAM architectureabstractThe capacity and cost-per-bit of DRAM have historically scaled to satisfy the needs of increasingly large and complex computer systems. However, DRAM latency has remained almost constant, making memory latency the performance bottleneck in today's systems. We observe that the high access latency is not intrinsic to DRAM, but a trade-off made to decrease cost-per-bit. To mitigate the high area overhead of DRAM sensing structures, commodity DRAMs connect many DRAM cells to each sense-amplifier through a wire called a bitline. These bitlines have a high parasitic capacitance due to their long length, and this bitline capacitance is the dominant source of DRAM latency. Specialized low-latency DRAMs use shorter bitlines with fewer cells, but have a higher cost-per-bit due to greater sense-amplifier area overhead. In this work, we introduce Tiered-Latency DRAM (TL-DRAM), which achieves both low latency and low cost-per-bit. In TL-DRAM, each long bitline is split into two shorter segments by an isolation transistor, allowing one segment to be accessed with the latency of a short-bitline DRAM without incurring high cost-per-bit. We propose mechanisms that use the low-latency segment as a hardware-managed or software-managed cache. Evaluations show that our proposed mechanisms improve both performance and energy-efficiency for both single-core and multi-programmed workloads. Donghyuk Lee, Yoongu Kim, Vivek Seshadri, Jamie Liu, Lavanya Subramanian, Onur Mutlu |
HPCA | 2 |
| 2013 | MISE: Providing performance predictability and improving fairness in shared main memory systemsabstractApplications running concurrently on a multicore system interfere with each other at the main memory. This interference can slow down different applications differently. Accurately estimating the slow down of each application in such a system can enable mechanisms that can enforce quality-of-service. While much prior work has focused on mitigating the performance degradation due to inter-application interference, there is little work on estimating slow down of individual applications in a multi-programmed environment. Our goal in this work is to build such an estimation scheme. To this end, we present our simple Memory-Interference-induced Slowdown Estimation (MISE) model that estimates slowdowns caused by memory interference. We build our model based on two observations. First, the performance of a memory-bound application is roughly proportional to the rate at which its memory requests are served, suggesting that request-service-rate can be used as a proxy for performance. Second, when an application's requests are prioritized over all other applications' requests, the application experiences very little interference from other applications. This provides a means for estimating the uninterfered request-service-rate of an application while it is run alongside other applications. Using the above observations, our model estimates the slowdown of an application as the ratio of its uninterfered and interfered request service rates. We propose simple changes to the above model to estimate the slowdown of non-memory-bound applications. We demonstrate the effectiveness of our model by developing two new memory scheduling schemes: 1) one that provides soft quality-of-service guarantees and 2) another that explicitly attempts to minimize maximum slowdown (i.e., unfairness) in the system. Evaluations show that our techniques perform significantly better than state-of-the-art memory scheduling approaches to address the above problems. Lavanya Subramanian, Vivek Seshadri, Yoongu Kim, Ben Jaiyen, Onur Mutlu |
HPCA | 3 |
| 2013 | An experimental study of data retention behavior in modern DRAM devices: implications for retention time profiling mechanismsabstractDRAM cells store data in the form of charge on a capacitor. This charge leaks off over time, eventually causing data to be lost. To prevent this data loss from occurring, DRAM cells must be periodically refreshed. Unfortunately, DRAM refresh operations waste energy and also degrade system performance by interfering with memory requests. These problems are expected to worsen as DRAM density increases. Jamie Liu, Ben Jaiyen, Yoongu Kim, Chris Wilkerson, Onur Mutlu |
ISCA | 3 |
| 2013 | Linearly compressed pages: a low-complexity, low-latency main memory compression frameworkabstractData compression is a promising approach for meeting the increasing memory capacity demands expected in future systems. Unfortunately, existing compression algorithms do not translate well when directly applied to main memory because they require the memory controller to perform non-trivial computation to locate a cache line within a compressed memory page, thereby increasing access latency and degrading system performance. Prior proposals for addressing this performance degradation problem are either costly or energy inefficient. Gennady Pekhimenko, Vivek Seshadri, Yoongu Kim, Hongyi Xin, Onur Mutlu, Phillip B. Gibbons, Michael A. Kozuch, Todd C. Mowry |
MICRO | 3 |
| 2013 | RowClone: fast and energy-efficient in-DRAM bulk data copy and initializationabstractSeveral system-level operations trigger bulk data copy or initialization. Even though these bulk data operations do not require any computation, current systems transfer a large quantity of data back and forth on the memory channel to perform such operations. As a result, bulk data operations consume high latency, bandwidth, and energy--degrading both system performance and energy efficiency. Vivek Seshadri, Yoongu Kim, Chris Fallin, Donghyuk Lee, Rachata Ausavarungnirun, Gennady Pekhimenko, Onur Mutlu, Phillip B. Gibbons, Michael A. Kozuch, Todd C. Mowry |
MICRO | 2 |
| 2012 | A case for exploiting subarray-level parallelism (SALP) in DRAMabstractModern DRAMs have multiple banks to serve multiple memory requests in parallel. However, when two requests go to the same bank, they have to be served serially, exacerbating the high latency of off-chip memory. Adding more banks to the system to mitigate this problem incurs high system cost. Our goal in this work is to achieve the benefits of increasing the number of banks with a low cost approach. To this end, we propose three new mechanisms that overlap the latencies of different requests that go to the same bank. The key observation exploited by our mechanisms is that a modern DRAM bank is implemented as a collection of subarrays that operate largely independently while sharing few global peripheral structures. Our proposed mechanisms (SALP-1, SALP-2, and MASA) mitigate the negative impact of bank serialization by overlapping different components of the bank access latencies of multiple requests that go to different subarrays within the same bank. SALP-1 requires no changes to the existing DRAM structure and only needs reinterpretation of some DRAM timing parameters. SALP-2 and MASA require only modest changes (<;0.15% area overhead) to the DRAM peripheral structures, which are much less design constrained than the DRAM core. Evaluations show that all our schemes significantly improve performance for both single-core systems and multi-core systems. Our schemes also interact positively with application-aware memory request scheduling in multi-core systems. Yoongu Kim, Vivek Seshadri, Donghyuk Lee, Jamie Liu, Onur Mutlu |
ISCA | 1 |
| 2010 | ATLAS: A scalable and high-performance scheduling algorithm for multiple memory controllersabstractModern chip multiprocessor (CMP) systems employ multiple memory controllers to control access to main memory. The scheduling algorithm employed by these memory controllers has a significant effect on system throughput, so choosing an efficient scheduling algorithm is important. The scheduling algorithm also needs to be scalable - as the number of cores increases, the number of memory controllers shared by the cores should also increase to provide sufficient bandwidth to feed the cores. Unfortunately, previous memory scheduling algorithms are inefficient with respect to system throughput and/or are designed for a single memory controller and do not scale well to multiple memory controllers, requiring significant finegrained coordination among controllers. This paper proposes ATLAS (Adaptive per-Thread Least-Attained-Service memory scheduling), a fundamentally new memory scheduling technique that improves system throughput without requiring significant coordination among memory controllers. The key idea is to periodically order threads based on the service they have attained from the memory controllers so far, and prioritize those threads that have attained the least service over others in each period. The idea of favoring threads with least-attained-service is borrowed from the queueing theory literature, where, in the context of a single-server queue it is known that least-attained-service optimally schedules jobs, assuming a Pareto (or any decreasing hazard rate) workload distribution. After verifying that our workloads have this characteristic, we show that our implementation of least-attained-service thread prioritization reduces the time the cores spend stalling and significantly improves system throughput. Furthermore, since the periods over which we accumulate the attained service are long, the controllers coordinate very infrequently to form the ordering of threads, thereby making ATLAS scalable to many controllers. We evaluate ATLAS on a wide variety of multiprogrammed SPEC 2006 workloads and systems with 4-32 cores and 1-16 memory controllers, and compare its performance to five previously proposed scheduling algorithms. Averaged over 32 workloads on a 24-core system with 4 controllers, ATLAS improves instruction throughput by 10.8%, and system throughput by 8.4%, compared to PAR-BS, the best previous CMP memory scheduling algorithm. ATLAS's performance benefit increases as the number of cores increases. Yoongu Kim, Dongsu Han, Onur Mutlu, Mor Harchol-Balter |
HPCA | 1 |
| 2010 | Thread Cluster Memory Scheduling: Exploiting Differences in Memory Access BehaviorabstractIn a modern chip-multiprocessor system, memory is a shared resource among multiple concurrently executing threads. The memory scheduling algorithm should resolve memory contention by arbitrating memory access in such a way that competing threads progress at a relatively fast and even pace, resulting in high system throughput and fairness. Previously proposed memory scheduling algorithms are predominantly optimized for only one of these objectives: no scheduling algorithm provides the best system throughput and best fairness at the same time. This paper presents a new memory scheduling algorithm that addresses system throughput and fairness separately with the goal of achieving the best of both. The main idea is to divide threads into two separate clusters and employ different memory request scheduling policies in each cluster. Our proposal, Thread Cluster Memory scheduling (TCM), dynamically groups threads with similar memory access behavior into either the latency-sensitive (memory-non-intensive) or the bandwidth-sensitive (memory-intensive) cluster. TCM introduces three major ideas for prioritization: 1) we prioritize the latency-sensitive cluster over the bandwidth-sensitive cluster to improve system throughput, 2) we introduce a ``niceness'' metric that captures a thread's propensity to interfere with other threads, 3) we use niceness to periodically shuffle the priority order of the threads in the bandwidth-sensitive cluster to provide fair access to each thread in a way that reduces inter-thread interference. On the one hand, prioritizing memory-non-intensive threads significantly improves system throughput without degrading fairness, because such ``light'' threads only use a small fraction of the total available memory bandwidth. On the other hand, shuffling the priority order of memory-intensive threads improves fairness because it ensures no thread is disproportionately slowed down or starved. We evaluate TCM on a wide variety of multiprogrammed workloads and compare its performance to four previously proposed scheduling algorithms, finding that TCM achieves both the best system throughput and fairness. Averaged over 96 workloads on a 24-core system with 4 memory channels, TCM improves system throughput and reduces maximum slowdown by 4.6%/38.6% compared to ATLAS (previous work providing the best system throughput) and 7.6%/4.6% compared to PAR-BS (previous work providing the best fairness). Yoongu Kim, Michael Papamichael, Onur Mutlu, Mor Harchol-Balter |
MICRO | 1 |
| 2004 | S-COI : The Secure Conflicts of Interest Model for Multilevel Secure Database Systems
Chanjung Park, Seog Park, Yoongu Kim |
DASFAA | 3 |