EDBT 2026 Demo / reviewers in the wild / expert
Kyle Kuan
dblp:218/1135
· DBLP profile ↗
5ranked-venue papers
5as first author
0since 2021 · last 2020
0000-0001-8964-1346ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Memory systems · 62% Energy-efficient computing · 38% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
cache |
0.8 | 2 | 2020 | Energy-Efficient Runtime Adaptable L1 STT-RAM Cache Design · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 HALLS: An Energy-Efficient Highly Adaptable Last Level STT-RAM Cache for Multicore Systems · IEEE Trans. Computers 2019 |
Energy-efficient computing › power management › memory power management
cache energy reduction |
0.8 | 2 | 2020 | Energy-Efficient Runtime Adaptable L1 STT-RAM Cache Design · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 HALLS: An Energy-Efficient Highly Adaptable Last Level STT-RAM Cache for Multicore Systems · IEEE Trans. Computers 2019 |
Energy-efficient computing
power management |
0.8 | 2 | 2020 | Energy-Efficient Runtime Adaptable L1 STT-RAM Cache Design · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 HALLS: An Energy-Efficient Highly Adaptable Last Level STT-RAM Cache for Multicore Systems · IEEE Trans. Computers 2019 |
Memory systems
non-volatile memory |
0.5 | 2 | 2020 | HALLS: An Energy-Efficient Highly Adaptable Last Level STT-RAM Cache for Multicore Systems · IEEE Trans. Computers 2019 Energy-Efficient Runtime Adaptable L1 STT-RAM Cache Design · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Memory systems › non-volatile memory › magnetic random access memory
STT-MRAM |
0.5 | 2 | 2020 | HALLS: An Energy-Efficient Highly Adaptable Last Level STT-RAM Cache for Multicore Systems · IEEE Trans. Computers 2019 Energy-Efficient Runtime Adaptable L1 STT-RAM Cache Design · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Memory systems › cache
STT-RAM cache |
0.4 | 1 | 2020 | Energy-Efficient Runtime Adaptable L1 STT-RAM Cache Design · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Memory systems › memory hierarchy › cache hierarchy
last-level cache |
0.4 | 1 | 2019 | HALLS: An Energy-Efficient Highly Adaptable Last Level STT-RAM Cache for Multicore Systems · IEEE Trans. Computers 2019 |
Methods — techniques the papers use, named apart from their topics
runtime adaptation · 0.4retention time tradeoff · 0.4simulation · 0.4runtime tuning · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | A Study of Runtime Adaptive Prefetching for STTRAM L1 CachesabstractSpin- Transfer Torque RAM (STTRAM) is a promising alternative to SRAM in on-chip caches due to several advantages. These advantages include non-volatility, low leakage, high integration density, and CMOS compatibility. Prior studies have shown that relaxing and adapting the STTRAM retention time to runtime application needs can substantially reduce overall cache energy without significant latency overheads, due to the lower STTRAM write energy and latency in shorter retention times. In this paper, as a first step towards efficient prefetching across the STTRAM cache hierarchy, we study prefetching in reduced retention STTRAM L1 caches. Using SPEC CPU 2017 benchmarks, we analyze the energy and latency impact of different prefetch distances in different STTRAM cache retention times for different applications. We show that expired _unused _prefetches← the number of unused prefetches expired by the reduced retention time STTRAM cache-can accurately determine the best retention time for energy consumption and access latency. This new metric can also provide insights into the best prefetch distance for memory bandwidth consumption and prefetch accuracy. Based on our analysis and insights, we propose Prefetch-Aware Retention time Tuning (PART) and Retention time-based Prefetch Control (RPC). Compared to a base STTRAM cache, PART and RPC collectively reduced the average cache energy and latency by 22.24 % and 24.59 %, respectively. When the base architecture was augmented with the state-of-the-art near-side prefetch throttling (NST), PART+RPC reduced the average cache energy and latency by 3.50 % and 3.59 %, respectively, and reduced the hardware overhead by 54.55 %. Kyle Kuan, Tosiron Adegbija |
ICCD | 1 |
| 2020 | Energy-Efficient Runtime Adaptable L1 STT-RAM Cache DesignabstractMuch research has shown that applications have variable runtime cache requirements. In the context of the increasingly popular spin-transfer torque RAM (STT-RAM) cache, the retention time, which defines how long the cache can retain a cache block in the absence of power, is one of the most important cache requirements that may vary for different applications. In this paper, we propose a logically adaptable retention time STT-RAM (LARS) cache that allows the retention time to be dynamically adapted to applications' runtime requirements. LARS cache comprises of multiple STT-RAM units with different retention times, with only one unit being used at a given time. LARS dynamically determines which STT-RAM unit to use during runtime, based on executing applications' needs. As an integral part of LARS, we also explore different algorithms to dynamically determine the best retention time based on different cache design tradeoffs. Our experiments show that by adapting the retention time to different applications' requirements, LARS cache can reduce the average cache energy by 25.31%, compared to prior work, with minimal overheads. Kyle Kuan, Tosiron Adegbija |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2019 | MirrorCache: An Energy-Efficient Relaxed Retention L1 STTRAM CacheabstractSpin-Transfer Torque RAM (STTRAM) is a promising alternative to SRAMs in on-chip caches, due to several advantages, including non-volatility, low leakage, high integration density, and CMOS compatibility. However, STTRAMs' wide adoption in resource-constrained systems is impeded, in part, by high write energy and latency. A popular approach to mitigating these overheads involves relaxing the STTRAM's retention time, in order to reduce the write latency and energy. However, this approach usually requires a dynamic refresh scheme to maintain cache blocks' data integrity beyond the retention time, and typically requires an external refresh buffer. In this paper, we propose mirrorCache---an energy-efficient, buffer-free refresh scheme. MirrorCache leverages the STTRAM cell's compact feature size, and uses an auxiliary segment with the same size as the logical cache size to handle the refresh operations without the overheads of an external refresh buffer. Our experiments show that, compared to prior work, mirrorCache can reduce the average cache energy by at least 39.7% for a variety of systems. Kyle Kuan, Tosiron Adegbija |
ACM Great Lakes Symposium on VLSI | 1 |
| 2019 | HALLS: An Energy-Efficient Highly Adaptable Last Level STT-RAM Cache for Multicore SystemsabstractSpin-Transfer Torque RAM (STT-RAM) is widely considered a promising alternative to SRAM in the memory hierarchy due to STT-RAM's non-volatility, low leakage power, high density, and fast read speed. The STT-RAM's small feature size is particularly desirable for the last-level cache (LLC), which typically consumes a large area of silicon die. However, long write latency and high write energy still remain challenges of implementing STT-RAMs in the CPU cache. An increasingly popular method for addressing this challenge involves trading off the non-volatility for reduced write speed and write energy by relaxing the STT-RAM's data retention time. However, in order to maximize energy saving potential, the cache configurations, including STT-RAM's retention time, must be dynamically adapted to executing applications' variable memory needs. In this paper, we propose a highly adaptable last level STT-RAM cache (HALLS) that allows the LLC configurations and retention time to be adapted to applications' runtime execution requirements. We also propose low-overhead runtime tuning algorithms to dynamically determine the best (lowest energy) cache configurations and retention times for executing applications. Compared to prior work, HALLS reduced the average energy consumption by 60.57 percent in a quad-core system, while introducing marginal latency overhead. Kyle Kuan, Tosiron Adegbija |
IEEE Trans. Computers | 1 |
| 2018 | LARS: Logically adaptable retention time STT-RAM cache for embedded systemsabstractSTT-RAMs have been studied as a promising alternative to SRAMs in embedded systems' caches and main memories. STT-RAMs are attractive due to their low leakage power and high density; STT-RAMs, however, also have drawbacks of long write latency and high dynamic write energy. A popular solution to this drawback relaxes the retention time to lower both write latency and energy, and uses a dynamic refresh scheme that refreshes data blocks to prevent them from prematurely expiring. However, the refreshes can incur overheads, thus limiting optimization potential. In addition, this solution only provides a single retention time, and cannot adapt to applications' variable retention time requirements. In this paper, we propose LARS (Logically Adaptable Retention Time STT-RAM) cache as a viable alternative for reducing the write energy and latency. LARS cache comprises of multiple STT-RAM units with different retention times, with only one unit on at a given time. LARS dynamically determines which STT-RAM unit to power on during runtime, based on executing applications' needs. Our experiments show that LARS cache is low-overhead, and can reduce the average energy and latency by 35.8% and 13.2%, respectively, as compared to the dynamic refresh scheme. Kyle Kuan, Tosiron Adegbija |
DATE | 1 |