Heewoo Kim

dblp:301/3787 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0001-6748-2890ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 AutoSSD: CXL-Enhanced Autonomous SSDs for Low Tail Latency
abstract
All-SSD RAID arrays offer high performance but suffer from long tail latencies due to background processes like garbage collection. Globally scheduling accesses to avoid busy SSDs may mitigate this problem, but such a solution presents challenges in collecting realtime data about SSD performance and tracking the location of redirected blocks.
Mingyao Shen, Suyash Mahar, Heewoo Kim, Joseph Izraelevitz, Steven Swanson
HPDC3
2025 NMP-PaK: Near-Memory Processing Acceleration of Scalable De Novo Genome Assembly
abstract
De novo assembly enables investigations of unknown genomes, paving the way for personalized medicine and disease management.However, it faces immense computational challenges arising from the excessive data volumes and algorithmic complexity.While state-of-the-art de novo assemblers utilize distributed systems for extreme-scale genome assembly, they demand substantial computational and memory resources.They also fail to address the inherent challenges of de novo assembly, including a large memory footprint, memory-bound behavior, and irregular data patterns stemming from complex, interdependent data structures.Given these challenges, de novo assembly merits a custom hardware solution, though existing approaches have not fully addressed the limitations.We propose NMP-PaK, a hardware-software co-designed system that accelerates scalable de novo genome assembly through near-memory processing (NMP).Our channel-level NMP architecture addresses memory bottlenecks while providing sufficient scratchpad space for processing elements.Customized processing elements maximize parallelism while efficiently handling large data structures that are both dynamic and interdependent.Software optimizations include customized batch processing to reduce the memory footprint and hybrid CPU-NMP processing to address hardware underutilization caused by irregular data patterns.NMP-PaK conducts the same genome assembly while incurring a 14× smaller memory footprint compared to the state-of-the-art de novo assembly.Moreover, NMP-PaK delivers 16× and 5.7× performance improvements over the CPU and GPU baselines, respectively, with a 2.4× reduction in memory operations.Consequently, NMP-PaK achieves 8.3× greater throughput than state-of-the-art
Heewoo Kim, Sanjay Sri Vallabh Singapuram, Haojie Ye, Joseph Izraelevitz, Trevor N. Mudge, Ronald G. Dreslinski, Nishil Talati
ISCA1
2023 RecPIM: A PIM-Enabled DRAM-RRAM Hybrid Memory System For Recommendation Models
abstract
The performance of modern recommendation models is limited because of the memory bandwidth-hungry embedding layer reductions. We propose RecPIM-a novel hybrid memory system with DRAM and RRAM with PIM capability. The performance of traditional RRAM PIM is limited by the latency of bit-serial computation. RecPIM presents a comprehensive optimization approach that includes access-pattern-aware mapping, compute complexity reduction, and selective PIM reduction to offset this computation latency. Our evaluation shows that RecPIM offers significant performance, energy, and EDP improvement of 2.6×, 1.7×, and 4.4×, on average, compared to a CPU baseline. We also co-design wear-leveling techniques and demonstrate a practical lifetime of more than 12 years.
Heewoo Kim, Haojie Ye, Trevor N. Mudge, Ronald G. Dreslinski, Nishil Talati
ISLPED1
2022 SiC Processors for Extreme High- Temperature Venus Surface Exploration
abstract
Being the ‘sister planet’ of the Earth, surface explo-ration of Venus is expected to provide valuable scientific insights into the history and the environment of the Earth. Despite the benefits, the surface temperature of Venus, at 450°C, poses a large challenge for any surface exploration. In particular, conventional Silicon electronics do not properly function under such high temperatures. Due to this constraint, the most prolonged previous surface exploration lasted only for 2 hours. Silicon Carbide (SiC) electronics, which can endure and function properly in high-temperature environments, is proposed as a strong candidate to be used in Venus surface explorations. However, this technology is still immature and associated with limiting factors, such as slower speed, power constraint, limited die area, and approximately 1,000 times longer channel than the state-of-the-art Si transistors. In this paper, we configure a computing infrastructure for high-temperature SiC-based technology, conduct design space explo-ration, and evaluate the performance of different SiC processors when used in Venus surface landers. Our evaluation shows that the SiC processor has an average 16.6× lower throughput than the RAD6000 Si processor used in the previous Mars rover. The Venus rover with SiC processor is expected to have a moving speed of 0.6 meters per hour and visual odometry processing time of 50 minutes. Lastly, we provide the design guidelines to improve the SiC processors at the microarchitecture and the instruction set architecture levels.
Heewoo Kim, Javad Bagherzadeh, Ronald G. Dreslinski
DATE1
2021 A Survey Describing Beyond Si Transistors and Exploring Their Implications for Future Processors
abstract
The advancement of Silicon CMOS technology has led information technology innovation for decades. However, scaling transistors down according to Moore’s law is almost reaching its limitations. To improve system performance, cost, and energy efficiency, vertical-optimization in multiple layers of the computing stack is required. Technological awareness in terms of devices and circuits could enable informed system-level decisions. For example, graphene is a promising material for extremely scaled high-speed transistors because of its remarkably high mobility, but it can not be used in integrated circuits as a result of the high leakage current from its zero bandgap. In this article, we discuss the fundamental physics of transistors and their ramifications on system design to assist device-level technology consideration during system design. Additionally, various emerging devices and their utilization on a vertically-optimized computing stack are introduced. This article serves as a survey of emerging device technologies that may be relevant in these areas, with an emphasis on making the descriptions approachable by system and software designers to understand the potential solutions. A basic vocabulary will be built to understand how to digest technical content, followed by a survey of devices, and finally a discussion of the implications for future processing systems.
Heewoo Kim, Aporva Amarnath, Javad Bagherzadeh, Nishil Talati, Ronald G. Dreslinski
ACM J. Emerg. Technol. Comput. Syst.1