Kyu-Hyoun Kim

dblp:287/1276 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
1since 2021 · last 2021
0000-0002-5624-8465ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Memory systems · 55% Hardware accelerators and domain-specific architectures · 23% Hardware reliability and fault tolerance · 10%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
non-volatile memory
0.522020
Temperature Aware Adaptations for Improved Read Reliability in STT-MRAM Memory Subsystem · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Contutto: a novel FPGA-based prototyping platform enabling innovation in the memory subsystem of a server class processor · MICRO 2017
Memory systems › non-volatile memory › magnetic random access memory
STT-MRAM
0.522020
Temperature Aware Adaptations for Improved Read Reliability in STT-MRAM Memory Subsystem · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Contutto: a novel FPGA-based prototyping platform enabling innovation in the memory subsystem of a server class processor · MICRO 2017
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN training accelerator
0.512021
RaPiD: AI Accelerator for Ultra-low Precision Training and Inference · ISCA 2021
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.512021
RaPiD: AI Accelerator for Ultra-low Precision Training and Inference · ISCA 2021
Hardware reliability and fault tolerance › memory reliability
read reliability
0.412020
Temperature Aware Adaptations for Improved Read Reliability in STT-MRAM Memory Subsystem · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Memory systems
memory compression
0.312018
Attaché: Towards Ideal Memory Compression by Mitigating Metadata Bandwidth Overheads · MICRO 2018
Memory systems
DRAM
0.322020
Understanding and mitigating refresh overheads in high-density DDR4 DRAM systems · ISCA 2013
Temperature Aware Adaptations for Improved Read Reliability in STT-MRAM Memory Subsystem · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Reconfigurable computing and FPGAs
FPGA prototyping
0.312017
Contutto: a novel FPGA-based prototyping platform enabling innovation in the memory subsystem of a server class processor · MICRO 2017
Memory systems › DRAM › DRAM architecture
DDR4
0.212013
Understanding and mitigating refresh overheads in high-density DDR4 DRAM systems · ISCA 2013
Energy-efficient computing › energy-efficient architecture
energy-efficient accelerator
0.112021
RaPiD: AI Accelerator for Ultra-low Precision Training and Inference · ISCA 2021
Memory systems
memory bandwidth
0.112018
Attaché: Towards Ideal Memory Compression by Mitigating Metadata Bandwidth Overheads · MICRO 2018
Memory systems
emerging memory technologies
0.112017
Contutto: a novel FPGA-based prototyping platform enabling innovation in the memory subsystem of a server class processor · MICRO 2017
Memory systems › non-volatile memory › non-volatile main memory
NVDIMM
0.112017
Contutto: a novel FPGA-based prototyping platform enabling innovation in the memory subsystem of a server class processor · MICRO 2017

Methods — techniques the papers use, named apart from their topics

performance modeling · 0.5read disturb analysis · 0.4bit error rate estimation · 0.4sub-ranking · 0.3hardware prediction · 0.3FPGA prototyping · 0.3simulation · 0.2
YearPublicationVenuePosition
2021 RaPiD: AI Accelerator for Ultra-low Precision Training and Inference
abstract
The growing prevalence and computational demands of Artificial Intelligence (AI) workloads has led to widespread use of hardware accelerators in their execution. Scaling the performance of AI accelerators across generations is pivotal to their success in commercial deployments. The intrinsic error-resilient nature of AI workloads present a unique opportunity for performance/energy improvement through precision scaling. Motivated by the recent algorithmic advances in precision scaling for inference and training, we designed RaPiD1, a 4-core AI accelerator chip supporting a spectrum of precisions, namely, 16 and 8-bit floating-point and 4 and 2-bit fixed-point. The 36mm2RaPiD chip fabricated in 7nm EUV technology delivers a peak 3.5 TFLOPS/W in HFP8 mode and 16.5 TOPS/W in INT4 mode at nominal voltage. Using a performance model calibrated to within 1% of the measurement results, we evaluated DNN inference using 4-bit fixed-point representation for a 4-core 1 RaPiD chip system and DNN training using 8-bit floating point representation for a 768 TFLOPs AI system comprising 4 32-core RaPiD chips. Our results show INT4 inference for batch size of 1 achieves 3 - 13.5 (average 7) TOPS/W and FP8 training for a mini-batch of 512 achieves a sustained 102 - 588 (average 203) TFLOPS across a wide range of applications.
Swagath Venkataramani, Vijayalakshmi Srinivasan, Wei Wang 0333, Sanchari Sen, Ankur Agrawal, Monodeep Kar, Shubham Jain 0004, Alberto Mannari, Hoang Tran, Eri Ogawa, Kazuaki Ishizaki, Hiroshi Inoue, Marcel Schaal, Mauricio J. Serrano, Jungwook Choi, Xiao Sun 0013, Naigang Wang, Chia-Yu Chen, Allison Allain, James Bonanno, Nianzheng Cao, Robert Casatuta, Matthew Cohen, Bruce M. Fleischer, Michael Guillorn, Howard Haynie, Jinwook Jung, Mingu Kang, Kyu-Hyoun Kim, Siyu Koswatta, Sae Kyu Lee, Martin Lutz, Silvia M. Müller, Jinwook Oh, Ashish Ranjan 0001, Zhibin Ren, Scot Rider, Kerstin Schelm, Michael Scheuermann, Joel Silberman, Vidhi Zalani, Xin Zhang 0025, Ching Zhou, Matthew M. Ziegler, Vinay Shah, Moriyoshi Ohara, Pong-Fei Lu, Brian W. Curran, Sunil Shukla, Leland Chang, Kailash Gopalakrishnan
ISCA31
2020 Temperature Aware Adaptations for Improved Read Reliability in STT-MRAM Memory Subsystem
abstract
Spin-transfer torque magneto-resistive random-access memory (STT-MRAM) is an exciting new emerging technology, being considered as a strong candidate to fill the gaps in the existing memory hierarchy between DRAM and the secondary memory. STT-MRAM has adequate endurance. However, unresolved write switching and read reliability issues still exist at the functional operating temperature corners. One biggest challenge is that the read bit error rate (RBER) is not at an acceptable level for system reliability across the wide operating temperature range. We present an STT-MRAM memory subsystem that is fully compatible with existing DDR-based DIMM designs and evaluate read disturb and read sense bit-error rate (BER) under various operating temperature conditions. We propose temperature aware adaptive techniques for reliable reads at the rank level. The proposed temperature adaptation technique improves overall reliability of the DDR4 STT-MRAM-based memory subsystem with an optimal read current considering an acceptable 64-byte cacheline BER. Our full system simulations show 1000× order of improvements toward a cell raw read disturb BER along with 5% reduction in memory power and less than 1% impact on overall system performance.
Saravanan Sethuraman, T. Venkata Kalyan, Karthick Rajamani, Chitra K. Subramanian, Kyu-Hyoun Kim, Hillery C. Hunter, M. B. Srinivas
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2018 Attaché: Towards Ideal Memory Compression by Mitigating Metadata Bandwidth Overheads
abstract
Memory systems are becoming bandwidth constrained and data compression is seen as a simple technique to increase their effective bandwidth. However, data compressionrequires accessing Metadata which incurs additional bandwidth overheads. Even after using a Metadata-Cache, the bandwidth overheads of Metadata can reduce the benefits of compression. This paper proposes Attaché, a framework that reduces the overheads of Metadata accesses. The Attaché framework consists of two components. The first component, called the Blended Metadata Engine (BLEM), enables data and its Metadata to be accessed together. BLEM incurs additional Metadata accesses only 0.003% times and removes almost all Metadata bandwidth overheads. The second component, called theCompression Pre-dictor(COPR), predicts if the memory block is compressed. TheCOPR predictor uses a fine-grained line-level predictor, a coarse-grained page-level predictor, and a global indicator. This enables Attaché to predict the compressibility of the memory block before sending a memory read request. We implement Attaché on a memory system that uses Sub-Ranking. On average, Attaché achieves 15.3% speedup (ideal 17%) and saves 22% energy consumption (ideal 23%) when compared to a baseline system that does not employ data compression. Attaché is completely hardware-based and uses only 368KB of SRAM.
Seokin Hong, Prashant J. Nair, Bülent Abali, Alper Buyuktosunoglu, Kyu-Hyoun Kim, Michael B. Healy
MICRO5
2017 Contutto: a novel FPGA-based prototyping platform enabling innovation in the memory subsystem of a server class processor
abstract
We demonstrate the use of an FPGA as a memory buffer in a POWER8® system, creating a novel prototyping platform that enables innovation in the memory subsystem of POWER-based servers. Our platform, called ConTutto, is pin-compatible with POWER8 buffered memory DIMMs and plugs into a memory slot of a standard POWER8 processor system, running at aggregate memory channel speeds of 35 GB/s per link. ConTutto, which means "with everything", is a platform to experiment with different memory technologies, such as STT-MRAM and NAND Flash, in an end-to-end system context. Enablement of STT-MRAM and NVDIMM using ConTutto shows up to 12.5x lower latency and 7.5x higher bandwidth compared to the respective technologies when attached to the PCIe bus. Moreover, due to the unique attach-point of the FPGA between the processor and system memory, ConTutto provides a means for in-line acceleration of certain computations on-route to memory, and enables sensitivity analysis for memory latency while running real applications. To the best of our knowledge, ConTutto is the first ever FPGA platform on the memory bus of a server class processor.
Bharat Sukhwani, Chuck Haymes, Kyu-Hyoun Kim, Adam J. McPadden, Daniel M. Dreps, Dean Sanner, Jan van Lunteren, Sameh W. Asaad
MICRO4
2013 Understanding and mitigating refresh overheads in high-density DDR4 DRAM systems
abstract
Recent DRAM specifications exhibit increasing refresh latencies. A refresh command blocks a full rank, decreasing available parallelism in the memory subsystem significantly, thus decreasing performance. Fine Granularity Refresh (FGR) is a feature recently announced as part of JEDEC's DDR4 DRAM specification that attempts to tackle this problem by creating a range of refresh options that provide a trade-off between refresh latency and frequency.
Janani Mukundan, Hillery C. Hunter, Kyu-Hyoun Kim, Jeffrey Stuecheli, José F. Martínez
ISCA3