Yangwook Kang

dblp:36/7454 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
4since 2021 · last 2026
0009-0007-9536-1894ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 5 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Patterns Behind Chaos: Forecasting Data Movement for Efficient Large-Scale Moe LLM Inference
abstract
Large-scale Mixture of Experts (MoE) Large Language Models (LLMs) have recently become the frontier open-weight models, achieving remarkable model capability similar to proprietary ones. But their random expert selection mechanism introduces significant data movement overhead that becomes the dominant bottleneck in multi-unit LLM serving systems. To understand the patterns underlying this data movement, we conduct comprehensive data-movement-centric profiling across four state-of-the-art large-scale MoE models released in 2025 (200B-1000B) using over 24,000 requests spanning diverse workloads. We perform systematic analysis from both temporal and spatial perspectives and distill six key insights to guide the design of diverse serving systems. We verify these insights on both future wafer-scale GPU architectures and existing GPU systems. On wafer-scale GPUs, lightweight architectural modifications guided by our insights yield a 6.6$\times$ average speedup across four 200B--1000B models. On existing GPU systems, our insights drive the design of a prefill-aware expert placement algorithm that achieves up to 1.25$\times$ speedup on MoE computation. Our work presents the first comprehensive data-centric analysis of large-scale MoE models together with a concrete design study applying the learned lessons. Our profiling traces are publicly available at \href{https://huggingface.co/datasets/core12345/MoE_expert_selection_trace}{\textcolor{blue}{https://huggingface.co/datasets/core12345/MoE\_expert\_selection\_trace}}.
Zhongkai Yu, Yue Guan 0003, Zhengding Hu, Shuyi Pei, Yangwook Kang, Yufei Ding 0001, Po-An Tsai
ISCA7
2025 BitWeaver: Read-Time Truncation in Memory
abstract
Large language models (LLMs) have demonstrated remarkable capabilities in generating contextually relevant responses to prompts, but their inference performance is often constrained by a severe memory bottleneck in the self-attention stage.This bottleneck, which is inherently memory-bound, has led to extensive research into strategies for reducing the size of the Key-Value (KV) cache.Many existing approaches employ quantization to lower the data precision and reduce data volume.However, these methods are constrained by memory technology, which requires the data read from memory to match the data written into it.As a result, such strategies must apply reductions at write-time, limiting either model performance or achievable speedup.We identify that enabling precision scaling at read-time -after data has been stored in memory -offers a unique opportunity to simultaneously reduce memory traffic and retain model accuracy.To this end, we develop a read-time precision-scaling mechanism and introduce BitWeaver, a hardware-enabled solution for in-memory truncation.BitWeaver dynamically reduces data precision during memory reads, achieving up to a 3× increase in memory throughput and execution speedups of up to 80% compared to baseline.Additionally, BitWeaver enhances sparse KV cache strategies by improving the efficiency of state-of-the-art sparsity techniques.
Garrett Gagnon, Srikanth Malla, Yangwook Kang, Liu Liu 0017
ICS3
2025 TRACI: Network Acceleration of Input-Dynamic Communication for Large-Scale Deep Learning Recommendation Model
abstract
Large-scale deep learning recommendation models (DLRMs) rely on embedding layers with terabyte-scale embedding tables, which present significant challenges to memory capacity.In addition, these embedding layers exhibit sparse and random data access patterns, which demand high memory bandwidth.Multi-GPU systems provide a promising solution, allowing for the scaling of both memory and aggregated bandwidth.However, network communication bandwidth becomes a bottleneck for multi-GPU DLRM systems.Overcoming the communication bottleneck is crucial to unlocking the potential of multi-GPU systems for efficient and high-performance DLRM training.This paper introduces TRACI, an in-network acceleration architecture designed to optimize the communication operator in embedding layers: Aggregation.While in-network acceleration has proven successful for the All-Reduce communication collective, existing solutions do not directly apply to Aggregation due to two key challenges.Firstly, existing multi-GPU shared memory operations are designed for point-to-point communication and do not allow the network to proactively optimize communication.Secondly, in Aggregation, data transfer patterns are dynamic and dependent on input, demanding the network to dynamically discover and exploit message connections on-the-fly.To address these challenges, we propose a solution that involves a novel network transaction and switch hardware design.We introduce a new network transaction that augments messages with input reuse and output reuse identifications, and can empower the network to proactively reduce
Guyue Huang, Hao Li 0120, Jiayi Huang 0001, Yangwook Kang, Yufei Ding 0001, Yuan Xie 0001
ISCA5
2023 ECSSD: Hardware/Data Layout Co-Designed In-Storage-Computing Architecture for Extreme Classification
abstract
With the rapid growth of classification scale in deep learning systems, the final classification layer becomes extreme classification with a memory footprint exceeding the main memory capacity of the CPU or GPU. The emerging in-storage-computing technique offers an opportunity on account of the fact that SSD has enough storage capacity for the parameters of extreme classification. However, the limited performance of naive in-storage-computing schemes is insufficient to support the heavy workload of extreme classification.
Siqi Li 0013, Fengbin Tu, Liu Liu 0017, Jilan Lin, Zheng Wang 0075, Yangwook Kang, Yufei Ding 0001, Yuan Xie 0001
ISCA6
2019 Towards building a high-performance, scale-in key-value storage system
abstract
Key-value stores are widely used as storage backends, due to their simple, yet flexible interface for cache, storage, file system, and database systems. However, when used with high performance NVMe devices, their high compute requirements for data management often leave the device bandwidth under-utilized. This leads to a performance mismatch of what the device is capable of delivering and what it actually delivers, and the gains derived from high speed NVMe devices is nullified. In this paper, we introduce KV-SSD (Key-Value SSD) as a key technology in a holistic approach to overcome such performance imbalance. KV-SSD provides better scalability and performance by simplifying the software storage stack and consolidating redundancy, thereby lowering the overall CPU usage and releasing the memory to user applications. We evaluate the performance and scalability of KV-SSDs over state-of-the-art software alternatives built for traditional block SSDs. Our results show that, unlike traditional key-value systems, the overall performance ofKV-SSD scales linearly, and delivers 1.6 to 57x gains depending on the workload characteristics.
Yangwook Kang, Rekha Pitchumani, Pratik Mishra, Yang-Suk Kee, Francisco Londono, Sangyoon Oh 0002, Jongyeol Lee, Daniel D. G. Lee
SYSTOR1
2014 Muninn: a Versioning Flash Key-Value Store Using an Object-based Storage Model
abstract
While non-volatile memory (NVRAM) devices have the potential to alleviate the trade-off between performance, scalability, and energy in storage and memory subsystems, a block interface and storage subsystems designed for slow I/O devices make it difficult to efficiently exploit NVRAMs in a portable and extensible way.
Yangwook Kang, Rekha Pitchumani, Thomas Marlette, Ethan L. Miller
SYSTOR1
2014 Random Slicing: Efficient and Scalable Data Placement for Large-Scale Storage Systems
abstract
The ever-growing amount of data requires highly scalable storage solutions. The most flexible approach is to use storage pools that can be expanded and scaled down by adding or removing storage devices. To make this approach usable, it is necessary to provide a solution to locate data items in such a dynamic environment. This article presents and evaluates the Random Slicing strategy, which incorporates lessons learned from table-based, rule-based, and pseudo-randomized hashing strategies and is able to provide a simple and efficient strategy that scales up to handle exascale data. Random Slicing keeps a small table with information about previous storage system insert and remove operations, drastically reducing the required amount of randomness while delivering a perfect load distribution.
Alberto Miranda, Sascha Effert, Yangwook Kang, Ethan L. Miller, Ivan Popov, André Brinkmann, Tom Friedetzky, Toni Cortes
ACM Trans. Storage3
2013 Enabling cost-effective data processing with smart SSD
abstract
This paper explores the benefits and limitations of in-storage processing on current Solid-State Disk (SSD) architectures. While disk-based in-storage processing has not been widely adopted, due to the characteristics of hard disks, modern SSDs provide high performance on concurrent random writes, and have powerful processors, memory, and multiple I/O channels to flash memory, enabling in-storage processing with almost no hardware changes. In addition, offloading I/O tasks allows a host system to fully utilize devices' internal parallelism without knowing the details of their hardware configurations. To leverage the enhanced data processing capabilities of modern SSDs, we introduce the Smart SSD model, which pairs in-device processing with a powerful host system capable of handling data-oriented tasks without modifying operating system code. By isolating the data traffic within the device, this model promises low energy consumption, high parallelism, low host memory footprint and better performance. To demonstrate these capabilities, we constructed a prototype implementing this model on a real SATA-based SSD. Our system uses an object-based protocol for low-level communication with the host, and extends the Hadoop MapReduce framework to support a Smart SSD. Our experiments show that total energy consumption is reduced by 50% due to the low-power processing inside a Smart SSD. Moreover, a system with a Smart SSD can outperform host-side processing by a factor of two or three by efficiently utilizing internal parallelism when applications have light trafic to the device DRAM under the current architecture.
Yangwook Kang, Yang-Suk Kee, Ethan L. Miller, Chanik Park
MSST1
2012 Emulating a Shingled Write Disk
abstract
Shingled Magnetic Recording technology is expected to play a major role in the next generation of hard disk drives. But it introduces some unique challenges to system software researchers and prototype hardware is not readily available for the broader research community. It is crucial to work on system software in parallel to hardware manufacturing, to ensure successful and effective adoption of this technology. In this work, we present a novel Shingled Write Disk (SWD) emulator that uses a hard disk utilizing traditional Perpendicular Magnetic Recording (PMR) and emulates a Shingled Write Disk on top of it. We implemented the emulator as a pseudo block device driver and evaluated the performance overhead incurred by employing the emulator. The emulator has a slight overhead which is only measurable during pure sequential reads and writes. The moment disk head movement comes into picture, due to any random access, the emulator overhead becomes so insignificant as to become immeasurable.
Rekha Pitchumani, Andy Hospodor, Ahmed Amer, Yangwook Kang, Ethan L. Miller, Darrell D. E. Long
MASCOTS4
2011 Reliable and randomized data distribution strategies for large scale storage systems
abstract
The ever-growing amount of data requires highly scalable storage solutions. The most flexible approach is to use storage pools that can be expanded and scaled down by adding or removing storage devices. To make this approach usable, it is necessary to provide a solution to locate data items in such a dynamic environment. This paper presents and evaluates the Random Slicing strategy, which incorporates lessons learned from table-based, rule-based, and pseudo-randomized hashing strategies and is able to provide a simple and efficient strategy that scales up to handle exascale data. Random Slicing keeps a small table with information about previous storage system insert and remove operations, drastically reducing the required amount of randomness while delivering a perfect load distribution.
Alberto Miranda, Sascha Effert, Yangwook Kang, Ethan L. Miller, André Brinkmann, Toni Cortes
HiPC3
2011 Object-based SCM: An efficient interface for Storage Class Memories
abstract
Storage Class Memory (SCM) has become increasingly popular in enterprise systems as well as embedded and mobile systems. However, replacing hard drives with SCMs in current storage systems often forces either major changes in file systems or suboptimal performance, because the current block-based interface does not deliver enough information to the device to allow it to optimize data management for specific device characteristics such as the out-of-place update. To alleviate this problem and fully utilize different characteristics of SCMs, we propose the use of an object-based model that provides the hardware and firmware the ability to optimize performance for the underlying implementation, and allows drop-in replacement for devices based on new types of SCM. We discuss the design of object-based SCMs and implement an object-based flash memory prototype. By analyzing different design choices for several subsystems, such as data placement policies and index structures, we show that our object-based model provides comparable performance to other flash file systems while enabling advanced features such as object-level reliability.
Yangwook Kang, Jingpei Yang, Ethan L. Miller
MSST1
2011 Design and evaluation of Oasis: An active storage framework based on T10 OSD standard
abstract
In this paper, we present the design and performance evaluation of Oasis, an active storage framework for object-based storage systems that complies with the current T10 OSD standard. In contrast with previous work, Oasis has the following advantages. First, Oasis enables users to transparently process the OSD object and supports different processing granularity (from the single object to all the objects in the OSD) by extending the OSD object attribute page defined in the T10 OSD standard. Second, Oasis provides an easy and efficient way for users to manage the application functions in the OSD by using the existing OSD commands. Third, Oasis can authorize the execution of the application function in the OSD by enhancing the T10 OSD security protocol, allowing only authorized users to use the system. We evaluate the performance and scalability of our system implementation on Oasis by running three typical applications. The results indicate that active storage far outperforms the traditional object-based storage system in applications that filter data on the OSD. We also experiment with Java based applications and C based applications. Our experiments indicate that Java based applications may be bottlenecked for I/O-intensive applications, while for applications that do not heavily rely on the I/O operations, both Java based applications and C based applications achieve comparable performance. Our microbenchmarks indicate that Oasis implementation overhead is minimal compared to the Intel OSD reference implementation, between 1.2% to 5.9% for Read commands and 0.6% to 9.9% for Write commands.
Yulai Xie 0002, Kiran-Kumar Muniswamy-Reddy, Dan Feng 0001, Darrell D. E. Long, Yangwook Kang, Zhongying Niu
MSST5
2010 Efficient Storage Management for Object-based Flash Memory
abstract
Flash memory has become increasingly popular in today's storage systems. However, replacing hard drives with flash memory in current systems often either requires major file system changes or causes performance degradation due to the limitations of block-based interface and out-of-place updates required by flash. To alleviate this problem, we propose an object-based model for flash memory that gives the hardware and firmware the ability to optimize performance for the underlying implementation. Based on this model, we propose two new data placement policies that exploit richer information from an object-based interface. Using simulation, we show that cleaning overhead can be reduced by up to 9% by separating data and metadata. Segregating the access time from metadata can further reduce the cleaning overhead by up to 23%.
Yangwook Kang, Jingpei Yang, Ethan L. Miller
MASCOTS1
2009 Adding aggressive error correction to a high-performance compressing flash file system
abstract
While NAND flash memories have rapidly increased in both capacity and performance and are increasingly used as a storage device in many embedded systems, their reliability has decreased both because of increased density and the use of multi-level cells (MLC). Current MLC technology only specifies the minimum requirement for an error correcting code (ECC), but provides no additional protection in hardware. However, existing flash file systems such as YAFFS and JFFS2 rely upon ECC to survive small numbers of bit errors, but cannot survive the larger numbers of bit errors or page failures that are becoming increasingly common as flash file systems scale to multiple gigabytes.
Yangwook Kang, Ethan L. Miller
EMSOFT1