Joonsung Kim 0001

dblp:216/7152 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-5432-7813ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 4 first-author · 10 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021
YearPublicationVenuePosition
2026 SMTcheck: Accurate SMT Interference Prediction to Improve Scheduling Efficiency in Datacenters
abstract
Simultaneous multithreading (SMT) is widely used in modern x86 processors to improve core utilization by sharing hardware resources between co-located threads. However, such resource sharing often leads to severe performance interference, making efficient workload co-scheduling difficult, especially given the complexity and diversity of modern x86 CPUs. Our analysis reveals that SMT-aware workload scheduling can significantly improve system throughput and reduce tail latency for datacenter workloads, but identifying optimal thread combinations is challenging due to the lack of visibility into platform-specific resource sharing behaviors. In this paper, we present SMTcheck, a lightweight, accurate, and platformindependent methodology for predicting SMT interference for diverse x86 processors. SMTcheck uses carefully designed code snippets (Diags) to extract hidden microarchitectural features of performance-critical shared resources. With these extracted features, SMTcheck builds per-resource microbenchmarks (Injectors) to apply pinpoint pressure to specific target resources in order to capture workload-specific contention characteristics. SMTcheck then constructs a hardware-aware contention model to predict performance interference between arbitrary workload pairs without requiring exhaustive profiling. We evaluate SMTcheck on six x86 desktop processors and five x86 server processors from Intel and AMD across different generations and show that it achieves high prediction accuracy by up to 95.5 % (94.6 % on average). We further demonstrate its effectiveness by implementing a contention-aware scheduler in the Linux kernel. Compared to the default Linux scheduler, our contention-aware scheduler significantly reduces tail latency for latency-critical workloads (e.g., database, key-value store) by up to 36.09 %, and improves the overall system throughput by up to$1.072 \times$. Finally, using real-world cluster traces from Alibaba and Google, we demonstrate that SMTcheck incurs negligible profiling overheads ($\approx 0.113 \%$), making it practical for deployment in productionscale datacenter environments.
Jinhyeok Oh, Gyutae Kim, Youngsok Kim, Jae-Hyun Hwang, Joonsung Kim 0001
HPCA7
2025 ScaleMoE: A Fast and Scalable Distributed Training Framework for Large-Scale Mixture-of-Experts Models
abstract
The size of pre-trained models has continuously increased to support growing demands for solving more complex problems. Especially, mixture-of-experts (MoE) model has become the most popular approach, enabling systems to easily train extremely large-scale models with relatively lower computational requirements. However, the current distributed training frameworks cannot achieve scalable performance for these large-scale MoE models due to substantial communication overheads. In this paper, we propose ScaleMoE, a fast and scalable distributed training framework for large-scale MoE models. We first identify three problems in state-of-the-art distributed training frameworks: high all-to-all communication overheads, severe load imbalance in expert selection, and insufficient consideration of heterogeneous networks. We propose three novel optimizations to resolve these problems. First, to reduce communication volumes, we propose adaptive all-to-all communication that eliminates unnecessary zeros caused by zero padding. Second, to address the load imbalance in expert selection, we propose dynamic expert clustering that rebalances experts using a novel clustering methodology. Lastly, to further minimize communication overheads, we propose topology-aware expert remapping that carefully maps experts to GPU devices while considering heterogeneous network bandwidths. Our evaluations show that ScaleMoE achieves scalable performance, reducing all-to-all communication overheads by up to $\mathbf{8 1 \%}$. In general, ScaleMoE significantly improves system performance, achieving a speedup of up to $3.3 \times$ compared to the state-of-the-art framework.
Seohong Choi, Huize Hong, Tae Hee Han, Joonsung Kim 0001
PACT4
2025 DMO-DB: Mitigating the Data Movement Bottlenecks of GPU-Accelerated Relational OLAP
abstract
Graphics Processing Units (GPUs) offer high computational throughput and memory bandwidth, making them promising accelerators for relational OnLine Analytical Processing (OLAP). GPU-accelerated relational OLAP executes the relational operations of a Structured Query Language (SQL) query on a GPU instead of the host Central Processing Unit (CPU). Depending on where input columns and their values reside in, a GPU-accelerated SQL query execution can be classified into two scenarios: 1) in-GPU, in which all input columns fit in the GPU memory, or 2) in-host, in which the input columns reside in the host memory and get transferred to the GPU memory when needed. However, both scenarios incur significant intraGPU and host-to-GPU data movement overheads, respectively. In-GPU executions incur excessive GPU cache misses and thus frequent off-chip GPU memory accesses. In-host executions suffer from the limited host-to-GPU data transfer bandwidth. This paper presents DMO-DB, a Data Movement-Optimized GPU-accelerated relational OLAP engine. Since modern GPUaccelerated relational OLAP decomposes SQL queries into multiple pipelines-sequences of relational operations that can be executed on input columns from the same table, DMO-DB leverages inter-pipeline dependencies to overcome the two data movement bottlenecks. DMO-DB introduces two key ideas: cache-fit bloom filtering and Ahead-of-Time value Discarding (AoTD), which preemptively eliminate unnecessary input values before their movement across the memory hierarchies. For in-GPU execution, GPU L1 data cache-fit filters discard non-contributing values before triggering costly off-chip DRAM accesses. For in-host execution, host CPU last level cache-fit filters strategically prune unnecessary input values, minimizing PCIe transfer overhead. After that, AoTD exploits multiple inter-pipeline dependencies by collecting these cache-fit bloom filters to earlier pipeline execution stages. Our evaluation using NVIDIA RTX A4000 and TITAN RTX GPUs shows that DMO-DB achieves speedups of $\mathbf{1. 5 3 x}$ over in-GPU Crystal-Opt and 6.10x over in-host HeavyDB.
Chaemin Lim, Suhyun Lee 0002, Jinwoo Choi 0003, Joonsung Kim 0001, Jinho Lee 0001, Youngsok Kim
PACT4
2025 GCStack+GCScaler: Fast and Accurate GPU Performance Analyses Using Fine-Grained Stall Cycle Accounting and Interval Analysis
abstract
To design next-generation Graphics Processing Units (GPUs), GPU architects rely on GPU performance analyses to identify key GPU performance bottlenecks and explore GPU design spaces.Unfortunately, the existing GPU performance analysis mechanisms make it difficult for GPU architects to conduct fast and accurate GPU performance analyses.The existing mechanisms can provide misleading
Hanna Cha, Sungchul Lee, Jounghoo Lee, Yeonan Ha, Joonsung Kim 0001, Youngsok Kim
ISCA5
2025 LATPC: Accelerating GPU Address Translation Using Locality-Aware TLB Prefetching and MSHR Compression
Yeonan Ha, Hanna Cha, Jiwon Lee 0001, Joonsung Kim 0001, Won Woo Ro, Youngsok Kim
MICRO5
2023 A Fast and Flexible FPGA-based Accelerator for Natural Language Processing Neural Networks
abstract
Deep neural networks (DNNs) have become key solutions in the natural language processing (NLP) domain. However, the existing accelerators customized for their narrow target models cannot support diverse NLP models. Therefore, naively running complex NLP models on the existing accelerators often leads to very marginal performance improvements. For these reasons, architects are now in dire need of a new accelerator that can run various NLP models while taking its full performance potential. In this article, we propose FlexRun, an FPGA-based modular accelerator to efficiently support diverse and complex NLP models. First, we identify key components commonly used by NLP models and implement them on top of a current state-of-the-art FPGA-based accelerator. Next, FlexRun conducts an in-depth design space exploration to find the best accelerator architecture for a target NLP model. Last, FlexRun automatically reconfigures the accelerator based on the exploration results. Our FlexRun design outperforms the current state-of-the-art FPGA-based accelerator by 1.21×–2.73× and 1.15×–1.50× for BERT and GPT2, respectively. Compared to Nvidia’s V100 GPU, FlexRun achieves 2.69× higher performance on average for various BERT and GPT2 models.
Suyeon Hur, Seongmin Na, Dongup Kwon, Joonsung Kim 0001, Andrew Boutros, Eriko Nurvitadhi, Jangwoo Kim
ACM Trans. Archit. Code Optim.4
2022 3D-FPIM: An Extreme Energy-Efficient DNN Acceleration System Using 3D NAND Flash-Based In-Situ PIM Unit
abstract
The crossbar structure of the nonvolatile memory enables highly parallel and energy-efficient analog matrix-vector-multiply (MVM) operations. To exploit its efficiency, existing works design a mixed-signal deep neural network (DNN) accelerator, which offloads low-precision MVM operations to the memory array. However, they fail to accurately and efficiently support the low-precision networks due to their naive ADC designs. In addition, they cannot be applied to the latest technology nodes due to their premature RRAM-based memory array.In this work, we present 3D-FPIM, an energy-efficient and robust mixed-signal DNN acceleration system. 3D-FPIM is a full-stack 3D NAND flash-based architecture to accurately deploy low-precision networks. We design the hardware stack by carefully architecting a specialized analog-to-digital conversion method and utilizing the three-dimensional structure to achieve high accuracy, energy efficiency, and robustness. To accurately and efficiently deploy the networks, we provide a DNN retraining framework and a customized compiler. For evaluation, we implement an industry-validated circuit-level simulator. The result shows that 3D-FPIM achieves an average of 2.09x higher performance per area and 13.18x higher energy efficiency compared to the baseline 2D RRAM-based accelerator.
Hunjun Lee, Minseop Kim, Dongmoon Min, Joonsung Kim 0001, Jongwon Back, Honam Yoo, Jong-Ho Lee 0002, Jangwoo Kim
MICRO4
2021 NLP-Fast: A Fast, Scalable, and Flexible System to Accelerate Large-Scale Heterogeneous NLP Models
abstract
Emerging natural language processing (NLP) models have become more complex and bigger to provide more sophisticated NLP services. Accordingly, there is also a strong demand for scalable and flexible computer infrastructure to support these large-scale, complex, and diverse NLP models. However, existing proposals cannot provide enough scalability and flexibility as they neither identify nor optimize a wide spectrum of performance-critical operations appearing in recent NLP models and only focus on optimizing specific operations. In this paper, we propose NLP-Fast, a novel system solution to accelerate a wide spectrum of large-scale NLP models. NLP-Fast mainly consists of two parts: (1) NLP-Perf: an in-depth performance analysis tool to identify critical operations in emerging NLP models and (2) NLP-Opt: three end-to-end optimization techniques to accelerate the identified performance-critical operations on various hardware platforms (e.g., CPU, GPU, FPGA). In this way, NLP-Fast can accelerate various types of NLP models on different hardware platforms by identifying their critical operations through NLP-Perf and applying the NLP-Opt's holistic optimizations. We evaluate NLP-Fast on CPU, GPU, and FPGA, and the overall throughputs are increased by up to 2.92×, 1.59×, and 4.47× over each platform's baseline. We release NLP-Fast to the community so that users are easily able to conduct the NLP-Fast's analysis and apply NLP-Fast's optimizations for their own NLP applications.
Joonsung Kim 0001, Suyeon Hur, Eunbok Lee, Jangwoo Kim
PACT1
2021 UC-Check: Characterizing Micro-operation Caches in x86 Processors and Implications in Security and Performance
abstract
The modern x86 processor (e.g., Intel, AMD) translates CISC-style x86 instructions to RISC-style micro operations (uops) as RISC pipelines are more efficient than CISC pipelines. However, this x86 decoding process requires complex hardware logic (i.e., x86 decoder) to identify variable-length x86 instructions, which incurs high translation overhead. To avoid this overhead, the x86 processors adopt a micro-operation cache (uop cache) to bypass the expensive x86 decoder by caching the decoded uops.
Joonsung Kim 0001, Hamin Jang, Hunjun Lee, Jangwoo Kim
MICRO1
2021 Performance Modeling and Practical Use Cases for Black-Box SSDs
abstract
Modern servers are actively deploying Solid-State Drives (SSDs) thanks to their high throughput and low latency. However, current server architects cannot achieve the full performance potential of commodity SSDs, as SSDs are complex devices designed for specific goals (e.g., latency, throughput, endurance, cost) with their internal mechanisms undisclosed to users. In this article, we propose SSDcheck , a novel SSD performance model to extract various internal mechanisms and predict the latency of next access to commodity black-box SSDs. We identify key performance-critical features (e.g., garbage collection, write buffering) and find their parameters (i.e., size, threshold) from each SSD by using our novel diagnosis code snippets. Then, SSDcheck constructs a performance model for a target SSD and dynamically manages the model to predict the latency of the next access. In addition, SSDcheck extracts and provides other useful internal mechanisms (e.g., fetch unit in multi-queue SSDs, background tasks triggering idle-time interval) for the storage system to fully exploit SSDs. By using those useful features and the performance model, we propose multiple practical use cases. Our evaluations show that SSDcheck’s performance model is highly accurate, and proposed use cases achieve significant performance improvement in various scenarios.
Joonsung Kim 0001, Kanghyun Choi, Wonsik Lee, Jangwoo Kim
ACM Trans. Storage1
2019 Enforcing Last-Level Cache Partitioning through Memory Virtual Channels
abstract
Ensuring fairness or providing isolation between multiple workloads with different characteristics that are colocated on a single, shared-memory system is a challenge. Recent multicore processors provide last-level cache (LLC) hardware partitioning to provide hardware support for isolation, with the cache partitioning often specified by the user. While more LLC capacity often results in higher performance, in this work we identify that a workload allocated more LLC capacity result in worse performance on real-machine experiments, which we refer to as MiW (more is worse). Through various controlled experiments, we identify that another workload with less LLC capacity causes more frequent LLC misses. The workload stresses the main-memory system shared by both workloads and degrades the performance of the former workload even if the LLC partitioning is used (a balloon effect). To resolve this problem, we propose virtualizing the datapath of main-memory controllers and dedicating the memory virtual channels (mVCs) to each group of applications, grouped for LLC partitioning. mVC can further fine-tune the performance of groups by differentiating buffer sizes among mVCs. It can reduce the total system cost by executing latency-critical and throughput-oriented workloads together on shared machines, of which performance criteria can be achieved only on dedicated machines if mVCs are not supported. Experiments on a simulated chip multiprocessor show that our proposals effectively eliminate the MiW phenomenon, hence providing additional opportunities for workload consolidation in a datacenter. Our case study demonstrates potential savings of machine count by 21.8% with mVC, which would otherwise violate a service level objective (SLO).
Jongwook Chung, Yuhwan Ro, Joonsung Kim 0001, Jaehyung Ahn, Jangwoo Kim, John Kim 0001, Jae W. Lee, Jung Ho Ahn
PACT3
2019 μLayer: Low Latency On-Device Inference Using Cooperative Single-Layer Acceleration and Processor-Friendly Quantization
abstract
Emerging mobile services heavily utilize Neural Networks (NNs) to improve user experiences. Such NN-assisted services depend on fast NN execution for high responsiveness, demanding mobile devices to minimize the NN execution latency by efficiently utilizing their underlying hardware resources. To better utilize the resources, existing mobile NN frameworks either employ various CPU-friendly optimizations (e.g., vectorization, quantization) or exploit data parallelism using heterogeneous processors such as GPUs and DSPs. However, their performance is still bounded by the performance of the single target processor, so that realtime services such as voice-driven search often fail to react to user requests in time. It is obvious that this problem will become more serious with the introduction of more demanding NN-assisted services.
Youngsok Kim, Joonsung Kim 0001, Dongju Chae, Dae-Hyun Kim 0003, Jangwoo Kim
EuroSys2
2019 CIDR: A Cost-Effective In-Line Data Reduction System for Terabit-Per-Second Scale SSD Arrays
abstract
An SSD array, a storage system consisting of multiple SSDs per node, has become a design choice to implement a fast primary storage system, and modern storage architects now aim to achieve terabit-per-second scale performance with the next-generation SSD array. To reduce the storage cost and improve the device endurability, such SSD array must employ data reduction schemes (i.e., deduplication, compression), which provide high data reduction capability at minimum costs. However, existing data reduction schemes do not scale with the fast increasing performance of an SSD array, due to inhibitive amount of CPU resources (e.g., in software-based schemes) or low data reduction ratio (e.g., in SSD device wide deduplication) or being cost ineffective to address workload changes in datacenters (e.g., in ASIC-based acceleration). In this paper, we propose CIDR, a novel FPGA-based, cost-effective data reduction system for an SSD array to achieve the terabit-per-second scale storage performance. Our key ideas are as follows. First, we decouple data reduction related computing tasks from the unscalable host CPUs by offloading them to a scalable array of FPGA boards. Second, we employ a centralized, node-wide metadata management scheme to achieve an SSD array-wide, high data reduction. Third, our FPGA-based reconfiguration adapts to different workload patterns by dynamically balancing the amount of software and hardware tasks running on CPUs and FPGAs, respectively. For evaluation, we built our example CIDR prototype achieving up to 12.8 GB/s (0.1 Tbps) on one FPGA. CIDR outperforms the baseline for a write-only workload by up to 2.47x and a mixed read-write workload by an expected 3.2x, respectively. We showed CIDR's scalability to achieve Tbps-scale performance by measuring a two-FPGA CIDR and projecting the performance impacts for more FPGAs.
Mohammadamin Ajdari, Pyeongsu Park, Joonsung Kim 0001, Dongup Kwon, Jangwoo Kim
HPCA3
2019 MnnFast: a fast and scalable system architecture for memory-augmented neural networks
abstract
Memory-augmented neural networks are getting more attention from many researchers as they can make an inference with the previous history stored in memory. Especially, among these memory-augmented neural networks, memory networks are known for their huge reasoning power and capability to learn from a large number of inputs rather than other networks. As the size of input datasets rapidly grows, the necessity of large-scale memory networks continuously arises. Such large-scale memory networks provide excellent reasoning power; however, the current computer infrastructure cannot achieve scalable performance due to its limited system architecture.
Hanhwi Jang, Joonsung Kim 0001, Jae-Eon Jo, Jangwoo Kim
ISCA2
2019 FIDR: A Scalable Storage System for Fine-Grain Inline Data Reduction with Efficient Memory Handling
abstract
Storage systems play a critical role in modern servers which run highly data-intensive applications. To satisfy the high performance and capacity demands of such applications, storage systems now deploy an array of fast SSDs per server. To reduce the storage cost of employing many SSDs per server, storage systems actively perform inline data reduction (e.g., data deduplication, compression). Existing inline data reduction studies can achieve high performance and scalability by offloading computation-intensive data-reduction operations to dedicated hardware accelerators. However, such existing studies suffer from limited workload support and scalability. For example, they reduce only large data blocks, which incur many IO requests, leading to low data reduction rates, and their offloading overlooks memory-intensive operations, leading to the unoptimal scalability.
Mohammadamin Ajdari, Wonsik Lee, Pyeongsu Park, Joonsung Kim 0001, Jangwoo Kim
MICRO4
2018 SSDcheck: Timely and Accurate Prediction of Irregular Behaviors in Black-Box SSDs
abstract
Modern servers are actively deploying Solid-State Drives (SSDs). However, rather than just a fast storage device, SSDs are complex devices designed for device-specific goals (e.g., latency, throughput, endurance, cost) with their internal mechanisms undisclosed to users as the proprietary asset, which leads to unpredictable, irregular inter/intra-SSD access latencies. This unpredictable irregular access latency has been a fundamental challenge to server architects aiming to satisfy critical quality-of-service requirements and/or achieve the full performance potential of commodity SSDs. In this paper, we propose SSDcheck, a novel SSD performance model to accurately predict the latency of next access to commodity black-box SSDs. First, after analyzing a wide spectrum of real-world SSDs, we identify key performance-critical features (e.g., garbage collection, write buffering) required to construct a general SSD performance model. Next, SSDcheck runs diagnosis code snippets to extract static feature parameters (e.g., size, threshold) from the target SSD, and constructs its performance model. Finally, during runtime, SSDcheck dynamically manages the performance model to predict the latency of the next access. Our evaluations show that SSDcheck achieves up to 98.96% and 79.96% on-average prediction accuracy for normal-latency and high-latency predictions, respectively. Next, we show the effectiveness of SSDcheck by implementing a new volume manager improving the throughput by up to 4.29x with the tail latency reduction down to 6.53%, and a new I/O request handler improving the throughput by up to 44.0% with the tail latency reduction down to 26.9%. We then show how to further improve the results of scheduling with the help of an emerging Non-Volatile Memory (e.g., PCM). SSDcheck does not require any hardware modifications, which can be harmlessly disabled for any SSDs uncovered by the performance model.
Joonsung Kim 0001, Pyeongsu Park, Jaehyung Ahn, Jihun Kim 0002, Jong Kim 0001, Jangwoo Kim
MICRO1
2018 DynaMix: Dynamic Mobile Device Integration for Efficient Cross-device Resource Sharing
Dongju Chae, Joonsung Kim 0001, Gwangmu Lee, Hanjun Kim 0001, Kyung-Ah Chang, Hyogun Lee, Jangwoo Kim
USENIX ATC2
2016 CloudSwap: A Cloud-Assisted Swap Mechanism for Mobile Devices
abstract
Application caching is a key feature to enable fast application switches for mobile devices by caching the entire memory pages of applications in the device's physical memory. However, application caching requires a prohibitive amount of memory unless a swap feature is employed to maintain only the working sets of the applications in memory. Unfortunately, mobile devices often disable the invaluable swap feature as it can severely decrease the flash-based local storage device's already marginal lifespan due to the increased writes to the device. As a result, modern mobile devices suffering from the insufficient memory space end up killing memory-hungry applications and keeping only a few applications in the memory. In this paper, we propose CloudSwap, a fast and robust swap mechanism for mobile devices to enable the memory-oblivious application caching. The key idea of CloudSwap is to use the fast local storage as a cache of read-intensive swap pages, while storing prefetch-enabled, write-intensive swap pages in a cloud storage. To preserve the lifespan of the local storage, CloudSwap minimizes the number of writes to the local storage by storing the modified portions of the locally swapped pages in a cloud. To reduce the remote swap-in latency, CloudSwap exploits two cloud-assisted prefetch schemes, the app-aware read-ahead scheme and the access pattern-aware prefetch scheme. Our evaluation shows that the performance of CloudSwap is comparable to a fast, but lifespan-critical local swap system, with only 18% lifespan reduction, compared to the local swap system's 85% lifespan reduction.
Dongju Chae, Joonsung Kim 0001, Youngsok Kim, Jangwoo Kim, Kyung-Ah Chang, Sang-Bum Suh, Hyogun Lee
CCGrid2