EDBT 2026 Demo / reviewers in the wild / expert
Changdae Kim 0001
dblp:26/9796-1
· DBLP profile ↗
12ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0002-9895-5125ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 4 first-author · 4 since 2021Software engineering, systems software and programming languages · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Disaggregated Memory for File-backed PagesabstractTo explore the opportunity of expanding the page cache using disaggregated memory for file-backed pages, this study presents BalloonStasher, an RDMA-based disaggregated memory for data-intensive applications. Utilizing the ephemeral nature of the page cache, BalloonStasher dynamically adapts to the changing page cache demands of multiple clients. BalloonStasher supports one-sided RDMA-based memory pooling or two-sided RDMA-based memory sharing when using memory nodes. Additionally, it also supports peer memory mode, which utilizes the idle memory of peer nodes. Our extensive performance study compares the benefits and limitations of the three cache modes, and shows that BalloonStasher can mitigate the memory underutilization problem and improve the performance of data-intensive applications by a large margin. Daegyu Han, Jaeyoon Nam, Hokeun Cha, Changdae Kim 0001, Kwangwon Koh, Taehoon Kim 0001, Sang-Hoon Kim, Beomseok Nam |
ACM Trans. Storage | 4 |
| 2023 | DEHype: Retrofitting Hypervisors for a Resource-Disaggregated EnvironmentabstractResource disaggregation has been proposed as a solution for resource under-utilization in data centers. However, host virtualization technologies, which are the basic building blocks for constructing data centers, are implemented without considering the disaggregated resources. In addition, we discover that a RDMA I/O unit plays a significant role in the performance of a disaggregated resource environment. In this study, we propose DEHype, which alleviates the inefficiency of the hypervisors utilized in a disaggregated environment by investigating host virtualization technologies that are suitable for disaggregated memory systems. Specifically, DEHype aims to identify and improve the performance issues associated with virtual machines through KVM/QEMU in a disaggregated resource environment. The results demonstrate the effectiveness of the proposed optimizations in improving the performance of disaggregated memory systems. DEHype achieves up to a 351% improvement over the state-of-the-art disaggregated memory system. Taehoon Kim 0001, Kwangwon Koh, Changdae Kim 0001, Eunji Pak, Yeonjeong Jeong, Sang-Hoon Kim |
CLUSTER | 3 |
| 2023 | FusionFlow: Accelerating Data Preparation for Machine Learning with Hybrid CPU-GPU ProcessingabstractData augmentation enhances the accuracy of DL models by diversifying training samples through a sequence of data transformations. While recent advancements in data augmentation have demonstrated remarkable efficacy, they often rely on computationally expensive and dynamic algorithms. Unfortunately, current system optimizations, primarily designed to leverage CPUs, cannot effectively support these methods due to costs and limited resource availability. To address these issues, we introduce FusionFlow, a system that cooperatively utilizes both CPUs and GPUs to accelerate the data preprocessing stage of DL training that runs the data augmentation algorithm. FusionFlow orchestrates data preprocessing tasks across CPUs and GPUs while minimizing interference with GPU-based model training. In doing so, it effectively mitigates the risk of GPU memory overflow by managing memory allocations of the tasks within the GPU-wide free space. Furthermore, FusionFlow provides a dynamic scheduling strategy for tasks with varying computational demands and reallocates compute resources on the fly to enhance training throughput for both single and multi-GPU DL jobs. Our evaluations show that FusionFlow outperforms existing CPU-based methods by 16--285% in single-machine scenarios and, to achieve similar training speeds, requires 50--60% fewer CPUs compared to utilizing scalable compute resources from external servers. Mansur Mukimbekov, Heelim Hong, Ze Jin, Changdae Kim 0001, Ji-Yong Shin, Myeongjae Jeon |
Proc. VLDB Endow. | 7 |
| 2022 | BWA-MEM-SCALE: Accelerating Genome Sequence Mapping on Commodity ServersabstractAs advances in Next-Generation Sequencing have made genome sequence data generation faster and cheaper, the acceleration of genome sequence mapping to the reference genome becomes an increasingly important problem. Much effort has been made to improve the performance of the sequence mapping process. Changdae Kim 0001, Kwangwon Koh, Taehoon Kim 0001, Daegyu Han, Jiwon Seo 0002 |
ICPP | 1 |
| 2021 | Failure-Atomic Byte-Addressable R-tree for Persistent MemoryabstractIn this article, we propose Failure-atomic Byte-addressable R-tree (FBR-tree) that leverages the byte-addressability, persistence, and high performance of persistent memory while guaranteeing the crash consistency. We carefully control the order of store and cacheline flush instructions and prevent any single store instruction from making an FBR-tree inconsistent and unrecoverable. We also develop a non-blocking lock-free range query algorithm for FBR-tree. Since FBR-tree allows read transactions to detect and ignore any transient inconsistent states, multiple read transactions can concurrently access tree nodes without using shared locks while other write transactions are making changes to them. Our performance study shows that FBR-tree successfully reduces the legacy logging overhead and the lock-free range query algorithm shows up to 2.6x higher query processing throughput than the shared lock-based crabbing concurrency protocol. Soojeong Cho, Wonbae Kim, Sehyeon Oh, Changdae Kim 0001, Kwangwon Koh, Beomseok Nam |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2020 | AniFilter: parallel and failure-atomic cuckoo filter for non-volatile memoriesabstractApproximate Membership Query (AMQ) data structures are widely used in databases, storage systems, and other domains. Recent advances in Non-Volatile Memory (NVM) technologies made possible byte-addressable, high performance persistent memories. This paper presents an optimized persistent AMQ for NVM, called AniFilter (AF). Based on Cuckoo Filter (CF), AF improves insertion throughput on NVM with Spillable Buckets and Lookahead Eviction, and lookup throughput with Bucket Primacy. To analyze the effect of our optimizations, we design a probabilistic model to estimate the performance of CF and AF. For failure atomicity, AF writes minimum amount of logging to maintain its consistency. We evaluated AF and four other AMQs - CF, Morton Filter (MF), Rank-and-Select Quotient Filter (RSQF), and Bloom Filter (BF) - on NVM. We use Intel Optane DC Persistent Memory and Quartz emulation for the evaluation. AF and CF are generally much faster than the other filters for sequential runs. However, in high load factors CF's insertion throughput becomes prohibitively low due to the eviction overhead. With our optimizations, AF's insertion throughput is fastest even in high load factors. In parallel evaluation, AF's performance is substantially higher than CF for both insertion and lookup. Our optimizations reduce the bandwidth consumption, making AF's parallel performance much faster than CF's on bandwidth-limited NVM. For parallel insertion AF is up to 10.7X faster than CF (2.6X faster on average) and for parallel lookup AF is up to 1.2X faster (1.1X faster on average). Hyungjun Oh, Bongki Cho, Changdae Kim 0001, Heejin Park, Jiwon Seo 0002 |
EuroSys | 3 |
| 2020 | Charge-Aware DRAM Refresh Reduction with Value TransformationabstractAs the memory capacity in a system has been growing, refresh operations consume increasing ratios of the total DRAM power. To reduce the power consumption of such refresh operations, this paper proposes a novel value-aware refresh reduction technique called ZERO - REFRESH which exploits zero values in memory contents. A DRAM cell can retain the discharged state without refresh operations, and ZERO - REFRESH skips refresh operations on rows with all discharged cells. For abundant unallocated memory pages in typical systems, the operating system fills them with zeros to clean the contents. For those idle pages, ZERO - REFRESH can eliminate refresh operations in an OS-transparent way without any new interface to DRAM. However, for allocated memory pages, memory contents may not have many consecutive zero values to match the refresh granularity of DRAM. To increase the frequency of zero values and to arrange them to match the refresh granularity, ZERO - REFRESH transforms the value of memory blocks to the base and delta values, inspired by the prior BDI (Base-Delta-Immediate) compression technique. Once values are converted, bits are transposed to be stored as consecutive discharged bits at the refresh granularity. Such value transformation and rearrangement can make the memory contents friendly to refresh reduction based on discharged cells. The experimental results based on simulation show that the DRAM refresh operations are reduced by 37% on average for a set of benchmark applications, if the entire memory is allocated for the applications. If the memory usage statistics collected from three data center traces are applied, the DRAM refresh operations can be reduced by 46%, 57%, and 83% respectively for the three scenarios. Seikwon Kim, Wonsang Kwak, Changdae Kim 0001, Daehyeon Baek, Jaehyuk Huh 0001 |
HPCA | 3 |
| 2019 | GVTS: Global Virtual Time Fair Scheduling to Support Strict Fairness on Many CoresabstractProportional fairness in CPU scheduling has been widely adopted to fairly distribute CPU shares corresponding to their weights. With the emergence of cloud environments, the proportionally fair scheduling has been extended to groups of threads or nested groups to support virtual machines or containers. Such proportional fairness has been supported by popular schedulers, such as Linux Completely Fair Scheduler (CFS) through virtual time scheduling. However, CFS, with a distributed runqueue per CPU, implements the virtual time scheduling locally. Across different queues, the virtual times of threads are not strictly maintained to avoid potential scalability bottlenecks. The uneven fluctuation of CPU shares caused by the limitations of CFS not only violates the fairness support for CPU assignments, but also significantly increases the tail latencies of latency-sensitive applications. To mitigate the limitations of CFS, this paper proposes a global virtual-time fair scheduler (GVTS), which enforces global virtual time fairness for threads and thread groups, even if they run across many physical cores. The new scheduler employs the hierarchical enforcement of target virtual time to enhance the scalability of schedulers, which is aware of the topology of CPU organization. We implemented GVTS in Linux kernel 4.6.4 with several optimizations to provide global virtual time efficiently. Our experimental results show that GVTS can almost eliminate the fairness violation of CFS for both non-grouped and grouped executions. Furthermore, GVTS can curtail the tail latency when latency-sensitive applications are co-running with batch tasks. Changdae Kim 0001, Seungbeom Choi, Jaehyuk Huh 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2018 | Exploring the Design Space of Fair Scheduling Supports for Asymmetric Multicore SystemsabstractAlthough traditional CPU scheduling efficiently utilizes multiple cores with equal computing capacity, the advent of multicores with diverse capabilities pose challenges to CPU scheduling. For such asymmetric multi-core systems, scheduling is essential to exploit the efficiency of core asymmetry, by matching each application with the best core type. However, in addition to the efficiency, an important aspect of CPU scheduling is fairness in CPU provisioning. Such uneven core capability is inherently unfair to threads and causes performance variance, as applications running on fast cores receive higher capability than applications on slow cores. Depending on co-running applications and scheduling decisions, the performance of an application may vary significantly. This study investigates the fairness problem in asymmetric multi-cores, and explores the design space of OS schedulers supporting multiple fairness constraints. In this paper, we consider two fairness-oriented constraints, minimum fairness for the minimum guaranteed performance and uniformity for performance variation reduction. This study proposes four scheduling policies which guarantee a minimum performance bound while improving the overall throughput and reducing performance variation too. The proposed fairness-oriented schedulers are implemented for the Linux kernel with an online application monitoring technique. Using an emulated asymmetric multi-core with frequency scaling and a real asymmetric multi-core with the big.LITTLE architecture, the paper shows that the proposed schedulers can effectively support the specified fairness while improving overall system throughput. Changdae Kim 0001, Jaehyuk Huh 0001 |
IEEE Trans. Computers | 1 |
| 2017 | Configuration Guidance Framework for Molecular Dynamics Simulations in Virtualized ClustersabstractWith the advancement of cloud computing, there has been a growing interest in exploiting demand-based cloud resources for parallel scientific applications. To satisfy different needs for computing resources, cloud providers provide many different types of virtual machines (VMs) with various numbers of computing cores and amounts of memory. The cost and execution time of a scientific application vary depending on the types of VMs, number of VMs, and current status of the cloud due to interference among VMs. However, currently, cloud users are solely responsible for selecting the most effective VM configuration for their needs, but often end up with sub-optimal selections. In this paper, using molecular dynamics simulations as a case study, we propose a framework to guide users to select the optimal VM configurations that satisfy their requirements for scientific parallel computing in virtualized clusters. For molecular dynamics computation on a cluster of VMs, the guidance framework uses artificial neural networks which are trained to predict its execution times for various inputs, VM configurations, and status of interference among VMs. Using our performance prediction mechanisms, the guidance framework helps users choose an optimal or near-optimal VM cluster configuration under cost and runtime constraints. Jaeung Han, Changdae Kim 0001, Jaehyuk Huh 0001, Gil-Jin Jang, Young-ri Choi |
IEEE Trans. Serv. Comput. | 2 |
| 2016 | Fairness-oriented OS Scheduling Support for Multicore SystemsabstractAlthough traditional CPU scheduling efficiently utilizes multiple cores with equal computing capacity, the advent of multicores with diverse capabilities pose challenges to CPU scheduling. For the multi-cores with uneven computing capability, scheduling is essential to exploit the efficiency of core asymmetry, by matching each application with the best core type. However, in addition to the efficiency, an important aspect of CPU scheduling is fairness in CPU provisioning. Such uneven core capability is inherently unfair to threads and causes performance variance, as applications running on fast cores receive higher capability than applications on slow cores. Depending on co-running applications and scheduling decisions, the performance of an application may vary significantly. This study investigates the fairness problem in multi-cores with uneven capability, and explores the design space of OS schedulers supporting multiple fairness constraints. In this paper, we consider two fairness-oriented constraints, minimum fairness for the minimum guaranteed performance and uniformity for performance variation reduction. This study proposes three scheduling policies which guarantee a minimum performance bound while improving the overall throughput and reducing performance variation too. The three proposed fairness-oriented schedulers are implemented for the Linux kernel with an online application monitoring technique. Using an emulated asymmetric multi-core with frequency scaling and a real asymmetric multi-core with the big.LITTLE architecture, the paper shows that the proposed schedulers can effectively support the specified fairness while improving overall system throughput. Changdae Kim 0001, Jaehyuk Huh 0001 |
ICS | 1 |
| 2011 | Virtualizing performance asymmetric multi-core systemsabstractPerformance-asymmetric multi-cores consist of heterogeneous cores, which support the same ISA, but have different computing capabilities. To maximize the throughput of asymmetric multi-core systems, operating systems are responsible for scheduling threads to different types of cores. However, system virtualization poses a challenge for such asymmetric multi-cores, since virtualization hides the physical heterogeneity from guest operating systems. In this paper, we explore the design space of hypervisor schedulers for asymmetric multi-cores, which do not require asymmetry-awareness from guest operating systems. The proposed scheduler characterizes the efficiency of each virtual core, and map the virtual core to the most area-efficient physical core. In addition to the overall system throughput, we consider two important aspects of virtualizing asymmetric multi-cores: performance fairness among virtual machines and performance scalability for changing availability of fast and slow cores. Youngjin Kwon, Changdae Kim 0001, Seung Ryoul Maeng, Jaehyuk Huh 0001 |
ISCA | 2 |