VLDB 2026 Research / reviewers in the wild / expert
Hyokyung Bahn
dblp:36/1121
· DBLP profile ↗
48ranked-venue papers
3as first author
7since 2021 · last 2027
0000-0002-7188-3889ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 since 2021Software engineering, systems software and programming languages · 6 · 1 first-author · 1 since 2021Computer networks · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3Artificial intelligence and machine learning · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Adaptive write-cache management for alternating deep learning and non-AI workloads on mobile devices
Jeongha Lee, Hyokyung Bahn |
Future Gener. Comput. Syst. | 2 |
| 2026 | Secure and Energy-Efficient Resource Optimization Framework for Real-Time Task Execution in Industrial IoTabstractThis paper presents a secure and energy-efficient resource optimization framework for Industrial IoT (IIoT) task scheduling. Although prior studies have explored various aspects of cost, performance, energy efficiency, compatibility, and security, few have incorporated all these factors into a unified framework. To consider computational cost while satisfying real-time deadline constraints, we adopt public edge servers and ensure secure execution against potential threats from the host OS and privileged administrators. Unlike conventional application-level TEEs (Trusted Execution Environments), our framework extends security to the entire virtual machine level, enabling seamless compatibility between local execution and offloaded tasks. Efficient task scheduling requires the simultaneous optimization of processor voltage scaling, memory placement, and offloading decisions based on workload intensity and resource conditions. Through a detailed analysis, we identify strong interdependencies among system resources and propose a task model that accounts for these interactions. By incorporating fine-grained execution time and energy models, which precisely capture processor-memory interactions, parallelism between local and remote executions, and TEE-added offloading overheads, our framework guarantees real-time deadline compliance while minimizing energy consumption. Additionally, it enables secure and compatible IIoT deployment, effectively balancing cost, performance, security, and energy efficiency. Yeonsu Do, Kyungwoon Cho, Yeonsu Park 0002, Hyokyung Bahn |
IEEE Internet Things J. | 4 |
| 2026 | Mutual Edge: A New Alternative Beyond Private and Public Edge in IIoT Resource ProvisioningabstractThis paper presents mutual edge, a new enterprise-driven edge resource provisioning paradigm for Industrial IoT systems, which addresses the limitations of both private and public edge models. Traditional resource provisioning often relies on overallocated private hardware or costly public edge rentals, which impose significant financial burdens. As an alternative, we propose the mutual edge model, where each industrial enterprise provisions a fraction of workloads using its own hardware and leases or lends resources across non-overlapping peak periods. To support this paradigm, we formulate a cost model that jointly considers hardware depreciation, electricity cost for each resource component, and leasing revenue from idle resources. We then formalize the problem as a constrained optimization, where the objective is to minimize execution cost while strictly meeting real-time task deadlines. Since this paradigm shift introduces challenges such as safe execution on mutually untrusted systems, we design and integrate security- and economics-aware mechanisms within a unified framework to address these challenges. Extensive simulations show that the proposed mutual edge framework reduces average computing cost by 78% and 56% compared to public and private edge models, respectively, under dynamic task sets and realistic IIoT scenarios. Gahyeon Kwon, Kyungwoon Cho, Hyokyung Bahn |
IEEE Internet Things J. | 3 |
| 2025 | Analyzing Data Access Characteristics of AIoT Workloads for Efficient Write Buffer ManagementabstractAs deep learning technologies increasingly influence various aspects of human life, the emerging paradigm of Artificial Intelligence of Things (AIoT) is gaining significant attention. AIoT, which relies on large datasets and complex models, poses significant challenges to the limited resources of mobile systems. Many solutions have been proposed to address these challenges, including computation offloading to edge or cloud servers. In this article, we demonstrate that analyzing data access patterns and managing them efficiently are also crucial for further enhancing the performance of AIoT systems. Specifically, we propose an efficient write buffer management scheme tailored to mobile AIoT workloads. Our approach is based on an extensive analysis of data access patterns, revealing that the conventional buffer cache architectures and algorithms are inefficient for AIoT workloads due to access characteristics such as write dominance and long repetitive loops. To address these issues, we adopt a non-volatile write buffer at the storage layer and judiciously manage write loops. The key contributions of our study are as follows: 1) Despite relying solely on limited flush information at the storage layer, our scheme achieves performance comparable to full loop detection at the host system. 2) Our scheme requires no modifications to host OS semantics or buffer cache interfaces, ensuring high adaptability. 3) Experimental results demonstrate that our scheme significantly reduces storage traffic, outperforms representative algorithms such as LRU, LFU, LIRS, CLOCK-Pro, and LeCaR, and enhances data access efficiency for AIoT workloads. Jeongha Lee, Soojung Lim, Hyokyung Bahn |
IEEE Internet Things J. | 3 |
| 2023 | Co-Optimizing CPU Voltage, Memory Placement, and Task Offloading for Energy-Efficient Mobile SystemsabstractEnergy saving is one of the most important missions in the design of battery-based mobile systems. Many ideas have been suggested for saving energy in different system layers. Specifically: 1) lowering the supply voltage for idle CPU time slots; 2) using hybrid memory to save DRAM refresh power; and 3) task offloading to edge/cloud servers are well-acknowledged techniques used in CPU, memory, and network subsystems. In this article, we show that co-optimizing these three techniques is necessary for further reducing the energy consumption of mobile real-time systems. To this end, we present an extended task model and formulate the effect of dynamic voltage/frequency scaling (DVFS), hybrid memory allocation, and task offloading problems as a unified measure. We then present a new real-time task scheduling scheme, called Co-TOMS, to co-optimize the energy-saving techniques in CPU, memory, and network subsystems by considering the given task set and resource conditions. The main contributions of our study can be summarized as follows. First, we optimize three energy-saving techniques across different system layers and find that they have significant influence on each other. For example, the effect of DVFS alone is limited in mobile systems, but combining it with offloading greatly amplifies its efficiency. Second, previous studies on offloading usually define a deadline as the maximum allowable latency at the application level, but we focus on hard real-time systems that must meet task-level deadlines. Third, we design a steady-state genetic algorithm that allows fast convergence with reasonable computation overhead under various resource and workload conditions. Soomin Ki, Gyuri Byun, Kyungwoon Cho, Hyokyung Bahn |
IEEE Internet Things J. | 4 |
| 2022 | Classification and Characterization of Memory Reference Behavior in Machine Learning WorkloadsabstractWith the recent penetration of artificial intelligence (AI) technologies into many areas of computing, machine learning is being incorporated into modern software design. As the in-memory data of AI workloads increasingly grows, it is important to characterize memory reference behaviors in machine learning workloads. In this paper, we perform a characterization study for memory references in machine learning workloads as the learning types (i.e., supervised vs. unsupervised) and the problem domains (i.e., classification, regression, and clustering) are varied. From this study, we uncover the following five characteristics. First, machine learning workloads exhibit significantly different memory reference patterns from traditional workloads, but they are similar regardless of learning types and problem domains. Second, in all workloads, memory reads and writes continue to appear for a wide range of memory addresses, but there is a specific time period where only reads appear. Third, among references to memory areas (i.e., code, data, heap, stack, library), library accounts for about 90% of total memory references. Fourth, there is a low popularity bias between memory pages referenced in machine learning workloads, especially for writes. Fifth, when estimating the likelihood of re-referencing, temporal locality is dominant in top 100 memory pages, but access frequency provides better information after that ranking. It is expected that the characterization of memory references conducted in this paper will be helpful in the design of memory management policies for machine learning workloads. Seokmin Kwon, Hyokyung Bahn |
SNPD | 2 |
| 2022 | Integrated Scheduling of Real-Time and Interactive Tasks for Configurable Industrial SystemsabstractWith the recent advances in Internet of Things and cyber-physical systems technologies, smart industrial systems support configurable processes consisting of human interactions as well as hard real-time functions. This implies that irregularly arriving interactive tasks and traditional hard real-time tasks coexist. As the characteristics of the tasks are heterogeneous, it is not an easy matter to schedule them all at once. To cope with this situation, this article presents a new task scheduling policy that uses the notion of “virtual real-time task” and two-phase scheduling. As hard real-time tasks must keep their deadlines, we perform offline scheduling based on genetic algorithms beforehand. This determines the processor's voltage level and memory location of each task and also reserves the virtual real-time tasks for interactive tasks. When interactive tasks arrive during the execution, online scheduling is performed on the time slot of the virtual real-time tasks. As interactive workloads evolve over time, we monitor them and periodically update the offline scheduling. Experimental results show that the proposed policy reduces the energy consumption by 66.8% on average without deadline misses and also supports the waiting time of less than 3 (s) for interactive tasks. Suhyeon Yoo, Yewon Jo, Hyokyung Bahn |
IEEE Trans. Ind. Informatics | 3 |
| 2020 | Student Session: Power-Saving Integrated Task Scheduling in Multicore and Hybrid Memory EnvironmentabstractA new task scheduling algorithm that schedules mixed task set consisting of real-time and interactive tasks is presented. Our algorithm aims at minimizing the power consumption of the system with the reasonable response time of interactive tasks as well as the deadline guarantees of real-time tasks. Experimental results show that the proposed algorithm improves the power consumption by 23% on average. Yewon Jo, Suhyeon Yoo, Hyokyung Bahn |
RTCSA | 3 |
| 2017 | Combining memory allocation and processor volatage scaling for energy-efficient IoT task schedulingabstractAs IoT (Internet-of-things) technologies grow rapidly for emerging applications such as smart living and health care, reducing power consumption in battery-based IoT devices becomes an important issue. An IoT device is a kind of real-time systems, of which power-savings have been widely studied in terms of processor's dynamic voltage/frequency scaling. However, recent research has shown that memory subsystems are getting reached to a significant portion of power consumption in such systems. In this paper, we show that power consumption of real-time systems can be further reduced by combining voltage/frequency scaling with task allocation in hybrid memory. If a task set is schedulable in a low voltage mode of a processor, we can expect that the task set will still be schedulable even with slow memory. By considering this, we adopt non-volatile memory technologies that consume less power than DRAM but provide relatively slow access latency. Our aim is to allocate tasks in non-volatile memory if it does not violate the deadline constraint of real-time tasks, thereby reducing the power consumption of the system further. To do so, we incorporate the memory allocation problem into the problem model of processor's voltage scaling, and evaluate the effectiveness of the combined approach. Sunhwa Annie Nam, Kyungwoon Cho, Hyokyung Bahn |
ICIS | 3 |
| 2017 | Exploiting write-only-once characteristics of file data in smartphone buffer cache management
Hyokyung Bahn |
Pervasive Mob. Comput. | 2 |
| 2017 | Exploiting I/O Reordering and I/O Interleaving to Improve Application Launch PerformanceabstractApplication prefetchers improve application launch performance through either I/O reordering or I/O interleaving. However, there has been no proposal to combine the two techniques together, missing the opportunity for further optimization. We present a new application prefetching technique to take advantage of both the approaches. We evaluated our method with a set of applications to demonstrate that it reduces cold start application launch time by 50%, which is an improvement of 22% from the I/O reordering technique. Yongsoo Joo, Sangsoo Park, Hyokyung Bahn |
ACM Trans. Storage | 3 |
| 2017 | Reducing Write Amplification of Flash Storage through Cooperative Data Management with NVMabstractWrite amplification is a critical factor that limits the stable performance of flash-based storage systems. To reduce write amplification, this article presents a new technique that cooperatively manages data in flash storage and nonvolatile memory (NVM). Our scheme basically considers NVM as the cache of flash storage, but allows the original data in flash storage to be invalidated if there is a cached copy in NVM, which can temporarily serve as the original data. This scheme eliminates the copy-out operation for a substantial number of cached data, thereby enhancing garbage collection efficiency. Simulated results show that the proposed scheme reduces the copy-out overhead of garbage collection by 51.4% and decreases the standard deviation of response time by 35.4% on average. Measurement results obtained by implementing the proposed scheme in BlueDBM, 1 an open-source flash development platform developed by MIT, show that the proposed scheme reduces the execution time and increases IOPS by 2--21% and 3--18%, respectively, for the workloads that we considered. This article is an extended version of Lee et al. [2016], which was presented at the 32nd International Conference on Massive Data Storage Systems and Technology in 2016. Julie Kim, Hyokyung Bahn, Sungjin Lee 0001, Sam H. Noh |
ACM Trans. Storage | 3 |
| 2016 | Reducing write amplification of flash storage through Cooperative Data Management with NVMabstractWrite amplification is a critical factor that limits the stable performance of flash-based storage systems. To reduce write amplification, this paper presents a new technique that cooperatively manages data in flash storage and nonvolatile memory (NVM). Our scheme basically considers NVM as the cache of flash storage, but allows the original data in flash storage to be invalidated if there is a cached copy in NVM, which can temporarily serve as the original data. This scheme eliminates the copy-out operation for a substantial number of cached data, thereby enhancing garbage collection efficiency. Experimental results show that the proposed scheme reduces the copy-out overhead of garbage collection by 51.4% and decreases the standard deviation of response time by 35.4% on average. Julie Kim, Hyokyung Bahn, Sam H. Noh |
MSST | 3 |
| 2016 | An Energy-Efficient Positioning Scheme for Location-Based Services in a SmartphoneabstractAs Location-Based Services (LBSs) are widely used in smartphone applications, Global Positioning System (GPS) becomes one of the main sources of a smartphone's energy consumption. This paper presents an energy-efficient positioning scheme for smartphones called EEPS (Energy-Efficient Positioning Scheme). EEPS adaptively performs the positioning of a smartphone considering the interesting regions of a user, accuracy requirement of each application executed within the smartphone, and the battery level. Simulations with various real applications and scenarios show that EEPS reduces 50.1% of energy consumption compared to GPS. Nevertheless, it satisfies the accuracy requirement of each application. Soyoon Lee, Hyokyung Bahn |
RTCSA | 3 |
| 2016 | Design and Implementation of Kernel Binder Cache to Accelerate Android IPCabstractAndroid supports a variety of service functions via user-level daemons. As applications invoke these service functions through IPC (Inter-Process Communication), accelerating IPC performance is critical to the responsiveness in Android. However, Android provides long IPC latency of more than 240us due to complicated software stacks between the kernel Binder and the user-level process Context Manager. Specifically, Android supports IPC via a virtual device driver called Binder, but actual tasks for IPC such as service function search and management are performed by a user-level process called Context Manager. This separation is advantageous in terms of modularity and flexibility, but degrades the responsiveness of services significantly due to additional context switching and inefficient request handling. To resolve this issue, we analyze the end-to-end path of Android IPC mechanisms and observe that 55% of IPC latency is due to the communication overhead between Binder and Context Manager. Based on this observation, we design and implement a kernel Binder cache that maintains a hot subset of service function mappings, thereby reducing requests transferred to Context Manager. The proposed Binder cache is implemented on Android 5.0 Lollipop. Measurement studies on Nexus 5 show that our Binder cache accelerates IPC performance twofold compared to current Android. Jesung Yeon, Jinseok Im, Byungjun Jeon, Hyokyung Bahn |
RTCSA | 6 |
| 2016 | An efficient page replacement algorithm for PCM-based mobile embedded systemsabstractTraditional embedded systems do not use virtual memory swapping but load the entire footprint into memory as they are usually single task special-purpose machines. Recently, as mobile embedded systems support multi-tasking, the necessity of swapping is becoming increasingly important. However, current mobile systems such as smartphones do not support swapping because flash memory has weaknesses to be a swap device in such environments. More recently, phase-change memory (PCM) emerges as an alternative medium for the swap device of mobile embedded systems. In this paper, we present a new page replacement algorithm for a mobile embedded system that uses PCM as a swap device. Although PCM provides high performance and byte-accessibility, its write operation is slow and it accommodates only limited endurance cycles. To cope with this situation, our algorithm tracks the dirtiness of a page at the granularity of a sub-page and replaces the least dirty page among pages not recently used, leading to reduced write traffic to PCM. Experimental results with various mobile workloads show that the proposed algorithm reduces the amount of data written to PCM by 24% on average and up to 74% compared to the well-known CLOCK algorithm. It also extends the lifetime of PCM by 50% on average. Seunghoon Yoo, Hyokyung Bahn |
SNPD | 2 |
| 2016 | Reducing Journaling Harm on Virtualized I/O SystemsabstractThis paper analyzes the host cache effectiveness in full virtualization, particularly associated with journaling of guests. We observe that the journal access of guests degrades cache performance largely due to the write-once access pattern and the frequent sync operations. To remedy this problem, we design and implement a novel caching policy, called PDC (Pollution Defensive Caching), that detects the journal accesses and prevents them from entering the host cache. The proposed PDC is implemented in QEMU-KVM 2.1 on Linux 4.14 and provides 3-32% performance improvement for various file and I/O benchmarks. Hyokyung Bahn, Minseong Jeong, Jesung Yeon, Seunghoon Yoo, Sam H. Noh, Kang G. Shin |
SYSTOR | 2 |
| 2016 | An Adaptive Location Detection scheme for energy-efficiency of smartphones
Soyoon Lee, Hyokyung Bahn |
Pervasive Mob. Comput. | 3 |
| 2016 | Eliminating Periodic Flush Overhead of File I/O with Non-Volatile Buffer CacheabstractFile I/O buffer caching plays an important role to narrow the wide speed gap between main memory and secondary storage. However, data loss or inconsistency may occur if the system crashes before updated data in the buffer cache is flushed to storage. Thus, most operating systems adopt a daemon that periodically flushes dirty data to secondary storage. This periodic flush degrades the caching efficiency seriously because most write requests lead to direct storage accesses. We show that periodic flush accounts for 30-70 percent of the total write traffic to storage. To remove this inefficiency, this paper presents a new buffer cache architecture that adopts a small amount of non-volatile memory to maintain modified data. This novel buffer cache architecture removes almost all storage accesses due to periodic flush operations without any loss of reliability. It also improves the buffer cache performance through space-efficient management techniques, such as delta-write and fragment-grouping. Our experimental results show that the proposed buffer cache reduces the storage write traffic by 44.3 percent and also improves the throughput by 23.6 percent on average. Hyojung Kang, Hyokyung Bahn, Kang G. Shin |
IEEE Trans. Computers | 3 |
| 2015 | Management of Virtual Memory Systems under High Performance PCM-based Swap DevicesabstractAs high performance NVM storage such as PCM or STT-RAM emerges, legacy software layers optimized for HDDs should be revisited. This paper explores the performance of a system that uses PCM as the swap device of virtual memory and discusses how this system can be managed efficiently. Specifically, we explore the challenges and implications of using PCM swap devices with a broad range of experiments in comparison with traditional HDD swap devices. For PCM swap devices, we evaluate the performance by attaching PCM on DIMM slots, thereby eliminating the software stack overhead of block I/O and the context switch time. To assess the potential benefit of PCM swap devices, we change various configurations and perform experiments to quantify the effects of memory management mechanisms and parameters such as page size, read-ahead options, page replacement algorithms, and total memory capacity. The results show that reducing the page size and turning off the read-ahead option improve the performance of PCM-based swap systems where the page fault handling time is sufficiently small. We also show that the performance is not degraded even with a small DRAM memory under a PCM swap device, this leads to the reduction of DRAM's energy consumption significantly compared to HDD-based swap systems. We anticipate that our results will provide directions in system software development in presence of ever faster swap devices. Yunjoo Park, Hyokyung Bahn |
COMPSAC | 2 |
| 2015 | Design and Implementation of a Journaling File System for Phase-Change MemoryabstractJournaling file systems are widely used in modern computer systems as they provide high reliability at reasonable cost. However, existing journaling file systems are not efficient for emerging PCM (phase-change memory) storage because they are optimized for hard disks. Specifically, the large amount of data that they write during journaling degrades the performance of PCM storage seriously as it has a long write latency. In this paper, we present a new journaling file system for PCM, called Shortcut-JFS, that reduces write traffic to PCM by more than half of existing journaling file systems running on block I/O interfaces. To do this, we devise two novel schemes that can be used under byte-addressable I/O interfaces: 1) differential logging that journals only the modified part of a block and 2) in-place checkpointing that eliminates the overhead of block copying. We implement Shortcut-JFS on Linux 2.6.32 and measure the performance of Shortcut-JFS compared to those of existing journaling and log-structured file systems. The results show that the performance improvement of Shortcut-JFS against Ext4 and LFS is 54 and 96 percent, respectively, on average. Seunghoon Yoo, Hyokyung Bahn |
IEEE Trans. Computers | 3 |
| 2014 | Empirical Study of NVM Storage: An Operating System's Perspective and ImplicationsabstractAs high performance NVM storage such as PCM and STT-RAM emerge, legacy software layers optimized for HDDs should be revisited. Specifically, as storage performance approaches DRAM performance, existing I/O mechanisms and software configurations should be reassessed. This paper explores the challenges and implications of using NVM storage with a broad range of experiments. We measure the performance of a system with NVM storage emulated by DRAM with proper timing parameters and compare it with that of HDD storage environments under various configurations. Our experimental results show that even with storage as fast as DRAM, the performance gain is not large for read operations as current I/O mechanisms do a good job of hiding the slow performance of HDD. To assess the potential benefit of fast storage media, we change various I/O configurations and perform experiments to quantify the effects of existing I/O mechanisms such as buffer caching, read-ahead, synchronous I/O, direct I/O, block I/O, and byte-addressable I/O on systems with NVM storage. We also investigate some unique performance characteristics of NVM in comparison with HDD by changing the number of accesses and the amount of data to be transferred. We anticipate that our results will provide directions in system software development in presence of ever faster storage devices. Hyokyung Bahn, Seunghoon Yoo, Sam H. Noh |
MASCOTS | 2 |
| 2014 | CLOCK-DWF: A Write-History-Aware Page Replacement Algorithm for Hybrid PCM and DRAM Memory ArchitecturesabstractPhase change memory (PCM) has emerged as one of the most promising technologies to incorporate into the memory hierarchy of future computer systems. However, PCM has two critical weaknesses to substitute DRAM memory in its entirety. First, the number of write operations allowed to each PCM cell is limited. Second, write access time of PCM is about 6–10 times slower than that of DRAM. To cope with this situation, hybrid memory architectures that use a small amount of DRAM together with PCM have been suggested. In this paper, we present a new memory management technique for hybrid PCM and DRAM memory architecture that efficiently hides the slow write performance of PCM. Specifically, we aim to estimate future write references accurately and then absorb frequent memory writes into DRAM. To do this, we analyze the characteristics of memory write references and find two noticeable phenomena. First, using write history alone performs better than using both read and write history in estimating future write references. Second, the frequency characteristic is a better estimator than temporal locality in predicting future memory writes. Based on these two observations, we present a new page replacement algorithm called CLOCK-DWF (CLOCK with Dirty bits and Write Frequency) that significantly reduces the number of write operations that occur on PCM and also increases the lifespan of PCM memory. Soyoon Lee, Hyokyung Bahn, Sam H. Noh |
IEEE Trans. Computers | 2 |
| 2014 | Caching Strategies for High-Performance Storage MediaabstractDue to the large access latency of hard disks during data retrieval in computer systems, buffer caching mechanisms have been studied extensively in database and operating systems. By storing requested data into the buffer cache, subsequent requests can be directly serviced without accessing slow disk storage. Meanwhile, high-speed storage media like PCM (phase-change memory) have emerged recently, and one may wonder if the traditional buffer cache will be still effective for these high-speed storage media. This article answers the question by showing that the buffer cache is still effective in such environments due to the software overhead and the bimodal data access characteristics. Based on this observation, we present a new buffer cache management scheme appropriately designed for the system where the speed gap between cache and storage is narrow. To this end, we analyze the condition that caching will be effective and find the characteristics of access patterns that can be exploited in managing buffer cache for high performance storage like PCM. Hyokyung Bahn |
ACM Trans. Storage | 2 |
| 2014 | A Unified Buffer Cache Architecture that Subsumes Journaling Functionality via Nonvolatile MemoryabstractJournaling techniques are widely used in modern file systems as they provide high reliability and fast recovery from system failures. However, it reduces the performance benefit of buffer caching as journaling accounts for a bulk of the storage writes in real system environments. To relieve this problem, we present a novel buffer cache architecture that subsumes the functionality of caching and journaling by making use of nonvolatile memory such as PCM or STT-MRAM. Specifically, our buffer cache supports what we call the in-place commit scheme. This scheme avoids logging, but still provides the same journaling effect by simply altering the state of the cached block to frozen. As a frozen block still provides the functionality of a cache block, we show that in-place commit does not degrade cache performance. We implement our scheme on Linux 2.6.38 and measure the throughput and execution time of the scheme with various file I/O benchmarks. The results show that our scheme improves the throughput and execution time by 89% and 34% on average, respectively, compared to the existing Linux buffer cache with ext4 without any loss of reliability. Hyokyung Bahn, Sam H. Noh |
ACM Trans. Storage | 2 |
| 2013 | Unioning of the buffer cache and journaling layers with non-volatile memory
Hyokyung Bahn, Sam H. Noh |
FAST | 2 |
| 2013 | On-Demand Snapshot: An Efficient Versioning File System for Phase-Change MemoryabstractVersioning file systems are widely used in modern computer systems as they provide system recovery and old data access functions by retaining previous file system snapshots. However, existing versioning file systems do not perform well with the emerging PCM (phase-change memory) storage, because they are optimized for hard disks. Specifically, a large amount of additional writes incurred by maintaining snapshot degrades the performance of PCM seriously as write operations are the performance bottleneck of PCM. This paper presents a novel versioning file system, designed for PCM, that reduces the writing overhead of a snapshot significantly. Unlike existing versioning file systems that incur cascade writes up to the file system root, our scheme breaks the recursive update chain at the immediate parent level. The proposed file system is implemented on Linux 2.6 as a prototype. Measurement studies with various I/O benchmarks show that the proposed file system improves the I/O throughput by 144 percent on average, compared to ZFS, a representative versioning file system. Jee-Eun Jang, Taeseok Kim, Hyokyung Bahn |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2012 | Shortcut-JFS: A write efficient journaling file system for phase change memoryabstractJournaling file systems are widely used in modern computer systems as it provides high reliability with reasonable performance. However, existing journaling file systems are not efficient for emerging PCM (Phase Change Memory) storage. Specifically, a large amount of write operations performed by journaling incur serious performance degradation of PCM storage as it has long write latency. In this paper, we present a new journaling file system for PCM, called Shortcut-JFS, that reduces write amount of journaling by more than a half exploiting the byte-accessibility of PCM. Specifically, Shortcut-JFS performs two novel schemes, 1) differential logging that performs journaling only for modified bytes and 2) in-place checkpointing that removes unnecessary block copy overhead. We implemented Shortcut-JFS on Linux 2.6, and measured the performance of Shortcut-JFS and legacy journaling schemes used in ext 3. The results show that the performance improvement of Shortcut-JFS against ext 3 is 40% on average. Seunghoon Yoo, Jee-Eun Jang, Hyokyung Bahn |
MSST | 4 |
| 2011 | Is Buffer Cache Still Effective for High Speed PCM (Phase Change Memory) Storage?abstractRecently, PCM (phase change memory) emerges as a new storage media and there is a bright prospect that PCM will be used as a storage device in the near future. Since the optimistic access time of PCM is expected to be almost identical to that of DRAM, we can make a question that the traditional buffer cache will be still effective for high speed secondary storage such as PCM. This paper answers it by showing that the buffer cache is still effective in such environments due to the software overhead and the bimodal block reference characteristics. Based on this observation, we present a new buffer cache management scheme appropriately for the system where the speed gap between the cache and storage is small. To this end, we analyze the condition that caching gains and find the characteristics of I/O traces that can be exploited in managing buffer cache for PCM-storage. Daeja Jin, Kern Koh, Hyokyung Bahn |
ICPADS | 4 |
| 2011 | Characterizing Memory Write References for Efficient Management of Hybrid PCM and DRAM MemoryabstractIn order to reduce the energy dissipation in main memory of computer systems, phase change memory (PCM) has emerged as one of the most promising technologies to incorporate into the memory hierarchy. However, PCM has two critical weaknesses to substitute DRAM memory in its entirety. First, the number of write operations allowed to each PCM cell is limited. Second, write access time of PCM is about 6-10 times slower than that of DRAM. To cope with this situation, hybrid memory architectures that use a small amount of DRAM together with PCM memory have been suggested. In this paper, we present a new memory management technique for hybrid PCM and DRAM memory architecture that efficiently hides the slow write performance of PCM. Specifically, we aim to estimate future write references accurately and then absorb most memory writes into DRAM. To do this, we analyze the characteristics of memory write references and find two noticeable phenomena. First, using write history alone performs better than using both read and write history in estimating future write references. Second, the frequency characteristic is a better estimator than temporal locality but combining these two properties appropriately leads to even better results. Based on these two observations, we present a new page replacement algorithm called CLOCK-DWF (CLOCK with Dirty bits and Write Frequency) that significantly reduces the number of write operations that occur on PCM. Soyoon Lee, Hyokyung Bahn, Sam H. Noh |
MASCOTS | 2 |
| 2011 | A Demand-Based FTL Scheme Using Dualistic Approach on Data Blocks and Translation BlocksabstractUsing NAND flash memory as a storage device is in the limelight due to its many attractive features, but it also has vulnerable points. Specifically, as NAND flash memory does not allow the overwrite of data in the same place, it performs out-place-update, which requires the address translation table between logical and physical addresses. Due to the ever growing size of NAND flash memory, keeping the whole address translation table in SRAM is becoming increasingly a serious problem. In this paper, we present three management schemes to reduce the SRAM space in address translation but also guarantee the performance. First, we store data in NAND flash memory by using a page level mapping scheme. A page level mapping scheme allows NAND flash memory to store data in any place, and thus we can improve the storage efficiency. Second, we keep only a small amount of address translation entries in the page address translation cache (PATC) to reduce the size of SRAM. The other address translation entries that are in NAND flash memory will be loaded in SRAM on demand. Furthermore, we manage an address translation table in NAND flash memory by using a hybrid mapping scheme to reduce the size of translation block mapping directory (TBMD). Third, we take advantage of PATC to identify data whether they are hot or cold. By separating hot data from cold data using PATC, we prolong NAND flash memory's lifespan and reduce garbage collection time without any additional cost. Integrating these three schemes leads to the improved read response time compared to the state-of-the-art FTL algorithm, DFTL, by up to 56.9% though it uses only 10% of SRAM. Moreover, if the proposed scheme uses the same amount of SRAM, the response time is improved and the average number of valid pages in a victim block also decreases by up to 67% by efficiently separating hot data from cold data. Sehwan Lee, Bitna Lee, Kern Koh, Hyokyung Bahn |
RTCSA (1) | 4 |
| 2011 | FeGC: An efficient garbage collection scheme for flash memory based storage systems
Ohhoon Kwon, Kern Koh, Hyokyung Bahn |
J. Syst. Softw. | 4 |
| 2010 | Unifying Buffer Replacement and Prefetching with Data Migration for Heterogeneous Storage DevicesabstractWith the good properties of NAND flash memory such as small size, shock resistance, and low-power consumption, large capacity SSD (Solid State Disk) is anticipated to replace hard disk in high-end systems. However, the cost of NAND flash memory is still high to substitute for hard disk entirely. Using hard disk and NAND flash memory together as secondary storage is an alternative solution to provide relatively low response time, large capacity, and reasonable cost. In this paper, we present a new buffer cache management scheme with data migration that is optimized to use both NAND flash memory and hard disk together as secondary storage. The proposed scheme has three salient features. First, it detects I/O access patterns from each storage, and allocates the buffer cache space for each pattern by computing marginal gain adaptively considering the I/O cost of storage. Second, it prefetches data selectively according to their access pattern and storage devices. Third, it moves the evicted data from the buffer cache to hard disk or NAND flash memory considering the access patterns of block references on the reclamation. Trace-driven simulations show that the proposed scheme improves the I/O performance significantly. It enhances the buffer cache hit ratio by up to 29.9% and reduces the total I/O elapsed time by up to 49.5% compared to the well-acknowledged UBM scheme. Sehwan Lee, Kern Koh, Hyokyung Bahn |
ICPADS | 3 |
| 2009 | A cost-aware page replacement algorithm for NAND flash based mobile embedded systemsabstractNAND flash memory is widely used as secondary storage in mobile embedded systems such as cellular phones and digital cameras. These systems usually employ a compressed file system (CFS) to store system files which are fixed during the design phase, in combination with a normal file system to store data files. Since retrieving pages from a CFS requires additional decompression time, it is reasonable to grant them higher priorities when making a page replacement decision. In this paper, we present a new page replacement algorithm for NAND flash memory based embedded systems that considers asymmetric operation cost of each page. The proposed algorithm considers the decompression cost of a page from CFS as well as the asymmetric I/O costs of reads and writes in flash memory. To do this, the algorithm partitions the memory space into a read area, a write area, and a compressed area depending on different operation costs. The size of each area is then dynamically adjusted based on the change of access patterns and the contribution to reducing the I/O costs. Trace-driven simulations show that the proposed algorithm improves the I/O performance of mobile embedded systems significantly. Specifically, it reduces I/O time by 4.8-53.3% compared to widely acknowledged algorithms such as CLOCK, CAR, and CFLRU. Hyejeong Lee, Seunghwan Hyun, Kern Koh, Hyokyung Bahn |
EMSOFT | 5 |
| 2009 | Characterizing virtual memory write references for efficient page replacement in NAND flash memoryabstractRecently, NAND flash memory is being used as the swap space of virtual memory as well as the file storage of embedded systems. Since temporal locality is dominant in page references of virtual memory, LRU and its approximated algorithms are widely used. However, we show that this is not true for write references. We analyze the characteristics of virtual memory read and write references separately, and find that the temporal locality of write references is weak and irregular. Based on this observation, we present a new page replacement algorithm that uses different strategies for read and write operations in predicting the re-reference likelihood of pages. For read operations, temporal locality alone is used, but for write operations, write frequency as well as temporal locality is used. The algorithm partitions the memory space into a read area and a write area to keep track of their reference patterns precisely, and then adjusts their sizes dynamically based on their reference patterns and I/O costs. Though the algorithm has no external parameter to tune, it performs better than CLOCK, CAR, and CFLRU by 20-66%. It also supports optimized implementations for virtual memory systems. Hyejeong Lee, Hyokyung Bahn |
MASCOTS | 2 |
| 2009 | Buffer Cache Management for Combined MLC and SLC Flash Memories Using both Volatile and Nonvolatile RAMsabstractThis paper presents a new buffer cache management scheme called DABC-NV for mixed MLC and SLC flash memories as the secondary storage and both byte-accessible NVRAM and conventional volatile RAM as their buffer caches. DABC-NV has four salient features. First, it allocates buffer cache space to MLC and SLC flash memories based on their I/O costs and then dynamically adjusts the allocated size according to the evolution of workloads. Second, it separately exploits read and write histories of block references, and thus it estimates future references of each operation more precisely. Third, it guarantees the complete consistency of write I/Os since all dirty data are cached in nonvolatile buffer caches. Fourth, metadata lists are maintained separately from cached blocks. This allows more efficient management of volatile and nonvolatile buffer caches based on read and write histories, respectively. Trace-driven simulations show that DABC-NV improves the I/O performance of embedded systems significantly. Specifically, it reduces I/O time by 24% on average compared to the CLOCK-NV algorithm. Hyokyung Bahn, Kern Koh |
RTCSA | 2 |
| 2009 | P/PA-SPTF: Parallelism-aware request scheduling algorithms for MEMS-based storage devicesabstractMEMS-based storage is foreseen as a promising storage media that provides high-bandwidth, low-power consumption, high-density, and low cost. Due to these versatile features, MEMS storage is anticipated to be used for a wide range of applications from storage for small handheld devices to high capacity mass storage servers. However, MEMS storage has vastly different physical characteristics compared to a traditional disk. First, MEMS storage has thousands of heads that can be activated simultaneously. Second, the media of MEMS storage is a square structure which is different from the platter structure of disks. This article presents a new request scheduling algorithm for MEMS storage called P-SPTF that makes use of the aforementioned characteristics. P-SPTF considers the parallelism of MEMS storage as well as the seek time of requests on the two dimensional square structure. We then present another algorithm called PA-SPTF that considers the aging factor so that starvation resistance is improved. Simulation studies show that PA-SPTF improves the performance of MEMS storage by up to 39.2% in terms of the average response time and 62.4% in terms of starvation resistance compared to the widely acknowledged SPTF algorithm. We also show that there exists a spectrum of scheduling algorithms that subsumes both the P-SPTF and PA-SPTF algorithms. Hyokyung Bahn, Soyoon Lee, Sam H. Noh |
ACM Trans. Storage | 1 |
| 2008 | Vector Read: Exploiting the Read Performance of Hybrid NAND FlashabstractOneNAND flash is a NAND based hybrid flash memory which offers much better read/write performance and other functionalities than standard NAND flash memory. Because of its superior read performance, OneNAND flash became the most promising storage solution for implementing high performance mobile device which adopts demand paging memory architecture. However, unfortunately, existing general purpose operating systems, such as Linux, fail to exploit the read performance of OneNAND flash because of a restriction imposed by their I/O architecture and device driver interface. This paper investigated the cause of such inability, and proposed the design of vectored read scheme for solving that problem. Implementation studies on Linux kernel shows that vectored read scheme reduces the block level read latency of OneNAND flash by up to 41% compared with commercial FTL device. It also proved its effectiveness in real-world applications by reducing page fault handling time by 28%. Seunghwan Hyun, Sehwan Lee, Sungyong Ahn, Hyokyung Bahn, Kern Koh |
RTCSA | 4 |
| 2007 | Memory-Efficient Compressed Filesystem Architecture for NAND Flash-Based Embedded Systems
Seunghwan Hyun, Sungyong Ahn, Sehwan Lee, Hyokyung Bahn, Kern Koh |
ICCSA (1) | 4 |
| 2007 | Page Replacement Algorithms for NAND Flash Memory Storages
Yun-Seok Yoo, Hyejeong Lee, Yeonseung Ryu, Hyokyung Bahn |
ICCSA (1) | 4 |
| 2006 | B-PIC: A Novel Caching Scheme for Multimedia Streaming Servers
Ohhoon Kwon, Taeseok Kim, Hyokyung Bahn, Kern Koh |
HiPC | 3 |
| 2006 | A New Address Mapping Scheme for High Parallelism MEMS-Based Storage Devices
Soyoon Lee, Hyokyung Bahn |
HPCC | 2 |
| 2006 | A New Key Management Scheme for Distributed Encrypted Storage Systems
Myungjin Lee, Hyokyung Bahn, Kijoon Chae |
ICCSA (1) | 2 |
| 2006 | Parallelism-Aware Request Scheduling for MEMS-based Storage DevicesabstractMEMS-based storage is being developed as a new storage media. Due to its attractive features such as high-bandwidth, low-power consumption, and low cost, MEMS storage is anticipated to be used for a wide range of applications. However, MEMS storage has vastly different physical characteristics compared to a traditional disk. First, MEMS storage has thousands of heads that can be activated simultaneously. Second, the media of MEMS storage is a square structure which is different from the platter structure of disks. This paper presents a new request scheduling algorithm for MEMS storage that makes use of the aforementioned characteristics. This new algorithm considers the parallelism of MEMS storage as well as the seek time on the two dimensional square structure. We then extend this algorithm to consider the aging factor so that starvation resistance is improved. Simulation studies show that the proposed algorithms improve the performance of MEMS storage by up to 39.2% in terms of the average response time and 62.4% in terms of starvation resistance compared to the widely acknowledged SPTF (Shortest Positioning Time First) algorithm. Soyoon Lee, Hyokyung Bahn, Sam H. Noh |
MASCOTS | 2 |
| 2005 | Web cache management based on the expected cost of web objects
Hyokyung Bahn |
Inf. Softw. Technol. | 1 |
| 2004 | A scalable Web cache sharing scheme
Yong H. Shin, Hyokyung Bahn |
Inf. Process. Lett. | 2 |
| 2002 | Replica-aware caching for Web proxies
Hyokyung Bahn, Hyunsook Lee, Sam H. Noh, Sang Lyul Min, Kern Koh |
Comput. Commun. | 1 |
| 2000 | Pareto-based soft real-time task scheduling in multiprocessor systemsabstractWe develop a new method to map (i.e. allocate and schedule) real-time applications into certain multiprocessor systems. Its objectives are: the minimization of the number of processors used; and the minimization of the deadline missing time. Given a parallel program with real time constraints and a multiprocessor system, our method finds schedules of the program in the system which satisfy all the real time constraints with minimum number of processors. The minimization is carried out through a Pareto-based genetic algorithm which independently considers the both goals, because they are non-commensurable criteria. Experimental results show that our scheduling algorithm achieved better performance than previous ones. The advantage of our method is that the algorithm produces not a single solution but a family of solutions known as the Pareto-optimal set, out of which designers can select optimal solutions appropriate for their environmental conditions. Jaewon Oh, Hyokyung Bahn, Chris Wu, Kern Koh |
APSEC | 2 |